# Underneath

> Underneath is a generative engine optimization agency. We help businesses become easier for AI-powered search and answer systems to understand, retrieve, describe and reference, including ChatGPT, Google AI Overviews and AI Mode, Gemini, Claude, Perplexity and Microsoft Copilot.

Canonical entity record for Underneath (Underneath Agency). Updated 2026-10-11. HTML: https://underneath.agency/agent · JSON-LD: https://underneath.agency/.well-known/entity.json · Site index: https://underneath.agency/llms.txt

## Identity

- Name: Underneath
- Legal name: Underneath Agency
- Also known as: Underneath Agency, Underneath GEO agency, underneath.agency
- Category: Generative engine optimization (GEO) agency
- Website: https://underneath.agency
- Founded: 2026
- Serves: United States, United Kingdom, Australia, Canada
- Headquarters: 2846 Simons Hollow Road, Bloomsburg, PA 17815, United States
- Contact: hello@underneath.agency. Email only; no phone number is published. Replies within one business day (Monday to Friday, US business hours).
- Organization @id: https://underneath.agency/#organization

## What Underneath does

Underneath finds out why people choose a business, then makes sure AI assistants and AI search give that answer: the pages, structured data, entity facts and third-party sources that AI engines retrieve, trust and cite. It sells generative engine optimization only, as one practice under one plan, and reports mentions, citations, share of voice and leads from AI referrals.

Mission: We help businesses, from single locations to enterprises, get mentioned, cited and recommended by AI assistants and AI search, and chosen because of it, with senior specialists, measurable methods and no shortcuts.

### What Underneath does not sell

- SEO as a standalone service
- paid advertising (PPC, social ads)
- social media management
- web design or development
- branding or creative production

### AI engines tracked in every program

ChatGPT, Claude, Gemini, Google AI Overviews, Google AI Mode, Perplexity, Microsoft Copilot, Bing.

## Services

### [Generative Engine Optimization](https://underneath.agency/services/generative-engine-optimization)

GEO services: get your business understood, cited and recommended in ChatGPT, AI Overviews, AI Mode, Gemini, Perplexity and Copilot answers. (updated 2026-09-27)


## Who it is for

Audiences: Single-location businesses; Franchises and multi-location brands; Growing companies; Enterprises; Ecommerce and Shopify merchants; SaaS companies.

Industries: Home and local services; Franchises and multi-location brands; Healthcare and dental; Legal and professional services; B2B software and technology; Financial services and insurance; Retail and ecommerce; Hospitality and travel.

## Pricing

Programs are priced in three monthly bands, set by scope: $5,000 to $10,000, $11,000 to $20,000, and $20,000 and above. Projects and audits are priced to scope. The Marketing & AI Visibility Audit starts every program and its fee is credited if you continue.

## How an engagement starts

1. Describe the problem through the contact form.
2. A free 30-minute strategy call, then a short written read: what we heard, what we would look at first, whether we are the right fit.
3. One of three ways of working: a monthly program (month one is the Marketing & AI Visibility Audit and the plan built from it), a fixed-scope project, or the audit on its own.

## How the work stays visible: The weekly 30

Every program has a 30-minute review in the same slot every week, with a decision-maker from your side in the room. Four things, always in this order.

1. The work completed, with links
2. The evidence across AI visibility, search and business metrics
3. What the data and research are telling us
4. The next priorities and decisions required

A written note follows within one business day. If a decision-maker cannot join the weekly review, we may not be the right fit.

## Evidence policy

No client results are published until a client has approved them. Scenarios on the site are labeled illustrative. No guaranteed rankings or AI mentions by a date; accounts, data and domains stay in the client’s name.

## Research

- [Do websites serve Markdown to AI agents? 2026 data](https://underneath.agency/research/agent-readable-web-study): 3.2% of top websites serve Markdown when an AI agent asks for it, at a median 96.1% smaller than the HTML; 45.8% of homepages carry JSON-LD. (updated 2026-10-08)
- [Do ChatGPT, Gemini, Perplexity and Claude agree on brands?](https://underneath.agency/research/ai-assistants-brand-agreement-study): 80 buyer questions, four AI assistants: a third of picks shared, the same top pick 10% of the time. They mostly choose differently, not read differently. (updated 2026-10-08)
- [Do AI answers match a business’s Google profile?](https://underneath.agency/research/ai-business-facts-accuracy-study): We asked 4 AI engines for the address, phone, website and hours of 159 local businesses. 18.9% of answers had a fact that differed from the Google profile. (updated 2026-10-08)
- [Do ChatGPT, Gemini, Perplexity and Claude cite pages that rank?](https://underneath.agency/research/ai-citations-google-rankings-study): 8.3% of ChatGPT’s citations rank in Google’s top 10 for the question, Claude’s 25.7%. We trace where the gap opens, stage by stage, over 80 buyer questions. (updated 2026-10-08)
- [Which AI crawlers do top websites block? 10,000 sites, 2026](https://underneath.agency/research/ai-crawler-blocking-study): 15.2% of top sites block GPTBot and 7.4% OpenAI’s search crawler. A same-day check shows ChatGPT cited no page its crawler was barred from. (updated 2026-10-08)
- [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study): ChatGPT ran 3.7 searches per buyer question before answering. When a search named a source such as NerdWallet or Avvo, the answer cited it 44.0% of the time. (updated 2026-10-08)
- [AI Mode vs AI Overviews: how different are the sources?](https://underneath.agency/research/ai-mode-vs-ai-overviews-study): Same 400 searches: AI Mode and AI Overviews shared 13.9% of cited URLs, and AI Mode cited 16.3% of top-10 pages to the AI Overview’s 29.7%. (updated 2026-10-08)
- [Do AI Overviews cite the pages that rank? 4,051 citations](https://underneath.agency/research/ai-overview-citations-study): Only 28.7% of 4,051 AI Overview citations are page-one Google results. A page-one result’s citation rate falls from 49.5% at position 1 to 15.5% at 9. (updated 2026-10-08)
- [What pages cited by AI Overviews have in common: 3,096 pages](https://underneath.agency/research/ai-overview-cited-pages-study): 3,096 top-10 pages on the same searches: rank predicted AI Overview citation. Compared within each search, only a machine-readable date held up. (updated 2026-10-08)
- [The YouTube videos Google’s AI cites are small](https://underneath.agency/research/ai-overview-youtube-videos-study): Within the same search, the YouTube videos Google’s AI cited came from smaller channels than the ones it showed and skipped. 800 US searches. (updated 2026-10-08)
- [When does Google show an AI Overview? 1,248 US searches](https://underneath.agency/research/ai-overviews-frequency-study): 1,248 US searches: AI Overviews followed the search more than the industry. Local packs cut the rate to 30.7%; rewording a topic lifted it from 59.4% to 93.8%. (updated 2026-10-08)
- [How faithfully do AI assistants quote software prices?](https://underneath.agency/research/ai-pricing-accuracy-study): 61.9% of 840 plan prices four AI assistants quoted for 45 software products were fully faithful to the official page. Most of the rest were real variants. (updated 2026-10-08)
- [Does rewording a question change AI brand recommendations?](https://underneath.agency/research/ai-prompt-phrasing-study): In 1,560 AI answers, a tight budget kept the original’s first brand 15.3% of the time, against 68.0% when the same question was simply asked again. (updated 2026-10-08)
- [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study): Same 20 buyer questions, five runs each: a quarter of ChatGPT’s brands appeared every time. AI visibility is a distribution, and one answer is a sample. (updated 2026-10-08)
- [Same question, four countries: do AI recommendations change?](https://underneath.agency/research/ai-recommendations-by-country-study): Five runs per country, 40 questions: ChatGPT’s brand lists overlapped 0.594 within one country but 0.429 across the US, UK, Canada and Australia. (updated 2026-10-08)
- [When does a Reddit thread become evidence in Google’s AI?](https://underneath.agency/research/ai-reddit-citations-study): Google’s AI cited Reddit in 17.9% of AI Overviews. Cited threads had twice the comments of skipped ones, and 20.8% of cited claims were not supported. (updated 2026-10-08)
- [How fresh are the pages AI engines cite?](https://underneath.agency/research/ai-source-freshness-study): Held to the same questions, AI assistants cited pages first published half as long ago as Google’s top 10. AI Overviews showed no pull toward recent pages. (updated 2026-10-08)
- [Do Wikipedia and schema make AI assistants recommend a brand?](https://underneath.agency/research/brand-entity-ai-recommendations-study): Brands on Wikipedia are named by more AI assistants, but within the same question most of the gap goes once brand prominence is taken into account. (updated 2026-10-08)
- [Which Google Maps businesses does ChatGPT recommend?](https://underneath.agency/research/chatgpt-local-picks-google-profile-study): Across 898 ChatGPT answers over two days, more reviews than the local median raised a Maps business’s chance of being listed by about 20 points. (updated 2026-10-08)
- [ChatGPT local recommendations: stable details, shifting lists](https://underneath.agency/research/chatgpt-local-recommendations-study): ChatGPT’s local listings match Google Maps field by field, but which businesses it lists changes from run to run, and cited pages often do not name them. (updated 2026-10-08)
- [“Is this brand legit?” How AI assistants build a reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study): Asked “Is this brand legit?”, four AI engines said yes in every complete answer, and 99.7% raised a problem. How they build a reputation from evidence. (updated 2026-10-08)
- [How many websites have an llms.txt file? 2026 adoption data](https://underneath.agency/research/llms-txt-adoption-study): We checked /llms.txt on 5,902 top websites: 11.5% publish a valid file, 2.1% an llms-full.txt, and 15.7% return an HTML page at the address instead. (updated 2026-10-08)
- [How many “best of” lists cited by AI rank their own brand first?](https://underneath.agency/research/self-promoting-best-lists-study): Of 269 AI-cited numbered “best X” lists with an identifiable publisher, 24.2% ranked their publisher first. Such lists were 1.1% of all citations. (updated 2026-10-08)

## Guides

- [How can a CPA or tax firm win better clients through AI search?](https://underneath.agency/resources/accounting-firms-clients-ai-search): By being the firm AI assistants tie to a specific niche and place. Early studies show a few large firms dominate answers, but focused smaller firms break in. (updated 2026-10-10)
- [Are AI assistants now choosing accounting software for small businesses?](https://underneath.agency/resources/accounting-software-ai-search): Partly. AI assistants now shape which accounting tools owners and their accountants consider, so vendors must be named, priced and described correctly there. (updated 2026-10-10)
- [How do AI agent companies get onto buyers’ shortlists in AI search?](https://underneath.agency/resources/ai-agent-companies-customers-from-ai-search): By making their agent easy to verify: clear use cases, public security and governance evidence, honest pricing and independent proof that assistants can find. (updated 2026-10-10)
- [Do people click fewer websites when they use AI answer engines?](https://underneath.agency/resources/ai-answer-engines-fewer-website-clicks): Yes. In a small 2024 lab study, people clicked about one source per question with AI answer engines and about four on Google. Larger studies agree. (updated 2026-10-08)
- [What do lost clicks to AI answers mean for pipeline and revenue?](https://underneath.agency/resources/ai-answers-pipeline-revenue): Fewer search clicks mean fewer prospects entering your funnel, and AI answers appear most on the research and comparison searches where buying starts. (updated 2026-10-10)
- [How often do AI answers say things their sources do not support?](https://underneath.agency/resources/ai-answers-unsupported-claims): Often: audits find roughly one in nine to one in three statements in AI search answers are not backed by the sources cited beside them. (updated 2026-10-08)
- [How often do AI assistant users never visit a website?](https://underneath.agency/resources/ai-assistant-sessions-without-website-visits): About a third of the time: in a 2026 panel, 34.1% of AI assistant sessions had no observed web visit, against 19.5% of the same people’s search sessions. (updated 2026-10-10)
- [Can our server logs show what content AI bots are looking for on our site?](https://underneath.agency/resources/ai-bot-server-logs-content-demand): Partly. Logs show which pages AI bots request, including ones you lack, but not whether any answer used or cited them. One site acted on that signal. (updated 2026-10-08)
- [Can an AI search engine cite my page for something my page does not say?](https://underneath.agency/resources/ai-citing-pages-for-claims-they-do-not-make): Yes. Audits find AI answers attach citations to pages that do not back the claim, from about 3% of Google AI Overview claims to half of one engine’s citations. (updated 2026-10-08)
- [How do AI coding tools win developers through AI search?](https://underneath.agency/resources/ai-coding-tools-developers-ai-search): By earning developer trust where AI answers look: docs, benchmarks, community threads and honest comparisons, then growing one developer into a team of seats. (updated 2026-10-10)
- [Can we trust an AI assistant’s explanation of why it recommended a competitor?](https://underneath.agency/resources/ai-explanations-for-recommendations): Only partly. In a 12-model hotel test, AI assistants’ stated reasons matched their real drivers imperfectly and left out list order and review count. (updated 2026-10-08)
- [How do AI image generators get recommended by AI assistants?](https://underneath.agency/resources/ai-image-generators-growth-from-ai-search): By being clear on what they do best and on commercial rights, then getting that into the independent reviews, galleries and lists AI assistants read. (updated 2026-10-10)
- [How much traffic would AI Mode as Google’s default cost you?](https://underneath.agency/resources/ai-mode-default-traffic-loss): In a 2026 field experiment, making AI Mode Google’s default cut clicks to outside websites by 18.8 points per search. What that means for your traffic. (updated 2026-10-08)
- [If AI Overviews cite my site, does that make up for lost clicks?](https://underneath.agency/resources/ai-overview-citations-vs-lost-clicks): Not on current evidence. Only 1% of visits to AI Overview pages led to a click on a cited source, and Wikipedia lost traffic despite heavy citation. (updated 2026-10-10)
- [Are people more likely to stop browsing after seeing an AI Overview?](https://underneath.agency/resources/ai-overviews-end-browsing-sessions): Yes. In a Pew panel of 900 US adults, browsing ended on 26% of Google pages with an AI Overview, against 16% of pages without one. (updated 2026-10-10)
- [How do AI research tools get found and trusted through AI search?](https://underneath.agency/resources/ai-research-tools-users-ai-search): By being named, with checkable proof of citation accuracy, when students, researchers and librarians ask AI which research assistant to trust. (updated 2026-10-10)
- [Do a few websites win most AI Mode and Perplexity citations?](https://underneath.agency/resources/ai-search-citation-concentration): Yes. A small group of sites takes most AI citations, most of all in Google AI Mode in one Swiss study, least in Perplexity, with a long tail behind them. (updated 2026-10-10)
- [Will AI search reduce our dependence on marketplaces and aggregators?](https://underneath.agency/resources/ai-search-marketplace-dependence): Not yet. Booking sites still supplied most of Gemini’s hotel citations, but their share fell for experience-led questions, where brands’ own pages can compete. (updated 2026-10-08)
- [What should our AI search strategy be for the next five years?](https://underneath.agency/resources/ai-search-strategy-next-five-years): Nobody can forecast five years of AI search. Invest in what has held across engines, measure properly, and keep options open on ads and Google’s AI Mode. (updated 2026-10-07)
- [How can AI software companies stand out in AI search when every rival claims AI?](https://underneath.agency/resources/ai-software-companies-in-ai-search): By earning independent, current proof of what the product does, because AI assistants and buyers both discount the hype that fills the AI software category. (updated 2026-10-08)
- [How can an AI startup win customers through AI search?](https://underneath.agency/resources/ai-startups-customers-from-ai-search): By earning independent coverage and reviews that assistants can find, so the startup reaches shortlists that turn into trials, team use and contracts. (updated 2026-10-08)
- [How do AI video companies win customers through AI search?](https://underneath.agency/resources/ai-video-generator-customers-from-ai-search): By being the tool AI assistants name when creators, marketers and training teams ask what to use, then turning that intent into free plans and upgrades. (updated 2026-10-08)
- [How can AI voice companies turn AI search into revenue?](https://underneath.agency/resources/ai-voice-software-revenue-from-ai-search): By being named, with clear consent and pricing facts, when developers, creators and contact center leaders ask AI which voice tool to try. (updated 2026-10-08)
- [Should our reputation priorities for AI assistants differ from those for human customers?](https://underneath.agency/resources/ai-vs-human-reputation-priorities): Partly. AI assistants share customers’ focus on rating and price, but in testing they overweighted eco-certification and ignored replies to reviews. (updated 2026-10-08)
- [How can an AI writing tool win users when ChatGPT also writes?](https://underneath.agency/resources/ai-writing-tools-users-from-ai-search): By owning a specific writing job general assistants do poorly, and getting that strength into the independent lists and reviews assistants cite. (updated 2026-10-08)
- [How do analytics and BI software companies get shortlisted when buyers ask AI?](https://underneath.agency/resources/analytics-bi-software-ai-search): By being named on the short lists AI answers build for BI comparisons and alternatives, with public proof on governance, accuracy and cost. (updated 2026-10-08)
- [Can we monitor ChatGPT and Gemini visibility through their APIs?](https://underneath.agency/resources/api-ai-visibility-monitoring): Not reliably. In a 2026 test, ChatGPT’s app and its API shared only 12.0% of cited websites for the same question asked a fraction of a second apart. (updated 2026-10-08)
- [How do API companies win developer demand when AI agents choose the integration?](https://underneath.agency/resources/api-companies-ai-search): By being the API that AI assistants and coding agents choose and can integrate on their own, so usage starts in code and grows into enterprise contracts. (updated 2026-10-08)
- [How can an apparel brand win new customers through AI search?](https://underneath.agency/resources/apparel-brands-customers-ai-search): By publishing the fit, size, fabric, care and returns facts AI tools need to match clothes to a shopper, and earning independent mentions that confirm them. (updated 2026-10-08)
- [How do application security vendors win demand when developers and CISOs ask AI?](https://underneath.agency/resources/application-security-demand-ai-search): AppSec vendors earn AI-driven demand by being verifiable to two buyers at once: developers who try tools and security leaders consolidating them. (updated 2026-10-08)
- [Are ads coming to ChatGPT and other AI assistants?](https://underneath.agency/resources/are-ads-coming-to-ai-assistants): They are already arriving. Researchers report OpenAI and Google have put ads into their AI products, and ChatGPT ads reached 31 European markets in August. (updated 2026-10-08)
- [Are AI search engines citing AI-generated content?](https://underneath.agency/resources/are-ai-search-engines-citing-ai-content): Yes. A 2026 audit of four AI search engines flagged about 16% of readable cited pages as AI-generated, from 7.3% for ChatGPT to 27.8% for Copilot. (updated 2026-10-08)
- [How does an asset manager get its funds considered when investors and advisers ask AI?](https://underneath.agency/resources/asset-managers-fund-demand-ai-search): By being the fund family assistants find in trusted fund research, with fees and facts consistent everywhere. Five firms already hold 58% of US fund assets. (updated 2026-10-08)
- [Will AI search point plant leaders to our firm when they plan an automation project?](https://underneath.agency/resources/automation-integrators-industrial-buyers-ai-search): It can, if assistants find proof of your platforms, industries, region and results, so plant leaders put you on the integrator longlist early. (updated 2026-10-08)
- [How can B2B SaaS companies generate revenue from AI search?](https://underneath.agency/resources/b2b-saas-revenue-from-ai-search): By getting onto the shortlists AI assistants build for category, alternatives and comparison questions, then turning that intent into trials and renewals. (updated 2026-10-08)
- [How do baby product brands get recommended when parents ask AI?](https://underneath.agency/resources/baby-product-brands-ai-recommendations): By giving AI assistants accurate, safety-complete product facts and independent proof that parents trust, so the brand is named when registries are built. (updated 2026-10-08)
- [How does a banking technology vendor get onto a bank’s shortlist when the first research happens in AI?](https://underneath.agency/resources/banking-technology-enterprise-deals-ai-search): By being on the first list a bank builds before its RFP. Core, payments and fraud deals are rare and long, so early AI answers carry unusual weight. (updated 2026-10-08)
- [How do beauty retailers turn AI product recommendations into sales?](https://underneath.agency/resources/beauty-retailers-ai-search): By being the store AI names when shoppers ask what to buy and where: accurate prices and stock, clear loyalty value, and content that answers routine questions. (updated 2026-10-08)
- [Which pages should we target to show up in AI recommendations?](https://underneath.agency/resources/best-of-lists-ai-recommendations): Ranked “best of” lists on other sites are the page type AI engines cite most when naming brands; your own site is a small share of citations. (updated 2026-10-11)
- [How does a beverage brand get its drinks recommended by AI assistants?](https://underneath.agency/resources/beverage-brands-ai-recommendations): By being the clear answer to need-state questions: 27% of US consumers already took AI drink advice, and shoppers check ingredients, price and reviews. (updated 2026-10-08)
- [Can AI search help a biotech company find partners, licensees and investors?](https://underneath.agency/resources/biotech-partnering-ai-search): Partly: AI tools now help investors and scientists screen biotech deals, so assets with public, cited science are easier to find before partnering talks. (updated 2026-10-08)
- [Will architects and contractors find our products when they ask AI what to specify?](https://underneath.agency/resources/building-materials-specification-ai-search): Increasingly, yes: product research is moving online and into AI, so specifiers can only name materials whose performance data AI can find and check. (updated 2026-10-08)
- [How does a business advisory firm get found when owners ask AI for help?](https://underneath.agency/resources/business-advisory-firms-clients-ai-search): Owners now ask AI about exits, cash and growth before calling anyone. Advisory firms get named by being specific, credentialed and independently covered. (updated 2026-10-08)
- [Could AI search engines filter out manipulative GEO content?](https://underneath.agency/resources/can-ai-search-filter-manipulative-geo): In lab tests, yes: one defense cut manipulative GEO’s success from 50.32% to 6.20% while keeping most useful sources. No live engine is known to use it yet. (updated 2026-10-08)
- [Can a competitor game AI assistants into recommending their product first?](https://underneath.agency/resources/can-competitors-game-ai-recommendations): In lab tests, yes: hidden text on a product page pushed a product to the top of AI recommendations. Stronger, defended AI engines resist far better. (updated 2026-10-08)
- [Can optimizing for AI search backfire on your brand?](https://underneath.agency/resources/can-geo-backfire-on-your-brand): Yes. In controlled tests, common rewriting tricks often lowered rankings and hype got flagged, but careful optimization did not harm answer quality. (updated 2026-10-08)
- [Can GEO be used to push false information into AI answers?](https://underneath.agency/resources/can-geo-push-false-information-into-ai-answers): Yes. In tests, one planted page made AI systems recommend a fake product in up to 27% of cases, and fabricated evidence was treated as real. (updated 2026-10-08)
- [Can smaller, lower-traffic websites get cited by AI search engines?](https://underneath.agency/resources/can-small-websites-get-cited-by-ai): Yes. AI engines cite a long tail of little-known sites, but being cited is not being recommended, and well-known brands still win most answers. (updated 2026-10-10)
- [Can sellers game AI shopping rankings with manipulative product copy?](https://underneath.agency/resources/can-you-game-ai-shopping-rankings): Sometimes, on weaker AI systems. In tests, GPT-5 and Claude flagged and demoted manipulative product copy, and any gain vanished once rivals copied it. (updated 2026-10-08)
- [What do commercial health sites cited by ChatGPT have in common?](https://underneath.agency/resources/chatgpt-cited-commercial-health-sites): In a 615-source ChatGPT audit, cited commercial health sites mostly showed medical review (71.1%), schema (86.8%) and pages over 1,500 words (68.4%). (updated 2026-10-11)
- [Will optimizing content for ChatGPT hurt our Google rankings?](https://underneath.agency/resources/chatgpt-optimization-google-rankings): Not if you protect pages already earning Google clicks. In one field test, a guarded rewrite lifted ChatGPT referrals with no Google loss beyond the site trend. (updated 2026-10-08)
- [Our ChatGPT referral traffic is growing; can we credit our GEO work?](https://underneath.agency/resources/chatgpt-referral-growth-and-geo): Not from growth alone. In one site’s logs, ChatGPT referrals grew 5.7 times, but untouched pages grew 3.5 times. Credit needs a comparison group. (updated 2026-10-08)
- [Where does ChatGPT fit in the customer journey compared with Google?](https://underneath.agency/resources/chatgpt-vs-google-customer-journey): Later than Google: in a 2026 US and British panel, search usually opened the journey, while people more often browsed before using an AI assistant than after. (updated 2026-10-10)
- [When a formulator asks AI who supplies a chemical, does our company come up?](https://underneath.agency/resources/chemical-suppliers-b2b-buyers-ai-search): Chemical suppliers reach B2B buyers in AI search when grades, specifications, safety data sheets and regulatory status are public and easy to verify. (updated 2026-10-08)
- [Does my brand show up differently in Chinese AI models than in Western ones?](https://underneath.agency/resources/chinese-vs-western-ai-brand-visibility): Likely yes. One study found Chinese AI models named brands far more often than Western ones, though its authors have a stake. Language clearly matters. (updated 2026-10-08)
- [How can a cloud provider win enterprise deals when buyers ask AI first?](https://underneath.agency/resources/cloud-infrastructure-enterprise-demand-ai-search): By being the cloud AI answers name for real architecture, cost and sovereignty questions, with proof architects can check before they commit. (updated 2026-10-08)
- [How do cloud security vendors get onto enterprise shortlists when buyers ask AI first?](https://underneath.agency/resources/cloud-security-enterprise-buyers-ai-search): Cloud security vendors win AI shortlists with proof assistants can check: clear category fit, cloud coverage, marketplace listings and third-party evidence. (updated 2026-10-08)
- [Can AI search bring our CNC shop the RFQs that now go to instant-quote platforms?](https://underneath.agency/resources/cnc-machine-shops-rfqs-ai-search): Yes, if assistants can verify your machines, materials, tolerances, certifications and location; instant-quote platforms are already built for that. (updated 2026-10-08)
- [How do collaboration software companies win teams that ask AI first?](https://underneath.agency/resources/collaboration-software-ai-search): By being named when teams ask AI which tool to try, then turning those free signups into seat growth and company-wide contracts that renew. (updated 2026-10-08)
- [Can we combine ChatGPT, Gemini and Perplexity into one AI visibility score?](https://underneath.agency/resources/combine-ai-engines-visibility-score): Only with care. Score each engine separately first, then combine with stated weights; pooling raw counts lets the most talkative engine dominate. (updated 2026-10-08)
- [Do contractors find our construction software when they ask AI what to use?](https://underneath.agency/resources/construction-software-customers-ai-search): Often they do: contractors use AI at work, and the shortlist for estimating, bidding and project software forms before your demo request arrives. (updated 2026-10-08)
- [Can AI search put our jobsite technology on a contractor’s shortlist?](https://underneath.agency/resources/construction-technology-demand-ai-search): It can, when AI can verify your accuracy, compliance and service record, because contractors now shortlist drones, robots and sensors before calling. (updated 2026-10-08)
- [How does a specialist consulting firm get on the shortlist when clients ask AI for advice?](https://underneath.agency/resources/consulting-firms-clients-ai-search): By being the firm AI can tie to a specific problem, industry and place, and confirm in outside sources, because clients now research consultants before calling. (updated 2026-10-08)
- [Can AI assistants win new subscribers for a consumer app?](https://underneath.agency/resources/consumer-app-subscribers-from-ai-search): Yes. One in ten US adults first heard of their latest app from an AI assistant, but the app store listing and the first hour of use still close the sale. (updated 2026-10-08)
- [How do consumer electronics brands and retailers win shoppers who compare specs in AI?](https://underneath.agency/resources/consumer-electronics-sales-from-ai-search): By being the model AI names when shoppers compare specs and prices: AI answers lean on independent lab reviews, retailer data and accurate prices. (updated 2026-10-08)
- [How do consumer fintech apps win customers when people ask AI about money?](https://underneath.agency/resources/consumer-fintech-apps-customers-ai-search): By being named, and described accurately, when people ask AI which app to use and whether it is safe. In money questions, trust facts decide the answer. (updated 2026-10-08)
- [Will AI assistants put our devices on the shortlist when buyers choose a brand and ecosystem?](https://underneath.agency/resources/consumer-tech-brands-ai-search): By being named, with correct compatibility and launch facts, when households ask AI which device or ecosystem to buy into, a choice that lasts years. (updated 2026-10-08)
- [Is my content business at risk from Google’s AI summaries?](https://underneath.agency/resources/content-business-risk-from-ai-search): It depends on your content. Factual pages lose visits to Google’s AI summaries; opinion and experience content gained, until AI Mode eroded the gain. (updated 2026-10-10)
- [Will AI search put our plant on the supplier shortlist before the RFQ goes out?](https://underneath.agency/resources/contract-manufacturers-rfqs-from-ai-search): It can, when your capabilities, certifications and location are written where AI can verify them, because engineers now use AI to build supplier longlists. (updated 2026-10-08)
- [How do makeup brands get recommended when shoppers ask AI?](https://underneath.agency/resources/cosmetics-brands-product-discovery-ai-search): By giving AI assistants shade, undertone and finish facts they can match to a shopper’s question, backed by reviews and independent coverage. (updated 2026-10-08)
- [Why do AI assistants keep naming the same CRMs, and how do we join them?](https://underneath.agency/resources/crm-software-ai-search): Because the leaders dominate the sources AI search reads. Challenger CRMs get named by owning specific needs: an industry, a team size, a budget. (updated 2026-10-08)
- [How do customer support software companies get chosen when buyers ask AI?](https://underneath.agency/resources/customer-support-software-ai-search): By being named while support leaders re-shop for AI agents, with clear outcome pricing and proof of resolution rates that buyers and AI answers can check. (updated 2026-10-08)
- [How can cybersecurity software companies generate revenue from AI search?](https://underneath.agency/resources/cybersecurity-software-revenue-from-ai-search): By getting named in the AI answers that now build security shortlists, backed by third-party proof that CISOs and AI assistants can both check. (updated 2026-10-08)
- [How do data platforms get shortlisted when data leaders ask AI?](https://underneath.agency/resources/data-infrastructure-enterprise-demand-ai-search): By being named, accurately, in the AI answers data leaders use to frame architecture choices, then winning the proof of concept on the buyer’s own data. (updated 2026-10-08)
- [How do data science platforms win enterprise buyers who research with AI?](https://underneath.agency/resources/data-science-platforms-ai-search): By being named and described accurately in the AI answers that frame platform comparisons, governance checks and cost estimates, then winning the pilot. (updated 2026-10-08)
- [How does a database company get chosen when AI helps pick the stack?](https://underneath.agency/resources/database-companies-developers-ai-search): By being the database AI coding tools and assistants reach for by default, then backing it with docs, benchmarks and proof enterprise architects can check. (updated 2026-10-08)
- [Can a site with deep content but no technical SEO get cited by Gemini?](https://underneath.agency/resources/deep-content-without-technical-seo-gemini): Sometimes. In a 14-hotel Gemini audit, an independent hotel with no schema but a 33-question FAQ was cited directly; a polished rival was not. (updated 2026-10-11)
- [Can AI search put your dental technology on a dentist’s shortlist?](https://underneath.agency/resources/dental-technology-ai-search): Dentists already use AI at work. To make their shortlist, a dental technology company needs clear, checkable facts on cost, clearance and fit. (updated 2026-10-08)
- [How do developer tool companies win users when developers ask AI first?](https://underneath.agency/resources/developer-tools-ai-search): By being the tool AI coding agents and assistants pick for a developer’s stack, then turning that free signup into team plans and enterprise contracts. (updated 2026-10-08)
- [How do DevOps companies get recommended when engineers ask AI which tool to use?](https://underneath.agency/resources/devops-platforms-ai-search): By being the tool AI answers and coding agents find in docs, GitHub and engineer communities, then converting team adoption into platform deals. (updated 2026-10-08)
- [How do we know whether a GEO campaign actually improved our AI citations?](https://underneath.agency/resources/did-geo-improve-ai-citations): Only with repeated measurement before and after, a margin of error, and an untreated comparison. A 3-point gain in AI citation share can be pure noise. (updated 2026-10-08)
- [Our AI visibility went up: was it because of our GEO work?](https://underneath.agency/resources/did-geo-work-raise-ai-visibility): Maybe. A rise can also come from chance, AI platform growth, engine updates, competitors or a changed prompt set, so you need a baseline and a control. (updated 2026-10-08)
- [How does a digital bank win depositors when savers ask AI where to put their money?](https://underneath.agency/resources/digital-banks-depositors-ai-search): By making rates, conditions and deposit insurance facts current and checkable, because savers now ask AI where to put money and AI often gets rates wrong. (updated 2026-10-08)
- [How do digital health apps win new users when people ask AI about their health first?](https://underneath.agency/resources/digital-health-apps-users-ai-search): By giving assistants proof they can check: plain privacy terms, honest evidence and app store trust, now that 32% of US adults ask AI about health. (updated 2026-10-08)
- [Do AI agents make up facts about my business, or just leave them out?](https://underneath.agency/resources/do-ai-agents-invent-or-omit-business-facts): Mostly they leave facts out. Answering from the web instead of a business’s site, agents left out 45% of facts, up from 29%; wrong facts rose only 4% to 6%. (updated 2026-10-08)
- [Do AI agents recommend businesses whose websites they can read more often?](https://underneath.agency/resources/do-ai-agents-recommend-readable-websites): In one large vendor study, AI agents clearly recommended businesses with agent-readable websites 20% of the time, against 11% for sites they struggled to read. (updated 2026-10-11)
- [Do AI assistants still answer from training data or from live web pages?](https://underneath.agency/resources/do-ai-assistants-answer-from-training-data): Mostly from live pages. In a 2026 test of four AI agents, memory supplied 7–10% of answers, though some AI answers still cite no source at all. (updated 2026-10-10)
- [Do AI assistants favor big brands over smaller competitors?](https://underneath.agency/resources/do-ai-assistants-favor-big-brands): Mostly yes when products look alike: AI assistants default to market leaders, but in tests a small, clearly stated advantage beat a famous name. (updated 2026-10-08)
- [Do pages need author bylines to be cited by ChatGPT?](https://underneath.agency/resources/do-ai-cited-pages-need-author-bylines): No. In a ChatGPT health audit, 64.7% of cited pages named no author and only 7.0% gave a name with credentials; publisher authority mattered more. (updated 2026-10-11)
- [Do AI engines cite local-language sources in other languages?](https://underneath.agency/resources/do-ai-engines-cite-local-language-sources): Mostly yes: ChatGPT and Perplexity switch to local-language sources, while Claude kept citing English sites in one study. Translation alone is not enough. (updated 2026-10-10)
- [Do ChatGPT and other AI engines cite my own website or third-party reviews?](https://underneath.agency/resources/do-ai-engines-cite-your-own-website): Mostly third-party sources. Brand-owned sites are a small share of AI citations in most studies, though the share rises for buying-ready questions. (updated 2026-10-10)
- [Do Gemini, GPT and Claude prefer the same kind of content?](https://underneath.agency/resources/do-ai-engines-prefer-same-content): Mostly yes on content quality: Gemini, GPT and Claude share most preferences. But they cite different sources, so one good page will not show up everywhere. (updated 2026-10-11)
- [Do Google AI Overviews downplay negative content?](https://underneath.agency/resources/do-ai-overviews-downplay-negative-content): In one audit of 11,000 Google questions, AI Overviews drew less from negatively toned sources. Brand reviews were not tested, and criticism still appears. (updated 2026-10-08)
- [Do Google AI Overviews reduce clicks to my website?](https://underneath.agency/resources/do-ai-overviews-reduce-clicks): Yes. In a 2026 field experiment, hiding Google AI Overviews raised clicks to outside sites by 8.8 points; browsing data and Wikipedia traffic agree. (updated 2026-10-10)
- [Do AI search engines just agree with however a question is phrased?](https://underneath.agency/resources/do-ai-search-engines-agree-with-leading-questions): Often, on loaded questions. In a 2024 audit, AI answer engines gave one-sided answers that agreed with a charged question 50% to 80% of the time. (updated 2026-10-08)
- [Do AI search engines cite the same websites as Google?](https://underneath.agency/resources/do-ai-search-engines-cite-the-same-sites-as-google): Mostly not. Studies find AI engines and Google share 18% to 38% of sources, and only 8.3% of ChatGPT’s citations ranked in Google’s top 10 in our test. (updated 2026-10-10)
- [Do ChatGPT, Copilot, Google and Perplexity cite the same sources?](https://underneath.agency/resources/do-ai-search-engines-cite-the-same-sources): No. Studies find ChatGPT, Copilot, Google and Perplexity cite largely different pages for the same question, so success on one engine rarely carries over. (updated 2026-10-10)
- [Do backlinks still matter for visibility in Perplexity?](https://underneath.agency/resources/do-backlinks-matter-for-perplexity-visibility): Probably, but the evidence is thin. In a test of 112 startups, linking websites were the strongest predictor of Perplexity visibility; no study proves cause. (updated 2026-10-10)
- [Does being cited by ChatGPT boost our Google rankings?](https://underneath.agency/resources/do-chatgpt-citations-boost-google-rankings): No evidence says so. The one field test found ChatGPT gains left Google clicks flat to falling, and ChatGPT mostly cites pages Google does not rank. (updated 2026-10-08)
- [Do citations make people trust AI answers more, even when the links are wrong?](https://underneath.agency/resources/do-citations-make-ai-answers-more-trusted): Yes. In a US experiment, reference links raised trust in AI answers, and broken or irrelevant links raised it just as much as valid ones. (updated 2026-10-08)
- [Do comparison pages help B2B brands get cited by AI?](https://underneath.agency/resources/do-comparison-pages-help-b2b-ai-citations): Partly. Comparison content gets used more once a page is cited, but brands’ own sites drew only 2.9% of AI citations in one large vendor dataset. (updated 2026-10-08)
- [Do FAQ pages help you get cited by AI search engines?](https://underneath.agency/resources/do-faq-pages-help-ai-citations): Not by format alone. In the research so far, FAQ formatting did not make AI engines use a page more; the evidence inside the page did. (updated 2026-10-08)
- [Do GEO content tactics like adding statistics, quotes or citations actually get you cited more by AI?](https://underneath.agency/resources/do-geo-content-tactics-work): Mostly not. A 2025 benchmark found most AI rewriting tactics did nothing or lowered a page’s rank; adding statistics hurt in 19 of 24 settings. (updated 2026-10-08)
- [Do people actually check the sources AI assistants cite?](https://underneath.agency/resources/do-people-check-ai-sources): Rarely. In a Pew panel of 900 US adults, only 1% of visits to Google pages with an AI Overview led to a click on a cited source. (updated 2026-10-08)
- [Do people trust AI search answers less than regular Google results?](https://underneath.agency/resources/do-people-trust-ai-search-less): Slightly. In US experiments, the same text was trusted a little less when labeled AI, and a week of Google’s AI Mode lowered trust in Google. (updated 2026-10-08)
- [Does advertising spend help our brand get recommended by AI, or does something else matter more?](https://underneath.agency/resources/does-ad-spend-help-ai-brand-recommendations): Not directly, on current evidence: search interest and online conversation tracked AI brand prominence far more closely than advertising spend did. (updated 2026-10-08)
- [Does AI search actually use the pages it cites?](https://underneath.agency/resources/does-ai-search-use-the-pages-it-cites): Not evenly. AI search engines lean heavily on some cited pages and barely use others, so a citation count overstates how much a page shaped the answer. (updated 2026-10-07)
- [Does showing up in AI answers actually drive business results?](https://underneath.agency/resources/does-ai-visibility-drive-business-results): Possibly, but not yet proven. A 2026 review of 45 studies found that claims about GEO’s return outstrip the evidence; none tied AI visibility to sales. (updated 2026-10-08)
- [Does blocking Google-Extended in robots.txt reduce our visibility in AI Overviews?](https://underneath.agency/resources/does-blocking-google-extended-hurt-ai-overviews): Possibly. A large 2025 study linked blocking Google-Extended to fewer AI Overview citations; our 2026 check saw no clear effect. Weigh it as a real risk. (updated 2026-10-10)
- [Can restructuring existing content, without changing what it says, increase AI citations?](https://underneath.agency/resources/does-content-structure-increase-ai-citations): Possibly, but the evidence is split. One lab test found a 17.3% lift from restructuring alone; a 252,000-trial test found formatting had no effect. (updated 2026-10-08)
- [Is it true that GEO increases AI visibility by 40%?](https://underneath.agency/resources/does-geo-increase-ai-visibility-by-40-percent): Not as usually stated. The 40% comes from a 2023 lab test in which a page already given to the AI won more of the answer’s text, not more traffic. (updated 2026-10-08)
- [Does keyword stuffing still work in AI search?](https://underneath.agency/resources/does-keyword-stuffing-work-in-ai-search): No. In published tests, repeating query keywords did nothing or lowered a page’s share of AI answers; relevance and search ranking mattered more. (updated 2026-10-08)
- [Does list order change what an AI assistant recommends?](https://underneath.agency/resources/does-list-order-change-ai-recommendations): Yes, in controlled tests: options shown first are picked more often. One hotel study valued first place at $11.7 a night; the size varies widely by AI model. (updated 2026-10-10)
- [Does the way customers phrase a question change which sources AI search cites?](https://underneath.agency/resources/does-question-phrasing-change-ai-sources): Yes. In a Gemini hotel test, experience-style questions drew 55.9% of citations from non-booking sites, against 30.8% for booking-style ones. (updated 2026-10-10)
- [If we rank on Google, will we show up in AI Overviews?](https://underneath.agency/resources/does-ranking-on-google-get-you-into-ai-overviews): Not reliably. In our test only 28.7% of AI Overview citations were page-one results, though the top result was cited in 49.5% of AI Overviews. (updated 2026-10-10)
- [Does readable writing help your content appear in AI answers?](https://underneath.agency/resources/does-readable-writing-help-ai-visibility): Somewhat. AI engines cite easier-to-read pages, and fluent rewrites helped in early lab tests, but later tests found polishing text alone rarely helps. (updated 2026-10-08)
- [Does Reddit and forum content actually shape Google AI Overview answers?](https://underneath.agency/resources/does-reddit-shape-google-ai-overviews): Less than citations suggest. Google’s AI Overviews cite Reddit, but one study found they drew 22.1 points less content from forums than from other sources. (updated 2026-10-10)
- [Does social media content help your brand show up in AI search?](https://underneath.agency/resources/does-social-media-help-ai-search-visibility): Mostly on Google. AI Overviews often cite Facebook, Instagram and YouTube; ChatGPT and Claude rarely cite social platforms at all. (updated 2026-10-07)
- [Can AI recommendations lower what a DTC brand pays for each new customer?](https://underneath.agency/resources/dtc-brands-ai-search-sales): Possibly. AI-referred shoppers bring a high share of new buyers and convert well, but the channel is still small next to paid social and search. (updated 2026-10-08)
- [How does an ecommerce brand get its products chosen by AI shopping assistants?](https://underneath.agency/resources/ecommerce-brands-ai-shopping-assistants): By making every listing, retailer page and review source agree, so ChatGPT, Google, Rufus and Perplexity can find, compare and trust your products. (updated 2026-10-08)
- [How do ecommerce platforms and apps win merchants who ask AI?](https://underneath.agency/resources/ecommerce-platform-ai-search): By being named when merchants and their agencies ask AI which platform or app to use, and by proving you can put merchants’ products in AI shopping. (updated 2026-10-08)
- [Can AI search bring an EdTech company more schools, teachers and families?](https://underneath.agency/resources/edtech-ai-search): Partly. Teachers, students and families already ask AI; districts still buy on evidence, privacy and state lists, so AI answers must find that proof. (updated 2026-10-08)
- [How do endpoint security vendors win enterprise deals when CISOs ask AI which EDR to buy?](https://underneath.agency/resources/endpoint-security-enterprise-customers-ai-search): Endpoint security vendors earn AI shortlist places through public test results, incident transparency and consistent facts, not self-ranked lists. (updated 2026-10-08)
- [Will clients find our engineering firm when they ask AI who can design their project?](https://underneath.agency/resources/engineering-firms-project-inquiries-ai-search): Engineering firms get named by AI when their project types, licenses, people and locations are stated clearly and confirmed by sources clients already trust. (updated 2026-10-08)
- [When engineers ask AI which CAD, simulation or PLM tool to use, will they hear our name?](https://underneath.agency/resources/engineering-software-ai-search): Engineering software gets recommended by AI when its workflows, file formats, plans and limits are stated plainly and confirmed by sources engineers trust. (updated 2026-10-08)
- [Is an English-only AI visibility check enough if you sell abroad?](https://underneath.agency/resources/english-only-ai-visibility-audits): No. In a 12-language European test, AI named local brands far more often when asked in their home language, so English-only checks undercount them. (updated 2026-10-08)
- [Are AI assistants now shaping which enterprise software gets shortlisted?](https://underneath.agency/resources/enterprise-software-shortlists-ai-search): Increasingly, yes. Enterprise buyers research vendors with AI assistants, often private ones at work, before any seller hears about the deal. (updated 2026-10-08)
- [How do ERP vendors get on the shortlist when buyers ask AI?](https://underneath.agency/resources/erp-software-ai-search): By being on the long list AI assistants draw from analyst, consultant and comparison sources, with clear industry fit, partners and implementation facts. (updated 2026-10-08)
- [Will AI search change how boards choose an executive search firm?](https://underneath.agency/resources/executive-search-firms-clients-ai-search): Referrals still decide most mandates, but directors now use AI to check firms and partners. Public, verifiable track records decide what those checks find. (updated 2026-10-08)
- [Can fake reviews on the web make AI assistants recommend a fake brand over ours?](https://underneath.agency/resources/fake-reviews-fake-brands-ai-recommendations): Yes. In tests of 12 AI models, planted review pages got a non-existent brand recommended in up to 73.8% of products, and no defense tested fixed it. (updated 2026-10-08)
- [How does a fashion brand get recommended when shoppers ask AI what to wear?](https://underneath.agency/resources/fashion-ecommerce-ai-recommendations): By describing products in shoppers’ style and occasion language, with clear images and consistent details everywhere AI shopping tools and visual search look. (updated 2026-10-08)
- [How can a financial consulting firm win CFO, deal and restructuring work through AI search?](https://underneath.agency/resources/financial-consulting-firms-clients-ai-search): By being the firm AI names for a specific finance problem, company size and deal stage, with a reputation, credentials and current commentary that hold up. (updated 2026-10-08)
- [Are AI assistants changing how banks and finance teams choose fintech software?](https://underneath.agency/resources/fintech-software-customers-from-ai-search): Yes. Finance teams and bank staff research vendors with AI assistants, and in finance those answers lean on trusted, regulated and well-ranked sources. (updated 2026-10-08)
- [How can a fintech startup get recommended by AI when established brands own the answers?](https://underneath.agency/resources/fintech-startups-customers-ai-search): By owning the category explainer, earning independent coverage early and stating fees plainly. AI favors known brands until it finds a clear reason not to. (updated 2026-10-08)
- [How do you fix wrong information about your brand in AI answers?](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers): Trace the error to the page it came from and fix that page. AI answers are built mostly from live web sources, often your own old pages. (updated 2026-10-08)
- [How can food ecommerce brands turn AI meal planning into orders?](https://underneath.agency/resources/food-ecommerce-sales-from-ai-search): By being the products AI names when shoppers plan meals and build lists: complete food data, compliant claims, reviews and presence where carts are filled. (updated 2026-10-08)
- [How can a footwear brand get its shoes found when shoppers ask AI what to buy?](https://underneath.agency/resources/footwear-brands-ai-search): By making each shoe model easy for AI to match to a fit, use and price question, so it lands on the shortlist before shoppers pick a size and a store. (updated 2026-10-08)
- [How do freight carriers, brokers and forwarders win shippers who research with AI?](https://underneath.agency/resources/freight-companies-shippers-ai-search): Shippers check freight providers through rankings, reviews and quotes; AI answers now sit on that path, so carriers and brokers must be easy to verify by lane. (updated 2026-10-08)
- [How can a furniture brand turn AI shopping answers into more high-ticket orders?](https://underneath.agency/resources/furniture-brands-sales-from-ai-search): By being named when shoppers ask AI to narrow a sofa or bedroom search, with product facts, reviews and delivery terms an assistant can read and check. (updated 2026-10-08)
- [GEO agency vs SEO agency](https://underneath.agency/resources/geo-agency-vs-seo-agency): An SEO agency ranks pages in search results; a GEO agency gets a brand cited and recommended in AI answers. What each does, what they cost, and which you need. (updated 2026-09-28)
- [How should global brands approach GEO across languages?](https://underneath.agency/resources/geo-strategy-for-global-brands-across-languages): Plan market by market: AI answers change with language and country, so earn local coverage, write native content and measure each market separately. (updated 2026-10-08)
- [How much control do we have when Google changes how search results look?](https://underneath.agency/resources/google-search-design-changes-traffic-risk): Very little: Google can shift traffic with design changes publishers cannot control. AI Overviews cut English Wikipedia’s search visits by about 5%. (updated 2026-10-08)
- [How can a health insurer win new members when shoppers ask AI to explain their plan choices?](https://underneath.agency/resources/health-insurers-members-ai-search): By making plan facts, networks and ratings accurate and public before enrollment opens, because shoppers now ask AI tools to explain and compare coverage. (updated 2026-10-08)
- [How can a hospital, clinic or practice win new patients from AI search?](https://underneath.agency/resources/healthcare-providers-patients-ai-search): By making providers easy for AI to name and describe correctly: full physician profiles, consistent facts, strong reviews. Patients now ask AI who to see. (updated 2026-10-08)
- [Can AI search help a healthcare software company win health system deals?](https://underneath.agency/resources/healthcare-software-ai-search): Yes, if AI answers can find proof of EHR integration, HIPAA security and peer results, because hospitals now narrow long vendor lists before any demo. (updated 2026-10-08)
- [Can AI search help a healthcare startup compete with established health brands?](https://underneath.agency/resources/healthcare-startups-demand-ai-search): It can, if AI answers find independent proof of outcomes; buyers now pay for results and patients ask AI first, but new companies are rarely named unprompted. (updated 2026-10-08)
- [How does a healthtech company get named when employers and health plans ask AI for vendors?](https://underneath.agency/resources/healthtech-employers-payers-ai-search): With independent evidence, plain outcome and cost facts, and coverage buyers trust. Employers are cutting vendors, and AI is an early place they look. (updated 2026-10-08)
- [How do home decor brands get found when shoppers ask AI for a look, not a product?](https://underneath.agency/resources/home-decor-brands-ai-product-discovery): By showing up in the visual product sets AI tools build from style questions, with images, colors, materials and reviews an assistant can read. (updated 2026-10-08)
- [How do brands build authority that AI search recognizes?](https://underneath.agency/resources/how-brands-build-authority-for-ai-search): By being written about by sources AI engines already trust, in the outlets each engine reads, and by keeping a site AI agents can read. (updated 2026-10-07)
- [How many prompts do we need to track to measure AI visibility?](https://underneath.agency/resources/how-many-prompts-to-track-ai-visibility): There is no fixed number. Studies suggest about 40 to 150 or more prompts per topic per AI engine, each asked several times, before a ranking is trustworthy. (updated 2026-10-08)
- [How many sources does each AI search engine cite per answer?](https://underneath.agency/resources/how-many-sources-ai-search-engines-cite): Anywhere from about 3 to 40 sources per answer, depending on the engine, the topic and how it was measured. The engine ranking flips between studies. (updated 2026-10-11)
- [How much web content is already optimized for AI search?](https://underneath.agency/resources/how-much-web-content-is-optimized-for-ai-search): About 9% of pages Google and Gemini returned for 1,000 real searches showed signs of AI-search optimization, rising to 16.36% for pages updated in 2026. (updated 2026-10-08)
- [How can a small brand get recommended by AI assistants?](https://underneath.agency/resources/how-small-brands-get-recommended-by-ai): Small brands start far behind in AI answers. The best evidence says to get written about by independent sites and make your advantages easy to verify. (updated 2026-10-08)
- [How to choose a GEO agency](https://underneath.agency/resources/how-to-choose-a-geo-agency): How to choose a generative engine optimization (GEO) agency in 2026: eight criteria, a scorecard, the red flags, and the questions to ask before you sign. (updated 2026-09-28)
- [How should you design prompts and runs to track AI visibility?](https://underneath.agency/resources/how-to-design-ai-visibility-tracking): Spread your budget across languages, assistants and wordings before buying more repeats of one prompt, and fix the question set and run count in advance. (updated 2026-10-08)
- [How can an HR consulting firm win clients through AI search?](https://underneath.agency/resources/hr-consulting-firms-clients-ai-search): By being the firm AI answers connect to a named people problem, such as pay transparency or work redesign, backed by research and coverage others cite. (updated 2026-10-08)
- [Is AI search deciding which HR software gets the demo?](https://underneath.agency/resources/hr-software-ai-search): Increasingly it shapes who gets considered. HR software vendors win demos when AI answers can verify their fit, security and compliance. (updated 2026-10-08)
- [Will AI assistants name your identity platform when a credential breach sends buyers looking?](https://underneath.agency/resources/iam-enterprise-demand-ai-search): Identity vendors get named in AI answers when their standards support, security record and facts are public, consistent and confirmed by independent sources. (updated 2026-10-08)
- [Will AI help put our pumps, compressors and valves on the bid list?](https://underneath.agency/resources/industrial-equipment-leads-ai-search): Equipment makers reach more bid lists when AI can verify their duty ranges, efficiency data, standards and local service, the facts engineers specify on. (updated 2026-10-08)
- [When an engineer asks AI for a part, will it name ours?](https://underneath.agency/resources/industrial-manufacturers-ai-search): By making every part number, spec and equivalent easy for AI to verify, on your site and in distributor catalogs, before engineers design it in. (updated 2026-10-08)
- [Will plant leaders find our MES or IIoT platform when they ask AI for options?](https://underneath.agency/resources/industrial-software-plants-ai-search): When plant teams research MES, maintenance or industrial data platforms with AI, the vendors named are those whose plant-level facts are easy to verify. (updated 2026-10-08)
- [How can an insurance carrier or agency win more quotes when shoppers ask AI?](https://underneath.agency/resources/insurance-companies-customers-ai-search): By being named, correctly and state by state, when shoppers ask AI to explain and compare coverage. 32% of auto shoppers already use AI tools. (updated 2026-10-08)
- [How can an insurtech win customers from incumbent insurers when shoppers ask AI?](https://underneath.agency/resources/insurtech-customers-ai-search): By being the clear, well-reviewed alternative AI can explain when shoppers ask about switching. Insurify says AI-sourced new customers rose from 0.15% to 5.7%. (updated 2026-10-08)
- [How do investing apps and robo-advisors get chosen when people ask AI which platform to use?](https://underneath.agency/resources/investment-platforms-customers-ai-search): By being the clear answer for a specific use case, with fees and safety facts stated plainly and confirmed by the comparison sites AI cites. (updated 2026-10-08)
- [Is a GEO agency worth it?](https://underneath.agency/resources/is-a-geo-agency-worth-it): Is a GEO agency worth it? When it pays back, when it does not, what it costs in 2026, and how the AI engines themselves answer the question. (updated 2026-09-27)
- [Is ChatGPT replacing Google search?](https://underneath.agency/resources/is-chatgpt-replacing-google): Not on current evidence. People use ChatGPT alongside Google: search appeared in 46.0% of assistant sessions in a 2026 US and UK browsing panel. (updated 2026-10-10)
- [Is traditional SEO still important for visibility in AI search?](https://underneath.agency/resources/is-seo-still-important-for-ai-search): Yes. Pages that rank higher are cited more often, but Google rankings explain only part of AI citations, and less for ChatGPT than for Google’s own AI. (updated 2026-10-11)
- [Is tracking ChatGPT alone enough to measure our AI search visibility?](https://underneath.agency/resources/is-tracking-chatgpt-enough): No. AI engines cite largely different sources and pick different brands for the same question, so ChatGPT alone shows less than half the picture. (updated 2026-10-08)
- [How can an IT consulting firm win more advisory clients when mid-size companies ask AI first?](https://underneath.agency/resources/it-consulting-firms-clients-ai-search): By being the independent adviser AI names for a company’s size, industry and decision, such as an ERP choice or a fractional CIO, with proof buyers can check. (updated 2026-10-08)
- [How can an IT services firm win more projects and support calls from AI search?](https://underneath.agency/resources/it-service-firms-leads-ai-search): By being the local or specialist firm AI names for a specific problem, place and project, backed by reviews, listings and clear proof of what you fix. (updated 2026-10-08)
- [How can a jewelry brand get recommended when shoppers ask AI?](https://underneath.agency/resources/jewelry-brands-ai-product-recommendations): By being the jeweler AI names for ring and gift questions, with current facts on certification, sourcing, price and where to buy, before peak seasons. (updated 2026-10-08)
- [How can legal software companies win law firms and in-house teams through AI search?](https://underneath.agency/resources/legal-software-ai-search): By being named, accurately, when lawyers ask AI which tools fit their practice, with public proof of accuracy, confidentiality and ethics that buyers can check. (updated 2026-10-08)
- [What separates legitimate GEO from manipulation?](https://underneath.agency/resources/legitimate-geo-vs-manipulation): Truth and transparency. Legitimate GEO makes real, verifiable facts easier for AI to find; manipulation fakes evidence, hides commands or hides motives. (updated 2026-10-08)
- [How do lending platforms reach borrowers who ask AI where to get a loan?](https://underneath.agency/resources/lending-platforms-borrowers-ai-search): By being named, with accurate rates, terms and reputation, when borrowers ask AI where to get a loan. Visibility brings applications, never approvals. (updated 2026-10-08)
- [When a shipper asks AI for a 3PL, does our warehouse network make the list?](https://underneath.agency/resources/logistics-providers-shipper-contracts-ai-search): 3PL and fulfillment providers reach more shippers in AI search when locations, capabilities, integrations and outside proof are easy to verify. (updated 2026-10-08)
- [How can logistics software companies win more customers through AI search?](https://underneath.agency/resources/logistics-software-ai-search): By being named, accurately, when shippers, carriers and 3PLs ask AI assistants which TMS, WMS or fleet tool fits, then turning that into demos. (updated 2026-10-08)
- [How do luxury brands win the high spenders who now ask AI what to buy?](https://underneath.agency/resources/luxury-brands-ai-search): By shaping the sources AI cites when high spenders ask open questions: most luxury AI prompts name no brand, and official sites are a minority of citations. (updated 2026-10-08)
- [How do machine learning companies get found when enterprise buyers ask AI who to hire?](https://underneath.agency/resources/machine-learning-companies-ai-search): By earning the independent proof AI assistants draw on, so they name you when enterprises ask who can build or supply a machine learning solution. (updated 2026-10-08)
- [Will AI steer a plant’s next machine purchase toward us or a rival?](https://underneath.agency/resources/machinery-companies-buyers-ai-search): Machinery makers get shortlisted when AI can verify their applications, cycle times, prices, financing and dealer support, the facts a capital case rests on. (updated 2026-10-08)
- [When a CEO asks AI which consulting firm to hire, will ours be named?](https://underneath.agency/resources/management-consulting-firms-clients-ai-search): Only if its reputation is visible in sources AI reads: rankings, client perception studies, case evidence and independent coverage, not just its own claims. (updated 2026-10-08)
- [How do marketing software companies win buyers who ask AI first?](https://underneath.agency/resources/marketing-software-ai-search-growth): By being named when marketing leaders ask AI to map a category: most CMOs now start software searches in AI tools, and they shortlist about three vendors. (updated 2026-10-08)
- [How does an online marketplace win more buyers and sellers when people ask AI?](https://underneath.agency/resources/marketplace-buyers-sellers-ai-search): By being the answer on both sides: where buyers should shop and where sellers should list. AI traffic is still small for marketplaces, but it converts. (updated 2026-10-08)
- [How can a mattress brand win customers who ask AI which mattress to buy?](https://underneath.agency/resources/mattress-brands-customers-from-ai-search): By being on the short list when shoppers ask AI to compare mattresses, with verifiable facts, real reviews, clear trial terms and no unsupported sleep claims. (updated 2026-10-08)
- [How can MDR and managed SOC providers win qualified leads from AI search?](https://underneath.agency/resources/mdr-providers-qualified-leads-ai-search): By being named, with defined response metrics and independent proof, when stretched IT and security leaders ask AI which MDR or managed SOC fits them. (updated 2026-10-08)
- [How should we measure share of citations in AI search tools?](https://underneath.agency/resources/measuring-share-of-citations-in-ai-search): Count every answer, including those that never searched the web, report each engine separately, and use enough repeated answers to tell real gaps from noise. (updated 2026-10-08)
- [Can AI search build demand for a medical device before the value analysis committee meets?](https://underneath.agency/resources/medical-device-demand-ai-search): It can shape what clinicians and hospital committees read about your device, if peer-reviewed evidence and clear FDA facts are easy for AI tools to find. (updated 2026-10-08)
- [How can a medical equipment company win more hospital RFQs when buyers ask AI first?](https://underneath.agency/resources/medical-equipment-leads-ai-search): Hospital and clinic buyers are adopting AI fast. Equipment sellers win RFQs when assistants can verify their prices, contracts, service and evidence. (updated 2026-10-08)
- [Can AI search bring independent practices and specialty clinics to a medical software vendor’s demo?](https://underneath.agency/resources/medical-practice-software-ai-search): Yes, if assistants find proof of specialty fit, certification, pricing and peer ratings; 23% of medical groups expect to switch or overhaul their EHR in a year. (updated 2026-10-08)
- [How do mortgage lenders and marketplaces win qualified borrowers when buyers start with AI?](https://underneath.agency/resources/mortgage-platforms-leads-ai-search): By being named, with accurate loan facts, when borrowers ask AI to estimate payments and compare lenders, and by keeping rate claims compliant. (updated 2026-10-08)
- [Which AI engine gives the most consistent answers about brands?](https://underneath.agency/resources/most-consistent-ai-engine-for-brands): No AI engine is steadiest on every measure. Perplexity repeated brand lists most in two tests; Gemini and ChatGPT described brands most steadily in another. (updated 2026-10-08)
- [How can a managed service provider win more clients from AI search?](https://underneath.agency/resources/msps-customers-from-ai-search): By being named when owners and IT leads describe their size, region and needs to AI, with pricing, coverage and security proof they can check. (updated 2026-10-08)
- [How does a neobank turn AI answers into accounts that receive a paycheck?](https://underneath.agency/resources/neobanks-account-openings-ai-search): By answering the safety and fee questions AI users ask about app-only banks, so the account they open becomes the one their paycheck goes to. (updated 2026-10-08)
- [How do network security vendors win buyers who ask AI about zero trust?](https://underneath.agency/resources/network-security-zero-trust-customers-ai-search): By being named, in NIST and CISA terms, in AI answers about firewall refreshes, SASE and zero trust access, with a public security record buyers can check. (updated 2026-10-08)
- [Do off-site mentions help once an AI agent is reading about you?](https://underneath.agency/resources/off-site-mentions-vs-site-readability-ai-agents): Less than you might think. Once an AI agent researches a named business, whether it can read the site mattered far more than off-site mentions. (updated 2026-10-08)
- [What on-page signals are linked to citations in Google AI Overviews and Perplexity?](https://underneath.agency/resources/on-page-signals-linked-to-ai-citations): Dates, clean HTML structure and structured data track with AI citations, but ranking matters more, and only a machine-readable date held up on Google. (updated 2026-10-08)
- [Can I trust a one-time AI visibility report for my brand?](https://underneath.agency/resources/one-time-ai-visibility-report): Not on its own. AI answers vary between runs, so a one-off report can mistake noise for a lead. Reliable figures need repeated runs over two to four weeks. (updated 2026-10-08)
- [How does an outdoor brand get onto the shortlist when people ask AI what gear to buy?](https://underneath.agency/resources/outdoor-brands-ai-search): By matching gear to specific activities and conditions, and earning the expert reviews and retailer guides AI draws on when it builds an outdoor shortlist. (updated 2026-10-08)
- [When a brand asks AI for a packaging supplier, will our name come up?](https://underneath.agency/resources/packaging-suppliers-quote-requests-ai-search): Brand owners now use AI in packaging sourcing. Packaging suppliers win quote requests when assistants can verify their materials, compliance and minimums. (updated 2026-10-08)
- [When a new business asks AI how to take payments, how does a payment company get chosen?](https://underneath.agency/resources/payment-companies-merchants-ai-search): By being named, with the right fees, when a new business asks AI how to take payments. Owners choose on cost and simplicity, and rarely switch later. (updated 2026-10-08)
- [How does a payment processor get onto the shortlist when merchants research with AI?](https://underneath.agency/resources/payment-processors-merchant-demand-ai-search): By being named and described accurately when payments teams, engineers and procurement research processors with AI, from pricing models to local acquiring. (updated 2026-10-08)
- [How does a payroll provider get named when businesses ask AI?](https://underneath.agency/resources/payroll-software-ai-search): By being the clear, verifiable answer for a business’s size, states and countries. AI answers on tax-bound questions shift by location and misquote prices. (updated 2026-10-08)
- [How can a penetration testing firm win more scoped leads from AI search?](https://underneath.agency/resources/pentest-firms-leads-ai-search): By being the firm AI names when a buyer describes their audit, scope and deadline, with credentials, methods and pricing logic others can verify. (updated 2026-10-08)
- [How do pet product brands get recommended when owners ask AI what to buy?](https://underneath.agency/resources/pet-product-brands-ai-search): By being the brand AI names, with accurate facts and claims, when owners ask what to buy for their pet; one recommendation can start years of repeat orders. (updated 2026-10-08)
- [How can a drugmaker keep its brands visible and accurate in AI answers?](https://underneath.agency/resources/pharma-brand-visibility-ai-search): By making the label, the safety story and the brand’s facts easy for AI to find and repeat, within FDA rules, since patients and doctors now ask AI first. (updated 2026-10-08)
- [How do you sell procurement software to buyers who run RFPs for a living?](https://underneath.agency/resources/procurement-software-ai-search): Professional buyers still run RFPs, but AI answers now shape which procurement tools reach the longlist, so vendors need proof buyers and AI can check. (updated 2026-10-08)
- [What kind of product content do AI shopping assistants prefer?](https://underneath.agency/resources/product-content-ai-shopping-assistants-prefer): Factual, specific, easy-to-compare content matched to what buyers ask. In tests, hard product facts drove 82.4% of AI rankings; hype and gimmicks backfired. (updated 2026-10-08)
- [How do productivity tools win users when AI assistants compete?](https://underneath.agency/resources/productivity-software-ai-search-growth): By being the tool AI assistants name for one job done better, then converting on day one. AI already sends productivity apps real traffic; suites raise the bar. (updated 2026-10-08)
- [How do professional services firms win clients when buyers ask AI who to hire?](https://underneath.agency/resources/professional-services-firms-clients-ai-search): Clients of law, accounting, design and consulting firms now research with AI first. Firms get named when directories, reviews and coverage confirm expertise. (updated 2026-10-08)
- [How do project management tools win customers from AI answers?](https://underneath.agency/resources/project-management-software-customers-ai-search): By being named for the use-case and team-size questions buyers ask AI, then turning free signups into paid seats that expand. Here is the evidence. (updated 2026-10-08)
- [How do we prove that GEO caused a change in sales or conversions?](https://underneath.agency/resources/prove-geo-caused-sales): You need a comparison group that got no GEO work, ideally chosen at random, because AI traffic grows for everyone. Before-and-after numbers overstate it. (updated 2026-10-08)
- [Can AI search bring a recruiting agency more employer clients?](https://underneath.agency/resources/recruiting-agencies-employer-clients-ai-search): Yes, if employers can see what roles you fill, where, how fast and how well. AI answers lean on local listings, reviews and specialty pages to name agencies. (updated 2026-10-08)
- [What are the risks of native ads blended into AI chatbot answers?](https://underneath.agency/resources/risks-of-native-ads-in-ai-answers): Native ads in AI answers are hard to tell apart from advice, so they can mislead trusting users and damage the advertiser’s brand once spotted. (updated 2026-10-08)
- [Will AI assistants put our robots on a manufacturer’s shortlist before an integrator is called?](https://underneath.agency/resources/robotics-companies-customers-ai-search): It can, if assistants find your applications, specs, integrators and real deployments in sources they trust, as robot buying spreads beyond automotive. (updated 2026-10-08)
- [How can sales software companies turn AI answers into pipeline?](https://underneath.agency/resources/sales-software-pipeline-from-ai-search): By being named when sales leaders ask AI to shortlist tools, then answering the demo requests that follow fast, with proof on data, security and fit. (updated 2026-10-08)
- [How do security awareness training vendors win customers through AI search?](https://underneath.agency/resources/security-awareness-training-customers-ai-search): By being named, with checkable proof, when IT leads, MSPs and CISOs ask AI which training meets their insurer, auditor and budget. (updated 2026-10-08)
- [How do parcel carriers and shipping platforms win small businesses that ask AI how to ship?](https://underneath.agency/resources/shipping-companies-customers-ai-search): Small sellers now ask AI how to ship cheaply, so parcel carriers and shipping platforms need current, citable rates, coverage and reviews to be named. (updated 2026-10-08)
- [Should ads in AI search be clearly labeled, and does labeling help the advertiser?](https://underneath.agency/resources/should-ai-search-ads-be-labeled): Yes. Clear labels protect users, and research argues they protect advertisers too, since disguised ads breed distrust once people spot them. (updated 2026-10-08)
- [How do SIEM vendors get onto enterprise shortlists through AI search?](https://underneath.agency/resources/siem-enterprise-pipeline-ai-search): By being named, with checkable proof, when SOC leaders use AI to build a SIEM long list during a migration, a consolidation or a cost review. (updated 2026-10-08)
- [How do skincare brands get their products found when shoppers ask AI about their skin?](https://underneath.agency/resources/skincare-brands-ai-search): By answering concern and ingredient questions with claims you can substantiate, real expert backing and honest reviews, the signals AI and shoppers both check. (updated 2026-10-08)
- [How do sportswear brands win the sale when shoppers ask AI what to train in?](https://underneath.agency/resources/sportswear-brands-ai-search): By being the product AI names for specific activity and attribute questions, with data clean enough to buy from your site, a retailer or the chat itself. (updated 2026-10-08)
- [How do subscription brands win new subscribers through AI search?](https://underneath.agency/resources/subscription-ecommerce-customers-ai-search): Be the clear, honest answer when shoppers compare plans, prices and cancellation terms; 55% of consumers would welcome AI reordering everyday items. (updated 2026-10-08)
- [How do supplement brands win new customers from AI assistants?](https://underneath.agency/resources/supplement-brands-customers-ai-search): By being the brand assistants can verify: third-party testing, honest labels and strong reviews, now that 32% of US adults ask AI chatbots about health. (updated 2026-10-08)
- [How does a supply chain consultancy get on the shortlist when executives research with AI?](https://underneath.agency/resources/supply-chain-consulting-firms-ai-search): AI answers now help shape which supply chain consultancies get invited to bid, so firms need proof that independent sources repeat, not just a good website. (updated 2026-10-08)
- [Can AI search put our supply chain software on the RFP longlist?](https://underneath.agency/resources/supply-chain-software-ai-search): Partly: AI answers now shape the vendor longlist before an RFP is written, so supply chain vendors must be named and verifiable there first. (updated 2026-10-08)
- [How does a technology consultancy get onto enterprise shortlists when buyers ask AI first?](https://underneath.agency/resources/technology-consulting-firms-ai-search): By being the firm AI answers tie to proven production results in a named platform and industry, backed by analyst, partner and client evidence. (updated 2026-10-08)
- [How do telehealth companies get chosen when patients ask AI where to get care?](https://underneath.agency/resources/telehealth-patients-from-ai-search): By being the service assistants can verify on price, insurance, licensing and reputation, as 58% of AI health users follow up with a provider. (updated 2026-10-10)
- [How do toy brands get found when shoppers ask AI for gift ideas?](https://underneath.agency/resources/toy-brands-product-discovery-ai-search): By making age, play value and safety facts easy for AI to match to gift questions, and earning the reviews and gift-guide coverage AI answers draw on. (updated 2026-10-10)
- [How does AI search bring a virtual care platform new contracts and more patient visits?](https://underneath.agency/resources/virtual-care-platforms-ai-search): By getting the platform onto health system, plan and employer shortlists, and by helping clients’ patients find the virtual front door they paid for. (updated 2026-10-10)
- [Are AI assistants deciding which warehouse automation vendors make the enterprise shortlist?](https://underneath.agency/resources/warehouse-automation-enterprise-leads-ai-search): Partly: buyers now use AI to learn options and build longlists years before an RFP, so AS/RS, AMR and sortation vendors must be named and verifiable there. (updated 2026-10-10)
- [How do wealth management firms win new clients when prospects check them with AI before calling?](https://underneath.agency/resources/wealth-management-firms-clients-ai-search): Referrals still lead, but prospects increasingly check a firm with AI. Firms win by making their niche, fees, fiduciary status and reviews easy to verify. (updated 2026-10-10)
- [Are websites hiding instructions to manipulate how AI search presents them?](https://underneath.agency/resources/websites-hiding-instructions-for-ai-search): Yes. A scan of 1.2 billion web addresses found 1,521 hidden instructions trying to make AI systems promote, cite or praise a page. Most rarely work. (updated 2026-10-08)
- [Will AI assistants recommend our wellness products when people ask how to sleep, train or recover better?](https://underneath.agency/resources/wellness-brands-customers-ai-search): By being the product assistants can describe and verify when people ask how to sleep, train or recover better, with wellness claims that stay lawful. (updated 2026-10-10)
- [What does content that AI engines prefer look like?](https://underneath.agency/resources/what-content-do-ai-engines-prefer): AI engines favor pages that state the conclusion first, cover the topic fully, explain how and why, and pack checkable facts such as definitions and figures. (updated 2026-10-08)
- [What does a GEO agency do?](https://underneath.agency/resources/what-does-a-geo-agency-do): A GEO agency gets a brand cited, mentioned and recommended in AI answers. The nine things it does, what it costs, and how the AI engines describe the job. (updated 2026-09-27)
- [What actually drives which product an AI assistant recommends?](https://underneath.agency/resources/what-drives-ai-product-recommendations): When AI assistants can see ratings, prices and reviews, that data decides the pick. Brand name mostly breaks ties when products look the same. (updated 2026-10-08)
- [What GEO practices are actually supported by research?](https://underneath.agency/resources/what-geo-practices-does-research-support): Research supports a conservative GEO playbook: relevant, complete, verifiable, well-structured pages that AI can reach, measured stage by stage. Tricks fail. (updated 2026-10-08)
- [What happens to brands that skip GEO when competitors adopt it?](https://underneath.agency/resources/what-happens-if-you-skip-geo): In controlled tests, brands that did nothing while rivals optimized got zero AI recommendations. The early lead shrinks fast, so copying rivals is not enough. (updated 2026-10-08)
- [What is GEO, and how is it different from SEO?](https://underneath.agency/resources/what-is-generative-engine-optimization): GEO is the work of getting your information found, used and cited in AI-written answers. SEO competes for a rank; GEO competes for a share of the answer. (updated 2026-10-08)
- [What should we measure to know whether we are visible in AI answers?](https://underneath.agency/resources/what-to-measure-ai-visibility): Measure how often each AI engine names and cites you in final answers, over repeated runs, and whether what it says is right, not rankings or one screenshot. (updated 2026-10-08)
- [If AI agents can’t read my website, where does their answer about us come from?](https://underneath.agency/resources/when-ai-agents-cant-read-your-site): From web searches and other people’s pages. In one test, 42% of an AI agent’s answer came from outside a business whose site it could not read. (updated 2026-10-11)
- [Which Google searches trigger an AI Overview?](https://underneath.agency/resources/which-searches-trigger-ai-overviews): Long, specific questions trigger Google AI Overviews most; local, brand and breaking-news searches rarely do. Overall rates run from 13.7% to 67%. (updated 2026-10-10)
- [Which types of websites do AI search engines rely on most?](https://underneath.agency/resources/which-websites-do-ai-search-engines-cite): It depends on the question. Official, reference and news sites lead on public-interest topics; review publishers lead on buying questions. (updated 2026-10-11)
- [Who in my organization should own AI search visibility?](https://underneath.agency/resources/who-should-own-ai-search-visibility): No single team: SEO owns being found, content and product own what pages say, PR owns third-party coverage, and one person should own measurement. (updated 2026-10-08)
- [Why does ChatGPT give a different answer about my brand each time?](https://underneath.agency/resources/why-ai-answers-about-your-brand-change): AI assistants pick each answer partly by chance, so a brand named once may vanish on the next run. Judge visibility by rates over many runs. (updated 2026-10-08)
- [What decides whether an AI engine cites my page over a competitor’s?](https://underneath.agency/resources/why-ai-cites-competitor-page-first): In a 252,000-trial test, topic match, list position, a stated price and a recent date decided which page AI cited first. Formatting barely mattered. (updated 2026-10-08)
- [What makes ChatGPT or Gemini recommend one hotel over another?](https://underneath.agency/resources/why-ai-recommends-one-hotel-over-another): Guest rating and price decide most AI hotel picks; eco-certification and review count help, list order matters, and replying to reviews did nothing. (updated 2026-10-08)
- [Why do AI search engines cite different sources every time I check?](https://underneath.agency/resources/why-ai-search-citations-change): Because AI search engines pick sources partly at random each time. In one 45-day study, about 65% of cited sources changed from one day to the next. (updated 2026-10-08)
- [Why can’t our marketing analytics see our visibility in AI answers?](https://underneath.agency/resources/why-analytics-miss-ai-visibility): Analytics log clicks and paid impressions. Most AI answers are read without a click, so brand mentions in them leave little or no trace in your data. (updated 2026-10-08)
- [Why doesn’t ChatGPT mention our newly launched product?](https://underneath.agency/resources/why-chatgpt-misses-new-products): ChatGPT often answers from training data that predates your launch, and even assistants that search rarely surface new products. Here is what helps. (updated 2026-10-08)
- [Why is our well-known brand missing from AI product recommendations?](https://underneath.agency/resources/why-well-known-brands-miss-ai-recommendations): AI assistants often know a famous brand but fail to call it up for a plain category question. Research shows why, and how to check your own brand. (updated 2026-10-08)
- [Why does Wikipedia matter so much for AI search visibility?](https://underneath.agency/resources/why-wikipedia-matters-for-ai-search): Wikipedia is the most cited site in many AI search studies and shapes answers beyond its citations. For buyer questions its share is small. (updated 2026-10-11)

## Machine-readable formats

- Entity page (HTML): https://underneath.agency/agent (text/html)
- Entity page (Markdown): https://underneath.agency/agent.md (text/markdown)
- Entity record (JSON-LD): https://underneath.agency/.well-known/entity.json (application/ld+json)
- Site index for language models: https://underneath.agency/llms.txt (text/plain)
- Every page, concatenated as Markdown: https://underneath.agency/llms-full.txt (text/markdown)
- Sitemap: https://underneath.agency/sitemap.xml (application/xml)
- Practice page (HTML): https://underneath.agency/services/generative-engine-optimization (text/html)
- Every HTML page has a Markdown twin at the same path with `.md` appended (for example https://underneath.agency/about.md), and answers `Accept: text/markdown` with it.

Tagline: Underneath every answer is a better question.



---
title: "Underneath | GEO Agency: Get Cited and Recommended by AI"
description: "Underneath is a generative engine optimization (GEO) agency. We get brands mentioned, cited and recommended by ChatGPT, AI Overviews, Gemini and Perplexity."
canonical: "https://underneath.agency"
published: 2026-09-26
updated: 2026-09-28
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Generative engine optimization (GEO) agency

# The answer is only the surface. *We work underneath it.*

Underneath is a generative engine optimization agency. We help businesses become easier for AI-powered search and answer systems to understand, retrieve, describe and reference, including ChatGPT, Google AI Overviews and AI Mode, Gemini, Claude, Perplexity and Microsoft Copilot.

When someone asks an AI assistant which company to choose, the answer can look simple. But an answer is built from information: claims, sources, entities, relationships, evidence, pages, reviews, the questions people ask, the information competitors have, and the things that are current, missing or conflicting.

**Underneath investigates that information system, changes what matters, and measures what happens next.**

[Talk to us about your business](https://underneath.agency/contact)[See the Underneath Method](https://underneath.agency/approach)

Illustrative example · AI answer · 3 sources

What’s the best accounting software for a growing business?

For growing businesses, YourBrand is frequently recommended, citing its transparent pricing1 and wide range of integrations2. Owners on forums also highlight its onboarding support3

- 1yourbrand.com/pricing
- 2yourbrand.com/integrations
- 3reddit.com/r/smallbusiness

- ChatGPT
- Gemini
- Perplexity
- AI Overviews

The AI assistants and AI search we get you mentioned, cited and recommended in

- ChatGPT
- Google AI Overviews
- Google AI Mode
- Gemini
- Claude
- Perplexity
- Microsoft Copilot
- Bing

Underneath an AI answer

## What happens underneath an AI *answer*?

A buyer asks:

> “What’s the best accounting software for a growing business?”

An AI system may return three recommendations. The customer sees the answer. We look underneath it.

- ### What is the system trying to understand?

  Who are the companies? What category are they in? Who are they for?
- ### What is being said?

  What claims are being made about each company?
- ### What supports those claims?

  Which pages, reviews, publications, communities or other sources provide the evidence?
- ### What is missing or conflicting?

  Where is information incomplete, outdated or inconsistent?
- ### Why does one company appear while another does not?

  The answer tells us where to look. The information underneath tells us what to change.

*Illustrative example. Not a representation of any one platform’s internal process.*

Our method

## Understand first. Then change what *matters*.

We don’t start with a list of pages to publish. We start with the questions your customers ask and the answers they receive. Then we work backwards.

- ### Diagnose

  **Find out what the market, and the machines, actually understand about you**.

  We establish a baseline across AI visibility, search, customer questions, competitors, entities, sources, customer language and technical foundations. We look beyond mentions and rankings to understand the claims, evidence and relationships underneath the answer.
- ### Architect

  **Decide what needs to become true**.

  The diagnosis becomes a prioritized plan across content, technical foundations, entities, sources, evidence and measurement. Every recommendation has a reason. Nothing is added simply to create activity.
- ### Deploy

  **Change the information system around the business**.

  Depending on what the diagnosis finds, that can mean changing important pages, improving information architecture, strengthening technical foundations, clarifying entities, improving external sources or making important claims easier to support and reference. The work follows the problem, not the other way around.
- ### Compound

  **Measure. Learn. Change. Test**.

  AI answers shift. Search results move. Competitors change. Businesses change. So the program does too. Every cycle starts with better evidence than the last.

[Explore the Underneath Method](https://underneath.agency/approach)

What makes us different

## We don’t optimize one page in *isolation*.

Your website is one source. Your reviews are another. Your profiles are another. Your customers are another. Publishers, communities, directories and other sources all contribute to the information environment around your business.

We look at how those pieces fit together. Then we identify where the system is helping you, where it is letting you down, and what deserves attention first.

### We ask seven questions.

- ### Understanding

  Why do your best customers choose you?
- ### Recognition

  Do search engines, AI systems and the sources they read recognize who you are?
- ### Retrieval

  Can the right information be found when it matters?
- ### Evidence

  Are the important claims supported?
- ### Competition

  Where do competitors have a stronger information footprint?
- ### Measurement

  What actually changed?
- ### Learning

  What do we know now that we didn’t know before?

**The questions are simple. Answering them well is the work.**

[See how we work](https://underneath.agency/approach)

Research

## We don’t just follow the industry. We *study* it.

Underneath publishes original research into how AI search and answer systems actually behave. Not opinion pieces dressed up as research. Our studies publish the question, sample, method, limitations and underlying data so the numbers can be checked and reused.

### 23 studies. One principle.

**If we make a claim about how AI systems behave, we want evidence behind it.** Our research has examined:

AI crawlers
:   Which crawlers websites block, how sites serve information to AI agents, and how common formats such as `llms.txt` and Markdown actually are.

AI search
:   How often AI Overviews appear, how their citations relate to traditional rankings, how AI Mode differs from AI Overviews, and which sources appear in answers.

AI assistants
:   How ChatGPT, Gemini, Perplexity and Claude agree, or disagree, on recommendations.

Recommendation stability
:   How answers change when the same question is repeated, reworded or asked in different countries.

Sources and citations
:   Which pages assistants cite, how fresh those sources are, and where those citations come from.

Local AI search
:   How AI recommendations compare with Google Maps and how accurately assistants represent local businesses.

The findings are not a collection of assumptions about what AI “likes.” They are observations from defined datasets.

[Explore Underneath Research](https://underneath.agency/research)

What we’ve learned

## A few things we’ve *learned*.

- ### AI recommendations are not fixed rankings.

  Across 80 buyer questions, all four assistants we studied recommended the same first pick for only 10.0% of questions. [The study](https://underneath.agency/research/ai-assistants-brand-agreement-study)
- ### Rewording the question can change the recommendation.

  Adding “on a tight budget” kept the same first brand only 15.3% of the time across three engines in our study, against 68.0% when the same question was simply asked again. [The study](https://underneath.agency/research/ai-prompt-phrasing-study)
- ### AI citations don’t simply mirror Google rankings.

  In our study of 4,051 AI Overview citations, only 28.7% of cited URLs ranked in the top 10 for the same search. [The study](https://underneath.agency/research/ai-overview-citations-study)
- ### AI answers can use sources outside the traditional top 10.

  17.9% of the AI Overviews we studied cited Reddit, and most of the cited threads were outside Google’s top 10 for the same search. [The study](https://underneath.agency/research/ai-reddit-citations-study)
- ### AI answers can contradict a business’s Google profile.

  Across 159 local businesses and four AI engines, 18.9% of answers contained a fact that differed from the business’s Google profile. [The study](https://underneath.agency/research/ai-business-facts-accuracy-study)

These are not universal laws. They are measurements from defined studies, with published methods and limitations. **That distinction matters.**

[Read the research](https://underneath.agency/research)

How we measure

## We measure the work in four *layers*.

**Visibility → Evidence → Behavior → Business.** An AI mention is not automatically a business result, so we don’t collapse everything into one “AI visibility score.”

- ### Visibility

  **What is being said?** Mentions, recommendations, descriptions, citations, competitive presence and persistence across important questions.
- ### Evidence

  **What supports it?** Claims, sources, entities, corroboration, conflicts, freshness and information coverage.
- ### Behavior

  **What did people do?** AI referrals, engagement, inquiries, applications, bookings and conversions.
- ### Business

  **Did it matter?** Pipeline, customers, revenue, acquisition cost or the business metric that actually matters to you.

A visibility increase is not automatically a revenue increase. A citation is not automatically a lead. A lead is not automatically a customer. **We measure each step on its own terms.**

[See how we measure](https://underneath.agency/results)

Results

## When the work produces results, we show the *evidence*.

**No invented scores. No convenient stories.** When client results are available and approved, we show the baseline, intervention, measurement and outcome, not just a screenshot of an AI answer. Until a result is substantiated and approved, we label it accordingly.

Because a good case study should answer four questions:

- What was the problem?
- What changed?
- What happened afterward?
- What can we actually attribute to the work?

Our standard is simple: **a result is only proof if you can check it.**

[How we report results](https://underneath.agency/results)

What we do

## Generative engine *optimization*.

GEO is the work of making a business easier for AI-powered search and answer systems to understand, retrieve, describe and reference. It builds on SEO, but asks a broader question:

> Given everything published about this business, on its own site and elsewhere, what will an answer system understand, retrieve and say?

Depending on the diagnosis, our work can include

- AI visibility and answer analysis
- Buyer-question and query research
- Claim and evidence analysis
- Content and information architecture
- Technical GEO
- Entity and information architecture
- Source and evidence strategy
- Monitoring and experimentation
- Team enablement

We don’t prescribe the work before we understand the problem.

[See what’s included](https://underneath.agency/services/generative-engine-optimization#deliverables)

Who we help

## Built for businesses where being chosen *matters*.

- ### Home and local services

  Businesses where the recommendation can become a call, booking or visit.
- ### Franchises and multi-location brands

  Businesses that need consistent information across markets and locations.
- ### B2B software and technology

  Businesses competing for consideration, comparison and shortlist inclusion.
- ### Financial services and insurance

  Businesses where accuracy, trust and compliance matter.
- ### Healthcare and dental

  Businesses where reputation and local recommendation matter.
- ### Legal and professional services

  Businesses where expertise and trust shape the decision.
- ### Retail and ecommerce

  Businesses where product information, comparison and recommendation matter.
- ### Hospitality and travel

  Businesses where local discovery and recommendation can directly influence demand.

Engagement model

## One team. *One* accountable lead.

Every engagement has one senior lead responsible for the strategy, priorities, decisions and overall program. Specialists join where the work requires them.

You know who is responsible. You know what changed. You know what we learned. And you know what happens next.

Every program runs on the weekly 30, a 30-minute review in the same slot every week, with a decision-maker from your side in the room. If a decision-maker cannot join the weekly review, we may not be the right fit.

[How we work](https://underneath.agency/approach)

Why Underneath

## Because the visible answer is not the whole *problem*.

A customer sees a recommendation. We want to understand what produced it.

A system cites a source. We want to understand why that source was there.

A competitor appears in the shortlist. We want to understand what information supports their position.

And when your business is missing, we want to know whether the problem is recognition, retrieval, evidence, representation, or something else entirely.

**The answer is only the surface.** That’s why we’re called Underneath.

[About Underneath](https://underneath.agency/about)

Free strategy call

## The market is already asking.

The question is whether your business is part of the answer. Tell us what you’re trying to change. We’ll take a first look at how your business appears across search and AI answers, what information sits underneath that visibility, and where we would start.

[Get a first look at my business](https://underneath.agency/contact)[Read the research](https://underneath.agency/research)

---

This is the Markdown twin of https://underneath.agency. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How We Work: The Underneath Method and First 90 Days | Underneath"
description: "How Underneath works: diagnose, architect, deploy and compound; what we look beneath the answer, the first 90 days, the weekly review and reporting."
canonical: "https://underneath.agency/approach"
published: 2026-09-26
updated: 2026-09-27
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Approach · The Underneath Method

# Understand first. Then say it *everywhere*.

Publishing more pages does not make a business more visible in AI answers.

When an AI system recommends a company, describes a product or cites a source, there is an information system underneath that answer: the entities involved, the claims made about them, the sources and evidence behind those claims, and what the system can retrieve. We work backwards from the answer to understand that system, then decide what needs to change.

**Diagnose → Architect → Deploy → Compound.** That is the Underneath Method.

1. 01 · Weeks 1 to 3

   ### Diagnose

   A baseline of how search and AI answers describe you, and why.
2. 02 · Weeks 2 to 4

   ### Architect

   A prioritized plan, with owners and measures.
3. 03 · Month 2 onward

   ### Deploy

   The changes ship, on your site and beyond it.
4. 04 · Ongoing

   ### Compound

   Measure, learn and set the next cycle.

Diagnose · Weeks 1 to 3

## Find out what the market, and the machines, actually *understand* about you.

We start with your customers, not with a list of pages to write.

We examine how your business appears across search and AI surfaces, the questions buyers ask, the competitors they compare you with, and the information available about your company across the web.

We look beyond mentions and rankings. We investigate the **claims, sources, entities and relationships underneath the answer.**

### We examine

AI visibility
:   Where you appear, where competitors appear, how descriptions change across questions, and which answers and citations matter.

Search visibility
:   How your existing search presence supports or contradicts the information AI systems encounter.

Entity understanding
:   How your company, products, services, people and locations are represented and connected.

Source ecosystem
:   Which first-party and third-party sources contribute to the story around your business.

Customer language
:   What customers say in reviews, sales conversations, questions and comparisons, and how that differs from the language your business uses.

Competitive information
:   What competitors are consistently associated with, supported by, cited by and understood for.

Technical foundation
:   Whether the information you want discovered can actually be found, interpreted and measured.

The result is a baseline and a set of priorities, not a generic SEO audit or a list of content ideas: a clear view of **what is understood, what is missing, what conflicts and what matters next.**

Architect · Weeks 2 to 4

## Decide what needs to become *true*.

Diagnosis tells us where the gaps are. Architecture determines what we do about them.

We turn the findings into a prioritized plan across the website, content, technical foundation, entities, sources and measurement. Every recommendation comes with the reason for it.

### We identify

- What should be clearer
- What needs stronger evidence
- What information is missing
- Which relationships need to be established
- Which sources deserve attention
- Which pages matter most
- Which technical issues are blocking discovery
- Which opportunities are worth pursuing now
- What should wait

The output is a practical roadmap with goals, priorities, owners and measures.

**No activity for activity’s sake.**

Deploy · Month 2 onward

## Change the information system around the *business*.

The work follows the diagnosis, so it looks different for every company.

### We may work across

Technical
:   The systems that allow important information to be discovered, understood and measured.

Content
:   Pages and information designed around the questions, decisions and claims that matter.

Entities
:   The way your company, products, services, people and locations are represented and connected.

Sources
:   The wider information environment that supports the understanding of your business.

Evidence
:   The material that makes important claims clearer, more credible and more referenceable.

Measurement
:   The signals that tell us whether the changes are actually affecting visibility and business performance.

We don’t prescribe the work before we understand the problem.

Compound · Ongoing

## Every cycle makes the next one *smarter*.

AI answers shift, search results move, competitors publish and your own business keeps changing. A one-time optimization project can’t keep up with that.

So we measure what changed, investigate why, test our assumptions and use what we learn to decide what comes next:

**Measure → Learn → Change → Test → Compound.**

Each cycle starts from better evidence than the last. Over time, what builds up is not just more mentions but a stronger, more accurate information footprint around the business.

The first 90 days

## From signature to shipped work, without a long *silence*.

The clock starts when the program does. You see the audit in the first week, a plan within the month, and the first real work live in the second.

1. Day 0

   ### Agreement and access

   Goals, markets, competitors, measurement and access are agreed.
2. Days 1 to 7

   ### Audit delivered

   A baseline across search, AI visibility, the website, sources, entities, customer language and measurement.
3. By day 30

   ### Plan agreed

   Priorities, pages, sources, technical work, owners and measures.
4. Days 31 to 60

   ### First work live

   The first changes ship, with the baseline kept so they can be measured.
5. Day 90

   ### First full review

   What changed, what we learned and what the next quarter should focus on.

At no point does the work vanish into a strategy deck.

Beneath the answer

## The answer is only the *surface*.

An AI answer can look simple. The information behind it isn’t. We investigate these layers together.

- ### The question

  What are buyers actually asking?
- ### The entity

  Who or what is being discussed?
- ### The claim

  What is being said about it?
- ### The source

  Where does that information come from?
- ### The evidence

  What supports the claim?
- ### The relationship

  What other entities, products, people, categories or concepts are connected to it?
- ### The answer

  How does all of that become the description, comparison, recommendation or citation a customer sees?

In practice, we take the answers to your most important customer questions and work backwards: which companies appear and how they are described, which sources are referenced, what information shows up consistently, what is missing, where competitors have stronger evidence and where your own information disagrees with itself.

**The answer tells us where to look.** The deeper analysis tells us what to change.

Engagement model

## One team. *One* accountable lead.

The team is shaped around the work: one plan, one accountable lead, and specialists brought in where the work needs them. Some specialist work is done by vetted partners, under the same lead and the same standards.

- ### Strategy lead

  Owns priorities, decisions and the overall program.
- ### GEO specialist

  Owns AI visibility research, analysis and ongoing experimentation.
- ### Content specialist

  Turns the opportunities the diagnosis finds into clear, useful pages that other sources can reference.
- ### Technical GEO specialist

  Handles the technical and information architecture that makes that information discoverable.

Nothing in a black box

## The weekly 30.

Every program has a 30-minute review in the same slot every week, with a decision-maker from your side in the room. Four things, always in this order.

1. 01

   ### What shipped

   The work completed, with links.
2. 02

   ### What changed

   The evidence across AI visibility, search and business metrics.
3. 03

   ### What we learned

   What the data and research are telling us.
4. 04

   ### What happens next

   The next priorities and decisions required.

A written note follows within one business day. You should never have to wait for a monthly report to find out what your agency has been doing. If a decision-maker cannot join the weekly review, we may not be the right fit.

Reporting

## A rhythm you can plan *around*.

Weekly
:   The weekly 30 and its written note.

Monthly
:   Performance against agreed goals, including visibility, citations, traffic, leads and other relevant business measures.

Quarterly
:   A deeper strategy review and a revised plan based on what the previous quarter taught us.

Always on
:   Shared visibility into the metrics and technical signals that matter.

How we think

## Seven questions, asked in *order*.

- ### Understanding

  Why do your best customers choose you?
- ### Recognition

  Do search engines, AI systems and the sources they read recognize who you are?
- ### Retrieval

  Can the right information be found when it matters?
- ### Evidence

  Are the important claims supported?
- ### Competition

  Where do competitors have a stronger information footprint?
- ### Measurement

  What actually changed?
- ### Learning

  What do we know now that we didn’t know before?

The questions are simple. Answering them well is the work.

Free strategy call

## Nothing is optimized in *isolation*.

Your website is one source. Your reviews are another. Your customers, publishers, directories, communities, search results and other sources all contribute to the information environment around your business.

Our job is to understand how those pieces fit together, then strengthen the ones that are letting you down. Bring the problem to a free 30-minute call and we’ll tell you which pieces we would look at first.

[Talk through my problem](https://underneath.agency/contact)[See what’s included](https://underneath.agency/services/generative-engine-optimization)

---

This is the Markdown twin of https://underneath.agency/approach. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Results: How We Measure GEO and AI Visibility | Underneath"
description: "How Underneath measures AI visibility in layers: what AI answers say, the evidence underneath, what people do and what the business gains. No invented scores."
canonical: "https://underneath.agency/results"
published: 2026-09-26
updated: 2026-09-27
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
How we measure

# A result is only proof *if you can check it.*

AI visibility is easy to make sound impressive: a percentage, a ranking, a screenshot of an answer. None of those means much without a baseline, a defined question, a consistent method and a connection to the business.

We measure the work in layers: **Visibility → Evidence → Behavior → Business.** That keeps three things apart: what changed in an AI answer, what changed in the information underneath it, and what changed for the company.

Each report shows what moved, what did not, what we believe caused it and what we would investigate next.

Visibility

## What does the market *see*?

We build a representative set of questions around the decisions your buyers make, not just searches for your brand name.

- Category questions
- Comparison questions
- Problem questions
- Product questions
- Location questions
- Alternative questions
- High-intent questions

We track how your business and competitors appear across those questions and how the answers change over time.

### We look at

Presence
:   Whether the business appears at all.

Position
:   Where and how prominently it appears when the answer involves multiple companies.

Description
:   What the system actually says about the business.

Recommendation
:   Whether the business is presented as an option when the question calls for one.

Citation
:   Which sources are being used to support the answer.

Persistence
:   Whether the pattern survives changes in wording and related questions.

**A single answer is an observation. A repeated pattern is a signal.**

Evidence

## What is underneath the *answer*?

Visibility tells us what happened. Evidence helps us understand why it may be happening.

For important answers, we trace the information underneath them. We look at the claims being made, the sources associated with those claims, the entities involved and the consistency of that information across the wider web.

### We investigate

Claims
:   What is being said about the company, product or service?

Sources
:   Where does the information come from?

Corroboration
:   Is important information supported beyond the company’s own website?

Freshness
:   Is the information current?

Competition
:   Where do competitors have a stronger or more complete information footprint?

Without this layer, a visibility change is hard to act on: you can see that something moved, but not what to do about it. Accuracy, consistency and coverage are measured in their own right, below.

Information health

## Accuracy is a measurement, *too*.

A business can be highly visible and still be badly represented.

AI answers can contain

- Outdated products
- Old pricing
- Incorrect locations
- Incomplete descriptions
- Outdated company information
- Incorrect relationships
- Unsupported claims
- Conflicting information

So we measure whether the facts that matter most are right.

### We track

Accuracy
:   Are important facts represented correctly?

Conflicts
:   Where do authoritative sources disagree?

Outdated claims
:   Which old information continues to surface?

Coverage
:   Which important facts are missing?

Consistency
:   Does the business describe itself consistently across important sources?

Appearing more often is only half the job. The other half is being described correctly when you do.

Behavior

## Did visibility change what people actually *did*?

An AI mention is not a business outcome. Wherever the data allows, we connect AI visibility to what people actually did.

### We measure

AI referrals
:   Sessions arriving from AI systems and AI-driven search experiences where attribution is available.

Engagement
:   What those visitors do after arriving.

Leads
:   Calls, forms, bookings, applications and other agreed conversion events.

Qualified demand
:   Which AI-referred inquiries become meaningful opportunities.

Revenue
:   Where the available attribution supports connecting AI-driven demand to revenue.

Assisted influence
:   Where AI or search visibility contributes to a journey without being the final recorded source.

We distinguish what we can directly attribute from what we can only reasonably observe.

Business

## The board doesn’t need an AI *score*.

It needs to know whether the work is paying off. Every engagement starts by agreeing what the business is trying to improve.

Depending on the company, that might be

- Qualified leads
- Pipeline
- Revenue
- Applications
- Bookings
- Cost per acquisition
- New customers
- Geographic demand
- Product adoption

AI visibility is then treated as one layer of the measurement system, not the final objective.

The question we answer is not whether a number on a dashboard went up. It is:

> Did the work produce a meaningful change in the outcomes we agreed to improve?

Measurement without false precision

## We don’t pretend the system is more observable than it *is*.

Models are updated, answers shift from one day to the next, sources come and go, and attribution is never complete. So we are careful about what the data can actually prove.

### We distinguish between

Observed
:   Something directly measured.

Correlated
:   Two things changed together, but causation is not established.

Attributed
:   The available tracking supports a connection between the change and the business outcome.

Hypothesized
:   A plausible explanation that needs further testing.

A report should tell you what we know and, just as plainly, what we don’t know yet.

Experiments

## Measure the intervention, not just the *outcome*.

When we make a meaningful change, we want to know why we made it and what we expect to learn. A typical cycle looks like this.

- ### Observation

  Something in the information environment stands out.
- ### Hypothesis

  We propose a reason it may be happening.
- ### Intervention

  We change the relevant part of the system.
- ### Retest

  We return to the relevant questions and measurements.
- ### Interpretation

  We determine what changed and how confident we should be in the explanation.
- ### Next decision

  We either build on the finding, revise the hypothesis or move somewhere else.

This prevents the program from becoming a collection of disconnected deliverables.

Our measurement model

## Four layers. One *picture*.

- ### Visibility

  **What is being said?** Mentions, recommendations, descriptions, citations and competitive presence.
- ### Evidence

  **What supports it?** Claims, sources, entities, corroboration, conflicts and information coverage.
- ### Behavior

  **What did people do?** AI referrals, engagement, inquiries, applications, bookings and conversions.
- ### Business

  **Did it matter?** Pipeline, customers, revenue, acquisition cost and the business metric that matters to you.

The layers should connect, but they should not be confused.

A visibility increase is not automatically a revenue increase. A citation is not automatically a lead. A lead is not automatically a customer.

We measure each step on its own terms.

Illustrative scenarios

## Different businesses require different *evidence*.

We don’t use the same dashboard for every client. What we measure follows the problem.

Scenario 01 · B2B software

### Assistants recommend three competitors when buyers ask for the best tools in the category.

Illustrative scenario, not a client result

Trace which comparison pages, reviews and analyst summaries the answers lean on, then earn a credible place in those sources instead of publishing more blog posts.

Share of AI answers mentioning the brand
:   Mentions

Citations of owned pages
:   Citations

Pipeline from AI referrals
:   Pipeline

- Generative engine optimization (GEO)
- Enterprise GEO
- What we would measure

Scenario 02 · Financial services

### Product pages sit below aggregators, and every new page needs compliance review.

Illustrative scenario, not a client result

Build an answer-first content model with a review workflow compliance can live with, and roll structured data out at template level.

Answers extracted by search engines
:   Answers

Product applications from AI referrals
:   Applications

Time from draft to approved page
:   Speed

- Enterprise GEO
- What we would measure

Scenario 03 · Enterprise software after a rebrand

### AI answers still describe discontinued products and old pricing, citing years-old community threads.

Illustrative scenario, not a client result

Correct the record at the source: clear, current brand facts on the site, and a disclosed expert presence in the communities models read.

Factual accuracy of AI answers
:   Accuracy

Outdated claims in answers
:   Claims

Community sentiment
:   Sentiment

- Community citations
- Generative engine optimization (GEO)
- What we would measure

Scenario 04 · Single-location service business

### Calls are flat, and AI assistants name three competitors when someone asks for a plumber nearby.

Illustrative scenario, not a client result

Fix call and form tracking first, so every lead has a source. Then correct the Google Business Profile and location-page facts, build them around the reasons customers give in reviews, and work toward being the business local AI answers name for the services that turn into booked jobs.

Tracked calls and form fills
:   Calls

Cost per booked job
:   Cost

Local AI answers naming the business for core services
:   AI answers

- Local GEO
- Google Business Profile
- What we would measure

What goes in front of your board

## A report should make decisions *easier*.

The cadence depends on the engagement. The questions don’t.

Weekly

- What shipped?
- What changed?
- What did we learn?
- What happens next?

Monthly

- How are the agreed metrics moving?
- Which experiments are producing useful evidence?
- What should change in the work?

Quarterly

- What changed in the information environment?
- What changed in visibility?
- What changed in customer behavior?
- What changed in the business?
- What did we learn that changes the next quarter?

The objective is not a prettier dashboard. It is a clearer decision.

What we won’t do

## We don’t publish numbers we can’t *defend*.

- No invented “AI visibility score.”
- No impressive percentage without a baseline.
- No case-study language around illustrative scenarios.
- No claim that an AI mention caused revenue simply because the two appeared in the same quarter.
- No screenshot presented as proof of a durable result.

And no client result is published until the client has approved it. Until then, every scenario on this site is labeled illustrative.

The point of measurement

## Make the invisible *discussable*.

AI visibility feels hard to measure because every answer is generated on the spot and the systems behind it keep changing. That doesn’t make measurement impossible. It makes discipline essential.

1. Define the question.
2. Establish the baseline.
3. Observe the answer.
4. Trace the evidence.
5. Measure the behavior.
6. Connect it to the business.
7. State what is known.
8. State what isn’t.

Then decide what to do next.

Free strategy call

## Your baseline is the first result worth knowing.

On a free 30-minute call we take a first look at how you show up in search and AI answers and what your site gives them to cite. Then we tell you plainly whether a full audit is worth it.

[See where you stand](https://underneath.agency/contact)[How we work](https://underneath.agency/approach)

---

This is the Markdown twin of https://underneath.agency/results. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Research: Original AI Search Studies and Data | Underneath"
description: "Original studies of how AI assistants and AI search choose sources and brands, with the data, method and limitations published for every number."
canonical: "https://underneath.agency/research"
published: 2026-09-26
updated: 2026-09-28
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Research

# Original data on how AI *finds and cites.*

Twenty-three studies of AI crawlers, Google AI Overviews and AI Mode, the brands ChatGPT, Gemini, Perplexity and Claude recommend, and what they say about businesses. Every study publishes its method, its limitations and the per-item data behind each number, free to reuse with attribution.

Featured study

## [ChatGPT local recommendations vs Google Maps](https://underneath.agency/research/chatgpt-local-recommendations-study)

We compared 1,405 local business listings from ChatGPT in the United States, United Kingdom, Canada and Australia with Google Maps. For the same business at the same location, 99.7% of names, 98.1% of website links and 99.6% of phone numbers were identical, including tracking tags that exist only on Google Business Profiles.

Research · Local AI search

Studies

## Twenty-three studies, *all the data published.*

- Research · AI crawlers

  ### [Which AI crawlers do top websites block?](https://underneath.agency/research/ai-crawler-blocking-study)

  The robots.txt of the top 10,000 sites: 15.2% block GPTBot, and half of those still allow ChatGPT’s search crawler.

  Underneath Research
- Research · Local AI search

  ### [Which Google Maps businesses does ChatGPT recommend?](https://underneath.agency/research/chatgpt-local-picks-google-profile-study)

  898 ChatGPT answers over two days: more reviews than the local median raised a Maps business’s chance of being listed by about 20 points.

  Underneath Research
- Research · AI crawlers

  ### [How many websites have an llms.txt file?](https://underneath.agency/research/llms-txt-adoption-study)

  11.5% of live top sites serve a valid llms.txt; more return an HTML page at the address instead.

  Underneath Research
- Research · AI crawlers

  ### [Do websites serve Markdown to AI agents?](https://underneath.agency/research/agent-readable-web-study)

  3.2% of top sites answer an agent’s request for Markdown, with pages a median 96.1% smaller than the HTML.

  Underneath Research
- Research · AI search

  ### [When does Google show an AI Overview?](https://underneath.agency/research/ai-overviews-frequency-study)

  1,248 US searches: local maps, wording and query mix explain more than industry; most cited sources do not rank on page one.

  Underneath Research
- Research · AI search

  ### [Do AI Overviews cite the pages that rank?](https://underneath.agency/research/ai-overview-citations-study)

  4,051 citations: only 28.7% rank in the top 10 for the same search, and YouTube appears in 64.2% of AI Overviews.

  Underneath Research
- Research · AI search

  ### [AI Mode vs AI Overviews: how different are the sources?](https://underneath.agency/research/ai-mode-vs-ai-overviews-study)

  For the same search, Google’s two AI answers share a mean 13.9% of cited URLs.

  Underneath Research
- Research · AI search

  ### [What pages cited by AI Overviews have in common](https://underneath.agency/research/ai-overview-cited-pages-study)

  Among top-10 pages on the same search, ranking position matters far more than schema or other markup.

  Underneath Research
- Research · AI search

  ### [How many “best of” lists cited by AI rank their own brand first?](https://underneath.agency/research/self-promoting-best-lists-study)

  269 numbered lists cited by six AI surfaces: 24.2% were published by a company that ranked itself first.

  Underneath Research
- Research · AI search

  ### [The YouTube videos Google’s AI cites are small](https://underneath.agency/research/ai-overview-youtube-videos-study)

  469 videos cited in AI Overviews and AI Mode: a median of 4,283 views, and 60.9% under 10,000.

  Underneath Research
- Research · AI search

  ### [When does a Reddit thread become evidence in Google’s AI?](https://underneath.agency/research/ai-reddit-citations-study)

  17.9% of AI Overviews cite Reddit. Cited threads have more discussion than the ones Google shows and the AI skips, and a fifth of the sentences citing them are not supported.

  Underneath Research
- Research · AI search

  ### [How fresh are the pages AI engines cite?](https://underneath.agency/research/ai-source-freshness-study)

  Against Google’s top 10 for the same questions, AI assistants cited pages first published about half as long ago.

  Underneath Research
- Research · AI assistants

  ### [Do ChatGPT, Gemini, Perplexity and Claude agree on brands?](https://underneath.agency/research/ai-assistants-brand-agreement-study)

  80 buyer questions: all four assistants recommended the same first pick for 10.0% of questions.

  Underneath Research
- Research · AI assistants

  ### [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)

  About a quarter of the brands ChatGPT named appeared in all five answers to the same question, and one answer showed about half of them.

  Underneath Research
- Research · AI assistants

  ### [Same question, four countries: do AI recommendations change?](https://underneath.agency/research/ai-recommendations-by-country-study)

  Asked five times in each of four countries, ChatGPT’s brands overlapped 0.594 within a country and 0.429 across countries.

  Underneath Research
- Research · AI assistants

  ### [Does rewording a question change AI brand recommendations?](https://underneath.agency/research/ai-prompt-phrasing-study)

  Adding “on a tight budget” kept the same first brand 15.3% of the time across three engines, against 68.0% for the same question asked again.

  Underneath Research
- Research · AI assistants

  ### [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)

  ChatGPT ran 3.7 searches per answer, and added a year to them in 52.5% of answers.

  Underneath Research
- Research · AI assistants

  ### [Do ChatGPT, Gemini, Perplexity and Claude cite pages that rank?](https://underneath.agency/research/ai-citations-google-rankings-study)

  Only 8.3% of ChatGPT’s citations ranked in Google’s top 10 for the same question.

  Underneath Research
- Research · AI assistants

  ### [How faithfully do AI assistants quote software prices?](https://underneath.agency/research/ai-pricing-accuracy-study)

  61.9% of 840 quoted plan prices were fully faithful to the pricing page; most differing prices still appear on the vendor’s own site.

  Underneath Research
- Research · AI assistants

  ### [“Is this brand legit?” How AI assistants build a reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study)

  79 brands, four engines: every complete answer called the brand legitimate, and 99.7% then raised a problem.

  Underneath Research
- Research · Local AI search

  ### [Do AI answers match a business’s Google profile?](https://underneath.agency/research/ai-business-facts-accuracy-study)

  159 local businesses, four engines: 18.9% of answers stated a fact that differed from the business’s Google profile, often one the business publishes elsewhere.

  Underneath Research
- Research · AI assistants

  ### [Do Wikipedia and schema make AI assistants recommend a brand?](https://underneath.agency/research/brand-entity-ai-recommendations-study)

  Brands on Wikipedia are named by more assistants, but within the same question most of the gap goes once brand prominence is taken into account.

  Underneath Research
- Research · Local AI search

  ### [ChatGPT local recommendations vs Google Maps](https://underneath.agency/research/chatgpt-local-recommendations-study)

  1,405 listings in four countries matched against Google Maps, field by field.

  Underneath Research

How we research

## Every number *traces to data.*

Each study states its question, sample, method and limitations, and publishes the per-item data behind every figure. Every statistic is computed directly from that dataset, so anyone who downloads it can reproduce the number, and each figure in the text is checked against the data before publication. Comparisons with other studies quote the original publisher. Corrections are logged in each study’s changelog.

Free strategy call

## Want these numbers for *your own category?*

We run the same measurements for a business’s own buyer questions and competitors. On a free 30-minute call we’ll take a first look and send you a short written read afterward.

[Get these numbers for my business](https://underneath.agency/contact)[How we work](https://underneath.agency/approach)

---

This is the Markdown twin of https://underneath.agency/research. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Resources: GEO Guides and the AI Search Glossary | Underneath"
description: "Guides on generative engine optimization agencies: what they do, how to choose one, whether one is worth it, and GEO vs SEO, plus the AI search glossary."
canonical: "https://underneath.agency/resources"
published: 2026-09-26
updated: 2026-09-27
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Resources

# Field notes on how AI assistants *find and choose.*

Guides to generative engine optimization, written from our own measurements of what the AI engines cite. Our figures come from repeated runs, and you can ask us for the data behind any of them.

Featured guide

## [What does a GEO agency do?](https://underneath.agency/resources/what-does-a-geo-agency-do)

A GEO agency gets a brand cited, mentioned and recommended in AI answers. The nine things it does, what it costs in 2026, and how the AI engines themselves describe the job, measured across 22 runs.

Guide · AI search

Key takeaways

1. A GEO agency works on the pages, the structured data, the brand’s entity and the third-party sources the engines cite.
2. Published 2026 price guides put agency retainers between about $1,500 and $50,000 a month; the cheapest useful engagement is a repeated audit.
3. On Google’s AI surfaces, a top-three ranking roughly doubles a page’s chance of a citation compared with positions 7 to 10; on Perplexity and ChatGPT, ranking counts for much less.

Guides

## Four guides to *hiring a GEO agency*

- Guide · AI search

  ### [How to choose a GEO agency](https://underneath.agency/resources/how-to-choose-a-geo-agency)

  Eight criteria, a scorecard, the red flags, and the questions to ask before you sign, benchmarked against how the AI engines answer the same question.

  Underneath team
- Guide · AI search

  ### [Is a GEO agency worth it?](https://underneath.agency/resources/is-a-geo-agency-worth-it)

  When hiring one pays back, when it does not, what it costs, and the three numbers that decide it.

  Underneath team
- Guide · AI search

  ### [GEO agency vs SEO agency](https://underneath.agency/resources/geo-agency-vs-seo-agency)

  What each does, where they overlap, where measurement says they differ, and whether you need one, the other or both.

  Underneath team
- Guide · AI search

  ### [What does a GEO agency do?](https://underneath.agency/resources/what-does-a-geo-agency-do)

  The nine things a generative engine optimization agency does, what it costs, and how the engines describe the job.

  Underneath team

Reference

## The AI search *glossary.*

26 terms, from canonical URLs to query fan-out, defined in a sentence or two each.

[Open the glossary](https://underneath.agency/resources/glossary)

- AEO (Answer Engine Optimization)
- AI Mode
- AI Overviews
- AI share of voice
- AI visibility
- Answer accuracy
- Buyer prompt set
- Canonical URL
- Citation rate
- Crawl budget
- Entity
- Featured snippet
- GEO (Generative Engine Optimization)
- Grounding
- Hallucination
- hreflang
- Knowledge graph
- llms.txt
- LLMO (Large Language Model Optimization)
- Local citation
- People Also Ask
- Query fan-out
- RAG (Retrieval-Augmented Generation)
- Schema markup
- SEO (Search Engine Optimization)
- Zero-click search

Free strategy call

## Some questions are easier to answer about your own business.

Bring the one that matters most. On a free 30-minute call we’ll take a first look at it and send you a short written read afterward.

[Ask about my business](https://underneath.agency/contact)[How we work](https://underneath.agency/approach)

---

This is the Markdown twin of https://underneath.agency/resources. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "AI Search Glossary: SEO, AEO, GEO & LLMO | Underneath"
description: "Plain-language definitions of search and AI search terms: SEO, AEO, GEO, LLMO, AI Overviews, AI share of voice, citations, grounding, RAG and llms.txt."
canonical: "https://underneath.agency/resources/glossary"
published: 2026-09-26
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Resources · Glossary

# The AI search glossary, *defined plainly.*

Some of these terms are established. Some are new and still argued over. Each is defined the way we use it.

- [A](#letter-A)
- [B](#letter-B)
- [C](#letter-C)
- [E](#letter-E)
- [F](#letter-F)
- [G](#letter-G)
- [H](#letter-H)
- [K](#letter-K)
- [L](#letter-L)
- [P](#letter-P)
- [Q](#letter-Q)
- [R](#letter-R)
- [S](#letter-S)
- [Z](#letter-Z)

AEO (Answer Engine Optimization)
:   An industry label for structuring content so search-results features extract and credit your answer: AI Overviews and AI Mode, featured snippets, People Also Ask and voice assistants. At Underneath this work is part of every GEO program.

AI Mode
:   Google’s conversational AI search mode: it answers a query with a generated response and cited links, and takes follow-up questions. It often cites different sources from AI Overviews for the same search.

AI Overviews
:   Google’s AI-generated summaries shown at the top of some search results, typically citing several source pages.

AI share of voice
:   Your brand’s mentions as a share of all tracked brand mentions across the buyer prompt set.

AI visibility
:   How often, and how accurately, a business appears in AI-generated answers: Google AI Overviews and AI Mode, ChatGPT, Gemini, Perplexity and Copilot. Measured with a fixed buyer prompt set run on a schedule.

Answer accuracy
:   The share of statements an AI system makes about a brand that are factually correct and current.

Buyer prompt set
:   The fixed list of real buyer questions we run across the tracked engines on a stated schedule to measure visibility in AI answers. Also called a prompt universe. Asking to see one is a useful test when [choosing a GEO agency](https://underneath.agency/resources/how-to-choose-a-geo-agency).

Canonical URL
:   The preferred version of a page declared to search engines when similar or duplicate URLs exist.

Citation rate
:   How often AI-generated answers cite a brand’s owned pages as sources.

Crawl budget
:   The number of URLs a search engine will crawl on a site within a given period. It mostly matters for very large sites.

Entity
:   A distinct, identifiable thing (an organization, product, person or concept) that search engines and models can recognize and connect.

Featured snippet
:   A highlighted excerpt shown above traditional results that answers a query directly.

GEO (Generative Engine Optimization)
:   Structuring content and managing a brand’s presence on the web so that AI answer engines (ChatGPT, Google AI Overviews and AI Mode, Gemini, Perplexity, Copilot, Claude) retrieve, cite and recommend it, and describe it correctly. Works on the brand’s own pages and on the third-party sources the engines draw on. Our guide to [what a GEO agency does](https://underneath.agency/resources/what-does-a-geo-agency-do) lists that work in detail. See [generative engine optimization services](https://underneath.agency/services/generative-engine-optimization).

Grounding
:   When an AI system retrieves live information from the web or documents to base its answer on, rather than relying only on training data.

Hallucination
:   An AI-generated statement that is presented confidently but is false or unsupported.

hreflang
:   An attribute that tells search engines which language and regional version of a page to show to which users.

Knowledge graph
:   A structured network of entities and the relationships between them, which search engines use to connect facts about people, organizations, places and things.

llms.txt
:   A proposed Markdown file at a site’s root that summarizes key content and links for large language models. Underneath publishes one; our [adoption study](https://underneath.agency/research/llms-txt-adoption-study) found a valid file on 11.5% of top websites, and no evidence yet that AI systems rely on it.

LLMO (Large Language Model Optimization)
:   An industry label for making what AI models say about a business correct and consistent. At Underneath this work is covered by generative engine optimization (GEO).

Local citation
:   A listing of a business’s name, address and phone number on a directory or map service. Distinct from an AI citation, which is a source link in an AI-generated answer.

People Also Ask
:   An expandable set of related questions and answers shown within Google search results.

Query fan-out
:   How an AI search system splits one question into several related sub-searches, runs them, and combines the results into one answer. Pages can be cited for a sub-search they rank for even when they do not rank for the original question.

RAG (Retrieval-Augmented Generation)
:   A technique in which a model retrieves relevant documents before generating an answer, often with citations.

Schema markup
:   Structured data using the Schema.org vocabulary that describes page content in a machine-readable way.

SEO (Search Engine Optimization)
:   Improving how a business ranks and converts in search results, including map results and AI Overviews. For how the two jobs differ, see [how GEO and SEO agencies compare](https://underneath.agency/resources/geo-agency-vs-seo-agency).

Zero-click search
:   A search where the user’s need is met on the results page or in an AI answer without clicking through to a website.

Free strategy call

## The vocabulary is the easy part. Next, find out what AI assistants say about you.

On a free 30-minute call we’ll take a first look at how AI assistants and search engines describe your business today.

[Check what AI says about you](https://underneath.agency/contact)[How we work](https://underneath.agency/approach)

---

This is the Markdown twin of https://underneath.agency/resources/glossary. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "About Underneath: Who We Are and What We Believe | Underneath"
description: "Underneath is a GEO agency founded in 2026. Why it exists, what we believe about AI answers, how we think, the standards we hold to, and who does the work."
canonical: "https://underneath.agency/about"
published: 2026-09-26
updated: 2026-09-27
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
About Underneath

# We started with a simple question: why do people *choose*?

The answer is changing. People still search. They still compare. They still visit websites. But increasingly, they ask.

They ask AI assistants to explain a category, compare options, recommend a provider, find a product, or help them decide what comes next. And the answer they receive can shape what they consider.

That answer is the visible part. **Underneath it is an information system.** Claims. Sources. Entities. Relationships. Pages. Reviews. References. Evidence.

Some of it is accurate. Some of it is incomplete. Some of it is outdated. And sometimes, the information simply does not connect in the way it needs to.

**Underneath exists to understand what is shaping the answer, and change what matters.**

The shift

## Why *Underneath*.

For a long time, the path to being found was relatively clear. You appeared in search results. People clicked. They visited your site. You made the case.

Now the journey can begin somewhere else. A customer can ask:

> Who should I consider? Which one is right for me? What are the differences? Who do people trust?

The system answers. That changes the problem.

It is no longer enough to ask whether a business can be found. We need to understand **how the business is understood**:

- What is being said?
- What supports it?
- Which sources are being used?
- What information is missing or conflicting?
- How does the system connect the business to the things it is being asked about?

That is the work underneath the answer. And that is where we work.

What we believe

## Five beliefs that decide how we *work*.

- ### The answer is only the surface.

  A visible answer tells you what happened. It does not necessarily tell you why. We work backwards from the answers that matter to understand the information underneath them.
- ### Being understood comes before being visible.

  Visibility is an outcome. Before we try to increase it, we want to understand how a business is represented: what is clear, what is supported, what is missing, and where the information does not hold together.
- ### Evidence matters more than volume.

  More pages do not automatically create more trust. More mentions do not automatically create more authority. We care about the claims that matter, the evidence behind them, and whether the information supporting a business is accurate and consistent.
- ### Measurement should make us more honest.

  Not every change can be attributed to one action. Not every correlation is proof. We distinguish what we observed from what we can attribute, and what we believe from what we have tested.

  The purpose of measurement is not to make a report look better. **It is to make the next decision better.**
- ### Durable work beats clever tricks.

  Platforms change. Models change. Interfaces change. We build around things that should remain valuable through those changes: clear information, strong evidence, sound technical foundations and a business that is represented accurately.

  No shortcuts disguised as strategy.

Why we exist

## Good work should be represented *accurately*.

Businesses spend years building products, expertise, reputations and customer relationships. We believe that work should be represented accurately when someone asks about it.

Purpose
:   To help businesses become easier to understand, discover and choose.

Mission
:   To understand the information underneath AI answers, improve what needs improving, and measure what changes.

Not to chase every new feature. Not to manufacture visibility. Not to promise that we can control an answer we do not control.

**To understand the system, change what matters, and learn from what happens next.**

How we think

## Start with the question. Then work *backwards*.

**Question → Answer → Evidence → Information system → Intervention → Measurement**

That is the foundation of our work. The deeper investigation can involve the claims being made, the sources supporting them, the entities being recognized, the relationships connecting them and the pages and information systems behind them.

But the principle stays simple: **understand first, change second, measure third.** And then do it again.

Who does the work

## One accountable *lead*.

Underneath is intentionally small. Every engagement has one senior lead responsible for the work from strategy through execution and measurement. Specialists join where their expertise is needed.

You know who is thinking about the problem. You know who is responsible for the work. And you know who to call when something changes.

We believe accountability works better when there are fewer layers between the client and the person doing the thinking.

Standards

## Six standards we *hold* to.

- ### Evidence over opinion

  Recommendations should have a reason behind them. We separate evidence, observation and hypothesis.
- ### Senior by default

  Experienced people do the work. We do not sell senior expertise and hand delivery to junior teams.
- ### Durable over clever

  We build for changing platforms, not temporary loopholes.
- ### Say it plainly

  If something is uncertain, we say so. If something is complicated, we explain it.
- ### Honest measurement

  We do not turn correlation into causation or activity into business impact.
- ### Own the outcome

  We care about what changes for the business, not simply how much work we can report.

Who we work with

## Businesses where trust, reputation and consideration *matter*.

- Home & local services
- Franchises & multi-location brands
- Healthcare & dental
- Legal & professional services
- B2B software & technology
- Financial services & insurance
- Retail & ecommerce
- Hospitality & travel

Different industries. The same underlying question:

> When a customer asks, what will they be told?

Where we work

## Founded in 2026. Deliberately *focused*.

Underneath was founded in 2026. We work remotely with businesses across the United States, United Kingdom, Australia and Canada, by email and video call. Headquarters: 2846 Simons Hollow Road, Bloomsburg, PA 17815, United States.

Our model is deliberately focused: a senior accountable lead, supported by specialist expertise when the work calls for it. Every program tracks the same engines: ChatGPT, Claude, Gemini, Google AI Overviews and AI Mode, Perplexity, Microsoft Copilot and Bing.

We do not want to become a large agency with more layers between the problem and the person solving it. We want to stay close to the work.

What we are building

## Understand the system. Change what *matters*.

The way people discover and evaluate businesses is changing. We believe the companies that understand that change early will have an advantage, not because they found a trick, but because they understood how they are represented when someone asks.

That is what Underneath is here to do. **Understand the system. Change what matters. Measure what changed. Repeat.**

Free strategy call

## What does AI say about your *business*?

We think the more important question is: why? That is where we start.

[Start a conversation](https://underneath.agency/contact)[How we work](https://underneath.agency/approach)

---

This is the Markdown twin of https://underneath.agency/about. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Contact: Book a Free Strategy Call | Underneath"
description: "Tell us what you’re trying to change. Book a free 30-minute conversation and get a short written read on whether we’re the right fit."
canonical: "https://underneath.agency/contact"
published: 2026-09-26
updated: 2026-09-27
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Contact

# Tell us what you’re trying to *change.*

Tell us what you’re seeing, what isn’t working, or what you’re trying to understand.

We’ll use the first conversation to understand the problem before recommending the work.

How it works

### Start with the *problem*.

1. 01

   ### Tell us what you’re seeing

   Use the form to describe the problem, the business, and what you’re trying to change. We read every submission and reply by email within one business day, Monday to Friday, US business hours. We work with clients across the United States, United Kingdom, Australia and Canada.
2. 02

   ### A free 30-minute conversation

   We spend 30 minutes understanding the problem, what may be underneath it, and where we’d look first. **It’s a first look, not an audit.**
3. 03

   ### A short written read

   After the call, we’ll send a concise summary of what we heard, what we’d investigate first, and whether we think we’re the right fit.
4. 04

   ### Decide what comes next

   If there’s a fit, we’ll recommend a scope and price. If there isn’t, we’ll tell you. No pressure. No generic proposal sent before we’ve understood the problem.

Prefer email

[hello@underneath.agency](mailto:hello@underneath.agency)

The weekly 30

- Every program runs on the same 30-minute review call, at the same time each week, with a decision-maker from your side.
- What shipped. The work completed, with links.
- What changed. The evidence across AI visibility, search and business metrics.
- What we learned. What the data and research are telling us.
- What happens next. The next priorities and decisions required.
- A written note follows within one business day.
- If a decision-maker cannot join the weekly review, we may not be the right fit.

Office

- Headquarters: 2846 Simons Hollow Road, Bloomsburg, PA 17815, United States

The form

- Full name
- Work email
- Business name
- Your role
- Website
- Business size: 1–10 employees; 11–50 employees; 51–200 employees; 201–1,000 employees; 1,000+
- What prompted this conversation? Leads or sales from AI answers and search have slowed; Competitors appear in Google or AI answers, but we don’t; AI assistants describe us incorrectly or not at all; We’re opening locations or entering new markets; We’re planning a site migration or redesign; We want to understand our AI visibility; Something else
- Where do you want help? AI visibility and recommendations; Incorrect or missing information; Competitor visibility; Local or multi-location visibility; International visibility; Technical or website changes; Measuring AI-driven demand; Not sure yet
- Expected investment: $5k–$10k / month; $11k–$20k / month; $20k+ / month; Project-based; Not sure yet
- What problem are you trying to solve? The more context you give us, the more useful the first conversation can be.
- A decision-maker on our side can join a 30-minute review call every week.

We use these details to reply and prepare for the conversation. See our [privacy policy](https://underneath.agency/privacy).

Send message

Before the call

## You don’t need to prepare a *presentation*.

If you have them, bring

- The questions your customers ask most often
- Examples of AI answers you’re concerned about
- Competitors you see appearing ahead of you
- Recent changes to your website, business or market
- The business outcome you’re trying to improve

**We’ll take it from there.**

---

This is the Markdown twin of https://underneath.agency/contact. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Privacy Policy | Underneath"
description: "How Underneath collects, uses, shares and protects personal information from inquiries and visits to underneath.agency, and how to exercise your rights."
canonical: "https://underneath.agency/privacy"
published: 2026-09-26
updated: 2026-09-27
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Legal

# Privacy policy

Last updated: September 2026

Underneath Agency, trading as **Underneath** (“Underneath,” “we,” “us,” or “our”), operates this website and provides generative engine optimization services. Our headquarters are at 2846 Simons Hollow Road, Bloomsburg, PA 17815, United States.

We work with clients in the United States, United Kingdom, Australia and Canada.

If you have questions about this Privacy Policy or how we handle personal information, contact us at [hello@underneath.agency](mailto:hello@underneath.agency).

## Information we collect

We collect information that you choose to provide to us when you contact Underneath or submit an inquiry through our website. This may include:

- your name
- work email address
- business name
- role or job title
- business website
- business size
- what prompted your inquiry
- the type of help you are looking for
- your expected investment (budget range)
- whether a decision-maker can join a weekly review call
- information included in your message
- other information you voluntarily provide to us

We ask for this information so that we can understand your situation, respond to your inquiry and prepare for a useful conversation. We do not ask you to provide sensitive personal information through our contact form.

## Information collected when you visit the website

Our website does not use cookies for advertising or behavioral tracking and does not run advertising or analytics scripts.

When you visit the website, our hosting provider, **Vercel**, may process technical information associated with the request so that the website can be delivered, maintained and secured. This may include information such as:

- IP address
- browser or user-agent information
- requested page
- request time
- technical information associated with the connection

This information is part of normal web infrastructure and may be processed by our hosting provider as necessary to operate and secure the website.

## AI crawler visits

Because Underneath researches how AI systems access and use information on the web, we separately maintain a limited log of visits from known AI crawlers. The log records:

- the crawler name
- the page requested
- the time of the visit

The crawler log is not intended to identify individual people. We use this information to understand how AI crawlers access our public website and to support our research and technical work.

## How we use information

We use information we collect to:

- respond to inquiries
- communicate with prospective and existing clients
- prepare for calls and other conversations
- understand the business problem you have asked us to address
- manage our business relationships
- operate, maintain and secure our website
- understand visits from known AI crawlers
- meet applicable legal or administrative obligations where necessary

We do not sell personal information. We do not use information submitted through our contact form for advertising purposes.

## How we share information

We do not sell or rent personal information.

We may share information with service providers that help us operate our website, communicate with you or provide our business services. Those providers may process information on our behalf as necessary to provide their services.

We may also disclose information where necessary to comply with applicable law, legal process or a legitimate request from a governmental authority, or where necessary to protect our rights, security or property.

## How long we keep information

We retain inquiry information for as long as reasonably necessary to respond to your inquiry, maintain an ongoing business relationship where one exists, and maintain appropriate business records. The length of time information is retained depends on the nature of the information and our relationship with you.

When information is no longer needed for these purposes, we take reasonable steps to delete or otherwise dispose of it.

## Security

We take reasonable measures to protect personal information against unauthorized access, loss, misuse or disclosure. Access to information is limited according to the needs of the business and the sensitivity of the information.

No method of transmitting or storing information over the internet can be guaranteed to be completely secure, so we cannot guarantee absolute security.

## Your privacy rights

Depending on where you live and the laws that apply to you, you may have rights concerning your personal information. These may include the right to:

- request access to personal information we hold about you
- request correction of inaccurate information
- request deletion of personal information in certain circumstances
- object to or restrict certain processing
- withdraw consent where processing is based on consent
- request information about how your personal information is handled

To exercise a privacy right or ask a question about your information, contact [hello@underneath.agency](mailto:hello@underneath.agency). We may need to verify your identity before completing certain requests. Your applicable rights may depend on your location and the circumstances in which we process your information.

## International visitors

Underneath works with clients in multiple countries, including the United States, United Kingdom, Australia and Canada.

If you access our website or communicate with us from outside the United States, your information may be processed in the United States or in another country where Underneath or one of our service providers operates.

Privacy laws differ between countries. Where applicable, we take reasonable steps to handle personal information in accordance with the requirements that apply to our processing.

## Third-party websites

Our website may contain links to other websites, including research sources, external resources and third-party services. This Privacy Policy applies only to information handled by Underneath.

We are not responsible for the privacy practices, security or content of third-party websites. We encourage you to review the privacy policies of those websites before providing them with personal information.

## Changes to this policy

We may update this Privacy Policy from time to time to reflect changes to our website, services, information practices or applicable requirements. When we make changes, we will update the “Last updated” date at the top of this page.

We encourage you to review this page periodically so you remain aware of how we handle personal information.

## Contact us

If you have a question about this Privacy Policy, want to exercise a privacy right, or want to understand how we handle your information, contact us at [hello@underneath.agency](mailto:hello@underneath.agency), or by post to Underneath Agency, 2846 Simons Hollow Road, Bloomsburg, PA 17815, United States.

---

This is the Markdown twin of https://underneath.agency/privacy. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Terms of Use | Underneath"
description: "Terms for using underneath.agency: website content, research licensing, intellectual property, acceptable use, liability and how to contact us."
canonical: "https://underneath.agency/terms"
published: 2026-09-26
updated: 2026-09-27
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Legal

# Terms of use

Last updated: September 2026

These Terms of Use govern your use of the Underneath website at **underneath.agency**.

The website is operated by **Underneath Agency**, trading as **Underneath** (“Underneath,” “we,” “us,” or “our”), headquartered at 2846 Simons Hollow Road, Bloomsburg, PA 17815, United States.

By accessing or using this website, you agree to these Terms of Use. If you do not agree with them, please do not use the website.

These Terms apply to use of the website only. Services provided to clients are governed by separate written agreements.

## Website content

The content on this website is provided for general informational purposes. It is intended to explain Underneath, our services, our research and our approach to generative engine optimization. It is not legal, financial, medical, technical or other professional advice for your particular circumstances.

You should not rely on information on this website as a substitute for advice appropriate to your situation.

We try to keep the information on the website accurate and current, but we do not guarantee that every statement, example, statistic, research finding or other piece of content will always be complete, current or error-free. Information about AI systems, search platforms and their behavior can change as those systems change.

## No guarantee of results

References on this website to visibility, citations, recommendations, traffic, leads, conversions, revenue or other outcomes are not guarantees of future results.

AI systems, search engines and other platforms operate independently of Underneath and may change their systems, interfaces, ranking or retrieval behavior without notice. Client results depend on the circumstances of the business, the work performed, external conditions and other factors.

Where an example is identified as illustrative, it is not a representation of a client result. Where we publish client results or case studies, we aim to describe the relevant evidence and limitations accurately.

## Research and published data

Underneath publishes original research into AI search, AI assistants, citations, recommendations, crawling and related subjects. Unless a particular study states otherwise, the research and accompanying materials remain subject to these Terms.

Where a research study or dataset is expressly published under the **Creative Commons Attribution 4.0 International (CC BY 4.0)** license, that study and dataset may be reused in accordance with that license, including with appropriate attribution. The applicable license for each research publication governs that publication.

You may not assume that every piece of content on the Underneath website is covered by the CC BY 4.0 license merely because some of our research is.

## Intellectual property

Unless otherwise stated, the content, design, layout, branding, graphics, logos, trademarks, text and other materials on this website are owned by or licensed to Underneath Agency.

You may view, download or print reasonable portions of the website for your personal or internal business use, provided that you do not remove proprietary notices or use the material in a misleading way.

You may not reproduce, republish, distribute, modify, sell, license or commercially exploit website content without our prior written permission, except where applicable law or an expressly stated license permits you to do so.

Nothing in these Terms transfers ownership of Underneath intellectual property to you.

## Our trademarks and brand

“Underneath,” the Underneath name, logos and related branding are the property of Underneath Agency unless otherwise stated.

You may not use our trademarks or branding in a way that suggests sponsorship, endorsement, partnership or affiliation without our prior written permission.

## Links to other websites

The website may contain links to third-party websites, publications, platforms or resources. These links are provided for convenience or reference.

Underneath does not control third-party websites and is not responsible for their content, availability, security, privacy practices or terms. A link does not necessarily mean that Underneath endorses the third party or its services.

## Acceptable use

You agree not to use the website:

- in violation of applicable law or regulation
- to interfere with the operation or security of the website
- to attempt to gain unauthorized access to systems or information
- to introduce malicious code or harmful software
- to impersonate another person or organization
- to use automated methods to interfere with the website or its operation
- to reproduce or systematically extract website content in a way that is not permitted by these Terms or an applicable license

Nothing in this section restricts legitimate research, accessibility tools, ordinary search-engine crawling, or other activity permitted by law and consistent with the website’s published policies.

## Availability of the website

We may change, suspend or discontinue any part of the website at any time. We do not guarantee that the website or any particular page, resource or feature will always be available or uninterrupted. We may also update, remove or replace content without notice.

## Errors and corrections

Despite our efforts, the website may occasionally contain typographical errors, outdated information, technical errors or other inaccuracies. We may correct or update information when we become aware of an error.

The existence of an error does not create an obligation for Underneath to maintain or update historical versions of website content.

## Limitation of liability

To the extent permitted by applicable law, Underneath and its directors, employees, contractors and service providers will not be responsible for indirect, incidental, special, consequential or punitive losses arising from or related to your use of, or inability to use, this website.

This includes losses arising from reliance on website content, interruption of access, technical problems, third-party websites or changes to information published on the website.

Nothing in these Terms excludes or limits liability that cannot lawfully be excluded or limited under applicable law.

## Indemnity

To the extent permitted by applicable law, you agree to be responsible for losses, claims, liabilities and reasonable costs arising from your misuse of the website or your violation of these Terms.

This provision does not apply to the extent that the relevant loss was caused by Underneath’s own unlawful conduct or by circumstances for which liability cannot lawfully be transferred to you.

## Client services

These Terms govern use of the website only. If you become a client of Underneath, the scope of services, fees, deliverables, responsibilities, confidentiality, intellectual property and other terms of that engagement will be governed by a separate written agreement.

If there is a conflict between these Terms and a signed client agreement, the client agreement will govern the client relationship to the extent of that conflict.

## Privacy

Our handling of personal information is described in our [Privacy Policy](https://underneath.agency/privacy). By using the website, you acknowledge that information may be handled as described in that policy.

## Changes to these Terms

We may update these Terms from time to time. When we make changes, we will update the “Last updated” date at the top of this page.

Your continued use of the website after updated Terms are posted means that the updated Terms will apply to your future use of the website, to the extent permitted by applicable law.

## Governing law

These Terms are governed by the laws applicable in the jurisdiction where Underneath Agency is organized, without regard to conflict-of-law principles, except where applicable law requires otherwise.

Any dispute relating to these Terms or your use of the website will be subject to the jurisdiction of the courts that have authority over the relevant matter, unless applicable law provides otherwise.

## Contact

Questions about these Terms can be sent to [hello@underneath.agency](mailto:hello@underneath.agency), or by post to Underneath Agency, 2846 Simons Hollow Road, Bloomsburg, PA 17815, United States.

---

This is the Markdown twin of https://underneath.agency/terms. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How can a CPA or tax firm win better clients through AI search?"
description: "By being the firm AI assistants tie to a specific niche and place. Early studies show a few large firms dominate answers, but focused smaller firms break in."
canonical: "https://underneath.agency/resources/accounting-firms-clients-ai-search"
published: 2026-10-09
updated: 2026-10-10
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How can a CPA or tax firm win better clients through AI search?

By becoming the firm AI assistants clearly associate with a specific kind of client, service and place. The first studies of AI answers about accounting firms find a small group of very large firms named most of the time, yet several smaller firms beat their size because the web ties them firmly to an industry, a region or a problem. For a firm that is already at capacity, that matters less for volume than for winning the clients it actually wants.

## The short version

1. Referrals still win most clients: in TaxDome’s 2025 survey of 350 US businesses, [57% found their current accountant through a peer referral](https://www.cpapracticeadvisor.com/2025/08/19/survey-of-smbs-shows-how-they-choose-and-evaluate-their-accountant-firm/167532/) and only 3% through advertising.
2. Tax questions are moving into chat: OpenAI said [tax-related ChatGPT searches in the first quarter of 2026 were four times](https://www.govtech.com/question-of-the-day/how-often-did-americans-turn-to-ai-for-help-with-their-taxes-this-year) those of the first quarter of 2025.
3. AI answers about firms are concentrated: in one study reported by [Accounting Today](https://www.accountingtoday.com/news/ai-recommendations-favoring-narrow-band-of-big-firms), 25 firms took 87.2% of all mentions across 3,480 answers, and 79 of the Top 100 Firms shared 3%.
4. Size is not destiny: in a second study, the four busiest single-audit firms in North Carolina’s Research Triangle were recommended 24 times, against 184 times for the six least busy.
5. Capacity makes fit the goal: 73% of 500 firms in [Advancetrack’s 2026 Accounting Talent Index](https://www.consultancy.uk/news/44502/three-in-four-accounting-firms-struggle-to-take-on-new-work-amid-hiring-challenges) said they were turning away potential clients for lack of staff.

A note before you read: this article is about how accounting firms appear in AI answers. Nothing below is tax, accounting or legal advice. If you sell accounting software rather than accounting services, read [our article on accounting software in AI search](https://underneath.agency/resources/accounting-software-ai-search) instead.

## Who hires an accounting firm, and how do they find one now?

Business owners, finance leaders and individuals with complex taxes, and most still start from a trusted person’s referral.

The individual market is large. In the 2026 filing season, the [IRS](https://www.irs.gov/newsroom/filing-season-statistics-for-week-ending-april-17-2026) received 72,821,000 electronically filed returns from tax professionals, out of 137,618,000 e-filed returns in all. Business clients are smaller in number but worth more each year, and they stay.

The best evidence on how businesses choose comes from the 2025 Niche Business Accounting Report, commissioned by TaxDome, a practice management software company, and reported by CPA Practice Advisor. It surveyed 350 US businesses with revenue between $1M and $100M:

- **Referrals dominate.** 57% found their current accountant through a peer referral; 3% chose one through advertising.
- **Specialists keep clients.** 98% of businesses that left a specialist moved to another specialist, not back to a generalist.
- **Specialization earns a premium.** Businesses said they would pay up to 25% more for specific offerings, and companies above $1M in revenue were 2× more likely to hire niche specialists.
- **Bigger clients check the firm’s tools.** Among businesses paying $10K+ per year, 83% said a firm’s use of technology is a key factor.

TaxDome sells to accounting firms, so treat this as vendor research. The pattern still matters for AI search. A referral gives a buyer a name, not a decision. We infer that the buyer then checks that name somewhere, and an AI assistant is now one of the places they check. The same habit is spreading across [law, consulting and other professional firms](https://underneath.agency/resources/professional-services-firms-clients-ai-search).

## Where do AI assistants already sit in that search?

At the question stage today, and increasingly at the “who should I hire?” stage that follows.

People bring tax questions to chat in volume. According to [GovTech’s summary of OpenAI’s figures](https://www.govtech.com/question-of-the-day/how-often-did-americans-turn-to-ai-for-help-with-their-taxes-this-year), one-third of tax-related ChatGPT queries were about earnings and withholdings, 30 percent were for help filing forms and using tax software, and 10 percent concerned investments and retirement reporting. OpenAI’s note on the chart said ChatGPT “is not intended to replace professional advice.”

That reminder matches what filers say. In a [Qlik study reported by MoneyLion](https://moneylion.com/trending/money/how-ai-changing-way-americans-file-taxes-what-it-still-gets-wrong), 11% of people who filed a return used AI to some degree, rising to 23% of 18- to 24-year-olds against 2% of those 55 and older. In a [Censuswide survey for Invoice Home](https://www.cpapracticeadvisor.com/2026/02/19/most-taxpayers-trust-tax-pros-over-ai-for-tax-preparation-survey-finds/178412/) of more than 2,000 filers, only 37% would consider using AI to file over hiring a tax professional, down from 43% in 2025.

Our reading: AI is not replacing the accountant. It is becoming the place people work out whether they need one, and which kind. A business owner who asks how to handle a multistate sales tax notice is one follow-up question away from “who can help me with this?”

Local questions are already answered with named firms. When we asked ChatGPT “Who is the best accountant in {city}?” and similar questions for five professions in four countries, [it answered with lists of named local businesses](https://underneath.agency/research/chatgpt-local-recommendations-study), each with an address, phone number, rating and link.

## What do prospective clients ask AI about accounting firms?

Questions that combine a service, a type of client and often a place. We wrote the sample prompts in this table to show how such questions tend to be phrased; none was taken from a real client’s search.

| Buyer | Illustrative prompt |
|---|---|
| Business owner | “Which CPA firms in Columbus specialize in dental practices?” |
| Startup CFO | “Outsourced accounting firm for a venture-backed SaaS company, under 50 employees” |
| Nonprofit director | “Who does single audits for charter schools near Raleigh?” |
| Individual | “I got an IRS levy notice. Do I need a CPA, an enrolled agent or a tax attorney?” |
| Investor | “Tax preparer who understands crypto trading and K-1s” |
| Any buyer | “Is [firm name] any good? What do clients say?” |

Two features stand out. First, the questions are specific: industry, service and location in one sentence. Second, the last one is a reputation check, the kind a referred buyer runs before calling. Our [study of “is it legit?” questions](https://underneath.agency/research/is-it-legit-ai-reputation-study) found 88.0% of answers cited a review or complaint platform, so what clients say in public becomes part of the answer.

## How does an AI mention turn into a client engagement?

Through a short path: the answer names the firm, the buyer reads its niche page, then requests a consultation.

The path looks like a referral, which is why it can be valuable: an AI answer → the firm’s page on that industry or service → a call or consultation request → an engagement letter → recurring tax, advisory or audit work. Because accounting relationships last for years and specialists rarely lose clients to generalists, one well-matched client is worth many years of fees.

The more useful question for most firms is which clients, not how many. Advancetrack’s survey found 73% of firms turning away potential clients, and 45% said the talent shortage is worse than three years ago. The AICPA’s [2025 National MAP Survey summary](https://www.aicpa-cima.com/resources/article/key-small-firm-insights-in-the-2025-national-map-survey) points the same way: a median 6.7% growth rate among participating firms, a push to fix underpricing, and fewer firms shedding unsuitable clients, which it reads as firms having narrowed their focus to the most attractive ones.

The services firms want to grow are named in the profession’s own data. Among the Accounting Today Top 100, [client advisory services were the most common source of growth](https://www.accountingtoday.com/list/inside-the-2026-top-100-now-leaving-the-station), with 75 firms expanding their practices, and tax made up roughly two-fifths of revenue for firms below $1 billion. A reasonable expectation is that AI visibility pays most when it brings inquiries for those higher-value services from the industries a firm already serves well. If advisory is your growth line, see [how advisory firms get named by owners](https://underneath.agency/resources/business-advisory-firms-clients-ai-search).

## Which accounting firms do AI assistants name today?

Mostly the largest firms, but smaller firms with a clear specialty and place can outperform their size.

These findings are observed in studies, not documented by any platform. Accounting Today reported two in October 2026, both run by marketing advisory firms.

**The concentration study.** Researchers asked 100 buyer questions about tax, audit, advisory and risk services, five times each across seven assistants, and collected 3,480 responses. Twenty-five firms accounted for 87.2% of all mentions; the first 10 took 61.3%, and RSM alone appeared 29.48% of the time. Eight Top 100 Firms were never mentioned at all. Several smaller firms beat what their revenue would predict through an established association with a particular industry, service, place or problem.

**The single-audit study.** Researchers took 10 firms named as auditor on three or more single audits in the Research Triangle, using the public federal record, and asked five assistants 10 buyer questions, five times each: 250 answers in August 2026 and 250 more in September. The four busiest firms, with 45 of 75 federal engagements, were recommended 24 times in August; the six least busy were recommended 184 times. The firms’ own websites were the most cited source, especially pages that named the region, described the buyer’s kind of work at length and had been updated recently. The busiest firm had 1,821 pages, none of which mentioned its place. Across 2,725 cited pages, the Federal Audit Clearinghouse, the authoritative record of who did the work, was cited only four times.

Both studies were produced by firms that sell marketing services, and the single-audit study covers one service in one region. Read them as early observations, not rules.

Our own local studies add the map. Of the businesses Google Maps ranked 1 to 3 for a local search, [ChatGPT listed 67.7%](https://underneath.agency/research/chatgpt-local-picks-google-profile-study), and businesses with more reviews than the local median were 19.5 points more likely to be listed, after adjusting for rank and other signals. Accountants were one of the five professions in that sample.

## What does it cost a firm to be missing or misdescribed?

Lost chances at the clients a firm wants most, and wrong details that misdirect prospects. No one has measured the revenue yet.

**Being absent.** In the concentration study, 79 of the Top 100 Firms together received only 3% of mentions. For a regional firm, the risk is that a referred prospect asks an assistant about its specialty and hears only national names. That is our inference; no study has yet measured lost engagements.

**Being misdescribed.** When we asked four assistants for the address, phone, website and hours of real local businesses, [answers about accountants differed from the Google profile 17.7% of the time](https://underneath.agency/research/ai-business-facts-accuracy-study). Most differences traced to the business’s own sources disagreeing with each other.

**Competing with capital.** The largest firms are consolidating and investing. The Top 100 reported 225 mergers in 2025, up from 122 in 2024, and in June 2026 the [Journal of Accountancy reported](https://www.journalofaccountancy.com/news/2026/jun/crowe-partners-with-private-equity-baker-tilly-on-the-move/) that KKR would make a “significant equity investment” in Crowe. Larger, better-funded firms naturally produce more of the web material that assistants learn from. The wider pattern is covered in [do AI assistants favor big brands over smaller competitors?](https://underneath.agency/resources/do-ai-assistants-favor-big-brands)

## How does GEO work for an accounting firm?

Generative engine optimization (GEO) makes a firm’s specialties, places and credentials easy for assistants to find and repeat.

1. **One page per niche and place.** A page for each industry and service the firm wants to grow, naming the region it serves, describing the client’s actual work in their words, and dated when it changes. The single-audit study suggests this is the kind of page assistants cite.
2. **Consistent facts everywhere.** The same name, address, phone, hours and partner names on the website, Google Business Profile, CPA directories and association listings. If AI answers already get something wrong, [here is how to correct wrong brand information](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers).
3. **Reviews and reputation.** Ask satisfied clients to review the firm on the platforms people already check, and answer criticism calmly. In our local study, review volume was the signal that held everywhere.
4. **Credible third-party mentions.** Industry association talks, trade press in the client’s industry, state society committees, rankings and lists. These place the firm’s name next to the industry it serves.
5. **People pages.** Partner bios with credentials, licenses and the industries each partner serves, so assistants can connect expertise to the firm. Medical practices do the same with [physician profiles that AI can describe](https://underneath.agency/resources/healthcare-providers-patients-ai-search).
6. **Search basics.** Pages that rank for the niche still matter; see [does ranking on Google get you into AI Overviews?](https://underneath.agency/resources/does-ranking-on-google-get-you-into-ai-overviews)
7. **Regular checks.** Ask the niche, local and reputation questions across assistants each quarter and note who is named and which pages are cited.

Every claim on these pages must also meet the profession’s rules on advertising, which bar false or misleading statements. That limit is useful: plain, verifiable facts are what assistants can repeat accurately. GEO cannot guarantee that an assistant names a firm; it makes the evidence about the firm accurate, specific and easy to find.

## What can’t the current research tell a CPA firm?

It shows who gets named, not how often an AI mention becomes a signed engagement.

- **No link to engagements yet.** We found no public data connecting a firm’s AI visibility to consultations or new clients.
- **Few studies, from interested parties.** The two accounting studies came from marketing firms and cover limited questions, services and places.
- **Answers move.** Our local study found that two answers to the same question shared only part of their business lists; one check is a snapshot.
- **Vendor surveys.** The client-behavior figures come largely from software companies and their survey partners.

## Where should a CPA or tax firm start?

Start with the two or three niches you most want to grow, and ask assistants who serves them locally.

Write down the questions a well-matched client would ask: the industry, the service, the place and the problem. Type each one into ChatGPT, Gemini, Perplexity and Google AI Mode the way a business owner or controller would. Note which firms are named, which pages are cited, and whether your firm’s details are right. The gap between that list and the clients you want is your plan.

If your partners want more of the right advisory, tax or audit clients without adding to the work you already turn away, [ask us to review your firm’s AI visibility and plan the work by niche](https://underneath.agency/contact). We will test how assistants answer your niche, local and reputation questions, show which sources they rely on, and plan the pages, listings and coverage that tie your firm to the clients you want. Our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page explains how that niche-by-niche work on pages, directory listings and reviews is diagnosed, carried out and measured.

## Frequently asked questions

### Does ChatGPT recommend local accountants?

Yes. In our study, ChatGPT answered questions such as “Who is the best accountant in {city}?” with named local firms, and listed 67.7% of the businesses Google Maps ranked in its top 3.

### Can a smaller CPA firm appear in AI answers ahead of large firms?

Sometimes. In a single-audit study, the six least busy firms were recommended 184 times against 24 for the four busiest, mostly through pages tied to the region and the buyer’s work.

### Will AI replace the need for a tax professional?

Filers do not think so yet. Only 37% would consider AI over a tax professional, down from 43%, and OpenAI says ChatGPT is not intended to replace professional advice.

### Do online reviews matter for AI recommendations of accounting firms?

They appear to. In our local study, businesses with more reviews than the local median were 19.5 points more likely to be listed by ChatGPT, after adjusting for map rank.

## Sources

- CPA Practice Advisor (2025-08-19), [Survey of SMBs Shows How They Choose and Evaluate Their Accounting Firm](https://www.cpapracticeadvisor.com/2025/08/19/survey-of-smbs-shows-how-they-choose-and-evaluate-their-accountant-firm/167532/)
- CPA Practice Advisor (2026-02-19), [Most Taxpayers Trust Tax Pros Over AI for Tax Preparation, Survey Finds](https://www.cpapracticeadvisor.com/2026/02/19/most-taxpayers-trust-tax-pros-over-ai-for-tax-preparation-survey-finds/178412/)
- GovTech (2026-04-15), [How often did Americans turn to AI for help with their taxes this year?](https://www.govtech.com/question-of-the-day/how-often-did-americans-turn-to-ai-for-help-with-their-taxes-this-year)
- MoneyLion (2026-08-25), [How AI Is Changing the Way Americans File Taxes — and What It Still Gets Wrong](https://moneylion.com/trending/money/how-ai-changing-way-americans-file-taxes-what-it-still-gets-wrong)
- Accounting Today (2026-10-07), [AI recommendations favoring narrow band of big firms](https://www.accountingtoday.com/news/ai-recommendations-favoring-narrow-band-of-big-firms)
- Accounting Today (2026-03-09), [Inside the 2026 Top 100: Now leaving the station](https://www.accountingtoday.com/list/inside-the-2026-top-100-now-leaving-the-station)
- Consultancy.uk (2026-06-16), [Three-in-four accounting firms struggle to take on new work amid hiring challenges](https://www.consultancy.uk/news/44502/three-in-four-accounting-firms-struggle-to-take-on-new-work-amid-hiring-challenges)
- AICPA & CIMA (2025-11-28), [Key small firm insights in the 2025 National MAP Survey](https://www.aicpa-cima.com/resources/article/key-small-firm-insights-in-the-2025-national-map-survey)
- Journal of Accountancy (2026-06), [Crowe partners with private equity; Baker Tilly on the move](https://www.journalofaccountancy.com/news/2026/jun/crowe-partners-with-private-equity-baker-tilly-on-the-move/)
- Internal Revenue Service (2026-04), [Filing season statistics for week ending April 17, 2026](https://www.irs.gov/newsroom/filing-season-statistics-for-week-ending-april-17-2026)
- Underneath (2026), [ChatGPT local recommendations: stable details, shifting lists](https://underneath.agency/research/chatgpt-local-recommendations-study)
- Underneath (2026), [Which Google Maps businesses does ChatGPT recommend?](https://underneath.agency/research/chatgpt-local-picks-google-profile-study)
- Underneath (2026), [Do AI answers match a business’s Google profile?](https://underneath.agency/research/ai-business-facts-accuracy-study)
- Underneath (2026), [“Is this brand legit?” How AI assistants build a reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study)

---

This is the Markdown twin of https://underneath.agency/resources/accounting-firms-clients-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Are AI assistants choosing small business accounting software?"
description: "Partly. AI assistants now shape which accounting tools owners and their accountants consider, so vendors must be named, priced and described correctly there."
canonical: "https://underneath.agency/resources/accounting-software-ai-search"
published: 2026-10-09
updated: 2026-10-10
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Are AI assistants now choosing accounting software for small businesses?

Not on their own, but they increasingly shape the shortlist, both for business owners and for the accountants who advise them. A small business owner who asks ChatGPT which accounting software fits a two-person landscaping company gets a short list of names, and so may the bookkeeper who sets that business up. For an accounting software company, that answer now sits in front of a subscription worth years of recurring revenue.

## The short version

1. The prize is large and sticky: Xero reported 4.92 million customers paying an average of NZ$55.44 a month in fiscal 2026, with average monthly churn of 1.14%, and Intuit’s QuickBooks Online Accounting revenue grew 22 percent in fiscal 2025.
2. Accountants are a second gatekeeper: more than 80% of mid-market QuickBooks Online customers work with an external firm, Intuit said at its 2025 investor day, and accountants in Intuit’s survey of 700 professionals use AI daily at nearly double the rate of small businesses (46% against 28%).
3. Switching events create waves of new buyers: about 780,000 UK sole traders and landlords must use compatible software for Making Tax Digital from April 2026, and a further 970,000 from April 2027.
4. Accounting buyers often pick wrong: in Capterra’s 2026 survey of buyers in the accounting profession, 44% ended up with software that disappointed, and around a third (36%) of purchases were followed by finding a better match.
5. AI answers about this category are not neutral or fixed: in our country study, ChatGPT put QuickBooks first in the US and Canada and Xero first in the UK and Australia, and in our pricing study only 61.9% of software plan prices that four assistants quoted were fully faithful to the vendor’s page.

## Who picks the books software, and how long does a subscriber stay?

Owners of very small businesses buy most of it, often with an accountant or bookkeeper choosing or approving the product.

The market is wide. The [US Small Business Administration’s Office of Advocacy](https://advocacy.sba.gov/?p=30239) counts 36.2 million small businesses in the United States, and almost every one of them has to keep books. Most accounting software is sold to them as a monthly subscription, often after a free trial, which makes the first choice unusually valuable.

The two leading cloud platforms show what a customer is worth. [Xero’s fiscal 2026 results](https://announcements.asx.com.au/asxpdf/20260514/pdf/06zkqzlkq0xnph.pdf) report 4.92 million customers, average revenue per customer of NZ$55.44 a month (up 23%), and average monthly churn of 1.14%. Xero puts the total lifetime value of its customers at NZ$21.0 billion. At [Intuit](https://investors.intuit.com/_assets/_940994c421caf4d8d251c405ed856b0d/intuit/news/2025-08-21_Intuit_Reports_Strong_Fourth_Quarter_and_Full_1266.pdf), QuickBooks Online Accounting revenue grew 22 percent in fiscal 2025, driven by “higher effective prices, customer growth, and mix shift”, inside a business solutions group with $11.1 billion of revenue.

The buyer is rarely only the owner. According to a [report on Intuit’s 2025 investor day](https://report.woodard.com/articles/intuit-investor-day-2025-game-changers-accountants-need-to-know-fpwr), more than 80% of mid-market QuickBooks Online customers already work with external firms. Accountants and bookkeepers recommend a platform, set clients up on it and often move whole client books from one product to another. That gives accounting software two journeys to win: the owner’s and the advisor’s.

## Where do AI assistants show up in that buying journey?

At both ends: owners ask AI assistants for recommendations, and accountants use AI tools daily in their own work.

On the owner side, the evidence is early and comes from small samples. A [Revenued survey](https://www.revenued.com/small-business-ai-usage) of nearly 300 small business owners found over 90% using at least one AI tool by September 2025, and nearly 70% of those AI users relied on ChatGPT or other OpenAI models. Revenued sells financial products to small businesses and is not a research firm, so read this as a direction, not a measurement. Across all software categories, [G2’s August 2025 survey](https://learn.g2.com/ai-search-surging-for-b2b-buyers) of more than 1,000 buyers found 87% saying AI chatbots are changing the way they research.

On the advisor side, the signal is stronger. Intuit’s [2025 Accountant Technology Survey](https://investors.intuit.com/_assets/_6a9ee074199c0a511699c988d19108db/intuit/news/2025-07-30_Accountants_Embrace_AI_and_Strategic_Advisory_1263.pdf) of 700 US accounting professionals found 46% using AI daily, nearly double the 28% rate among small businesses, and 64% planning to invest in or upgrade AI over the next year. The people who recommend accounting software are among the heaviest AI users in the small business world.

Google is part of this too. An owner who types a bookkeeping software query into Google will usually see an AI Overview, the AI summary at the top of the results: in [our study of 800 US searches](https://underneath.agency/research/ai-overviews-frequency-study), B2B software and technology keywords triggered one on 96.0% of searches, the highest of the eight industries we tested.

The platforms themselves are moving into the assistants. Intuit signed a [multi-year contract worth more than $100 million](https://techcrunch.com/2025/11/18/intuit-signs-100m-deal-with-openai-to-bring-its-apps-to-chatgpt) with OpenAI so that QuickBooks and its other apps can work inside ChatGPT, and Xero says it built a connector into Claude.ai. The category leaders are treating AI assistants as a place where customers will meet their products. For the wider software picture, see [our B2B SaaS article](https://underneath.agency/resources/b2b-saas-revenue-from-ai-search).

## What do owners and bookkeepers ask assistants about books software?

Fit-for-my-business questions, switching questions, compliance questions and price questions, usually with the business size or trade attached.

We wrote these prompts to mirror situations owners and accountants face, such as a new tax rule or a discontinued product; we did not observe them:

- **Fit:** “What is the best accounting software for a cleaning business with five employees and job costing?”
- **Switching:** “What should I move to now that I can’t buy a new QuickBooks Desktop Pro subscription?”
- **Compliance:** “Which software is compatible with Making Tax Digital for a landlord with two properties?”
- **Integration:** “Which accounting software connects best to Shopify and Stripe and handles sales tax?”
- **Outgrowing:** “When should a $10 million distributor move off QuickBooks Online?”
- **Accountant-side:** “Which platform makes it easiest to manage 60 bookkeeping clients?”

Several of these are driven by real events. Intuit stopped selling new subscriptions to QuickBooks Desktop Pro Plus, Premier Plus and Mac Plus in the US after [September 30, 2024](https://insightfulaccountant.com/accounting-tech/general-ledger/desktop-product-discontinuation-september-30-2024), and encouraged accountants to move clients online. In the UK, [HMRC](https://www.gov.uk/government/news/one-year-until-making-tax-digital-for-income-tax-launches) requires sole traders and landlords with qualifying income above £50,000 to keep digital records and use compatible software from April 2026, about 780,000 people, with 970,000 more from April 2027. Each event sends a wave of buyers looking for an answer at the same moment.

Geography changes the answer. In [our country study](https://underneath.agency/research/ai-recommendations-by-country-study), ChatGPT named both QuickBooks and Xero for small business accounting software in every country, but put QuickBooks first in every US and Canadian run and Xero first in every UK and Australian run. Across all 40 questions in that study, ChatGPT put the same brand first in all four countries for only 25.4% of questions. Location matters just as much for [payroll software in AI answers](https://underneath.agency/resources/payroll-software-ai-search), where tax rules differ by state and country.

## How does an AI recommendation turn into subscription revenue?

Through a trial for owners, through the accountant for advised businesses, and increasingly inside the assistant itself.

**The owner path.** The owner asks, gets three or four names, opens one or two trials and subscribes to the product that imports their bank feed and invoices without friction. The value is the monthly fee multiplied by how long they stay. With Xero’s monthly churn of 1.14%, we infer that a typical new customer stays for years, so one recommendation can carry several years of revenue, plus payroll, payments and add-ons sold later.

**The accountant path.** The accountant asks an assistant about a niche need, such as construction job costing or multi-currency, compares options and then standardizes clients on one platform. One firm’s decision can move dozens of client subscriptions. Accountants in Capterra’s [2026 buying survey](https://www.capterra.com/resources/accounting-trends-technology-strategy), based on 139 buyers in the accounting profession, named product identification as a challenge 33% of the time, so help in finding the right product has real value to them.

**The in-assistant path.** When a product works inside ChatGPT, discovery and use can happen in the same conversation. Intuit said its apps in ChatGPT would help it reach new audiences beyond its [approximately 100 million customers](https://www.cpapracticeadvisor.com/2025/11/18/intuit-reaches-multiyear-deal-with-openai-to-bring-its-apps-to-chatgpt/173359/), on a platform with 700 million-plus weekly users. How many paying QuickBooks customers this produces has not been published.

## Why does an assistant put QuickBooks, Xero or a challenger first?

Platforms document how they search, not how they choose; studies point to geography, consistent facts and independent coverage.

**Documented by the platforms.** Before answering, AI Overviews and AI Mode may use what [Google calls](https://developers.google.com/search/docs/appearance/ai-features) a “query fan-out” technique, issuing several related searches across subtopics. A single question about the best software for a landlord can therefore pull in pages about tax rules, pricing and reviews at the same time. Google also says there are no extra requirements for appearing in these features beyond normal search eligibility.

**Observed in studies.** Our country study shows the user’s location changes which accounting brand comes first. [Our four-assistant study](https://underneath.agency/research/ai-assistants-brand-agreement-study) found B2B software had the highest agreement between ChatGPT, Gemini, Perplexity and Claude of eight industries, with an overlap score of 0.543 on a 0-to-1 scale. We infer that once a category’s leaders are fixed in answers, challengers face a crowded shortlist.

**Our inference about trust factors specific to accounting.** A reasonable expectation is that assistants lean on the signals accountants themselves trust: listings in tax authority software directories, such as HMRC’s list of compatible products; partner and certification programs; integration directories; reviews from accountants and bookkeepers; and plain statements of which countries, tax rules and business sizes a product supports. No platform has confirmed that any of these accounting-specific signals affects which product it names.

## What does a books platform lose when assistants leave it off the list?

A customer who signs up elsewhere is usually gone for years, because changing books is painful.

The economics above make the cost clear: a subscription that lasts years, sold to a business that rarely switches without a trigger. When the triggers come, such as a discontinued desktop product or a new tax rule, buyers decide quickly. If the assistant’s shortlist leaves you out at that moment, you miss the switching window, not just one sale.

Buyers also struggle to find the right fit without help. In Capterra’s survey of accounting-profession buyers, 44% ended up choosing software that disappointed them, and around a third (36%) of purchases were followed by finding a better match later. We infer that a product that answers fit questions clearly has a chance to be the better match before the purchase, not after it.

Errors cost money too. In [our pricing study](https://underneath.agency/research/ai-pricing-accuracy-study), 61.9% of the plan prices four assistants gave for 45 software products were fully faithful to the official pricing page; 7.6% differed, usually because an older figure was still published somewhere. For a category that competes on monthly price, an outdated plan price in an AI answer can lose the comparison.

## How does GEO work for accounting software?

Generative engine optimization (GEO) gives assistants accurate, checkable facts about your plans, markets and integrations. No accounting vendor can be promised a recommendation.

For an accounting software company it usually means:

1. **One clear story everywhere.** State the business sizes, trades, countries, tax rules and integrations you support in the same words on your site, app marketplaces, review profiles and partner directories.
2. **Country pages that stand on their own.** Since answers change by country, publish separate, current pages for each market’s compliance needs, such as Making Tax Digital in the UK or sales tax in the US.
3. **Advisor-facing proof.** Accountants and bookkeepers write reviews, answer forum questions and publish comparisons. Their independent coverage is the third-party evidence assistants can cite.
4. **Migration and comparison content.** Publish fair guides for buyers leaving desktop software or outgrowing a starter tool, including where your product is not the right fit.
5. **One current pricing page.** Retire old plan pages and figures so assistants find one price per plan. When an assistant still quotes a retired plan, [fixing wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) shows how to find the page it came from.
6. **Measurement by market.** Track a fixed set of owner and accountant questions in each country you sell in, across ChatGPT, Gemini, Perplexity, Copilot and Google, and compare it with trial signups. Expect answers to vary run to run; see [why AI answers about your brand change](https://underneath.agency/resources/why-ai-answers-about-your-brand-change).

## What don’t we know yet about AI answers and accounting software sales?

How many books subscriptions begin with an AI answer; no published study has counted them.

- The owner surveys are small or come from companies with a stake in AI, and the accountant data comes from Intuit, the market leader.
- No accounting software vendor has published AI referral or trial data, so we cannot quote conversion rates for this category.
- Our country finding comes from one day of five runs per country on ChatGPT and Gemini; answers may differ on other days or assistants.
- Whether being named by an AI assistant matters more than an accountant’s recommendation is unknown. Most likely they work together, but that is our inference.
- AI visits often leave no clear trace in analytics; see [why analytics miss AI visibility](https://underneath.agency/resources/why-analytics-miss-ai-visibility).

## How can an accounting software company see whether AI answers win it subscribers and accountant partners?

Check, market by market, what the main assistants say about you for your best customers’ questions.

Run the fit, switching, compliance and price questions your owners and accountants actually ask, note who is named first and whether your plans and prices are right, then compare the gaps with trial and partner signup data. For an outside reading of those answers, weighted toward subscriptions and accountant adoption, [ask us for an accounting software visibility review](https://underneath.agency/contact). The [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page describes how we turn that review into country-by-country fixes for plan facts, comparison pages and accountant-facing coverage.

## Frequently asked questions

### Does ChatGPT recommend QuickBooks or Xero?

Both, and the order depends on where the user is. In our country study, ChatGPT named both everywhere but led with QuickBooks in the US and Canada and with Xero in the UK and Australia. That was one study day, so treat it as a snapshot.

### Do accountants still influence which accounting software small businesses buy?

Yes. Intuit said more than 80% of its mid-market QuickBooks Online customers work with external firms. Accountants also use AI more than their clients, so they may meet AI answers about your product before the owner does.

### Can a smaller accounting software company get named against QuickBooks and Xero?

It can for specific needs, such as a trade, a country or an integration, more easily than for “best accounting software” in general. Our article on [how small brands get recommended by AI](https://underneath.agency/resources/how-small-brands-get-recommended-by-ai) summarizes the research.

### Are AI assistants accurate about accounting software prices?

Often, but not always. In our test of 45 software products, 61.9% of quoted plan prices were fully faithful to the vendor’s page, so buyers should confirm on the pricing page and vendors should keep only one current price per plan online.

## Sources

- U.S. SBA Office of Advocacy (2025), [New Advocacy Report Shows the Number of Small Businesses in the U.S. Exceeds 36 million](https://advocacy.sba.gov/?p=30239)
- Xero (2026), [FY26 annual results market release](https://announcements.asx.com.au/asxpdf/20260514/pdf/06zkqzlkq0xnph.pdf)
- The Motley Fool Australia (2026), [Xero FY26 result: Revenue surges 31% but profit dips due to Melio acquisition costs](https://www.fool.com.au/2026/05/14/xero-fy26-result-revenue-surges-31-but-profit-dips-due-to-melio-acquisition-costs/)
- Intuit (2025), [Intuit Reports Strong Fourth Quarter and Full Year Fiscal 2025 Results](https://investors.intuit.com/_assets/_940994c421caf4d8d251c405ed856b0d/intuit/news/2025-08-21_Intuit_Reports_Strong_Fourth_Quarter_and_Full_1266.pdf)
- Intuit (2025), [2025 Intuit QuickBooks Accountant Technology Survey release](https://investors.intuit.com/_assets/_6a9ee074199c0a511699c988d19108db/intuit/news/2025-07-30_Accountants_Embrace_AI_and_Strategic_Advisory_1263.pdf)
- Woodard Report (2025), [Intuit Investor Day 2025: Game-Changers Accountants Need to Know](https://report.woodard.com/articles/intuit-investor-day-2025-game-changers-accountants-need-to-know-fpwr)
- Capterra (2026), [Accounting Software Buyer Trends: Action Areas for 2026](https://www.capterra.com/resources/accounting-trends-technology-strategy)
- HM Revenue & Customs (2025), [One year until Making Tax Digital for Income Tax launches](https://www.gov.uk/government/news/one-year-until-making-tax-digital-for-income-tax-launches)
- Insightful Accountant (2024), [Desktop Product Discontinuation: September 30, 2024](https://insightfulaccountant.com/accounting-tech/general-ledger/desktop-product-discontinuation-september-30-2024)
- Revenued (2025), [AI Usage Among Small Businesses](https://www.revenued.com/small-business-ai-usage)
- G2 (2025), [How AI Chat is Rewriting B2B Software Buying](https://learn.g2.com/ai-search-surging-for-b2b-buyers)
- TechCrunch (2025), [Intuit signs $100M+ deal with OpenAI to bring its apps to ChatGPT](https://techcrunch.com/2025/11/18/intuit-signs-100m-deal-with-openai-to-bring-its-apps-to-chatgpt)
- CPA Practice Advisor (2025), [Intuit Reaches Multiyear Deal with OpenAI to Bring Its Roster of Apps to ChatGPT](https://www.cpapracticeadvisor.com/2025/11/18/intuit-reaches-multiyear-deal-with-openai-to-bring-its-apps-to-chatgpt/173359/)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- Underneath (2026), [Same question, four countries](https://underneath.agency/research/ai-recommendations-by-country-study), [When does Google show an AI Overview?](https://underneath.agency/research/ai-overviews-frequency-study), [software pricing accuracy](https://underneath.agency/research/ai-pricing-accuracy-study) and [four-assistant agreement](https://underneath.agency/research/ai-assistants-brand-agreement-study)

---

This is the Markdown twin of https://underneath.agency/resources/accounting-software-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How AI agent companies get onto buyers’ AI search shortlists"
description: "By making their agent easy to verify: clear use cases, public security and governance evidence, honest pricing and independent proof that assistants can find."
canonical: "https://underneath.agency/resources/ai-agent-companies-customers-from-ai-search"
published: 2026-10-09
updated: 2026-10-10
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How do AI agent companies get onto buyers’ shortlists in AI search?

By being easy to tell apart from the hype: a clearly named use case, public evidence of what the agent does and how it is governed, and independent proof that AI assistants can find. Buyers of AI agents are curious but wary, and a growing share of their research now runs through ChatGPT, Gemini, Perplexity and Google’s AI features. The agent companies that get named are, we expect, the ones that answer the buyer’s safety questions before they are asked.

## The short version

1. [Gartner](https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027) expects 33% of enterprise software applications to include agentic AI by 2028, up from less than 1% in 2024, but estimates that only about 130 of the thousands of agentic AI vendors are real.
2. The same firm predicts over 40% of agentic AI projects will be canceled by the end of 2027, because of escalating costs, unclear business value or inadequate risk controls.
3. Trust is falling as adoption rises: in [Capgemini’s survey](https://www.capgemini.com/news/press-releases/trust-and-human-ai-collaboration-set-to-define-the-next-era-of-agentic-ai-unlocking-450-billion-opportunity-by-2028/) of 1,500 large-company executives, confidence in fully autonomous AI agents dropped from 43% to 27% in a year.
4. Security is the sharpest worry: 80% of organizations told [SailPoint](https://nhimg.org/wp-content/uploads/2025/09/SailPoint-The-Rising-Risk-of-AI-Agents-Expanding-the-Attack-Surface-report-SP2648-.pdf) their AI agents had performed unintended actions, and only 44% had governance policies for them.
5. Spending is still early: [Menlo Ventures](https://menlovc.com/perspective/2025-the-state-of-generative-ai-in-the-enterprise/) puts agent platforms at $750 million in 2025, 10% of horizontal AI application spend, with copilots taking 86%.

## Which companies are buying AI agents, and how much can one account grow?

Mostly large companies moving from pilots to production; a customer’s value grows with every task the agent takes on.

Most buyers are early. Capgemini found that nearly a quarter of large organizations had launched agent pilots and 14% had begun implementation, while only 2% had fully scaled deployment. A [Gartner poll of 147 CIOs](https://www.gartner.com/en/newsroom/press-releases/2025-06-11-gartner-predicts-that-guardian-agents-will-capture-10-15-percent-of-the-agentic-ai-market-by-2030) found 24% had deployed a few agents and 50% were researching and experimenting. Intent is strong: in [Microsoft’s 2025 Work Trend Index](https://www.microsoft.com/en-us/worklab/work-trend-index/2025-the-year-the-frontier-firm-is-born), 81% of leaders expected agents to be moderately or extensively integrated into their AI strategy within 12–18 months.

The buyer depends on the job. Asked about use cases, a majority of 125 leaders in the same Gartner poll said their agents focus on internal administration such as IT, HR and accounting, and 23% on customer-facing work. Developers are a buying group of their own: among developers who use agents at work, 84% use them for software development, per [Stack Overflow’s 2025 survey](https://survey.stackoverflow.co/2025/ai). Smaller companies move faster. In [LangChain’s survey](https://www.langchain.com/stateofaiagents), 51% of respondents had agents in production, rising to 63% at mid-sized companies. Customer-facing voice agents for contact centers follow their own path, covered in [our guide to AI voice companies](https://underneath.agency/resources/ai-voice-software-revenue-from-ai-search).

Pricing makes the customer’s value elastic. Many agent vendors charge per task rather than per seat. Intercom prices its Fin customer service agent at [$0.99 per outcome](https://fin.ai/pricing), and Salesforce sells [Agentforce Flex Credits](https://www.salesforce.com/agentforce/pricing/) at $500 per 100,000 credits, paid per action. Revenue therefore depends less on the signature than on how far the agent spreads once it works. Capgemini projects that organizations with scaled agent deployments will generate about $382 million each over three years, against about $76 million for others, which is the size of prize buyers are weighing.

## How do buyers evaluate an AI agent before they buy?

Through pilots and proofs of concept, judged on reliability, cost and control.

Agents are judged harder than ordinary software because they act. LangChain found performance quality was the top concern for small companies, cited by 45.8%, against 22.4% for cost. Many teams still have humans checking agent output by hand. Gartner warns that most agentic projects today are “early stage experiments or proof of concepts that are mostly driven by hype,” which is why so many will be canceled.

Buyers are also learning to discount labels. Gartner describes “agent washing,” the rebranding of chatbots, assistants and robotic process automation as agents without real agentic capabilities. [Menlo Ventures’ data](https://menlovc.com/perspective/2025-the-state-of-generative-ai-in-the-enterprise/) suggests the skepticism is warranted: only 16% of enterprise deployments it studied qualified as true agents that plan, act, observe and adapt.

Software buyers in general now start that evaluation with AI. In [G2’s July 2026 survey](https://sell.g2.com/2026-buyer-behavior-report) of 1,038 software decision-makers, more than 80% had sourced software recommendations from an AI chatbot in the last two years, and 39% named IT security review as the biggest delay between choosing a vendor and buying. That survey covers all software, not agents alone, and comes from a review platform with a stake in the topic.

## Which questions do agent buyers ask AI assistants?

Mostly use-case, comparison and safety questions, because the category is new and crowded.

We wrote the example prompts below to show how an operations, IT or support leader might shop for an agent; none was observed in real use:

- Use case: “Which AI agents can resolve tier-one support tickets in Zendesk without a human?”
- Build or buy: “Should we build an internal IT help desk agent or buy one?”
- Comparison: “Salesforce Agentforce vs Sierra vs Decagon for retail customer service.”
- Pricing: “How much does an AI SDR agent cost per meeting booked?”
- Governance: “Which AI agent platforms support audit logs, role-based permissions and human approval steps?”
- Legitimacy: “Is [vendor] a real agent or a chatbot with workflows?”

Governance and legitimacy questions carry unusual weight here. In [our reputation study](https://underneath.agency/research/is-it-legit-ai-reputation-study), assistants asked whether a brand was legitimate said yes in every complete answer, but 99.7% raised at least one problem. A buyer asking about an agent vendor will hear whatever concerns the public record contains.

## How does a shortlist mention become usage revenue for an agent vendor?

Through a pilot that reaches production, then usage that grows as the agent takes on more work.

From the pricing and adoption data above, we infer that most agent deals move through four steps:

1. **Shortlist.** A buyer asks an assistant for agents that handle a specific task in their stack. Three to five names come back.
2. **Pilot.** The buyer tests one or two on real data. This is where most deals are lost, because the agent must perform on the buyer’s own workflows.
3. **Production.** Security, legal and finance approve. Here the buyer checks the vendor’s public evidence again, often with an assistant.
4. **Expansion.** With per-outcome or per-action pricing, revenue grows as volumes rise and new use cases are added.

Because revenue scales after deployment, a vendor named early in a category gets more than one deal from each mention: it gets the chance to grow inside the account. The cost of being absent is the reverse. A buyer who pilots two competitors rarely adds a third, and an AI answer that leaves you out produces no visit to measure. For how lost answers show up in pipeline, see [our article on AI answers and pipeline](https://underneath.agency/resources/ai-answers-pipeline-revenue).

## Why does an assistant name one agent vendor over another that looks the same?

The assistant makers explain little; research points to independent coverage, consistent facts and visible proof of what an agent does.

**What OpenAI has published.** Its [ChatGPT search help page](https://help.openai.com/en/articles/9237897-chatgpt-search) says a question is typically rewritten into “one or more targeted queries” for search providers, sometimes followed by more specific ones. Nothing there explains how an agent vendor gets picked. We counted the searches ourselves: in [our hidden searches study](https://underneath.agency/research/ai-hidden-searches-study), ChatGPT averaged 3.7 per buyer question, enough for a single governance question to pull in an agent vendor’s reviews, pricing and security documentation together.

**Research findings.** When [Chen and colleagues](https://arxiv.org/abs/2509.08919) looked at US software questions, 72.7% of the sources AI search used were independent “earned” sites; for Google, the share was 45.4%. [Our brand entity study](https://underneath.agency/research/brand-entity-ai-recommendations-study) found the same pull toward outside sources: each tenfold increase in independent sites naming a brand was associated with 4.7 times the odds of a recommendation. Prices are a weak spot: of the software plan prices assistants quoted in [our pricing study](https://underneath.agency/research/ai-pricing-accuracy-study), only 61.9% fully matched the vendor’s page. Usage-based agent pricing, with credits and outcomes, is harder still to quote correctly.

**Trust factors specific to agents.** Security and governance come first. SailPoint found 23% of organizations reported AI agents being coaxed into revealing access credentials, and 92% said governing AI agents is paramount to enterprise security. Capgemini reports that organizations are prioritizing transparency about how agents make decisions. Gartner predicts that “guardian agents,” which monitor other agents, will account for at least 10 to 15% of agentic AI markets by 2030. We infer that agent vendors with public documentation of permissions, audit trails, human approval steps, evaluation results and certifications give both buyers and assistants something concrete to cite. For a close parallel in a security-led market, see [our cybersecurity article](https://underneath.agency/resources/cybersecurity-software-revenue-from-ai-search).

## Why are agent companies so hard for assistants to tell apart?

Because the category is crowded, new and full of look-alike claims.

Supply is exploding. In Y Combinator’s Spring 2025 batch, 67 of 144 startups were [building AI agents or tools for creating them](https://getcoai.com/news/nearly-50-of-y-combinators-spring-2025-batch-builds-ai-agents/). New companies also start with little for assistants to draw on: in a test of 112 Product Hunt startups, [Sharma](https://arxiv.org/abs/2601.00912) found ChatGPT surfaced them in only 3.32% of discovery questions, even though it recognized almost all of them by name. The causes are set out in [why ChatGPT misses new products](https://underneath.agency/resources/why-chatgpt-misses-new-products), and they apply to any agent startup launched recently.

When many vendors describe themselves with the same words (“autonomous,” “agentic,” “AI workforce”), we infer an assistant has little to separate them except what independent sources say. That is the opening for a vendor that names a narrow job, publishes results for it and earns coverage that repeats them. Comparison pages are one way to make that separation explicit; [our article on comparison pages](https://underneath.agency/resources/do-comparison-pages-help-b2b-ai-citations) reviews the evidence.

## Which GEO work helps an agent vendor stand out from look-alikes?

Making your agent’s use case, proof and safeguards findable and consistent, with no promise of a mention.

For an agent vendor, generative engine optimization (GEO) breaks into six tasks:

1. **A precise entity.** Say what the agent does, for whom, in which systems and with what level of autonomy, the same way on your site, marketplaces (Salesforce AppExchange, Microsoft, AWS), review profiles, GitHub and LinkedIn.
2. **Use-case pages with evidence.** Publish one page per job the agent does, with measured results, limits and the human steps that remain. Vague “autonomous AI workforce” pages give an assistant nothing to quote.
3. **Public trust documentation.** A trust center covering permissions, data handling, audit logs, human approval, evaluation methods and certifications answers the questions buyers and security teams ask first.
4. **Clear pricing.** State your unit (per outcome, per action, per seat) and a worked example, and keep one current page so assistants do not quote stale credits or plans.
5. **Independent proof.** Earn analyst mentions, practitioner reviews, case studies told by customers, integration partner listings and community discussion. Earning that standing step by step is covered in [building authority for AI search](https://underneath.agency/resources/how-brands-build-authority-for-ai-search).
6. **Stage-by-stage tracking.** Re-ask use-case, comparison, pricing and governance questions in ChatGPT, Gemini, Perplexity, Claude, Copilot and Google, and link what you find to incoming pilot requests.

## Where is the evidence on AI search for agent vendors still missing?

Mostly at the money end: no study ties AI visibility to pilots, production deployments or usage revenue for agent companies.

The adoption figures come from consultancies, analyst firms, investors and vendors, and they disagree: 51% of LangChain’s respondents, many from technology firms, had agents in production, while Capgemini found 2% of large organizations fully scaled. Definitions of “agent” differ between surveys, which is part of the agent washing problem. The security figures come from SailPoint, which sells identity security. Gartner’s cancellation and adoption figures are predictions, not measurements. And no published study shows how assistants weigh security documentation or certifications when naming an agent vendor; our trust-factor points above are inference.

## How can an agent vendor find out whether AI answers are costing it pilots?

Check whether assistants name you, describe your agent accurately and answer safety questions from your own evidence.

Put your use-case, comparison, pricing and governance questions to the main assistants, then set the answers beside the pilots and production deployments you won and lost, so you can see which gaps cost you shortlists and, later, per-outcome usage. We can run that comparison with you and plan the trust documentation and independent proof that would close the gaps: [talk to us about an agent visibility audit](https://underneath.agency/contact). Our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page sets out how that work runs, from mapping the governance questions buyers ask to measuring whether new evidence changes the answers.

## Frequently asked questions

### Do AI assistants recommend AI agents by name?

Yes, when asked category questions, they typically name products. How they choose is not documented by the platforms. Studies of software questions find assistants lean on independent sites, so reviews, analyst coverage and case studies matter.

### What do enterprise buyers check before buying an AI agent?

Reliability on their own data, cost and control. Surveys point to security and governance as the main brakes: SailPoint found 80% of organizations had seen unintended agent actions, and Capgemini found trust in fully autonomous agents falling.

### Should an agent company publish its security and governance details?

Yes. Buyers ask about permissions, audit logs and human approval early, and an assistant can only repeat what is public. A gated or missing trust page leaves the answer to others.

### Does “agent washing” hurt real agent companies in AI search?

It makes them harder to tell apart. When many vendors use the same words, specific, verifiable claims and independent proof are what separate a real agent from a rebranded chatbot.

## Sources

- Gartner (2025), [Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027](https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027)
- Gartner (2025), [Gartner Predicts that Guardian Agents will Capture 10-15% of the Agentic AI Market by 2030](https://www.gartner.com/en/newsroom/press-releases/2025-06-11-gartner-predicts-that-guardian-agents-will-capture-10-15-percent-of-the-agentic-ai-market-by-2030)
- Capgemini (2025), [Trust and human-AI collaboration set to define the next era of agentic AI](https://www.capgemini.com/news/press-releases/trust-and-human-ai-collaboration-set-to-define-the-next-era-of-agentic-ai-unlocking-450-billion-opportunity-by-2028/)
- SailPoint (2025), [AI agents: The new attack surface](https://nhimg.org/wp-content/uploads/2025/09/SailPoint-The-Rising-Risk-of-AI-Agents-Expanding-the-Attack-Surface-report-SP2648-.pdf)
- Menlo Ventures (2025), [2025: The State of Generative AI in the Enterprise](https://menlovc.com/perspective/2025-the-state-of-generative-ai-in-the-enterprise/)
- Microsoft (2025), [2025: The Year the Frontier Firm Is Born](https://www.microsoft.com/en-us/worklab/work-trend-index/2025-the-year-the-frontier-firm-is-born)
- LangChain (2024), [State of AI Agents](https://www.langchain.com/stateofaiagents)
- Stack Overflow (2025), [2025 Developer Survey: AI](https://survey.stackoverflow.co/2025/ai)
- G2 (2026), [2026 Buyer Behavior Report: The Evaluation Maze](https://sell.g2.com/2026-buyer-behavior-report)
- Intercom (2026), [Fin pricing](https://fin.ai/pricing)
- Salesforce (2026), [Agentforce pricing](https://www.salesforce.com/agentforce/pricing/)
- CO/AI (2025), [Nearly 50% of Y Combinator’s spring 2025 batch builds AI agents](https://getcoai.com/news/nearly-50-of-y-combinators-spring-2025-batch-builds-ai-agents/)
- OpenAI (2025), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Chen and colleagues (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919)
- Sharma (2025), [The Discovery Gap: How Product Hunt Startups Vanish in LLM Organic Discovery Queries](https://arxiv.org/abs/2601.00912)
- Underneath (2026), [hidden searches](https://underneath.agency/research/ai-hidden-searches-study), [brand entity](https://underneath.agency/research/brand-entity-ai-recommendations-study), [software pricing accuracy](https://underneath.agency/research/ai-pricing-accuracy-study) and [“is it legit” reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study) studies

---

This is the Markdown twin of https://underneath.agency/resources/ai-agent-companies-customers-from-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Do people click fewer websites when they use AI answer engines?"
description: "Yes. In a small 2024 lab study, people clicked about one source per question with AI answer engines and about four on Google. Larger studies agree."
canonical: "https://underneath.agency/resources/ai-answer-engines-fewer-website-clicks"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Do people click fewer websites when they use AI answer engines?

Yes. In a 2024 lab study of 21 highly educated participants, people using AI answer engines tended to click one source, against an average of four on Google. Larger studies of Google’s AI features and of AI assistants point the same way, so each question now sends visits to fewer websites.

## The short version

1. In a 2024 lab study, people clicked about one source per question with answer engines and an average of 4 on Google (Narayanan Venkit and colleagues).
2. They also looked at far fewer sources, hovering over an average of 2 with answer engines against 12 on Google.
3. In a 2026 US experiment with 1,100 people, sending every search to Google’s AI Mode cut clicks to outside websites by 18.8 percentage points.
4. In a 2026 panel, page visits followed 38.9% of AI assistant sessions, against 60.2% of the same people’s search sessions.
5. The one source people click may not be a Google winner: only 8.3% of ChatGPT’s citations in our study ranked in Google’s top 10.

## Do answer engine users click fewer sources than Google users?

Yes, in the one lab study that compared them directly: about one source against four.

[Narayanan Venkit and colleagues](https://arxiv.org/abs/2410.22349) asked 21 participants to research questions in their own field and debate topics. Each used one answer engine, an AI tool that writes an answer and cites sources, and Google for the same questions. The engines were YouChat, Bing Copilot and Perplexity, as they were in 2024.

| Engine | Sources clicked per question (average) |
|---|---|
| Google | 3.80 |
| Perplexity | 1.29 |
| YouChat | 1.07 |
| Bing Copilot | 0.76 |

The authors sum it up: “With answer engines, users tend to click on a single source and trust the model’s selection.” On Google, participants clicked and read an average of 4 sources.

The gap was wider for sources people merely looked at. Participants hovered over an average of 2 sources with answer engines and 12 on Google.

## How much weight should you put on that study?

Some, but not much on its own: it was small, expert-heavy and used 2024 engines. The 21 participants were mostly people with or working toward a PhD, chosen so they could judge answers in their field. Each session lasted 90 minutes and followed a think-aloud format, which can change how people behave.

Several of the authors work at Salesforce AI Research, an industry lab. The answer engines tested have changed a lot since 2024, and none was ChatGPT. Treat the one-versus-four figure as a clear direction, not a precise rate for your buyers. For ChatGPT itself, see [how people use ChatGPT alongside Google](https://underneath.agency/resources/is-chatgpt-replacing-google).

## Do larger studies show the same pattern?

Yes: field data on Google’s AI features and on AI assistants also shows fewer clicks out. The studies differ in method, but they point the same way.

| Study | Who | Finding |
|---|---|---|
| [Chapekis and colleagues](https://arxiv.org/abs/2608.04831), Pew Research Center | 900 US adults, March 2025 | A search result was clicked on 8% of Google pages with an AI Overview, against 15% without |
| [Wang and colleagues](https://arxiv.org/abs/2608.18352) | 1,100 US users, March 2026 | Forcing every search into AI Mode cut outside clicks by 18.8 points |
| [Iannelli and Ai](https://arxiv.org/abs/2607.04282) | US and British panel, February 2026 | Page visits followed 38.9% of assistant sessions, against 60.2% of search sessions |

An AI Overview is the AI summary at the top of Google’s results; AI Mode is Google’s conversational search. In the Pew panel, people clicked a source cited inside the AI Overview on only 1% of visits. In the AI Mode experiment, fewer people clicked through to news sites, Reddit and Wikipedia alike.

The assistant panel study comes from authors at Scrunch AI, a commercial company, and was not yet peer reviewed. In that panel, 34.1% of assistant sessions [had no observed web visit at all](https://underneath.agency/resources/ai-assistant-sessions-without-website-visits), against 19.5% of search sessions.

## Why do people click fewer sources with AI answers?

Mostly because they trust the engine’s pick and are shown fewer sources to choose from. In the lab study, the answer engines displayed an average of 4.31 sources per answer and cited 3.0 of them.

Agreement also cut clicking. When a question matched the participant’s own view, they clicked an average of 0.48 sources and [often accepted the answer with little checking](https://underneath.agency/resources/do-people-check-ai-sources). They checked more when the question challenged their view.

People notice the narrower choice. When Wang and colleagues asked AI Mode users for their impressions after a week, 13.4% mentioned limited links or a narrow range of sources. Another 15.3% described difficulty getting to specific websites.

## Which websites lose the clicks?

Every source except the one the AI puts in front of the reader. On Google, people in the lab study opened about four results, so several sites could win a visit from one question. With an answer engine, the visit goes to the single source the person trusts, if any.

That source is often not the one that ranks. In [our rankings study](https://underneath.agency/research/ai-citations-google-rankings-study), only 8.3% of the pages ChatGPT cited ranked in Google’s top 10 for the question. [Our comparison of Google’s AI surfaces](https://underneath.agency/research/ai-mode-vs-ai-overviews-study) found AI Mode cited 16.3% of top-10 pages, against 29.7% for AI Overviews.

And [Grossman and colleagues](https://arxiv.org/abs/2604.27790) found that on average only 18% of sources were shared between an AI Overview and Google’s ordinary results for the same search.

## What should you do about it?

Aim to be the one source the AI names, and expect fewer visits from everyone else. Practical steps:

1. Check who gets cited. Run your buyers’ key questions through ChatGPT, Perplexity, Google AI Mode and others, and note which sources appear.
2. Be the obvious primary source. Original data, clear definitions and plainly stated facts give an AI engine a reason to cite you rather than a summary of you.
3. Earn mentions on cited sites. If the engines favor review sites, publications or forums in your category, your presence there matters as much as your own pages.
4. Value mentions, not only clicks. A citation that wins no click can still shape a buyer’s shortlist.
5. Check what is said about you. In the lab study, around 30% of statements from You.com and Perplexity were not supported by their listed sources, so errors about your brand are possible.

If you want help becoming the source AI answers cite, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The direct comparison rests on one small 2024 study, so much is still open. The main gaps:

- Scale. No large study has counted per-question clicks for ChatGPT, Perplexity or Gemini users against Google users.
- Buyers. No study covers business buyers or high-value purchases, where people may check more sources.
- Current tools. The engines in the lab study have changed, and ChatGPT was not among them.
- Click value. Nobody has tested whether the one click an answer engine sends is worth more than a typical search click.

## Frequently asked questions

### Do people click the links in AI answers?

Less often than on Google. In a 2026 panel, page visits followed 38.9% of AI assistant sessions against 60.2% of search sessions by the same people.

### How many sources do people check when using AI search?

About one per question, in a 2024 lab study, against four on Google. On Google’s own AI Overviews, only 1% of visits led to a click on a cited source.

### Does ranking on Google still matter for AI answers?

Yes, but it does not guarantee a citation. AI Overviews in our study cited 29.7% of top-10 pages, and ChatGPT’s citations ranked in Google’s top 10 only 8.3% of the time.

### Why do people trust AI answers without clicking?

They tend to accept the engine’s choice of source. In one lab study, people asking questions that matched their own view clicked an average of 0.48 sources.

## Sources

- Narayanan Venkit, P., Laban, P., Zhou, Y., Mao, Y. and Wu, C.-S. (2024), [Search Engines in an AI Era: The False Promise of Factual and Verifiable Source-Cited Responses](https://arxiv.org/abs/2410.22349), arXiv:2410.22349.
- Chapekis, A., Lieb, A., Shah, S. and Smith, A. (2026), [Investigating Click Behaviors On Google Search Result Pages That Produce an AI Overview](https://arxiv.org/abs/2608.04831), arXiv:2608.04831.
- Wang, S. T., Gleason, J., Bart, Y., Wilson, C. and Metaxa, D. (2026), [AI in Search Reduces Publisher Referrals Without Improving User Experience: Experimental Evidence](https://arxiv.org/abs/2608.18352), arXiv:2608.18352.
- Iannelli, M. and Ai, A. (2026), [The New Shape of Search: How Conversational AI Recomposes Information Seeking](https://arxiv.org/abs/2607.04282), arXiv:2607.04282.
- Grossman, R., Liu, S., Chen, M. K., Smith, M., Borcea, C. and Chen, Y. (2026), [How Generative AI Disrupts Search: An Empirical Study of Google Search, Gemini, and AI Overviews](https://arxiv.org/abs/2604.27790), arXiv:2604.27790.
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude cite pages that rank?](https://underneath.agency/research/ai-citations-google-rankings-study)
- Underneath (2026), [AI Mode vs AI Overviews: how different are the sources?](https://underneath.agency/research/ai-mode-vs-ai-overviews-study)

---

This is the Markdown twin of https://underneath.agency/resources/ai-answer-engines-fewer-website-clicks. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "What do lost clicks to AI answers mean for pipeline and revenue?"
description: "Fewer search clicks mean fewer prospects entering your funnel, and AI answers appear most on the research and comparison searches where buying starts."
canonical: "https://underneath.agency/resources/ai-answers-pipeline-revenue"
published: 2026-10-08
updated: 2026-10-10
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# What do lost clicks to AI answers mean for pipeline and revenue?

They mean fewer prospects at the top of your funnel, and AI answers show up most on the research-stage searches where buying starts. Studies of Google’s AI features find drops from about 5% of Wikipedia’s search visits to nearly half of result clicks on pages with an AI summary. No study has yet traced those lost clicks through to a company’s leads, pipeline or revenue.

## The short version

1. Researchers estimate Google’s AI Overviews cut search visits to English Wikipedia by 5.45%, about 100.27 million visits a month (Khosravi and Yoganarasimhan, 2026).
2. In a Pew Research Center panel of 900 US adults, people clicked a result on 8% of pages with an AI Overview and 15% without.
3. In a 2025 test, Google showed AI Overviews on 92% of product question searches but only 17.4% of plain product searches, so research-stage traffic is most exposed.
4. In a 2026 US experiment with 1,100 people, making AI Mode the default cut clicks to outside websites by 18.8 percentage points.
5. Not all content loses: Reddit discussion communities gained 12.0% more comments after AI Overviews launched, though AI Mode later erased much of the gain.

## How do lost search clicks reach your pipeline?

Every search visit you lose is one fewer person who can enter your funnel.

That is the plain logic set out by [Khosravi and Yoganarasimhan](https://arxiv.org/abs/2602.18455), who measured what Google’s AI Overviews did to Wikipedia. Many sites, they note, depend on search visits for advertising, subscriptions, lead generation and commerce. For them, “a decline in search traffic reduces both immediate monetization opportunities and the number of users entering the top of the conversion funnel.”

Losses can compound: on Wikipedia, roughly 40% of all traffic comes from clicks between articles. If a search visit never happens, the pages that visitor would have read next are lost too. The authors say the total effect could therefore be larger than they measured. A related question is [whether AI Overviews end browsing sessions](https://underneath.agency/resources/ai-overviews-end-browsing-sessions) altogether.

For a business site, the equivalent is the visitor who lands on a guide and then reads a product, pricing or demo page. No study has measured that chain for company websites.

## How large are the click losses so far?

Between a few percent and about half, depending on the site, the measure and the AI feature.

| Study | Who and when | What they found |
|---|---|---|
| Khosravi and Yoganarasimhan | English Wikipedia, May to December 2024 | Search visits fell 5.45% against German editions and 4.82% against French |
| Chapekis and colleagues, Pew Research Center | 900 US adults, March 2025 | A result was clicked on 8% of AI Overview pages against 15% of others |
| Wang and colleagues | 1,100 US users, March 2026 | Hiding AI Overviews raised outside clicks by 8.8 points; forcing AI Mode cut them by 18.8 |

The Wikipedia figure is an average over all searches, including many that showed no AI Overview. The authors say the effect on searches that did show one is likely larger. Projected over a year, their estimate comes to about 1.20 billion fewer search visits. We weigh all the evidence in [how much AI Overviews cut clicks](https://underneath.agency/resources/do-ai-overviews-reduce-clicks).

The [Pew study](https://arxiv.org/abs/2608.04831) is observational, so it shows an association, not proof of cause. The [field experiment by Wang and colleagues](https://arxiv.org/abs/2608.18352) is the strongest causal evidence. It ran for one week, though, and its sample skewed young: 75% were under 45.

## Which searches put your pipeline most at risk?

Research-stage searches: comparisons and questions draw AI answers far more often than plain product searches.

[Grossman and colleagues](https://arxiv.org/abs/2604.27790) ran product searches through Google in December 2025. An AI Overview, the AI summary at the top of Google’s results, appeared on 17.4% of plain product searches. It appeared on 88.2% of product comparisons and 92% of product questions, which were written by AI from real searches.

The authors conclude that AI search matters more in the “consideration or research stages than the final purchase stage.”

[Our own study](https://underneath.agency/research/ai-overviews-frequency-study) of US commercial keywords found the same tilt. 60.8% of 800 commercial keywords showed an AI Overview, rising to 98.1% of keywords phrased as questions. Searches that showed a map of local businesses were the exception: their predicted chance was 30.7%, against 78.6% for searches without one.

For a business, the risky searches are the early ones, such as “how do I choose a…” or “X vs Y”. Those are often where a buyer first met your brand.

## Can you put a dollar figure on the loss?

Not reliably yet: the published dollar figures are illustrations, not measured losses.

Khosravi and Yoganarasimhan priced their Wikipedia estimate as if every lost visit were a page carrying ads. That gave $0.901 million to $3.090 million a month. Wikipedia carries no ads, so this shows scale, not money anyone lost.

[Xu and colleagues](https://arxiv.org/abs/2605.14021) looked at the pages AI Overviews cite, across 55,393 trending searches in March and April 2026. 50.63% of the cited pages displayed ads. They add that sites earning through subscriptions or direct conversions lose in the same way when the answer arrives without a click.

To estimate your own exposure, multiply the search visits at risk by your visit-to-lead rate and your lead value. Use your own numbers; no study supplies them.

## Does every kind of content lose clicks?

No: discussion and advice content gained engagement from AI Overviews, at least until AI Mode arrived. [Zhang, Cui and Zhang](https://arxiv.org/abs/2605.16428) tracked 105,012 Reddit communities through Google’s rollout. Communities that AI Overviews could cite saw daily comments rise by 12.0% and commenters by 12.4%, compared with communities Google bars from AI Overviews.

The gain was largest for experience content, such as opinions, advice and personal stories. There the effect on comments was 2.3 times larger than for factual communities. The authors’ explanation: a summary can replace a fact, but it points people toward experiences they want to read in full.

That advantage faded once Google launched AI Mode, its conversational search. For experience communities, the extra gain in commenters fell by 59%, and for comments it disappeared. These are Reddit comments, not visits or sales, so treat this as a pattern, not a forecast for your site.

## What if AI Mode becomes Google’s default?

Then the losses would likely grow: forcing AI Mode cut outside clicks by 18.8 points in a one-week experiment. In that experiment, the share of people clicking through fell for news sites, Reddit and Wikipedia alike. People also ran 0.92 fewer search sessions a day, and 11.2 percentage points more of them tried a rival search engine.

In [our comparison of Google’s two AI surfaces](https://underneath.agency/research/ai-mode-vs-ai-overviews-study), AI Mode answered all 400 searches we tested. We look at the traffic risk in more detail in [our article on AI Mode as Google’s default](https://underneath.agency/resources/ai-mode-default-traffic-loss).

## What should you do about it?

Re-plan your funnel for fewer research-stage visits, and work to be named inside the AI answer itself. Practical steps:

1. Find your exposed traffic. List the question and comparison searches that feed your pipeline, and check which already show an AI Overview.
2. Size the risk in your own numbers. Lost visits times conversion rate times lead value gives a range to plan around.
3. Aim to be cited, not just ranked. In [our citation study](https://underneath.agency/research/ai-overview-citations-study), 28.7% of AI Overview citations were page-one results, so ranking alone does not get you into the answer.
4. Invest in what a summary cannot replace. First-hand experience, advice and opinion drew more engagement in the Reddit study than plain facts did.
5. Track influence beyond clicks. Branded searches, direct visits and what new leads say about how they found you all capture effects that referral data misses.

If you want help working out where AI answers touch your pipeline, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

Nobody has yet measured the effect on business pipeline or revenue directly. The main gaps:

- Pipeline effects. No study follows lost AI clicks through to leads, deals or revenue for B2B or consumer companies.
- Dollar values. The only dollar figures are illustrations built on assumed ad rates.
- Click quality. Google argues that AI features improve the quality of clicks that remain, but the studies here did not test that.
- AI assistants. Most evidence covers Google’s AI features; how ChatGPT and other assistants change visits to business sites is barely measured.
- Time. The studies cover short windows between 2024 and 2026, and these products change monthly.

## Frequently asked questions

### Do AI Overviews reduce website traffic?

Yes, on the best evidence. In a 2026 field experiment with 1,100 US users, hiding AI Overviews raised clicks to outside websites by 8.8 percentage points.

### How much revenue do websites lose to AI search?

No study has measured it. The closest is an illustration for Wikipedia of $0.901 million to $3.090 million a month in ad value, for a site that carries no ads.

### Which marketing content is most at risk from AI answers?

Factual answers to research questions. Google showed AI Overviews on 88.2% of product comparison searches in one test, while experience-based discussion content gained engagement.

### Are the clicks that remain worth more?

Possibly, but it is unproven. Google says AI features improve click quality, and none of the studies reviewed here has tested that claim.

## Sources

- Khosravi, M. and Yoganarasimhan, H. (2026), [Impact of AI Search Summaries on Website Traffic: Evidence from Google AI Overviews and Wikipedia](https://arxiv.org/abs/2602.18455), arXiv:2602.18455.
- Chapekis, A., Lieb, A., Shah, S. and Smith, A. (2026), [Investigating Click Behaviors On Google Search Result Pages That Produce an AI Overview](https://arxiv.org/abs/2608.04831), arXiv:2608.04831.
- Wang, S. T., Gleason, J., Bart, Y., Wilson, C. and Metaxa, D. (2026), [AI in Search Reduces Publisher Referrals Without Improving User Experience: Experimental Evidence](https://arxiv.org/abs/2608.18352), arXiv:2608.18352.
- Grossman, R., Liu, S., Chen, M. K., Smith, M., Borcea, C. and Chen, Y. (2026), [How Generative AI Disrupts Search: An Empirical Study of Google Search, Gemini, and AI Overviews](https://arxiv.org/abs/2604.27790), arXiv:2604.27790.
- Xu, H., Iqbal, U. and Montgomery, J. M. (2026), [Measuring Google AI Overviews: Activation, Source Quality, Claim Fidelity, and Publisher Impact](https://arxiv.org/abs/2605.14021), arXiv:2605.14021.
- Zhang, P., Cui, R. and Zhang, D. J. (2026), [The Impact of AI Search on the Online Content Ecosystem: Evidence from Google and Reddit](https://arxiv.org/abs/2605.16428), arXiv:2605.16428.
- Underneath (2026), [When does Google show an AI Overview? 1,248 US searches](https://underneath.agency/research/ai-overviews-frequency-study)
- Underneath (2026), [AI Mode vs AI Overviews: how different are the sources?](https://underneath.agency/research/ai-mode-vs-ai-overviews-study)
- Underneath (2026), [Do AI Overviews cite the pages that rank? 4,051 citations](https://underneath.agency/research/ai-overview-citations-study)

---

This is the Markdown twin of https://underneath.agency/resources/ai-answers-pipeline-revenue. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How often do AI answers say things their sources do not support?"
description: "Often: audits find roughly one in nine to one in three statements in AI search answers are not backed by the sources cited beside them."
canonical: "https://underneath.agency/resources/ai-answers-unsupported-claims"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How often do AI answers say things their sources do not support?

Often enough to matter: independent audits find that somewhere between one in nine and one in three statements in AI search answers are not backed by the sources listed beside them. The rate depends on the engine, the topic and how strictly “support” is judged. For a brand, it means an answer can describe you inaccurately while still looking well sourced.

## The short version

1. In a 2024 audit of 303 questions, 30.8% of You.com’s statements, 31.6% of Perplexity’s and 23.1% of BingChat’s were not supported by any source the engine listed ([Narayanan Venkit and colleagues](https://arxiv.org/abs/2410.22349)).
2. In Google’s AI Overviews, 11.0% of 98,020 claims checked over 40 days of US searches were not supported by the pages they cited ([Xu and colleagues](https://arxiv.org/abs/2605.14021)).
3. The main failure is omission, not contradiction: an unsupported claim was about 2.6 times more likely to be missing from the cited pages than contradicted by them (Xu and colleagues).
4. In our own check of AI answers about whether brands are legit, a readable cited page supported the claim fully or in part 72.4% of the time ([our reputation study](https://underneath.agency/research/is-it-legit-ai-reputation-study)).

## How often are AI answers unsupported by their own sources?

In the best-known audit, roughly a quarter to a third of statements had no support in the listed sources. [Narayanan Venkit and colleagues](https://arxiv.org/abs/2410.22349), from Penn State and Salesforce AI Research, ran 303 questions through You.com, BingChat and Perplexity. They split each answer into single statements and had an automated judge check each one against the full text of every listed source.

| Engine (August 2024) | Statements not supported by any listed source | Citations that pointed to a supporting source |
|---|---|---|
| You.com | 30.8% | 68.3% |
| BingChat | 23.1% | 65.8% |
| Perplexity | 31.6% | 49.0% |

Two caveats matter. The automated judge agreed only moderately with human checkers, and these are 2024 versions of products that have changed since. Even so, the authors concluded that none of the three engines reached acceptable performance on most of their measures.

## Are Google’s AI Overviews more faithful?

Yes, by a clear margin, but about one claim in nine was still unsupported. [Xu and colleagues](https://arxiv.org/abs/2605.14021) at Washington University in St. Louis ran 55,393 trending US searches between March 13 and April 21, 2026. They checked every claim in the AI Overviews that appeared (an AI Overview is the AI summary at the top of Google’s results).

Of 98,020 claims, 11.0% were not supported by the pages the Overview cited. Only 4.1% were contradicted by or in conflict with their cited content; a further 7.0% were simply not addressed by any cited page. Most Overviews were largely sound: 41.9% had every claim supported, and only 2.74% had fewer than half supported.

The authors call 11% a ceiling, not a precise rate. Their checker could not read social media and video pages, and if every claim tied to those pages were in fact supported, the rate would fall to about 5.3%. Accuracy also varied by topic, from 94.77% of claims consistent in health to 76.85% in jobs and education, partly because fast-changing facts such as school closings had moved on before pages were checked.

## What do these errors actually look like?

Mostly missing support and misattribution, rather than facts that contradict the source. In the AI Overview data, an unsupported claim was about 2.6 times more likely to be something no cited page mentions than something a cited page contradicts. The answer states a fact, and [the citation beside it does not contain it](https://underneath.agency/resources/ai-citing-pages-for-claims-they-do-not-make).

Misattribution is the second pattern. Venkit’s team found that even when one listed source did support a statement, engines often cited a different one. All 21 experts in their user study flagged answers that cited sources which did not say what the answer claimed, and one called the citations a way for the system “to legitimize itself.”

The citation list can also look broader than the answer really is. [Huang and colleagues](https://arxiv.org/abs/2603.16138) compared 11,000 real search questions across several systems. Google’s AI Overviews cited Reddit and Quora on some questions, yet [drew 22.1% less content from social platforms](https://underneath.agency/resources/do-ai-overviews-downplay-negative-content) than from other sources they cited.

Answers also tend to sound surer than their sources. In the same study, adding web search to an AI assistant [cut hedging language by up to 60%](https://underneath.agency/resources/do-ai-search-engines-agree-with-leading-questions) while keeping confident wording.

## Does citing better sources fix the problem?

No: in the largest audit, how credible a source was had little bearing on whether claims matched it. Xu and colleagues found AI Overviews cited domains that were, on average, more credible than Google’s own first-page results. Yet source quality and claim accuracy were “largely independent,” so the authors argue the problem cannot be fixed by better sourcing alone.

More sources do not guarantee more support either. In the 2024 audit, BingChat listed the most sources, but 36% of the sources it listed [were never cited in the answer text](https://underneath.agency/resources/does-ai-search-use-the-pages-it-cites). The authors warn this can create “a false sense of factual backing from multiple sources.”

## What does this mean for how AI describes a brand?

The same gap shows up in brand questions, and much of it starts with the web’s own inconsistencies. In [our study of “Is this brand legit?” answers](https://underneath.agency/research/is-it-legit-ai-reputation-study) from four assistants about 79 brands, we checked a random sample of 240 cited claims.

Where the cited page could be read, it supported the claim fully or in part 72.4% of the time. Another 25.9% were not supported and 1.7% were contradicted. Many review sites block automated reading, and all coding was done by AI models, so treat these as rough bounds.

The clearest gap concerned how common a problem is. 71.4% of claims saying a complaint was common rested on evidence showing individual reports, or nothing about frequency.

Prices show where many wrong facts come from. In [our pricing study](https://underneath.agency/research/ai-pricing-accuracy-study), 64 quoted software prices differed from the official pricing page. For 39 of them, the same figure was still published on another page of the vendor’s own site.

## What should you do about it?

Treat being cited and being described correctly as two separate goals, and check the second one directly.

1. Ask the main AI assistants the questions your buyers ask about you, and read each claim, not just whether your site is cited.
2. Retire or update old pricing, plan and policy pages. Assistants can repeat any figure you still publish, even on a forgotten help page.
3. State key facts in plain, specific sentences on pages that load without logins or scripts. Since omission is the main failure, a fact that is hard to find invites the answer to fill the gap from somewhere else.
4. Watch the review and forum sites that assistants cite about you, because that is where many negative claims originate.
5. Re-check on a schedule. Every audit here is a snapshot, and the engines change often.

If you want help running these checks across engines, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The research measures unsupported claims well, but not how often they are false or how they affect buyers.

- Unsupported is not the same as wrong. Many unsupported claims may be true but drawn from a page the engine did not cite.
- Most audits rely on automated judges. Agreement with human checkers ranged from moderate in the 2024 audit to high in the AI Overview study.
- The big audits cover trending searches, debates and expert questions. Very little independent research looks at product or brand questions specifically.
- Results date quickly. The 2024 audit predates current versions of every engine it tested, and no one has rerun the same test across today’s assistants.
- No study we found measures how an unsupported claim about a company changes what buyers think or do.

## Frequently asked questions

### Are AI search answers accurate?

Mostly, but not reliably. In Google’s AI Overviews, 88.97% of 98,020 checked claims were consistent with the cited pages, but two of three chat-style engines in a 2024 audit left around 30% of their statements unsupported.

### Which AI search engine is most faithful to its sources?

No study ranks today’s engines on the same test. In the 2024 audit, BingChat had the fewest unsupported statements (23.1%) and Perplexity the most (31.6%), but all three products have changed since.

### Can a citation be wrong even when the fact is right?

Yes. Experts in the 2024 user study found true statements attached to sources that never mentioned them, and Perplexity’s citations pointed to a supporting source only 49.0% of the time.

### Why would an AI assistant state something about my company that is not on my website?

Usually because another page says it. In our pricing study, 21 of the 64 prices that differed from the official page also appeared on a third-party page cited for the product.

## Sources

- Narayanan Venkit, Laban, Zhou, Mao and Wu (2024), [Search Engines in an AI Era: The False Promise of Factual and Verifiable Source-Cited Responses](https://arxiv.org/abs/2410.22349), arXiv:2410.22349.
- Xu, Iqbal and Montgomery (2026), [Measuring Google AI Overviews: Activation, Source Quality, Claim Fidelity, and Publisher Impact](https://arxiv.org/abs/2605.14021), arXiv:2605.14021.
- Huang, Goyal, Saha and Chandrasekharan (2026), [Answer Bubbles: Information Exposure in AI-Mediated Search](https://arxiv.org/abs/2603.16138), arXiv:2603.16138.
- Underneath (2026), [“Is this brand legit?” How AI assistants build a reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study)
- Underneath (2026), [How faithfully do AI assistants quote software prices?](https://underneath.agency/research/ai-pricing-accuracy-study)

---

This is the Markdown twin of https://underneath.agency/resources/ai-answers-unsupported-claims. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How often do AI assistant users never visit a website?"
description: "About a third of the time: in a 2026 panel, 34.1% of AI assistant sessions had no observed web visit, against 19.5% of the same people’s search sessions."
canonical: "https://underneath.agency/resources/ai-assistant-sessions-without-website-visits"
published: 2026-10-08
updated: 2026-10-10
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How often do AI assistant users never visit a website?

About a third of the time, according to the one large study that has measured it. In a 2026 US and British panel, 34.1% of AI assistant sessions showed no observed web visit, against 19.5% of the same people’s search sessions. The share swings widely with how a session is counted, and a session with no visit is not proof that the question was answered.

## The short version

1. 34.1% of AI assistant sessions had no observed web step in a February 2026 panel of US and British users (Iannelli and Ai, 2026).
2. For the same people’s search sessions the figure was 19.5%, a gap of 13.0 percentage points that held at 13.6 in March 2026.
3. The share depends heavily on where one session ends and the next begins: it ran from 46.9% to 29.2% across the researchers’ three choices.
4. It also varied by market and by assistant: 50.9% in the United States against 20.7% in Great Britain, and 48.8% for Gemini against 27.3% for ChatGPT.
5. No visit does not mean no web: in our study, every assistant answer that ran a search cited at least one page.

## How many AI assistant sessions involve no website visit?

About a third, in the one large study that has measured it.

[Iannelli and Ai](https://arxiv.org/abs/2607.04282) linked the prompts and responses of an opt-in panel to the same people’s searches and page visits in February 2026. The assistants were standalone ones, mainly ChatGPT and Gemini with a small share of Perplexity. A session ended after 30 minutes with no activity.

Counting each person equally, 34.1% of sessions with an assistant had no observed outside web step: no search and no page visit. Counting every session equally, the share was 37.8%.

The authors work at Scrunch AI, a commercial company, so this is vendor research. The paper was under review and not yet peer reviewed.

| Where outside web activity fell | Share of assistant sessions |
|---|---|
| Nowhere: none observed | 34.1% |
| Only before the assistant | 18.3% |
| Only after the assistant | 10.5% |
| On both sides, or between prompts | 37.1% |

Only about one session in ten matched the familiar picture of [asking the AI first and then clicking out](https://underneath.agency/resources/chatgpt-vs-google-customer-journey).

## Is that more than with ordinary Google search?

Yes: only 19.5% of the same people’s search sessions had no visit to a page beyond the search engine. The researchers built search sessions the same way and checked whether any non-search page was visited. Comparing each person’s two kinds of sessions, assistant sessions stayed self-contained 13.0 percentage points more often. We compare clicks per question in [our article on answer engines and website clicks](https://underneath.agency/resources/ai-answer-engines-fewer-website-clicks).

The finding repeated a month later. In March 2026, 35.3% of assistant sessions and 20.4% of search sessions had no visit, a gap of 13.6 points.

Some of those searches may have been answered by Google’s own AI summary on the results page. The authors expect that to shrink the gap, not widen it.

Google is not immune to sessions that stop on its own page. In a [Pew Research Center panel](https://arxiv.org/abs/2608.04831) of 900 US adults in March 2025, people ended their browsing on 26% of Google pages showing an AI Overview. That is the AI summary at the top of results; the figure was 16% for pages without one.

## Why does the number move so much?

Because it depends on how a session is defined and how much browsing the panel captured.

The researchers tried three gaps of inactivity to separate one session from the next. With a 15-minute gap, 46.9% of assistant sessions had no web step; with a 60-minute gap, 29.2% did. The shares with web activity only before, or only after, the assistant barely moved.

Coverage matters too. Among people with at least one day of recorded browsing, 37.7% of sessions had no web step. Among people with browsing on all 28 days of the month, it was 26.7%.

Assistant use in mobile apps and through APIs was not observed at all. So the honest range is roughly a quarter to nearly half of assistant sessions. The point the authors stand behind is the ordering: assistant sessions stay self-contained more often than search sessions, under every definition they tried.

## Does it differ by country or by assistant?

Sharply: US sessions stayed inside the assistant far more often than British ones. Looking at sessions that used a single assistant, 50.9% were self-contained in the United States against 20.7% in Great Britain. By assistant, the share was 48.8% for Gemini and 27.3% for ChatGPT.

In every split, assistant sessions were still more self-contained than search sessions, which stood at 21.6% and 17.9% in the two markets. The authors conclude that the levels “are clearly population- and provider-dependent,” while the ordering holds.

## Does a session with no visit mean the question was answered?

No: the data shows that no visit happened, not that the need was met. The authors put it plainly: “Assistant-contained also does not mean resolved.” Their data held timestamps only, never the text of prompts or answers.

The self-contained sessions break down into three patterns, as shares of all assistant sessions:

| Pattern | Share |
|---|---|
| One prompt and one response | 19.0% |
| Several exchanges in one conversation | 7.4% |
| More than one conversation | 7.7% |

Many people went back to the web later. Of the self-contained sessions, 12.0% were followed by a web visit within one hour and 49.5% within six hours. The authors warn that those later visits may have nothing to do with the earlier question.

## What does this mean for measuring AI’s influence on your brand?

It means referral clicks from AI assistants will undercount their influence. A user who never leaves the assistant can still be reading about you, because the assistant searches the web on their behalf.

In [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study), ChatGPT ran a mean of 3.7 searches per buyer question before answering. Every answer that ran a search cited at least one page, and none of the 42 answers without a search cited anything. Your page can shape an answer that sends you no visitor.

Google’s conversational AI Mode shows a similar trade. In a [2026 field experiment](https://arxiv.org/abs/2608.18352) with 1,100 US users, sending every search to AI Mode cut clicks to outside websites by 18.8 percentage points. Yet it raised the time people spent per search session by 0.43 minutes.

We cover that experiment in [our article on AI Mode and traffic](https://underneath.agency/resources/ai-mode-default-traffic-loss).

## What should you do about it?

Measure AI visibility directly, because traffic reports will miss much of it. Practical steps:

1. Ask the assistants. Run the questions your buyers ask through ChatGPT, Gemini and others, and record whether you are named and what is said.
2. Do not judge AI by referral traffic alone. A low click count from an assistant can sit alongside frequent mentions.
3. Watch indirect signals. Branded searches, direct visits and “how did you hear about us?” answers can pick up influence that never arrives as a click.
4. Make your key facts easy to quote. If the answer is all the buyer reads, prices, features and proof points need to be stated plainly on pages assistants can find.

If you want help tracking how assistants describe your brand, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

Much remains open, because one panel study carries most of the evidence. The main gaps:

- Breadth. The figures come from one AI-forward opt-in panel in two countries over two months, February and March 2026.
- Mobile apps. Assistant use in native mobile apps and through APIs was not observed, so quick phone lookups are missing.
- Outcomes. With no prompt or answer text, nobody can say whether self-contained sessions ended satisfied, confused or interrupted.
- Business buyers. No study has measured self-contained sessions for B2B purchase research.
- Cause. The design cannot tell whether assistants reduce website visits overall or simply attract tasks that need fewer visits. For the wider question, see [whether ChatGPT is taking Google’s place](https://underneath.agency/resources/is-chatgpt-replacing-google).

## Frequently asked questions

### What share of ChatGPT sessions end without a website visit?

In one 2026 panel, 27.3% of single-assistant ChatGPT sessions had no observed web step, counting each person equally. Gemini’s figure was 48.8%, so the level depends on the assistant.

### Do AI assistants send traffic to websites?

Sometimes, but less than search does. In the same panel, web activity came only after the assistant in 10.5% of sessions, and on both sides or between prompts in 37.1%.

### Are zero-click sessions new with AI?

No. In the same panel, 19.5% of ordinary search sessions also ended without a page visit, though assistant sessions did so more often.

### How do I measure AI influence if users never click?

Check the answers themselves. Assistants search the web before answering, ChatGPT 3.7 times per buyer question in our study, so they can rely on your pages without sending a visit.

## Sources

- Iannelli, M. and Ai, A. (2026), [The New Shape of Search: How Conversational AI Recomposes Information Seeking](https://arxiv.org/abs/2607.04282), arXiv:2607.04282.
- Chapekis, A., Lieb, A., Shah, S. and Smith, A. (2026), [Investigating Click Behaviors On Google Search Result Pages That Produce an AI Overview](https://arxiv.org/abs/2608.04831), arXiv:2608.04831.
- Wang, S. T., Gleason, J., Bart, Y., Wilson, C. and Metaxa, D. (2026), [AI in Search Reduces Publisher Referrals Without Improving User Experience: Experimental Evidence](https://arxiv.org/abs/2608.18352), arXiv:2608.18352.
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [How much traffic would AI Mode as Google’s default cost you?](https://underneath.agency/resources/ai-mode-default-traffic-loss)

---

This is the Markdown twin of https://underneath.agency/resources/ai-assistant-sessions-without-website-visits. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "What Server Logs Reveal About AI Bot Content Demand"
description: "Partly. Logs show which pages AI bots request, including ones you lack, but not whether any answer used or cited them. One site acted on that signal."
canonical: "https://underneath.agency/resources/ai-bot-server-logs-content-demand"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Can our server logs show what content AI bots are looking for on our site?

Partly: your server logs record every page AI bots request, including addresses that do not exist, and those failed requests can hint at content people are asking AI assistants about. What logs cannot show is whether a page an AI bot read was used in an answer, cited, or seen by anyone. One published case used AI bot 404 errors to decide which pages to create, but it changed several things at once, so the tactic’s own effect is unknown.

## The short version

1. One website treated AI bot requests for missing pages as demand signals and built new pages from the most frequent patterns; the study could not separate that tactic from three others applied at the same time.
2. On that site, ChatGPT referrals to changed pages grew 6.1 times in four months, but untouched pages grew 3.5 times, and the best estimate of the real effect, about 1.8 to 2.3 times, was only suggestive.
3. A bot visit is not a citation: when ChatGPT’s developer interface searched for shopping answers it returned 37.38 pages on average but cited 3.22.
4. Bots make 62.4% of requests for web pages on one major network, so AI demand signals sit inside a lot of other bot traffic.
5. Robots.txt rules do not map neatly onto citations: Perplexity cited 209 of 517 pages (40.4%) that sites had closed to its crawler.

## What can AI bot requests in our logs tell us?

They show which pages AI systems ask for, how often, and which addresses they request that do not exist.

The volume is large and mixed. A 2026 study by [Finder and colleagues at ora](https://arxiv.org/abs/2609.34951), an agent-readiness vendor, cites Cloudflare data showing bots made 62.4% of requests for web pages on its network in September 2026. Much of that is old search machinery. The fastest-growing part is AI agents that fetch a page because a person has just asked a question.

Those bots come in different kinds. [Our crawler study](https://underneath.agency/research/ai-crawler-blocking-study) separates training crawlers such as GPTBot, search crawlers such as OAI-SearchBot, and fetchers such as ChatGPT-User that load a page when someone asks. Sites treat them differently: 15.2% of top sites block GPTBot, but only 7.4% block OAI-SearchBot. Reading your logs by bot type tells you whether a request feeds a future model, a search index or a live answer.

## Has anyone used AI bot errors to decide what pages to create?

Yes, one published case did, as part of a wider set of changes on a single website.

[Watanabe and Nakayashiki](https://arxiv.org/abs/2606.04362) work at Glasp, a site with hundreds of thousands of question-and-answer pages about YouTube videos. In January 2026 they treated AI bot requests for missing pages as demand: “the highest-frequency patterns were used to generate new pages.” They also merged duplicate addresses, rewrote titles as questions, and ranked pages for rewriting by how often AI bots requested them.

A fourth rule protected search traffic. Pages already earning Google clicks were locked against rewriting. The authors are explicit that the four changes were made together, so they estimate “its combined effect, not the marginal effect of any single tactic.”

## Did acting on those log signals increase AI traffic?

Possibly, but the gain was far smaller than the raw numbers suggest and did not pass the strictest test.

The raw growth was striking. ChatGPT referrals to the changed pages grew 6.1 times between January and May 2026. But pages the team never touched grew 3.5 times over the same months, [because ChatGPT itself was growing](https://underneath.agency/resources/chatgpt-referral-growth-and-geo). Comparing the two groups week by week, the authors estimate the changes lifted referrals by about 1.8 to 2.3 times.

Even that estimate is labelled suggestive. When they checked whether similar jumps happened by chance before the changes, the chance came out at 0.16, too high to rule out luck. Google clicks to the changed pages fell about 25%, close to the 20% fall across the whole site, so [the rewrites did not obviously hurt search](https://underneath.agency/resources/chatgpt-optimization-google-rankings).

| What the Glasp data shows | Figure |
|---|---|
| Growth in ChatGPT referrals, changed pages | 6.1 times |
| Growth in ChatGPT referrals, untouched pages | 3.5 times |
| Estimated effect of all four changes together | about 1.8 to 2.3 times |
| Effect of the 404-based pages alone | not measured |

## What can’t server logs tell us?

They show what a bot read, not whether an answer used it, cited it or reached a buyer.

The ora team, who logged 37,927 AI agent journeys across 1,056 businesses, put it simply: “Logs show what the agent read, not what the answer used.” [Uberti-Bona Marin and colleagues](https://arxiv.org/abs/2609.18729) measured the gap for ChatGPT’s developer interface on shopping questions. Its searches returned 37.38 pages per answer, but only 3.22 appeared as citations. Just 27.5% of the websites it found made it into the final answer.

The request itself is also one step removed from the buyer. In [our hidden searches study](https://underneath.agency/research/ai-hidden-searches-study), ChatGPT ran 3.7 searches per answer on average, and none of 509 searches repeated the user’s question word for word. A bot’s request reflects the assistant’s own rewritten search, not the buyer’s words.

Logs can also miss citations entirely. On the day of our crawler study, Perplexity cited 209 of 517 pages (40.4%) on top sites that had closed those pages to its crawler. So the crawler named in your rules is not the only route by which a page reaches an answer.

## Do agent-facing files like llms.txt show up as demand?

Rarely, by the evidence so far, so a quiet llms.txt in your logs is normal rather than a warning sign.

[Our llms.txt study](https://underneath.agency/research/llms-txt-adoption-study) reports that Ahrefs found 97% of the llms.txt files it monitors received no traffic in May 2026, and AI retrieval bots made 1.1% of the requests that did arrive. The ora authors report that only one in three agents that read an llms.txt went on to fetch a page it listed. Few sites offer agents a lighter format either: 3.2% of top sites return Markdown when an agent asks for it, in [our agent-readable web study](https://underneath.agency/research/agent-readable-web-study).

## What should you do about it?

Use AI bot logs as a cheap source of content ideas, and test any pages you build against a control.

1. Separate AI bot traffic by type: training crawlers, search crawlers and fetchers that act for a person.
2. Pull the missing addresses AI bots request most often and group them into topics. Treat them as hypotheses about buyer questions, not proof of demand.
3. Check those topics against what assistants actually answer for your buyers’ questions before building pages.
4. When you add or rewrite pages, leave a similar group untouched, as Glasp did, so platform growth is not mistaken for your effect.
5. Protect pages that already earn search traffic before rewriting anything.
6. Make sure your bot rules do not block the fetchers that act for buyers, unless you mean to. Weigh [what blocking Google-Extended may cost](https://underneath.agency/resources/does-blocking-google-extended-hurt-ai-overviews) as well.

If you want help turning log signals into a measured content program, see [our generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

It has not tested whether pages built from AI bot 404 errors earn citations or traffic on their own.

- The only published case is one website, one engine (ChatGPT) and four changes made together.
- No study we reviewed explains why AI bots request addresses that do not exist.
- No study links bot request volume for a page to how often it is cited.
- Log evidence on llms.txt comes from industry sources, not controlled tests.
- The 62.4% bot share is one network’s figure, quoted by a vendor.

## Frequently asked questions

### Can I see ChatGPT crawling my website in server logs?

Yes, OpenAI publishes separate names for its bots, such as GPTBot, OAI-SearchBot and ChatGPT-User, and each does a different job; 15.2% of top sites block GPTBot while only 7.4% block OAI-SearchBot.

### Do AI bot 404 errors show what customers want?

They may hint at it, but the evidence is thin. One site built pages from the most frequent AI bot 404 patterns, as one of four changes, and could not measure that change on its own.

### If an AI bot reads my page, will it cite it?

Not necessarily. On shopping questions, ChatGPT’s developer interface returned 37.38 pages per answer from its searches but cited 3.22.

### How do I know whether AI traffic growth is from my changes?

Compare changed pages with similar pages you left alone. On one site, untouched pages grew 3.5 times, which would otherwise have been credited to the changes.

## Sources

- Watanabe and Nakayashiki (2026), [Disentangling Answer Engine Optimization from Platform Growth: A Log-Based Natural Experiment on ChatGPT Referral Traffic](https://arxiv.org/abs/2606.04362), arXiv:2606.04362.
- Finder, Elovic, Shalev and Yosef (2026), [AX is the New AEO](https://arxiv.org/abs/2609.34951), arXiv:2609.34951.
- Uberti-Bona Marin, Bertaglia, Astante, Rijsbosch, van Dijck, Hannák, Spanakis and Kollnig (2026), ["If I Had to Buy Just ONE: Galaxy S26 Ultra": Auditing AI-Generated Product Recommendations](https://arxiv.org/abs/2609.18729), arXiv:2609.18729.
- Underneath (2026), [Which AI crawlers do top websites block? 10,000 sites, 2026](https://underneath.agency/research/ai-crawler-blocking-study)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [How many websites have an llms.txt file? 2026 adoption data](https://underneath.agency/research/llms-txt-adoption-study)
- Underneath (2026), [Do websites serve Markdown to AI agents? 2026 data](https://underneath.agency/research/agent-readable-web-study)

---

This is the Markdown twin of https://underneath.agency/resources/ai-bot-server-logs-content-demand. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Can AI Cite Your Page for Something It Does Not Say?"
description: "Yes. Audits find AI answers attach citations to pages that do not back the claim, from about 3% of Google AI Overview claims to half of one engine’s citations."
canonical: "https://underneath.agency/resources/ai-citing-pages-for-claims-they-do-not-make"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Can an AI search engine cite my page for something my page does not say?

Yes, and it happens often enough to plan for. Independent audits find AI answers regularly put a link next to a claim the linked page does not support. How often depends heavily on the engine: from a few percent of claims in Google’s AI Overviews to about half the citations of one 2024 answer engine.

## The short version

1. In a 2024 audit of 909 answers, You.com and BingChat cited a page that supported the statement roughly two-thirds of the time; for Perplexity it was 49.0% ([Narayanan Venkit and colleagues, 2024](https://arxiv.org/abs/2410.22349)).
2. Across 98,020 claims in Google’s AI Overviews in spring 2026, 2.66% were contradicted by a page the Overview itself cited ([Xu, Iqbal and Montgomery, 2026](https://arxiv.org/abs/2605.14021)).
3. In our study of “Is this brand legit?” answers, a readable cited page supported the claim fully or in part 72.4% of the time ([our reputation study](https://underneath.agency/research/is-it-legit-ai-reputation-study)).
4. In a test with 4,927 US adults, adding reference links raised trust in AI answers even when the links were wrong or invented ([Li and Aral, 2025](https://arxiv.org/abs/2504.06435)).

## How often do AI citations point to a page that does not back the claim?

Often in older answer engines, much less often in Google’s current AI Overviews.

[Narayanan Venkit and colleagues](https://arxiv.org/abs/2410.22349) ran 303 real questions through You.com, Perplexity and BingChat in 2024, scoring all 909 answers. For each cited statement, an AI judge checked whether the cited page actually supported it.

| Engine (2024) | Citations that pointed to a supporting page |
|---|---|
| You.com | 68.3% |
| BingChat | 65.8% |
| Perplexity | 49.0% |

The authors note the problem went beyond [statements with no support at all](https://underneath.agency/resources/ai-answers-unsupported-claims). Even when one of the listed sources did support a statement, engines often cited a different source that did not. That is the case most relevant to you: a real claim, attributed to the wrong page.

Two cautions apply. These were 2024 engines, and ChatGPT and Google’s AI Overviews were not tested. The AI judge also agreed with human checkers only moderately, so the figures are estimates.

## Is it better in Google’s AI Overviews?

Much better, though not zero.

[Xu, Iqbal and Montgomery](https://arxiv.org/abs/2605.14021) captured Google’s AI Overviews (the AI summary at the top of Google’s results) for 55,393 trending searches over 40 days in spring 2026. They split the Overviews into 98,020 single claims and checked each one against the pages it cited.

- 2.66% of claims were contradicted by a page the Overview cited.
- 6.98% were not mentioned by any cited page at all.
- 41.9% of Overviews had every claim backed by their sources.
- 2.74% of Overviews had fewer than half their claims backed.

Some unbacked claims may come from social and video pages the researchers could not read. Even assuming all of those were fine, the unbacked rate would fall only from 11.0% to roughly 5.3%. For a business, that means a small but real share of claims arrive with a link that does not say what the summary says.

## What does misattribution look like in practice?

A correct-sounding claim, a real link, and a page that says something different or nothing on the point.

In the 2024 user study, all 21 expert participants spotted misattribution at some point. One said a statement “doesn’t seem to be in the source,” even though the statement itself was true. Others described [engines picking only one side](https://underneath.agency/resources/do-ai-search-engines-agree-with-leading-questions) of a source that discussed both.

Our own [reputation study](https://underneath.agency/research/is-it-legit-ai-reputation-study) found a commercial version of this. We asked four AI engines whether 79 brands were legitimate and checked 240 cited claims against the pages they cited:

- Where a cited page could be read, it supported the claim fully or in part 72.4% of the time.
- Of the checkable claims saying a problem was common, 71.4% cited evidence showing only individual reports, or nothing about frequency.
- 119 of the 240 claims could not be checked at all, because review sites block automated reading.

That second point is the subtle risk. Your page, or a review page about you, can be cited for a broad claim such as “customers often complain” when it only shows a few cases.

## Do readers notice when a citation is wrong?

Mostly not, and a citation can make a wrong answer more convincing.

[Li and Aral](https://arxiv.org/abs/2504.06435) ran a randomized experiment with 4,927 adults, chosen to represent the US population. Adding reference links to AI search answers raised trust. The increase was the same whether the links were valid or invented and broken.

People also rarely click through. A Pew Research study, reported by [Huang and colleagues](https://arxiv.org/abs/2603.16138), found users clicked sources cited inside Google’s AI summary in only 1% of visits. So a misattributed claim with your name on the link will usually be read, not checked.

## Why would an AI attach your URL to a claim you never made?

Because an answer blends several pages, and engines often link a sentence to the wrong one.

The research points to three patterns:

- **Linking the wrong source.** Narayanan Venkit and colleagues found engines often cited a page that did not support a statement even when another listed page did.
- **Uneven use of sources.** [Huang and colleagues](https://arxiv.org/abs/2603.16138) studied 11,000 real searches. Google’s AI Overviews drew 13.8 percentage points less content from negatively worded sources than average, while still citing them. Being cited is not the same as being used; see [how unevenly AI search uses cited pages](https://underneath.agency/resources/does-ai-search-use-the-pages-it-cites).
- **Overstating what a page shows.** In our reputation study, claims that a problem was common rested on single reports.

An earlier study cited by Huang and colleagues found only 51.5% of generated statements were fully supported by their citations. The pattern recurs across engines and years.

## What should you do about it?

Monitor what AI answers attribute to you, make your own claims unmistakable, and keep evidence of what your pages said.

1. **Search AI engines for your brand and key claims.** Click the citations to your pages and check that each linked page says what the answer says.
2. **State key facts once, plainly and specifically.** Clear, quotable sentences give an engine less room to blend your page with someone else’s.
3. **Keep dated copies of important pages.** If an answer attributes a statement to you, you can show what the page said at the time.
4. **Watch review sites, not just your own pages.** Broad claims about your reputation are often tied to review pages showing individual cases.
5. **Report serious misattributions to the platform.** Where a wrong claim carries legal or safety risk, raise it through the engine’s feedback channel and get legal advice.

If you want help tracking how AI engines describe and cite you, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The research measures how often citations miss, but not how often that harms the business being cited.

- The engine-by-engine citation accuracy figures come from 2024 engines. Comparable current figures for ChatGPT and Perplexity are scarce in the papers we reviewed.
- Most audits use AI judges to decide whether a page supports a claim, which adds error.
- The AI Overview audit used trending searches, not the commercial searches buyers make about vendors.
- Our reputation study could not check half its sampled claims because review sites block automated reading.
- No study measures the reputational or legal consequences for a company wrongly cited.

## Frequently asked questions

### Can Google’s AI Overviews misquote my website?

Occasionally. In a 2026 audit of 98,020 claims, 2.66% were contradicted by a page the Overview cited, and 6.98% were not mentioned by any cited page.

### Which AI search engine cites sources most accurately?

No current head-to-head answer exists. In a 2024 audit, Perplexity’s citations supported their statements 49.0% of the time, against 68.3% for You.com.

### Do people check the sources AI answers cite?

Rarely. Pew found users clicked sources cited in Google’s AI summary in only 1% of visits, and links raised trust even when invalid.

### What should I do if an AI answer attributes a false claim to my company?

Record the answer and your page as it stood, then report it to the platform. For serious cases, get legal advice.

## Sources

- Narayanan Venkit, Laban, Zhou, Mao and Wu (2024), [Search Engines in an AI Era: The False Promise of Factual and Verifiable Source-Cited Responses](https://arxiv.org/abs/2410.22349), arXiv:2410.22349.
- Xu, Iqbal and Montgomery (2026), [Measuring Google AI Overviews: Activation, Source Quality, Claim Fidelity, and Publisher Impact](https://arxiv.org/abs/2605.14021), arXiv:2605.14021.
- Li and Aral (2025), [Human Trust in AI Search: A Large-Scale Experiment](https://arxiv.org/abs/2504.06435), arXiv:2504.06435.
- Huang, Goyal, Saha and Chandrasekharan (2026), [Answer Bubbles: Information Exposure in AI-Mediated Search](https://arxiv.org/abs/2603.16138), arXiv:2603.16138.
- Underneath (2026), [“Is this brand legit?” How AI assistants build a reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study)

---

This is the Markdown twin of https://underneath.agency/resources/ai-citing-pages-for-claims-they-do-not-make. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How do AI coding tools win developers through AI search?"
description: "By earning developer trust where AI answers look: docs, benchmarks, community threads and honest comparisons, then growing one developer into a team of seats."
canonical: "https://underneath.agency/resources/ai-coding-tools-developers-ai-search"
published: 2026-10-09
updated: 2026-10-10
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How do AI coding tools win developers through AI search?

By being the tool an AI assistant names when a developer asks what to use, backed by the public evidence developers check before they trust it: documentation, benchmarks, security facts and community threads. Developers already work inside AI tools all day, so the recommendation often happens in the same window as the code. Most revenue then comes from seats, as one developer’s trial turns into a team and then a company-wide license.

## The short version

1. Use is near-universal but trust is falling. In the [2025 Stack Overflow survey](https://survey.stackoverflow.co/2025/ai), 84% of respondents use or plan to use AI tools, yet more developers distrust their accuracy (46%) than trust it (33%).
2. The market is concentrated. Among developers who use or build AI agents, 82% had used ChatGPT and 68% GitHub Copilot as ready-made assistants; [Copilot](https://techcrunch.com/2025/07/30/github-copilot-crosses-20-million-all-time-users/) passed 20 million all-time users, and [Cursor](https://cursor.com/blog/series-d) crossed $1 billion in annualized revenue.
3. Revenue grows by seats: [Cursor’s plans](https://cursor.com/pricing) run from $20 a month for an individual to $40 per user per month for teams, and [Gartner expects](https://www.gartner.com/en/newsroom/press-releases/2024-04-11-gartner-says-75-percent-of-enterprise-software-engineers-will-use-ai-code-assistants-by-2028) 75% of enterprise software engineers to use AI code assistants by 2028.
4. AI answers already send developers to products. ChatGPT accounted for around 10% of new signups at [Vercel](https://vercel.com/blog/how-were-adapting-seo-for-llms-and-ai-search), a developer platform, up from 1% six months earlier.
5. A coding tool that ChatGPT knows by name still may not be suggested. In [Sharma’s test](https://arxiv.org/abs/2601.00912) of 112 Product Hunt startups, using questions such as “What are the best AI coding assistants launched in 2025?”, ChatGPT recognized 99.4% by name but surfaced only 3.32% in discovery questions.

## Who pays for AI coding assistants, and how big can an account get?

Developers choose them first; engineering leaders and procurement buy the seats once a team depends on them.

The buying motion starts bottom-up. An individual developer tries a tool on a free or personal plan, such as [GitHub Copilot’s](https://github.com/features/copilot/plans) Free, Pro ($10 a month), Pro+ ($39) or Max ($100) tiers, or Cursor’s $20 individual plan. If it sticks, the team asks for a shared plan, and the company eventually signs for every engineer, adding single sign-on, data controls and central billing.

That is why a single developer is worth far more than one subscription. Cursor charges teams $40 per user per month, so a 200-engineer organization is a different deal from one developer on a personal plan. Microsoft reported that Copilot is used by 90% of the Fortune 100 and that its growth among enterprise customers rose about 75% from the previous quarter. [Anthropic](https://www.anthropic.com/news/anthropic-raises-series-f-at-usd183b-post-money-valuation) said Claude Code was generating over $500 million in run-rate revenue within months of its full launch in May 2025. AI voice companies see a similar pattern, where a developer starts small on an API and spending grows with usage, as [how voice companies turn AI answers into revenue](https://underneath.agency/resources/ai-voice-software-revenue-from-ai-search) shows.

The enterprise wave is still forming. Gartner’s survey of 598 respondents found 63% of organizations piloting, deploying or already using AI code assistants, and it forecasts adoption among enterprise engineers rising from less than 10% in early 2023 to 75% by 2028. Each organization that standardizes on one tool decides where hundreds or thousands of seats go.

## Where do AI assistants sit in how developers find tools?

Everywhere: developers ask AI for answers daily, then check documentation, GitHub and community threads before they commit.

Developers are heavy users of the assistants that now recommend tools. Stack Overflow found 51% of professional developers use AI tools daily, and that developers who mostly use AI in their workflow frequently use it to search for answers or learn new concepts. [JetBrains’ 2025 survey](https://blog.jetbrains.com/research/2025/10/state-of-developer-ecosystem-2025/) found 85% of developers regularly use AI tools for coding and 62% rely on at least one AI coding assistant, agent or code editor.

They do not stop at the answer. The [Stack Overflow results summary](https://stackoverflow.blog/2025/07/29/developers-remain-willing-but-reluctant-to-use-ai-the-2025-developer-survey-results-are-here/) shows developers rely on a portfolio of community resources: Stack Overflow (84%), GitHub (67%) and YouTube (61%). People learning to code still use technical documentation more than any other resource (68%). And 75% said they would still ask another person for help when they do not trust AI’s answers.

Some vendors can already see AI answers in their signups. Vercel itself traced around 10% of its new signups to ChatGPT, up from 4.8% the previous month. Developers are also asking more than one assistant, and ChatGPT’s lead is narrowing: its share of generative AI website visits went from about 76% in June 2025 to roughly 53% by May 2026 in [Similarweb’s 2026 data](https://aisearch.similarweb.com/blog/gen-ai-stats/), as Gemini and Claude grew.

Distribution matters as much as discovery. GitHub’s [Octoverse 2025](https://github.blog/news-insights/octoverse/octoverse-a-new-developer-joins-github-every-second-as-ai-leads-typescript-to-1/) reports that nearly 80% of new developers on GitHub use Copilot within their first week. A tool built into the platform where developers already work starts with an advantage that a challenger must earn through recommendations. The same contest plays out for [DevOps tools engineers ask AI about](https://underneath.agency/resources/devops-platforms-ai-search).

## Which questions do developers ask about AI coding tools?

Technical ones: which tool fits a stack, how it scores, its cost per seat, and what happens to the code.

These coding-tool prompts are our own examples of what developers and engineering managers ask; they were not observed in real logs:

- Comparison: “Cursor vs GitHub Copilot vs Claude Code for a large TypeScript monorepo.”
- Workflow fit: “Best AI code review tool that comments on GitHub pull requests.”
- Evidence: “Which coding agent has the highest SWE-bench Verified score, and how was it run?”
- Security: “Which AI coding assistants don’t train on our code and support self-hosting?”
- Cost: “What would AI coding assistants cost per seat for 200 engineers with single sign-on?”
- Migration: “Is it worth switching from Copilot to Cursor for a Python data team?”

Answers to these questions hinge on hard facts that assistants can get wrong. Coding assistants rename plans and change usage credits often, and in [our pricing study](https://underneath.agency/research/ai-pricing-accuracy-study), only 61.9% of the plan prices four assistants quoted for 45 software products were fully faithful to the vendor’s pricing page. Benchmark claims age quickly too, as each model release reshuffles the [public leaderboards](https://www.swebench.com/).

## How does one developer’s trial grow into a company-wide seat license?

Through a trial that spreads: one developer tries the tool, the team adopts it, and procurement buys seats.

**Named.** An answer to a comparison or workflow question puts your tool on a developer’s list. With free tiers common, the cost of trying is a few minutes. AI video tools run a similar [path from free plan to upgrade](https://underneath.agency/resources/ai-video-generator-customers-from-ai-search).

**Individual trial.** The developer installs the tool, often from an editor marketplace, and tests it on real work. This is where the product either earns trust or is dropped; Stack Overflow found 66% of developers frustrated by AI solutions that are “almost right, but not quite.”

**Team plan.** Developers who keep using a tool ask for it at work. Seat prices such as Cursor’s $40 per user per month turn a personal choice into a budget line.

**Enterprise license.** Security, legal and procurement review the vendor. Data handling is a deciding question. GitHub, for example, documents that it does not use Copilot Business or Enterprise data to train its models, while individual subscribers’ interaction data may be used unless they opt out.

AI search can enter at more than one step. A developer may first hear of a tool in an answer, and an engineering manager may later ask an assistant to compare the three tools the team is already using. We infer that the second question matters as much as the first, because it is where hundreds of seats are decided.

## Why does an assistant suggest Copilot, Cursor or a newcomer?

Platforms document little; studies point to independent coverage and developer communities; developers add hard evidence and security.

**Documented by the platform.** ChatGPT, Claude, Gemini and Copilot say nothing public about how they pick which code assistant to suggest. GitHub’s own plans page describes Copilot as “the competitive advantage developers ask for by name,” a vendor claim, not an observed result.

**Observed in studies.** In Sharma’s Product Hunt study, startups whose own sites scored higher on GEO-style content checks were no more likely to be surfaced, while the number of other sites linking to them and community discussion predicted visibility in Perplexity. Across US software ranking questions in the work of [Chen and colleagues](https://arxiv.org/abs/2509.08919), AI search drew 72.7% of its sources from independent “earned” sites, against 45.4% for Google, so outside coverage counts for more in AI answers than a coding tool’s own site. In our [Reddit study](https://underneath.agency/research/ai-reddit-citations-study), Google’s AI Overviews cited Reddit in 35.4% of B2B software answers, and specialist communities supplied 58.7% of the Reddit threads shown in software searches. In [our freshness study](https://underneath.agency/research/ai-source-freshness-study), four assistants cited pages first published about half as long ago as Google’s top 10 for the same questions (a ratio of 0.50), which matters in a category that ships new models every few weeks.

**Trust factors specific to coding tools.** Developers are skeptical by habit, and the evidence gives them reasons:

- **Measured productivity.** In a [randomized trial by METR](https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/) with 16 experienced open-source developers working on 246 real issues, developers took 19% longer with AI tools, though they expected a 24% speedup and afterward believed AI had sped them up by 20%. Results from one setting in early 2025, but developers know this study.
- **Security of generated code.** [Veracode](https://www.veracode.com/blog/genai-code-security-report/) tested over 100 models and found 45% of code samples failed security tests. Vendors that sell code security face their own version of this, set out in [our article on application security vendors](https://underneath.agency/resources/application-security-demand-ai-search).
- **Overall sentiment.** [Google’s 2025 DORA report](https://cloud.google.com/blog/products/ai-machine-learning/announcing-the-2025-dora-report) found 90% of respondents use AI at work, while 30% report little or no trust in the code it generates.
- **Transparent benchmarks and data policies.** We infer that public, reproducible benchmark results, clear training and retention policies and honest limits give both developers and assistants something checkable to cite.

## What does a coding assistant lose if AI answers skip it?

The trial, and with it the team and enterprise seats that grow from it.

The market leaders already have distribution that does not depend on answers: Copilot inside GitHub, ChatGPT as the general assistant most developers use. Among developers who use or build AI agents, 82% had used ChatGPT and 68% Copilot as ready-made assistants. A challenger that the assistants do not name has fewer ways to reach the developer who is about to try something new.

Seat economics raise the stakes. Once a team has standardized on a tool, configured it and built habits around it, expansion revenue flows to that vendor, and switching means a new security review. That is our inference from how these products are sold; no study has yet measured AI-driven switching between coding tools. For the general mechanics of missing new products, see [our article on why ChatGPT misses new products](https://underneath.agency/resources/why-chatgpt-misses-new-products).

## What does GEO look like for a code assistant or coding agent?

It gives assistants and skeptical developers evidence they can check; no coding tool can be promised a recommendation.

For a coding assistant, generative engine optimization (GEO) breaks into seven jobs:

1. **Documentation that machines can read.** Public, current docs and quick-starts, available as clean text. Our [llms.txt study](https://underneath.agency/research/llms-txt-adoption-study) found only 11.5% of top websites publish a valid llms.txt file, and our [agent-readable web study](https://underneath.agency/research/agent-readable-web-study) found 3.2% return Markdown when an agent asks. Our article on [what AI agents do when they cannot read your site](https://underneath.agency/resources/when-ai-agents-cant-read-your-site) explains why it matters.
2. **Benchmarks with methods.** Publish results on public benchmarks with the setup, date and model version, and link the raw runs. Developers and assistants can check reproducible numbers; they discount bare claims. Research tools face a similar test on [proof of citation accuracy](https://underneath.agency/resources/ai-research-tools-users-ai-search).
3. **Community presence on the record.** Real answers in GitHub issues and discussions, Stack Overflow, specialist Reddit communities and developer YouTube, from your engineers, not marketing accounts.
4. **Honest comparison and migration pages.** Fair comparisons with named alternatives, including where they are better, and step-by-step migration guides. We look at whether such pages earn citations in [our summary on comparison pages](https://underneath.agency/resources/do-comparison-pages-help-b2b-ai-citations).
5. **Security and data pages.** What you train on, what you retain, where data is processed, self-hosting options and certifications, in plain words on public pages.
6. **A pricing page an assistant can quote.** Seat prices, usage credits and overage rules in one current place, with old plans retired.
7. **Repeated checks where developers ask.** Rerun a fixed set of comparison, workflow, security and price questions in ChatGPT, Claude, Gemini, Perplexity, Copilot and Google, and ask each new workspace how it heard of you. The broader software picture is in [our B2B SaaS article](https://underneath.agency/resources/b2b-saas-revenue-from-ai-search).

## What is still unknown about AI answers and coding-tool adoption?

Whether an assistant’s answer decides which coding tool a company standardizes on; nobody has measured it.

The adoption figures come from developer surveys with different samples and questions, and the revenue figures are company-reported. Vercel’s signup share is a single developer-platform example, not a coding-assistant figure. The productivity and security studies measure the tools, not how assistants recommend them. No published study follows an AI recommendation from a developer’s first trial to an enterprise seat count, and we have not tested how assistants answer coding-tool comparisons ourselves. Ways to judge whether that work pays off in seats are covered in [our article on business results](https://underneath.agency/resources/does-ai-visibility-drive-business-results).

## How can a coding tool vendor see whether AI answers are costing it seats?

Check what assistants say when developers and engineering leaders compare you with tools they already use.

A practical first step is an audit of comparison, workflow, security and pricing questions across the main assistants, matched against where your trials, team plans and enterprise seats come from, so the gaps that cost the most seats are fixed first. To see that comparison mapped against your trials and team plans, [ask us for a coding-tool answer review](https://underneath.agency/contact). What the follow-on work covers, from docs and benchmark pages to developer community presence, is laid out on our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page.

## Frequently asked questions

### Do benchmark scores decide which coding tool an AI assistant recommends?

Not on their own, as far as anyone has shown. No assistant documents how it chooses. Benchmarks help when they are public, recent and reproducible, because they give independent sites and developers something to cite.

### Should coding tool companies publish comparison pages against Copilot or Cursor?

Yes, if they are fair and specific about where each tool is better. Developers distrust one-sided pages, and assistants draw mostly on independent sources, so your own page works best alongside third-party comparisons.

### Do developer communities like Reddit and Stack Overflow affect AI answers?

They appear in them. In our research, Google’s AI Overviews cited Reddit in about a third of software answers, mostly from specialist communities. How much any community shapes ChatGPT or Claude answers is not documented.

### How can a developer tool see whether signups come from AI search?

Read three signals side by side: referrals from AI assistants, a question at signup or in the editor extension about where the developer heard of you, and regular checks of what assistants say for your main comparisons. Vercel’s experience shows the share can rise quickly once you start measuring it.

## Sources

- Stack Overflow (2025), [2025 Developer Survey: AI](https://survey.stackoverflow.co/2025/ai) and [Developers remain willing but reluctant to use AI](https://stackoverflow.blog/2025/07/29/developers-remain-willing-but-reluctant-to-use-ai-the-2025-developer-survey-results-are-here/)
- JetBrains (2025), [State of Developer Ecosystem 2025](https://blog.jetbrains.com/research/2025/10/state-of-developer-ecosystem-2025/)
- Google Cloud (2025), [Announcing the 2025 DORA Report](https://cloud.google.com/blog/products/ai-machine-learning/announcing-the-2025-dora-report)
- GitHub (2025), [Octoverse: A new developer joins GitHub every second](https://github.blog/news-insights/octoverse/octoverse-a-new-developer-joins-github-every-second-as-ai-leads-typescript-to-1/)
- GitHub (2026), [Copilot plans and pricing](https://github.com/features/copilot/plans)
- TechCrunch (2025), [GitHub Copilot crosses 20 million all-time users](https://techcrunch.com/2025/07/30/github-copilot-crosses-20-million-all-time-users/)
- Cursor (2025), [Series D](https://cursor.com/blog/series-d) and (2026) [Pricing](https://cursor.com/pricing)
- Anthropic (2025), [Anthropic raises Series F at $183B post-money valuation](https://www.anthropic.com/news/anthropic-raises-series-f-at-usd183b-post-money-valuation)
- Gartner (2024), [Gartner Says 75% of Enterprise Software Engineers Will Use AI Code Assistants by 2028](https://www.gartner.com/en/newsroom/press-releases/2024-04-11-gartner-says-75-percent-of-enterprise-software-engineers-will-use-ai-code-assistants-by-2028)
- METR (2025), [Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity](https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/)
- Veracode (2025), [2025 GenAI Code Security Report](https://www.veracode.com/blog/genai-code-security-report/)
- SWE-bench (2026), [Leaderboards](https://www.swebench.com/)
- Vercel (2025), [How we’re adapting SEO for LLMs and AI search](https://vercel.com/blog/how-were-adapting-seo-for-llms-and-ai-search)
- Similarweb (2026), [AI Search Stats 2026: Market Share, Referral, and Citation Data](https://aisearch.similarweb.com/blog/gen-ai-stats/)
- Chen and colleagues (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919)
- Sharma (2025), [The Discovery Gap: How Product Hunt Startups Vanish in LLM Organic Discovery Queries](https://arxiv.org/abs/2601.00912)
- Underneath (2026), [Reddit citations](https://underneath.agency/research/ai-reddit-citations-study), [source freshness](https://underneath.agency/research/ai-source-freshness-study), [software pricing accuracy](https://underneath.agency/research/ai-pricing-accuracy-study), [llms.txt adoption](https://underneath.agency/research/llms-txt-adoption-study) and [agent-readable web](https://underneath.agency/research/agent-readable-web-study)

---

This is the Markdown twin of https://underneath.agency/resources/ai-coding-tools-developers-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Can You Trust an AI’s Reason for Recommending a Rival?"
description: "Only partly. In a 12-model hotel test, AI assistants’ stated reasons matched their real drivers imperfectly and left out list order and review count."
canonical: "https://underneath.agency/resources/ai-explanations-for-recommendations"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Can we trust an AI assistant’s explanation of why it recommended a competitor?

Only partly: an AI assistant’s stated reason is a rough guide to its choice, not a faithful account of it. In a large controlled test with synthetic hotels, the reasons models gave lined up with what actually drove their picks only imperfectly, and they almost never mentioned two factors they acted on.

## The short version

1. In a test of 12 AI models choosing between hotels, the factors named in the models’ reasons matched the factors that really drove their choices only partly, scoring 0.59 to 0.85 on a 0-to-1 scale ([Baig and colleagues, 2026](https://arxiv.org/abs/2606.16344)).
2. Where a hotel appeared in the list drove 4.1% of the choices, yet the models mentioned it in no more than 0.7% of their reasons (same study).
3. Brand affiliation appeared in up to 55% of one model’s reasons while accounting for only 2.0% of what moved its choices (same study).
4. When planted web pages fooled AI recommenders into backing a fake product, their answers contained 1.5 to 11 times more invented praise than answers that resisted ([Luo and Chen, 2026](https://arxiv.org/abs/2606.13610)).

## Why does it matter what an AI says about its own choice?

Because teams use the AI’s own words to diagnose lost visibility, and those words can point at the wrong cause.

When an assistant recommends a rival, it usually adds a short justification: better reviews, a stronger brand, a lower price. It is tempting to treat that sentence as the diagnosis and fix whatever it names. If the explanation leaves out what actually tipped the decision, the fix targets the wrong thing.

The research below separates two questions. What does the assistant say drove its choice? And what, measured by experiment, actually did?

## What did the hotel experiment test?

It randomly varied hotel details, then compared what moved each model’s picks with what each model said moved them.

[Baig and colleagues](https://arxiv.org/abs/2606.16344) showed 12 AI models, including OpenAI, Google and Anthropic models and four freely downloadable ones, sets of five made-up hotels. Rating, review count, review recency, price, chain status, an eco-label and list position were all shuffled at random. Each model picked one hotel and gave a one- or two-sentence reason. Across the whole project the team made 61,459 model calls.

Because every detail was assigned at random, the team could measure how much each one changed a hotel’s chances. [A top rating mattered most](https://underneath.agency/resources/why-ai-recommends-one-hotel-over-another): a 4.7-star hotel was picked 31.6 percentage points more often than a 3.9-star one. Then the team coded all 35,223 usable reasons to see which details the models named.

Two caveats matter. The hotels were synthetic, and two of the three authors work for a travel company, which the paper discloses. The test also covers only the final pick from a ready-made shortlist, not how an assistant finds candidates on the web.

## How closely did the stated reasons match the real drivers?

They matched in broad strokes but missed specific factors in a consistent way.

On a 0-to-1 scale, where 1 means the reasons ranked the factors exactly as the choices did, every model scored between 0.59 and 0.85. That is real agreement, but far from a faithful account. The gaps were not random:

| Factor | Share of what drove the choices | How often the reasons mentioned it |
|---|---|---|
| Position in the list | 4.1% | At most 0.7% of reasons |
| Number of reviews | 9.4% | Rarely |
| Chain or brand affiliation | 2.0% | Up to 55% of one model’s reasons |

The authors conclude that model-written explanations “cannot be relied upon, on their own” as an honest disclosure of what drove a decision. Put simply, the assistant talks about brand more than brand matters to it, and talks about list order and review count less than they matter.

## What do AI assistants act on without mentioning?

Mostly things outside the product itself, such as the order in which options were shown.

List position is the clearest case. In the hotel study, [being listed first](https://underneath.agency/resources/does-list-order-change-ai-recommendations) was worth $11.7 per night in the models’ choices, without changing anything about the hotel. One model, Gemini 2.0 Flash, showed a first-place advantage of about 26 percentage points. No hotel manager reading the model’s reason would learn this, because the reasons almost never mention order.

A second study shows a worse version of the same gap. [Luo and Chen](https://arxiv.org/abs/2606.13610) rewrote real web pages to promote fake products across 225 real products, then asked 12 AI models for recommendations, mostly in Chinese with an English repeat. A single planted page fooled models up to 27% of the time. When fooled, the models did not say “one page told me so.” Their answers contained 1.5 to 11 times more social-proof language, such as invented community praise, than answers that resisted. The stated reason was not just incomplete; it was made up after the fact.

Our own [prompt phrasing study](https://underneath.agency/research/ai-prompt-phrasing-study) shows why reasons sound convincing. When a budget changed which brands an assistant named, new brands were given a price or value reason 55.9% of the time. A reason was almost always supplied. Whether it caused the switch is a separate question the answer cannot settle.

## Do the sources an AI cites explain its answer either?

Not reliably: citations show what an assistant chose to display, not what it drew on most.

[Huang and colleagues](https://arxiv.org/abs/2603.16138) compared Google’s AI Overviews (the AI summary at the top of Google’s results) with the pages they cited. Overviews cited Reddit and Quora, yet drew 22.1% less of their content from those social sources than from other cited sources. The citation list suggested a breadth the summary did not use.

[Narayanan Venkit and colleagues](https://arxiv.org/abs/2410.22349) found the same pattern in 2024 answer engines. BingChat listed sources it never cited in its answer text, 36% of them on average. In their user study, 15 of 21 expert participants raised the lack of clarity about why some sources were chosen over others.

So neither the assistant’s sentence of reasoning nor its source list is a full record of why a rival won.

## What should you do about it?

Treat the AI’s explanation as one clue, then test the cause yourself.

1. **Log the reason, but do not act on it alone.** Record what assistants say about why a rival was chosen, then check it against evidence before changing anything.
2. **Test the factors the reasons tend to skip.** Ask the same question several times and in several orders or wordings. The hotel study shows order and review volume can matter without ever being mentioned.
3. **Fix what is measurable on your side.** [Ratings and price carried the most weight](https://underneath.agency/resources/ai-vs-human-reputation-priorities) in the hotel test. Visible facts like these are worth getting right whatever the assistant says.
4. **Watch for invented praise around rivals.** If an answer credits a competitor with community buzz you cannot find, the cause may be a planted or low-quality page rather than real reputation.
5. **Re-check after model updates.** The hotel weights describe specific model versions, and the authors warn they can drift.

If you want help running these checks across assistants, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The research shows reasons and real drivers diverge, but it has not mapped that gap for most business categories.

- The main evidence is one hotel experiment with synthetic listings, single-turn questions and English prompts. It may not carry over to software, finance or local services.
- It tested only the final pick from a fixed shortlist. How well assistants explain which candidates they found in the first place is unstudied.
- The authors call their reason-versus-choice comparison exploratory, and two of them are affiliated with a travel company.
- No study yet tests whether assistants explain choices more honestly when asked directly, or across a longer conversation.
- Models change often, so today’s gaps may shrink, grow or move.

## Frequently asked questions

### Why did ChatGPT recommend my competitor instead of me?

The answer itself will not reliably tell you. In a 12-model hotel experiment, stated reasons matched real drivers only partly, scoring 0.59 to 0.85 on a 0-to-1 scale, and skipped list order entirely.

### Can I ask an AI assistant to explain its recommendation?

You can, but treat the answer as a hypothesis. In the hotel study, brand was mentioned in up to 55% of one model’s reasons while driving only 2.0% of its choices.

### Does the order of options affect AI recommendations?

In controlled tests, yes. Being listed first was worth $11.7 per night in hotel choices, and one model showed a first-place advantage of about 26 percentage points.

### Are AI citations a reliable guide to why an answer says what it says?

Not on their own. Google’s AI Overviews drew 22.1% less content from cited social sources than from other cited pages, so a citation list overstates what was used.

## Sources

- Baig, Gillani and Ali (2026), [Whose hotel does the AI recommend? An algorithm audit of reputation signals in LLM-assisted hotel selection](https://arxiv.org/abs/2606.16344), arXiv:2606.16344.
- Luo and Chen (2026), [One Polluted Page Is Enough: Evaluating Web Content Pollution in LLM Recommenders](https://arxiv.org/abs/2606.13610), arXiv:2606.13610.
- Huang, Goyal, Saha and Chandrasekharan (2026), [Answer Bubbles: Information Exposure in AI-Mediated Search](https://arxiv.org/abs/2603.16138), arXiv:2603.16138.
- Narayanan Venkit and colleagues (2024), [Search Engines in an AI Era: The False Promise of Factual and Verifiable Source-Cited Responses](https://arxiv.org/abs/2410.22349), arXiv:2410.22349.
- Underneath (2026), [Does rewording a question change AI brand recommendations?](https://underneath.agency/research/ai-prompt-phrasing-study)

---

This is the Markdown twin of https://underneath.agency/resources/ai-explanations-for-recommendations. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How do AI image generators get recommended by AI assistants?"
description: "By being clear on what they do best and on commercial rights, then getting that into the independent reviews, galleries and lists AI assistants read."
canonical: "https://underneath.agency/resources/ai-image-generators-growth-from-ai-search"
published: 2026-10-09
updated: 2026-10-10
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How do AI image generators get recommended by AI assistants?

By being specific about what they do best and clear about commercial rights, then getting that story into the independent reviews, galleries, tutorials and “best of” lists that AI assistants read. Image generation is crowded, the leading chat assistants now make images themselves, and business buyers ask about licensing before quality. A generator that answers the rights question plainly gives assistants, and buyers, a reason to name it.

## The short version

1. ChatGPT is now an image tool too: multimedia requests grew from 2% to just over 7% of its messages, with a spike after its new image generation launched in 2025, per a [study by OpenAI and Harvard economists](https://www.nber.org/system/files/working_papers/w34255/w34255.pdf).
2. Paid growth is expensive in this category: top-quartile Photo & Video apps pay above $14 per install on iOS, per [RevenueCat’s](https://www.revenuecat.com/state-of-subscription-apps-2025/) 2025 benchmarks.
3. Business demand is real: 53% of B2B marketers whose organizations use AI use creative asset tools for images and video, per the [Content Marketing Institute](https://contentmarketinginstitute.com/b2b-research/b2b-content-marketing-trends-research).
4. Rights terms differ sharply: [Midjourney](https://docs.midjourney.com/hc/en-us/articles/32083055291277-Terms-of-Service) requires companies with more than $1,000,000 a year in revenue to be on a Pro or Mega plan to own their images, while [Getty Images](https://newsroom.gettyimages.com/en/getty-images/getty-images-launches-commercially-safe-generative-ai-offering) offers uncapped indemnification.
5. Rules are tightening: the [US Copyright Office](https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-2-Copyrightability-Report.pdf) concluded that prompts alone do not make users authors, and the [EU AI Act’s Article 50](https://artificialintelligenceact.eu/article/50/), which comes into force on 2 August 2026, requires machine-readable marking of generated images.

## Who uses AI image generators, and what is a customer worth?

Individual creators on monthly plans, plus marketing teams and brands that pay more for rights and control.

There are three broad buyers. Individual creators and hobbyists want quality and style. Marketers want speed: in the Content Marketing Institute’s survey, creative asset tools were the second most used kind of AI tool, at 53% of marketers whose organizations use AI. Brands, agencies and media companies want images they can use without legal risk, and they will pay for that.

Pricing reflects those tiers. Midjourney’s [plans page](https://docs.midjourney.com/hc/en-us/articles/27870484040333-Comparing-Midjourney-Plans) lists four subscriptions from $10 to $120 a month, with a 20% discount for paying a year upfront. Its terms tie ownership of images to plan level for larger companies, which pushes business users to the higher tiers. Stock providers sell rights as the product: Getty says customers get “representations and warranties, uncapped indemnification” on images from its generator.

The consumer economics are attractive but uneven. For AI apps in general, RevenueCat put typical earnings above $0.63 per install at the 60-day mark, against a $0.31 median for all apps. In Photo & Video apps, top apps earn 5–7 times more than the median, and 27.57% of apps reached $1,000 in revenue within two years, the highest success rate of any category. The catch is acquisition cost: RevenueCat attributes Photo & Video’s high cost per install to “the competitive landscape.” A user who arrives because an AI assistant named you costs nothing per install, which is why the channel matters here.

## What happens to image generators now that chat assistants make pictures too?

They are now image generators themselves, and they are also where users ask which tool to use.

The OpenAI and Harvard study of ChatGPT messages found multimedia requests grew from 2% to just over 7% of usage, “with a large spike in April 2025 after ChatGPT released new image-generation capabilities,” and the level stayed elevated. [OpenAI’s announcement](https://openai.com/index/introducing-4o-image-generation/) of that feature also says all generated images “come with C2PA metadata” identifying them as AI-made. For a casual user, the assistant may be all the image generator they need. Writing tools face the same squeeze; see [how AI writing tools still win users](https://underneath.agency/resources/ai-writing-tools-users-from-ai-search).

Specialists still hold ground. In [Andreessen Horowitz’s fifth Top 100 list](https://a16z.com/100-gen-ai-apps-5/) of consumer AI products, Midjourney and Leonardo are among the fourteen “All Stars” that have appeared in all five editions of its web list, and Midjourney is “famously bootstrapped.” We infer that the durable positions are style and quality leadership, control (editing, consistency, brand styles), workflow integration and commercial safety, not simple text-to-image.

## Which questions do image-generator users ask AI assistants?

Style, use-case, comparison, price and rights questions, with rights dominating for business users.

We made up the prompts below to show how creators and marketers tend to phrase image-tool requests; they are not logged queries:

- Style: “Best AI image generator for photorealistic product shots on white backgrounds.”
- Use case: “AI tool that keeps the same character across a 20-page children’s book.”
- Comparison: “Midjourney vs Adobe Firefly vs ChatGPT images for a marketing team.”
- Price: “Cheapest AI image generator with no watermark and unlimited images.”
- Rights: “Which AI image generators are safe for commercial use and offer indemnification?”
- Ownership: “Can I copyright images I make with an AI image generator?”

Rights questions are where a wrong or vague answer costs the most. They also have documented answers: the Copyright Office concludes that “prompts alone do not provide sufficient human control to make users of an AI system the authors of the output,” while creative human arrangements or modifications of the output can be protected. A vendor that explains this plainly, alongside its own license terms, gives an assistant something accurate to repeat.

## Why do licensing and provenance decide trust in this category?

Because business buyers carry the legal risk, and lawsuits and new rules have made that risk visible.

Litigation is active. [NPR reports](https://www.npr.org/2025/06/12/nx-s1-5431684/ai-disney-universal-midjourney-copyright-infringement-lawsuit) that Disney and Universal filed a 110-page lawsuit against Midjourney alleging that it trained on “countless” copyrighted works, and that Getty Images has sued Stability AI. The Copyright Office report drew on more than 10,000 public comments, a sign of how contested the topic is. For a marketing or legal team, the question “is this safe to use?” comes before “is this the best image?”

Vendors answer it in different ways, and the differences are concrete:

| Rights question | Example of a documented answer |
|---|---|
| Who owns the output? | Midjourney: users own assets, but companies over $1,000,000 in revenue must be on Pro or Mega |
| Is it protected if someone sues? | Getty: uncapped indemnification, trained solely on its own library |
| Can others tell it is AI-made? | OpenAI: C2PA metadata on all generated images |
| Is it copyrightable? | Copyright Office: not on the basis of prompts alone |

Provenance is becoming a legal requirement as well as a trust signal. Article 50 of the EU AI Act requires providers of systems that generate images to ensure outputs “are marked in a machine-readable format and detectable as artificially generated or manipulated.” Google DeepMind describes [SynthID](https://deepmind.google/science/synthid/) as “a tool to watermark and identify content generated through AI.” We infer that a generator with clear, public answers on ownership, indemnification, training data and marking is easier for both buyers and assistants to recommend for business use.

## How does an AI recommendation turn into revenue for an image generator?

Through a free or low-cost first session that becomes a subscription, and for businesses, a plan bought for rights.

We infer two paths:

1. **Creators.** A user asks an assistant for a tool that does a specific style or task, tries one or two named tools, and subscribes if the first results are good. RevenueCat’s data shows how much rides on that first session: across categories, most trial-to-paid conversions happen immediately.
2. **Businesses.** A marketer or designer asks which tools are safe for commercial use, shortlists those with clear terms, and the company buys a team or enterprise plan, sometimes with indemnification or an API. Here the assistant’s description of your licensing can win or lose the deal before anyone visits your site.

Neither path shows up neatly in analytics. A creator who heard of you in ChatGPT often arrives by typing your name. For how to measure it anyway, see [why analytics miss AI visibility](https://underneath.agency/resources/why-analytics-miss-ai-visibility).

## What decides which image generators an AI assistant names?

The platforms say little, but citation studies point to ranked lists, independent reviews and recently published pages.

**Documented by the platforms.** A request like “best image generator for logos” is not searched as typed: [OpenAI explains](https://help.openai.com/en/articles/9237897-chatgpt-search) that ChatGPT search typically rewrites it into “one or more targeted queries,” and [Google reports](https://blog.google/products/search/ai-mode-search/) that AI Mode fans out into “multiple related searches” across subtopics. Neither company explains how a particular generator ends up in the answer.

**Observed in studies.** “Best AI image generator” lists are central to this category, and ranked lists are what assistants cite most: in [Kumar’s](https://arxiv.org/abs/2606.20065) data from an AI visibility platform, the “best-of” listicle was the most-cited content format, at about 21% of all citations. Many such lists are written by vendors. In [our study](https://underneath.agency/research/self-promoting-best-lists-study) of numbered lists cited by AI, 24.2% of those with an identifiable publisher ranked their own publisher first. Outside coverage was the strongest signal in [our brand entity study](https://underneath.agency/research/brand-entity-ai-recommendations-study), where every tenfold increase in independent sites naming a brand came with 4.7 times the odds of being recommended. And assistants that search favor fresh pages: in [our freshness study](https://underneath.agency/research/ai-source-freshness-study), 17.4% to 22.6% of their dated citations were under 90 days old, against 6.9% of Google’s top 10, which suits a category where models change every few months.

**Our inference for image generators.** Assistants answer in text, so a generator’s visual quality reaches them only through what others write: reviews with side-by-side tests, tutorials, creator community threads and comparison articles. Rights and pricing reach them through your own pages and how consistently others repeat them. When those disagree, answers can go wrong; see [how to fix wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers).

## What does GEO involve for an image generation product?

It gets your best styles and your licensing terms stated clearly, repeated consistently and backed by outside reviewers. No one can guarantee an assistant will name you.

In practice, generative engine optimization (GEO) for an image tool comes down to six tasks:

1. **A clear entity.** Say what you are best at (styles, consistency, editing, product shots, illustration), who you are for and where you work (web, app, API, plugins), the same way on your site, app stores, review profiles and social accounts.
2. **A plain-language rights page.** Cover ownership, commercial use by plan, indemnification, training data, content marking and what users may not generate. Keep it current and dated, because these are the questions business buyers ask.
3. **Independent tests.** Earn side-by-side reviews from design publications, creators and testing sites, rather than relying on your own “best of” page. [Why ranked lists drive AI recommendations](https://underneath.agency/resources/best-of-lists-ai-recommendations) covers how those lists get used.
4. **Text that describes images.** Publish galleries and use-case pages with written explanations of style, prompts and settings, so text-based answers have something to quote.
5. **Creator community presence.** Tutorials, community challenges and genuine participation in the forums where creators compare tools.
6. **Measurement.** Run style, use-case, comparison, price and rights prompts through ChatGPT, Gemini, Perplexity, Claude, Copilot and Google at regular intervals, and ask each new subscriber how they found you.

For challengers facing bigger names, see [do AI assistants favor big brands](https://underneath.agency/resources/do-ai-assistants-favor-big-brands).

## What is still unknown about AI search and image generator growth?

We have found no study of how often assistants recommend a specific image generator, or how those users convert.

The figures on image requests inside ChatGPT were produced by researchers at OpenAI, working with Harvard economists. RevenueCat’s benchmarks cover mobile Photo & Video and AI apps broadly, not image generators alone. Individual artists and hobbyists are outside the Content Marketing Institute’s sample, which is B2B marketers only. The legal picture is unsettled: lawsuits are ongoing, the Copyright Office left training data and liability to a separate part of its report, and EU marking rules are only now taking effect. Studies of AI citations cover many categories, and how they apply to image generators is our inference. We have also seen no study of how often assistants suggest their own image tools instead of a specialist.

## How can an image generator tell whether assistants are costing it subscribers and business plans?

Test whether assistants name you for your strongest styles, and whether they get your licensing terms right.

Put your style, use-case, comparison, price and commercial-use prompts to each major assistant and compare what comes back with your subscription and team-plan numbers; that shows whether the gap is with creators, business buyers or both. For help getting your rights terms and independent reviews into those answers, [talk to us about an image generator audit](https://underneath.agency/contact). You can see how we scope that work, including the rights page, independent tests and recurring prompt checks, on our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page.

## Frequently asked questions

### Do AI assistants recommend their own image tools over specialists?

That has not, to our knowledge, been measured. Asked for the best tool in the category, assistants do list outside generators, though no platform documents how the selection is made.

### Can images from AI generators be copyrighted?

In the US, the Copyright Office concluded that prompts alone do not make a user the author, while creative human arrangements or modifications of AI output can be protected. This is general information, not legal advice.

### What makes an image generator “commercially safe”?

There is no single standard. Vendors use the term for different things: licensed training data, ownership terms, indemnification or content marking. Buyers should read the terms, and vendors should state them plainly.

### Should image generators publish comparisons with Midjourney or ChatGPT?

Yes, as long as the comparison is fair. Users ask for exactly these match-ups, and an honest page with real side-by-side examples and clear differences gives assistants something specific to repeat.

## Sources

- Chatterji and colleagues, NBER (2025), [How People Use ChatGPT](https://www.nber.org/system/files/working_papers/w34255/w34255.pdf)
- OpenAI (2025), [Introducing 4o Image Generation](https://openai.com/index/introducing-4o-image-generation/)
- RevenueCat (2025), [State of Subscription Apps 2025](https://www.revenuecat.com/state-of-subscription-apps-2025/)
- Content Marketing Institute (2025), [B2B Content Marketing Benchmarks, Budgets, and Trends](https://contentmarketinginstitute.com/b2b-research/b2b-content-marketing-trends-research)
- Midjourney (2026), [Terms of Service](https://docs.midjourney.com/hc/en-us/articles/32083055291277-Terms-of-Service) and [Comparing Midjourney Plans](https://docs.midjourney.com/hc/en-us/articles/27870484040333-Comparing-Midjourney-Plans)
- Getty Images (2023), [Getty Images Launches Commercially Safe Generative AI Offering](https://newsroom.gettyimages.com/en/getty-images/getty-images-launches-commercially-safe-generative-ai-offering)
- US Copyright Office (2025), [Copyright and Artificial Intelligence, Part 2: Copyrightability](https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-2-Copyrightability-Report.pdf) and [press release](https://www.copyright.gov/newsnet/2025/1060.html)
- EU Artificial Intelligence Act, [Article 50: Transparency obligations](https://artificialintelligenceact.eu/article/50/)
- NPR (2025), [Disney and Universal sue AI firm Midjourney for copyright infringement](https://www.npr.org/2025/06/12/nx-s1-5431684/ai-disney-universal-midjourney-copyright-infringement-lawsuit)
- Google DeepMind (2026), [SynthID](https://deepmind.google/science/synthid/)
- Andreessen Horowitz (2025), [The Top 100 Gen AI Consumer Apps, 5th Edition](https://a16z.com/100-gen-ai-apps-5/)
- OpenAI (2025), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Google (2025), [Expanding AI Overviews and introducing AI Mode](https://blog.google/products/search/ai-mode-search/)
- Kumar (2026), [Generative Engine Optimization at Scale: Measuring Brand Visibility Across AI Search Engines](https://arxiv.org/abs/2606.20065)
- Underneath (2026), [self-ranking best lists](https://underneath.agency/research/self-promoting-best-lists-study), [brand entity](https://underneath.agency/research/brand-entity-ai-recommendations-study) and [source freshness](https://underneath.agency/research/ai-source-freshness-study) studies

---

This is the Markdown twin of https://underneath.agency/resources/ai-image-generators-growth-from-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How much traffic would AI Mode as Google’s default cost you?"
description: "In a 2026 field experiment, making AI Mode Google’s default cut clicks to outside websites by 18.8 points per search. What that means for your traffic."
canonical: "https://underneath.agency/resources/ai-mode-default-traffic-loss"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How much traffic would AI Mode as Google’s default cost you?

On the best evidence so far, a lot. In a preregistered 2026 field experiment, people whose Google searches were all sent to AI Mode for a week made about 2.3 fewer clicks a day to outside websites, against a pre-study average of 4.1. Nobody has measured AI Mode as the real default for a whole market, so treat this as the size of the risk, not a forecast.

## The short version

1. Sending every search to AI Mode cut clicks to outside websites by 18.8 percentage points per search and by about 2.3 clicks per person per day, in a US experiment with 1,100 people (Wang and colleagues, 2026).
2. That loss comes on top of today’s Google, which already includes AI Overviews; hiding AI Overviews in the same experiment raised outside clicks by 8.8 points.
3. AI Mode answered all 400 searches in our test, while AI Overviews appear on 36.0% to 96.0% of searches depending on the industry, so legal, healthcare, local and multi-location businesses would see the most searches newly answered by AI.
4. The loss was broad: fewer people clicked through to news sites, Reddit and Wikipedia, and opinion and advice communities lost most of the extra engagement AI Overviews had sent them.
5. The figures come from a forced one-week switch, a sample younger and more educated than the US population, and an AI Mode with no ads yet. Measure your own exposure before you plan around them.

## What AI Mode is

AI Mode is Google’s conversational search. Powered by Gemini, it answers a search directly and takes follow-up questions. It is separate from AI Overviews, the AI summary at the top of a normal results page.

| Date | Event |
|---|---|
| March 2026 | Wang and colleagues test AI Mode as the default for 1,100 US users |
| August 21, 2025 | AI Mode rolls out globally |
| May 2025 | AI Mode launches, with early access in the United States |
| May 2024 | Google launches AI Overviews |

## What the experiment tested

Researchers at the University of Pennsylvania and Northeastern University ran a preregistered field experiment on Google Search in March 2026 ([Wang, Gleason, Bart, Wilson and Metaxa](https://arxiv.org/abs/2608.18352)). 1,100 US adults who search with Google in Chrome installed a browser extension. After three ordinary days, each was randomly assigned to one of three versions of Google for seven days.

| Group | What changed |
|---|---|
| No AI Search | AI Overviews hidden; AI Mode searches sent to the normal results |
| Current Search | Nothing: Google as it is |
| AI Mode Search | Every search redirected to AI Mode |

The AI Mode group is the closest test yet of AI Mode as the default. 94.7% of that group’s searches reached AI Mode. Before the switch, participants used AI Mode for only 0.6% of their searches, and 36% of their searches showed an AI Overview.

## What happened to clicks

Compared with people on normal Google, the AI Mode group changed as follows.

| Measure | Change with AI Mode as default | 95% interval |
|---|---|---|
| Outside clicks per search (click-through rate) | −18.8 points | −22.2 to −15.3 |
| Outside clicks per person per day | −2.3 | −2.8 to −1.8 |
| Search sessions per day | −0.92 | −1.30 to −0.55 |
| Share of people who clicked a news site | −12.5 points | −18.7 to −6.3 |
| Share who clicked Reddit | −21.2 points | −27.5 to −14.8 |
| Share who clicked Wikipedia | −9.9 points | −14.7 to −5.0 |
| Share who searched on Bing, DuckDuckGo or Yahoo | +11.2 points | +6.4 to +16.0 |
| Minutes per session | +0.43 | +0.20 to +0.66 |

People spent longer inside Google and visited the rest of the web less. They also [trusted what they found less](https://underneath.agency/resources/do-people-trust-ai-search-less): trust fell 0.34 points on a seven-point scale, and satisfaction, usefulness and sense of control all dropped. The fall in daily sessions was larger for heavy Google users.

## What it means for your traffic

The study defines click-through rate as outside clicks divided by Google searches, so −18.8 points means about 19 fewer visits to other websites for every 100 searches. Each follow-up question in an AI Mode conversation counted as a search, which pulls the rate per search down; the per-day count does not have that problem.

The per-day count is the clearest traffic number. Before the switch, participants averaged 4.1 outside clicks and 9.2 searches a day. A drop of about 2.3 clicks a day is roughly 55% of that average, or 43% to 68% across the study’s interval. That share is our arithmetic, not a figure the authors report, and it assumes the comparison group stayed near its pre-study average.

How much of that reaches a given site depends on two things we measured in [our AI Mode and AI Overviews study](https://underneath.agency/research/ai-mode-vs-ai-overviews-study):

- **Coverage.** On 400 US searches run on September 26, 2026, an AI Overview appeared on 63.2%, while AI Mode answered all 400 and cited a source on 94.0%. As the default, AI Mode would also answer the searches that show no AI Overview today.
- **Ranking does not carry over.** Of pages in the organic top 10, AI Overviews cited 29.7% and AI Mode 16.3%; for positions 1 to 3, 42.3% against 26.0%. On the 147 searches with no AI Overview, 69.3% of AI Mode’s citations were Google’s own pages.

Being cited also does not mean being visited. A [Pew Research Center panel](https://arxiv.org/abs/2608.04831) found users clicked a source cited in an AI Overview on 1% of visits to pages that showed one. For chat-style tools beyond Google, see [how answer engines change website clicks](https://underneath.agency/resources/ai-answer-engines-fewer-website-clicks).

## Which industries have the most new exposure

AI Mode answered all 400 of our test searches, 50 in each of eight industries. AI Overviews do not appear that evenly. In our [AI Overview frequency study](https://underneath.agency/research/ai-overviews-frequency-study) of 800 US commercial keywords, searched on September 26, 2026, the share that showed an AI Overview ranged from 96.0% to 36.0% by industry. The lower the share today, the more of an industry’s searches AI Mode as the default would newly answer with AI.

| Industry | Keywords with an AI Overview today | 95% interval |
|---|---|---|
| B2B software and technology | 96.0% | 85.9% to 100.0% |
| Financial services and insurance | 88.0% | 79.5% to 92.5% |
| Retail and ecommerce | 68.0% | 39.5% to 85.5% |
| Hospitality and travel | 62.0% | 41.4% to 90.9% |
| Home and local services | 53.0% | 34.0% to 69.8% |
| Healthcare and dental | 43.0% | 26.9% to 75.8% |
| Legal and professional services | 40.0% | 25.9% to 56.6% |
| Franchises and multi-location brands | 36.0% | 18.6% to 64.6% |

The biggest driver is the local pack, the map with three businesses. Holding industry, intent and other factors constant, a search with a local pack had a 30.7% predicted chance of an AI Overview, against 78.6% without one. Local and multi-location businesses are among the most sheltered from AI answers today, so they would see the largest share of their searches newly answered by AI. Shares are by keyword, not weighted by search volume.

## Who loses most

Every kind of site the experiment tracked lost visitors: fewer people clicked through to news sites, Reddit and Wikipedia. The authors conclude that the impact “generalizes well beyond Wikipedia and news.”

Even people who knew where they were going had trouble. Asked about their week with AI Mode, 15.3% of participants said it was harder to reach a specific website, and 13.4% mentioned fewer links or less source diversity.

Opinion and advice content loses a cushion it had. [Zhang, Cui and Zhang](https://arxiv.org/abs/2605.16428) found that AI Overviews raised daily comments by 12.0% in Reddit communities they can cite, relative to communities they cannot, mostly in discussions of advice and personal experience. After AI Mode launched globally on August 21, 2025, that extra engagement largely disappeared: for comments it went from 9.54 a day to −0.17, and for comment authors it fell by 59%. The authors read this as a conversational interface absorbing the depth that a static summary used to send onward.

## What Google says

Google disputes that its AI features cost websites traffic. As the experiment’s authors summarize it, Google argues that AI features are complements that have kept overall referral traffic steady, improved click quality and opened new chances to be surfaced through longer, more complex queries. Google has also said that AI features increase search engagement.

The experiment tested part of this. [Hiding AI Overviews raised outside clicks](https://underneath.agency/resources/do-ai-overviews-reduce-clicks), AI Mode reduced them, and AI Mode cut daily search sessions by 0.92 rather than adding to them. Click quality was not measured, there or in any other study cited here, so that claim remains open.

## What it means for paid search

The share of people who clicked an ad fell by 42.7 points in the AI Mode group, but only because AI Mode showed no ads at the time. The authors call their results a pre-monetization baseline and expect satisfaction and other measures to shift once Google adds advertising, sponsored results or other mechanisms. For a paid search budget, how AI Mode carries ads is the open question: this experiment cannot say whether paid clicks would hold up.

## What the evidence does not tell you yet

- **Not a real default.** The experiment forced every search into AI Mode for seven days. The authors note that gradual, opt-in adoption may behave differently, and that people may adapt over a longer period.
- **Not a representative sample.** Participants skewed younger, more educated and more left-leaning than the US population.
- **Before ads.** AI Mode carried no ads during the study, and the authors expect satisfaction and other measures to shift once Google monetizes it.
- **Not by industry.** The study measured everyday searches together, not your category, and did not separate high-stakes topics such as health, legal or financial questions.
- **One market, one engine.** US users on Google in March 2026.

## How to estimate your own exposure

1. Export 90 days of Google organic clicks by query from Search Console.
2. Mark the queries that show an AI Overview today; part of the loss is already priced into those.
3. For the rest, assume an AI answer will appear: AI Mode answered every search in our test.
4. Check whether AI Mode cites your pages on your most valuable queries, and repeat the check: in [our tests](https://underneath.agency/research/ai-overviews-frequency-study), when an AI Overview appeared on two days, the URLs it cited two days apart had a median overlap of only 0.43.
5. Apply the study’s range as a stress test, not a forecast: about 19 fewer outside clicks per 100 searches, or roughly half of daily outside clicks, and see what that does to pipeline.
6. Treat queries where people want to land on your site, such as brand, login or pricing searches, separately; in the study, people found it harder to reach a specific website in AI Mode.

If the stress test hurts, the work is to become a source that AI answers use and name, alongside protecting the demand that comes to you directly. That is what [generative engine optimization](https://underneath.agency/services/generative-engine-optimization) is for.

## Frequently asked questions

### Will Google make AI Mode the default?

None of the research cited here says it will. The experiment’s authors describe Google as positioning AI Mode “as the future of search,” and AI Mode rolled out globally on August 21, 2025. This guide treats AI Mode as the default as a scenario to plan for, not a prediction.

### How much traffic would we lose if AI Mode became the default?

In the only randomized test so far, outside clicks fell by 18.8 percentage points per search and by about 2.3 clicks per person per day, against a pre-study average of 4.1. Your own loss depends on how many of your queries already show AI Overviews and whether AI Mode cites you.

### Does being cited in AI Mode make up for lost clicks?

Not on current evidence. Pew Research Center found users clicked a source cited in an AI Overview on 1% of visits to pages that showed one, and a study of Wikipedia found that traffic from AI Overview citations did not offset the decline AI Overviews caused.

### Are AI Overviews already costing us traffic?

Probably some. In the same experiment, hiding AI Overviews raised outside clicks by 8.8 points per search. Pew found that users clicked a search result on 8% of visits to pages with an AI Overview, against 15% on pages without one.

### Which industries would be hit hardest if AI Mode became the default?

The ones AI Overviews reach least today. In our sample, AI Overviews appeared on 36.0% of franchise and multi-location keywords, 40.0% of legal and 43.0% of healthcare keywords, against 96.0% in B2B software. AI Mode answered every search, so those industries would see the most searches newly answered by AI.

### Which websites lose the most in AI Mode?

The experiment found losses across news sites, Reddit and Wikipedia, and the authors expect the effect to reach well beyond them. Research on Reddit found that advice and opinion communities lost most of the extra engagement AI Overviews had given them once AI Mode launched.

## Sources

- Wang, Gleason, Bart, Wilson and Metaxa (2026), [AI in Search Reduces Publisher Referrals Without Improving User Experience: Experimental Evidence](https://arxiv.org/abs/2608.18352), arXiv:2608.18352.
- Zhang, Cui and Zhang (2026), [The Impact of AI Search on the Online Content Ecosystem: Evidence from Google and Reddit](https://arxiv.org/abs/2605.16428), arXiv:2605.16428.
- Chapekis, Lieb, Shah and Smith, Pew Research Center (2026), [Investigating Click Behaviors On Google Search Result Pages That Produce an AI Overview](https://arxiv.org/abs/2608.04831), arXiv:2608.04831.
- Khosravi and Yoganarasimhan (2026), [Impact of AI Search Summaries on Website Traffic: Evidence from Google AI Overviews and Wikipedia](https://arxiv.org/abs/2602.18455), arXiv:2602.18455.
- Underneath (2026), [AI Mode vs AI Overviews: how different are the sources?](https://underneath.agency/research/ai-mode-vs-ai-overviews-study)
- Underneath (2026), [When does Google show an AI Overview? 1,248 US searches](https://underneath.agency/research/ai-overviews-frequency-study)

---

This is the Markdown twin of https://underneath.agency/resources/ai-mode-default-traffic-loss. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "If AI Overviews cite my site, does that make up for lost clicks?"
description: "Not on current evidence. Only 1% of visits to AI Overview pages led to a click on a cited source, and Wikipedia lost traffic despite heavy citation."
canonical: "https://underneath.agency/resources/ai-overview-citations-vs-lost-clicks"
published: 2026-10-08
updated: 2026-10-10
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# If AI Overviews cite my site, does that make up for lost clicks?

Not on current evidence. In a Pew panel of US adults, just 1% of visits to Google pages with an AI Overview ended in a click on a cited source, and Wikipedia, one of the most cited sources of all, still lost search traffic. A citation can be worth having, but it is a different thing from a visit, and the two should be measured separately.

## The short version

1. Just 1% of visits to Google pages with an AI Overview led to a click on a cited source, in Pew’s March 2025 browsing panel of 900 US adults.
2. Wikipedia, among the sources AI Overviews cite most, lost about 5% of its English search traffic, and the study’s authors found citations did not offset the decline.
3. An AI Overview cites a median of 8 sources, but a desktop screen shows only about three before the reader clicks “Show all”.
4. 29.8% of the domains AI Overviews cite do not appear anywhere on the first page of results, so a citation and a ranking are separate wins (Xu and colleagues, 2026).
5. Discussion content is the exception: Reddit communities that AI Overviews can cite gained 12.0% more daily comments.

## How often do people click the sources in an AI Overview?

Rarely: about 1% of visits to an AI Overview page led to a click on a cited source.

[Chapekis, Lieb, Shah and Smith](https://arxiv.org/abs/2608.04831) at Pew Research Center tracked the browsing of 900 US adults through March 2025. For every Google search with an AI Overview, they checked whether the next page the person visited was one of the sources the summary cited. Just 1% of visits ended that way.

Compare that with [the clicks AI Overviews take away](https://underneath.agency/resources/do-ai-overviews-reduce-clicks). On pages without an AI Overview, people clicked a regular result on 15% of visits. On pages with one, they clicked a regular result on 8% of visits. Adding the 1% of citation clicks to those 8% still leaves the total well short of the 15% seen without a summary.

| What happened next | Page with an AI Overview | Page without one |
|---|---|---|
| Clicked a regular search result | 8% | 15% |
| Clicked a source cited in the summary | 1% | not applicable |
| Ended the browsing session | 26% | 16% |

Pew matched only the first three cited links, which is what a desktop screen showed at the time. Clicks on sources further down the list would not be counted, so the true figure may be slightly higher. We look at the last row of the table in [whether AI Overviews end browsing sessions](https://underneath.agency/resources/ai-overviews-end-browsing-sessions).

## Did being cited protect Wikipedia’s traffic?

No: Wikipedia is heavily cited by AI Overviews and still lost search traffic after they launched.

[Khosravi and Yoganarasimhan](https://arxiv.org/abs/2602.18455) compared English Wikipedia articles with the same articles in German and French during 2024, when only US users had AI Overviews by default. English search traffic fell about 5.45% relative to German and 4.82% relative to French: roughly 100.27 million fewer search visits a month. It is a clear case of [Google design changes moving traffic](https://underneath.agency/resources/google-search-design-changes-traffic-risk) that publishers cannot control.

The authors address citations directly: they may send some users to Wikipedia, “but our estimates indicate that any resulting traffic did not offset the average decline.” Wikipedia is cited unusually often. [Huang and colleagues](https://arxiv.org/abs/2603.16138) found Wikipedia cited for 28% of the queries they ran through AI Overviews, more than any other domain. A business site cited far less often has less reason to expect a better result.

Wikipedia carries no ads, so its loss is attention, not income. To show the stakes for others, the authors estimate that an ad-supported site losing the same traffic would forgo roughly $0.901M to $3.090M of ad revenue a month.

## Why do citations send so few visitors?

Most citations sit out of sight, and the summary has often answered the question already.

[Xu, Iqbal and Montgomery](https://arxiv.org/abs/2605.14021) collected 7,583 AI Overviews in spring 2026. They cited a median of 8 sources each, with up to 32. But Google shows only the first three on a desktop screen, and on a phone the sources may not appear until the reader taps “Show more”, as the Pew team notes. Most cited sites are therefore listed where few people look.

Being listed is also not the same as supplying the answer. Huang and colleagues measured how much of each cited page made it into the summary. In the answers they analyzed, AI Overviews cited nearly twice as many sources as ChatGPT search, 5.0 against 2.7 on average, but drew less content from each. Forum pages such as Reddit and Quora were cited, yet were under-represented in the summaries by 22.1 percentage points.

Placement among the citations matters too. In [our study of AI Overview citations](https://underneath.agency/research/ai-overview-citations-study), the first organic result was cited in 49.5% of AI Overviews. YouTube was cited in 64.2%, mostly with videos that did not rank on page one.

## Is a citation worth anything without the click?

Possibly, for visibility and credibility, but no study we reviewed measures its business value.

A citation can put a site in front of searchers it would not otherwise reach. Xu and colleagues found that 29.8% of the domains AI Overviews cited did not appear anywhere on the first page of results. In our own study, only 28.7% of AI Overview citations were page-one organic results for the same search.

Citations also shape trust. In a randomized experiment, [Li and Aral](https://arxiv.org/abs/2504.06435) found that adding reference links made people trust AI search answers more, even when the links were wrong. That is a reason to care what the summary says about you, not evidence that a citation brings customers.

Accuracy is a real issue. Xu and colleagues checked 98,020 individual claims in AI Overviews against the pages cited for them, and 11.0% were not supported by those pages. Being cited does not guarantee being quoted correctly.

## Which sites do gain from being surfaced?

Opinion, advice and experience content can gain; factual content mostly loses.

[Zhang, Cui and Zhang](https://arxiv.org/abs/2605.16428) compared Reddit communities that AI Overviews can cite with adult communities Google keeps out of them. After AI Overviews expanded in August 2024, the citable communities gained 12.0% more daily comments and 12.4% more daily commenters. The effect was 2.3 to 2.8 times larger for experience-based communities than for fact-based ones.

The authors’ explanation: a summary cannot replace a discussion, so it works like an advertisement for it. A summary can replace a fact. The gain is also fragile. After Google’s conversational AI Mode rolled out globally in 2025, the extra gain for experience communities fell by 59% in commenters and vanished for comments.

## What should you do about it?

Track citations and clicks as two separate numbers, and do not count one as a substitute for the other.

1. Report AI Overview citations as a visibility measure, not a traffic measure. On the Pew evidence, about 1 visit in 100 to an AI Overview page becomes a click on a cited source.
2. Keep traffic forecasts honest. If informational searches now show AI Overviews, plan for fewer clicks even if your citation count rises.
3. Check what AI Overviews say about your business, since 11.0% of claims in one large study were not supported by the cited pages.
4. Invest in [content a summary cannot replace](https://underneath.agency/resources/content-business-risk-from-ai-search): firsthand experience, tools, data and discussion that people need to see in full.
5. Watch AI Mode, which removed much of the gain discussion communities had won from AI Overviews.

If you want help tracking both numbers, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The research shows citations send few clicks, but not what else a citation is worth.

- The Pew data is from March 2025 and matched only the first three cited links on desktop.
- No study we reviewed measures whether a citation leads to later brand searches, direct visits or sales.
- The Wikipedia result covers one nonprofit publisher in 2024.
- The Reddit result measures comments and commenters, not visits to commercial sites.
- No study tests whether appearing among the first three citations earns more clicks than appearing lower.

## Frequently asked questions

### Do people click on links in Google AI Overviews?

Rarely. In Pew’s March 2025 panel of 900 US adults, just 1% of visits to pages with an AI Overview led to a click on a cited source.

### Is being cited in an AI Overview the same as ranking on Google?

No. In our study, only 28.7% of AI Overview citations were page-one organic results for the same search, so the two are related but separate.

### How many sources does an AI Overview cite?

A median of 8, in a 2026 study of 7,583 AI Overviews by Xu and colleagues. A desktop screen shows about three of them before the reader expands the list.

### Should I stop Google from using my pages in AI Overviews?

Only with care. During spring 2026, the main controls that kept pages out of AI Overviews, such as “nosnippet”, also affected normal search results, according to Xu and colleagues.

## Sources

- Chapekis, Lieb, Shah and Smith (2026), [Investigating Click Behaviors On Google Search Result Pages That Produce an AI Overview](https://arxiv.org/abs/2608.04831), arXiv:2608.04831.
- Khosravi and Yoganarasimhan (2026), [Impact of AI Search Summaries on Website Traffic: Evidence from Google AI Overviews and Wikipedia](https://arxiv.org/abs/2602.18455), arXiv:2602.18455.
- Xu, Iqbal and Montgomery (2026), [Measuring Google AI Overviews: Activation, Source Quality, Claim Fidelity, and Publisher Impact](https://arxiv.org/abs/2605.14021), arXiv:2605.14021.
- Huang and colleagues (2026), [Answer Bubbles: Information Exposure in AI-Mediated Search](https://arxiv.org/abs/2603.16138), arXiv:2603.16138.
- Li and Aral (2025), [Human Trust in AI Search: A Large-Scale Experiment](https://arxiv.org/abs/2504.06435), arXiv:2504.06435.
- Zhang, Cui and Zhang (2026), [The Impact of AI Search on the Online Content Ecosystem: Evidence from Google and Reddit](https://arxiv.org/abs/2605.16428), arXiv:2605.16428.
- Underneath (2026), [Do AI Overviews cite the pages that rank? 4,051 citations](https://underneath.agency/research/ai-overview-citations-study)

---

This is the Markdown twin of https://underneath.agency/resources/ai-overview-citations-vs-lost-clicks. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Do people stop browsing after seeing an AI Overview?"
description: "Yes. In a Pew panel of 900 US adults, browsing ended on 26% of Google pages with an AI Overview, against 16% of pages without one."
canonical: "https://underneath.agency/resources/ai-overviews-end-browsing-sessions"
published: 2026-10-08
updated: 2026-10-10
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Are people more likely to stop browsing after seeing an AI Overview?

Yes, on the evidence so far. In a March 2025 Pew Research Center panel of 900 US adults, browsing ended on 26% of Google pages with an AI Overview and 16% of pages without one. That is an association from real browsing data, but a field experiment and a traffic study point the same way.

## The short version

1. Browsing ended on 26% of Google pages with an AI Overview and 16% without, across a month of browsing by 900 US adults (Pew Research Center).
2. The same people clicked a search result on 8% of AI Overview pages against 15% of other pages. They clicked a source cited in the AI Overview on just 1%.
3. In a 2026 US experiment with 1,100 people, hiding AI Overviews raised clicks to outside websites by 8.8 percentage points.
4. AI Overview rates depend on the searches counted: 18% of the Pew panel’s searches, 13.7% of trending searches elsewhere, 60.8% of commercial keywords in ours.
5. Nobody has shown whether people who stop after an AI Overview were satisfied, or simply stopped.

## Do searches end more often when an AI Overview appears?

Yes: in Pew’s panel, browsing ended after 26% of AI Overview pages, against 16% of other pages. An AI Overview is the AI-written summary Google shows at the top of many results pages. [Chapekis, Lieb, Shah and Smith](https://arxiv.org/abs/2608.04831), researchers at Pew Research Center, studied a month of real browsing from March 1 to 31, 2025.

The 900 US adults came from a panel recruited by address-based sampling, chosen to be as representative of the country as possible. Over the month they made 2,457,176 web page visits. 91,121 of those visits were to Google search pages, covering 68,879 different searches.

For each Google page, the researchers recorded what the person did next.

| Next action after a Google results page | With an AI Overview | Without one |
|---|---|---|
| Ended the browsing session | 26% | 16% |
| Clicked a search result | 8% | 15% |
| Clicked a source cited in the AI Overview | 1% | Not applicable |

## How did the researchers define stopping?

A session counted as ended when the person left the web browser for five seconds or more. Other next steps were a click on an AI Overview source, a click on a search result, another Google search, or a visit elsewhere.

Clicks on AI Overview sources counted only the first three cited links. In April 2025, three was the most a desktop browser showed without the user asking for more.

Five seconds is a short pause. Some people who “ended” a session may have returned minutes later. The study also cannot see what they did elsewhere, such as in an app.

## Is the AI Overview the cause, or the kind of search?

Probably partly the AI Overview itself: an experiment and a traffic study both point that way.

The worry is fair. Pew found that longer searches, questions and searches with both a noun and a verb were [more likely to show an AI Overview](https://underneath.agency/resources/which-searches-trigger-ai-overviews). Those searches might end sooner anyway.

Pew’s researchers adjusted for those features of each search and for each person’s own habits, and the gaps held. Their data is still observational, so it cannot prove cause.

Stronger evidence comes from a [field experiment by Wang and colleagues](https://arxiv.org/abs/2608.18352) with 1,100 US Google users in March 2026. People whose AI Overviews were hidden [clicked through to outside websites](https://underneath.agency/resources/do-ai-overviews-reduce-clicks) 8.8 percentage points more often. A Google page change meant the tool hid only 51.1% of AI Overviews, and the researchers adjusted their estimate for that.

Those people’s search sessions also became shorter, by 0.59 minutes. A [study of Wikipedia traffic](https://arxiv.org/abs/2602.18455) adds a real-world estimate. After AI Overviews launched in the US in 2024, English Wikipedia’s search visits fell 5.45% relative to German editions.

## How often will your buyers see an AI Overview?

It depends heavily on the search: published rates range from about one search in seven to nearly all of them. The differences come from which searches each study counted.

| Study | Searches counted | Showed an AI Overview |
|---|---|---|
| Pew panel, March 2025 | Real searches by 900 US adults | 18% |
| [Xu and colleagues](https://arxiv.org/abs/2605.14021), March to April 2026 | 55,393 trending searches | 13.7% |
| [Grossman and colleagues](https://arxiv.org/abs/2604.27790), December 2025 | Representative real-user searches | 51.5% |
| [Our frequency study](https://underneath.agency/research/ai-overviews-frequency-study), September 2026 | 800 US commercial keywords | 60.8% |

Questions are the clearest signal. Xu and colleagues found AI Overviews on 64.7% of trending searches phrased as questions, against 9.5% of the rest. In our study, 98.1% of commercial keywords phrased as questions showed one.

Those are the research-style searches where buyers often first meet a brand.

## Does stopping mean the searcher got what they needed?

Unknown: the data shows people stopped, not whether they were satisfied. Researchers call a search that ends because the answer was on the page “good abandonment.” Neither study separates that from giving up.

The experiment offers a clue. Hiding AI Overviews produced no detectable change in how much people trusted information on Google or rated their search experience. The authors read this as a sign that any benefits of AI Overviews do not show up in those measures.

The answers people stop at are not always fully backed by their sources. Across 98,020 statements in AI Overviews, Xu and colleagues found 11.0% were not supported by the pages cited. Most were points the source did not mention at all.

Some of that may come from social media pages they could not read; even so, they put the floor at roughly 5.3%. In [a small pilot of our own](https://underneath.agency/research/ai-overviews-frequency-study), judged by AI coders rather than people, 12.7% of 71 cited sentences were not supported by the page cited. Someone who stops there leaves with that answer, and your page gets no chance to correct it.

## What should you do about it?

Assume fewer searches will reach your site, and work to be the source the AI Overview names. Practical steps:

1. Find your exposed searches. Check which of your important searches show an AI Overview, starting with question-style and longer searches.
2. Aim to be cited, not only ranked. In [our citation study](https://underneath.agency/research/ai-overview-citations-study), the first organic result was cited in 49.5% of AI Overviews and the ninth in 15.5%. Ranking helps but does not guarantee a mention.
3. State key facts plainly. If the summary is all people read, your prices, features and claims need to be easy to lift accurately.
4. Measure more than clicks. Track how often your brand appears in AI Overviews, alongside branded searches and direct visits.
5. Watch AI Mode. Google’s conversational search cut outside clicks further in the same experiment, as we explain in [our article on AI Mode and traffic](https://underneath.agency/resources/ai-mode-default-traffic-loss).

If you want help finding which of your searches end on Google’s page, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

Only one study has measured session endings directly, so several questions remain open. The main gaps:

- Breadth. The direct evidence is one US panel over one month in early 2025, and AI Overviews have changed since.
- Cause. The Pew findings are associations; the experiment measured clicks and session length, not session endings.
- What happens next. A five-second exit may hide a return visit, a phone call or a switch to another app.
- Business searches. None of the studies reports session endings for B2B or high-value purchase searches.
- Satisfaction. No study has asked people who stopped whether they got what they needed.

## Frequently asked questions

### Do AI Overviews reduce clicks to websites?

Yes. In Pew’s panel, people clicked a result on 8% of AI Overview pages against 15% of others. In an experiment, hiding AI Overviews raised outside clicks by 8.8 percentage points.

### Do people click the links inside AI Overviews?

Rarely. In Pew’s panel of 900 US adults, just 1% of visits to AI Overview pages led to a click on a cited source.

### Why do some searches show an AI Overview and others don’t?

Longer searches and questions trigger them most. One 2026 study of trending searches found AI Overviews on 64.7% of questions against 9.5% of other searches.

### Is a search that ends on Google bad for my brand?

Not if your brand is named in the answer. In our study, the top organic result was cited in 49.5% of AI Overviews, so ranking well still helps you appear where the search ends.

## Sources

- Chapekis, A., Lieb, A., Shah, S. and Smith, A. (2026), [Investigating Click Behaviors On Google Search Result Pages That Produce an AI Overview](https://arxiv.org/abs/2608.04831), arXiv:2608.04831.
- Wang, S. T., Gleason, J., Bart, Y., Wilson, C. and Metaxa, D. (2026), [AI in Search Reduces Publisher Referrals Without Improving User Experience: Experimental Evidence](https://arxiv.org/abs/2608.18352), arXiv:2608.18352.
- Khosravi, M. and Yoganarasimhan, H. (2026), [Impact of AI Search Summaries on Website Traffic: Evidence from Google AI Overviews and Wikipedia](https://arxiv.org/abs/2602.18455), arXiv:2602.18455.
- Xu, H., Iqbal, U. and Montgomery, J. M. (2026), [Measuring Google AI Overviews: Activation, Source Quality, Claim Fidelity, and Publisher Impact](https://arxiv.org/abs/2605.14021), arXiv:2605.14021.
- Grossman, R., Liu, S., Chen, M. K., Smith, M., Borcea, C. and Chen, Y. (2026), [How Generative AI Disrupts Search: An Empirical Study of Google Search, Gemini, and AI Overviews](https://arxiv.org/abs/2604.27790), arXiv:2604.27790.
- Underneath (2026), [When does Google show an AI Overview? 1,248 US searches](https://underneath.agency/research/ai-overviews-frequency-study)
- Underneath (2026), [Do AI Overviews cite the pages that rank? 4,051 citations](https://underneath.agency/research/ai-overview-citations-study)

---

This is the Markdown twin of https://underneath.agency/resources/ai-overviews-end-browsing-sessions. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How do AI research tools get found and trusted through AI search?"
description: "By being named, with checkable proof of citation accuracy, when students, researchers and librarians ask AI which research assistant to trust."
canonical: "https://underneath.agency/resources/ai-research-tools-users-ai-search"
published: 2026-10-09
updated: 2026-10-10
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How do AI research tools get found and trusted through AI search?

By being the research assistant an AI answer names when a student, researcher or librarian asks which tool to use, and by publishing the evidence on citation accuracy, sources and data use that this audience checks before trusting anything. AI research tools face an unusual problem: the general assistants that recommend them, such as ChatGPT and Gemini with their own deep research modes, are also their closest substitutes. The tools that win make the difference easy to verify.

## The short version

1. Researchers have moved fast. In [Elsevier’s 2025 survey](https://www.elsevier.com/insights/confidence-in-research/researcher-of-the-future) of more than 3,200 researchers, 58% use AI tools in their work, up from 37% in 2024, and 61% use them to find and summarize the latest research.
2. Trust has not kept up: only 22% of those researchers believe AI tools are currently trustworthy, and 59% name transparency and clear citations as the marker that would build confidence.
3. Citation accuracy is the open wound. In a [2024 test for systematic reviews](https://pmc.ncbi.nlm.nih.gov/articles/PMC11153973/), 28.6% of the references GPT-4 produced were hallucinated; the [Tow Center](https://www.cjr.org/tow_center/we-compared-eight-ai-search-engines-theyre-all-bad-at-citing-news.php) found eight AI search tools gave incorrect answers to more than 60% of queries asking them to identify the source of news excerpts.
4. The specialist tools have real scale and simple prices: [Consensus](https://consensus.app/home/about-us/) reports more than 10 million users from more than 12,500 universities, and [Elicit’s](https://elicit.com/pricing) paid plans run from $11 to $89 per user per month before custom plans for companies and schools.
5. Institutions are next. [Clarivate’s 2026 survey](https://clarivate.com/pulse-of-the-library/) of 1,876 library responses found 46% of academic libraries at some stage of AI implementation, while students’ AI use for assignments rose from 45% in 2024 to 95% in 2026.

## Who uses AI research tools, and what is a user worth?

Students, academic researchers, professionals who review evidence, and the libraries and companies that license tools for them.

**Students** are the largest group and the fastest adopters. Clarivate’s user research found AI use for school assignments rising from 45% to 95% of students surveyed in two years. They mostly start free.

**Academic researchers** use AI for literature work. Elsevier found researchers use AI tools to find and summarize research (61%), perform literature reviews (51%) and draft grant proposals (41%). [Oxford University Press](https://corp.oup.com/news/how-are-researchers-responding-to-ai/), in a survey of over 2,000 researchers, found 76% had used some form of AI tool in their research.

**Professionals** in fields such as pharma, healthcare and policy run systematic reviews and evidence summaries, often with budgets. Elicit, for example, sells a Pro plan “for systematic reviews” and says researchers report up to 80% time savings on systematic reviews with its product. Biotech partnering is one place this shows: scientists vetting a potential partner use AI tools for literature reviews, as [how AI search shapes biotech partnering](https://underneath.agency/resources/biotech-partnering-ai-search) explains.

**Institutions** license tools for everyone. Elicit’s Enterprise plan is described as “for companies & schools,” [scite](https://scite.ai/pricing) sells a team plan at $250 a month that includes 8 seats, and Elsevier’s [Scopus AI](https://www.elsevier.com/products/scopus/scopus-ai) is offered to institutions alongside a paid Scopus license.

What a user is worth follows that ladder. A student on a free plan is worth little today and may be worth a lot later. A researcher on Elicit’s Pro plan pays $39 per user per month billed annually. A lab, a company or a university library that licenses a tool for hundreds of people is a contract that renews each year if the tool earns its place.

## Where does AI search sit in how researchers find tools?

Near the start: researchers already ask AI about the literature, so asking which tool to use is a short step.

No survey we found asks researchers how they discovered the research assistant they use. The surveys do show where they look for guidance: OUP found 54% would look to academic societies for guidance on AI, 43% to their own institution and 27% to publishers. Those are the kinds of independent, trusted sources that AI answers also draw on.

The bigger shift is that general assistants now do research themselves. [OpenAI](https://openai.com/index/introducing-deep-research/) says deep research in ChatGPT can take 5 to 30 minutes to find and synthesize hundreds of online sources into a report, and [Google](https://blog.google/products/gemini/google-gemini-deep-research/) offers a similar Deep Research mode in Gemini. For a specialist research tool, the assistant is both the channel (where a user may ask “what is the best AI tool for a literature review?”) and the alternative (where the same user may simply ask the literature question directly). Writing tools face the same double role; see [how AI writing tools compete with ChatGPT](https://underneath.agency/resources/ai-writing-tools-users-from-ai-search).

Students and researchers are also spread across several assistants. By [Similarweb’s 2026 count](https://aisearch.similarweb.com/blog/gen-ai-stats/), ChatGPT’s slice of visits to generative AI websites slid from about 76% in June 2025 to roughly 53% by May 2026, as Gemini and Claude picked up share. Libraries are also moving content into those assistants: Clarivate found 19% of libraries overall pursue a balanced approach that includes making library content and services available directly in external services such as ChatGPT, Gemini or Claude.

## Which questions do people ask about AI research tools?

Questions about accuracy, fit and access: which tool cites real papers, suits a method, and is licensed.

These sample prompts are our own, written to show how students, researchers and librarians might phrase the question; they are not recorded queries:

- Task fit: “Best AI tool for screening papers for a systematic review in medicine.”
- Comparison: “Elicit vs Consensus vs Scite for a PhD literature review.”
- Accuracy: “Which AI research assistant doesn’t make up citations?”
- Substitution: “Is ChatGPT deep research good enough for a literature review, or do I need a specialist tool?”
- Access: “Does my university have a license for an AI research assistant?”
- Data: “Which AI research tools don’t train on my unpublished manuscripts?”

The accuracy and substitution questions decide the category. If a user believes a general assistant is good enough, a specialist tool loses the sale before it is considered. For how AI answers handle unsupported statements generally, see [our article on unsupported claims](https://underneath.agency/resources/ai-answers-unsupported-claims).

## How does an AI recommendation turn into users and licenses?

Through a free account that becomes a habit, then a paid plan, then an institutional license that renews.

**Named.** An answer to a task or comparison question names your tool. Because research users verify by habit, they often open several tools to compare.

**Free account.** Elicit and Consensus both let users start free. The test is whether results hold up when the user checks the cited papers.

**Paid individual or team plan.** Heavier work, such as systematic reviews, pushes users to paid tiers. Elicit’s plans are Plus ($11), Pro ($39) and Scale ($89) per user per month, billed annually; scite’s Basic plan is $20 a month.

**Institutional license.** Libraries and companies buy for whole groups, with security, data and training controls. Elicit lists “No training on your data by default” as an Enterprise feature. This step can take a long time: Clarivate found 33% of libraries still in the exploration and evaluation stage of AI.

We infer that AI visibility matters most at the first and last steps: it brings in individual users, and it shapes the reputation that a librarian or research office checks before signing a license.

## Why would an assistant send a researcher to one specialist tool over another?

Platforms document little; studies point to independent sources; researchers add citation accuracy, coverage and data use.

**Documented by the platform.** ChatGPT, Gemini, Claude and the rest say nothing public about how they pick a literature-review or citation tool to suggest. OpenAI and Google document what their own research modes do, which is the competition specialist tools must answer.

**Observed in studies.** In a test of 112 Product Hunt startups, [Sharma](https://arxiv.org/abs/2601.00912) found ChatGPT recognized 99.4% when asked by name but surfaced only 3.32% in discovery questions, and that the number of other sites linking to a product and community discussion predicted visibility in Perplexity. On software ranking questions more broadly, AI search took 72.7% of its sources from independent “earned” sites in work by [Chen and colleagues](https://arxiv.org/abs/2509.08919), against 45.4% for Google. We infer that coverage in academic, library and scholarly-publishing sources counts for more than claims on a vendor’s own site. Our study of [Wikipedia, outside coverage and AI recommendations](https://underneath.agency/research/brand-entity-ai-recommendations-study) tests that idea for brands in general.

**Trust factors specific to research tools.** This audience is trained to check sources, and the record gives it reasons:

- **Citation accuracy.** In the 2024 systematic review test, published in the Journal of Medical Internet Research, GPT-3.5 hallucinated 39.6% of references, GPT-4 28.6% and Bard 91.4%. Those models are now old, but the finding shaped how researchers judge every AI tool.
- **Source attribution.** The Tow Center’s test of eight AI search tools found error rates ranging from 37% for Perplexity to 94% for Grok 3 when asked to identify the source of news excerpts.
- **Data use.** OUP found only 8% of researchers trust AI companies not to use their research data without permission, and only 6% trust them to meet their privacy and security needs.
- **What researchers ask for.** Elsevier found the trust markers researchers want are transparency and clear citations (59%), recency and up-to-date literature (55%) and training on high-quality peer-reviewed content (55%).

A specialist tool’s advantage is that it can show its corpus, its citation method and its accuracy tests in public, in a way a general assistant usually does not. Elicit, for example, states that it searches over 138 million academic papers. Our article on [whether citations make AI answers more trusted](https://underneath.agency/resources/do-citations-make-ai-answers-more-trusted) explains why showing sources alone is not proof.

## What does it cost a research tool to be left out?

The first account, and the reputation that later decides an institutional license.

When the question is “which tool should I use for my literature review?” and the answer lists two competitors and a general assistant’s research mode, the user rarely goes looking for a fourth option. The scale of the leaders shows how much is at stake: Consensus reports more than 10 million users, and Elicit says it is trusted by over 5 million researchers. A tool that is missing from answers also misses the word of mouth that follows, because students and researchers recommend tools to each other.

The license stage magnifies the loss. Institutional decisions are slow and deliberate, and a reasonable expectation is that a librarian evaluating options will look at what is said about each tool, including in AI answers. That last point is our inference; no study has measured how libraries use AI answers in procurement.

## How does GEO work for an AI research tool?

It makes your tool easy for assistants to describe accurately and for researchers to verify; it cannot promise a recommendation.

For a research assistant aimed at scholars, generative engine optimization (GEO) tends to involve seven pieces of work:

1. **A clear statement of what you search.** Name your corpus, its size and update schedule, the fields it covers and what it leaves out, on a public page assistants can cite.
2. **Published accuracy evidence.** Describe how you check citations and how often references are wrong, with method and date. Specific, repeatable tests are what this audience trusts.
3. **Independent coverage where scholars look.** Library guides, academic society resources, scholarly-publishing newsletters, methods papers and independent reviews. [How brands build authority for AI search](https://underneath.agency/resources/how-brands-build-authority-for-ai-search) sets out how that kind of third-party standing is built.
4. **An honest answer to “why not just use ChatGPT?”.** A fair page on what a specialist tool does that a general research mode does not, and where the general tool is good enough.
5. **Data and privacy pages.** What happens to uploaded manuscripts, whether you train on user data, and what institutions control, in plain words.
6. **Plain pricing and licensing.** Individual, team and institutional options in one current place, so an assistant can tell a user whether their library may already pay for it.
7. **Repeated checks across assistants.** Put a fixed set of task, comparison, accuracy and access questions to ChatGPT, Gemini, Claude, Perplexity, Copilot and Google on a schedule, and ask new sign-ups how they heard of you. The wider software picture, including subscription revenue, is in [our B2B SaaS article](https://underneath.agency/resources/b2b-saas-revenue-from-ai-search).

## What don’t we know yet about how researchers pick AI tools?

Nobody has published how researchers settle on a research assistant, or how often an AI answer starts it.

The adoption figures come from publishers and information companies (Elsevier, OUP, Clarivate) that sell research tools themselves. The user counts come from the vendors. The citation-accuracy tests measured general assistants and older models, not the specialist tools, and no independent study we found compares specialist research tools’ citation accuracy with the latest deep research modes. No published study follows an AI recommendation through to a paid plan or an institutional license. For how often people actually open the sources AI cites, see [our article on checking sources](https://underneath.agency/resources/do-people-check-ai-sources).

## How can a research tool see what assistants tell students, researchers and librarians?

Ask the assistants the comparison and “is ChatGPT enough?” questions your users ask, and read what comes back.

Then audit task, comparison, accuracy and licensing questions across the main assistants and match the results against where your free sign-ups, paid upgrades and library licenses come from, so the gaps that cost the most users and renewals get fixed first. For a second pair of eyes on that audit, [talk to us about your research tool’s AI visibility](https://underneath.agency/contact). The [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page explains how we trace which sources shape those answers, then strengthen the corpus, accuracy and licensing evidence librarians check.

## Frequently asked questions

### Do AI assistants recommend specialist research tools or their own research modes?

There is no published data on how often each happens. General assistants document their own research modes in detail, so specialist tools need equally clear public evidence of what they do better, especially on citation accuracy and coverage.

### Does publishing citation-accuracy tests help an AI research tool?

It gives researchers and assistants something specific to check, and researchers say transparency and clear citations are what would raise their trust. Whether it changes how often assistants recommend a tool has not been measured.

### Should research tools target students or institutions?

Both, through different evidence. Students respond to recommendations and free access; libraries and research offices need data, security and licensing information. AI answers can reach both, so both kinds of information should be public.

### How can a research tool tell whether users come from AI search?

Use three signals together: AI referral traffic, a sign-up question asking how each new researcher heard of you, and regular checks of what assistants say to your key questions. Referral data alone undercounts, because many AI-driven visits arrive as direct traffic.

## Sources

- Elsevier (2025), [Researcher of the Future: Confidence in Research](https://www.elsevier.com/insights/confidence-in-research/researcher-of-the-future)
- Oxford University Press (2024), [How are researchers responding to AI?](https://corp.oup.com/news/how-are-researchers-responding-to-ai/)
- Clarivate (2026), [Pulse of the Library](https://clarivate.com/pulse-of-the-library/)
- Chelli and colleagues, Journal of Medical Internet Research (2024), [Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews](https://pmc.ncbi.nlm.nih.gov/articles/PMC11153973/)
- Columbia Journalism Review, Tow Center (2025), [AI Search Has A Citation Problem](https://www.cjr.org/tow_center/we-compared-eight-ai-search-engines-theyre-all-bad-at-citing-news.php)
- Consensus (2026), [About Consensus](https://consensus.app/home/about-us/)
- Elicit (2026), [Pricing](https://elicit.com/pricing) and [Elicit](https://elicit.com/)
- scite (2026), [Pricing](https://scite.ai/pricing)
- Elsevier (2026), [Scopus AI](https://www.elsevier.com/products/scopus/scopus-ai)
- OpenAI (2025), [Introducing deep research](https://openai.com/index/introducing-deep-research/)
- Google (2024), [Try Deep Research and our new experimental model in Gemini](https://blog.google/products/gemini/google-gemini-deep-research/)
- Similarweb (2026), [AI Search Stats 2026: Market Share, Referral, and Citation Data](https://aisearch.similarweb.com/blog/gen-ai-stats/)
- Chen and colleagues (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919)
- Sharma (2025), [The Discovery Gap: How Product Hunt Startups Vanish in LLM Organic Discovery Queries](https://arxiv.org/abs/2601.00912)

---

This is the Markdown twin of https://underneath.agency/resources/ai-research-tools-users-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Do a few websites win most AI Mode and Perplexity citations?"
description: "Yes. A small group of sites takes most AI citations, most of all in Google AI Mode in one Swiss study, least in Perplexity, with a long tail behind them."
canonical: "https://underneath.agency/resources/ai-search-citation-concentration"
published: 2026-10-08
updated: 2026-10-10
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Do a few websites win most AI Mode and Perplexity citations?

Yes: in every study we reviewed, a small group of websites collects a large share of AI search citations, while thousands of others are cited once or twice. How unequal it is depends on the engine and the industry. In one Swiss study, Google AI Mode was the most top-heavy engine and Perplexity the least, so targets should be set engine by engine.

## The short version

1. In a 2026 Swiss study of four engines over 45 to 46 days, AI Mode scored 0.782 on a 0-to-1 inequality scale and Perplexity 0.671, the lowest of the four.
2. In a 2026 audit of four engines, the 25 most cited websites took 23.8% of all citations, while 59.1% of cited websites appeared only once.
3. In ChatGPT’s answers to 100 health questions, 10 organizations took 52.8% of the 615 citations, led by Wikipedia, Mayo Clinic and Cleveland Clinic.
4. For news, the top 20 outlets took 67.3% of OpenAI models’ news citations in 2025, against 31.9% for Google’s and 28.5% for Perplexity’s.
5. In our own US data, AI Mode spread its citations across more websites than AI Overviews did, so the engine ranking is not settled.

## How concentrated are AI search citations?

Very: in most studies a handful of websites takes a large share, followed by a long tail. The exact degree depends on what is measured.

The clearest test comes from [Schulte and colleagues](https://arxiv.org/abs/2604.07585) at the University of St. Gallen; one author is also affiliated with a company that supplied the data. They asked ChatGPT, Gemini, Google AI Mode and Perplexity eight Swiss-German questions in each of four industries, daily from 24 January to 20 March 2026. They scored inequality on a scale where 0 means every website gets an equal share and 1 means one website gets them all.

The average across all engines and industries was 0.715. The authors describe this as “a highly unequal citation landscape in which a handful of domains captures most of the AI-generated visibility.” Reference sites such as Wikipedia are often in that handful; see [Wikipedia’s outsized role in AI search](https://underneath.agency/resources/why-wikipedia-matters-for-ai-search).

## Which engine concentrates citations the most?

In the Swiss study, Google AI Mode was the most concentrated and Perplexity the least. Other studies rank engines differently.

| Engine | Inequality score (0 = equal, 1 = one site gets all) |
|---|---|
| Google AI Mode | 0.782 |
| Gemini | 0.723 |
| ChatGPT | 0.684 |
| Perplexity | 0.671 |

Source: Schulte and colleagues, Swiss-German questions, January to March 2026.

Other measurements disagree on the order. [Allaham and Diakopoulos](https://arxiv.org/abs/2605.23684) audited the web versions of four engines on politics, health and the environment. On the same scale, ChatGPT was the most concentrated at 0.648 and Copilot the least at 0.492. Our own [comparison of AI Mode and AI Overviews](https://underneath.agency/research/ai-mode-vs-ai-overviews-study), on 400 US searches in September 2026, found AI Mode spreading its citations more widely than AI Overviews: the equivalent of 93.5 equally cited websites against 58.8. Market, language, topic and method all change the answer.

## Which websites tend to win?

The winners are usually large reference, institutional or platform sites, and they change by topic. There is no single list for every industry.

In health, [Jacques and colleagues](https://arxiv.org/abs/2601.17109) put 100 consumer health questions to ChatGPT 5.2 Pro. Ten organizations took 52.8% of all citations: Wikipedia had 10.7%, Mayo Clinic 9.9% and Cleveland Clinic 9.8%. The same study asked [whether cited pages need author bylines](https://underneath.agency/resources/do-ai-cited-pages-need-author-bylines). In news, [Yang](https://arxiv.org/abs/2507.05301) analyzed more than 366,000 citations from a public AI search comparison site in 2025. OpenAI models cited reuters.com in 22.8% of their news citations and apnews.com in 12.2%.

In commercial searches, video can lead. In [our AI Overview study](https://underneath.agency/research/ai-overview-citations-study) of 481 US AI Overviews, YouTube was cited in 64.2% of them and accounted for 15.7% of all 4,051 citations. Industry also matters within one market: in the Swiss study, telecommunications scored 0.750 and sporting goods 0.680.

## Is there still room for smaller websites?

Yes: the long tail is large, even though the top is crowded. Most cited websites appear only rarely, but they do appear.

In the Allaham and Diakopoulos audit, the 25 most cited websites took 23.8% of citations and the other 76.2% were spread across a long tail. Within that tail, 59.1% of websites were cited only once. A [2026 study across twelve European languages](https://arxiv.org/abs/2606.23165) found a similar shape: 80% of citations came from 3,878 of 21,077 websites, or 18.4%.

Google’s AI Overviews show the same shape. [Xu and colleagues](https://arxiv.org/abs/2605.14021) tracked 55,393 trending US searches over 40 days in 2026. The 10 most cited sites took 29.7% of AI Overview citations, while 56.2% of the sites cited appeared exactly once. The authors conclude that breadth, not concentration, is the main shape of AI Overview citations, and that the regular first page of results is more concentrated.

Breadth also differs by engine. In Yang’s news data, Perplexity models cited 1,430 different news sources, against 881 for Google’s and 707 for OpenAI’s. A long tail means a smaller site can be cited, but rarely for any single question.

## Do the same few websites win every day?

No: even the leading sources move from run to run. Concentration describes the pattern over many answers, not any single answer.

In repeated runs of the same questions within 24 hours, the Swiss study found that only 32 to 43% of cited sources overlapped between runs, on average. [Sielinski](https://arxiv.org/abs/2603.08924), who works for an AI visibility company, warns that a single run gives a misleadingly precise picture of how websites rank. In practice, a website’s share of citations is an average that needs many runs to estimate.

## What should you do about it?

Find out which websites dominate your category on each engine, then earn a place on them. Concretely:

1. List the most cited websites for your buyers’ questions on each engine you care about, from repeated runs, not one.
2. Set separate targets per engine; a share that is good on Perplexity may be weak on AI Mode.
3. Prioritize coverage on the sites that already dominate your category, such as [major reference, review or institutional sites](https://underneath.agency/resources/which-websites-do-ai-search-engines-cite).
4. Keep publishing your own useful pages; the long tail is real, but each page wins only occasionally.
5. Recheck the list regularly, because the leaders move.

If you want help mapping the dominant sources in your category, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The research shows that concentration is high, but not how stable or universal the ranking of engines is. The gaps:

- The engine ranking conflicts between studies, and no study compares the same questions across several markets.
- The Swiss study covers four industries in one market and language; the US picture may differ.
- Health and news studies cover single topics and dates.
- No study shows whether winning a top source makes an engine more likely to recommend your brand.
- Several datasets come from monitoring companies, and two key papers have authors linked to them.

## Frequently asked questions

### Which AI search engine relies on the fewest websites?

In one Swiss study, Google AI Mode did: it scored 0.782 on a 0-to-1 inequality scale. Our US data found the opposite for AI Mode against AI Overviews, so check your own market.

### Is Perplexity less concentrated than other AI engines?

In the Swiss study it was the least concentrated of four engines, at 0.671. In Yang’s 2025 news data it also cited the most different news sources, 1,430.

### What share of AI citations go to the top websites?

It varies by topic. In one health study, 10 organizations took 52.8% of ChatGPT’s citations; in a broader audit, the top 25 websites took 23.8%.

### Can a small website still get cited by AI search?

Yes, but rarely for any one question. In one audit, 59.1% of cited websites appeared only once.

## Sources

- Schulte and colleagues (2026), [Don’t Measure Once: Measuring Visibility in AI Search (GEO)](https://arxiv.org/abs/2604.07585), arXiv:2604.07585.
- Allaham and Diakopoulos (2026), [Synthetic Sources?: Auditing Generative Search Engine Citations for Evidence of AI-Generated Sources](https://arxiv.org/abs/2605.23684), arXiv:2605.23684.
- Jacques and colleagues (2026), [Authority Signals in AI Cited Health Sources: A Framework for Evaluating Source Credibility in ChatGPT Responses](https://arxiv.org/abs/2601.17109), arXiv:2601.17109.
- Yang (2025), [News Source Citing Patterns in AI Search Systems](https://arxiv.org/abs/2507.05301), arXiv:2507.05301.
- Żatuchin (2026), [The Language Blind Spot: How Query Language and Brand Recognition Tier Shape AI-Constructed Brand Reputation Across Twelve European Languages](https://arxiv.org/abs/2606.23165), arXiv:2606.23165.
- Xu and colleagues (2026), [Measuring Google AI Overviews: Activation, Source Quality, Claim Fidelity, and Publisher Impact](https://arxiv.org/abs/2605.14021), arXiv:2605.14021.
- Sielinski (2026), [Quantifying Uncertainty in AI Visibility: A Statistical Framework for Generative Search Measurement](https://arxiv.org/abs/2603.08924), arXiv:2603.08924.
- Underneath (2026), [AI Mode vs AI Overviews: how different are the sources?](https://underneath.agency/research/ai-mode-vs-ai-overviews-study)
- Underneath (2026), [Do AI Overviews cite the pages that rank?](https://underneath.agency/research/ai-overview-citations-study)

---

This is the Markdown twin of https://underneath.agency/resources/ai-search-citation-concentration. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Will AI search reduce your reliance on marketplaces?"
description: "Not yet. Booking sites still supplied most of Gemini’s hotel citations, but their share fell for experience-led questions, where brands’ own pages can compete."
canonical: "https://underneath.agency/resources/ai-search-marketplace-dependence"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Will AI search reduce our dependence on marketplaces and aggregators?

Not yet: in the one study that measured it, booking sites still supplied 55.3% of the sources Google’s Gemini cited for hotel questions, and 69.2% for booking-style questions. Their share fell sharply when travelers asked about the experience rather than the price. That gap is where a brand’s own content can compete, but nobody has yet shown it moves bookings.

## The short version

1. In a Gemini audit of Tokyo hotel questions, online travel agencies supplied 55.3% of cited sources overall and 69.2% for booking-style questions ([Zhu and Chang](https://arxiv.org/abs/2603.20062)).
2. For experience-led questions, the agencies’ share fell to 44.1%, a swing of 25.1 points that opens room for other sources.
3. Hotels with deep, question-answering websites were cited directly: they scored 8.6 out of 15 on content depth against 3.4 for hotels that were not.
4. In shopping answers, Google’s AI Overviews drew 17.8% of their source websites from retailers and marketplaces ([Uberti-Bona Marin and colleagues](https://arxiv.org/abs/2609.18729)).
5. Engines differ: Perplexity cited the brand’s own site in 94.9% of our “is this brand legit?” answers, Google AI Mode in 17.7% ([our reputation study](https://underneath.agency/research/is-it-legit-ai-reputation-study)).

## Do AI answers still send buyers through marketplaces and aggregators?

Mostly, yes: for booking-style questions, intermediaries still supply most of the sources AI cites.

The clearest evidence comes from hotels, an industry built on paying intermediaries for demand. Commission rates commonly run from 15% to 25% of booking value. For independent hotels, bookings from online travel agencies can exceed 61% of room revenue.

[Zhu and Chang](https://arxiv.org/abs/2603.20062), researchers at Blossom AI, a company, put 156 Tokyo hotel questions to Gemini 2.5 Flash with Google Search in March 2026, in English and Japanese. Online travel agencies such as Booking.com, Expedia and Jalan supplied 55.3% of all citations. For booking-style questions such as “cheap hotel in Shinjuku”, they supplied 69.2%. The authors’ own verdict: disintermediation is not imminent.

## Where can a brand’s own content compete?

In experience-led questions, where the intermediaries’ share fell to 44.1%.

The study paired each booking-style question with an experience version, such as “good value hotel with local charm in Shinjuku”. That [one change in question framing](https://underneath.agency/resources/does-question-phrasing-change-ai-sources) cut the agencies’ share of citations by 25.1 points. The gap held in every category tested:

| Kind of question | Non-agency share, booking-style | Non-agency share, experience-led |
|---|---|---|
| Budget | 15.5% | 42.5% |
| Rating and quality | 22.1% | 46.7% |
| Business travel | 40.1% | 68.5% |
| Convenience and location | 40.1% | 59.2% |

Budget questions were the agencies’ stronghold, because price comparison is what they do best. Business travel was the most open, thanks to content about workspaces and quiet rooms that listings rarely carry. In Japanese, 62.1% of citations for experience-led questions came from sources other than agencies.

Most of that freed-up share went to blogs, editorial sites and other third parties, not to hotels. Hotels’ own sites made up a steady 19–24% of the non-agency citations.

## Do shopping answers lean on retailers and marketplaces too?

It depends on the engine: Google’s AI answers draw on retailers, while chat assistants lean on review sites.

[Uberti-Bona Marin and colleagues](https://arxiv.org/abs/2609.18729) put 117 real product questions to ChatGPT, Gemini and Google’s AI Overviews, the AI summaries at the top of Google results, from the Netherlands in September 2026. In AI Overviews, retailers and marketplaces made up 17.8% of the source websites shown, and manufacturers and brands 16.1%. ChatGPT and Gemini leaned mainly on editorial and product-review sites. The authors treat these source categories as exploratory.

Our own [AI Overview study](https://underneath.agency/research/ai-overviews-frequency-study) of 1,248 US Google searches found marketplaces and retailers were a small share of citations overall but 7.9% on transactional searches. On those searches, Google’s own pages took 37.9% of citations, the largest single source type. For buying questions, Google itself is becoming one of the intermediaries.

## How often do AI answers cite a brand’s own website?

It varies by engine, from under a fifth of answers to nearly all of them.

In [our study](https://underneath.agency/research/is-it-legit-ai-reputation-study) of 79 brands asked “Is this brand legit?” on four engines in September 2026, Perplexity cited the brand’s own site in 94.9% of answers and Google AI Mode in 17.7%. Brand-owned pages are rarely the main evidence for a recommendation, though. [Chen and colleagues](https://arxiv.org/abs/2509.08919) found that for US consumer electronics questions, a search-enabled GPT drew 92.1% of its sources from independent media and review sites.

Assistants also go looking for directories. In [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study), ChatGPT ran a search aimed at a platform or directory, such as Yelp, Avvo or G2, in 31.2% of its answers to buyer questions. Leaving those platforms is not an option yet.

## What makes a brand’s own site get cited instead of the middleman?

Depth: in one small audit, hotels with deep, question-answering pages were cited directly, and shallow ones were not.

Zhu and Chang scored 14 hotel websites on content depth: FAQ, area guide, blog, access information and unique content. Hotels Gemini cited directly averaged 8.6 out of 15; hotels it did not cite averaged 3.4. Every hotel scoring 6 or more was cited; every hotel below 6 was not.

The contrast between two hotels makes the point. Kadoya Hotel, [an independent with no technical search work](https://underneath.agency/resources/deep-content-without-technical-seo-gemini), was cited through a 33-question FAQ and a 13-attraction sightseeing guide. Hotel K5, with a 9.6 rating on Booking.com and full technical setup, had only brief pages. Gemini mentioned K5 but drew its facts from Expedia, Hotels.com and an editorial site. The hotel was found, but through intermediaries.

With 14 hotels, this is an association, not proof. Better-known hotels may both write more and get cited more. Other signals, such as [guest rating and price](https://underneath.agency/resources/why-ai-recommends-one-hotel-over-another), also shape which hotel an assistant recommends.

## What should you do about it?

Keep the marketplace channel healthy while building the content that competes for experience-led questions.

1. **Keep marketplace and directory listings complete.** They still supply most citations for price and booking questions.
2. **Build deep answers on your own site.** Long FAQ pages, area or use-case guides and specific details beat brief pages ticking every box.
3. **Target the questions intermediaries answer badly.** Experience, fit and specialist needs were the most open in the research.
4. **Earn coverage in editorial and review sources.** They took much of the share the intermediaries lost.
5. **Measure share of citations by question type and engine.** Track how often answers cite you, an intermediary or a third party.
6. **Do not cut marketplace spend on citation data alone.** No study yet shows AI citations turning into direct bookings.

To set up that tracking, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The research shows who is cited when a buyer is deciding, not who captures the sale.

- The intermediary evidence covers one engine, one city and one month, and comes from researchers at a company.
- Citation is not booking. A hotel cited through Booking.com may still be booked directly, and the reverse.
- The content-depth finding rests on 14 hotels and cannot rule out that famous hotels simply write more.
- The retail source categories were sorted by an AI model and checked only by spot inspection.
- No study here covers Amazon or other product marketplaces from the seller’s side, or tracks changes over time.

## Frequently asked questions

### Will ChatGPT replace Booking.com and Expedia for hotel discovery?

Not on current evidence. In a Gemini audit of Tokyo hotel questions, online travel agencies still supplied 55.3% of cited sources, and the authors concluded disintermediation is not imminent.

### Can AI search help a hotel get more direct bookings?

It may help with discovery, but bookings have not been measured. Hotels with deep FAQ and area-guide pages were cited directly by Gemini, while the study measured citations, not reservations.

### Do AI shopping answers link to retailers or to brand sites?

Both, depending on the engine. Google’s AI Overviews drew 17.8% of source websites from retailers and marketplaces and 16.1% from manufacturers and brands in one audit.

### What kind of content helps a brand’s own site get cited by AI?

Deep, specific answers to buyer questions. In one hotel audit, every site scoring 6 or more out of 15 for content depth was cited directly, and every site below 6 was not.

## Sources

- Zhu and Chang (2026), [The End of Rented Discovery: How AI Search Redistributes Power Between Hotels and Intermediaries](https://arxiv.org/abs/2603.20062), arXiv:2603.20062.
- Uberti-Bona Marin et al. (2026), ["If I Had to Buy Just ONE: Galaxy S26 Ultra": Auditing AI-Generated Product Recommendations](https://arxiv.org/abs/2609.18729), arXiv:2609.18729.
- Chen, Wang, Chen and Koudas (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919), arXiv:2509.08919.
- Underneath (2026), [When does Google show an AI Overview? 1,248 US searches](https://underneath.agency/research/ai-overviews-frequency-study)
- Underneath (2026), [“Is this brand legit?” How AI assistants build a reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)

---

This is the Markdown twin of https://underneath.agency/resources/ai-search-marketplace-dependence. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "What should our AI search strategy be for the next five years?"
description: "Nobody can forecast five years of AI search. Invest in what has held across engines, measure properly, and keep options open on ads and Google’s AI Mode."
canonical: "https://underneath.agency/resources/ai-search-strategy-next-five-years"
published: 2026-10-07
updated: 2026-10-07
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# What should our AI search strategy be for the next five years?

Plan for a channel that is growing fast but whose rules are still being written. Put most of your effort into what has held across engines and years so far: credible third-party coverage, accurate and findable information, and honest measurement. Keep options open on what is unsettled, such as ads in AI answers, Google’s AI Mode and which assistants win. No study forecasts five years ahead; the research shows a direction and a set of open questions.

## The short version

1. Google answered 67% of the same US searches with an AI Overview in 2025, up from 42% in 2024 ([Aral, Li and Zuo](https://arxiv.org/abs/2602.13415)).
2. AI assistants already pick products: ChatGPT stated a first-person product preference in 79% of product-recommending answers in a 2026 audit ([Uberti-Bona Marin and colleagues](https://arxiv.org/abs/2609.18729)).
3. The channel is fragmented: for the same question, ChatGPT and Gemini shared only 5.4% of the websites they displayed.
4. In a US field experiment with 1,100 people, making AI Mode the only Google experience cut clicks to outside websites by 18.8 percentage points ([Wang and colleagues](https://arxiv.org/abs/2608.18352)).
5. A 2026 review of 45 studies found no optimization technique with a proven, lasting effect across engines on discovery or buyer behavior ([Martinez](https://arxiv.org/abs/2607.14035)).

## Can anyone predict what AI search will look like in five years?

No: the best evidence covers days and months, not years, and the platforms change faster than studies run. Plan for uncertainty, not a forecast.

The main randomized experiment on AI in Google search, by [Wang and colleagues](https://arxiv.org/abs/2608.18352), measured seven days of use. A Swiss study by [Schulte and colleagues](https://arxiv.org/abs/2604.07585) tracked German-language prompts on four AI engines for about 45 days. Roughly 65% of cited sources changed from one day to the next. A [critical review](https://arxiv.org/abs/2607.14035) of the field rates one claim with high confidence: commercial AI engines differ from one another and change over time.

Platform decisions can also flip overnight. [Aral, Li and Zuo](https://arxiv.org/abs/2602.13415) ran the same Google searches in 2024 and 2025. AI answers on Covid questions went from 1% of answers to 66% globally, which the authors attribute to a policy change. A five-year plan has to survive changes like that.

## How fast is AI search actually spreading?

Fast, on every measure we found, though some figures come from industry sources cited secondhand. The trend is clear even where the exact size is not.

In the Aral study, 67% of identical US searches showed an AI answer in 2025, against 42% in 2024. Shopping searches started from a low base but moved fastest: in the seven early countries, AI answers on shopping searches rose 222% in a year. Still, only 13% of shopping searches worldwide showed one in 2025, so commercial searches remain less covered than general knowledge.

Usage is rising outside Google too. The St. Gallen researchers cite ChatGPT growth to about 780 million weekly users by September 2025, from about 100 million in January 2024. They also cite clickstream research in which traditional search queries fell by more than 20% after people adopted AI search. A [2026 position paper](https://arxiv.org/abs/2606.12439) cites an AP-NORC poll of 1,437 US adults, in which 60% said they use AI to find information at least some of the time. We could not check these secondhand figures against their original sources.

## Is generative search a new distribution channel for companies?

Yes, increasingly: AI assistants now shortlist and recommend products, not just links. But it is a fragmented channel you cannot yet buy into reliably.

In the audit by [Uberti-Bona Marin and colleagues](https://arxiv.org/abs/2609.18729), ChatGPT stated a first-person product preference in 79% of answers that recommended a product. The same paper cites the largest published measurement of ChatGPT use, in which product and service recommendations were roughly 2% of conversations in 2025. The position paper cites a Salesforce survey in which 39% of 8,350 shoppers across 21 countries used AI for product discovery and related tasks.

Unlike search, the channel has no single front door. ChatGPT and Gemini shared only 5.4% of displayed websites for the same question, with no website in common in 76.7% of comparisons. In [our four-assistant study](https://underneath.agency/research/ai-assistants-brand-agreement-study), 66.3% of the options recommended for a question came from one assistant only. For where assistants sit in the buying journey, see [our guide to ChatGPT and Google in the customer journey](https://underneath.agency/resources/chatgpt-vs-google-customer-journey).

## What happens to website traffic as AI answers spread?

It falls in the best evidence so far, and Google’s design choices drive the size more than yours. Plan for fewer, later visits.

In the Wang experiment with 1,100 US participants, forcing every search into AI Mode cut clicks to outside websites by 18.8 percentage points. Outside Google, [Iannelli and Ai](https://arxiv.org/abs/2607.04282) studied a browsing panel; both work for an AI visibility company. They found that 34.1% of sessions involving an AI assistant had no visit to any outside website. For the same people, 19.5% of search-led sessions had none.

So the value of an AI answer increasingly sits in the answer itself: whether you are named, and how. For an estimate of what a default AI Mode would mean for your own traffic, see [our AI Mode traffic analysis](https://underneath.agency/resources/ai-mode-default-traffic-loss).

## Which parts of the future are genuinely uncertain?

Three are wide open: how AI answers get paid for, whether users accept AI Mode, and which assistants lead. Each could change your plan.

**Ads.** OpenAI expanded advertising in ChatGPT to 31 European markets in August, according to the product-recommendation audit. A Carnegie Mellon model by [Zhang and colleagues](https://arxiv.org/abs/2603.29071) suggests that competition between AI engines pushes them toward fewer ads. In their simulations, ad-heavy policies earned more at first but shrank the user base. That is a model, not observed behavior. Our guide on [ads in AI assistants](https://underneath.agency/resources/are-ads-coming-to-ai-assistants) covers what is known.

**User acceptance.** Google presents AI Mode as the future of search, yet in the Wang study it was used for only 0.6% of searches before the experiment. Forcing it on people raised the share who searched on a rival engine by 11.2 percentage points and lowered their trust in Google’s information.

**Which assistants lead.** With so little overlap between assistants, a plan built around one of them is a bet on that one surviving and staying the same.

## Which bets hold up whatever happens?

Credible third-party coverage, accurate information that search systems can find, and measurement you can trust. Each is useful under every scenario above.

Earned media is the most consistent pattern across engines. In a 2025 study by [Chen and colleagues](https://arxiv.org/abs/2509.08919), ChatGPT drew 93.5% of its cited sources for well-known brands from earned media, such as reviews and publishers, and Claude 87.3%. The critical review rates two levers with high confidence: how closely content matches the question, and where it sits among the material the AI reads.

Tactics are the weak bet. The same review of 45 studies found that no technique shows a stable, long-run effect across engines on discovery or buyer behavior. It rates the link from AI citations to clicks, conversions or revenue as very low confidence. A framework by [Kato and colleagues](https://arxiv.org/abs/2609.11915) starts from the same gap: standard marketing data do not record how often buyers see and notice a firm’s name in AI answers.

## What should you do about it?

Treat AI answers as a lasting channel with an owner and a yearly review. Build durable assets rather than chasing tactics. In practice:

1. **Give the channel an owner.** Someone should track AI visibility, ads and platform changes across engines and markets.
2. **Invest in durable assets.** Earn coverage in the publications and review sites AI engines cite, keep facts about your company accurate everywhere, and keep pages easy for search systems to find.
3. **Measure as a range, not a rank.** Track several assistants and countries, repeat questions over time, and report how often you appear.
4. **Build business measurement now.** Use comparison groups and stated start dates so you can tell your own effect from the market’s growth.
5. **Set triggers, not forecasts.** Decide in advance what you will change if Google makes AI Mode its default, if assistants open self-serve ads, or if one assistant clearly leads your category.
6. **Keep search fundamentals.** Search still opens many buying journeys, and AI engines run searches of their own.

If you want help building a multi-year plan, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

It cannot say where AI search will be in five years, or which strategies pay off over that span. The main gaps:

- No long-run study exists: the causal evidence covers a week, and most audits cover days to months.
- Several growth figures, including weekly users, survey adoption and declines in traditional searching, come from industry sources cited secondhand.
- The ads analysis is a theoretical model; how ads will affect organic answers is unmeasured.
- No study links AI visibility to revenue with strong evidence.
- Two of the studies cited here come from authors who work for AI visibility or monitoring companies.

## Frequently asked questions

### Will AI search replace Google within five years?

The research cannot say. It shows AI answers on 67% of identical US Google searches in 2025. Clickstream research cited in one study found traditional searches fell by more than 20% after people adopted AI search.

### Is AI search a sales channel yet?

Partly. In one 2026 audit, ChatGPT stated a first-person product pick in 79% of product-recommending answers, but different assistants showed very different sources and picks.

### Should we move our SEO budget into GEO?

Not wholesale. AI engines run their own web searches, and the review of 45 studies found no GEO technique with proven lasting effects. Keep search strong while building AI visibility.

### How often should we revisit our AI search strategy?

At least yearly, with monthly tracking. In one study, roughly 65% of the sources AI engines cited changed from one day to the next.

## Sources

- Aral, Li and Zuo (2026), [The Rise of AI Search: Implications for Information Markets and Human Judgement at Scale](https://arxiv.org/abs/2602.13415), arXiv:2602.13415.
- Uberti-Bona Marin, Bertaglia, Astante, Rijsbosch and van Dijck (2026), ["If I Had to Buy Just ONE: Galaxy S26 Ultra": Auditing AI-Generated Product Recommendations](https://arxiv.org/abs/2609.18729), arXiv:2609.18729.
- Wang, Gleason, Bart, Wilson and Metaxa (2026), [AI in Search Reduces Publisher Referrals Without Improving User Experience: Experimental Evidence](https://arxiv.org/abs/2608.18352), arXiv:2608.18352.
- Martinez (2026), [Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026)](https://arxiv.org/abs/2607.14035), arXiv:2607.14035.
- Schulte, Bleeker and Kaufmann (2026), [Don’t Measure Once: Measuring Visibility in AI Search (GEO)](https://arxiv.org/abs/2604.07585), arXiv:2604.07585.
- Wen and colleagues (2026), [Position: Generative Engine Optimization Creates Underexamined Risks, Governance Must Target Concentration, Disclosure, and Academic Blind Spots](https://arxiv.org/abs/2606.12439), arXiv:2606.12439.
- Iannelli and Ai (2026), [The New Shape of Search: How Conversational AI Recomposes Information Seeking](https://arxiv.org/abs/2607.04282), arXiv:2607.04282.
- Zhang, Jiao, Li and Xiong (2026), [An Economic Framework for Generative Engines: Advertising or Subscription?](https://arxiv.org/abs/2603.29071), arXiv:2603.29071.
- Chen, Wang, Chen and Koudas (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919), arXiv:2509.08919.
- Kato, Honma and Kato (2026), [Generative Marketing Mix Modeling: A Causal Inference Framework Linking GEO and GEM to Business Impact](https://arxiv.org/abs/2609.11915), arXiv:2609.11915.
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude agree on brands?](https://underneath.agency/research/ai-assistants-brand-agreement-study)

---

This is the Markdown twin of https://underneath.agency/resources/ai-search-strategy-next-five-years. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How AI software companies can stand out in AI search"
description: "By earning independent, current proof of what the product does, because AI assistants and buyers both discount the hype that fills the AI software category."
canonical: "https://underneath.agency/resources/ai-software-companies-in-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How can AI software companies stand out in AI search when every rival claims AI?

By giving AI assistants and buyers independent, current and checkable proof of what the product actually does, because claims alone are cheap in a category where almost every product now says it uses AI. The prize is real: enterprises spent $37 billion on generative AI in 2025, and AI products are often found and tried by individual users before procurement is involved. But AI companies also face an odd problem: the assistants that recommend software are built by companies that sell competing AI products.

## The short version

1. The market is large and still being divided: [Menlo Ventures](https://menlovc.com/perspective/2025-the-state-of-generative-ai-in-the-enterprise/) estimates companies spent $37 billion on generative AI in 2025, $19 billion of it on applications, and that 76% of enterprise AI use cases are now bought rather than built.
2. AI tools spread from the bottom up: Menlo found 27% of AI application spend comes through product-led growth, where individual users adopt first, against 7% in traditional software.
3. Buyers are skeptical of AI claims: [Gartner](https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027) estimates only about 130 of the thousands of vendors selling “agentic AI” are real, and calls the rest “agent washing.”
4. New AI products are hard for assistants to surface: in a [study of 112 Product Hunt launches](https://arxiv.org/abs/2601.00912), ChatGPT named them in only 3.32% of open questions such as “What are the best AI tools launched this year?”, though it recognized them 99.4% of the time by name.
5. Recent pages carry weight in a fast-moving category: across each assistant’s dated citations in [our freshness study](https://underneath.agency/research/ai-source-freshness-study), 17.4% to 22.6% were pages published in the last 90 days, compared with 6.9% of Google’s top 10.

## Who pays for AI software, from a single seat to an enterprise contract?

Individual users often adopt first, then teams and enterprises buy, and the best-known products reach billions in revenue.

The category covers three kinds of company: AI-native applications (coding tools, writing and meeting assistants, support agents), generative AI model and platform providers, and established software companies adding AI features. Their buyers range from a single developer on a credit card to an enterprise committee. Firms that build custom systems for clients face a related task, covered in [our guide for machine learning companies](https://underneath.agency/resources/machine-learning-companies-ai-search).

Menlo Ventures’ 2025 enterprise report, based on a survey of about 500 US enterprise decision-makers and a market model, gives the clearest picture of spend:

| Measure | 2025 figure |
|---|---|
| Enterprise spend on generative AI | $37 billion, up from $11.5 billion in 2024 |
| Spend on AI applications | $19 billion |
| Share of AI use cases purchased rather than built | 76% |
| AI deals that reach production, vs. traditional software | 47% vs. 25% |
| Application spend via product-led growth, vs. traditional software | 27% vs. 7% |

Menlo counts at least 10 products earning over $1 billion in annual recurring revenue and 50 earning over $100 million. It reports that Cursor reached $200 million in revenue before hiring a single enterprise sales rep, and that [startups took 63% of the AI application market](https://underneath.agency/resources/ai-startups-customers-from-ai-search) in 2025.

Consumer AI is a second market. [Sensor Tower](https://sensortower.com/press/press-release-boosted-by-gen-ai-services-consumers-spent-more-money-in-apps-than-games-for-first-time) reports AI app downloads grew 148% in 2025, and [Menlo’s 2026 consumer report](https://menlovc.com/wp-content/uploads/2026/09/menlo_ventures_consumer_ai_report-2026.pdf) finds 55% of US AI users now pay for at least one AI product.

What a customer is worth depends on whether they stay. In [RevenueCat’s 2026 report](https://www.revenuecat.com/state-of-subscription-apps) on subscription apps, AI-powered apps earned 41% more revenue per payer but saw subscribers churn 30% faster. For AI companies, acquisition and retention are two separate problems.

## Where does AI search sit in how buyers find AI tools?

At discovery and comparison, and for AI tools it often replaces the Google search a buyer would once have run.

Using AI to research a purchase is close to universal among business buyers: in [Forrester’s](https://forrester.com/blogs/state-of-business-buying-2026) survey of nearly 18,000 of them, 94% reported using AI during their buying process. In [G2’s 2025 survey](https://research.g2.com/cmos-2025-buyer-behavior-report-research-g2), generative AI chatbots were the top influence on vendor shortlists at 17.1%, ahead of software review sites at 15.1%.

AI products are a large share of what those buyers purchase. In [TrustRadius’s 2026 survey](https://www.demandgenreport.com/industry-news/news-brief/trustradius-ai-has-changed-how-buyers-research-but-not-what-they-trust/53853/) of technology buyers, AI tools or tools with AI features made up 59% of the purchases buyers made in the past year, and 75% of buyers who purchased an AI tool said it lived up to their expectations. G2 found half of enterprise buyers at companies with 1,000 to 5,000 employees had switched vendors for better AI.

AI search also changes what “ranking” means. In [our rankings study](https://underneath.agency/research/ai-citations-google-rankings-study), only 8.3% of the pages ChatGPT cited ranked in Google’s top 10 for the question. A strong Google position for “best AI note-taker” does not guarantee a mention when the same question goes to an assistant.

## Which questions do buyers ask when comparing AI tools?

Questions about use cases, alternatives, head-to-head comparisons, data safety, accuracy and price. We wrote the example prompts below to mirror how buyers compare AI tools; they are not drawn from real logs.

| Buyer concern | Illustrative prompt |
|---|---|
| Use case | “What’s the best AI tool for turning sales calls into CRM notes in HubSpot?” |
| Category | “Which AI coding assistants work best for a large TypeScript monorepo?” |
| Alternatives | “Alternatives to Jasper for a B2B content team that needs brand voice controls” |
| Comparison | “Cursor vs GitHub Copilot for a 40-person engineering team” |
| Data and security | “Does this AI search tool respect SharePoint permissions, and is customer data used for training?” |
| Accuracy | “Which AI support agents have published resolution rates, and how were they measured?” |
| Price | “How much does ElevenLabs cost for a podcast network, and what counts as usage?” |

These are questions where the answer changes month to month. Google documents that its AI features [may use “query fan-out”](https://developers.google.com/search/docs/appearance/ai-features), running several related searches before answering, and OpenAI documents that [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search) rewrites a question into targeted queries for its search providers. Our inference: for AI tools, those searches pull in recent launch posts, reviews, benchmarks and community threads, and old pages lose out to newer ones.

## How does AI visibility turn into revenue for an AI company?

Through two paths: individual signups that spread inside companies, and enterprise pilots that convert unusually well.

**The bottom-up path.** A user asks an assistant for a tool, signs up, and brings it to their team. With 27% of AI application spend arriving this way, according to Menlo, an assistant’s answer to one person can seed an enterprise account. Menlo names n8n, ElevenLabs, Gamma and Wispr Flow among companies that scaled in this way.

**The enterprise path.** A buying group asks assistants to map a category, [builds a shortlist and runs a pilot](https://underneath.agency/resources/ai-agent-companies-customers-from-ai-search). Menlo found 47% of AI deals reach production against 25% for traditional software, so getting into the pilot matters a lot. These groups are larger than usual: Forrester found that purchases including generative AI features involve buying groups twice the size, 14 members against seven.

In both paths the assistant’s answer comes before the money, and the product has to keep the customer afterwards. We infer that for AI companies with high early churn, being recommended for the right use case matters more than being recommended often: a user who arrives for a job the product does well is more likely to stay.

## Are AI companies competing with the assistants that recommend them?

Often, yes; the largest assistant makers sell AI products in many of the same categories.

Menlo’s data shows the overlap. Horizontal AI, the largest application category at $8.4 billion, is 86% copilots, “led by ChatGPT Enterprise, Claude for Work, and Microsoft Copilot.” Menlo’s 2026 consumer report adds that once most people use AI for a task, a product’s competition “becomes any AI that can do the job—including the general AI assistant the consumer already has open.”

So an AI writing, search or meeting tool may need ChatGPT, Gemini, Claude or Copilot to recommend it over the assistant’s own built-in features. Whether assistants favor their makers’ products has not, to our knowledge, been measured in a public study. We have no evidence either way and do not assume it.

What we can say is practical. A specialist product will be named when its advantage is specific and documented, such as a named integration, a measured accuracy figure or a compliance status, rather than when it claims to be generally better than a general assistant. That is our inference, not a platform rule.

## Why do assistants recommend some AI products and pass over others?

Mostly independent, recent evidence about the product; the platforms do not publish how they choose.

What has been observed:

- **Recognition is not recommendation.** In Sharma’s Product Hunt study, new products were recognized by name almost every time but surfaced rarely in open questions. Perplexity found launches more often when they had more referring domains and Reddit presence. It is a single-author study run on a small ChatGPT model through a developer interface, so treat it as a signal rather than a verdict.
- **Independent coverage counts.** When more independent sites name a brand in the pages assistants cite, recommendations follow: in [our brand-entity study](https://underneath.agency/research/brand-entity-ai-recommendations-study), each tenfold increase went with 4.7 times the odds of being recommended.
- **Newer pages win more citations.** For the same questions, the pages the four assistants cited in our freshness study were about half as old as Google’s top 10, counting from first publication. It is a pattern, not proof that updating a page causes citations.
- **Self-promotion is common and visible.** In [our study of “best X” lists](https://underneath.agency/research/self-promoting-best-lists-study), 24.2% of numbered lists that AI surfaces cited ranked their own publisher first. In AI software, where nearly every vendor publishes “best AI tools” lists, that kind of evidence is easy to discount, we infer.

Credibility is the trust factor that matters most here. Gartner’s “agent washing” warning and the US Federal Trade Commission’s [Operation AI Comply](https://www.ftc.gov/news-events/news/press-releases/2024/09/ftc-announces-crackdown-deceptive-ai-claims-schemes) both target overstated AI claims. In that sweep, DoNotPay agreed to pay $193,000 over claims that it offered “the world’s first robot lawyer.” Buyers have reason to look for proof, and the assistants answering them search for reviews and named sources. For more on reputation, see [how brands build authority for AI search](https://underneath.agency/resources/how-brands-build-authority-for-ai-search).

## What does GEO involve for a company that sells AI products?

Generative engine optimization (GEO) gives assistants accurate, current, independently confirmed evidence of what your product does and for whom.

For a company whose product is itself AI, the work usually looks like this:

1. **A precise identity.** Say exactly which jobs the product does, for which users, on which models and integrations, with the same words on your site, docs, marketplaces and review profiles. Vague “AI platform” language is easy to confuse with a dozen rivals.
2. **Proof, not adjectives.** Publish how accuracy, speed or savings figures were measured, link to customer stories with names, and keep security, data-use and training-data policies public and plain.
3. **Freshness.** Date and update the pages that describe features, pricing and model support, so assistants do not repeat last year’s product.
4. **Independent coverage.** Earn reviews on software platforms, coverage in trade press, comparisons by independent writers and real discussion in developer and practitioner communities.
5. **Honest comparisons.** Explain where you differ from general assistants and named rivals, including where they are the better fit. Whether such pages earn citations is covered in [do comparison pages help B2B brands get cited by AI](https://underneath.agency/resources/do-comparison-pages-help-b2b-ai-citations).
6. **Presence across assistants.** Check ChatGPT, Gemini, Claude, Copilot, Perplexity and Google’s AI features, because each is also a potential competitor and none behaves the same way.

No one can promise that an assistant will recommend a given AI product over its rivals, or over itself. GEO raises the odds that the evidence it finds is strong, current and accurate.

## What can’t the research tell AI software companies yet?

Buyers clearly use AI to find AI tools; whether assistants treat their makers’ own products differently is still unknown.

- **No public test of self-preference.** We found no published study of whether ChatGPT, Gemini, Claude or Copilot favor their makers’ products in recommendations.
- **Market figures are estimates.** Menlo’s spend figures combine a survey with a market model, and Menlo invests in several companies it names.
- **Vendor and platform interests.** G2 and TrustRadius run review platforms; RevenueCat sells subscription software. Their data are useful but not neutral.
- **Revenue links are unproven.** No public study ties AI visibility to revenue for AI software companies. The broader question of whether AI visibility pays is reviewed in [does AI visibility drive business results](https://underneath.agency/resources/does-ai-visibility-drive-business-results).
- **Fast change.** Assistants, models and the products they recommend change within months, so any snapshot ages quickly.

## How should an AI software company test whether assistants are feeding its signups and pilots?

Ask the assistants your buyers use the same use-case questions your best customers asked before they bought.

That first check shows whether you are named for the jobs you do best, which rivals and general assistants appear instead, and whether your features, pricing and data policies are described correctly. Then fix the facts, publish the proof and earn the independent coverage that confirms it.

If your growth depends on signups that turn into teams, or on pilots that turn into contracts, [talk to us about an audit of how assistants present your product](https://underneath.agency/contact). We will show how AI assistants describe your product against competitors and built-in assistant features, and which evidence would most improve your chance of being named for the use cases that bring in revenue. For the full scope, our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page walks through how the diagnosis leads to proof pages, fresher product facts and independent coverage.

## Frequently asked questions

### Do AI assistants recommend AI startups or only big names?

They can recommend startups, but new products are rarely surfaced unprompted. When 112 Product Hunt launches were tested, ChatGPT recognized almost all of them by name yet put them forward in only 3.32% of open discovery questions.

### Does ChatGPT favor OpenAI products over competitors?

No public study has measured this, so nobody can say. What is documented is that the major assistant makers sell AI products that compete with many AI software companies.

### Why does my AI product show outdated features in AI answers?

Assistants draw on whatever pages they find, and older pages about your product may still rank. Dated, updated feature and pricing pages, plus current third-party coverage, give them newer evidence to use.

### Is “AI-powered” in our messaging enough to be recommended?

No. With “agent washing” common, according to Gartner, and the FTC acting against deceptive AI claims, specific and checkable proof of what the product does carries more weight than the label.

## Sources

- Menlo Ventures (2025-12-09), [2025: The State of Generative AI in the Enterprise](https://menlovc.com/perspective/2025-the-state-of-generative-ai-in-the-enterprise/)
- Menlo Ventures (2026-09-15), [2026: The State of Consumer AI](https://menlovc.com/wp-content/uploads/2026/09/menlo_ventures_consumer_ai_report-2026.pdf)
- Gartner (2025-06-25), [Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027](https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027)
- Federal Trade Commission (2024-09-25), [FTC Announces Crackdown on Deceptive AI Claims and Schemes](https://www.ftc.gov/news-events/news/press-releases/2024/09/ftc-announces-crackdown-deceptive-ai-claims-schemes)
- Sensor Tower (2026-01-21), [Boosted by Gen AI services, consumers spent more money in apps than games for first time](https://sensortower.com/press/press-release-boosted-by-gen-ai-services-consumers-spent-more-money-in-apps-than-games-for-first-time)
- RevenueCat (2026), [State of Subscription Apps 2026](https://www.revenuecat.com/state-of-subscription-apps)
- Forrester (2026-01), [The State Of Business Buying, 2026](https://forrester.com/blogs/state-of-business-buying-2026)
- G2 (2025), [Proving Value in the Age of AI: 2025 Buyer Behavior Report](https://research.g2.com/cmos-2025-buyer-behavior-report-research-g2)
- Demand Gen Report (2026-07-30), [TrustRadius: AI Has Changed How Buyers Research, But Not What They Trust](https://www.demandgenreport.com/industry-news/news-brief/trustradius-ai-has-changed-how-buyers-research-but-not-what-they-trust/53853/)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- OpenAI Help Center (2025), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Sharma (2025), [The Discovery Gap: How Product Hunt Startups Vanish in LLM Organic Discovery Queries](https://arxiv.org/abs/2601.00912), arXiv:2601.00912.
- Underneath (2026), [How fresh are the pages AI engines cite?](https://underneath.agency/research/ai-source-freshness-study)
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude cite pages that rank?](https://underneath.agency/research/ai-citations-google-rankings-study)
- Underneath (2026), [Do Wikipedia and schema make AI assistants recommend a brand?](https://underneath.agency/research/brand-entity-ai-recommendations-study)
- Underneath (2026), [How many “best of” lists cited by AI rank their own brand first?](https://underneath.agency/research/self-promoting-best-lists-study)

---

This is the Markdown twin of https://underneath.agency/resources/ai-software-companies-in-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How can an AI startup win customers through AI search?"
description: "By earning independent coverage and reviews that assistants can find, so the startup reaches shortlists that turn into trials, team use and contracts."
canonical: "https://underneath.agency/resources/ai-startups-customers-from-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How can an AI startup win customers through AI search?

By getting written about, reviewed and compared on independent sites that AI assistants search, so the startup reaches the shortlists buyers now build in ChatGPT, Gemini, Perplexity and Google’s AI features. AI startups are selling into the fastest-growing software market on record, but they are also the companies assistants know least about. The work is less about tuning a website and more about becoming visible in the places assistants look.

## The short version

1. AI took close to 50% of all global venture funding in 2025, up from 34% in 2024, per [Crunchbase](https://news.crunchbase.com/ai/big-funding-trends-charts-eoy-2025/). Money is not the scarce resource for most AI startups; attention is.
2. Startups earned 63% of enterprise spending on AI applications in 2025, up from 36% a year earlier, in [Menlo Ventures’](https://menlovc.com/perspective/2025-the-state-of-generative-ai-in-the-enterprise/) estimate, and 27% of AI application spend came through self-serve, product-led adoption, against 7% in traditional software.
3. The market is crowded: [G2](https://company.g2.com/news/how-g2-is-bringing-accountability-to-ai-software-claims-in-2026) now lists 30,274 products across 111 AI categories, up 630% in a year.
4. New products are nearly invisible to open questions. In a test of 112 Product Hunt startups, ChatGPT recognized 99.4% when asked by name but surfaced them in only 3.32% of discovery questions ([Sharma](https://arxiv.org/abs/2601.00912), 2025).
5. For a young AI company, outside coverage matters most among the factors we have measured: in [our brand entity study](https://underneath.agency/research/brand-entity-ai-recommendations-study), every tenfold increase in the independent sites naming a brand went with 4.7 times the odds of an AI recommendation.

## Who pays an AI startup first, and how big can that first account get?

Individual users first, then teams and enterprises, and each early user can become a large contract.

Spending is rising fast. Menlo Ventures, a venture firm that invests in AI companies, estimates that companies spent $37 billion on generative AI in 2025, and that $19 billion of it went to applications rather than models or infrastructure. In that application layer, startups pulled ahead of incumbents, taking 63% of the market. Enterprises also prefer buying to building: Menlo found 76% of AI use cases are purchased rather than built internally.

The buyer is often not a procurement team at first. Menlo reports that 27% of AI application spending came through product-led growth, where individuals sign up and use a tool before any contract exists, nearly four times the rate in traditional software. Its starkest example is Cursor, which reached $200 million in revenue before hiring a single enterprise sales rep. When a formal process does start, it moves: 47% of AI deals reached production, against 25% for traditional software.

Budgets are now permanent. In [Andreessen Horowitz’s survey of 100 CIOs](https://a16z.com/ai-enterprise-2025/), leaders expected their spending on AI models to grow about 75% over the next year, and innovation budgets, which had funded a quarter of that spending, had dropped to just 7%. AI is now a core IT and business-unit line item.

That is why one new customer can be worth so much. Stripe told [TechCrunch](https://techcrunch.com/2025/02/27/stripe-ceo-says-ai-startups-are-growing-faster-than-saas-ever-did-and-calling-them-wrappers-misses-the-point/) that the top 100 AI companies reached $5 million in annualized revenue in 24 months, against 37 months for the top SaaS companies of 2018. A single engineer who finds your tool can be the first seat in a company-wide rollout.

## How do buyers find and judge an AI company they have never heard of?

They shortlist quickly, often with an AI assistant, then test hard because AI claims are easy to make.

Software buyers already lean on chatbots to build lists. In [G2’s July 2026 survey](https://sell.g2.com/2026-buyer-behavior-report) of more than 1,000 software buyers, 82% had sourced software recommendations from an AI chatbot in the last two years, and AI chatbots (37%) were almost level with review sites (38%) as the top influence on shortlists. This is a survey of all software buyers, run by a review platform with a stake in the answer, but AI products are a large and growing share of what those buyers evaluate.

Then the scrutiny starts. Andreessen Horowitz found that enterprise AI procurement “now mirrors traditional software buying,” with checklists, security reviews and benchmark comparisons. Enterprises increasingly use external benchmarks as quasi “Magic Quadrants” to filter vendors, but one leader told the firm, “It’s hard to pick without really trialing things.” The same survey found the main reason buyers prefer AI-native vendors is their faster innovation rate.

Trust is the obstacle a young AI company has to clear. Among developers in [Stack Overflow’s 2025 survey](https://survey.stackoverflow.co/2025/ai), more actively distrust the accuracy of AI tools (46%) than trust it (33%). Regulators are watching claims too: US prosecutors and the SEC charged the founder of the shopping app Nate with telling investors it ran on AI when contract workers did much of the work, after it raised about $42 million ([DLA Piper](https://www.dlapiper.com/insights/publications/2025/04/doj-and-sec-send-warning-against-ai-washing-with-charges-against-technology-startup-founder)). G2 has responded with an “AI Verified” designation that requires 10 or more reviewers to confirm they used a product’s AI features.

We made up the prompts below to show how buyers of new AI tools tend to frame these questions; they are examples, not observed data:

- Category with a constraint: “Best AI meeting note taker that works with Microsoft Teams and keeps data in the EU.”
- Emerging category: “What new AI tools can automate accounts payable for a mid-size company?”
- Alternatives: “Alternatives to Intercom’s AI agent for a small support team.”
- Comparison: “Cursor vs GitHub Copilot vs newer AI coding tools for a 50-person team.”
- Trust check: “Is [startup name] legit? Who funds it, and does it train on customer data?”

The last kind of question matters more for a startup than for an incumbent. [Our reputation study](https://underneath.agency/research/is-it-legit-ai-reputation-study) shows how assistants assemble a verdict on a brand from whatever evidence they can find.

## Why are new AI companies so often missing from AI answers?

Because assistants recognize new products when named but rarely bring them up for open questions.

[Sharma](https://arxiv.org/abs/2601.00912) tested 112 startups from the 2025 Product Hunt leaderboard, drawn from developer tools, productivity and AI. Asked about each product by name, ChatGPT recognized 99.4% and Perplexity 94.3%. Asked discovery questions such as “What are the best AI tools launched this year?”, the success rates collapsed to 3.32% and 8.29%. The version of ChatGPT he tested answered from training data that predated the launches, so no website change could have helped it. We cover that “recency wall” in detail in [why ChatGPT misses new products](https://underneath.agency/resources/why-chatgpt-misses-new-products).

Crowding makes it worse. G2 added 10,259 new AI products in its current fiscal year alone. An assistant asked for “the best AI tools” for a task picks a handful from thousands, and in [one vendor’s tracking data](https://arxiv.org/abs/2606.20065), niche and small brands appeared in just 11% of relevant answers on their first run, against 73% for household names.

There is a real opening, though. Assistants that search the web favor recent pages. When [our freshness study](https://underneath.agency/research/ai-source-freshness-study) dated the pages each assistant cited, 17.4% to 22.6% had been published in the previous 90 days, compared with 6.9% of the pages in Google’s top 10. A startup that is written about now can appear in answers long before it ranks on Google.

## What path leads from a ChatGPT mention to an AI startup’s first enterprise contract?

Usually through a self-serve signup that grows into team use, then a contract.

The path differs from classic software because so much AI adoption starts with one user. We infer it typically runs in four steps:

1. **Shortlist.** A developer, marketer or operations lead asks an assistant for tools that do a job. The answer names three to five products.
2. **Signup or trial.** The user tries one or two. Menlo’s data on product-led adoption shows how often this is where AI spending begins.
3. **Team spread.** If the tool works, colleagues join. Menlo describes n8n formalizing contracts “only after hundreds of employees were already active users.”
4. **Contract.** Procurement, security and finance arrive. Here the startup’s public record (reviews, security pages, customer stories, press) is checked again, often by asking an assistant.

The AI answer rarely shows up cleanly in analytics. Much of it arrives as a direct visit, a branded search or a signup that says “found you in ChatGPT” in an onboarding survey. Why product analytics undercount these signups is explained in [why analytics miss AI visibility](https://underneath.agency/resources/why-analytics-miss-ai-visibility).

Buying the product also tends to work. MIT’s NANDA initiative, as reported by [Fortune](https://fortune.com/2025/08/18/mit-report-95-percent-generative-ai-pilots-at-companies-failing-cfo/), found that purchasing AI tools from specialized vendors succeeded about 67% of the time, while internal builds succeeded one-third as often. That gives a startup a strong story once it is on the list. Firms that build systems to order can use the same story; see [how machine learning firms get named](https://underneath.agency/resources/machine-learning-companies-ai-search). The hard part is getting on it.

## What decides whether an assistant names your AI startup?

OpenAI and Google say little about it; research favors independent coverage, community discussion and recent, specific pages.

**What the platforms publish.** According to [OpenAI](https://help.openai.com/en/articles/9237897-chatgpt-search), ChatGPT search typically rewrites a question “into one or more targeted queries” for its search providers, and may send further, more specific queries once it has seen the first results. [Google’s AI Mode announcement](https://blog.google/products/search/ai-mode-search/) describes a “query fan-out” technique that runs multiple related searches across subtopics. Neither explains how a new AI product makes the cut. Our own measurements add scale: ChatGPT averaged 3.7 searches per buyer question in [our hidden searches study](https://underneath.agency/research/ai-hidden-searches-study), enough for one question about an AI note taker to touch its pricing, reviews and comparisons together.

**What studies have found.** AI search leaned on independent “earned” sites for 72.7% of its sources on US software questions, according to [Chen and colleagues](https://arxiv.org/abs/2509.08919). In Sharma’s test, the assistant that searched surfaced more startups when they had links from other sites, a strong Product Hunt result and genuine Reddit discussion; a score for on-page optimization showed no link with discovery. In our brand entity study, independent coverage was the strongest signal we measured.

**Trust factors specific to AI products.** We infer that AI buyers look for evidence an incumbent does not need: proof that the AI does what the site says, clear data and training policies, security documentation, model and hosting choices, and named customers. G2’s new AI designations exist precisely because buyers “can’t distinguish real AI investment from marketing fluff.” A reasonable expectation is that assistants answering trust questions draw on the same public evidence. Established vendors face the same test, covered in [how AI software companies stand out](https://underneath.agency/resources/ai-software-companies-in-ai-search).

## Why does being early matter so much in AI categories?

Because AI categories form fast, and the brands named first tend to stay named.

Several findings point the same way. In [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), B2B software had the most stable recommended brands of any industry we tested, with a mean overlap of 0.708 across repeated runs. Kumar’s data shows small brands do not climb much on their own: the fastest movers travelled only 10–20 points in three months. Andreessen Horowitz reports that switching costs are rising as companies build workflows around a tool. We infer that once a new category’s shortlist settles, a startup outside it faces a long climb, while one inside it benefits from every buyer who sees it named.

The risk runs the other way too. A startup missing from the answer loses deals it never sees, because the buyer who asked an assistant and tried three other tools never visits its site. For a wider view of that cost, see [what happens if you skip GEO](https://underneath.agency/resources/what-happens-if-you-skip-geo).

## What GEO work fits a startup that assistants barely know yet?

Building the outside evidence assistants find and trust about a young AI product; nothing can promise a mention.

For a young AI company, generative engine optimization (GEO) splits into six jobs:

1. **Entity clarity.** Name the category you are in, the job you do, who it is for and what it integrates with, the same way on your site, launch platforms, review profiles, GitHub, LinkedIn and Crunchbase. New categories need a plain name buyers would actually type.
2. **Independent coverage.** Pursue the newsletters, “best AI tools for” roundups, podcasts, analyst notes and practitioner communities your buyers read. Our article on [best-of lists](https://underneath.agency/resources/best-of-lists-ai-recommendations) covers which pages tend to matter.
3. **Reviews and verification.** Earn early, specific reviews on the platforms your buyers use, including those that verify AI use, because buyers and assistants both lean on them when judging an unknown vendor.
4. **Proof pages, ungated.** Publish honest comparison and alternatives pages, benchmarks with methods, security and data-use pages, pricing and real customer examples, dated and kept current.
5. **Community presence.** Take part genuinely in the forums and subreddits where your users compare tools; Sharma found real discussion mattered, and spam does not count.
6. **Measurement from day one.** Re-run your category, alternatives, comparison and “is it legit” questions in ChatGPT, Gemini, Perplexity, Claude, Copilot and Google, and add a “where did you hear about us?” field to signup.

For the broader playbook for challengers, see [how small brands get recommended by AI](https://underneath.agency/resources/how-small-brands-get-recommended-by-ai) and [do AI assistants favor big brands](https://underneath.agency/resources/do-ai-assistants-favor-big-brands).

## What is still unproven about AI search for young AI companies?

Nobody has measured how often an AI answer turns into a paying customer for an AI startup.

The market data comes from investors, including Menlo Ventures and Andreessen Horowitz, who back some of the companies they describe. The buyer data comes from surveys by review platforms with a stake in the result. Sharma’s study is a single-author test of one assistant version without search and one with it, on 112 products; newer assistants that search more often may surface new products faster. No published study follows an AI startup from first AI mention to closed contract, and none shows how long a new AI brand takes to appear in unbranded answers. Treat any promised timeline with suspicion.

## How can an AI startup tell whether assistants are sending it signups?

Test whether assistants name you for the questions behind your best signups, trials and demos.

Put your category, alternatives, comparison and trust questions to the main assistants, then compare the answers with where your best self-serve users and largest contracts came from, so you know which gaps cost you trials, seat expansion and pipeline. We can map that with you and plan the independent coverage and proof that would close the gaps: [ask us for an AI visibility review](https://underneath.agency/contact). For a young company still building its outside record, the [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page shows how we sequence entity clarity, reviews, coverage and proof pages.

## Frequently asked questions

### How long does it take for a new AI startup to show up in ChatGPT?

No study has measured it for new companies. Assistants that answer from training data cannot know you until a newer model is trained; assistants that search can find you as soon as other sites write about you. Being named when someone types your name is easy; being named for an open category question is the hard part.

### Should an AI startup launch on Product Hunt for AI visibility?

It can help. In Sharma’s test, a strong Product Hunt result went with more discovery on the assistant that searched the web. A launch only helps if it also produces coverage, reviews and discussion elsewhere.

### Do AI assistants trust benchmarks published by the startup itself?

That is not documented. Enterprise buyers do use external benchmarks as a first filter, so publishing results with clear methods, and getting them repeated by independent testers, is a reasonable bet.

### Is AI search worth it for an AI startup selling to enterprises?

Yes, at the shortlist stage. Enterprise deals often start with one user trying a tool, and buyers increasingly use assistants to build their first list. The contract is still won through trials, security reviews and references.

## Sources

- Crunchbase News (2026), [6 Charts That Show The Big AI Funding Trends Of 2025](https://news.crunchbase.com/ai/big-funding-trends-charts-eoy-2025/)
- Menlo Ventures (2025), [2025: The State of Generative AI in the Enterprise](https://menlovc.com/perspective/2025-the-state-of-generative-ai-in-the-enterprise/)
- Andreessen Horowitz (2025), [How 100 Enterprise CIOs Are Building and Buying Gen AI in 2025](https://a16z.com/ai-enterprise-2025/)
- TechCrunch (2025), [Stripe CEO says AI startups are growing faster than SaaS ever did](https://techcrunch.com/2025/02/27/stripe-ceo-says-ai-startups-are-growing-faster-than-saas-ever-did-and-calling-them-wrappers-misses-the-point/)
- G2 (2026), [2026 Buyer Behavior Report: The Evaluation Maze](https://sell.g2.com/2026-buyer-behavior-report)
- G2 (2026), [How G2 Is Bringing Accountability to AI Software Claims in 2026](https://company.g2.com/news/how-g2-is-bringing-accountability-to-ai-software-claims-in-2026)
- Stack Overflow (2025), [2025 Developer Survey: AI](https://survey.stackoverflow.co/2025/ai)
- DLA Piper (2025), [DOJ and SEC send warning against AI washing with charges against technology startup founder](https://www.dlapiper.com/insights/publications/2025/04/doj-and-sec-send-warning-against-ai-washing-with-charges-against-technology-startup-founder)
- Fortune (2025), [MIT report: 95% of generative AI pilots at companies are failing](https://fortune.com/2025/08/18/mit-report-95-percent-generative-ai-pilots-at-companies-failing-cfo/)
- OpenAI (2025), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Google (2025), [Expanding AI Overviews and introducing AI Mode](https://blog.google/products/search/ai-mode-search/)
- Sharma (2025), [The Discovery Gap: How Product Hunt Startups Vanish in LLM Organic Discovery Queries](https://arxiv.org/abs/2601.00912)
- Kumar (2026), [Generative Engine Optimization at Scale: Measuring Brand Visibility Across AI Search Engines](https://arxiv.org/abs/2606.20065)
- Chen and colleagues (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919)
- Underneath (2026), [brand entity](https://underneath.agency/research/brand-entity-ai-recommendations-study), [source freshness](https://underneath.agency/research/ai-source-freshness-study), [hidden searches](https://underneath.agency/research/ai-hidden-searches-study), [recommendation consistency](https://underneath.agency/research/ai-recommendation-consistency-study) and [“is it legit” reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study) studies

---

This is the Markdown twin of https://underneath.agency/resources/ai-startups-customers-from-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How do AI video companies win customers through AI search?"
description: "By being the tool AI assistants name when creators, marketers and training teams ask what to use, then turning that intent into free plans and upgrades."
canonical: "https://underneath.agency/resources/ai-video-generator-customers-from-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How do AI video companies win customers through AI search?

By being one of the tools an AI assistant names when a creator, marketer or training team asks what to use, and by making the free plan, the price and the trust questions easy to check once they arrive. AI video is a crowded, fast-moving category where buyers try several tools and switch cheaply, so being on the shortlist matters more than being first in a list of links. No published data yet shows how much of any AI video company’s revenue starts in an AI answer, so treat this as a channel to measure, not a proven one.

## The short version

1. The category is large and concentrated at the top. [Synthesia](https://seedcamp.com/views/synthesia-raises-200m-series-e-at-4b-valuation-to-continue-the-next-era-of-video-creation/) reports $150 million in annual recurring revenue (ARR) and 60,000 business customers, including 90% of the Fortune 100; [HeyGen](https://www.heygen.com/blog/heygen-surpasses-200m-arr) says it passed $200 million in ARR in June 2026, doubling in eight months.
2. Use is now mainstream: 63% of video marketers in [Wyzowl’s 2026 survey](https://wyzowl.com/video-marketing-statistics/) had used AI video tools, up from 51% a year earlier, and 52% of learning and development (L&D) teams in [Synthesia’s 2026 report](https://webcdn.synthesia.io/reports/AI%20in%20Learning%20and%20Development%20Report%202026.pdf) use AI for video creation.
3. Buyers shop around. In [Adobe’s survey](https://news.adobe.com/news/2025/10/adobe-max-2025-creators-survey) of 16,000 creators, 60% used more than one generative AI tool in three months, and 58% find new tools through personal research.
4. The money path is short and self-serve: [Synthesia’s pricing](https://www.synthesia.io/pricing) runs from a free plan to $29 and $89 a month before enterprise contracts, and [Runway’s](https://runwayml.com/pricing) paid plans start at $12 a month billed annually.
5. An assistant can know a tool’s name without ever recommending it. In a test of 112 Product Hunt startups, [Sharma](https://arxiv.org/abs/2601.00912) found ChatGPT recognized 99.4% by name but surfaced only 3.32% in “what are the best tools” questions.

## Who pays for AI video tools, and how much is each kind of customer worth?

Three groups buy it: creators, marketing teams and enterprise training or communications teams, each worth very different amounts.

**Creators and solo businesses** buy on a card. Adobe’s Creators’ Toolkit Report, run with The Harris Poll in September 2025, found 86% of creators actively use creative generative AI, and 52% use it to generate new assets such as images and video. They are price-sensitive: the top barriers to adopting a tool were high cost (38%), unreliable output quality (34%) and uncertainty about how the AI model was trained (28%).

**Marketing teams** buy for volume. Wyzowl, which surveyed 266 respondents in late 2025, found 91% of businesses use video as a marketing tool, and 69% of video marketers have made social media videos, the most common use. Explainers (68%), product demos (39%) and sales videos (37%) follow.

**Enterprise training, HR and communications teams** buy for scale and languages. Synthesia’s 2026 L&D report, based on 421 responses, found 87% of respondents already using AI, with voice generation (63%) and video creation (52%) among the main uses. HeyGen says 85% of the Fortune 100 use it “for training, sales enablement, and localization.” Voice tools sold to the same teams are covered in [how AI voice companies win these buyers](https://underneath.agency/resources/ai-voice-software-revenue-from-ai-search).

What a customer is worth follows the same split. A creator may pay tens of dollars a month. A company that standardizes on one platform for training in dozens of languages signs an enterprise contract and then grows it. Synthesia said it [reached $100 million in ARR](https://www.synthesia.io/post/100-million-revenue-adobe-investment) through new customers “across all pricing plans” and that more than 70% of the Fortune 100 were customers by April 2025, up from 40% two years earlier. Tech.eu reported that the company credited [existing customers spending more](https://tech.eu/2025/04/15/synthesia-scores-adobe-investment-as-hits-100m-in-arr/) as one of its growth drivers.

## At what point do AI assistants enter the search for a video tool?

At the research step, where buyers look for a tool, and increasingly inside the assistant itself.

Adobe found creators scout and test new tools through personal research (58%), social media trends (57%) and recommendations from other creators (41%). The survey did not ask which share of that research happens in ChatGPT or Gemini, so we cannot give an AI-specific figure for AI video buyers. Outside video, [G2’s 2026 survey of software purchasing](https://company.g2.com/news/g2-research-the-answer-economy) found 51% of B2B software buyers now start research with an AI chatbot more often than with Google. Enterprise video platforms are bought the same way, so a reasonable expectation is that their buyers behave similarly.

Two developments make AI assistants more than a research tool in this category:

- **The assistant is becoming a place to make the video.** HeyGen documents a [ChatGPT app](https://www.heygen.com/chatgpt-app) that lets users describe a video in ChatGPT and have HeyGen’s agent produce it without leaving the chat. For a buyer, the first trial may now start inside the assistant.
- **The audience is spread across assistants.** [Similarweb’s 2026 data](https://aisearch.similarweb.com/blog/gen-ai-stats/) shows ChatGPT’s share of generative AI website visits falling from about 76% in June 2025 to roughly 53% by May 2026, while Gemini rose to around 27–28%. A video company visible only in ChatGPT reaches a shrinking share of researchers.

Google’s own AI answers matter too. In [our study of 1,248 US searches](https://underneath.agency/research/ai-overviews-frequency-study), B2B software and technology keywords showed an AI Overview on 96.0% of searches, and in [our YouTube study](https://underneath.agency/research/ai-overview-youtube-videos-study) the AI Overview cited a YouTube video on 91.0% of B2B software searches. These were general software searches, not AI video ones, but they suggest that tutorial and review videos about your tool are part of what Google’s answers draw on.

## Which questions do AI video buyers ask?

Mostly fit questions: which tool for a use, which alternative to a known brand, and the cost per minute.

We wrote these example prompts ourselves to show how creators, marketers and training teams tend to frame a video-tool search; they are not observed data:

- Use case: “Best AI avatar tool for compliance training videos in 20 languages.”
- Alternatives: “Cheaper alternatives to Synthesia for a five-person marketing team.”
- Head-to-head: “HeyGen vs Synthesia for personalized sales videos.”
- Model choice: “Which AI video generator makes the most realistic product b-roll: Runway, Kling or Veo?”
- Rights and risk: “Can we use an AI avatar of our CEO in ads that run in the EU?”
- Price: “How many minutes of video do I get on HeyGen’s cheapest paid plan?”

Vendors already compete on these questions in public. HeyGen’s own blog lists a post titled “15 Best Synthesia Alternatives in 2026, Ranked by Real Render Tests.” Price questions are where answers slip most easily. Credit-based plans are hard to summarize, and in [our pricing study](https://underneath.agency/research/ai-pricing-accuracy-study), only 61.9% of the plan prices four assistants quoted for 45 software products were fully faithful to the vendor’s pricing page.

## How does an assistant’s mention become a paid video plan?

Through a free plan: the assistant names you, the buyer signs up, the credits run out, and the team upgrades.

**Named.** The assistant puts your tool on a short list for the buyer’s use case. Creators rarely stop at one: 60% of Adobe’s respondents used more than one generative AI tool in the past three months.

**Free plan or trial.** The cost of trying is close to zero. Synthesia’s pricing page lists a $0 plan with 10 minutes of video a month; Runway offers 125 one-time credits to explore its tools.

**Paid self-serve.** Synthesia’s paid plans are $29 and $89 a month, and Runway’s are $12, $28 and $76 a month billed annually. Usage is metered in credits or minutes, so revenue grows with use.

**Team and enterprise.** Training and communications teams move to contracts for security, seats and languages. Synthesia’s report found budget approvals (44%) and procurement processes (19%) add friction at this step, which is where a buyer goes back to research and asks follow-up questions about security and compliance. AI coding tools follow a similar climb; see [how coding tools grow from one developer to a team](https://underneath.agency/resources/ai-coding-tools-developers-ai-search).

The click itself may not show up as AI traffic. Similarweb reports that roughly six in ten ChatGPT referrals now land on a homepage, and that traffic filed as “direct” may increasingly be AI-driven discovery. HeyGen credits much of its growth to a product that spreads “largely through word of mouth.” We infer that, for video companies, part of what looks like direct or word-of-mouth signups may begin in an AI answer, and only a signup survey question will show it.

## Why would an assistant suggest one video generator over another?

Platforms document little; studies point to independent coverage, community presence and fresh pages; buyers add consent and rights.

**Documented by the platform.** No major assistant publishes how it picks the tools it recommends for a creative use case. What is documented is the rule set around the output: under [Article 50 of the EU AI Act](https://artificialintelligenceact.eu/article/50/), which applies from 2 August 2026, providers of systems that generate video must mark output as artificially generated, and deployers who publish a deepfake must disclose it. [Lewis Silkin notes](https://www.lewissilkin.com/insights/2026/07/31/the-new-ai-labelling-rules-for-deployers-in-the-advertising-supply-chain) that this is likely to capture a lot of AI-generated advertising.

**What research has found.** On US software ranking questions, AI search took 72.7% of its sources from independent “earned” sites, against 45.4% for Google, in work by [Chen and colleagues](https://arxiv.org/abs/2509.08919). Sharma’s Product Hunt study, whose sample leaned toward developer, productivity and AI products, found that startups whose own sites scored higher on GEO-style content checks were no more likely to be surfaced, while the number of other sites linking to them and community discussion did predict Perplexity visibility. Freshness matters in a category that ships new models every few months: in [our freshness study](https://underneath.agency/research/ai-source-freshness-study), four assistants cited pages first published about half as long ago as Google’s top 10 for the same questions (a ratio of 0.50).

**Trust factors specific to AI video.** Buyers ask about things other software categories rarely face:

- **Consent and likeness.** Synthesia describes a framework that requires explicit consent for all avatar likenesses. Whether an avatar of a real person can be created and published is a buying question in itself.
- **Training data and rights.** In Adobe’s survey, 69% of creators were concerned about their content being used to train AI without permission. Image tools answer similar questions; see [how image generators handle commercial rights](https://underneath.agency/resources/ai-image-generators-growth-from-ai-search).
- **Security and governance.** Synthesia’s L&D respondents named security (58%), accuracy (52%) and legal constraints (41%) as major obstacles. Synthesia announced in September 2024 that it was the [first AI video company](https://www.synthesia.io/post/synthesia-world-first-iso-42001-compliant-ai-video-company) to achieve ISO/IEC 42001 certification for AI management.
- **Reviews.** HeyGen reports that it won 281 badges in G2’s Summer 2026 reports. A reasonable expectation is that review profiles feed both buyers and the answers they read.

## What does a video tool lose when assistants skip it?

Usually the free signup, and with it the chance to grow into a paid or enterprise account.

The category moves fast and punishes a missing name. In [a16z’s March 2026 ranking](https://a16z.com/100-gen-ai-apps-6/) of consumer AI products, video generation “saw the most movement,” with Kling AI, Hailuo and Pixverse building traction and Google’s Veo 3 driving traffic to Google Labs. The same report notes that where Google and OpenAI focus their creative efforts, including video, standalone traffic compresses. When the model makers’ own tools sit inside the assistant, an independent video company that the assistant does not name has little else to fall back on. AI research tools face the same pressure from built-in research modes; see [how research tools earn trust](https://underneath.agency/resources/ai-research-tools-users-ai-search).

Low switching costs cut both ways. A creator who tries three tools in a month can find you from a single answer, and leave for a rival named in the next one. In an enterprise deal, the cost is larger: once a training team standardizes on one platform and builds its library there, the account tends to stay, so missing the first shortlist can mean missing years of seat and usage revenue. That last point is our inference from how these contracts are sold; no study has measured it for AI video.

## What can GEO do for an AI video platform, and what can’t it?

It makes your tool easy for assistants to describe, compare and trust; it cannot promise a placement in any answer.

For a video platform, generative engine optimization (GEO) work usually falls into seven areas:

1. **A clear category and use case.** Say plainly whether you are an avatar platform, a generative model, an editor or a localization tool, and for whom. Assistants answer use-case questions, so “training videos in 140 languages” is easier to match than “the future of video.”
2. **Independent reviews and coverage.** Creator reviews, tutorials on YouTube, practitioner newsletters, L&D and marketing publications, and review platforms. Those outside voices are what AI search tends to draw on when it names video tools. Our article on [the pages AI engines cite when naming brands](https://underneath.agency/resources/best-of-lists-ai-recommendations) covers which ones matter.
3. **Fair comparison pages.** Publish honest comparison and alternatives pages with real test conditions, the way buyers already ask. Video vendors weighing a “HeyGen vs Synthesia” style page can check the citation evidence in [our summary on comparison pages](https://underneath.agency/resources/do-comparison-pages-help-b2b-ai-citations).
4. **A pricing page an assistant can quote.** State what each plan includes in minutes as well as credits, keep one current page, and retire old figures that assistants can still find.
5. **Trust pages in plain words.** Consent rules for avatars, how models were trained, commercial-use rights, EU labeling support and security certifications, on public pages rather than behind a sales call.
6. **Fresh pages for each release.** New models and features change the answer to “which tool is best for this” every few months. Dated, specific release pages and updated benchmarks give assistants something current to cite. Our article on [why ChatGPT misses new products](https://underneath.agency/resources/why-chatgpt-misses-new-products) explains the gap launches face.
7. **Measurement across assistants.** Track a fixed set of use-case, alternatives and price questions in ChatGPT, Gemini, Perplexity, Copilot and Google, and add “how did you hear about us?” to signup.

## What is still unknown about AI search and video tool sales?

Nobody knows how much AI video revenue starts in an AI answer; no company in the category has published it.

The revenue figures here are company-reported. The adoption surveys are small (Wyzowl’s 266 respondents) or run by vendors with an interest in the result (Synthesia’s L&D report, Adobe’s creator survey). The best evidence on how AI search picks software comes from general software questions and from Product Hunt startups, not from AI video specifically. No published study yet follows an AI recommendation through to a paid video plan or an enterprise contract, and none shows how assistants handle questions about consent or deepfake rules. Whether visibility work pays off at all is the subject of [our article on business results](https://underneath.agency/resources/does-ai-visibility-drive-business-results).

## How should an AI video company find the answers that cost it signups and upgrades?

Check which AI answers name you for the use cases that bring in your paying customers.

A practical first step is an audit of your use-case, alternatives, model-choice and pricing questions across the main assistants, matched against where your free signups, upgrades and enterprise deals come from, so you fix the gaps that cost the most revenue first. We can build that view with you and plan the fixes: [ask us to review what assistants say about your video tool](https://underneath.agency/contact). The work that follows, from consent and rights pages to fresh release pages and fair comparisons, is described on our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page.

## Frequently asked questions

### Do AI assistants recommend the AI video tools with the best models?

Not necessarily. No assistant documents how it chooses, and studies of software questions find AI search leans heavily on independent sites. A tool with a strong model but little independent coverage can be well known by name yet rarely suggested, as Sharma’s 99.4% versus 3.32% gap shows.

### Should AI video companies publish “alternatives” and comparison pages?

Yes, provided they are fair and state their test conditions. Buyers ask head-to-head and alternatives questions, and vendors in the category already publish such pages. Pages that rank yourself first without test conditions are a weak bet.

### Do the EU’s deepfake labeling rules affect AI visibility?

Not directly. They affect what buyers ask. From 2 August 2026, companies publishing deepfake video in the EU must disclose it, so buyers ask which tools support labeling. Clear public answers make it easier for an assistant to answer that question about you.

### How can a video company tell whether signups come from AI search?

Use three signals together: AI referral traffic, a “how did you find us?” question at signup, and regular checks of what assistants say for your main video use cases. Referrals by themselves miss part of it, since many AI-driven visits land on the homepage or show up as direct traffic.

## Sources

- Seedcamp (2026), [Synthesia raises $200M Series E at $4B valuation](https://seedcamp.com/views/synthesia-raises-200m-series-e-at-4b-valuation-to-continue-the-next-era-of-video-creation/)
- Synthesia (2025), [Synthesia surpasses $100 million in annual recurring revenue and secures strategic investment from Adobe Ventures](https://www.synthesia.io/post/100-million-revenue-adobe-investment)
- Tech.eu (2025), [Synthesia scores Adobe investment as hits $100m ARR “milestone”](https://tech.eu/2025/04/15/synthesia-scores-adobe-investment-as-hits-100m-in-arr/)
- HeyGen (2026), [HeyGen Doubles to $200M ARR in Eight Months](https://www.heygen.com/blog/heygen-surpasses-200m-arr)
- HeyGen (2026), [HeyGen ChatGPT App](https://www.heygen.com/chatgpt-app)
- Wyzowl (2026), [Video Marketing Statistics 2026](https://wyzowl.com/video-marketing-statistics/)
- Synthesia (2026), [AI in Learning and Development Report 2026](https://webcdn.synthesia.io/reports/AI%20in%20Learning%20and%20Development%20Report%202026.pdf)
- Adobe (2025), [Adobe Creators’ Toolkit Report](https://news.adobe.com/news/2025/10/adobe-max-2025-creators-survey)
- Synthesia (2026), [Pricing](https://www.synthesia.io/pricing)
- Runway (2026), [Pricing](https://runwayml.com/pricing)
- Synthesia (2024), [Our journey to becoming the world’s first ISO 42001-compliant AI video company](https://www.synthesia.io/post/synthesia-world-first-iso-42001-compliant-ai-video-company)
- G2 (2026), [In the Answer Economy, Don’t Win the Click — Win the Answer](https://company.g2.com/news/g2-research-the-answer-economy)
- Similarweb (2026), [AI Search Stats 2026: Market Share, Referral, and Citation Data](https://aisearch.similarweb.com/blog/gen-ai-stats/)
- Andreessen Horowitz (2026), [The Top 100 Gen AI Consumer Apps, 6th Edition](https://a16z.com/100-gen-ai-apps-6/)
- EU Artificial Intelligence Act (2024), [Article 50: Transparency Obligations](https://artificialintelligenceact.eu/article/50/)
- Lewis Silkin (2026), [The new AI labelling rules for deployers in the advertising supply chain](https://www.lewissilkin.com/insights/2026/07/31/the-new-ai-labelling-rules-for-deployers-in-the-advertising-supply-chain)
- Chen and colleagues (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919)
- Sharma (2025), [The Discovery Gap: How Product Hunt Startups Vanish in LLM Organic Discovery Queries](https://arxiv.org/abs/2601.00912)
- Underneath (2026), [When does Google show an AI Overview?](https://underneath.agency/research/ai-overviews-frequency-study), [AI Overview YouTube citations](https://underneath.agency/research/ai-overview-youtube-videos-study), [software pricing accuracy](https://underneath.agency/research/ai-pricing-accuracy-study) and [source freshness](https://underneath.agency/research/ai-source-freshness-study)

---

This is the Markdown twin of https://underneath.agency/resources/ai-video-generator-customers-from-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How can AI voice companies turn AI search into revenue?"
description: "By being named, with clear consent and pricing facts, when developers, creators and contact center leaders ask AI which voice tool to try."
canonical: "https://underneath.agency/resources/ai-voice-software-revenue-from-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How can AI voice companies turn AI search into revenue?

By being one of the voice tools an AI assistant names when a developer, a creator or a contact center leader asks what to use, and by publishing the latency, language, pricing and consent facts those buyers check next. Revenue in this category is mostly usage-based, so an AI answer that wins a trial can turn into a bill that grows with every minute of audio or every call. No voice company has published how much of its revenue starts in an AI answer, so treat this as a channel to measure, not a proven one.

## The short version

1. The market has a clear leader and a crowded field behind it. [ElevenLabs](https://elevenlabs.io/blog/series-d) closed 2025 with over $330 million in annual recurring revenue (ARR), and [a16z counts](https://a16z.com/ai-voice-agents-2025-update/) 90 voice agent companies in Y Combinator since 2020.
2. Contact centers are the large prize and are still deciding. In a [Gartner survey](https://www.gartner.com/en/newsroom/press-releases/2024-12-09-gartner-survey-reveals-85-percent-of-customer-service-leaders-will-explore-or-pilot-customer-facing-conversational-genai-in-2025) of 187 service leaders, 44% were exploring a generative AI voicebot, 11% were piloting one and only 5% had one deployed.
3. Buyers are unhappy with what they have: in [Deepgram’s survey](https://deepgram.com/learn/state-of-voice-ai-2025) of 400 business leaders, 80% use some form of voice agent but only 21% are “very satisfied,” and 84% plan to raise voice budgets.
4. Pricing is metered: [ElevenLabs’ plans](https://elevenlabs.io/pricing) run from free to $990 a month in credits, and voice agent platforms such as [Retell AI](https://www.retellai.com/pricing) charge $0.07 to $0.31 a minute.
5. Consent is a buying question. [Consumer Reports](https://www.consumerreports.org/media-room/press-releases/2025/03/consumer-reports-assessment-of-ai-voice-cloning-products/) found that four of six voice cloning products it tested let researchers clone a voice from public audio without any technical check for the speaker’s consent.

## Who pays for synthetic voice and voice agents, and how much?

Three groups: creators and media teams, developers building voice into products, and contact centers replacing phone work.

**Creators, publishers and media teams** buy text to speech, dubbing and narration on a card. ElevenLabs’ pricing page lists Free ($0), Starter ($6), Creator ($22), Pro ($99), Scale ($299) and Business ($990) plans, all drawing on one pool of credits, with Enterprise priced on request.

**Developers** buy an application programming interface (API) and pay per character or per minute. ElevenLabs says its developer API is used by companies including Meta, Epic Games and Salesforce. These buyers often start small and grow with their own product’s usage.

**Contact centers and customer service teams** buy voice agents to answer and make calls. The economics are large. [Gartner estimates](https://www.gartner.com/en/newsroom/press-releases/2022-08-31-gartner-predicts-conversational-ai-will-reduce-contac) about 17 million contact center agents worldwide, with labor representing up to 95% of contact center costs, and it predicts that by 2029 [agentic AI will resolve 80%](https://www.gartner.com/en/newsroom/press-releases/2025-03-05-gartner-predicts-agentic-ai-will-autonomously-resolve-80-percent-of-common-customer-service-issues-without-human-intervention-by-20290) of common customer service issues without a human, cutting operational costs by 30%. Voice agent platforms charge per minute (Retell’s pay-as-you-go range is $0.07 to $0.31; [Vapi](https://vapi.ai/pricing) lists $0.05 a minute for hosting), so a contract’s value tracks call volume. Support software vendors selling to the same leaders are covered in [how support software gets chosen through AI](https://underneath.agency/resources/customer-support-software-ai-search).

**Accessibility** is a smaller but real use. ElevenLabs reports that more than 1,000 people with speech impairments have reclaimed their voices through its [Impact Program](https://elevenlabs.io/blog/series-c), and 86% of Deepgram’s respondents see voice AI as a key driver of more accessible customer interactions.

What a customer is worth therefore ranges from a few dollars a month to a usage contract that grows with every call. ElevenLabs says its 2025 growth came from enterprise adoption by companies such as Deutsche Telekom, Square and Revolut for customer support, commerce, training and inbound sales.

## At what point do voice buyers turn to an AI assistant?

At the shortlist step, alongside developer documentation, demos and benchmarks; no voice-specific survey measures it yet.

We found no published survey that asks voice software buyers how often they use ChatGPT, Gemini or Perplexity to choose a vendor. The closest evidence is for software buying in general: [G2’s 2026 survey](https://company.g2.com/news/g2-research-the-answer-economy) found 51% of B2B software buyers now start research with an AI chatbot more often than with Google. Voice platforms are bought by the same software and customer experience teams, so a reasonable expectation is that many of their buyers do the same.

What we can document is how many choices buyers face. A16z reports that companies building with voice made up 22% of one recent Y Combinator class, per Cartesia, and that founders building voice agents concentrate on B2B (about 69%) and healthcare (about 18%). When a category grows that quickly, buyers lean on summaries, and an AI answer that lists five voice agent platforms does the first round of filtering for them. Agent vendors outside voice face a similar filter; see [how AI agent companies reach shortlists](https://underneath.agency/resources/ai-agent-companies-customers-from-ai-search).

Voice buyers are not all asking the same assistant, either. ChatGPT’s share of generative AI website visits slid from about 76% in June 2025 to roughly 53% by May 2026 in [Similarweb’s 2026 data](https://aisearch.similarweb.com/blog/gen-ai-stats/), while Gemini and Claude grew. A voice company that checks only ChatGPT sees a shrinking slice of how developers and contact center teams research it.

## Which questions do voice software buyers ask?

Fit questions: which voice tool for a use case, language or budget, and whether it is safe and legal.

We wrote these voice prompts ourselves to show the kinds of questions creators, developers and contact centers bring; they are not recorded queries:

- Use case: “Best AI voice agent for after-hours appointment booking at a dental group.”
- Technical fit: “Which text to speech API has the lowest latency for a real-time voice agent?”
- Languages: “AI voices that sound natural in Hindi, Arabic and Brazilian Portuguese.”
- Comparison: “ElevenLabs vs Cartesia vs Deepgram for a customer support bot.”
- Price: “What does a voice agent cost per minute, including telephony?”
- Consent and law: “Can we clone our voice actor’s voice for ads, and what do we need from them?”

The last two types carry the most risk of a wrong answer. Per-minute pricing combines several separate charges, and in [our pricing study](https://underneath.agency/research/ai-pricing-accuracy-study), only 61.9% of the plan prices four assistants quoted for 45 software products were fully faithful to the vendor’s pricing page. Legal questions draw on press coverage, regulators and consumer groups as much as on the vendor’s own pages.

## How does a voice tool go from an AI answer to a metered bill?

Through a test: the assistant names you, the buyer builds a prototype, and usage grows into a contract.

**Named.** An answer to a use-case or comparison question puts your product on a short list. In voice, a buyer can hear the difference in minutes, so the shortlist is usually followed by a test, not a meeting.

**Free tier or prototype.** ElevenLabs’ free plan includes 10,000 credits a month; voice agent platforms let developers wire up a test number. The first revenue event is small. AI video tools rely on a similar [free-plan-first path to revenue](https://underneath.agency/resources/ai-video-generator-customers-from-ai-search).

**Usage grows.** Creators move up credit tiers. Developers ship and pay for each character or minute their product uses. A contact center that moves a queue onto voice agents pays per call minute, although a16z notes that price-per-minute models are coming under pressure as model costs fall, and expects pricing to combine a platform fee with usage.

**Enterprise contract.** Large deployments add security review, custom voices and integration. Deepgram found compatibility with existing systems and performance quality were the top factors for selecting voice vendors, and 46% of respondents said the ability to customize models would speed adoption.

Tracing this path is the hard part. By Similarweb’s count, roughly six in ten ChatGPT referrals now land on a homepage, and visits logged as “direct” may increasingly be AI-driven discovery. We infer that a voice company relying only on referral data will undercount AI-sourced signups, especially from developers who copy a product name from an answer into a new tab.

## What decides whether an AI assistant names a voice company?

Platforms document little; studies point to independent coverage; buyers in this category add consent, safety and legal standing.

**Documented by the platform.** No major assistant publishes how it picks the vendors it lists for a voice question. What is documented is the legal frame buyers ask about. The US Federal Communications Commission [ruled on February 8, 2024](https://docs.fcc.gov/public/attachments/DOC-400393A1.pdf) that AI-generated voices in calls are “artificial” under the Telephone Consumer Protection Act, so the consent rules for robocalls apply. Tennessee’s [ELVIS Act](https://www.tn.gov/governor/news/2024/3/21/photos--gov--lee-signs-elvis-act-into-law.html) added “voice” to the likeness rights it protects. In the EU, [Article 50 of the AI Act](https://artificialintelligenceact.eu/article/50/), which applies from 2 August 2026, requires providers of systems that generate synthetic audio to mark the output as artificially generated.

**Observed in studies.** No study looks at voice tools alone, but when [Chen and colleagues](https://arxiv.org/abs/2509.08919) tested US software ranking questions, AI search drew 72.7% of its sources from independent “earned” sites, against 45.4% for Google. In a test of 112 Product Hunt startups, [Sharma](https://arxiv.org/abs/2601.00912) found ChatGPT recognized 99.4% when asked by name but surfaced only 3.32% in discovery questions such as “what are the best tools,” so being known is not the same as being recommended. A16z’s [March 2026 ranking](https://a16z.com/100-gen-ai-apps-6/) of consumer AI products adds a market view: ElevenLabs has appeared on every edition since September 2023, and voice has been more defensible than image or video because the model giants have not focused there. For the image side of that comparison, see [how AI image generators get recommended](https://underneath.agency/resources/ai-image-generators-growth-from-ai-search).

**Trust factors specific to voice.** Independent scrutiny is part of the record assistants can find:

- **Consent checks.** Consumer Reports tested six voice cloning products in March 2025 and named four (ElevenLabs, Speechify, PlayHT and Lovo) that required only a self-attestation of the right to clone a voice. It credited Descript and Resemble AI with steps that made non-consensual cloning harder.
- **Safety measures in public.** ElevenLabs’ [safety page](https://elevenlabs.io/safety) describes blocking the cloning of celebrity and other high-risk voices, verification for its professional cloning tool, C2PA content credentials and a public classifier that detects its own audio.
- **Fit for the stack.** Deepgram’s respondents put compatibility with existing systems first, so integration pages, latency figures and supported languages are the facts that answer comparison questions.

We infer that a voice company whose safety practices are only described by critics will find that critique in AI answers about it, because those answers lean on independent sources. Publishing what has changed, with dates, gives assistants something newer to cite. Our article on [fixing wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) covers how to trace an outdated claim to its source.

## What does a voice platform lose when assistants pass it over?

The test, and with it the usage that would have grown on your platform instead of a rival’s.

Buyers are actively shopping. Deepgram found 84% of respondents plan to increase voice budgets within a year and only 21% are very satisfied with their current voice technology. Gartner found 44% of service leaders exploring a voicebot but only 5% with one deployed, which means most contact center vendor choices in this category have not been made yet. A vendor missing from the answers at this stage misses the evaluation itself.

Usage pricing raises the stakes. A developer who builds on a rival’s API and ships pays that rival for every character or minute afterward, and switching means re-testing voices, latency and integrations. That switching cost is our inference from how these products are priced; no study has measured it for voice. For a large brand, the risk runs the other way: a16z’s ranking and ElevenLabs’ growth show how concentrated attention can become, and our article on [whether AI assistants favor big brands](https://underneath.agency/resources/do-ai-assistants-favor-big-brands) explains why challengers need a clearly stated advantage to break in.

## What does GEO involve for a text to speech or voice agent company?

It gives assistants clear, checkable facts about voices, latency, pricing and consent; it cannot secure a place in any answer.

For a voice company, generative engine optimization (GEO) tends to span seven pieces of work:

1. **One clear description per buyer.** Say plainly whether you sell text to speech, voice cloning, dubbing, speech recognition or voice agents, and for which use cases, the same way on your site, docs, marketplace listings and profiles.
2. **Independent coverage and benchmarks.** Developer write-ups, latency and quality comparisons by third parties, customer case studies, analyst and trade coverage in customer service publications. These are the sources AI search leans on. For how a voice brand earns that coverage, read [how brands build authority for AI search](https://underneath.agency/resources/how-brands-build-authority-for-ai-search).
3. **Comparison pages with real numbers.** Latency, languages, voice quality samples and per-minute cost against named alternatives, with test conditions. Whether such pages get cited is covered in [our summary on comparison pages](https://underneath.agency/resources/do-comparison-pages-help-b2b-ai-citations).
4. **A pricing page an assistant can quote.** Explain credits in plain units (characters, minutes, calls), show what telephony and model costs add, and retire old price tables.
5. **Consent, safety and legal pages.** How you verify consent for cloning, what you block, how output is watermarked, and how you support FCC, ELVIS Act and EU labeling obligations, on public pages with dates.
6. **Readable developer documentation.** Public, current API docs and quick-start guides, since developers and their coding assistants read them directly. Our [agent-readable web study](https://underneath.agency/research/agent-readable-web-study) found only 3.2% of top websites return Markdown when an AI agent asks for it.
7. **Tracking voice questions in every major assistant.** Track a fixed set of use-case, technical, price and consent questions in ChatGPT, Gemini, Perplexity, Copilot and Google, and ask new accounts how they found you.

## What is still unknown about how voice buyers use AI assistants?

How often they pick a voice vendor from an AI answer, and what usage follows: nobody has published it.

The market figures here come from vendors (ElevenLabs, Deepgram) and investors (a16z) with an interest in the category’s growth, and the Gartner figures are forecasts and survey results, not measured outcomes. What we know about how AI search chooses software rests on general software questions and Product Hunt startups; none of it isolates voice tools. No study yet shows how assistants answer consent or legal questions about voice cloning, or whether published safety changes alter those answers. To judge whether voice visibility work is paying off in usage, read [our article on business results](https://underneath.agency/resources/does-ai-visibility-drive-business-results).

## How can a voice company see which AI answers feed its usage revenue?

Check what assistants say about you for the use cases behind your biggest usage accounts.

A practical first step is an audit of your use-case, technical, pricing and consent questions across the main assistants, matched against where your signups, API usage and contact center deals come from, so the gaps that cost the most usage revenue get fixed first. To have us map those answers against your API usage and contact center pipeline, [ask us for a voice visibility audit](https://underneath.agency/contact). Our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page covers what happens after that audit: clearer consent, latency and pricing pages, readable developer docs and ongoing measurement.

## Frequently asked questions

### Do AI assistants recommend the voice tool with the best audio quality?

Not necessarily. Assistants do not say how they pick a voice tool, and research on software questions shows AI search leaning heavily on independent sources. A product with excellent audio but little independent coverage can be well known by name yet rarely suggested.

### Will negative coverage about voice cloning safety show up in AI answers?

It can. Answers draw on independent sources such as consumer groups and the press. The practical response is to publish dated, specific information on what your safeguards are now, so newer and clearer sources exist alongside older criticism.

### Is usage-based pricing a problem for AI answers?

It makes errors more likely. Credits, per-minute rates and telephony charges are hard to summarize, and assistants often quote prices imperfectly. A plain pricing page with worked examples reduces the room for error.

### How can a voice company tell whether customers come from AI search?

Use three signals together: AI referral traffic, a “how did you hear about us?” field on the signup or API key form, and repeated checks of what assistants say for your main voice use cases. Referrals on their own miss much of it, since many AI-driven visits show up as homepage or direct traffic.

## Sources

- ElevenLabs (2026), [ElevenLabs raises $500M Series D at $11B valuation](https://elevenlabs.io/blog/series-d)
- ElevenLabs (2025), [Series C announcement](https://elevenlabs.io/blog/series-c)
- ElevenLabs (2026), [Pricing](https://elevenlabs.io/pricing) and [Safety](https://elevenlabs.io/safety)
- Andreessen Horowitz (2025), [AI Voice Agents: 2025 Update](https://a16z.com/ai-voice-agents-2025-update/)
- Andreessen Horowitz (2026), [The Top 100 Gen AI Consumer Apps, 6th Edition](https://a16z.com/100-gen-ai-apps-6/)
- Deepgram (2025), [State of Voice AI 2025](https://deepgram.com/learn/state-of-voice-ai-2025)
- Gartner (2024), [Gartner Survey Reveals 85% of Customer Service Leaders Will Explore or Pilot Customer-Facing Conversational GenAI in 2025](https://www.gartner.com/en/newsroom/press-releases/2024-12-09-gartner-survey-reveals-85-percent-of-customer-service-leaders-will-explore-or-pilot-customer-facing-conversational-genai-in-2025)
- Gartner (2022), [Gartner Predicts Conversational AI Will Reduce Contact Center Agent Labor Costs by $80 Billion in 2026](https://www.gartner.com/en/newsroom/press-releases/2022-08-31-gartner-predicts-conversational-ai-will-reduce-contac)
- Gartner (2025), [Gartner Predicts Agentic AI Will Autonomously Resolve 80% of Common Customer Service Issues Without Human Intervention by 2029](https://www.gartner.com/en/newsroom/press-releases/2025-03-05-gartner-predicts-agentic-ai-will-autonomously-resolve-80-percent-of-common-customer-service-issues-without-human-intervention-by-20290)
- Retell AI (2026), [Pricing](https://www.retellai.com/pricing)
- Vapi (2026), [Pricing](https://vapi.ai/pricing)
- Consumer Reports (2025), [Consumer Reports’ Assessment of AI Voice Cloning Products](https://www.consumerreports.org/media-room/press-releases/2025/03/consumer-reports-assessment-of-ai-voice-cloning-products/)
- Federal Communications Commission (2024), [FCC Makes AI-Generated Voices in Robocalls Illegal](https://docs.fcc.gov/public/attachments/DOC-400393A1.pdf)
- State of Tennessee (2024), [Gov. Lee Signs ELVIS Act Into Law](https://www.tn.gov/governor/news/2024/3/21/photos--gov--lee-signs-elvis-act-into-law.html)
- EU Artificial Intelligence Act (2024), [Article 50: Transparency Obligations](https://artificialintelligenceact.eu/article/50/)
- G2 (2026), [In the Answer Economy, Don’t Win the Click — Win the Answer](https://company.g2.com/news/g2-research-the-answer-economy)
- Similarweb (2026), [AI Search Stats 2026: Market Share, Referral, and Citation Data](https://aisearch.similarweb.com/blog/gen-ai-stats/)
- Chen and colleagues (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919)
- Sharma (2025), [The Discovery Gap: How Product Hunt Startups Vanish in LLM Organic Discovery Queries](https://arxiv.org/abs/2601.00912)
- Underneath (2026), [software pricing accuracy](https://underneath.agency/research/ai-pricing-accuracy-study) and [agent-readable web](https://underneath.agency/research/agent-readable-web-study)

---

This is the Markdown twin of https://underneath.agency/resources/ai-voice-software-revenue-from-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Do AI assistants weigh reputation differently from customers?"
description: "Partly. AI assistants share customers’ focus on rating and price, but in testing they overweighted eco-certification and ignored replies to reviews."
canonical: "https://underneath.agency/resources/ai-vs-human-reputation-priorities"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Should our reputation priorities for AI assistants differ from those for human customers?

Partly: AI assistants put rating and price first, as customers do, but in a controlled hotel audit they gave eco-certification far more weight than human studies would predict and gave replies to reviews almost none. They also lean hard on review volume and on the complaint sites where reviews live. Keep serving customers first, then add the few signals the machine weighs differently.

## The short version

1. In an audit of twelve AI models choosing hotels, rating and price carried 35.7% and 33.8% of the total weight, the same top two that human research finds ([Baig and colleagues](https://arxiv.org/abs/2606.16344)).
2. Eco-certification ranked third with 13.1% of the weight, although human studies treat green labels as a minor factor.
3. A line saying management replies to reviews carried 0.1% of the weight, statistically nothing.
4. The models’ own explanations misled: one model named brand in up to 55% of its reasons, while brand carried only 2.0% of the weight.
5. In our “is it legit” study, 88.0% of AI answers about a brand cited a review or complaint platform ([our reputation study](https://underneath.agency/research/is-it-legit-ai-reputation-study)).

## Where do AI assistants and customers agree?

On the basics: both put rating and price first, and both value rating well above review count.

[Baig and colleagues](https://arxiv.org/abs/2606.16344) asked twelve AI models, including GPT-4o-mini, Gemini and Claude models accessed through developer access, to pick one of five invented hotels. Each hotel’s rating, review count, price, certification and other details were randomized. The researchers then measured how much of each model’s choice each detail explained.

Rating took 35.7% of the weight and price 33.8%. Rating outweighed review volume by nearly four to one. That matches decades of research on human guests. One hotel-industry review of studies found rating’s link to hotel performance was 0.888 against 0.055 for review count, more than ten times larger.

So the investments you already make for guests, [a strong rating and a defensible price](https://underneath.agency/resources/why-ai-recommends-one-hotel-over-another), are the same ones the AI rewards most.

## Where does the AI depart from what customers value?

In two places: it over-rewards eco-certification and ignores whether management replies to reviews.

| Signal | Share of the AI’s weight | What human research says |
|---|---|---|
| Guest rating | 35.7% | The strongest lever |
| Price | 33.8% | A core driver |
| Eco-certification | 13.1% | Ranks below price, rating and other core attributes |
| Review volume | 9.4% | Secondary to rating |
| Chain or independent | 2.0% | Chains expected to reassure guests |
| Management replies to reviews | 0.1% | Linked to better online reputation |

The eco-certification result is the surprise. Human studies find green labels do affect bookings, but modestly, and mostly among environmentally minded travelers. The AI panel placed it third, rivaling review volume.

Chain membership also behaved unexpectedly. Brand is often treated as a way to reassure an unsure guest, yet the models gave chains a small penalty rather than a bonus.

## Should you stop replying to reviews?

No: replies still matter to people and to later reviews, but do not expect them to move AI picks directly.

In the audit, a hotel card that said “Management responds to guest reviews” gained nothing. Human research cited by the authors finds that hotels which start replying see their online reputation improve, partly because later reviews change. That indirect route still works through the reviews themselves.

It matters because reviews are where AI assistants build a brand’s reputation. In [our study](https://underneath.agency/research/is-it-legit-ai-reputation-study) of 79 brands asked “Is this brand legit?” on four AI engines in September 2026, 88.0% of answers cited a review or complaint platform. Claims tied only to review platforms were negative 56.5% of the time. Trustpilot and the BBB alone made up 61.7% of review-platform citations.

So the reply line on a listing is not the lever. What customers end up writing on the main review sites is.

## Why does review volume count for more with AI than you might expect?

Because assistants treat a large review count as proof that the rating is real.

In the hotel audit, a high rating moved choices more when it was backed by many reviews. The authors read this as the models using volume to judge how far to trust the rating. They warn that this favors established properties over newcomers with thin review histories.

Our own data on live ChatGPT answers points the same way, for local services rather than hotels ([our local picks study](https://underneath.agency/research/chatgpt-local-picks-google-profile-study)). The median business ChatGPT listed had 235 reviews, against 133.5 for those it left out, while both had a median rating of 4.9. Among businesses listed at least once, 69.2% of those with more reviews than the local median were listed in all three repeat runs, against 44.1% of those with fewer. Those are associations, not proof of cause.

For a new location or product, the review count is the gap to close first.

## Can you trust what an assistant says about why it chose?

Not fully: the models’ stated reasons left out factors they acted on and stressed ones they ignored.

The hotel audit asked each model to explain its pick. The explanations broadly tracked the real weights, with telling gaps. List position was almost never named, despite a 4.1% share of the weight. Review volume was rarely named, despite a 9.4% share. Brand was named in up to 55% of one model’s reasons, while it carried 2.0% of the weight.

The practical lesson: do not set reputation priorities by [asking an assistant what it values](https://underneath.agency/resources/ai-explanations-for-recommendations). Measure what it does instead.

## How do AI assistants describe a brand’s reputation?

As legitimate but flawed: answers nearly always pair a yes with a list of problems.

In our reputation study, every complete answer said the brand was legitimate, and 99.7% also made at least one negative claim. Negative claims reached the first paragraph in 10.9% of answers, but in 31.6% of Perplexity’s. Claims tied only to editorial review sites were negative 19.9% of the time, and those tied only to the brand’s own website 6.4%.

A human shopper can choose to read your site or your reviews. The AI assembles both into one verdict, and the complaints come from third-party review platforms. Google’s summaries may weigh that criticism differently; see [how AI Overviews treat negative sources](https://underneath.agency/resources/do-ai-overviews-downplay-negative-content).

## What should you do about it?

Keep serving customers first, then add the few signals the AI weighs more heavily.

1. **Hold rating and price discipline.** Both audiences put them first.
2. **Grow review volume steadily.** Ask every satisfied customer, and start early for new locations and products.
3. **Earn and display real certifications.** In hotels, eco-certification counted far more with AI than with people. Claim only what you hold.
4. **Keep replying to reviews, for people.** Count it as customer service, not AI work.
5. **Watch the complaint platforms.** Resolve problems where Trustpilot, the BBB and similar sites record them.
6. **Measure behavior, not explanations.** Track what assistants recommend and say over repeated questions.

For help measuring what assistants say about your brand, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The causal evidence comes from one industry and a stylized test, so treat the specific weights as a starting point.

- The only controlled audit is of hotels, with synthetic cards. Whether eco-certification carries similar weight in other industries is untested.
- “Management responds” was a single line on a card. How assistants treat actual reply text on review sites has not been measured.
- The human comparison uses published studies, not the same hotel cards, so the authors compare order of importance, not exact sizes.
- Our reputation and local studies are observational, collected on one or two days, and coded by AI models rather than people.
- Model versions change often, and the authors warn the weights may drift.

## Frequently asked questions

### Do AI assistants care about responses to reviews?

Not directly, in the evidence so far. A line saying management replies to reviews carried 0.1% of the weight in an audit of twelve AI models, although replies may still improve the reviews that assistants read.

### Does eco-certification help with AI recommendations?

In hotels, yes. Eco-certification was the third most important signal in a controlled audit, with 13.1% of the weight, far more than human studies would suggest.

### Are AI assistants’ explanations of their recommendations accurate?

Only partly. In a hotel audit, models rarely mentioned list position or review volume even though both affected their picks, and one model cited brand in up to 55% of its reasons despite brand carrying 2.0% of the weight.

### Which review sites do AI assistants cite about a brand?

Mostly a few big ones. In our study of “Is this brand legit?” answers, Trustpilot and the BBB made up 61.7% of review-platform citations.

## Sources

- Baig, Gillani and Ali (2026), [Whose hotel does the AI recommend? An algorithm audit of reputation signals in LLM-assisted hotel selection](https://arxiv.org/abs/2606.16344), arXiv:2606.16344.
- Underneath (2026), [“Is this brand legit?” How AI assistants build a reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study)
- Underneath (2026), [Which Google Maps businesses does ChatGPT recommend?](https://underneath.agency/research/chatgpt-local-picks-google-profile-study)

---

This is the Markdown twin of https://underneath.agency/resources/ai-vs-human-reputation-priorities. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How can an AI writing tool win users when ChatGPT also writes?"
description: "By owning a specific writing job general assistants do poorly, and getting that strength into the independent lists and reviews assistants cite."
canonical: "https://underneath.agency/resources/ai-writing-tools-users-from-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How can an AI writing tool win users when ChatGPT also writes?

By owning a specific writing job that general assistants handle poorly, and making sure the independent lists, reviews and comparisons that AI assistants read say so. Writing is one of the main things people already do inside ChatGPT, Gemini and Copilot, so the assistant a user asks for a recommendation is also your largest competitor. The tools that grow from AI search will, we expect, be the ones with a reason to exist that a third party can verify.

## The short version

1. Writing is the most common work use of ChatGPT: 40% of work-related messages in June 2025, and about two-thirds of writing messages ask it to edit the user’s own text, per a [study by OpenAI and Harvard economists](https://www.nber.org/system/files/working_papers/w34255/w34255.pdf).
2. General suites now include writing help: [Google](https://workspace.google.com/blog/product-announcements/empowering-businesses-with-AI) cut the price of Workspace Business Standard with Gemini from $32 to $14 per user per month by folding AI into the plan.
3. Business demand is broad: 95% of B2B marketers told the [Content Marketing Institute](https://contentmarketinginstitute.com/b2b-research/b2b-content-marketing-trends-research) their organizations use AI applications, and 89% of those use AI content creation tools.
4. A writing subscriber can be worth more than a typical app user: in [RevenueCat](https://www.revenuecat.com/state-of-subscription-apps-2025/)’s subscription benchmarks, most AI apps passed $0.63 per install by day 60, twice the $0.31 median across all apps.
5. The lists AI cites are often written by vendors: in [our study](https://underneath.agency/research/self-promoting-best-lists-study), 24.2% of AI-cited numbered “best X” lists with an identifiable publisher ranked that publisher first.

## Who uses AI writing tools, and what is a user worth?

Students, professionals and marketing teams, and most pay a modest subscription after starting free.

The audience is wide and young. [Pew Research Center](https://www.pewresearch.org/short-reads/2025/06/25/34-of-us-adults-have-used-chatgpt-about-double-the-share-in-2023/) found 34% of US adults have used ChatGPT, and 28% of employed adults now use it for work. Among teens, 26% use ChatGPT for schoolwork, up from 13% in 2023, per a [separate Pew survey](https://www.pewresearch.org/short-reads/2025/01/15/about-a-quarter-of-us-teens-have-used-chatgpt-for-schoolwork-double-the-share-in-2023/). These are the same people who look for paraphrasers, grammar checkers, essay helpers and email assistants. Students are also a core audience for AI research assistants, covered in [how AI research tools get found by students](https://underneath.agency/resources/ai-research-tools-users-ai-search).

Businesses buy too. In the Content Marketing Institute’s research, content creation tools for marketing copy were by far the most used kind of AI tool, at 89% of marketers whose organizations use AI, ahead of creative asset tools at 53%. Marketing teams are the natural buyers of brand-voice, long-form and SEO writing tools, and they buy seats, not single subscriptions.

A single user is worth little; a habit is worth a lot. Grammarly’s [plans page](https://www.grammarly.com/plans) shows a free tier and a Pro plan at $12 a month, a typical structure for the category. RevenueCat’s data on subscription apps shows AI apps earning more than $0.63 per install after 60 days, matching only Health and Fitness. The same report finds “not enough usage” is the top reason people cancel, at 32% to 47% of cancellations across categories. For a writing tool, the value of an AI recommendation depends on whether the user makes it part of daily work.

## How do general AI assistants change the market for writing tools?

They are both the biggest competitor and the place users ask which tool to use.

The OpenAI and Harvard study of ChatGPT messages found writing covers drafting emails and documents but also “editing, critiquing, summarizing, and translating text provided by the user.” Writing is especially common for management and business users, at 52% of their work-related messages. The same study notes writing fell from 36% of all ChatGPT usage in July 2024 to 24% a year later as other uses grew, but it remains the leading work task.

The suites are absorbing basic writing help. Google’s Workspace announcement says Gemini now helps users “summarize, draft, and find information” across Gmail and Docs, for $2 more than Workspace without Gemini. A standalone tool that only rewrites a paragraph is competing with a feature the buyer already pays for. Productivity apps face the same pressure; see [how productivity tools compete with built-in AI](https://underneath.agency/resources/productivity-software-ai-search-growth).

Some independent tools still win at scale. In [Andreessen Horowitz’s fifth Top 100 list](https://a16z.com/100-gen-ai-apps-5/) of consumer AI products, QuillBot and Gamma are among the fourteen “All Stars” that have appeared in all five editions of its web list. Grammarly, which says it works across more than 1 million apps and websites, competes on being present wherever people already write. We infer the durable positions are specialization (academic, legal, technical, brand voice), workflow presence and trust, not raw text generation. Image generators face a similar test; see [how specialist image tools hold their ground](https://underneath.agency/resources/ai-image-generators-growth-from-ai-search).

## Which questions do writers ask AI assistants about tools?

Use-case, comparison, “free” and trust questions, often naming the general assistant as the alternative.

These example prompts are ours, written to show how students, professionals and marketers phrase writing-tool questions; they are not observed data:

- Use case: “Best AI tool to rewrite academic paragraphs without changing citations.”
- Free tier: “Free AI paraphrasing tool with no daily word limit.”
- Comparison with assistants: “Grammarly vs QuillBot vs just using ChatGPT for editing a thesis.”
- Team fit: “AI writing tool that keeps our brand voice for a 10-person marketing team.”
- Privacy: “Which AI writing assistants do not train on my documents?”
- Integrity: “Is it allowed to use an AI grammar checker on a university essay?”

Integrity questions matter in this category. Pew found only 18% of teens say it is acceptable to use ChatGPT to write essays, while 42% say it is not. A tool that is clear about what it does, editing rather than ghostwriting, has an answer ready when a student or teacher asks. Comparison questions that include “just use ChatGPT” are the most direct test of a tool’s reason to exist.

## How does an AI recommendation turn into paying users?

Through a free signup that becomes a daily habit, then a subscription or a team plan.

We infer the path for most writing tools runs like this:

1. **Named.** A user asks an assistant for a tool for a specific writing job and gets three to five names, sometimes alongside advice to use the assistant itself.
2. **Free use.** The user tries one, usually on a free tier or trial.
3. **Habit.** The tool earns a place in the browser, word processor or email client. If usage stays low, the user leaves; RevenueCat’s cancellation data shows how often that happens.
4. **Paid.** A paywall on volume, quality or features converts the habitual users; for business tools, one user brings in a team.

RevenueCat reports that across categories, most trial-to-paid conversions happen immediately and that lifetime value rises nearly 60% from month 1 to year 1. Those figures cover mobile subscription apps of all kinds, not writing tools alone, but they show why a recommendation that reaches the right user at the right moment is valuable: the first session decides a lot. The AI answer itself rarely appears in analytics; it shows up as direct visits, branded searches and signups. For more on that measurement gap, see [why analytics miss AI visibility](https://underneath.agency/resources/why-analytics-miss-ai-visibility).

## What decides which writing tools an AI assistant names?

Platforms document little; studies point to “best of” lists, independent coverage and clear, checkable claims.

**Documented by the platforms.** Per [OpenAI](https://help.openai.com/en/articles/9237897-chatgpt-search), a request such as “best paraphrasing tool” is typically turned into “one or more targeted queries” sent to search providers, and [Google](https://blog.google/products/search/ai-mode-search/) describes AI Mode running “multiple related searches” across subtopics. Neither says how writing tools are ordered within the answer that comes back.

**Observed in studies.** Ranked lists carry outsized weight. In [Kumar’s](https://arxiv.org/abs/2606.20065) data from an AI visibility platform, the “best-of” listicle was the most-cited content format, at about 21% of all citations. Writing tools are a category full of vendor-written “best AI writing tools” lists. In our study of numbered lists cited by AI, 24.2% ranked their own publisher first, and publishers that included themselves put themselves first 92.9% of the time. For software questions, [Chen and colleagues](https://arxiv.org/abs/2509.08919) found AI search drew 72.7% of its sources from independent “earned” sites, and in [our brand entity study](https://underneath.agency/research/brand-entity-ai-recommendations-study), each tenfold rise in independent sites naming a brand went with 4.7 times the odds of a recommendation.

**Trust factors specific to writing tools.** We infer users and assistants weigh: whether the tool trains on user text, how it handles confidential documents, quality on specific kinds of writing, academic integrity positioning, language coverage and integrations. Quality is not a given: in the Content Marketing Institute’s survey, 12% of marketers using AI for content said its quality had decreased. Specific evidence, such as published samples, independent reviews and clear data policies, gives an assistant something concrete to repeat. For how self-ranking lists fit with fair competition, see [can competitors game AI recommendations](https://underneath.agency/resources/can-competitors-game-ai-recommendations).

## What does a writing tool lose when the assistant names a rival, or itself?

The user often stays inside the assistant, so the lost signup is invisible.

When an assistant is asked which writing tool to use and names others, or suggests it can do the job itself, the user may never search for you at all. New tools start furthest behind: in a test of 112 Product Hunt startups, [Sharma](https://arxiv.org/abs/2601.00912) found ChatGPT surfaced them in only 3.32% of discovery questions. The reasons a young paraphraser or grammar app gets overlooked are set out in [why ChatGPT misses new products](https://underneath.agency/resources/why-chatgpt-misses-new-products). Assistants that search do favor recent pages: in [our freshness study](https://underneath.agency/research/ai-source-freshness-study), 17.4% to 22.6% of each assistant’s dated citations were under 90 days old, against 6.9% of Google’s top 10. That gives a tool with fresh, specific coverage a real chance, even against older names.

## What does GEO involve for a writing assistant or editing tool?

It makes your specific strength findable, consistent and verified by others; it cannot promise a mention.

For a writing product, generative engine optimization (GEO) tends to span six pieces of work:

1. **A sharp entity.** Say plainly what writing job you do best, for whom and where (browser, Docs, Word, email), and say it the same way on your site, app store listings, review profiles and social accounts.
2. **Honest comparison pages.** Publish fair comparisons with named competitors and with general assistants, including where they are better. Our [comparison-page research summary](https://underneath.agency/resources/do-comparison-pages-help-b2b-ai-citations) explains what tends to help.
3. **Independent lists and reviews.** Get tested by independent reviewers, educators, writing communities and marketing publications rather than relying on your own “best of” page. Our note on [how best-of lists shape AI recommendations](https://underneath.agency/resources/best-of-lists-ai-recommendations) explains why outside lists count for more.
4. **Public trust pages.** State your data and training policy, privacy, security and academic integrity position in plain language, so an assistant asked a trust question can find your answer.
5. **Proof of quality.** Publish before-and-after samples, published benchmarks with methods and real customer examples by use case.
6. **Measurement.** Put your use-case, comparison, free-tier and trust questions to ChatGPT, Gemini, Perplexity, Claude, Copilot and Google on a regular schedule, and add a “where did you hear about us?” question to signup.

## What don’t we know yet about AI assistants sending people to writing tools?

Nobody has published how often AI answers point people to a writing tool, or how many of them pay.

What we know about writing inside ChatGPT comes from a study co-written by OpenAI’s own researchers. The revenue benchmarks come from RevenueCat’s data on mobile subscription apps of all kinds, which may not match web-first writing tools. The Content Marketing Institute heard only from B2B marketers, so students and everyday writers are not represented. Studies of AI citations cover many categories, not writing tools specifically, and the role of self-ranking lists in writing tools is our inference from those studies. We also cannot yet say how often assistants recommend themselves instead of a tool; we have seen no study that measures it.

## How should an AI writing tool check whether assistants are sending it subscribers?

Ask assistants about the writing jobs you do best, and note which tools, or which assistant, they suggest.

Run your use-case, comparison, free-tier and trust questions through the main assistants, then line the answers up against signup and free-to-paid conversion data to see which missing mentions cost you subscribers. To plan the independent reviews and published proof that would close those gaps, [ask us for a review of your writing tool’s AI visibility](https://underneath.agency/contact). How we help a writing product state its one clear strength and get it confirmed by outside reviewers is set out on the [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page.

## Frequently asked questions

### Will ChatGPT recommend a competing writing tool?

It can. Asked for the best paraphraser or grammar checker, assistants do list third-party products, and no platform publishes how that list is chosen. Whether an assistant suggests itself instead has not, to our knowledge, been measured.

### Do “best AI writing tools” lists on our own site help?

They are cited, but in our study about a quarter of AI-cited numbered lists ranked their own publisher first, and readers and assistants can see who wrote them. Independent lists and reviews carry more weight with buyers, and are harder to dismiss.

### Should a writing tool compare itself with ChatGPT on its website?

Yes, honestly. Users ask that comparison, and a fair page that explains where a specialized tool is better, and where it is not, gives assistants a clear answer to quote.

### How do we know if AI search is sending us users?

Combine three signals: whether assistants name you for your core questions, signups from AI referrals and direct traffic, and a “how did you hear about us” question at signup. Then follow those users through to paid plans.

## Sources

- Chatterji and colleagues, NBER (2025), [How People Use ChatGPT](https://www.nber.org/system/files/working_papers/w34255/w34255.pdf)
- Google Workspace (2025), [The best of Google AI, now included in Workspace Business and Enterprise plans](https://workspace.google.com/blog/product-announcements/empowering-businesses-with-AI)
- Content Marketing Institute (2025), [B2B Content Marketing Benchmarks, Budgets, and Trends](https://contentmarketinginstitute.com/b2b-research/b2b-content-marketing-trends-research)
- RevenueCat (2025), [State of Subscription Apps 2025](https://www.revenuecat.com/state-of-subscription-apps-2025/)
- Pew Research Center (2025), [34% of U.S. adults have used ChatGPT, about double the share in 2023](https://www.pewresearch.org/short-reads/2025/06/25/34-of-us-adults-have-used-chatgpt-about-double-the-share-in-2023/)
- Pew Research Center (2025), [About a quarter of U.S. teens have used ChatGPT for schoolwork](https://www.pewresearch.org/short-reads/2025/01/15/about-a-quarter-of-us-teens-have-used-chatgpt-for-schoolwork-double-the-share-in-2023/)
- Andreessen Horowitz (2025), [The Top 100 Gen AI Consumer Apps, 5th Edition](https://a16z.com/100-gen-ai-apps-5/)
- Grammarly (2026), [Plans](https://www.grammarly.com/plans) and [About](https://www.grammarly.com/about)
- OpenAI (2025), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Google (2025), [Expanding AI Overviews and introducing AI Mode](https://blog.google/products/search/ai-mode-search/)
- Kumar (2026), [Generative Engine Optimization at Scale: Measuring Brand Visibility Across AI Search Engines](https://arxiv.org/abs/2606.20065)
- Chen and colleagues (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919)
- Sharma (2025), [The Discovery Gap: How Product Hunt Startups Vanish in LLM Organic Discovery Queries](https://arxiv.org/abs/2601.00912)
- Underneath (2026), [self-ranking best lists](https://underneath.agency/research/self-promoting-best-lists-study), [brand entity](https://underneath.agency/research/brand-entity-ai-recommendations-study) and [source freshness](https://underneath.agency/research/ai-source-freshness-study) studies

---

This is the Markdown twin of https://underneath.agency/resources/ai-writing-tools-users-from-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How BI and analytics vendors get shortlisted when buyers ask AI"
description: "By being named on the short lists AI answers build for BI comparisons and alternatives, with public proof on governance, accuracy and cost."
canonical: "https://underneath.agency/resources/analytics-bi-software-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How do analytics and BI software companies get shortlisted when buyers ask AI?

By being named, and described accurately, when data leaders ask AI assistants to compare platforms, suggest alternatives or solve a specific reporting problem, then winning the trial on the buyer’s own data. Analytics buyers now use AI every day and shortlist very few products, so the first list matters. It also matters that AI is changing the product itself: buyers increasingly ask whether a BI tool can answer questions in plain language accurately, and vendors need public proof that theirs can.

## The short version

1. Data teams already work through AI: 80% of analytics professionals used AI in their daily workflow in [dbt Labs’ 2025 survey](https://www.getdbt.com/resources/reports/state-of-analytics-engineering-2025) of 459 practitioners, up from 30% a year earlier, mostly through ChatGPT, Claude and Gemini.
2. Plain-language analytics is the new buying question: 30% of those teams use AI to answer data questions in natural language, and another 29% want to but don’t yet.
3. Incumbents bundle hard: Microsoft reported more than 40,000 paid [Fabric customers](https://www.microsoft.com/en-us/investor/events/fy-2026/earnings-fy-2026-q4), up more than 60% in a year, and says [Power BI](https://powerbi.microsoft.com/en-us/blog/microsoft-named-a-leader-in-the-2025-gartner-magic-quadrant-for-analytics-and-bi-platforms/) has 30 million monthly active users.
4. Data readiness decides deals: 63% of organizations lack, or are unsure they have, the right data management practices for AI, according to a [Gartner survey](https://www.gartner.com/en/newsroom/press-releases/2025-02-26-lack-of-ai-ready-data-puts-ai-projects-at-risk) of 1,203 data management leaders.
5. Winning accounts expand: the data platform [Snowflake](https://www.01net.it/snowflake-reports-financial-results-for-the-second-quarter-of-fiscal-2027/) reported 828 customers spending more than $1 million a year and a net revenue retention rate of 126% in mid-2026.

## Who signs for a BI platform, and how much is an account worth?

A data leader and a business sponsor buy it together, usually after a trial; good accounts grow.

The buyer is split in two. A chief data officer, head of analytics or data engineering lead judges the technology: connectors, the semantic layer, governance, performance and cost. A business sponsor, often in finance, sales operations or marketing, judges whether people will use it. Budgets are rising after a lean period: in the dbt survey, 30% of respondents reported budget increases for their data teams, against just 9% the year before, and 40% reported headcount increases. EdTech shows a similar split between the people who use a tool and the people who approve it, with teachers finding tools that districts then buy, as our guide to [EdTech buyers who ask AI](https://underneath.agency/resources/edtech-ai-search) explains.

Two market facts shape every deal:

- **Bundling.** Microsoft has been named a Leader in Gartner’s Magic Quadrant for analytics and BI platforms for the eighteenth consecutive year, and Power BI is now part of Fabric. Many buyers already own a BI tool through a bigger contract, so an independent vendor must win a comparison against something that looks free.
- **AI is moving the category.** In a 2024 survey of 1,000 data and business leaders by MIT SMR Connections, [sponsored by ThoughtSpot](https://www.thoughtspot.com/press-releases/nearly-70-of-leaders-prioritize-genai-for-data-and-analytics), 67% were already using generative AI for an analytics use case. Vendor-sponsored, but consistent with the dbt data.

What a customer is worth depends on how it pays. Seat-based BI grows with users; consumption-based platforms grow with data and queries. Snowflake’s 126% net revenue retention means, roughly, that its existing customers spent 26% more than a year before, and it added 692 net new customers in the quarter. Snowflake is a [data platform chosen by data leaders](https://underneath.agency/resources/data-infrastructure-enterprise-demand-ai-search) rather than a BI tool, but it shows how analytics accounts compound once they become part of daily work.

## Where does AI already sit in the analytics buying journey?

Inside the data team’s daily work, and increasingly at the start of software research.

Analytics buyers are heavy AI users. dbt found 70% of respondents use AI for analytics development in some form, mainly through general-purpose assistants such as ChatGPT, Claude and Gemini. A data engineer who asks an assistant to write a query today can just as easily ask it which tool to use for the next project. That makes the assistant both a research channel and, for simple questions, a competitor to the product, we infer.

Surveys of technology buyers beyond the data team point the same way. In [TrustRadius’s 2026 report](https://www.demandgenreport.com/industry-news/news-brief/trustradius-ai-has-changed-how-buyers-research-but-not-what-they-trust/53853/) on nearly 2,500 technology buyers and vendors, 63% of buyers used AI during their purchase journey and 94% of them fact-checked its responses at least some of the time. These figures cover technology purchases in general, not BI alone.

Google’s AI features sit on top of BI research as well. Of the eight industries in [our study of 1,248 US searches](https://underneath.agency/research/ai-overviews-frequency-study), B2B software and technology had the highest rate: an AI Overview appeared on 96.0% of its searches.

## Which questions do analytics buyers ask AI assistants?

Comparisons, alternatives, use cases, accuracy and cost. We wrote the data-team prompts below to illustrate the pattern; they are not records of real buyer questions.

| Stage | Illustrative prompt |
|---|---|
| Use case | “What is the best way to give our sales team self-serve dashboards on top of Snowflake?” |
| Alternatives | “What are the main alternatives to Tableau for a company standardizing on Google Cloud?” |
| Comparison | “Power BI or Looker for a 2,000-person company with a dbt semantic layer?” |
| Natural language | “Which BI tools can answer plain-English questions accurately without exposing raw data?” |
| Embedded | “What embedded analytics platforms work for a multi-tenant SaaS product?” |
| Cost | “How does consumption pricing compare with per-seat BI licensing at 500 users?” |
| Governance | “Which BI platforms support row-level security and a central metrics layer?” |

Behind each of these prompts the assistant may run searches of its own. Google says AI Overviews and AI Mode [may use a “query fan-out” technique](https://developers.google.com/search/docs/appearance/ai-features), which issues multiple related searches, and OpenAI’s help center says [ChatGPT search typically rewrites a question](https://help.openai.com/en/articles/9237897-chatgpt-search) into targeted queries. Those searches often go looking for outside judgment: [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study) found that 43.8% of ChatGPT’s answers involved a search aimed at a named publication, ranking or award. For BI, a reasonable expectation is that analyst evaluations, review sites and practitioner blogs are among those named sources.

## How does a mention in an AI answer become a BI trial, then a contract?

Through a short shortlist and a trial on the buyer’s data, then expansion as usage spreads.

1. A data leader or analyst asks an assistant a use-case, alternatives or comparison question.
2. The answer names a few platforms and cites sources: reviews, comparisons, documentation, analyst coverage.
3. The buyer narrows to a short list. Those lists are short: 83% of buyers shortlisted three or fewer products, TrustRadius reports. In the [Gartner Digital Markets](https://cdn-static.bizzabo.com/bizzabo.file.upload/yVkIqwmWRDKbEhqeE3aQ_Gartner-Digital-Markets_2025-Report_Making-The-List.pdf) survey of 3,500 software buyers, initial lists averaged 4.4 options, and 81% of buyers end up buying from that initial list most or all of the time.
4. A trial or proof of concept runs on real data. In the same Gartner Digital Markets survey, 62% of buyers said the trial is their top factor in the final decision.
5. A contract follows, priced by seats or consumption, and grows as more teams use it.

AI answers matter most at steps 1 to 3. For analytics vendors, step 4 is where a vendor proves the claims the answer repeated: accuracy, speed, governance. A vendor that is described wrongly, for example with outdated pricing or a missing connector, may never get to run that trial.

## Why does an assistant name one BI platform and not another?

No platform documents how it chooses; studies point to third-party sources, and data leaders reward proof they can check.

**What Google and OpenAI say.** Both describe AI answers that search the web and cite sources; neither says how a BI platform ends up on the list.

**What research has measured.** On US software questions, earned sites supplied 72.7% of AI search sources, against 45.4% for Google, according to [Chen and colleagues](https://arxiv.org/abs/2509.08919). In [our four-assistant study](https://underneath.agency/research/ai-assistants-brand-agreement-study), assistants agreed most on B2B software, with an overlap of 0.543 on a scale from 0 to 1, still far from full agreement. Lists also drift between runs of the same question: [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study) found only 25.2% of the brands ChatGPT named held their place across all five.

**What analytics buyers trust.** Several trust factors are specific to the category:

- **Accuracy of AI answers on data.** Of the dbt respondents using AI for natural-language questions, two-thirds rely on plain query generation and one-third on a semantic layer, which past research links to higher accuracy. 27% plan to increase investment in semantic layer tooling. Vendors with published accuracy methods give buyers something to test.
- **Data quality and readiness.** Poor data quality is the challenge data teams cite most, at over 56% in the dbt survey, and Gartner predicts organizations will abandon 60% of AI projects unsupported by AI-ready data through 2026. The same governance check shapes how [data science platforms win enterprise buyers](https://underneath.agency/resources/data-science-platforms-ai-search).
- **Reviews and peers over analyst reports.** TrustRadius found 74% of buyers use reviews, while analyst reports were used by only 13%, a 63% decrease since 2022. Analyst recognition still appears in vendor marketing, but buyers lean more on peers.
- **Governance and security.** Row-level security, auditability and where data is processed are standard checks for any platform that touches company data.

**Our inference.** The proof analytics buyers check is mostly public: documentation, connector lists, pricing pages, reviews, benchmarks and practitioner write-ups. We would expect a BI vendor whose connector lists, pricing examples and accuracy methods are public, specific and consistent to give AI answers more to cite, but nobody has tested that for analytics software.

## What does a BI vendor lose when assistants leave it out?

A place on a list of three or four, which is most of the deal.

- **Short lists leave little room.** With 81% of software buyers purchasing from their initial list most or all of the time, a vendor absent from AI-shaped lists has to displace a product the buyer already favors.
- **The bundled default wins by silence.** If an answer names only the tool the buyer already owns, an independent vendor loses without a comparison, we infer.
- **Wrong facts end trials early.** Outdated pricing models, connectors or AI features in an answer can remove a vendor from consideration. Correcting a stale connector list or pricing model is covered in [how to fix wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers).
- **Regret is a renewal risk and an opportunity.** Gartner Digital Markets found 59% of buyers regret at least one software purchase from the past 18 months. Buyers looking to replace a tool will ask for alternatives, and the answer decides who gets the call.

## How does GEO work for an analytics or BI company?

It makes your proof on accuracy, governance, connectors and cost easy for AI assistants to find and repeat. It cannot promise any BI vendor a spot on a data leader’s shortlist.

1. **One clear identity.** Say what the platform is for (self-serve BI, embedded analytics, natural-language analytics, a metrics layer) and for whom, the same way on your site, review profiles and partner marketplaces.
2. **Public technical proof.** Publish connector lists, security and governance documentation, pricing models with examples, and how your natural-language features stay accurate. Write for the data engineer who will check it.
3. **Honest comparison and alternatives pages.** Buyers ask “X or Y” and “alternatives to Z.” Answer with real trade-offs, including where the bundled option is good enough; see [whether comparison pages help](https://underneath.agency/resources/do-comparison-pages-help-b2b-ai-citations).
4. **Reviews and practitioner coverage.** Encourage detailed reviews from data teams and earn coverage in data engineering publications and communities. Third-party lists tend to outweigh your own pages, as we show in [which pages to target](https://underneath.agency/resources/best-of-lists-ai-recommendations).
5. **Authority beyond your site.** Talks, partner listings with data platforms, and independent benchmarks build the record assistants draw on; see [how brands build authority for AI search](https://underneath.agency/resources/how-brands-build-authority-for-ai-search).
6. **Measure across assistants and runs.** Repeat your BI questions in ChatGPT, Google AI Overviews and AI Mode, Gemini, Perplexity, Copilot and Claude; [how many prompts to track](https://underneath.agency/resources/how-many-prompts-to-track-ai-visibility) explains how many to use. The broader approach for subscription software sits in [B2B SaaS revenue from AI search](https://underneath.agency/resources/b2b-saas-revenue-from-ai-search).

## What is still unknown about AI search in BI buying?

Nobody has measured how often data leaders choose BI tools through AI, or whether it lifts revenue.

- **No BI-specific buyer study.** The TrustRadius and Gartner Digital Markets figures cover software buyers in general.
- **Many sources are vendors.** dbt Labs, Microsoft, Snowflake and ThoughtSpot all sell in this market; TrustRadius and Gartner Digital Markets sell vendor visibility.
- **AI may change what buyers buy.** If assistants answer simple data questions directly, demand may shift toward governed data and semantic layers rather than dashboards. That is our inference, not a finding.
- **Revenue effects are untested.** Whether a mention leads to more analytics trials or contracts is open; see [does AI visibility drive business results](https://underneath.agency/resources/does-ai-visibility-drive-business-results).

## What should a BI vendor check before buyers start their next trials?

Run your buyers’ comparison and alternatives questions through the assistants, then close the proof gaps you find.

Start from three people: the data leader, the analyst and the business sponsor. Ask their questions of each main assistant more than once, and log who is named, which sources are cited, and whether your pricing model, connectors, governance and natural-language features come out right. A miss usually traces back to proof that is not public, or to thin coverage from reviewers and practitioners.

That review is where we start with analytics vendors: [talk to us about an audit of your BI shortlist visibility](https://underneath.agency/contact). We will show which comparison and alternatives answers include or omit you, and which public proof gaps are most likely costing you trials, proofs of concept and seat or consumption growth. Our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page explains how that audit becomes a program of connector, governance and pricing proof, with comparison pages and practitioner coverage.

## Frequently asked questions

### Do data teams use ChatGPT to choose BI tools?

They use it heavily for work: 70% of dbt’s respondents use AI for analytics development. No study isolates vendor selection.

### Do Gartner Magic Quadrant placements still matter?

Less than before for buyers: only 13% of technology buyers used analyst reports in their decision, TrustRadius found.

### How do we compete with a BI tool bundled into a larger contract?

Be specific about where you are better, and say it on pages assistants can cite: comparisons, benchmarks and reviews.

### Should we publish our natural-language accuracy methods?

Yes, if you can back them. Only 30% of dbt’s respondents use AI for natural-language data questions; accuracy is the barrier.

### When would AI visibility show up in BI trials and contracts?

Expect trials first and contracts later; enterprise analytics deals follow a proof of concept on real data.

## Sources

- dbt Labs (2025), [2025 State of Analytics Engineering Report](https://www.getdbt.com/resources/reports/state-of-analytics-engineering-2025)
- Microsoft (2026-07-29), [Fiscal Year 2026 Fourth Quarter Earnings Conference Call](https://www.microsoft.com/en-us/investor/events/fy-2026/earnings-fy-2026-q4)
- Microsoft Power BI Blog (2025), [Microsoft named a Leader in the 2025 Gartner Magic Quadrant for Analytics and BI Platforms](https://powerbi.microsoft.com/en-us/blog/microsoft-named-a-leader-in-the-2025-gartner-magic-quadrant-for-analytics-and-bi-platforms/)
- Gartner (2025-02-26), [Lack of AI-Ready Data Puts AI Projects at Risk](https://www.gartner.com/en/newsroom/press-releases/2025-02-26-lack-of-ai-ready-data-puts-ai-projects-at-risk)
- Snowflake, via 01net (2026-08), [Snowflake Reports Financial Results for the Second Quarter of Fiscal 2027](https://www.01net.it/snowflake-reports-financial-results-for-the-second-quarter-of-fiscal-2027/)
- ThoughtSpot and MIT SMR Connections (2024-09-12), [Nearly 70% of Leaders Prioritize GenAI for Data and Analytics](https://www.thoughtspot.com/press-releases/nearly-70-of-leaders-prioritize-genai-for-data-and-analytics)
- Demand Gen Report (2026), [TrustRadius: AI Has Changed How Buyers Research, But Not What They Trust](https://www.demandgenreport.com/industry-news/news-brief/trustradius-ai-has-changed-how-buyers-research-but-not-what-they-trust/53853/)
- Gartner Digital Markets (2025), [Making the List: 2025 Software Buying Trends](https://cdn-static.bizzabo.com/bizzabo.file.upload/yVkIqwmWRDKbEhqeE3aQ_Gartner-Digital-Markets_2025-Report_Making-The-List.pdf)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- OpenAI Help Center (2025), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Chen, Wang, Chen and Koudas (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919), arXiv:2509.08919.
- Underneath (2026), [When does Google show an AI Overview? 1,248 US searches](https://underneath.agency/research/ai-overviews-frequency-study)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude agree on brands?](https://underneath.agency/research/ai-assistants-brand-agreement-study)
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)

---

This is the Markdown twin of https://underneath.agency/resources/analytics-bi-software-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Can we monitor ChatGPT and Gemini visibility through their APIs?"
description: "Not reliably. In a 2026 test, ChatGPT’s app and its API shared only 12.0% of cited websites for the same question asked a fraction of a second apart."
canonical: "https://underneath.agency/resources/api-ai-visibility-monitoring"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Can we monitor ChatGPT and Gemini visibility through their APIs?

Not reliably, if you want to know what customers see. In a 2026 audit, ChatGPT’s consumer app and its API shared only 12.0% of cited websites for identical questions sent moments apart, and Gemini’s only 14.8%. API data is useful for controlled tests, but it describes the API, not the app.

## The short version

1. For the same product question, ChatGPT’s app and API shared on average 12.0% of cited websites and 4.8% of exact pages (Uberti-Bona Marin and colleagues, 2026).
2. The split was often total: the app and API shared no website at all in 60.9% of ChatGPT pairs and 43.4% of Gemini pairs.
3. The differences had a pattern: YouTube and Reddit were 44.4 and 29.1 percentage points more likely to appear through Gemini’s API than its app.
4. Citations are a small slice of what an engine finds: ChatGPT’s API retrieved an average of 37.38 pages per question but cited only 3.22.
5. In our own API-based study, ChatGPT’s API reported 3.7 searches per answer, a useful detail, though the app may search differently ([our hidden searches study](https://underneath.agency/research/ai-hidden-searches-study)).

## Do APIs show the same sources as the ChatGPT and Gemini apps?

No. In the most direct test so far, the app and the API cited mostly different websites.

An API is the developer connection that software uses to send questions to an AI model without opening the chat app. It is attractive for tracking because it can be automated and lets you fix the model and settings. [Uberti-Bona Marin and colleagues](https://arxiv.org/abs/2609.18729), at universities in Maastricht, Utrecht and Zurich, tested whether it gives the same picture as the app. They took 117 real product questions written by users and sent each one to the logged-out ChatGPT and Gemini apps and to their APIs, from the Netherlands.

Each app request was paired with an API request for the same question. Across 702 pairs, the median gap between the two was 0.17 seconds. The results still diverged sharply.

| Same question, app vs API | ChatGPT | Gemini |
|---|---|---|
| Cited websites shared, on average | 12.0% | 14.8% |
| Exact pages shared, on average | 4.8% | 11.9% |
| Pairs sharing no website at all | 60.9% | 43.4% |

Most of the disagreement was about which websites appeared, not which page on a website. Websites seen on only one side made up 79.8% of all ChatGPT pages observed across the two.

## Why do the API and the app give different answers?

They are different setups, even under the same brand. The research has not pinned down exactly which difference matters most.

The logged-out ChatGPT app did not say which model it used, so the researchers picked the closest API model they could. Settings such as location also work differently: the OpenAI API accepts an approximate location, while Gemini’s API does not. The authors stress that they cannot attribute the gap to the access method alone.

Search behavior is part of it. When a tool calls an API, the tool’s settings decide whether and how often the model searches the web. One tracking platform let ChatGPT and Claude search the web only once a week per brand, to control cost. In between, they answered from memory ([Kumar](https://arxiv.org/abs/2606.20065), who co-founded that platform). A tool built that way measures something different from the app a customer uses. Querying through the API, [Sielinski](https://arxiv.org/abs/2607.10341) found OpenAI’s search model returned no citations for about 17% of questions on one topic.

APIs also expose different layers. ChatGPT’s API reported both the pages its searches returned and the pages it cited. On average it returned 37.38 pages per question and cited 3.22, and only 27.5% of returned websites made it into the citations. So a tool that reads search results is measuring a different thing from one that reads citations. Counts also vary widely by engine; see [how many sources each engine cites](https://underneath.agency/resources/how-many-sources-ai-search-engines-cite).

## Do the differences follow a pattern?

Yes. The API and the app leaned toward different kinds of sources, which can bias a tracking report.

For ChatGPT, the app favored technology publishers such as TechRadar and Tom’s Guide, while the API showed more manufacturer websites. For Gemini, YouTube was 44.4 percentage points and Reddit 29.1 points more likely to appear through the API. Gemini’s API also mixed far more types of sources: 69.9% of its answers combined three or more types, against 19.0% in the app.

The APIs also showed sources more often: by 9.7 percentage points for ChatGPT. A brand team reading API-based reports could therefore over- or underestimate how often review sites, forums or its own site reach real customers. Separate engines also cite largely different pages, as our guide to [whether AI engines share sources](https://underneath.agency/resources/do-ai-search-engines-cite-the-same-sources) shows.

## Is API data useless for tracking AI visibility?

No. It is useful for controlled, repeatable tests, as long as you do not present it as what customers see.

APIs let you fix the model and settings, and some report steps such as the searches an engine ran. In [our hidden searches study](https://underneath.agency/research/ai-hidden-searches-study), ChatGPT’s API reported running a mean of 3.7 searches per answer. And none of the 509 searches logged across the assistants repeated the user’s question word for word. That shows how an engine looks for sources. We note in that study that these are API models, and the consumer apps may search differently.

API results were not reliably steadier either. Across repeated requests, ChatGPT’s API shared more cited websites between runs than the app did (37.9% against 26.0%), but fewer exact pages. Repeats matter on any channel: in the ChatGPT app, the number of different websites seen for a question rose from 2.65 after one request to 5.56 after three. We explain that built-in churn in [why AI citations keep changing](https://underneath.agency/resources/why-ai-search-citations-change).

## How do researchers capture what consumers actually see?

By automating the consumer apps themselves, logged out, and recording the page a user would see. It is harder but closer to reality.

An audit of AI-generated sources by [Allaham and Diakopoulos](https://arxiv.org/abs/2605.23684) used the chat interfaces of ChatGPT, Copilot, Gemini and Perplexity for this reason, collecting 26,266 unique cited addresses. [Schulte and colleagues](https://arxiv.org/abs/2604.07585) kept their ChatGPT analysis to one data source after finding that mixing API and app data would be inconsistent. A spurious image-delivery address had made up 5.8% of ChatGPT citations in part of their data. A small study discussed by [Martinez](https://arxiv.org/abs/2609.06811) compared an interface and an API directly. Its 234 usable runs showed differences in which providers were named under some conditions, too few to generalize.

Our own studies mix both channels, and we label them. For example, [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study) collected ChatGPT and Gemini answers from their consumer apps through a data provider, and Perplexity through its API.

## What should you do about it?

Ask every vendor or team how the data is collected, and match the channel to the question. In practice:

1. **Ask whether the numbers come from the app or the API,** for each engine, and which model and settings.
2. **Use app-based data for “what do customers see?”** questions, such as share of voice and cited sources.
3. **Use API data for controlled experiments,** where you need fixed settings and repeatable runs, and label it as such.
4. **Check whether web search was on** for every API answer, and how often answers came back without sources.
5. **Never mix channels in one trend line.** A switch from app to API can look like a sudden change in visibility.
6. **Spot-check the app by hand** each month for your most important questions, and compare with the tool.
7. **Ask each question more than once** on either channel, since one request shows only part of the sources.

If you want help setting up monitoring this way, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

Only one study has compared app and API sources head to head for ChatGPT and Gemini. Open questions:

- **Why they differ.** Model version, search settings and hidden system instructions were not separated.
- **Other places and topics.** The test used 117 general product questions, asked from one country, the Netherlands, at one point in time.
- **Signed-in users.** Logged-out apps may differ from what people see with an account and chat history.
- **Brand mentions.** The head-to-head comparison focused on cited sources, not on which brands the app and API recommended.
- **Change over time.** Both apps and APIs update often, so the size of the gap may move.

## Frequently asked questions

### Do ChatGPT API results match the ChatGPT app?

Not closely, for cited sources. In a 2026 test with 117 product questions, the app and API shared on average 12.0% of cited websites, and nothing at all in 60.9% of pairs.

### Why do AI visibility tools use APIs?

Because APIs are easy to automate at scale and let you fix the model and settings, which makes runs comparable. The cost is that the results may not reflect the consumer app your customers use.

### Is Gemini’s API closer to the Gemini app than ChatGPT’s?

Somewhat, but still far apart. Gemini’s app and API shared 14.8% of cited websites on average and no website at all in 43.4% of pairs.

### Should we stop using API-based tracking?

No, but label it and do not mix it with app data. Use it for controlled tests, and confirm key findings against the consumer apps.

## Sources

- Uberti-Bona Marin, L. G., Bertaglia, T., Astante, G., Rijsbosch, B., van Dijck, G., Hannák, A., Spanakis, G. and Kollnig, K. (2026), ["If I Had to Buy Just ONE: Galaxy S26 Ultra": Auditing AI-Generated Product Recommendations](https://arxiv.org/abs/2609.18729), arXiv:2609.18729.
- Kumar, P. (2026), [Generative Engine Optimization at Scale: Measuring Brand Visibility Across AI Search Engines](https://arxiv.org/abs/2606.20065), arXiv:2606.20065.
- Sielinski, R. (2026), [From Stochastic to Stable: Rank Stability and Structural Sufficiency in AI Visibility Measurement](https://arxiv.org/abs/2607.10341), arXiv:2607.10341.
- Allaham, M. and Diakopoulos, N. (2026), [Synthetic Sources?: Auditing Generative Search Engine Citations for Evidence of AI-Generated Sources](https://arxiv.org/abs/2605.23684), arXiv:2605.23684.
- Schulte, J., Bleeker, M. and Kaufmann, P. (2026), [Don’t Measure Once: Measuring Visibility in AI Search (GEO)](https://arxiv.org/abs/2604.07585), arXiv:2604.07585.
- Martinez, O. (2026), [Measuring GEO Visibility: Prompt Corpora Define the Answer Market](https://arxiv.org/abs/2609.06811), arXiv:2609.06811.
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)

---

This is the Markdown twin of https://underneath.agency/resources/api-ai-visibility-monitoring. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How API companies win developers when AI agents pick the API"
description: "By being the API that AI assistants and coding agents choose and can integrate on their own, so usage starts in code and grows into enterprise contracts."
canonical: "https://underneath.agency/resources/api-companies-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How do API companies win developer demand when AI agents choose the integration?

By being the API that AI assistants name and coding agents can actually integrate without help, because for API companies the first integration decides the revenue that follows. Agents now write much of the integration code, read most of the documentation and increasingly call the APIs themselves. An API that is hard for a machine to find, read or test risks being replaced by a rival, or by custom code, before a developer ever compares the two.

## The short version

1. [Gartner](https://www.gartner.com/en/newsroom/press-releases/2024-03-20-gartner-predicts-more-than-30-percent-of-the-increase-in-demand-for-apis-will-come-from-ai-and-tools-using-llms-by-2026) predicted that more than 30% of the increase in demand for APIs would come from AI and tools built on large language models by 2026.
2. In [Postman’s 2025 survey](https://www.postman.com/state-of-api/2025/) of over 5,700 developers, architects and executives, 89% of developers used AI, but only 24% designed APIs with AI agents in mind; 70% knew of the Model Context Protocol (MCP), and only 10% used it regularly.
3. When [Amplifying](https://amplifying.ai/research/claude-code-picks/report) asked Claude Code open-ended questions such as “add user authentication,” it chose Stripe for 91% of payment picks, and for email it chose Resend 62.7% of the time against 6.9% for SendGrid.
4. Machines already read payments documentation more than people: on fintech and payments docs sites hosted by [Mintlify](https://mintlify.com/data.md/), agents made 70% of requests in August 2026, up from 20.5% in February.
5. APIs are a revenue line, not a side project: 65% of organizations in the Postman survey earned revenue from their API programs, and a quarter of those earned more than half their total revenue from them.

## Which developers and teams choose an API, and what does an account grow into?

Developers pick the API, usually on a free or sandbox account, and the company pays as usage grows.

Developers judge a technology by its API first. In the [2025 Stack Overflow Developer Survey](https://survey.stackoverflow.co/2025/work), developers ranked an easy-to-use API as the top reason to endorse a technology and a robust, complete API second, ahead of reputation and cost. The 2024 survey found 75% of developers were more likely to endorse a technology if it provided good API access. Their endorsement matters: 48% of developers said they endorsed or influenced the purchase of new technology in their organization in the past year.

The commercial model rewards that first choice for years. [Twilio](https://www.twilio.com/en-us/press/releases/Q4-full-year-2025-earnings) reported more than 402,000 active customer accounts at the end of 2025, counting any account with at least $5 of revenue in the month, and full-year revenue of $5.07 billion. Its dollar-based net expansion rate was 108% for the year, up from 104%: existing accounts spent more than they did a year before. Many accounts start small, and the money comes from the ones that grow.

For API platform vendors, the companies that sell API management, gateways and tooling, the buyer is a platform or architecture team, and cloud defaults weigh heavily. In the Postman survey, AWS API Gateway was used by 47% of respondents and Azure’s by 26%, and 31% of organizations ran more than one gateway. Postman found 82% of organizations had adopted some level of an API-first approach, with 25% fully API-first.

The stakes for the customer are rising too. Among organizations earning API revenue, 74% made at least 10% of their total revenue from APIs, Postman found. Choosing the wrong provider is a long-lived mistake, which is why developers test before they commit.

## At which points do assistants and agents already touch API selection?

In three places: answering developers’ questions, writing integration code, and calling APIs directly.

**Answering questions.** Developers ask assistants which API to use and how to use it. Postman found 41% of respondents used AI to generate API documentation, and its platform saw 7.53 million calls made to AI APIs in 12 months, up 40% year over year.

**Writing the integration.** Coding agents increasingly choose and install the provider themselves. In Amplifying’s test of 2,430 open-ended prompts, Claude Code recorded what it installed, configured and committed, not just what it suggested. For how AI shortlists work in software buying more broadly, see [our article on B2B SaaS](https://underneath.agency/resources/b2b-saas-revenue-from-ai-search).

**Reading the documentation.** Mintlify, a documentation platform, reports that agents made 61.87% of requests across the docs sites it hosts in August 2026. Fintech and payments was the most agent-read industry it tracks, at 70%. These are Mintlify’s own measurements of its own customers. The same agent readership shapes how [developer tool companies win users](https://underneath.agency/resources/developer-tools-ai-search) beyond APIs.

**Calling the API.** Gartner’s prediction and Postman’s findings point the same way: agents are becoming API consumers in their own right. Postman’s report puts it bluntly: “Agents are already calling your APIs, with or without MCP.” In the same survey, 51% of developers cited unauthorized agent access as a top security risk.

## Which questions lead developers and platform teams to an API?

Category, alternatives, comparison, pricing and compliance questions, plus requests that leave the choice to an agent.

The agent requests are observed. Amplifying used real prompts that name no provider, such as “add user authentication” and “i need a database, what should i use,” and recorded what the agent chose.

We wrote the developer and platform-team questions below as examples of how API choices get framed; they are not observed data:

- Category: “Best SMS API for sending two-factor codes in India and Brazil.”
- Alternatives: “Cheaper alternatives to Twilio for transactional SMS.”
- Comparison: “Stripe vs Adyen for marketplace payouts in Europe.”
- Pricing: “Which email API has the most generous free tier for a startup?”
- Compliance: “Which payments APIs are PCI DSS Level 1 and support 3-D Secure?”
- API platform: “API gateway that runs in our own Kubernetes cluster with OAuth and rate limiting.”

The first five are usually asked by a developer and settled in code. The last is asked by a platform team and settled in a procurement process, where we infer an AI answer helps decide which vendors get an evaluation slot.

## How does a coding agent’s choice turn into API usage and contracts?

Through integration: the agent picks the API, the code ships, usage grows, and the account becomes a contract.

**Pick and test.** The faster an agent can get working keys, the more likely the integration completes in one session. Stripe’s [documentation for AI agents](https://docs.stripe.com/building-with-llms) tells agents to run a command “to get working API keys without signing up,” then install Stripe’s MCP server and agent skills to build the integration. That is documented by Stripe for its own product; it does not say how agents choose between providers.

**Ship.** Once an API is in production code, replacing it means rewriting and retesting the integration. We infer that this switching cost is why the first pick is so valuable for API companies: the agent’s choice tends to persist.

**Grow.** Usage-based pricing means revenue follows the customer’s growth. Twilio’s 108% net expansion shows how much of an API company’s growth comes from accounts already integrated.

**Contract.** As volume rises, the account moves to committed pricing, security reviews and procurement. By then the comparison has usually already happened, often in an assistant or an agent, long before sales was involved.

## What decides whether an AI agent picks your API?

No platform documents it; studies point to an established position in the stack and documentation agents can use.

**Documented by the platform.** Anthropic describes MCP as “a universal, open standard for connecting AI systems with data sources,” which [it introduced in 2024](https://www.anthropic.com/news/model-context-protocol). No assistant or coding agent vendor publishes how it chooses between competing APIs.

**Observed in studies.** In Amplifying’s Claude Code study, no other payment processor was recommended as the primary pick, though Paddle, LemonSqueezy and PayPal appeared as second choices. Email was more open: Resend won 62.7% of primary picks, while SendGrid was offered as an alternative 55 times but chosen as the primary pick only 6.9% of the time. Custom code built by the agent took 21.6% of email picks. The project’s existing stack mattered more than how the request was worded.

A [2026 study of 37,927 agent journeys](https://arxiv.org/abs/2609.34951) by ora research, the company behind the agent-readiness score it tested, found that businesses whose sites agents could read were clearly recommended in 20 against 11 percent of runs, a 1.9 times gap. Its readiness score included whether pricing and API documentation were reachable. The study covered businesses in general, not API companies alone; our summary is in [our article on agent-readable websites](https://underneath.agency/resources/do-ai-agents-recommend-readable-websites).

**Trust factors specific to APIs.** We infer that agents, like developers, favor APIs with complete reference documentation, a published machine-readable specification, working code samples in common languages, sandbox access, clear rate limits and pricing, and a public status page. Postman found 55% of teams struggled with inconsistent documentation and 34% could not find existing APIs, even inside their own organizations. If teams cannot find an API, a reasonable expectation is that an agent will struggle too.

## What does an API company lose when the agent integrates a rival?

The integration, and every dollar of usage that would have followed it.

Being well known is not enough. SendGrid, one of the best-known email APIs, appeared as an alternative 55 times in Amplifying’s email answers, yet the code the agent wrote used Resend or custom code in most cases. An alternative pick is a mention; a primary pick is an integration.

Some categories also look closed to newcomers in agent answers. For payments, Stripe took 91% of primary picks and no rival took any, though the remaining picks were mock interfaces rather than other providers. We infer that challengers in such categories need to win on specific stacks, regions or use cases where the default does not fit, and to make that fit easy for agents and assistants to see.

The loss is invisible in most dashboards: an agent that wires a competitor’s SDK into a codebase leaves no visit and no lead in your analytics, a blind spot covered in [why analytics miss AI visibility](https://underneath.agency/resources/why-analytics-miss-ai-visibility).

## How does GEO work for an API company or API platform?

It makes your API easy to find, understand, test and trust for assistants and agents; it cannot guarantee a pick.

For an API business, whether it sells the endpoints or the platform that manages them, generative engine optimization (GEO) typically involves six pieces of work:

1. **Agent-readable documentation.** Serve reference docs and guides as clean text or Markdown, publish an llms.txt and an OpenAPI specification, and keep every endpoint, error code and limit current. In our study of 5,902 top websites, only [11.5% published a valid llms.txt](https://underneath.agency/research/llms-txt-adoption-study).
2. **A path agents can complete.** Offer sandbox keys, a command-line setup and, where it fits, an MCP server, so an agent can go from question to working call in one session.
3. **Quickstarts for the stacks that matter.** Since context drives agent picks, publish working examples for the frameworks and languages your customers use.
4. **Independent technical coverage.** Earn mentions in tutorials, comparison write-ups, community answers and integration marketplaces, where assistants that search the web look for evidence. Our guide to [building authority for AI search](https://underneath.agency/resources/how-brands-build-authority-for-ai-search) goes further on earning that coverage.
5. **Clear pricing and compliance facts.** State per-unit prices, free-tier limits, regions and certifications on public pages, so assistants quote them correctly; when [our pricing study](https://underneath.agency/research/ai-pricing-accuracy-study) checked four assistants on 45 software products, only 61.9% of the plan prices they quoted fully matched the official page.
6. **Measurement across assistants and agents.** Track how often you are the primary pick, an alternative or only a mention, by stack and region, over repeated runs, and compare with new API keys, first calls and expansion revenue.

For API platform vendors selling to architecture teams, the same work applies to comparison, deployment-model and compliance pages, since those are the questions that decide who gets evaluated.

## What don’t we know yet about how agents and assistants choose APIs?

Whether better documentation or an MCP server raises an API’s pick rate is not yet measured.

The agent-choice data comes mainly from one research group testing JavaScript and Python projects with one or two agents. Mintlify measures its own customers; ora research tested its own readiness score. Postman and Gartner describe demand and practice, not how agents choose. No published study yet links an API company’s visibility in AI answers or agent picks to new keys, usage or contract value. Treat those links as hypotheses to test against your own data.

## How can an API company see whether agents are choosing it or a rival?

Run the requests developers give coding agents for your category, and record which API ends up in the code.

The audit should put your category’s real developer questions and “add this feature” requests to the main assistants and agents, in the stacks and regions you sell to, and compare the results with new keys, first calls and expansion revenue. It shows where you are the default, where you are only an alternative and where an agent writes custom code instead. To scope that audit against your developer funnel and enterprise contract pipeline, [ask us for a review of your API’s visibility to agents and assistants](https://underneath.agency/contact). The [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page describes the technical side of that engagement, including agent-readable documentation, quickstarts for key stacks and tracking of primary picks.

## Frequently asked questions

### Do AI coding agents choose which API a developer uses?

Often, when the developer leaves the choice open. In Amplifying’s tests, agents installed and configured a specific provider in response to requests that named none, and in some categories one provider won almost every time.

### Should API companies build an MCP server?

It is worth testing. MCP gives agents a standard way to call your API, but Postman found only 10% of developers used it regularly in 2025, and no study yet shows that an MCP server raises how often an API is chosen.

### Is being listed as an alternative in AI answers valuable?

Less than it looks. In the email category Amplifying tested, a well-known provider appeared as an alternative 55 times but was the integrated choice in only 6.9% of picks. Track primary picks separately from mentions.

### How should an API platform vendor think about AI search differently?

API management is bought by architecture and platform teams through evaluations, so the questions that matter are comparisons, deployment options and compliance. Make those facts public, consistent and easy to quote.

## Sources

- Gartner (2024), [Gartner Predicts More Than 30% of the Increase in Demand for APIs Will Come From AI and Tools Using LLMs by 2026](https://www.gartner.com/en/newsroom/press-releases/2024-03-20-gartner-predicts-more-than-30-percent-of-the-increase-in-demand-for-apis-will-come-from-ai-and-tools-using-llms-by-2026)
- Postman (2025), [2025 State of the API Report](https://www.postman.com/state-of-api/2025/)
- Nordic APIs (2025), [A Deep Dive Into the State of the API 2025](https://nordicapis.com/a-deep-dive-into-the-state-of-the-api-2025/)
- Amplifying (2026), [What Claude Code Actually Chooses](https://amplifying.ai/research/claude-code-picks/report)
- Mintlify (2026), [Data: agents vs human traffic](https://mintlify.com/data.md/)
- Stack Overflow (2025), [2025 Developer Survey: Work](https://survey.stackoverflow.co/2025/work)
- Twilio (2026), [Twilio Announces Fourth Quarter and Full Year 2025 Results](https://www.twilio.com/en-us/press/releases/Q4-full-year-2025-earnings)
- Stripe (2026), [Agents and AI on Stripe](https://docs.stripe.com/building-with-llms)
- Anthropic (2024), [Introducing the Model Context Protocol](https://www.anthropic.com/news/model-context-protocol)
- Finder, Elovic, Shalev and Yosef (2026), [AX is the New AEO](https://arxiv.org/abs/2609.34951)
- Underneath (2026), [llms.txt adoption](https://underneath.agency/research/llms-txt-adoption-study) and [software pricing accuracy](https://underneath.agency/research/ai-pricing-accuracy-study)

---

This is the Markdown twin of https://underneath.agency/resources/api-companies-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How can an apparel brand win new customers through AI search?"
description: "By publishing the fit, size, fabric, care and returns facts AI tools need to match clothes to a shopper, and earning independent mentions that confirm them."
canonical: "https://underneath.agency/resources/apparel-brands-customers-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How can an apparel brand win new customers through AI search?

By making the facts that decide a clothing purchase, fit by size, fabric, care and returns, complete and consistent everywhere AI tools look, and by earning independent mentions that back them up. AI is not yet where most Americans start looking for clothes, but it is where many check fit, compare products and narrow a choice. For an apparel brand, the payoff is a new customer who keeps what they bought and comes back.

## The short version

1. Fit is the costly problem: [Coresight Research estimates](https://fashionunited.com/news/business/sizing-intelligence-is-strategic-priority-as-brands-prepare-for-ai-driven-commerce/2026063073231) the US online apparel return rate reached 23.4% in 2025, in an online apparel and footwear market worth $201.1 billion, and nearly 70% of shoppers who returned clothes cited size and fit.
2. Discovery through AI is still small in the US, help with fit is not: in a [YouGov poll](https://yougov.com/en-us/articles/54891-are-clothes-shoppers-ready-for-ai-in-apparel-retail) of 957 US clothes shoppers, only 6% wanted to use tools like ChatGPT or Gemini to discover clothing, but 25% wanted AI size and fit recommendations.
3. Product facts do not line up across channels: only 14% of UK shoppers in the [Connected Consumer 2026 report](https://athoscommerce.com/news/ai-fragmented-discovery-and-rising-consumer-expectations-are-reshaping-fashion-ecommerce/) said fashion product details always match across social platforms, marketplaces and retailer sites.
4. AI answers favor known names: [EMARKETER’s index](https://www.emarketer.com/content/ai-visibility-index-apparel-fashion-insights-q3-2026) of ChatGPT recommendations across 20 apparel and fashion categories found Nike mentioned in 12% of them in August 2026.
5. Specific brands can still win: [Shopify reports](https://www.shopify.com/news/agentic-holiday-2026) that Cottonique, a hypoallergenic apparel brand, grew AI-referred sales 276% year over year.

## Who buys everyday apparel, and what is a new customer worth?

Shoppers replacing basics they wear constantly; a customer who finds a good fit tends to buy the same item again.

Everyday apparel, the T-shirts, jeans, underwear, hoodies and work pants people wear every week, is bought differently from trend-led fashion. The shopper usually knows what they need. The questions are whether it will fit, how the fabric feels and lasts, how to care for it, what it costs and whether returning it is easy. [Trend-led style discovery](https://underneath.agency/resources/fashion-ecommerce-ai-recommendations) is a separate question; this article is about the product facts.

The value of a new customer is in repeat purchases of something that fits. We found no public lifetime value figure for apparel brands, so we do not invent one. What is public is the cost of getting fit wrong. Coresight’s 23.4% return rate means that, for online clothing, roughly one item in four comes back, and size and fit cause most of those returns. A brand that wins a customer with an accurate answer, and keeps the sale, earns twice: the first order and the next ones. Whether AI can also lower what you pay for that customer is the subject of [our guide to DTC acquisition costs](https://underneath.agency/resources/dtc-brands-ai-search-sales).

Bodies and needs are also changing. In a Coresight survey reported alongside its sizing study, 70% of US GLP-1 users said they had dropped at least one clothing size. Every size change is a moment when a shopper looks for something new and asks what will fit. Shoes raise the same fit questions; see [how footwear brands get found](https://underneath.agency/resources/footwear-brands-ai-search).

## Where does AI already sit in how people shop for clothes?

Mostly in research and fit questions, not yet as the main place Americans discover clothing.

The evidence is mixed, and the mix is useful. YouGov’s May 2026 poll found US clothes shoppers prefer to discover new styles and brands by browsing in stores (60%), on retailer websites or apps (46%), through friends or family (40%) and through search engines (37%). Interest in general AI tools for discovery was just 6%. But a quarter wanted AI help with size and fit, and 26% wanted it to check stock. Their biggest worries were privacy (51%) and the accuracy of recommendations (47%).

Other surveys find more use. In the UK, the Connected Consumer 2026 report, based on 2,000 consumers, found 60% use tools such as ChatGPT, Claude or Gemini at least occasionally while shopping for fashion, mostly to find deals, compare products and summarize reviews. Coresight found 58% of US consumers familiar with AI had used or intended to use AI tools for shopping. In Global Payments’ survey of more than 16,000 consumers, [reported by FashionUnited](https://fashionunited.com/news/retail/us-consumers-would-let-ai-agents-buy-clothes-but-not-without-a-say/2026092974919), 42% worried an AI agent could buy the wrong item, which the report notes is a particular risk for clothing because fit and sizing vary between brands.

The AI shopping tools themselves are built around clothing. Google says its [AI Mode shopping experience](https://blog.google/products-and-platforms/products/shopping/google-shopping-ai-mode-virtual-try-on-update/) draws on a Shopping Graph of more than 50 billion product listings, with more than 2 billion refreshed every hour, and lets shoppers virtually try on billions of apparel listings using their own photo. OpenAI documents a “Try on” button for clothes and accessories in ChatGPT. For accessories such as jewelry, shoppers ask about certification, sourcing and gifts instead, as covered in [how jewelry brands get recommended](https://underneath.agency/resources/jewelry-brands-ai-product-recommendations). Across all US retail, [Adobe’s data](https://techcrunch.com/2026/04/16/ai-traffic-to-us-retailers-rose-393-in-q1-and-its-boosting-their-revenue-too/) shows AI traffic rose 393% in the first quarter of 2026 compared with a year earlier.

Our reading: AI is a growing research and decision tool for apparel, used more for “will this fit and is it worth it” than for “show me something new.” That is exactly where product data decides the outcome.

## Which questions do apparel shoppers ask AI?

Questions about fit, fabric, care, specific problems and returns, often in the shopper’s own words.

We wrote the prompts below to illustrate the questions; they are not observed queries:

- Fit: “Jeans for athletic thighs and a small waist, 32-inch inseam, that don’t gap at the back.”
- Fabric: “Heavyweight 100% cotton T-shirt that won’t shrink or twist after washing.”
- Care: “Work pants I can machine wash and wear without ironing.”
- Problem: “Wire-free bra for sensitive skin that doesn’t itch.”
- Size change: “I dropped two sizes; which brands run true to size in women’s trousers?”
- Returns: “Which basics brands offer free returns and exchanges?”

The problem-style question matters most for smaller brands. Shopify quotes Cottonique’s co-founder: “Nobody with eczema types ‘hypoallergenic bra.’” Its report describes how an AI that can parse “my bra makes my skin burn” needs a brand that actually solves that problem, and Cottonique’s AI-referred sales grew alongside a 52% rise in total sales. That is one brand’s reported experience, not a general result, but it shows how AI answers can connect a plain-language need to a specific product.

## How does an AI answer turn into an apparel sale that sticks?

Through a named product, a listing with the right size and price, a try-on, a purchase, and no return.

**The answer.** The assistant or AI Mode panel names a few products. Google says each listing in its Shopping Graph carries details like reviews, prices, color options and availability.

**The listing and try-on.** The shopper sees sizes, prices and stock, and may try the item on virtually. OpenAI warns that try-on images “do not guarantee fit or size” and tells shoppers to check the merchant’s measurements, product details and return policy. Those measurements are your data.

**The purchase.** It may happen on your site, at a retailer or marketplace, or through agentic checkout: Google describes tracking a price and tapping “buy for me” to complete checkout on the merchant’s site with Google Pay.

**The keep or return.** This is where apparel differs from most categories. With size and fit behind most clothing returns, a reasonable expectation is that answers based on accurate measurements lead to fewer returns, and answers based on vague size labels lead to more. We found no published data on return rates for AI-referred apparel orders.

**The repeat.** A basic that fits is bought again. Aviator Nation’s ecommerce director told Shopify that someone might first meet the brand in person, then use its site, social media “or ChatGPT to find the product afterward.” AI is one touchpoint in a longer relationship, not the whole of it.

## What decides which clothing brands AI names?

Platforms document product data, reviews and price; studies show known brands and outside sources carry weight.

**Documented by the platforms.** OpenAI says ChatGPT considers structured metadata from first-party and third-party providers, such as price and product description, along with other third-party content and reviews, when choosing products. Google’s [Merchant Center rules](https://support.google.com/merchants/answer/6324492?hl=en) make the size attribute required for free listings of clothing products, and Google [recommends the material attribute](https://support.google.com/merchants/answer/6324410?hl=en) “if customers might search for your product by material or if they might decide to buy your product based on the material.” Google also says AI Mode uses query fan-out, running several searches at once to work out what makes a product right for a request.

**Observed in studies.** EMARKETER’s apparel index says AI recommendations “favor established brands,” that visibility varies sharply by category, and that retailer and publisher sources shape which brands surface. Our own [Reddit citations study](https://underneath.agency/research/ai-reddit-citations-study) found Google’s AI cited Reddit in 17.9% of AI Overviews; community discussion of fit and durability is part of what these systems read. A separate piece asks [whether Reddit shapes Google AI Overviews](https://underneath.agency/resources/does-reddit-shape-google-ai-overviews).

**Industry argument, not platform documentation.** Coresight’s report argues that brands with inconsistent or incomplete sizing data “risk becoming less visible” as AI agents recommend products. That is a research firm’s view, written with a sizing technology company, not a documented ranking rule. We think it is a reasonable expectation, because assistants can only match a fit question to fit data that exists. Makeup brands face the same matching problem with shade, as [makeup recommendations from AI](https://underneath.agency/resources/cosmetics-brands-product-discovery-ai-search) shows.

**Trust factors specific to apparel.** Accurate size charts with garment measurements, honest fit notes (runs small, relaxed through the hip), fabric composition and weight, stretch, shrinkage and care instructions, clear return and exchange terms, reviews that mention fit and how pieces hold up after washing, and editorial and community mentions that confirm all of it. Our article on [what drives AI product recommendations](https://underneath.agency/resources/what-drives-ai-product-recommendations) covers the wider evidence.

## What does an apparel brand lose when AI leaves it out or gets it wrong?

New customers at the moment of a fit question, and margin when wrong answers lead to returns.

If an answer names three brands for “jeans that don’t gap at the waist” and yours is not one of them, you lose a shopper who might have stayed for years. We infer this from how answers narrow choices; no public data measures it for apparel. If an answer names your product with wrong size or fabric facts, the cost arrives later, as a return. The inconsistency problem is real: with only 14% of UK shoppers seeing product details that always match across channels, a brand’s marketplace listings, retailer pages and own site often disagree, and an assistant may repeat whichever version it finds. AI-driven sales can also be hard to see; our article on [why analytics miss AI visibility](https://underneath.agency/resources/why-analytics-miss-ai-visibility) explains why.

## How does GEO work for an apparel brand?

It makes your fit, fabric and returns facts easy for AI to find and trust, on your site and elsewhere.

For a clothing label, generative engine optimization (GEO) means getting AI answers to name your pieces and describe their fit, fabric and care correctly. No one can buy or promise a recommendation. For apparel it usually covers:

1. **Fit data by size.** Garment measurements for every size, fit notes and model sizing, on product pages and in feeds, not only in a generic chart.
2. **Fabric and care facts.** Composition, weight, stretch, shrinkage and care, written the same way everywhere. Our guide to [what product content AI shopping assistants prefer](https://underneath.agency/resources/product-content-ai-shopping-assistants-prefer) explains why specific facts beat adjectives.
3. **Complete feeds.** Size, material, color and availability attributes filled in for Google and other channels, kept current.
4. **Consistency across channels.** Matching facts on your site, retailers and marketplaces, so assistants do not see conflicting versions. We weigh the marketplace side in [whether AI search reduces dependence on marketplaces](https://underneath.agency/resources/ai-search-marketplace-dependence).
5. **Answers to problem questions.** Plain pages that use the words customers use, such as itchy seams, gaping waistbands or shrinking hems, and say honestly which products solve them.
6. **Reviews and community.** Honest customer reviews that mention fit and wear, and genuine participation where people discuss basics. Never fake reviews or planted posts.
7. **Independent coverage.** Editorial roundups and product tests earned through digital PR, so the facts are confirmed by someone other than you. Our article on [which pages to target](https://underneath.agency/resources/best-of-lists-ai-recommendations) explains why.
8. **Measurement.** Tracking a fixed set of fit, fabric and problem questions across ChatGPT, Google AI Mode, Gemini and Perplexity over time.

## What don’t we know yet about AI and apparel sales?

How much apparel revenue AI drives today, and whether AI-referred orders are returned less often.

The surveys disagree on adoption: 6% interest in AI discovery among US clothes shoppers in YouGov’s poll, against 60% occasional use for fashion among UK consumers in a report from a company that sells search and merchandising software. The Coresight sizing report was produced with a sizing technology partner. EMARKETER’s index covers ChatGPT only. Cottonique’s and Aviator Nation’s experiences come from Shopify, which sells the tools behind them. Adobe’s traffic figures cover all retail. We found no independent data on conversion or return rates for AI-referred apparel orders, and no platform documentation that says size data affects ranking directly.

## Where should an apparel brand start?

With your best-selling basics: check what AI tools say about their fit, fabric and returns, and where they are wrong.

Pick 20 to 30 questions your customers ask about fit, fabric, care and specific problems, and run them across the main AI assistants and Google’s AI Mode. Record which brands are named, how your products are described, which sizes and measurements are quoted and which sources are cited. Then compare your own site, feeds and retailer listings for conflicts. If you would rather hand that off, [ask us to audit how AI describes your range](https://underneath.agency/contact); we will set out the product-data, review and coverage work most likely to bring in new customers who keep what they buy. Our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page describes how that work on fit data, product feeds and channel consistency is planned and then measured.

## Frequently asked questions

### Do AI shopping tools let customers try on our clothes?

Yes, on some surfaces. Google offers virtual try-on across billions of apparel listings in Search, and ChatGPT offers a “Try on” button for clothes and accessories. Both work from product images, so good images and accurate measurements matter.

### Should we focus on our own site or on marketplaces?

Both, with the same facts. Shoppers and assistants see your products in many places, and inconsistent details are common.

### Is this different from fashion trend discovery?

Yes. Trend discovery is about style and inspiration. Everyday apparel turns on fit, fabric, care and returns, which are matters of accurate product data.

### Can a smaller apparel brand compete with Nike in AI answers?

On broad questions it is hard. On specific needs, such as sensitive skin or unusual fit, a brand that clearly solves the problem has a better chance. Our look at [whether AI assistants favor big brands](https://underneath.agency/resources/do-ai-assistants-favor-big-brands) covers the wider pattern.

## Sources

- FashionUnited, reporting Coresight Research (June 30, 2026), [Sizing intelligence is strategic priority as brands prepare for AI-driven commerce](https://fashionunited.com/news/business/sizing-intelligence-is-strategic-priority-as-brands-prepare-for-ai-driven-commerce/2026063073231)
- YouGov (2026), [Are clothes shoppers ready for AI in apparel retail?](https://yougov.com/en-us/articles/54891-are-clothes-shoppers-ready-for-ai-in-apparel-retail)
- Athos Commerce and Drapers (June 4, 2026), [Athos Commerce Report Reveals AI, Fragmented Discovery, and Rising Consumer Expectations Are Reshaping Fashion Ecommerce](https://athoscommerce.com/news/ai-fragmented-discovery-and-rising-consumer-expectations-are-reshaping-fashion-ecommerce/)
- FashionUnited, reporting Global Payments (September 29, 2026), [US consumers would let AI agents buy clothes, but not without a say](https://fashionunited.com/news/retail/us-consumers-would-let-ai-agents-buy-clothes-but-not-without-a-say/2026092974919)
- EMARKETER (2026), [AI Visibility Index: Apparel and Fashion Insights Q3 2026](https://www.emarketer.com/content/ai-visibility-index-apparel-fashion-insights-q3-2026)
- Shopify (October 6, 2026), [Welcome to the first holiday season of the agentic era](https://www.shopify.com/news/agentic-holiday-2026)
- Google (May 20, 2025), [Shop with AI Mode, use AI to buy and try clothes on yourself virtually](https://blog.google/products-and-platforms/products/shopping/google-shopping-ai-mode-virtual-try-on-update/)
- Google Merchant Center Help (2026), [Size attribute](https://support.google.com/merchants/answer/6324492?hl=en)
- Google Merchant Center Help (2026), [Material attribute](https://support.google.com/merchants/answer/6324410?hl=en)
- OpenAI (2026), [Shopping with ChatGPT search](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search)
- TechCrunch, reporting Adobe data (April 16, 2026), [AI traffic to US retailers rose 393% in Q1, and it’s boosting their revenue too](https://techcrunch.com/2026/04/16/ai-traffic-to-us-retailers-rose-393-in-q1-and-its-boosting-their-revenue-too/)
- Underneath (2026), [Reddit citations in AI Overviews study](https://underneath.agency/research/ai-reddit-citations-study)

---

This is the Markdown twin of https://underneath.agency/resources/apparel-brands-customers-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How AppSec vendors win demand when developers ask AI"
description: "AppSec vendors earn AI-driven demand by being verifiable to two buyers at once: developers who try tools and security leaders consolidating them."
canonical: "https://underneath.agency/resources/application-security-demand-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How do application security vendors win demand when developers and CISOs ask AI?

By being the tool an AI answer can describe accurately to two different buyers: the developer who tries it and the security leader who consolidates the stack. Application security demand is rising with exploited vulnerabilities, open source malware and AI-written code. The vendors that win are the ones whose documentation, research and community proof are easy to find and hard to dispute.

## The short version

1. Developers are part of the buying decision: in the [2025 Stack Overflow Developer Survey](https://survey.stackoverflow.co/2025/work), 48% endorsed or influenced a new technology purchase in the past year, and security or privacy concerns were the top reason to reject a tool.
2. Those developers already work with AI: 84% use or plan to use AI tools, yet 46% do not trust the accuracy of their output, according to [Stack Overflow](https://stackoverflow.co/company/press/archive/stack-overflow-2025-developer-survey/).
3. Demand is driven by real attacks: [Verizon’s 2025 breach report](https://verizon.com/about/news/2025-data-breach-investigations-report) found exploitation of vulnerabilities was the entry point in 20% of breaches, up 34%, and third-party involvement doubled to 30%.
4. Buyers want fewer tools: in [Cycode’s 2025 survey](https://cycode.com/blog/introducing-the-state-of-aspm-2025-report/) of over 700 security leaders, organizations ran an average of 50 AppSec tools and 88% were willing to consolidate within 12 months.
5. Winning accounts expand: [JFrog](https://jfrog.com/press-room/jfrog-announces-second-quarter-2026-results/) reported 97 customers above $1 million in annual recurring revenue and 121% net dollar retention, which it credited partly to demand for software supply chain security.

## Who buys application security, and what is a customer worth?

Two groups buy it: developers who adopt tools in their workflow, and security leaders who fund and consolidate them.

Application security covers static code analysis (SAST), dynamic testing (DAST), open source and dependency scanning (SCA), secrets detection, container and pipeline security, and the posture platforms (ASPM) that pull findings together. Gartner’s 1Q26 forecast, as summarized by [Louis Columbus](https://softwarestrategiesblog.com/2026/04/01/top-10-fastest-growing-security-categories-gartner-2026-forecast/), sizes the application security subsegment at $8.6 billion in 2025 and $9.9 billion in 2026, growth of 13.0%.

The developer side matters more here than in most security categories. Stack Overflow’s survey of more than 49,000 developers found that 48% had endorsed or influenced a technology purchase in the past year. What turns developers away is telling: security or privacy concerns ranked first, prohibitive pricing second and “availability of better alternatives” third.

Pricing is built around developers too. [GitHub](https://github.blog/changelog/2025-03-04-introducing-github-secret-protection-and-github-code-security/) sells Secret Protection at $19 and Code Security at $30 per month per active committer, so contract size grows with the engineering team.

The security side controls the budget and wants consolidation. Cycode surveyed over 700 CISOs, AppSec directors and DevSecOps managers in the US, UK and Germany; 50% came from organizations with over 5,000 employees. 67% said managing their array of tools was a significant hurdle, and 61% had begun consolidating. Cycode sells a posture platform, so it has an interest in that finding.

A won customer can be large and long-lived. Snyk crossed [$300 million in annual recurring revenue](https://www.calcalistech.com/ctechnews/article/bymlg4mejl), with its static analysis product alone at $100 million. JFrog’s customers above $1 million rose to 97 from 61 a year earlier, and it was named a Leader in Gartner’s first Magic Quadrant for Software Supply Chain Security.

## What is pushing companies to buy application security now?

Exploited vulnerabilities, poisoned open source and AI-written code, each documented in 2025–2026 research.

- **Exploits.** Verizon analyzed 12,195 confirmed data breaches and found credential abuse (22%) and vulnerability exploitation (20%) were the leading ways in.
- **Open source risk.** [Black Duck’s 2026 audit](https://www.nasdaq.com/press-release/black-duck-research-shows-open-source-vulnerabilities-have-doubled-ai-accelerates) of 947 codebases found open source in 98% of them and mean vulnerabilities per codebase up 107%.
- **Malicious packages.** [Sonatype](https://www.infosecurity-magazine.com/news/454000-malicious-open-source/) found 454,648 new malicious packages in 2025, as developers downloaded components 9.8 trillion times.
- **AI-written code.** [Veracode](https://www.veracode.com/press-release/ai-generated-code-poses-major-security-risks-in-nearly-half-of-all-development-tasks-veracode-research-reveals/) tested more than 100 AI models on 80 coding tasks; they introduced security flaws in 45% of cases. Veracode, Black Duck and Sonatype all sell application security, so read their figures as vendor research.

These triggers become the questions buyers type, which is where AI search enters.

## When does an AppSec buyer meet an AI answer?

Inside the developer’s daily workflow, and increasingly in Google results for software questions; no study isolates AppSec buyers.

The developers who trial scanners already lean on AI for answers. AI tools were in use or planned by 84% of Stack Overflow’s respondents, up from 76% in 2024. Trust is low: 46% said they do not trust the accuracy of AI output, and 35% visited Stack Overflow after running into problems with AI responses. Developers check answers against communities and documentation; Stack Overflow (84%), GitHub (67%) and YouTube (61%) were the community platforms they used most.

Google’s AI answers lean on those same places for software searches. In [our Reddit study](https://underneath.agency/research/ai-reddit-citations-study), 35.4% of AI Overviews for B2B software and technology keywords cited a Reddit thread, one of the two highest rates of eight industries, and r/cybersecurity was among the most-cited communities. In [our YouTube study](https://underneath.agency/research/ai-overview-youtube-videos-study), AI Overviews for B2B software searches cited a YouTube video on 91.0% of searches, far more than any other industry.

AI answers about code are also fallible in ways that matter to this industry. Sonatype analyzed nearly 37,000 dependency upgrades suggested by AI tools and said 28% were hallucinations, recommending versions that do not exist. Developers have reason to double-check; a vendor whose facts are published clearly gives them, and the assistant, something to check against.

## Which questions do AppSec buyers ask AI assistants?

Developers ask how and which; security leaders ask what to consolidate and how to prove compliance. We wrote the AppSec prompts in this table to show typical needs; none comes from logged queries.

| Asked by | Need | Illustrative prompt |
|---|---|---|
| Developer | Fit | “What is the best SAST tool for a TypeScript and Go monorepo that runs in GitHub Actions?” |
| Developer | Noise | “Which code scanners have the fewest false positives for Java?” |
| Developer | Alternatives | “Open source alternatives to Snyk for dependency scanning?” |
| AppSec lead | Comparison | “Semgrep vs GitHub Code Security vs Checkmarx for a 400-developer team?” |
| AppSec lead | Consolidation | “Do we need an ASPM platform if we already have SAST, SCA and container scanning?” |
| CISO | Compliance | “Which tools generate SBOMs that will help with the EU Cyber Resilience Act?” |
| CISO | AI code | “How do we scan code written by AI coding assistants before it ships?” |

Each prompt can set off several searches. Google documents that AI Overviews and AI Mode [may use a “query fan-out” technique](https://developers.google.com/search/docs/appearance/ai-features) across subtopics, and OpenAI documents that [ChatGPT search rewrites questions](https://help.openai.com/en/articles/9237897-chatgpt-search) “into one or more targeted queries.” For AppSec, we infer those searches will reach language support pages, documentation, benchmarks and community threads, not only vendor home pages.

## How does a scanner named in an AI answer become a platform contract?

Through two linked motions: a developer tries the tool, then the security team standardizes on a platform.

1. **Bottom-up trial.** A developer asks an assistant which scanner suits their stack, installs a free tier or plugin, and sees results in their own code.
2. **Team purchase.** Usage spreads, and the team buys seats; per-committer pricing like GitHub’s ties revenue to how many developers commit code.
3. **Security evaluation.** The AppSec lead, facing dozens of tools, runs a comparison of a few platforms. AI answers and the sources they cite shape which ones make that list, we infer.
4. **Platform contract.** The winner replaces point tools as part of consolidation.
5. **Expansion.** Revenue grows with developers, repositories and modules; JFrog’s 121% net dollar retention shows what expansion looks like at a public vendor.

The AI answer can matter at steps 1 and 3, for different readers. A developer wants to know whether it works with their language and pipeline; a security leader wants coverage, compliance and consolidation. A vendor visible only to one of them, we infer, loses deals at the other step.

## Why does an assistant suggest one code scanner over another?

The platforms do not document vendor selection; studies show AI answers draw on communities, video and independent sources.

**What Google and OpenAI disclose.** Both say their AI answers search the web and link to supporting pages. Neither explains how it picks the SAST, SCA or posture tools it suggests.

**Observed in studies.** Our Reddit study found Google’s AI cites practitioner communities heavily for software questions. Our YouTube study found the text Google shows for a cited video came from what is said in it: for 97.9% of 616 cited videos with an excerpt, the excerpt was not in the video’s description. A recorded walkthrough of how a scanner handles a real vulnerability is, in other words, readable evidence. In [our brand study](https://underneath.agency/research/brand-entity-ai-recommendations-study), coverage on independent sites was the strongest predictor of being recommended that we measured.

**Our inference on AppSec trust signals.** The evidence developers and AppSec leads trust is mostly public:

- documentation that states supported languages, frameworks, package managers and pipelines;
- reproducible accuracy claims, such as published benchmark results or detection examples;
- vulnerability research and disclosures credited to the vendor’s team;
- open source projects, rules or plugins that developers already use;
- analyst placements, such as the 18 companies in Gartner’s software supply chain security report, which [The Stack](https://www.thestack.technology/oss-security-finally-gets-a-magic-quadrant/) reports named eight leaders;
- candid discussion in practitioner communities.

A reasonable expectation is that vendors with more of this in public view give both developers and assistants more to cite. No study has tested it for AppSec.

## What does an AppSec vendor lose when assistants name rivals instead?

Missed trials, which never show up in a pipeline report, and missed consolidation decisions, which last for years.

- **Developers drop tools fast.** “Availability of better alternatives” is the third-ranked reason developers reject a technology. If an assistant names a rival for a developer’s stack, the trial you never got is invisible.
- **Consolidation shrinks the field.** With organizations running an average of 50 tools and 88% willing to consolidate, each consolidation decision removes vendors, and the ones not on the shortlist are the first to go, we infer.
- **Wrong facts cost trials.** If an assistant says you do not support a language you do support, developers may never test you. Our guide to [fixing wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) shows how to trace and correct a claim like that.

No study has put a dollar figure on these losses.

## How does GEO work for an application security company?

It puts your language coverage and detection proof where developers and assistants can check them; nobody can promise a mention.

1. **Documentation as a public asset.** Keep language, framework and integration support on crawlable, current pages. This is the first thing developers and assistants check.
2. **Clear category language.** Say plainly whether you are SAST, SCA, DAST, ASPM, supply chain security or a platform, and use the same words everywhere.
3. **Proof developers can reproduce.** Publish detection examples, benchmark methods and false positive data, with enough detail to test.
4. **Original research.** Vulnerability and malware research earns coverage in security media and developer communities, which assistants can cite. More on that in [how brands build authority for AI search](https://underneath.agency/resources/how-brands-build-authority-for-ai-search).
5. **Video with spoken explanation.** Short walkthroughs of real findings, with captions, give Google’s AI something to quote.
6. **Honest community presence.** Answer questions in developer and security communities as identified staff; planted posts backfire, as we explain in [legitimate GEO versus manipulation](https://underneath.agency/resources/legitimate-geo-vs-manipulation).
7. **Comparison and alternatives pages.** Fair comparisons help buyers, though third-party sources carry more weight; see [whether comparison pages help B2B citations](https://underneath.agency/resources/do-comparison-pages-help-b2b-ai-citations).
8. **Measurement for both buyers.** Track developer and security-leader questions separately, across assistants and over repeated runs.

For the general security buying picture, see [cybersecurity software and AI search](https://underneath.agency/resources/cybersecurity-software-revenue-from-ai-search); for self-serve and seat-based revenue, see [B2B SaaS and AI search](https://underneath.agency/resources/b2b-saas-revenue-from-ai-search).

## What is still unmeasured about AI answers and AppSec purchases?

No one has measured how AppSec buyers use assistants, or whether being named lifts trials or platform wins.

- **No AppSec buyer study.** Stack Overflow covers developers in general, not AppSec purchases.
- **Vendor research dominates.** Most threat figures come from companies that sell AppSec tools; their methods differ.
- **Selection is opaque.** No platform publishes how it chooses which tools to name, and answers change between runs.
- **No study ties visibility to scanner revenue.** The broader evidence is weighed in [does AI visibility drive business results](https://underneath.agency/resources/does-ai-visibility-drive-business-results).

## How can an AppSec vendor tell whether AI answers bring it trials and platform deals?

Test what assistants tell developers and security leaders about your tool, using the questions each one actually asks.

Write two question sets: one for developers (stack fit, accuracy, setup, alternatives) and one for security leaders (consolidation, compliance, AI-written code). Put every question to ChatGPT, Gemini, Perplexity, Copilot and Google’s AI more than once, since the scanners named shift between runs. Note whether your tool appears, how its language support and false positive rate are described, and which pages are cited. Gaps usually point to thin documentation, missing research or little community presence.

For an outside view, [talk to us about an AppSec visibility audit](https://underneath.agency/contact). It shows where AI answers send developers and AppSec leaders instead of you, and which gaps most likely cost you free-tier installs, seat expansion and consolidation wins. The [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page explains how we then build the documentation, detection proof and research coverage that developers and security leaders each check.

## Frequently asked questions

### Do developers trust AI recommendations for security tools?

Not fully: 46% of developers in Stack Overflow’s 2025 survey said they do not trust the accuracy of AI output, so they check documentation and communities.

### Should AppSec vendors target developers or CISOs in AI search?

Both. Developers influence purchases (48% did last year), while security leaders decide consolidation. Their questions differ, so track them separately.

### Does publishing on YouTube help AppSec visibility in Google’s AI?

Possibly. In our study, Google’s AI cited a YouTube video on 91.0% of B2B software searches, quoting what was said in the video.

### Does an analyst report placement matter for AI answers?

Untested, but it is public, citable proof. Gartner’s first software supply chain security report covered 18 companies.

## Sources

- Stack Overflow (2025-07), [2025 Developer Survey: Work](https://survey.stackoverflow.co/2025/work)
- Stack Overflow (2025-07), [2025 Developer Survey: AI](https://survey.stackoverflow.co/2025/ai)
- Stack Overflow (2025-07-29), [2025 Developer Survey press release](https://stackoverflow.co/company/press/archive/stack-overflow-2025-developer-survey/)
- Verizon (2025-04-23), [2025 Data Breach Investigations Report news release](https://verizon.com/about/news/2025-data-breach-investigations-report)
- Cycode (2025-03-13), [Introducing the State of ASPM 2025 Report](https://cycode.com/blog/introducing-the-state-of-aspm-2025-report/)
- JFrog (2026-08-06), [JFrog Announces Second Quarter 2026 Results](https://jfrog.com/press-room/jfrog-announces-second-quarter-2026-results/)
- Louis Columbus, Software Strategies Blog (2026-04-01), [Gartner’s $246.2B Security Forecast shows 10 categories growing 2x to 3x the market](https://softwarestrategiesblog.com/2026/04/01/top-10-fastest-growing-security-categories-gartner-2026-forecast/)
- GitHub (2025-03-04), [Introducing GitHub Secret Protection and GitHub Code Security](https://github.blog/changelog/2025-03-04-introducing-github-secret-protection-and-github-code-security/)
- Calcalist (2024), [Snyk crosses $300 million in annual recurring revenue](https://www.calcalistech.com/ctechnews/article/bymlg4mejl)
- Black Duck, via Nasdaq (2026-02-25), [Black Duck Research Shows Open Source Vulnerabilities Have Doubled as AI Accelerates Code Creation](https://www.nasdaq.com/press-release/black-duck-research-shows-open-source-vulnerabilities-have-doubled-ai-accelerates)
- Infosecurity Magazine (2026-01-28), [Researchers Uncover 454,000+ Malicious Open Source Packages](https://www.infosecurity-magazine.com/news/454000-malicious-open-source/)
- Veracode (2025-07-30), [AI-Generated Code Poses Major Security Risks in Nearly Half of All Development Tasks](https://www.veracode.com/press-release/ai-generated-code-poses-major-security-risks-in-nearly-half-of-all-development-tasks-veracode-research-reveals/)
- The Stack (2026-06-22), [OSS security finally gets a Magic Quadrant](https://www.thestack.technology/oss-security-finally-gets-a-magic-quadrant/)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- OpenAI Help Center (2025), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Underneath (2026), [When does a Reddit thread become evidence in Google’s AI?](https://underneath.agency/research/ai-reddit-citations-study)
- Underneath (2026), [The YouTube videos Google’s AI cites are small](https://underneath.agency/research/ai-overview-youtube-videos-study)
- Underneath (2026), [Do Wikipedia and schema make AI assistants recommend a brand?](https://underneath.agency/research/brand-entity-ai-recommendations-study)

---

This is the Markdown twin of https://underneath.agency/resources/application-security-demand-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Are ads coming to ChatGPT and other AI assistants? | Underneath"
description: "They are already arriving. Researchers report OpenAI and Google have put ads into their AI products, and ChatGPT ads reached 31 European markets in August."
canonical: "https://underneath.agency/resources/are-ads-coming-to-ai-assistants"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Are ads coming to ChatGPT and other AI assistants?

They are already arriving. Researchers report that OpenAI and Google have integrated advertising into their AI products, while Anthropic has stayed ad-free. How far ads will spread inside the answers themselves is still open, and the research suggests competition may hold them back.

## The short version

1. A 2026 Carnegie Mellon paper reports that OpenAI and Google have integrated advertising into their AI products, while Anthropic remains ad-free ([Zhang and colleagues, 2026](https://arxiv.org/abs/2603.29071)).
2. In August, OpenAI expanded ChatGPT advertising to 31 European markets, according to a September 2026 audit that cites other work ([Uberti-Bona Marin and colleagues, 2026](https://arxiv.org/abs/2609.18729)).
3. In a 40-day audit of Google’s AI Overviews, 2.16% of results pages with an Overview also carried a Google sponsored ad, and no ads appeared inside the Overview itself ([Xu, Iqbal and Montgomery, 2026](https://arxiv.org/abs/2605.14021)).
4. US search advertising was a $102.9 billion business in 2024, nearly 40% of all US internet ad revenue, which is the revenue AI answers put at stake ([Khosravi and Yoganarasimhan, 2026](https://arxiv.org/abs/2602.18455)).

## Have AI assistants already started showing ads?

Yes. By 2026, researchers describe ads in OpenAI’s and Google’s AI products as a fact, not a forecast.

In mid-2025, [Erickson](https://arxiv.org/abs/2506.06447) wrote that it “may only be a matter of time” before conversational search carries ads, noting OpenAI was considering it. That was a prediction. By March 2026, [Zhang and colleagues](https://arxiv.org/abs/2603.29071) at Carnegie Mellon described OpenAI and Google as having integrated advertising into their AI products, with Anthropic strictly ad-free.

The spread continued. [Uberti-Bona Marin and colleagues](https://arxiv.org/abs/2609.18729) report that in August, OpenAI expanded ChatGPT advertising to 31 European markets. They cite other work suggesting these ads are most likely to show on product recommendation questions, in 10 to 14% of sessions. That estimate is secondhand, and the paper does not describe the ad format.

On Google’s side, [Khosravi and Yoganarasimhan](https://arxiv.org/abs/2602.18455) note that Google had expanded ad placements in AI Overviews (the AI summary at the top of Google’s results) across devices and countries by late 2025. Google’s AI Mode did not yet show ads when [Wang and colleagues](https://arxiv.org/abs/2608.18352) ran their 1,100-person experiment, and the authors expect that to change.

## Why do AI companies want ads at all?

Because search advertising is enormous, AI answers are expensive to produce, and buying questions are common.

US search advertising generated $102.9 billion in 2024, nearly 40% of all US internet ad revenue. AI answers threaten that model: when the answer appears on the page, fewer people click through to sites that carry ads. Zhang and colleagues add that AI engines face large running costs for every answer, which pushes them toward either ads or paid subscriptions.

Buying questions are already a real share of AI use. The largest published measurement of ChatGPT use from 2025 found product and service recommendations made up roughly 2% of conversations, as reported by Uberti-Bona Marin and colleagues.

## Will ads take over AI answers?

Not necessarily: an economic model suggests competition between AI engines limits how many ads each can show.

[Zhang and colleagues](https://arxiv.org/abs/2603.29071) built a model of an AI engine choosing, question by question, between an answer with ads and one without. They then simulated markets of 500 users over 20 periods, in 160 runs under different conditions. Their conclusions:

- Ads pay only when the immediate ad revenue beats the long-term value of keeping a user happy and moving them toward a paid plan.
- Ad-heavy policies earn more early but shrink the user base and slow subscriptions.
- When rival AI engines are strong, the best policy shifts further toward ad-free answers, because users can switch.

Real users do switch. In a field experiment with 1,100 people, forcing searches into Google’s AI Mode raised use of competing search engines by 11.2% ([Wang and colleagues, 2026](https://arxiv.org/abs/2608.18352)).

This is a theoretical model with simulated users, not data from any real AI engine. It shows the pressure, not the outcome.

## How do ads sit next to Google’s AI answers today?

Mostly beside them, not inside them, in the one large audit we found.

[Xu, Iqbal and Montgomery](https://arxiv.org/abs/2605.14021) captured 7,583 results pages with an AI Overview over 40 days in spring 2026:

- 2.16% also showed at least one Google sponsored ad on the same page.
- 0.51% placed a sponsored ad above the Overview.
- No sponsored ads appeared inside the Overview itself.
- 50.63% of the pages the Overviews cited ran visible display ads of their own.

The authors note their searches were mostly trending, informational ones, where ads are rarer. Commercial searches are likely to carry more ads. How often Overviews appear on commercial searches is measured in [our study of 1,248 US searches](https://underneath.agency/research/ai-overviews-frequency-study). The last point matters for publishers: more than half the cited pages depend on ad-supported visits that an AI answer may replace.

## Would buyers be able to tell an ad from an answer?

Not reliably, if ads are blended into the answer, and the research warns that is the risky case.

Erickson distinguishes clearly labeled banner ads from [“native” ads written into the answer](https://underneath.agency/resources/risks-of-native-ads-in-ai-answers) itself. Erickson cites earlier research finding users may not notice ads woven into conversational answers unless they are clearly disclosed. The paper calls the danger the “fake friend dilemma”: a user trusts the assistant as a helper while it serves other interests.

Presentation already shapes trust. In a September 2026 audit of 1,536 responses, ChatGPT expressed a first-person preference (“my pick would be…”) in 79% of product recommendations, against 7% for Gemini and 2% for AI Overviews. In a separate experiment with 4,927 US adults, adding reference links raised trust in AI answers even when the links were invalid ([Li and Aral, 2025](https://arxiv.org/abs/2504.06435)).

Regulators are moving. Uberti-Bona Marin and colleagues report that on 31 August 2026, the European Commission designated ChatGPT a Very Large Online Search Engine, bringing audit and consumer protection duties. [Wen and colleagues](https://arxiv.org/abs/2606.12439) argue US-style rules requiring “Ad” or “Sponsored” [labels should extend to AI answers](https://underneath.agency/resources/should-ai-search-ads-be-labeled).

## What should you do about it?

Plan for a paid channel inside AI answers, but keep investing in the earned visibility ads cannot replace.

1. **Assign an owner to track AI ad products.** Watch for self-serve ad programs from OpenAI, Google and others, including where they run and what formats they allow.
2. **Keep earned and paid AI visibility separate in reporting.** Once ads appear beside answers, a mention you earned and a placement you bought are different results.
3. **Protect your organic presence in answers.** Modeling suggests ads will be limited where rival engines compete, and Anthropic remains ad-free, so earned visibility will still carry much of the load. We weigh that bet in [planning AI search for five years](https://underneath.agency/resources/ai-search-strategy-next-five-years).
4. **Insist on clear labels if you buy.** Disguised placements risk buyer trust and regulatory attention.
5. **If you publish content, watch traffic.** More than half of pages cited in AI Overviews run display ads, so lost clicks hit ad revenue.

If you want help building earned visibility in AI answers, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The research confirms ads have arrived in some AI products but says little about how well they work.

- No paper we reviewed measures click or conversion rates for ads inside AI assistants.
- The European rollout figures and session estimate are reported secondhand, not measured by the researchers.
- The economic model that predicts limits on ads uses simulated users and assumptions, not platform data.
- The AI Overview ad audit covered mostly trending searches, so ad frequency on commercial searches is unmeasured.
- Whether ad-supported answers become less accurate or less balanced has not been tested.

## Frequently asked questions

### Does ChatGPT show ads?

According to researchers, yes in some markets. A September 2026 audit reports OpenAI expanded ChatGPT advertising to 31 European markets in August.

### Are there ads in Google’s AI Overviews?

Next to them, at times. A 40-day audit found 2.16% of pages with an Overview also showed a sponsored ad, and none inside the Overview itself.

### Will AI assistants become full of ads like search results?

Not necessarily. A Carnegie Mellon model, run over 160 simulations, found that strong rival AI engines push the best policy toward ad-free answers.

### Is Claude ad-free?

As of March 2026, researchers described Anthropic as remaining strictly ad-free while OpenAI and Google had integrated advertising.

## Sources

- Erickson (2025), [Fake Friends and Sponsored Ads: The Risks of Advertising in Conversational Search](https://arxiv.org/abs/2506.06447), arXiv:2506.06447.
- Zhang, Jiao, Li and Xiong (2026), [An Economic Framework for Generative Engines: Advertising or Subscription?](https://arxiv.org/abs/2603.29071), arXiv:2603.29071.
- Uberti-Bona Marin and colleagues (2026), ["If I Had to Buy Just ONE: Galaxy S26 Ultra": Auditing AI-Generated Product Recommendations](https://arxiv.org/abs/2609.18729), arXiv:2609.18729.
- Xu, Iqbal and Montgomery (2026), [Measuring Google AI Overviews: Activation, Source Quality, Claim Fidelity, and Publisher Impact](https://arxiv.org/abs/2605.14021), arXiv:2605.14021.
- Khosravi and Yoganarasimhan (2026), [Impact of AI Search Summaries on Website Traffic: Evidence from Google AI Overviews and Wikipedia](https://arxiv.org/abs/2602.18455), arXiv:2602.18455.
- Wang, Gleason, Bart, Wilson and Metaxa (2026), [AI in Search Reduces Publisher Referrals Without Improving User Experience: Experimental Evidence](https://arxiv.org/abs/2608.18352), arXiv:2608.18352.
- Li and Aral (2025), [Human Trust in AI Search: A Large-Scale Experiment](https://arxiv.org/abs/2504.06435), arXiv:2504.06435.
- Wen and colleagues (2026), [Position: Generative Engine Optimization Creates Underexamined Risks, Governance Must Target Concentration, Disclosure, and Academic Blind Spots](https://arxiv.org/abs/2606.12439), arXiv:2606.12439.

---

This is the Markdown twin of https://underneath.agency/resources/are-ads-coming-to-ai-assistants. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Are AI search engines citing AI-generated content? | Underneath"
description: "Yes. A 2026 audit of four AI search engines flagged about 16% of readable cited pages as AI-generated, from 7.3% for ChatGPT to 27.8% for Copilot."
canonical: "https://underneath.agency/resources/are-ai-search-engines-citing-ai-content"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Are AI search engines citing AI-generated content?

Yes. In a 2026 audit of ChatGPT, Copilot, Gemini and Perplexity on politics, health and environment questions, a detection tool flagged about 16% of the cited pages it could read as AI-generated, and none of the sampled pages disclosed it. Rates differed widely by engine. It is one study using one detector, so treat the figures as an estimate, probably on the low side.

## The short version

1. Across 712 real questions, about 16% of the unique sources four AI search engines cited were judged AI-generated ([Allaham and Diakopoulos](https://arxiv.org/abs/2605.23684)).
2. The share ranged from 7.3% of ChatGPT’s cited sources to 27.8% of Copilot’s, with Gemini at 14.7% and Perplexity at 9.4% (Allaham and Diakopoulos).
3. 97.1% of the AI-generated sources came from outside the 25 most-cited websites, so they sit in the long tail rather than on familiar domains (Allaham and Diakopoulos).
4. A separate detector flagged 8.90% of pages in Google and Gemini results as rewritten to win AI citations, rising to 16.36% among pages updated in 2026 ([Chu and colleagues](https://arxiv.org/abs/2608.16824)).
5. In a simulation, a single planted page at the top of the results led the most vulnerable AI assistants to recommend a fake product in 27% of cases ([Luo and Chen](https://arxiv.org/abs/2606.13610)).

## How much of what AI search cites is AI-generated?

About one in six cited pages, in the only audit that has measured it directly. [Allaham and Diakopoulos](https://arxiv.org/abs/2605.23684) at Northwestern University put 712 real user questions about US politics, health and the environment to ChatGPT, Copilot, Gemini and Perplexity, through each product’s normal interface.

The 2,848 answers cited 26,266 unique web pages. The team could extract text from 72.9% of them and ran that text through Pangram, a detection tool that estimates whether writing was produced by AI. About 16% of the readable cited sources were classed as likely or highly likely AI-generated.

None of it was labeled. The authors checked 200 flagged pages by hand for any statement that AI had been used, and found no such disclosure.

## Which engines cite the most AI-generated pages?

Copilot, by a wide margin, and ChatGPT the least. Each figure below is the share of that engine’s cited sources that the detector flagged.

| Engine | Cited sources flagged as AI-generated |
|---|---|
| Copilot | 27.8% |
| Gemini | 14.7% |
| Perplexity | 9.4% |
| ChatGPT | 7.3% |

ChatGPT had the lowest share despite citing the most pages, an average of 14.7 unique sources per answer. By topic, the paper reports shares of 16.3% for health, 13.3% for the environment and 11.1% for politics.

## Where do the AI-generated sources come from?

Mostly from obscure sites, not the big domains engines cite most. The 25 most-cited websites accounted for 23.8% of all citations, but only 2.9% of the AI-generated sources. The other 97.1% came from the long tail of rarely cited domains.

The problem is spread thin rather than concentrated. 28.0% of all cited domains, 1,754 sites, had at least one cited page flagged as AI-generated. They included niche advocacy sites, Reddit, some news sites and even government health pages. For the kinds of sites engines lean on overall, see [which websites AI search relies on](https://underneath.agency/resources/which-websites-do-ai-search-engines-cite).

One clear case: the eighth-largest source of flagged pages, with 28 cited pages, describes itself on its homepage as an “AI-powered research tool.”

## How reliable are these estimates?

Reasonably, as a lower bound, but they rest on one detector and three topics. The authors tested Pangram first: it classified 100% of 200 human-written test texts correctly. But on 105 articles from sites that journalists had reported as publishing AI-generated news, it missed 31.4%, so the true share of AI-generated citations is probably higher.

Other limits matter. About 27% of cited pages could not be read at all, often PDFs or videos. The questions covered only politics, health and the environment, in English, and the authors stress that AI-generated content is not necessarily low quality. The paper is a preprint.

## Is content written to win AI citations growing too?

It appears so, according to one detection study. [Chu and colleagues](https://arxiv.org/abs/2608.16824), at the CISPA Helmholtz Center for Information Security, built a detector for [pages rewritten to be picked up by AI search](https://underneath.agency/resources/how-much-web-content-is-optimized-for-ai-search), a practice called generative engine optimization (GEO). They ran it on 10,095 pages returned by Google Search and by Gemini for 1,000 real queries.

The detector flagged 8.90% of pages overall: 8.14% of Google Search pages and 9.09% of the pages Gemini cited in its answers. Among pages that declared when they were last updated, the rate rose from 7.02% for 2024 to 12.80% for 2025 and 16.36% for 2026. Only 19.57% of pages carried such a date, so the trend rests on a minority of pages.

The flagged pages’ own evidence was often weak: 69.34% of the citations inside them pointed to sources with little editorial accountability or that were hard to inspect. The authors call all of these estimates, not definitive measurements.

Self-serving content reaches AI answers in other forms too. In [our study of “best of” lists cited by AI](https://underneath.agency/research/self-promoting-best-lists-study), 24.2% of cited numbered lists with an identifiable publisher ranked that publisher first, though such lists made up only 1.1% of all citations.

## Why does this matter for brands?

Because planted or synthetic pages can shape what AI assistants recommend. In a simulation by [Luo and Chen](https://arxiv.org/abs/2606.13610), researchers rewrote real retrieved web pages to promote fake products and tested 12 AI models on 225 real products.

A single altered page at the top of the results fooled the most vulnerable models in 27% of cases. Replacing the top three pages raised the worst model’s rate to 73.8%. The same page placed lower in the results was nearly harmless, with fooled rates of 1–4%.

This was a controlled test, mostly in Chinese, not an audit of live engines. Still, together with the audits above, it shows that your accurate pages compete for citation with mass-produced ones, and that whatever ranks first can carry outsized weight. We look at deliberate misuse in [whether GEO can push false information](https://underneath.agency/resources/can-geo-push-false-information-into-ai-answers).

## What should you do about it?

Know which pages are cited for your category, and give engines better material than the synthetic alternatives.

1. Run your buyers’ key questions through the main assistants and look at who wrote the cited pages. Flag anonymous or obviously machine-written sources describing you or your category.
2. Publish original, checkable material: your own data, named experts and specific facts that a generic AI-written page cannot supply.
3. If you use AI to draft content, have a subject expert edit and fact-check it. Errors on your own pages can be repeated in AI answers with your name attached.
4. Watch for fake reviews, clone sites or misleading comparison pages about your products, and report them to the platforms hosting them.

For help monitoring the sources AI search uses about your brand, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The evidence is early: one citation audit, one page-level detector study and one simulation.

- No study has measured AI-generated citations for product, software or local-business questions.
- No one has tested whether cited AI-generated pages are less accurate than human-written ones.
- We found no evidence on whether engines favor or penalize AI-written pages when choosing what to cite.
- The rising trend in pages written for AI citation is based on self-declared update dates on a fifth of pages, not on citations over time.
- Detectors disagree and miss some AI text, so all shares here are estimates.

## Frequently asked questions

### Does ChatGPT cite AI-generated websites?

Sometimes. In the 2026 audit, 7.3% of ChatGPT’s cited sources were flagged as AI-generated, the lowest of the four engines tested.

### Can tools tell whether a cited page was written by AI?

Partly. The detector used had no false alarms on 200 human-written texts, but it missed 31.4% of articles from known AI content sites.

### Do websites disclose that their content is AI-generated?

Rarely, if ever, in this audit. Of 200 cited pages flagged as AI-generated and checked by hand, none disclosed the use of AI.

### Is AI-generated content more likely to be cited by AI search?

The research has not answered this. Studies so far measure how much AI-generated content is cited, not whether engines prefer it over human writing.

## Sources

- Allaham and Diakopoulos (2026), [Synthetic Sources?: Auditing Generative Search Engine Citations for Evidence of AI-Generated Sources](https://arxiv.org/abs/2605.23684), arXiv:2605.23684.
- Chu, Leng, Li, Shen, Shen and Zhang (2026), [GEO-Flag: Detecting and Measuring GEO-Optimized Web Content](https://arxiv.org/abs/2608.16824), arXiv:2608.16824.
- Luo and Chen (2026), [One Polluted Page Is Enough: Evaluating Web Content Pollution in LLM Recommenders](https://arxiv.org/abs/2606.13610), arXiv:2606.13610.
- Underneath (2026), [How many “best of” lists cited by AI rank their own brand first?](https://underneath.agency/research/self-promoting-best-lists-study)

---

This is the Markdown twin of https://underneath.agency/resources/are-ai-search-engines-citing-ai-content. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How fund managers get considered when investors ask AI"
description: "By being the fund family assistants find in trusted fund research, with fees and facts consistent everywhere. Five firms already hold 58% of US fund assets."
canonical: "https://underneath.agency/resources/asset-managers-fund-demand-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How does an asset manager get its funds considered when investors and advisers ask AI?

By making sure assistants find your fund family in the independent research they already trust, with fees, tickers and strategy facts that match everywhere. Investors and advisers now use AI to research funds, and the answers lean on third-party sources and familiar names. For a mid-size or newer fund sponsor, that turns AI visibility into a question of whether your strategies make the shortlist at all.

## The short version

1. The industry is concentrating: the [Investment Company Institute](https://icifactbook.org/pdf/2026-factbook-ch2.pdf) reports the five largest fund complexes held 58% of US mutual fund and ETF assets at the end of 2025, up from 35% in 2005, while firms ranked 11 to 25 fell from 21% to 14%.
2. New funds need to be found quickly: [Cerulli Associates](https://etfexpress.com/2026/06/16/new-etf-launches-far-exceed-closures-cerulli-associates/) counted 953 active ETF launches in 2025, and 92% of that year’s ETF closures were subscale products under $50 million.
3. Investors already research with AI: in an [Investing.com survey](https://hedgefundalpha.com/news/retail-investors-now-use-ai-to-inform-investment-decisions/) of 938 US investors, 54% had used chatbots such as ChatGPT for investing research.
4. Advisers, the gatekeepers for many funds, are adopting it too: 51% use generative AI, according to a [Broadridge and Financial Services Institute study](https://www.broadridge.com/jp/press-release/2026/advisors-signal-desire-for-more-technology-and-product-education).
5. Answers lean on independent sources: in a [study of banking questions](https://arxiv.org/abs/2509.08919), 64.6% of the sources AI engines cited were editorial or third-party sites, against 34.1% brand-owned.

A note before you read: this article is about how fund companies show up in AI answers. Nothing in it is investment, legal or compliance advice.

## Who chooses funds today, and what is a fund client worth?

Two groups choose: self-directed investors on platforms, and financial advisers who pick funds for client portfolios.

The market is large and still growing. The ICI’s [2026 Fact Book](https://icifactbook.org/pdf/2026-factbook-quick-facts-guide.pdf) puts US-registered fund assets at $45.1 trillion, with 76.0 million households, 56.4% of all US households, owning funds. Investors can choose from 16,829 investment companies, including 8,030 mutual funds and 4,813 ETFs, offered by 772 fund sponsors.

A fund client is worth recurring fees on the assets that stay. The revenue line is simple: assets multiplied by the expense ratio, every year. That ratio keeps shrinking. In 2025 the simple average expense ratio for index equity ETFs was 0.45%, but the asset-weighted average was 0.14%, because investors put most of their money into the cheapest, largest funds. When fees are thin, scale decides who earns a living, and scale comes from being considered.

The adviser channel adds a second buyer with different questions. Advisers want product education, not just product names. In the Broadridge and FSI survey of 428 advisers, 82% said better training and awareness of tools would help them grow, and 48% said they need to build their knowledge of crypto, a category many fund sponsors now package in ETFs.

## Where does AI already sit in fund research?

At the research stage: investors and advisers use AI to explore and compare before they buy, not usually to decide.

The evidence comes mostly from surveys, and the best available ones are run by companies with research products to sell, so treat them as directional:

- **US investors.** The Investing.com survey (March 2026, drawn from its own users) found 62% had used AI tools to help with investment decisions and 54% had used chatbots such as ChatGPT for investing research. 39% worried about incorrect or misleading recommendations.
- **Global retail investors.** [Finimize’s Modern Investor Pulse](https://finimize.com/business/press/modern-investor-pulse-2026-q3-40-per-cent-of-retail-investors-use-ai-for-research-weekly), a survey of 2,808 investors, found 40.9% use AI for research at least weekly. US investors were the most likely to never use it, at 24.4%.
- **New investors.** In the same survey, AI-generated research was the first catalyst for 4.5% of investors who started in the last two years, roughly double the 2.5% among longer-standing investors.
- **Advisers.** Broadridge and FSI found 67% of advisers under 45 use generative AI, against 43% of those 65 and over.

The pattern we infer from these numbers: AI is becoming the place where a fund category is first explained and a shortlist first forms. The purchase still runs through a platform, an adviser or a model portfolio, but the names on the shortlist were often set earlier.

## Which questions do investors and advisers ask about funds?

Category, cost, structure and fund-family questions, plus education questions from advisers. We drafted the sample prompts below to show the shape of these questions; none was logged from a real investor or adviser.

| Who asks | Illustrative prompt |
|---|---|
| Self-directed investor | “What’s the lowest-cost way to own the S&P 500 in an ETF?” |
| Self-directed investor | “Active bond ETF or bond mutual fund for a taxable account?” |
| Income investor | “Which dividend ETFs have the longest track records?” |
| Adviser | “Who issues buffer ETFs, and how do their caps and costs compare?” |
| Adviser | “Explain interval funds for private credit: liquidity, fees, risks” |
| Home-office research | “Fund families with the strongest active fixed income ETF lineups” |

These questions matter because the categories that are growing are the ones that need explaining. In the [MMI-Broadridge survey](https://s203.q4cdn.com/209299927/files/doc_news/AI-Product-Innovation-and-Next-Generation-Investors-Set-the-Course-for-the-Future-of-Asset-and-Wealth-Management-MMI-Broadridge-Surve-KR08K.pdf) of asset and wealth managers, 72% put active ETFs among their top three growth categories, and 78% of asset managers named them a key growth area. A new active ETF or interval fund is exactly what an adviser asks an assistant to explain.

## Why are the stakes higher for mid-size and newer fund sponsors?

Because the industry is concentrating, launches are at record levels, and AI answers tend to default to familiar names.

The ICI data shows the squeeze. The five largest fund complexes now hold 58% of mutual fund and ETF assets, and the ten largest hold 72%. From 2015 to 2025, 408 sponsors entered the US market and 515 left. Index funds, mostly run by the largest complexes, now hold 52% of long-term fund assets.

At the same time, the shelf is getting more crowded. [ETFGI](https://etfgi.com/node/28333) counted a record 2,759 new ETFs listed worldwide through November 2025, including 1,033 in the United States. Cerulli reports that 83% of ETF issuers intend to launch at least one active ETF in 2026. Most closures hit funds that never reached scale: Cerulli says they “did not attract adviser and end-investor interest.”

AI answers may add to the advantage of size. In a [controlled test](https://arxiv.org/abs/2606.17443) by Xi Chu and YuPeng Hou, assistants recommended a well-known brand 100% of the time when products were otherwise identical, though a rival’s small, visible quality edge broke that default. The test used consumer products, not funds. Still, a default toward the familiar name would fit what the ICI data already shows about fund assets: a few complexes dominate.

Our inference: for a sponsor outside the top tier, a fund that assistants cannot explain or place in its category is a fund that starts each conversation from zero.

## How does AI visibility turn into assets under management?

Through the shortlist: an answer names fund families or tickers, the investor or adviser checks the documents, and assets follow.

The path looks like this, and each step is our inference about how a documented buying process meets AI research:

1. **Category answer.** An investor asks how to get exposure to a theme, or an adviser asks how a structure works. The answer names a handful of fund families or tickers.
2. **Due diligence.** The investor opens fact sheets, ratings and the prospectus; the adviser checks the firm’s approved list and research notes.
3. **Purchase.** The trade happens on a platform or inside a model portfolio, often with no visible link to the AI answer. The platform itself may be picked through AI too, as our guide to [how investing apps get chosen](https://underneath.agency/resources/investment-platforms-customers-ai-search) explains.
4. **Recurring revenue.** The assets pay the expense ratio for as long as they stay, which is why one adviser adding a fund to a model can matter more than many single purchases.

This is why measuring AI’s effect through website visits undercounts it for fund companies. The adviser who first learned about a strategy from an assistant may never visit the sponsor’s site before placing it in a model. The same blind spot shows up well beyond funds, as our piece on [why analytics miss AI visibility](https://underneath.agency/resources/why-analytics-miss-ai-visibility) explains.

## What decides whether a fund family is named?

Mostly independent editorial coverage, plus the accuracy of the firm’s own data. Platforms document only part of this.

**Documented by platforms.** Google says its systems give even more weight to strong expertise and trust signals for “Your Money or Your Life” topics, which include financial stability, and that such content “must be highly accurate and consistent with established expert consensus.” Google also says [AI Overviews and AI Mode](https://developers.google.com/search/docs/appearance/ai-features) may use “query fan-out,” running several related searches across subtopics before answering.

**Observed in studies.**

- **Editorial sources lead in finance.** In the banking study by Mahe Chen and colleagues, editorial and third-party sites made up 64.6% of cited sources, brand-owned sites 34.1% and social sites 1.2%. Bankrate and NerdWallet led the domains. That study covered banks, not funds.
- **Third-party mentions travel with recommendations.** Across the brands in [our brand study](https://underneath.agency/research/brand-entity-ai-recommendations-study), every tenfold rise in the number of independent sites that named a brand went with 4.7 times the odds of being recommended. For a fund family, those sites are fund research databases, adviser trade press and financial media.
- **Stale pages travel.** In [our pricing study](https://underneath.agency/research/ai-pricing-accuracy-study), 61.9% of software prices quoted by assistants were fully faithful to the official page, and for 39 of 64 prices that differed, the same figure appeared on another page of the vendor’s own site.

**Our inference for fund companies.** The pricing study covered software, but the lesson carries: an old fact sheet with last year’s expense ratio, or a data vendor with an outdated strategy description, can become the version an assistant repeats. Fund facts live in many places (your site, prospectuses, data vendors, ETF databases, platform listings), and assistants can pick up any of them.

## What does GEO look like for a fund company inside the marketing rules?

For a fund sponsor, generative engine optimization (GEO) makes accurate, compliant facts and credible coverage easy for assistants to find.

1. **A clear entity.** Make the relationship between the parent firm, the fund family brand and each fund unambiguous, with consistent names and tickers across your site and data vendors.
2. **One version of the facts.** Expense ratios, inception dates, strategy descriptions and share classes should match across fact sheets, data vendors and platform listings. Retire stale PDFs rather than leaving them online.
3. **Category education.** Explain the structures you sell (active ETFs, buffer ETFs, interval funds) in plain language, including risks and who they do not suit. These are the pages advisers and assistants need. [How to fix wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) covers tracing errors back to their source.
4. **Independent coverage.** Portfolio manager commentary in trade and financial press, analyst coverage, and inclusion in research databases give assistants third-party evidence. Our guide on [how brands build authority for AI search](https://underneath.agency/resources/how-brands-build-authority-for-ai-search) covers the method.
5. **Compliance from the start.** The SEC’s [investment adviser marketing rule](https://www.law.cornell.edu/cfr/text/17/275.206%284%29-1) prohibits advertisements that include “any untrue statement of a material fact,” and fund advertising has its own SEC and FINRA rules. GEO content goes through the same review as any other marketing.
6. **Measure across engines.** Assistants differ in what they cite, so check several of them; see [why tracking ChatGPT alone is not enough](https://underneath.agency/resources/is-tracking-chatgpt-enough).

No one can promise that an assistant will name a particular fund. GEO improves the evidence assistants find; whether a strategy suits an investor stays a decision for the investor and their adviser.

## What can’t the research tell fund marketers yet?

Whether AI visibility moves fund flows, and how assistants treat individual funds as opposed to firms.

- **No fund-level study.** We found no independent study of how assistants choose among specific funds or ETFs. The closest evidence covers banks, consumer products and firm-level brand visibility.
- **Vendor data.** The investor surveys come from companies selling research tools, and their methods are their own.
- **No public link to flows.** No study we found connects a fund family’s AI visibility with net new assets.

## Where should a fund company start?

Start by asking assistants the category and structure questions your strategies answer, and record which firms and sources appear.

That first check shows whether assistants place your funds in the right categories, whether your fees and facts are quoted correctly, and which research sites shape the answers. It also shows where a competitor with a weaker product is named simply because it is better documented.

If your growth plan depends on new strategies reaching scale, or on getting onto advisers’ shortlists, [ask us for an AI visibility review of your fund lineup](https://underneath.agency/contact). We will test the investor and adviser questions behind your key strategies, trace the sources behind each answer, and plan the data fixes, coverage and education content, reviewed with your compliance team, that give your funds a fair chance of being considered. See our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) for how we keep fund facts consistent across data vendors and build adviser education content alongside your compliance review.

## Frequently asked questions

### Do AI assistants recommend specific funds?

They often name fund families or tickers when asked category questions, but no independent study yet shows how they choose among specific funds. Treat every answer as a starting point.

### Why would an assistant quote the wrong expense ratio?

Usually because an old page says so. In our pricing study, most differing software prices also appeared elsewhere on the vendor’s own site, which points to stale pages rather than invention.

### Is GEO allowed under fund marketing rules?

GEO content is marketing content. It goes through the same compliance review as any fact sheet or article, under the advertising rules that already apply to your firm and your funds.

### Can a small fund sponsor compete with the largest complexes in AI answers?

It can be considered, but it starts behind. In a controlled test, assistants defaulted to the known brand when products looked identical, and a small, visible advantage changed that. Clear, documented differences matter most.

## Sources

- Investment Company Institute (2026), [2026 Investment Company Fact Book, Chapter 2](https://icifactbook.org/pdf/2026-factbook-ch2.pdf)
- Investment Company Institute (2026), [2026 Investment Company Fact Book: Quick Facts Guide](https://icifactbook.org/pdf/2026-factbook-quick-facts-guide.pdf)
- ETF Express (2026-06-16), [New ETF launches far exceed closures: Cerulli Associates](https://etfexpress.com/2026/06/16/new-etf-launches-far-exceed-closures-cerulli-associates/)
- ETFGI (2025-12-31), [ETFGI reports the global ETFs industry had a record 2759 new products listed at the end of November 2025](https://etfgi.com/node/28333)
- Investing.com, via Hedge Fund Alpha (2026-04-09), [Nearly Two-Thirds Of Retail Investors Now Use AI To Inform Investment Decisions](https://hedgefundalpha.com/news/retail-investors-now-use-ai-to-inform-investment-decisions/)
- Finimize (2026-07-07), [40% of retail investors use AI for research weekly](https://finimize.com/business/press/modern-investor-pulse-2026-q3-40-per-cent-of-retail-investors-use-ai-for-research-weekly)
- Broadridge and Financial Services Institute (2026-02-02), [Advisors Signal Desire for More Technology and Product Education](https://www.broadridge.com/jp/press-release/2026/advisors-signal-desire-for-more-technology-and-product-education)
- Money Management Institute and Broadridge (2025-11-21), [AI, Product Innovation, and Next-Generation Investors Set the Course for the Future of Asset and Wealth Management](https://s203.q4cdn.com/209299927/files/doc_news/AI-Product-Innovation-and-Next-Generation-Investors-Set-the-Course-for-the-Future-of-Asset-and-Wealth-Management-MMI-Broadridge-Surve-KR08K.pdf)
- Chen, Wang, Chen and Koudas (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919)
- Xi Chu and YuPeng Hou (2026), [Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems](https://arxiv.org/abs/2606.17443)
- Google Search Central (2025), [Creating helpful, reliable, people-first content](https://developers.google.com/search/docs/fundamentals/creating-helpful-content)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- Legal Information Institute, Cornell Law School, [17 CFR 275.206(4)-1, Investment adviser marketing](https://www.law.cornell.edu/cfr/text/17/275.206%284%29-1)
- Underneath (2026), [Do Wikipedia and schema make AI assistants recommend a brand?](https://underneath.agency/research/brand-entity-ai-recommendations-study)
- Underneath (2026), [How faithfully do AI assistants quote software prices?](https://underneath.agency/research/ai-pricing-accuracy-study)

---

This is the Markdown twin of https://underneath.agency/resources/asset-managers-fund-demand-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How automation integrators win industrial buyers via AI search"
description: "It can, if assistants find proof of your platforms, industries, region and results, so plant leaders put you on the integrator longlist early."
canonical: "https://underneath.agency/resources/automation-integrators-industrial-buyers-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Will AI search point plant leaders to our firm when they plan an automation project?

It can, when an assistant can match your firm to the project’s platform, industry, region and goal, and confirm it in sources other than your own site. Manufacturers are short of people and under pressure to automate, and many of their engineers now start research with AI before they call an integrator.

This guide is for leaders of automation system integrators and controls solution providers: firms that design and commission PLC and SCADA systems, machine vision, motion and line automation, and that migrate aging control systems. Robot makers seeking manufacturer shortlists are covered in a separate article in this series; here robots are one technology among many that an integrator installs.

## The short version

1. Demand is spreading to new sectors: the [J.P. Morgan and Control System Integrators Association (CSIA) survey](https://www.inddist.com/technology-software/article/22975346/integrators-remain-key-resource-as-manufacturers-navigate-increasingly-complex-tech-landscape) of February 2026 named data centers, defense, semiconductors, power and energy, and life sciences as the sectors with the strongest momentum.
2. Labor is the business case: the [Manufacturing Institute and Deloitte](https://themanufacturinginstitute.org/manufacturers-need-as-many-as-3-8-million-new-employees-by-2033/) estimate US manufacturing may need 3.8 million workers from 2024 to 2033, and 1.9 million of those jobs could go unfilled.
3. Automation is a top spending priority: in [Deloitte’s survey of 600 manufacturing executives](https://www.deloitte.com/us/en/insights/industry/manufacturing/2025-smart-manufacturing-survey.html), 46% ranked process automation and 37% physical automation as a first or second investment priority, and respondents reported 10% to 20% gains in production output from smart manufacturing.
4. Most manufacturers are past pilots: in [Rockwell Automation’s 2026 study](https://www.rockwellautomation.com/en-ca/company/news/press-releases/90-of-Manufacturers-Say-Digital-Transformation-Is-Now-Essential-According-to-New-Global-Study.html) of 1,560 respondents, 59% were actively using smart manufacturing technology and only 18% remained in pilot mode.
5. Vision projects are growing: North American machine vision transactions rose 8.8% in the second quarter of 2026, [according to A3](https://www.automate.org/vault/a3-vision-and-imaging-statistics/a3-vision-and-imaging-2q-2026-abridged-report).

## Who hires automation integrators, and what is a client worth?

Operations and engineering leaders at manufacturers, and a good client brings repeat projects, upgrades and service work.

Deloitte found that smart manufacturing initiatives are owned by operations leaders, such as the chief operating officer, in 51% of companies, and by technology owners, such as the chief technology officer, in 38%. In practice a plant or controls engineer scopes the project, an operations leader owns the business case, and procurement runs the bid. [Forrester](https://www.forrester.com/press-newsroom/forrester-2026-the-state-of-business-buying/) reports that a typical business purchase now involves 13 internal stakeholders and nine external influencers. For an integrator, those outside influencers often include the automation vendor’s sales team, [a consulting engineering firm](https://underneath.agency/resources/engineering-firms-project-inquiries-ai-search) or a peer at another plant.

The reason clients call is usually people. In the National Association of Manufacturers’ outlook survey for the first quarter of 2024, cited in the Manufacturing Institute and Deloitte study, 65% of respondents named attracting and retaining talent as their primary business challenge. AMT, the Association For Manufacturing Technology, [links rising automation demand](https://www.automation.com/article/583.4-million-usd-new-machinery-orders-us-economic-strengths) to “nearly half a million current manufacturing job openings.”

There is no public figure for an average integration project. The CSIA article cites a MarketsandMarkets projection that the US integrator market will exceed $9.2 billion by 2029. Our inference about client value: the first project, often a migration of obsolete controls or a single new line, is the entry point. The value comes from the next lines, the plant standards, the service agreement and the sister plants that follow when the first project goes well.

## Where does AI search enter an automation project?

At the start, when an engineer frames the problem and builds a longlist of possible integrators.

Controls and plant engineers are a large part of the audience in the 2026 [State of Marketing to Engineers](https://advertising.globalspec.com/state-of-marketing-to-engineers-research-report) survey by TREW Marketing and GlobalSpec, with Elektor. Its figures show where they look before they shortlist anyone:

| Where engineers research a purchase (TREW and GlobalSpec, 2026) | Engineers |
|---|---|
| Use generative AI during the purchasing process | 69% |
| Routinely use generative AI platforms for purchase research (up 8 points) | 21% |
| Routinely research in online technical publications | 76% |
| Routinely research on supplier and vendor websites | 74% |
| Routinely use industry directory websites | 31% |
| Find an AI search summary usually enough, among those who notice one | 6% |

Forrester describes generative AI searches as the starting point for business buyers, followed by validation from their internal and external networks. That matches how integrators are usually chosen: an engineer gathers names from research, the automation vendor’s partner list and colleagues, then checks each one. Robot makers that sell through integrators face the same longlist; see [how robot makers reach manufacturer shortlists](https://underneath.agency/resources/robotics-companies-customers-ai-search).

Note one difference between the integrator and the client. Rockwell’s study found that 34% of operations are already AI-augmented. That is AI inside the plant, which integrators install. This article is about AI in the buying process, which decides who gets asked to install it. Plant software vendors face the same research stage, covered in [how MES and IIoT platforms get found](https://underneath.agency/resources/industrial-software-plants-ai-search).

## What do plant leaders ask when they look for an integrator?

Project questions that combine a platform, a process, an industry, a place and a business goal.

These examples are ours, written to show the pattern. They are not taken from buyer logs.

| Need | Illustrative question |
|---|---|
| Migration | “Integrators that migrate legacy PLC-5 controls to a current platform in Ohio, with minimal downtime” |
| Vision | “Machine vision inspection integrator for label and lot code checks in a pharmaceutical plant” |
| Credentials | “CSIA certified integrators near Houston for water treatment SCADA” |
| ROI | “What payback should I expect from automating a case packing line?” |
| New sector | “Controls integrators with data center electrical and cooling experience” |
| Risk | “How do I choose an integrator for an operational technology cybersecurity upgrade?” |

The ROI question is where many projects begin, before the buyer has decided to hire anyone. A firm that answers it clearly, with ranges, assumptions and the cases where automation does not pay back, gives an assistant something specific to cite.

## How does an AI answer turn into an integration contract?

Through a longlist, a site visit, a paid study or proposal, a project, then service and the next line.

1. **Longlist.** An engineer asks an assistant, a vendor’s partner directory or a colleague who can do the work. Firms named in AI answers, and in the pages those answers cite, join the list.
2. **Check.** The engineer reads case studies, platform credentials, certifications, industries served and locations.
3. **Scope.** A site visit and often a paid feasibility study or front-end engineering package. This is the first revenue.
4. **Proposal and project.** A fixed-price or time-and-materials proposal, an ROI case for the operations leader, then design, factory acceptance testing, installation and commissioning.
5. **Service and expansion.** Support agreements, training, upgrades and the next line or plant.

An assistant’s answer can shape the longlist and the first check, but nothing after the site visit. References, engineering depth, price and the vendor relationship decide the rest. Deloitte found that 65% of respondents ranked operational risk as a first or second concern about smart manufacturing projects, which is why proof of delivery matters more than claims.

## What makes an assistant name one integrator over another?

Independent proof it can find for the exact project; the platforms document how they search, not how they rank firms.

Documented by the platforms: [OpenAI says](https://help.openai.com/en/articles/9237897-chatgpt-search) ChatGPT search rewrites a question into one or more targeted queries sent to search providers, and that sites must allow OAI-SearchBot to be eligible. [Google says](https://blog.google/products/search/ai-mode-search/) AI Mode runs multiple related searches through a “query fan-out” technique, and its [guidance for site owners](https://developers.google.com/search/docs/appearance/ai-features) says there are “no additional requirements to appear in AI Overviews or AI Mode” beyond normal search best practices. None of them say how they choose a firm.

Observed in our studies:

- **Google rankings are not the whole story.** In [our study of AI citations and Google rankings](https://underneath.agency/research/ai-citations-google-rankings-study), only 8.3% of the pages ChatGPT cited ranked in Google’s top 10 for the question, and 68.5% were not in Google’s top 100 for the question or for any of ChatGPT’s own searches. Our article on [whether SEO still matters for AI search](https://underneath.agency/resources/is-seo-still-important-for-ai-search) explains what that means.
- **Independent mentions go with recommendations.** In [our brand entity study](https://underneath.agency/research/brand-entity-ai-recommendations-study), the odds of being recommended rose 4.7 times with each tenfold increase in independent sites naming a brand in the cited pages. For an integrator, those sites are trade journals, vendor case studies and association pages.
- **Assistants look for reviews.** In [our hidden searches study](https://underneath.agency/research/ai-hidden-searches-study), ChatGPT looked for reviews in 46.2% of answers. Integrators rarely collect public reviews, so case studies, references and trade coverage carry more of that weight.

Our inference for integrators: the trust signals are the ones a plant engineer already checks. Platforms you are certified on, industries and processes you know, regulated work you have done, the region you serve, independent certifications such as CSIA’s audited program, and documented results. The CSIA describes its certification as validating an integrator’s business and management practices “through an independent audit.” Vendor partner status and certifications should be stated exactly, because buyers check them with the vendor.

## What does an integrator lose when it is left out of AI answers?

Early project conversations, especially in new sectors where buyers have no integrator relationship yet.

We found no measurement of integration work lost to AI absence. The reasoning below is labeled:

- **New sectors mean new buyers.** The CSIA survey points to data centers, semiconductors and life sciences, and A3 data show robot orders from automakers fell 25% in the first half of 2026 while other industries grew. We infer that many of these buyers have no integrator on file and will research before they call. Warehouse projects show the same long research phase; see [how warehouse automation vendors make shortlists](https://underneath.agency/resources/warehouse-automation-enterprise-leads-ai-search).
- **The ROI conversation happens without you.** If an assistant answers the payback question with someone else’s case study, that firm frames the project.
- **Wrong facts filter you out.** An outdated platform list, an old office or a lapsed certification in an answer can remove you from a longlist. Our guide to [fixing wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) covers how to trace the source.
- **Talent and capacity are scarce on both sides.** Integrators face the same retirement gap as their clients, the CSIA says. Our inference: a firm that attracts better-matched inquiries spends its limited engineering hours on work it is most likely to win.

## How does GEO work for an automation integrator?

Generative engine optimization (GEO) helps AI assistants find, describe and verify your firm for the projects you want.

For integrators, the work usually covers:

1. **Capability pages by project type.** Control system migration, SCADA, machine vision, motion, packaging lines, batch control, OT cybersecurity: what you do, on which platforms, in which industries, with typical scope and timelines.
2. **Credentials stated exactly.** Vendor partner levels, CSIA certification, safety and quality certifications, licenses where required, matched to the issuer’s own listing.
3. **Case studies with numbers.** Before-and-after throughput, downtime, scrap or labor hours, published with client permission. These are what an assistant can quote.
4. **ROI explainers.** Honest pages on payback ranges, assumptions and when automation is the wrong answer.
5. **Independent coverage.** Bylined articles in trade publications, conference talks, association awards and vendor case studies naming your firm. Engineers rank online technical publications as their top research source.
6. **Consistent directory profiles.** Vendor partner directories, CSIA’s directory, LinkedIn and Google Business Profile with the same name, offices, platforms and industries. Machine shops follow the same rule, keeping one identity across Thomasnet, LinkedIn and Google Business Profile, as [CNC shops competing for RFQs](https://underneath.agency/resources/cnc-machine-shops-rfqs-ai-search) shows.
7. **Measurement.** Ask a fixed set of project, platform, industry and region questions across ChatGPT, Gemini, Perplexity, Copilot and Google’s AI features, repeatedly, and log who is named and what is cited. Our guide to [designing AI visibility tracking](https://underneath.agency/resources/how-to-design-ai-visibility-tracking) explains how to set this up, then compare with project inquiries.

No integrator can be promised a mention. The aim is to make your firm the easiest to verify when a project is a genuine fit.

## What can’t the data tell an integrator yet?

How many integration projects start with an AI answer, and how assistants rank integrators specifically.

- **No attribution data.** The surveys measure how engineers research. None links AI answers to integration contracts.
- **Interested sources.** Rockwell sells automation, CSIA represents integrators, A3 represents automation suppliers, TREW and GlobalSpec sell marketing to technical firms, and market-size projections come from commercial research firms.
- **Our studies were broad.** They tested buyer questions across several industries, not integrator searches, so carrying their results over to controls work is our own inference.
- **Selection data is private.** We found no public study, current and saveable, of how manufacturers choose between integrators.

## Where should an automation integrator start?

Pick the projects you most want to win and ask assistants who can deliver them in your region.

Write ten questions the way a plant engineer would, naming the platform, process, industry and location, plus two or three ROI questions. Run them across the main assistants and record whether your firm appears, what it is credited with, which case studies and directories are cited, and which firms are named instead. That shows which proof is missing and where it should live.

If your next few years depend on winning more paid feasibility studies rather than chasing every bid, [we can review how assistants present your integration firm](https://underneath.agency/contact). The review covers the migration, vision and line automation questions plant leaders ask, the integrators and partner directories that answer them today, and the case studies and credentials that would put your firm in those answers. How we then build capability pages, exact credential listings and trade coverage for an integrator is outlined on our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page.

## Frequently asked questions

### Do manufacturers really use ChatGPT to find integrators?

Many engineers use generative AI somewhere in purchasing: 69% in the 2026 TREW and GlobalSpec survey. We found no data on integrator selection specifically, and referrals and vendor partner lists still matter.

### Does being on a vendor’s partner list help in AI answers?

We have no direct evidence. Partner directories are outside sources that confirm your platforms and region, so keep them complete and consistent with your site.

### Should we publish ROI figures from client projects?

Only with permission and with the assumptions stated. Specific, sourced results are more useful to buyers, and to assistants, than general claims.

### Is CSIA certification worth stating prominently?

If you hold it, yes. It is an independently audited credential a buyer can check, which is the kind of fact assistants and engineers can verify.

## Sources

- Industrial Distribution (2026), [Integrators Remain Key Resource as Manufacturers Navigate Increasingly Complex Tech Landscape](https://www.inddist.com/technology-software/article/22975346/integrators-remain-key-resource-as-manufacturers-navigate-increasingly-complex-tech-landscape)
- The Manufacturing Institute and Deloitte (2024-04), [Manufacturers Need as Many as 3.8 Million New Employees by 2033](https://themanufacturinginstitute.org/manufacturers-need-as-many-as-3-8-million-new-employees-by-2033/)
- Deloitte (2025), [2025 Smart Manufacturing and Operations Survey](https://www.deloitte.com/us/en/insights/industry/manufacturing/2025-smart-manufacturing-survey.html)
- Rockwell Automation (2026-05-19), [90% of Manufacturers Say Digital Transformation Is Now Essential, According to New Global Study](https://www.rockwellautomation.com/en-ca/company/news/press-releases/90-of-Manufacturers-Say-Digital-Transformation-Is-Now-Essential-According-to-New-Global-Study.html)
- Association for Advancing Automation (2026), [A3 Vision & Imaging 2Q 2026 Abridged Report](https://www.automate.org/vault/a3-vision-and-imaging-statistics/a3-vision-and-imaging-2q-2026-abridged-report)
- Association for Advancing Automation (2026-08), [Robot Orders Increase in Q2 as Automation Demand Broadens Across Industries](https://www.automate.org/robotics/news/robot-orders-increase-in-q2-as-automation-demand-broadens-across-industries)
- AMT via Automation.com (2026-07-13), [$583.4 Million in New Machinery Orders Highlight US Economic Strengths](https://www.automation.com/article/583.4-million-usd-new-machinery-orders-us-economic-strengths)
- Forrester (2026-01-21), [Forrester’s 2026 Buyer Insights: GenAI Is Upending B2B Buying](https://www.forrester.com/press-newsroom/forrester-2026-the-state-of-business-buying/)
- TREW Marketing and GlobalSpec (2026), [State of Marketing to Engineers research report](https://advertising.globalspec.com/state-of-marketing-to-engineers-research-report)
- OpenAI Help Center (2026), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Google (2025-03-05), [Expanding AI Overviews and introducing AI Mode](https://blog.google/products/search/ai-mode-search/)
- Google Search Central (2026), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude cite pages that rank?](https://underneath.agency/research/ai-citations-google-rankings-study)
- Underneath (2026), [Do Wikipedia and schema make AI assistants recommend a brand?](https://underneath.agency/research/brand-entity-ai-recommendations-study)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)

---

This is the Markdown twin of https://underneath.agency/resources/automation-integrators-industrial-buyers-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How can B2B SaaS companies generate revenue from AI search?"
description: "By getting onto the shortlists AI assistants build for category, alternatives and comparison questions, then turning that intent into trials and renewals."
canonical: "https://underneath.agency/resources/b2b-saas-revenue-from-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How can B2B SaaS companies generate revenue from AI search?

By getting named on the shortlists AI assistants now build for software buyers, then converting that ready-made intent into trials, demos and contracts that renew. Buyer surveys say the shortlist is increasingly formed inside ChatGPT, Gemini, Perplexity and Google’s AI features, and a handful of SaaS companies report that AI referrals already bring in a tenth of their signups. What no one has yet published is a controlled link from AI visibility to annual recurring revenue (ARR), so treat the channel as real, fast-growing and still unproven at the contract level.

## The short version

1. A won SaaS customer tends to stay: private B2B SaaS companies in SaaS Capital’s 2025 survey reported median gross revenue retention of 95% on contracts above $250,000 a year and about 91% on smaller ones.
2. Existing customers drive growth too: in Benchmarkit’s data, 40% of total new ARR came from customers a company already had, so a buyer first won through an AI answer can keep adding revenue.
3. Shortlists are tiny: 2.6 products on average in [TrustRadius’s](https://solutions.trustradius.com/vendor-blog/bridging-the-trust-gap-b2b-tech-buying-in-the-age-of-ai) 2025 survey, and 70% of buyers purchased the product they had in mind when they drew it up.
4. Company-reported examples show the channel converts: at [Ahrefs](https://ahrefs.com/blog/ai-search-traffic-conversions-ahrefs/), 0.5% of visits from AI search produced 12.1% of signups; ChatGPT referred about 10% of new [Vercel](https://vercel.com/blog/how-were-adapting-seo-for-llms-and-ai-search) signups; at Webflow, ChatGPT traffic converted at 24%.
5. The economics favor winning early: private SaaS companies spent a median of $2.00 in sales and marketing for each $1 of new ARR from new customers in 2024, per [Benchmarkit’s data](https://thesaasbarometer.substack.com/p/what-are-the-latest-saas-metrics), and that revenue then recurs for years.

## Who buys B2B software today, and what is a new customer worth?

A committee buys it, mostly on its own research, and each new customer is a stream of renewals.

[Forrester’s 2026 buying study](https://www.forrester.com/press-newsroom/forrester-2026-the-state-of-business-buying/) counts 13 internal stakeholders and nine external influencers in a typical business purchase, with procurement among the decision-makers in 53% of buying cycles. Software committees are smaller at the core: [G2’s 2025 report](https://company.g2.com/news/buyer-behavior-in-2025) found groups of 3 to 4 members rising, with IT part of nearly 50% of purchase decisions. Those buyers want to work alone for as long as they can. A [Gartner survey of 632 B2B buyers](https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-sales-survey-finds-61-percent-of-b2b-buyers-prefer-a-rep-free-buying-experience) found 61% prefer a rep-free buying experience overall.

The cycle is shorter than many sales teams assume. Technology buyers in the TrustRadius survey reported an average deal cycle of 3.8 months, rising to 6.6 months for enterprise deals. [6sense’s 2025 study](https://customerthink.com/research-round-up-6sense-study-provides-critical-insights-on-b2b-buyer-behavior/), which covers all kinds of B2B purchases, put the average at 10.1 months globally.

What makes SaaS different from most industries is what happens after the signature. [SaaS Capital’s 2025 survey](https://www.saas-capital.com/wp-content/uploads/2025/09/RB32WS1-2025-B2B-SaaS-Retention-Benchmarks.pdf) of private B2B SaaS companies found median gross revenue retention of 95% for companies with contracts above $250,000 a year and about 91% below that. Benchmarkit found that 40% of total new ARR now comes from existing customers. So a customer who first met you in an AI answer is not one sale. It is an annual contract, likely renewals and, often, expansion. Acquiring that customer is expensive: the $2.00 of sales and marketing spend per $1 of new-customer ARR above is up from $1.76 the year before.

## At what point do SaaS buyers bring AI assistants into a purchase?

At the start and at the shortlist, alongside Google rather than instead of it.

G2’s 2026 survey found 71% of buyers rely on AI chatbots at some point in their research and 61% use them in tandem with Google. Some 41% use “Deep Research” tools for structured software evaluations, which return long reports ending in a recommended set of vendors. G2 also reports that AI chatbots are now the top source influencing buyer shortlists, ahead of review sites, analyst firms and vendor websites. Marketing leaders are one group where this shows; see [how marketing software wins AI-first buyers](https://underneath.agency/resources/marketing-software-ai-search-growth).

Google’s own surface is just as present. In [our study of 1,248 US searches](https://underneath.agency/research/ai-overviews-frequency-study), B2B software and technology keywords showed an AI Overview on 96.0% of searches, the highest of eight industries; after adjusting for the kind of searches each industry has, software stayed highest at 83.0%. TrustRadius found 72% of technology buyers had encountered AI Overviews during research, and 90% clicked through to the cited sources.

Not every survey agrees on the pace. The same TrustRadius report, fielded in early 2025, found only 7% of buyers using tools like ChatGPT as part of their buying process. Forrester describes generative AI search as the starting point for B2B buyers, but notes that buyers then validate with colleagues and outside influencers because the answers are often incomplete.

## What kinds of questions do SaaS evaluators put to AI assistants?

Mostly shortlist questions: best tool for a situation, alternatives to an incumbent, head-to-head comparisons, price and fit.

G2 describes buyers “running head-to-head vendor comparisons” rather than basic education prompts. Tally, a form builder, noticed the same thing: its most likely explanation is that default web browsing made ChatGPT “more likely to recommend alternatives based on user intent,” citing searches like “simple or clean form builders.”

We wrote the example prompts below to show how SaaS evaluation questions tend to be phrased; none come from real buyer logs:

- Category with constraints: “Best help desk software for a 30-person B2B support team on Salesforce.”
- Alternatives: “Alternatives to Jira for a small engineering team that hates complexity.”
- Comparison: “HubSpot vs Pipedrive for an outbound sales team: which is easier to adopt?”
- Price and packaging: “How much does Intercom cost per seat with the AI agent included?”
- Fit and risk: “Which data warehouse tools have SOC 2 Type II and EU data residency?”

Each type maps to a different stage. Category and “alternatives to” questions settle which SaaS products make the first list. Comparison, price and security questions settle which of them survive to a trial or demo. Prices are where answers slip most visibly: in [our pricing study](https://underneath.agency/research/ai-pricing-accuracy-study), only 61.9% of the plan prices four assistants quoted for 45 software products were fully faithful to the vendor’s pricing page. Whether vendor-written comparisons earn citations is covered in [our article on comparison pages](https://underneath.agency/resources/do-comparison-pages-help-b2b-ai-citations).

## How does a mention in an AI answer become ARR?

Through the shortlist: the assistant names you, the buyer checks you, then trials or demos you, then renews.

**Shortlist.** Buyers decide early and rarely change their minds. 6sense found buying groups filled four of the five spots on their vendor shortlist on Day 1, and buyers bought from their top-ranked vendor 77% of the time. TrustRadius found 82% of buyers had a top product in mind when they made the shortlist. Yet the assistant can change that plan: in [G2’s March 2026 survey of 1,076 software buyers](https://company.g2.com/news/g2-research-the-answer-economy), 69% said they chose a different vendor than planned because a chatbot recommended it, and 85% of buyers think more highly of a vendor when AI includes it in an answer. These are buyer-reported figures from a review platform with a stake in the topic.

**Visit.** Many buyers never click; those who do arrive ready. Ahrefs found its AI search visitors converted at a 23x higher rate than traditional organic search visitors, while estimating that people click links 75% less in AI assistants than in classic search. [Webflow](https://kylepoyar.substack.com/p/traffic-is-no-longer-reliable) told Growth Unhinged that 10% of its signups came from AI discovery, growing 4x year on year, with ChatGPT traffic converting at 6x the rate of Google. G2’s 2025 report put the gap at 40%: AI search-driven leads converted that much better than traditional search leads.

**Trial or demo.** Here the path splits by sales motion. Product-led companies see it in signups: [Tally](https://blog.tally.so/from-2-to-3m-arr-how-we-bootstrapped-tally-with-a-tiny-team/) says ChatGPT became its number one referral source, with over 2,000 tracked new users a week signing up via AI tools, and Vercel’s write-up credits AI search with helping Tally grow from $2M to $3M ARR in four months. Sales-led companies see it in demo requests and trials: Forrester found more than 60% of business buyers now use a trial, rising to 78% for purchases of $10 million or more, and TrustRadius found product demos the resource buyers consult most, at 55%.

**Revenue.** A signup is not a customer. Tally notes that about 2% of its free users eventually upgrade to its paid plan. Turning free signups into seat growth is the focus of [our collaboration software guide](https://underneath.agency/resources/collaboration-software-ai-search). For sales-led SaaS, the AI answer rarely shows up as a referral at all; we infer it appears as a branded search, a direct visit or a buyer who already knows your name on the first call. That is why the full value sits in contract value and retention, not in click counts. Sales tool vendors are one example, since a shortlist built in an AI answer reaches them as demo requests, as our guide to [turning AI answers into sales pipeline](https://underneath.agency/resources/sales-software-pipeline-from-ai-search) explains.

## Why does an assistant put one SaaS product on its shortlist and not another?

The platforms say little; studies point to review sites, independent coverage, user communities and plainly stated product facts.

**Documented by the platform.** According to Google, AI Mode relies on a [“query fan-out” technique](https://blog.google/products/search/ai-mode-search/): it runs multiple related searches across subtopics, then combines what comes back. A buyer’s single question can therefore pull in pages about pricing, integrations and reviews at once.

**Observed in studies.** When [Chen and colleagues](https://arxiv.org/abs/2509.08919) asked US software ranking questions, AI search took 72.7% of its sources from independent “earned” sites, against 45.4% for Google. In our AI Overview research on B2B software searches, Google cited Reddit in [35.4% of answers](https://underneath.agency/research/ai-reddit-citations-study), with communities such as r/projectmanagers and r/crmsoftware among the most cited, and cited a YouTube video on [91.0% of searches](https://underneath.agency/research/ai-overview-youtube-videos-study). Only 23.9% of the URLs those answers cited were [page-one results](https://underneath.agency/research/ai-overview-citations-study). Being known is not enough: in a test of 112 Product Hunt startups, [Sharma](https://arxiv.org/abs/2601.00912) found ChatGPT recognized 99.4% when asked by name but surfaced only 3.32% in discovery-style questions.

There is a stability advantage too. Of eight industries in [our four-assistant study](https://underneath.agency/research/ai-assistants-brand-agreement-study), B2B software showed the most agreement between assistants (an overlap score of 0.543 on a 0-to-1 scale), and [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study) found its recommended brands the most stable across repeated runs (0.708). We infer that once a software brand is established in a category’s answers, its position is more durable than in local or retail categories.

**Trust factors specific to SaaS.** G2 reports that buyers rank citations from software review sites as their top signal of confidence in a chatbot’s answer, and that they distrust answers that differ across assistants. G2’s 2025 report adds that about 8 in 10 buyers face stricter IT security, legal and compliance reviews for AI software. A reasonable expectation is that public security documentation, integration lists and pricing pages carry weight at the comparison stage.

## What does a SaaS company lose when the AI shortlist skips it?

Usually the deal itself, because buyers rarely add vendors after the shortlist forms.

TrustRadius found only 14% of buyers had a shortlist of four or more products, and 79% of buyers knew the product they purchased before they started researching. If the assistant builds a three-vendor list without you, there is little room to recover later in the cycle. G2’s finding that one in three buyers purchased from a vendor they had never heard of cuts both ways: challengers can win deals they could not have reached before, and category leaders can lose deals they never saw. AI video is a category where that cuts hardest, because buyers try several tools and switch cheaply, as our guide to [how AI video tools win customers](https://underneath.agency/resources/ai-video-generator-customers-from-ai-search) shows.

This loss rarely shows up in a SaaS dashboard. Ahrefs’ AI visitors were only 0.5% of its traffic, so a team watching visits would have missed that they brought in 12.1% of signups. [Our article on lost clicks and pipeline](https://underneath.agency/resources/ai-answers-pipeline-revenue) looks at what fewer search clicks do to the top of a SaaS funnel.

## What does GEO involve for a B2B SaaS vendor?

It improves what assistants can find and trust about your product; nobody can promise a shortlist slot.

For a SaaS company, generative engine optimization (GEO) usually spans six areas of work:

1. **Entity clarity.** Describe the product, category, ideal customer and integrations the same way on your site, review profiles, documentation, marketplace listings and LinkedIn, so every assistant has one consistent story.
2. **Review platforms.** Keep a steady flow of recent, specific reviews on the platforms your buyers use, because both buyers and answers lean on them at the decision stage.
3. **Independent coverage.** Pursue the independent “best of” lists, newsletters, podcasts and practitioner communities your category’s buyers read. Our article on [which pages to target](https://underneath.agency/resources/best-of-lists-ai-recommendations) covers this in detail.
4. **Decision-stage content, ungated.** Publish honest comparison and alternatives pages, integration guides, security and compliance pages and implementation timelines. TrustRadius notes that AI models can only learn from publicly available content, so gated material does not reach them.
5. **Accurate pricing.** Keep one current pricing page and retire old figures, since assistants can quote an older or adjacent figure from elsewhere on your site.
6. **Measurement tied to revenue.** Track a fixed set of buyer prompts by stage across ChatGPT, Gemini, Perplexity, Copilot and Google, then connect it to self-reported attribution. Tally’s onboarding survey showed far more AI-sourced users than referral data did.

## What can’t the evidence yet tell a SaaS company about AI-driven revenue?

Nobody has yet shown that AI visibility causes new ARR rather than simply arriving alongside it.

Most of the buyer data comes from surveys by companies that sell to software vendors, including review platforms whose own business benefits from the conclusion. Surveys also disagree: 51% said they start research with an AI chatbot more often than with Google in G2’s 2026 data, up from 29% a year earlier, against 7% using such tools in TrustRadius’s early 2025 data. The company examples come from developer and marketing tools whose buyers adopted AI early, and they are self-reported. No published study yet follows AI answers through to closed contracts and renewals. The field test we know of, covered in [our article on business results](https://underneath.agency/resources/does-ai-visibility-drive-business-results), found that unoptimized pages grew too, which is a warning against crediting every rise to your own work.

## How should a SaaS company check whether AI shortlists feed its trials and demos?

Map which assistants name you for the questions that come before your trials and demo requests.

Run your category, alternatives, comparison and pricing questions through the main assistants, then line the results up against trial, demo and pipeline data, so each gap is ranked by contract value and renewal potential rather than by mentions alone. To build that view together, [ask us for an audit of your SaaS shortlists](https://underneath.agency/contact). Our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page explains how the program runs after that audit, from review platforms and ungated decision-stage pages to measurement tied to trials and pipeline.

## Frequently asked questions

### Do AI assistants use G2 and other review sites when recommending software?

G2 reports that buyers see citations from software review sites as their strongest reason to trust a chatbot’s answer. Studies of software questions also find that AI search leans on independent sites more than Google does. How much any single review platform weighs in a given assistant is not documented.

### Does AI search traffic convert better for SaaS companies?

In the cases published so far, yes. Ahrefs, Webflow and G2 all report higher conversion from AI-referred visitors than from traditional search. These are company-reported figures, and part of the gap may come from fewer, more decided visitors rather than from AI itself.

### Should SaaS companies publish “alternatives to” and comparison pages?

Yes, as long as they treat competitors fairly. Buyers ask comparison and alternatives questions, and assistants look for pages that answer them. Pages that only praise you, or rank you first on your own site, are a weak bet; see [our comparison-page research summary](https://underneath.agency/resources/do-comparison-pages-help-b2b-ai-citations).

### How do you measure revenue from AI search in a SaaS business?

Combine three signals: how often assistants name you for your buyers’ questions, AI referral signups and demo requests, and what new customers say when asked how they found you. Then trace those accounts through to first-year contract value and each renewal.

## Sources

- G2 (2026), [In the Answer Economy, Don’t Win the Click — Win the Answer](https://company.g2.com/news/g2-research-the-answer-economy)
- G2 (2025), [New G2 Research: How AI is Redefining the Buyer Journey in 2025 and How Sellers Can Adapt](https://company.g2.com/news/buyer-behavior-in-2025)
- Forrester (2026), [Forrester’s 2026 Buyer Insights: GenAI Is Upending B2B Buying](https://www.forrester.com/press-newsroom/forrester-2026-the-state-of-business-buying/)
- Gartner (2025), [Gartner Sales Survey Finds 61% of B2B Buyers Prefer a Rep-Free Buying Experience](https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-sales-survey-finds-61-percent-of-b2b-buyers-prefer-a-rep-free-buying-experience)
- TrustRadius (2025), [Bridging the Trust Gap: B2B Tech Buying in the Age of AI](https://solutions.trustradius.com/vendor-blog/bridging-the-trust-gap-b2b-tech-buying-in-the-age-of-ai)
- CustomerThink (2025), [Research Round-Up: 6sense Study Provides Critical Insights on B2B Buyer Behavior](https://customerthink.com/research-round-up-6sense-study-provides-critical-insights-on-b2b-buyer-behavior/)
- SaaS Capital (2025), [2025 B2B SaaS Retention Benchmarks](https://www.saas-capital.com/wp-content/uploads/2025/09/RB32WS1-2025-B2B-SaaS-Retention-Benchmarks.pdf)
- The SaaS Barometer (2025), [What are the Latest SaaS Metrics Benchmarks Telling Us?](https://thesaasbarometer.substack.com/p/what-are-the-latest-saas-metrics)
- Ahrefs (2025), [Does AI Search Traffic Convert Better Than Traditional Search? For Ahrefs, Yes](https://ahrefs.com/blog/ai-search-traffic-conversions-ahrefs/)
- Vercel (2025), [How we’re adapting SEO for LLMs and AI search](https://vercel.com/blog/how-were-adapting-seo-for-llms-and-ai-search)
- Growth Unhinged (2025), [Traffic is no longer a reliable growth metric](https://kylepoyar.substack.com/p/traffic-is-no-longer-reliable)
- Tally (2025), [From 2 to $3M ARR: How We Bootstrapped Tally With a Tiny Team](https://blog.tally.so/from-2-to-3m-arr-how-we-bootstrapped-tally-with-a-tiny-team/)
- Google (2025), [Expanding AI Overviews and introducing AI Mode](https://blog.google/products/search/ai-mode-search/)
- Chen and colleagues (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919)
- Sharma (2025), [The Discovery Gap: How Product Hunt Startups Vanish in LLM Organic Discovery Queries](https://arxiv.org/abs/2601.00912)
- Underneath (2026), [When does Google show an AI Overview?](https://underneath.agency/research/ai-overviews-frequency-study), [AI Overview citations](https://underneath.agency/research/ai-overview-citations-study), [Reddit citations](https://underneath.agency/research/ai-reddit-citations-study), [YouTube videos](https://underneath.agency/research/ai-overview-youtube-videos-study), [software pricing accuracy](https://underneath.agency/research/ai-pricing-accuracy-study), [four-assistant agreement](https://underneath.agency/research/ai-assistants-brand-agreement-study) and [recommendation consistency](https://underneath.agency/research/ai-recommendation-consistency-study)

---

This is the Markdown twin of https://underneath.agency/resources/b2b-saas-revenue-from-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How do baby product brands get recommended when parents ask AI?"
description: "By giving AI assistants accurate, safety-complete product facts and independent proof that parents trust, so the brand is named when registries are built."
canonical: "https://underneath.agency/resources/baby-product-brands-ai-recommendations"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How do baby product brands get recommended when parents ask AI?

By being the brand whose facts, safety details and independent reviews are easiest for an assistant to find and trust. Expecting parents now use AI to sort through an overwhelming number of strollers, bassinets and pumps, and the answers draw on product data, review sites and parent communities. A baby brand wins those moments by being accurate and well-evidenced everywhere an assistant looks, not by louder claims.

## The short version

1. Parents are drowning in information: in a 2026 [NielsenIQ](https://nielseniq.com/global/en/insights/report/2026/2026-u-s-mother-baby-market-trends-understanding-moms-rethinking-consumption/) study with Momcozy, mothers consulted 3.3 sources per baby purchase, 58% said information overload creates pressure, and 39% named not knowing what to trust as a purchase barrier.
2. The registry is the sales engine: [Babylist](https://www.prnewswire.com/news-releases/babylist-tops-750-million-in-revenue-in-2025-increasing-45-as-it-expands-ecosystem-for-growing-families-302726386.html) says 22% of US families having a baby use its registry, more than 10 million people bought through it in 2025, and its AI registry assistant now [searches the web](https://help.babylist.com/hc/en-us/articles/52717506395931-What-s-new-with-Addie) for answers.
3. Safety is the trust test: the [Consumer Product Safety Commission](https://www.cpsc.gov/s3fs-public/Nursery-Products-Annual-Report-2025.pdf) estimates 70,000 emergency-room injuries among children under five in 2024 involved nursery products in use, and staff have reports of 537 deaths from 2020 through 2022.
4. Google already answers most baby questions with AI: one study cited in recent research found [AI Overviews](https://arxiv.org/abs/2608.04831) on 84% of baby care and pregnancy searches.
5. AI shoppers are valuable when they arrive: [Adobe’s data](https://techcrunch.com/2026/04/16/ai-traffic-to-us-retailers-rose-393-in-q1-and-its-boosting-their-revenue-too/) shows AI traffic to US retail sites rose 393% in early 2026, and AI visitors converted 42% better than other visitors in March.

## Who buys baby gear, and how does one family turn into revenue?

Mostly expecting mothers building a registry, with gift givers paying for much of the first wave.

The buyer is usually one parent doing a great deal of research in a short window. In NielsenIQ’s report, 91% of respondents identified themselves as the primary caregiver, and 58% of all mothers work full time. Nearly all, 98%, said their core caregiving tasks centered on feeding and sleep, which is why breast pumps (72%), bottle sterilizers (65%) and bassinets or cribs (61%) top their product lists.

The purchase is often split. The parent researches and adds products to a registry; friends and grandparents buy from it. In Babylist data from 2019, users spent about 40 hours building a registry, and only 11% felt they had researched products enough to feel prepared, as reported by [Tinybeans](https://tinybeans.com/babylist-report-baby-registry-trends/). The numbers are old, but the pattern is the same one Babylist still builds its business on. Gift buying for older children raises similar questions, covered in [how toy brands get found for gifts](https://underneath.agency/resources/toy-brands-product-discovery-ai-search).

The market is large and durable. Babylist puts the kids and baby category at $235 billion, with roughly 80% of spending considered essential, based on its own analysis of third-party reports. It reported more than $750 million in 2025 revenue, up 45%, and more than 100 million gifts given through its registry to date.

For a brand, the revenue path is therefore not a single click. It runs from being named during research, to being added to a registry, to being bought by someone else, and then to the products a family needs as the baby grows. NielsenIQ found the top purchase factor was long-term value, with 36% wanting a product to fit the baby’s needs across stages.

## Where do AI assistants already sit in a parent’s research?

At the start of the search, inside registry tools, and on Google results pages parents already use.

Three kinds of evidence point the same way:

- **Parents use AI daily.** In a July 2025 [BSM Media](https://www.webull.com/news/13243978330498048) survey of nearly 500 mothers, 85% said they use AI in daily life, while 73% named fear of AI misuse as their biggest hesitation. BSM Media helps brands market to mothers, and the sample is small, so read it as direction, not a precise rate.
- **Assistants are built for this exact task.** When [OpenAI launched shopping research](https://openai.com/index/chatgpt-shopping-research/) in ChatGPT in November 2025, one of its own example requests was a lightweight, compact stroller for a toddler that could handle city sidewalks. Babylist’s assistant, Addie, now sees what is on a parent’s registry, recommends items to fill the gaps, and can add products found on other sites. If it cannot find a source it trusts, Babylist says, it tells the parent instead of guessing.
- **Google answers baby questions with AI.** Research by Hu and colleagues, cited in a 2026 study of AI Overview clicks, found AI Overviews on 84% of baby care and pregnancy searches. A separate paper notes the same team found AI Overviews and Featured Snippets [disagree on 33%](https://arxiv.org/abs/2603.16138) of those queries.

Shopping-wide data shows how fast this has grown. [Adobe](https://news.adobe.com/news/2026/01/adobe-holiday-shopping-season) measured a 693.4% jump in AI-referred traffic to US retail sites over the 2025 holiday season. AI visitors in March 2026 also spent 48% longer on site, and their revenue per visit was 37% higher than other traffic.

## What do parents ask AI before they buy?

Questions about needs, comparisons, safety, fit with their home, and what a gift giver should buy.

These prompts are illustrative, written by us to show the stages. They are not captured from real parents, and we do not claim what any assistant answers to them today.

| Stage | Illustrative prompt |
|---|---|
| Need | “What do I actually need for a newborn in a one-bedroom apartment?” |
| Shortlist | “Best lightweight stroller that works with an infant car seat” |
| Comparison | “Compare two bassinets for a parent recovering from a C-section” |
| Safety | “Has this bassinet ever been recalled, and does it meet the federal sleep standard?” |
| Cost | “Which breast pumps does insurance usually cover?” |
| Gift | “Good registry gift under $50 for a first baby” |

Two features stand out for baby brands. First, the questions carry constraints (space, budget, recovery, travel), so an assistant needs specific facts to match a product to a family. Second, safety questions are common and high-stakes. A wrong answer about a recall or an age limit is not just lost revenue; it erodes the trust the whole category runs on.

## How does an AI answer become a sale for a baby brand?

Through a named shortlist, a registry add, a purchase by the parent or a gift giver, then repeat stage purchases.

1. **Named.** An assistant lists the brand for a need the parent described, such as a compact travel system.
2. **Checked.** Parents verify. NielsenIQ found mothers draw on social media (40%), ecommerce sites (37%), parenting forums (36%), healthcare professionals (32%) and baby registry platforms (31%).
3. **Added.** The product goes on a registry, where a gift giver may buy it weeks later. That delay is one reason we infer AI influence is easy to undercount in a brand’s own analytics.
4. **Bought and kept.** NielsenIQ’s model suggests that pairing real-parent word of mouth with endorsement from medical professionals can reach a 2.5 times multiplier on conversion for core baby durables. That is a modeled estimate from a research firm working with a brand partner, not an observed experiment.

Retail data supports the value of step one. Adobe found AI-referred visitors converted 42% better than other visitors in March 2026, a reversal from a year earlier. Salesforce reported that over the 2025 holidays, shoppers referred from AI search [converted nine times more often](https://www.salesforce.com/news/stories/2025-holiday-shopping-data/) than shoppers from social media. Neither figure is specific to baby products.

## What decides which baby brand an assistant names?

Product facts, independent reviews and safety information; only the platforms know the exact rules.

**Documented by the platform.** OpenAI says that when choosing products, ChatGPT [considers](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search) structured information such as price and product description from first-party and third-party providers, other third-party content, and OpenAI safety standards and product policies. It says its review summaries come from public websites and are not verified by OpenAI. OpenAI also says shopping research reads product pages directly and avoids low-quality or spammy sites, but can still make mistakes about price and availability.

**Observed in studies.**

- In a skincare test of three AI systems, product facts such as rating, price and reviews explained 82.4% of how products were ranked ([Chu and Hou](https://arxiv.org/abs/2606.17443)). We explain that study in [what actually drives AI product recommendations](https://underneath.agency/resources/what-drives-ai-product-recommendations).
- In shopping answers, Google’s AI Overviews drew 25.4% of their sources from editorial and review sites, 22.2% from community and social sites, 17.8% from retailers and marketplaces, and 16.1% from manufacturers and brands ([Uberti-Bona Marin and colleagues](https://arxiv.org/abs/2609.18729)).
- When ChatGPT researched buyer questions in [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study), it went looking for reviews or ratings on 46.2% of them. And when shoppers asked “is this brand legit?”, 88.0% of the answers in [our reputation study](https://underneath.agency/research/is-it-legit-ai-reputation-study) pointed to a review or complaint platform.
- In [our brand study](https://underneath.agency/research/brand-entity-ai-recommendations-study), each tenfold increase in independent sites naming a brand went with 4.7 times the odds of being recommended.

**Our inference for baby products.** The trust factors that matter most are likely the ones parents already use: independent testing and reviews from parenting publications, registry platform listings and reviews, real-parent discussion in forums, clear compliance with federal standards, and endorsements from clinicians where they are genuine. A brand’s own claims count for less than the evidence others publish about it.

## What does it cost a baby brand to be missing or misdescribed?

A missed registry slot, and in the worst case an assistant repeating outdated safety information about your product.

The commercial loss is easy to state but has not been measured: if a stroller is not on the shortlist when a parent builds a registry, it is unlikely to be the gift someone buys. We found no public data on how many registry adds start with an AI answer.

The safety risk is better documented, which is why accuracy matters more here than in most categories. The CPSC report says high chairs, cribs and mattresses, infant carriers and strollers were linked to 62% of nursery product injuries in 2024. The rules change often: the agency issued a final rule for nursing pillows, effective April 23, 2025, and banned inclined sleepers, after which deaths involving those products fell from 27 in 2019 to 4 in 2022. An assistant that leans on an old review or a stale product listing can describe a model, a version or a standard that no longer applies.

Answers also shift. Ask ChatGPT the same question five times and the list moves: in [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), just 25.2% of the brands it named came back in every repeat. And polluted web pages can mislead: in tests of 12 AI models, planted review pages got a fake brand recommended in up to 73.8% of products ([Luo and Chen](https://arxiv.org/abs/2606.13610)); we cover that in [can fake reviews make AI recommend a fake brand](https://underneath.agency/resources/fake-reviews-fake-brands-ai-recommendations).

## How does GEO work for a baby product brand?

Generative engine optimization (GEO) makes your products easy for assistants to identify, describe accurately and back with independent evidence.

For a baby brand, the work usually covers:

1. **Complete, consistent product facts.** Model names, age and weight limits, dimensions, folded size, compatible car seats, materials, price and availability, kept identical across your site, retailer listings, registry platforms and product feeds. Our guide to [what product content AI shopping assistants prefer](https://underneath.agency/resources/product-content-ai-shopping-assistants-prefer) lists the facts worth publishing.
2. **Safety information stated plainly.** Which federal standard a product meets, the current model year, and recall status with a link to the official notice. This helps parents and gives assistants something accurate to repeat.
3. **Independent evidence.** Earned reviews and testing from parenting publications, “best of” lists, and real parent discussion, which assistants draw on heavily. We explain why in [which pages to target for AI recommendations](https://underneath.agency/resources/best-of-lists-ai-recommendations).
4. **Registry and retailer presence.** Accurate listings and genuine reviews on the registries and stores parents use, since assistants such as Addie sit inside them.
5. **Answer-ready pages.** Plain guides that answer the questions parents ask, such as which stroller fits a small car trunk, written without overstated safety or health claims.
6. **Correction and monitoring.** Ask a fixed set of parent questions across ChatGPT, Gemini, Perplexity and Google’s AI features, repeat them over time, and fix wrong facts at the source. The steps for a stroller or bassinet listed with the wrong limits are in [how to fix wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers).

None of this guarantees a recommendation. It raises the odds that when an assistant looks, it finds evidence that is accurate, current and trusted.

## What does the evidence not tell baby brands yet?

How many registry adds or gift purchases start with an AI answer, for any brand.

- **Surveys are small or sponsored.** BSM Media’s survey covered about 500 mothers and comes from a firm that helps brands market to mothers; NielsenIQ’s report was produced with Momcozy, a baby brand.
- **Retail data is cross-category.** Adobe’s and Salesforce’s AI traffic figures cover all retail, not baby products.
- **Platform rules are unpublished.** OpenAI names the inputs it considers but not how it weighs them; no platform says how it handles recalled baby products in answers.
- **Studies are adjacent.** The product-ranking and fake-review tests used other categories, and the AI Overview findings concern baby care information, not product recommendations.

Treat any precise estimate of revenue from AI answers in this category as a guess until brands publish their own data.

## Where should a baby product brand start?

Start by asking the questions parents ask, and checking whether your products appear and are described correctly.

Run your core registry questions, comparisons and safety questions in the assistants parents use, several times each. Note which brands are named, which sources are cited, and whether your model names, limits and safety details are right. That shows where the gaps are: missing facts, thin independent coverage, or outdated information.

If registry adds and sales are what you need to grow, [talk to us about a review of your products in AI answers](https://underneath.agency/contact). We will show where your products appear when parents ask AI for help, why other brands are named instead, and which fixes are most likely to put your products on more registries. The [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page sets out how we keep limits, safety details and listings consistent across registries and retailers, and track parent questions over time.

## Frequently asked questions

### Do parents really use AI to choose baby products?

Many do. In BSM Media’s 2025 survey of nearly 500 mothers, 85% used AI in daily life, and Babylist built an AI registry assistant that recommends products and searches the web for answers.

### Can an AI assistant recommend a recalled baby product?

It is possible if it relies on outdated pages; no platform publishes how it screens recalls. OpenAI says its shopping answers can make mistakes, so brands should keep recall status and model details clear and current everywhere.

### Do baby brands need to be on registry platforms to show up in AI answers?

We infer it helps. Babylist’s assistant works inside its registry, and NielsenIQ found 31% of mothers use registry platforms as an information source before buying.

### Will health claims help a baby product get recommended?

No evidence says so, and overstated claims carry risk. In product-ranking tests, manipulative copy was often flagged and demoted, as we explain in [can sellers game AI shopping rankings](https://underneath.agency/resources/can-you-game-ai-shopping-rankings); accurate facts and independent reviews are a safer basis.

## Sources

- NielsenIQ (2026), [2026 U.S. Mother & Baby Market Trends: Understanding Moms, Rethinking Consumption](https://nielseniq.com/global/en/insights/report/2026/2026-u-s-mother-baby-market-trends-understanding-moms-rethinking-consumption/)
- Babylist via PR Newswire (2026-03-26), [Babylist Tops $750 Million in Revenue in 2025, Increasing 45% as it Expands Ecosystem for Growing Families](https://www.prnewswire.com/news-releases/babylist-tops-750-million-in-revenue-in-2025-increasing-45-as-it-expands-ecosystem-for-growing-families-302726386.html)
- Babylist (2026), [What’s new with Addie](https://help.babylist.com/hc/en-us/articles/52717506395931-What-s-new-with-Addie)
- Tinybeans (2019-05-29), [New Survey on Baby Registries Reveals How Much People Actually Spend](https://tinybeans.com/babylist-report-baby-registry-trends/)
- U.S. Consumer Product Safety Commission (2026-03), [Injuries and Deaths Associated with Nursery Products Among Children Younger than Age Five, 2025 Report](https://www.cpsc.gov/s3fs-public/Nursery-Products-Annual-Report-2025.pdf)
- BSM Media via PR Newswire (2025-07-29), [Moms and Artificial Intelligence: New Survey from BSM Media Reveals Shift in Comfort, Concerns, and Consumer Expectations](https://www.webull.com/news/13243978330498048)
- OpenAI (2025-11-24), [Introducing shopping research in ChatGPT](https://openai.com/index/chatgpt-shopping-research/)
- OpenAI (2026), [Shopping with ChatGPT Search](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search)
- TechCrunch (2026-04-16), [AI traffic to US retailers rose 393% in Q1, and it’s boosting their revenue too](https://techcrunch.com/2026/04/16/ai-traffic-to-us-retailers-rose-393-in-q1-and-its-boosting-their-revenue-too/)
- Adobe (2026-01-07), [Holiday Shopping Season Drove a Record $257.8 Billion Online with Consumers Embracing Generative AI Tools](https://news.adobe.com/news/2026/01/adobe-holiday-shopping-season)
- Salesforce (2026-01), [AI and agents account for $262 billion of 2025 holiday spend](https://www.salesforce.com/news/stories/2025-holiday-shopping-data/)
- Investigating Click Behaviors On Google Search Result Pages That Produce an AI Overview (2026), [Investigating Click Behaviors On Google Search Result Pages That Produce an AI Overview](https://arxiv.org/abs/2608.04831)
- Answer Bubbles: Information Exposure in AI-Mediated Search (2026), [Answer Bubbles: Information Exposure in AI-Mediated Search](https://arxiv.org/abs/2603.16138)
- Chu and Hou (2026), [Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems](https://arxiv.org/abs/2606.17443)
- Uberti-Bona Marin and colleagues (2026), ["If I Had to Buy Just ONE: Galaxy S26 Ultra": Auditing AI-Generated Product Recommendations](https://arxiv.org/abs/2609.18729)
- Luo and Chen (2026), [One Polluted Page Is Enough: Evaluating Web Content Pollution in LLM Recommenders](https://arxiv.org/abs/2606.13610)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [“Is this brand legit?” How AI assistants build a reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study)
- Underneath (2026), [Do Wikipedia and schema make AI assistants recommend a brand?](https://underneath.agency/research/brand-entity-ai-recommendations-study)
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)

---

This is the Markdown twin of https://underneath.agency/resources/baby-product-brands-ai-recommendations. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How banking technology vendors get on bank shortlists via AI"
description: "By being on the first list a bank builds before its RFP. Core, payments and fraud deals are rare and long, so early AI answers carry unusual weight."
canonical: "https://underneath.agency/resources/banking-technology-enterprise-deals-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How does a banking technology vendor get onto a bank’s shortlist when the first research happens in AI?

By being named, accurately, when a bank or credit union maps the market before it writes a request for proposal. Core banking, payments and fraud platforms are bought rarely, through long evaluations that regulators expect banks to document, so the few early answers a buyer reads can shape a decision that lasts a decade. For a banking technology vendor, AI visibility is built on checkable proof: named institutions, conversion records, resilience facts and independent coverage.

## The short version

1. The buyer pool is small and shrinking: the [FDIC](https://www.fdic.gov/quarterly-banking-profile/quarterly-banking-profile-second-quarter-2026.pdf) counted 4,238 insured institutions in mid-2026, down 41 in one quarter, and the [NCUA](https://ncua.gov/files/publications/analysis/quarterly-data-summary-2026-Q2.pdf) 4,214 federally insured credit unions.
2. Core decisions are rare: an Independent Community Bankers of America executive estimates just 2% to 3% of institutions convert to another core each year, and only 19% of institutions in an [American Bankers Association survey](https://www.pcbb.com/bid/2025-04-23-core-providers-beefed-up-tech-but-satisfaction-can-improve) said they were likely to switch at their next renewal.
3. Each win is large: [Jack Henry](https://pulse2.com/jack-henry-faster-payments-revenue-jumps-nearly-50-as-digital-transaction-revenue-rises-11-6/) called its 58 competitive core wins in fiscal 2026 a record, and 14 of them were institutions with more than $1 billion in assets.
4. Selection starts years ahead: choosing and converting a core can take one to three years, and [CUInsight](https://www.cuinsight.com/the-core-conversion-crunch-why-financial-institutions-must-act-now-to-secure-their-future/) reports some vendors are already booking conversion slots into 2029.
5. Budgets are rising: in [Bank Director’s 2025 Technology Survey](https://www.nasdaq.com/press-release/bank-directors-2025-technology-survey-banks-grapple-data-ai-maturity-2025-09-16), 71% of banks increased technology budgets, with a median increase of 10%.

This guide is about how banking technology vendors are found and described by AI assistants. It is not legal or regulatory advice; claims about compliance, resilience or certifications should go through your own legal and risk review.

## Who buys banking technology, and what is one customer worth?

A small executive group at a bank or credit union buys it; one core customer means years of revenue.

The buyer is an institution, not a person. In Bank Director’s survey of 141 directors and executives at US banks under $100 billion in assets, 54% said a management-level team or steering committee gives final approval for technology investments, 48% said a C-level executive other than the chief information or technology officer holds that authority, and just over a quarter said the board is directly involved in major technology decisions. A typical evaluation team, we infer, includes the CEO or COO, the technology lead, the chief risk officer and, for fraud and payments, the operations and fraud heads.

The market is concentrated in a few thousand institutions, and mergers keep shrinking it. In the second quarter of 2026, four banks opened while 36 institutions merged with other banks, according to the FDIC. The credit union count fell from 4,370 a year earlier. Every merger removes a potential customer and often forces a technology choice for the combined institution.

What a win is worth shows in the vendors’ own results. Jack Henry, which reports about 7,400 bank and credit union clients, recorded fiscal 2026 revenue of $2.54 billion. Its processing revenue rose to about $1.10 billion, and its faster-payments revenue grew 49.5% as institutions adopted real-time payment services. The economics, we infer, follow the same pattern across core, payments and fraud: the contract is won once and then earns for years through processing, transaction and add-on revenue.

That is a different sale from the finance-team software covered in [our guide to fintech software buyers](https://underneath.agency/resources/fintech-software-customers-from-ai-search), where controllers can start a trial in a week. Here, the trial is a conversion.

## How do banks choose a core, payments or fraud vendor today?

Slowly, through peers, consultants and formal proposals, often starting two or three years before a contract ends.

The trigger is usually the renewal date. The American Bankers Association’s 2024 Core Platforms Survey, with nearly 800 institutions responding, found satisfaction with core providers falls the longer a contract runs, reaching only 41% within two years of renewal. Aside from cost, the main reasons institutions gave for considering a change were disappointment with customer service (42%) and the relationship with their core provider (38%).

From there, the process is long and staged:

1. **Early market scan.** CUInsight advises institutions to “survey the market and identify high-level vendor fit before committing to a full RFP process.” It recommends starting evaluations “at least three years before contract expiration” because popular vendors run out of conversion capacity.
2. **Peers and advisers.** A Cornerstone Advisors partner, quoted by PCBB, recommends asking colleagues at other institutions how they feel about their core provider, “and not just those who offered testimonials.” [Cornerstone Advisors](https://www.crnrstone.com/bold-solutions/transformation/core-banking-systems) itself says it has run more than 700 core evaluations.
3. **Request for proposal and negotiation.** Formal proposals, demonstrations, reference calls and contract terms follow.
4. **Due diligence.** US regulators expect banks to manage the risks of every third-party relationship. In 2024 the Federal Reserve, FDIC and OCC issued a [guide for community banks](https://www.federalreserve.gov/newsevents/pressreleases/bcreg20240503b.htm) that walks through “each stage of the third-party relationship,” building on the agencies’ June 2023 guidance.
5. **Conversion.** Industry advisers quoted by PCBB say it can take one to three years to select a new core and convert, so institutions should start 24 to 30 months before the contract expires.

Fraud and payments decisions move faster than core decisions but follow the same logic. In [Alloy’s 2026 fraud report](https://www.alloy.com/blog/2026-state-of-fraud-report-key-takeaways), a survey of banks, credit unions and fintechs by a fraud-prevention vendor, 67% of respondents saw an increase in fraud attempts over the past year, and 71% said fraud most commonly occurred in online or mobile banking channels. Instant payments add urgency: the [FedNow Service](https://www.frbservices.org/resources/financial-services/fednow/volume-value-stats) settled 4,997,811 payments in the second quarter of 2026, with volume up 83.2% on the previous quarter.

## Where do AI assistants enter that buying process?

Mostly at the early market scan, the step before any vendor knows a bank is looking.

No public survey measures how bank technology buyers use AI assistants, so the evidence comes from wider business buying:

- In a buyer survey by [Responsive](https://www.responsive.io/news/buyer-intelligence-2025), a company that sells proposal software, 48% of US buyers said they use generative AI for vendor discovery. In [Digital Commerce 360’s coverage](https://www.digitalcommerce360.com/2026/01/05/ai-reshapes-b2b-buying-rfps/) of the same research, 90% of buyers said they research before first contact, and the weight placed on industry expertise was “especially strong in technology and financial services.”
- The [TrustRadius 2026 report](https://www.marketscale.com/industries/business-services/94-of-b2b-buyers-fact-check-ai-research-outputs-and-vendors-are-underestimating-how-far-trust-has-fallen), as summarized by MarketScale, found 63% of technology buyers used AI tools during their purchase journey, and the average shortlist held just 2.7 products.
- Banks are building their own AI habits. In Bank Director’s survey, 66% of banks had drafted an acceptable use policy for AI and 62% were experimenting with it in limited use cases.

Put together, a reasonable expectation is this: when a COO is asked to “see who else is out there” two years before renewal, an AI assistant is now one of the first places that question goes, alongside peer calls and a consultant. And with short shortlists, being absent from that first map is costly. Many professional buyers start with a favorite: 61% of Responsive’s respondents said they begin with a preferred vendor in mind, though 45% said they are open to switching.

## Which questions do bank buyers ask an assistant about technology vendors?

Questions about fit by institution size, conversion risk, integration, fraud coverage and regulatory readiness. The examples below are illustrative, written by us; they are not observed prompts.

| Stage | Illustrative prompt |
|---|---|
| Market scan | “Which core banking providers serve credit unions between $1 billion and $3 billion in assets?” |
| Comparison | “How do cloud-native cores compare with traditional providers for a community bank?” |
| Conversion risk | “How long does a core conversion take, and what usually goes wrong?” |
| Payments | “Which providers help a community bank receive instant payments through FedNow?” |
| Fraud | “Fraud platforms that cover check fraud and real-time payment scams for regional banks” |
| Due diligence | “What should we ask a core vendor about resilience and financial stability before signing?” |
| Alternatives | “Alternatives to our current core if we only want a sidecar for digital accounts” |

Two things stand out. First, many of these questions are about risk, not features, because a failed conversion is a board-level event; one industry association executive calls a core change “like open-heart surgery.” Second, several questions are really about capacity and timing, which few vendors state publicly.

## How does an AI answer turn into a signed bank contract?

By putting the vendor on the list that becomes the RFP invitation list. The rest is the vendor’s own sales work.

The path, as we see it:

1. **Named in the market scan.** An assistant, a peer or a consultant names three or four vendors.
2. **Checked.** Buyers verify what they read. In the TrustRadius report, 94% of buyers who used AI said they fact-check its outputs at least some of the time. For a bank, that means vendor pages, trade press, peer calls and, increasingly, a second AI question.
3. **Invited to the RFP.** In Responsive’s survey, the proposal response was the most important factor shaping the decision, cited by 81% of buyers, and industry expertise carried significant weight for 52%.
4. **Due diligence and contract.** Third-party risk review, negotiation and a multi-year agreement.
5. **Expansion.** Once the core is live, the vendor sells payments, card and digital services on top. Jack Henry’s results show the effect: full-year processing revenue rose 8.2%, faster than total revenue, which grew 7.1%.

The AI answer sits at step one, which is why its value is easy to miss. A vendor sees an RFP arrive and credits the relationship; the bank’s team may have first met the vendor’s name in an assistant two years earlier. We suggest asking prospects, at the first meeting, where they first heard of you, and logging which RFP invitations came without a prior relationship. More on that measurement gap is in [why analytics miss AI visibility](https://underneath.agency/resources/why-analytics-miss-ai-visibility).

## What decides whether a banking technology vendor is named?

Clear facts and independent evidence; platforms document how they search, not how they choose vendors.

**Documented by the platforms.** Google says AI Overviews and AI Mode [may use “query fan-out”](https://developers.google.com/search/docs/appearance/ai-features), issuing multiple related searches before answering. OpenAI says [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search) “typically rewrites your query into one or more targeted queries” that it sends to search providers. Neither publishes how it ranks banking technology vendors.

**Observed in our studies.**

- **Business software questions almost always get an AI answer.** In [our frequency study](https://underneath.agency/research/ai-overviews-frequency-study), 96.0% of US B2B software and technology keywords triggered an AI Overview.
- **They draw less on page-one results than finance does.** In [our citation study](https://underneath.agency/research/ai-overview-citations-study), 23.9% of B2B software citations were page-one Google results, so pages beyond the top ten, such as trade coverage and documentation, get cited too.
- **Named sources tend to be cited.** In [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study), when ChatGPT’s own search named a source, the answer cited it 44.0% of the time, against 8.1% when it did not.
- **Self-promotion is visible.** In [our study of “best” lists](https://underneath.agency/research/self-promoting-best-lists-study) cited by AI, 24.2% of numbered lists with an identifiable publisher ranked their own publisher first.

**Our inference, specific to banking technology.** The facts a bank needs most are the ones vendors often keep behind a sales call: which institutions of what size run the platform, how long conversions take and how many are booked, which payment rails and fraud types are covered, which independent audits exist, and how the company is funded. An assistant can only repeat what it can find and check. Trade association surveys, banking press, consultants’ public reports and customer announcements are the kind of independent sources we expect an assistant to rely on, and the kind a risk-minded buyer trusts.

## What does GEO look like for a vendor selling to banks and credit unions?

Generative engine optimization (GEO) makes your fit, track record and risk facts easy for assistants to find and repeat.

For a banking technology vendor, the work usually covers:

1. **A clear entity.** One plain description of what the platform is (full core, sidecar core, payments hub, fraud platform), which institutions it serves by type and asset size, and where it runs.
2. **Named proof.** Public customer announcements with permission, conversion case studies with dates, and the institutions’ own statements, so “who uses it” has a checkable answer.
3. **Due-diligence facts in public.** Resilience, security audits, data handling and ownership, written for the third-party risk questions banks must ask, reviewed by your legal team.
4. **Independent coverage.** Banking trade press, association research, conference sessions and consultant briefings, which an assistant can cite and a buyer can verify.
5. **Fair comparison pages.** Honest pages on how you differ from incumbents and from cloud-native challengers; [our comparison-page guide](https://underneath.agency/resources/do-comparison-pages-help-b2b-ai-citations) covers what the evidence says.
6. **Monitoring.** Ask assistants the questions your buyers ask each quarter and correct wrong facts at their source, following [how to fix wrong information about your brand in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers).

None of this guarantees a mention; no one can promise how an assistant will answer. GEO makes sure that when a bank’s team asks, the evidence about you is complete, current and accurate. How enterprise shortlists form more broadly is covered in [are AI assistants shaping which enterprise software gets shortlisted?](https://underneath.agency/resources/enterprise-software-shortlists-ai-search).

## What is still unknown about AI and bank technology buying?

Plenty: there is no public measure of how often bank buyers use AI, or of RFP invitations it produces.

- **No bank-specific data.** The AI-use figures above come from wider business buyer surveys, two of them run by vendors (Responsive sells proposal software, Alloy sells fraud prevention). We found no survey of bank technology buyers’ AI use.
- **Private assistants are invisible.** Many banks use assistants inside their own controlled tools. What those tools search, and which sources they can reach, is not public.
- **Our studies are snapshots.** They cover US searches collected in September 2026; answers change over time and between assistants.
- **The link to revenue is unmeasured.** With conversions this rare, no vendor has published how many RFP invitations trace back to AI answers. Our review of [whether AI visibility drives business results](https://underneath.agency/resources/does-ai-visibility-drive-business-results) sums up what other industries have measured so far.

## Where should a banking technology vendor start?

Start with the questions banks ask two to three years before renewal, and check what assistants say about you.

A useful first review covers the early market-scan questions in your category, the risk and conversion questions buyers ask about you by name, and the comparisons with incumbents. It shows whether you are named, which competitors and sources appear instead, and whether your customer base, conversion record and due-diligence facts are described correctly.

If your growth depends on being invited into core, payments or fraud evaluations before the RFP is written, [talk to us about a review of your standing in AI answers](https://underneath.agency/contact). We will compare how assistants describe you and your competitors for bank and credit union buyers, list what they miss or get wrong, and plan the content, coverage and proof that can earn you a place on the next shortlist. Our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page details how we put named proof, conversion records and due-diligence facts where bank and credit union buyers, and their assistants, can check them.

## Frequently asked questions

### Do bank executives really use ChatGPT to research core vendors?

No public survey measures it for banks. Wider surveys suggest many do: 48% of US business buyers in Responsive’s research use generative AI for vendor discovery, and most banks are now writing AI use policies.

### Why does AI visibility matter if banks buy through consultants and RFPs?

Because the shortlist forms first. Consultants and RFPs decide among the names already in view, and the average shortlist in the TrustRadius report held only 2.7 products.

### Can a newer cloud core compete with incumbents in AI answers?

It can be named if the evidence exists: live institutions, conversion records and independent coverage. Assistants cannot recommend what they cannot verify.

### Should we publish our conversion capacity and timelines?

Where your business allows, yes. Banks ask about timing early, CUInsight reports slots booking into 2029, and a vendor that states its capacity gives assistants and buyers a fact to repeat.

## Sources

- FDIC (2026-08), [Quarterly Banking Profile, Second Quarter 2026](https://www.fdic.gov/quarterly-banking-profile/quarterly-banking-profile-second-quarter-2026.pdf)
- NCUA (2026), [Quarterly Credit Union Data Summary, 2026 Q2](https://ncua.gov/files/publications/analysis/quarterly-data-summary-2026-Q2.pdf)
- PCBB BID Daily Newsletter (2025-04-23), [Core providers beefed up tech, but satisfaction can improve](https://www.pcbb.com/bid/2025-04-23-core-providers-beefed-up-tech-but-satisfaction-can-improve)
- Pulse 2.0 (2026-08), [Jack Henry: Faster payments revenue jumps nearly 50% as digital transaction revenue rises 11.6%](https://pulse2.com/jack-henry-faster-payments-revenue-jumps-nearly-50-as-digital-transaction-revenue-rises-11-6/)
- CUInsight (n.d.), [The core conversion crunch: why financial institutions must act now to secure their future](https://www.cuinsight.com/the-core-conversion-crunch-why-financial-institutions-must-act-now-to-secure-their-future/)
- Bank Director, via Nasdaq (2025-09-16), [Bank Director’s 2025 Technology Survey: Banks grapple with data, AI maturity](https://www.nasdaq.com/press-release/bank-directors-2025-technology-survey-banks-grapple-data-ai-maturity-2025-09-16)
- Cornerstone Advisors (n.d.), [Core banking systems](https://www.crnrstone.com/bold-solutions/transformation/core-banking-systems)
- Board of Governors of the Federal Reserve System, FDIC and OCC (2024-05-03), [Agencies issue guide to assist community banks to develop and implement third-party risk management practices](https://www.federalreserve.gov/newsevents/pressreleases/bcreg20240503b.htm)
- Alloy (2026), [2026 State of Fraud Report: key takeaways](https://www.alloy.com/blog/2026-state-of-fraud-report-key-takeaways)
- Federal Reserve Financial Services (2026-07-06), [FedNow Service volume and value statistics](https://www.frbservices.org/resources/financial-services/fednow/volume-value-stats)
- Responsive (2025), [GenAI overtakes search for a quarter of B2B buyers](https://www.responsive.io/news/buyer-intelligence-2025)
- Digital Commerce 360 (2026-01-05), [AI reshapes B2B buying and RFPs](https://www.digitalcommerce360.com/2026/01/05/ai-reshapes-b2b-buying-rfps/)
- MarketScale (2026), [94% of B2B buyers fact-check AI research outputs](https://www.marketscale.com/industries/business-services/94-of-b2b-buyers-fact-check-ai-research-outputs-and-vendors-are-underestimating-how-far-trust-has-fallen)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- OpenAI Help Center (2026), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Underneath (2026), [When does Google show an AI Overview? 1,248 US searches](https://underneath.agency/research/ai-overviews-frequency-study)
- Underneath (2026), [Do AI Overviews cite the pages that rank? 4,051 citations](https://underneath.agency/research/ai-overview-citations-study)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [How many “best of” lists cited by AI rank their own brand first?](https://underneath.agency/research/self-promoting-best-lists-study)

---

This is the Markdown twin of https://underneath.agency/resources/banking-technology-enterprise-deals-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How beauty retailers turn AI recommendations into sales"
description: "By being the store AI names when shoppers ask what to buy and where: accurate prices and stock, clear loyalty value, and content that answers routine questions."
canonical: "https://underneath.agency/resources/beauty-retailers-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How do beauty retailers turn AI product recommendations into sales?

By being the store an assistant names when a shopper asks what to buy and where to buy it, with prices, stock and loyalty value it can state correctly. Beauty shoppers already ask AI for product advice in large numbers, and the big retailers are building their own routes into ChatGPT and Google’s AI. This guide follows the path from an AI answer to a beauty order, and shows what a multi-brand retailer can do to be on it.

## The short version

1. Beauty shoppers already use AI: 49% have received beauty product recommendations from AI platforms such as ChatGPT, Claude, Gemini and Copilot ([NielsenIQ](https://nielseniq.com/global/en/insights/analysis/2026/the-ai-beauty-advisor-era-has-arrived-is-your-product-content-ready/)), and [Criteo found](https://www.criteo.com/blog/5-beauty-shopper-trends-2026/) that 57% of beauty shoppers who use AI assistants say the recommended products influence what they buy.
2. The big retailers have moved: Sephora launched an app inside ChatGPT with its loyalty perks, and Ulta Beauty began agentic checkout in Google’s AI Mode and the Gemini app, drawing on its 46+ million members ([Google Cloud and Ulta Beauty](https://www.googlecloudpresscorner.com/2026-04-22-Ulta-Beauty-and-Google-Introduce-Gemini-Enabled-Shopping-Experiences-That-Streamline-Beauty-Discovery-and-Purchase)).
3. Ulta reports “double the conversion and intent” from shoppers who find its products through Gemini and ChatGPT ([Reuters](https://www.zawya.com/en/business/americas/retailers-tap-ai-shopping-traffic-but-fight-to-keep-customer-data-424957)).
4. Amazon is the rival to beat: it sold an estimated $8.1 billion of beauty products in the first quarter of 2026, up 13%, and says shoppers who used its Rufus assistant to buy spent 80% more over the holidays ([Glossy](https://www.glossy.co/beauty/beauty-briefing-what-to-know-about-sephoras-and-ultas-ai-partnerships/)).
5. Beauty is a discovery category: up to 53% of purchases go to a brand the shopper had not bought in the previous 12 months, which is exactly the moment an AI answer can steer.

## Who buys beauty online, and what is a shopper worth to a retailer?

Repeat buyers who replenish and discover, so a shopper is worth a stream of baskets, not one order.

US beauty is large and still growing. [Circana reports](https://www.circana.com/post/us-prestige-and-mass-beauty-retail-deliver-a-positive-performance-in-2025-circana-reports) that prestige beauty retail sales grew 4% to $36 billion in 2025, while beauty sales in mass retail rose 5% to $72.7 billion. In the first quarter of 2026, prestige sales grew 6% to $8.1 billion and mass sales grew 7% to $18.1 billion. Online keeps gaining: McKinsey’s State of Beauty 2025, [as summarized by RetailBoss](https://retailboss.co/beauty-e-commerce-capture-31-percent-global-retail-sales-2030), puts ecommerce at 26% of global beauty retail sales in 2024, up from 15% in 2019, and expects 31% by 2030. Circana adds that hair is the only prestige category where online already accounts for the majority of retail sales.

The retailers that sell this category depend on loyal members. [Ulta Beauty](https://www.ulta.com/investor/news-events/press-releases/detail/234/ulta-beauty-announces-second-quarter-fiscal-2026-results-and) grew net sales 8.9% to $3.0 billion in its second quarter of fiscal 2026, helped by its acquisition of Space NK, and highlights its Ulta Beauty Rewards loyalty program in describing the business. Members return to replenish favorites, try new launches and redeem points, so the value of winning a new shopper sits in the baskets that follow the first one. Prestige houses face their own version of this, covered in [how luxury brands win high spenders](https://underneath.agency/resources/luxury-brands-ai-search).

## Where does AI already sit in the beauty shopping journey?

At the research stage, between a social video or a need and the choice of product and store.

The survey evidence is consistent. NielsenIQ reports that more than half of shoppers are interested in AI tools to help manage their shopping and that 49% have already received beauty recommendations from AI. Criteo’s Health & Beauty Pulse, based on surveys across six markets, found that 38% of shoppers already use AI assistants when shopping for beauty and personal care, and 57% of them say AI recommendations influence what they buy. NielsenIQ data cited by Glossy puts beauty-related searches on ChatGPT at over 1 billion a week.

Beauty journeys are long and wide. Criteo found that US shoppers browse an average of 19 makeup products and 12 fragrances before buying. Glossy’s reporters describe today’s path as fragmented: a TikTok, a Reddit post, a store visit to try products, then perhaps a checkout on Amazon. An assistant compresses that research into one conversation.

Across all US retail, [Adobe found](https://techcrunch.com/2026/04/16/ai-traffic-to-us-retailers-rose-393-in-q1-and-its-boosting-their-revenue-too/) that AI traffic converted 42% better than other traffic in March 2026, and in June that visitors referred by AI generated 41% higher revenue per visit. Those figures are not beauty-specific, but they match what Ulta reports for its own AI visitors. Outdoor gear shows the same contest between brand and specialty retailer for the store link in an answer, as our guide to [outdoor gear shortlists in AI answers](https://underneath.agency/resources/outdoor-brands-ai-search) shows.

## What do beauty shoppers ask AI before they buy?

Category, comparison and “where should I buy it” questions, often with a budget or hair type attached.

We wrote these sample beauty questions ourselves to show the range; none comes from real shopper logs:

- Category: “Best hair oil for frizzy, color-treated hair under $40.”
- Comparison: “Is this prestige setting spray worth it, or is there a drugstore alternative that works as well?”
- Where to buy: “Where can I get this fragrance in the 50ml size in stock today, with the best rewards?”
- Gifting: “A beauty gift set for a 30-year-old who loves clean fragrance, under $75.”
- Routine: “I bought this serum; what moisturizer and sunscreen go with it?”
- Loyalty and value: “Which beauty retailer gives the best birthday gift and points?”

For a multi-brand retailer, the third, fifth and sixth questions matter as much as the first. A brand only has to be named; a retailer has to be named as the place to buy, and then earn the add-on purchase. How brands get named for skin concerns is covered in [our skincare brand guide](https://underneath.agency/resources/skincare-brands-ai-search).

## How does an AI recommendation become a beauty retailer’s sale?

Through the merchant link: the assistant names a product, then a store, and the order lands where checkout is easiest.

The path has four steps, and the retailer can lose the shopper at each one:

1. **The product is named.** The assistant suggests a handful of products. Glossy’s hosts found the size of that list varies widely: Sephora’s assistant gave three or four products, Ulta’s gave 11 or 12, and ChatGPT gave 15. NielsenIQ warns that AI experiences “could reduce that visible choice to only one or two highly personalized recommendations.”
2. **A store is chosen.** [OpenAI documents](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search) that when a shopper clicks a product, ChatGPT may list merchants offering it, and “merchants are ranked based on factors like availability, price, quality, and whether they are the maker or primary seller.” The price shown first typically comes from the first merchant listed. A multi-brand retailer is competing with the brand’s own site and with Amazon for the same product. The brand side of that contest is covered in [how ecommerce brands get chosen by assistants](https://underneath.agency/resources/ecommerce-brands-ai-shopping-assistants).
3. **Checkout happens somewhere.** Sephora’s app inside ChatGPT lets shoppers use loyalty points and member perks such as free shipping and samples ([Cosmetics Business](https://cosmeticsbusiness.com/sephora-bets-on-ai-with-chatgpt-app-launch)). Ulta’s agentic checkout lets shoppers complete eligible purchases inside AI Mode and the Gemini app. Retailers still prefer the sale on their own site: Ulta’s head of digital told Reuters it would rather customers complete transactions there. Reuters also reported that OpenAI ended its Instant Checkout tool in March 2026; [FashionUnited](https://fashionunited.com/news/retail/us-consumers-would-let-ai-agents-buy-clothes-but-not-without-a-say/2026092974919) says the feature was scaled back, and OpenAI’s help page still describes it for some eligible merchants.
4. **The basket and the member.** On the retailer’s own site, the first order can grow into a routine, a loyalty sign-up and later replenishment. That is where a retailer’s advantage over a general-purpose assistant lies.

## What decides which beauty retailer an assistant sends shoppers to?

Platforms document product data, price and availability; the rest comes from surveys and our inference.

**Documented by the platforms.** OpenAI says ChatGPT considers structured product data from first-party and third-party providers, such as price and description, and other third-party content, and that review summaries are built from reviews on public websites. Its merchant ranking factors are availability, price, quality and whether the seller is the maker or primary seller. When merchants update prices, OpenAI notes there may be a delay before ChatGPT reflects them. Google and Ulta say the agentic checkout runs on the Universal Commerce Protocol, an open standard for agentic commerce.

**Observed by retailers and researchers.** Ulta reports double the conversion and intent from AI shoppers. NielsenIQ found that 84% of consumers are more likely to purchase when key product attributes can be easily compared, and 65% find personalized recommendations helpful.

**Our inference.** Because the brand’s own site is the “maker” of every product a retailer sells, a reasonable expectation is that retailers win the merchant slot on what they add: price and stock, exclusives, gift sets, samples, loyalty value and fast shipping, stated clearly enough for an assistant to read. Our article on [what drives AI product recommendations](https://underneath.agency/resources/what-drives-ai-product-recommendations) shows that hard facts such as ratings, prices and reviews tend to decide the pick when assistants can see them.

**Trust factors specific to beauty retail.** Authentic stock from authorized brands, honest reviews, [accurate shade and size options](https://underneath.agency/resources/cosmetics-brands-product-discovery-ai-search), return policies on opened products, and expert advice. Our article on [fake reviews and AI recommendations](https://underneath.agency/resources/fake-reviews-fake-brands-ai-recommendations) explains why review quality matters to assistants as well as to shoppers.

## What does a beauty retailer lose when AI sends shoppers elsewhere?

The first order, the basket that follows it and, often, the loyalty member.

Beauty is unusually open to switching. Criteo found that only 34% of shoppers say they rarely switch from their preferred beauty brands, and up to 53% of purchases come from shoppers buying a brand they had not bought in the previous 12 months. Each of those discovery moments is a chance for an assistant to send the shopper to a different store. Amazon’s own claim that Rufus users spent 80% more shows what a retailer gains when the assistant is its own, and what others risk when it is not.

There is also a data cost. Glossy’s reporting stresses that much of the value of AI shopping for retailers is knowing how a shopper started, compared and chose. When the research happens in someone else’s assistant and the checkout happens on Amazon, the retailer loses both the sale and the insight. We have not seen a study that measures how much beauty revenue retailers lose this way, so treat the size of the loss as unmeasured.

## How does GEO work for a beauty retailer?

It makes your store the easiest correct answer to “where should I buy this”; it cannot guarantee being named.

Generative engine optimization (GEO) for a multi-brand beauty retailer usually covers six pieces of work:

1. **Accurate, fast product data.** Keep prices, sizes, shades and stock current in the feeds assistants use, so the merchant list can name you. In our [study of AI pricing accuracy](https://underneath.agency/research/ai-pricing-accuracy-study), which covered software rather than beauty, only 61.9% of the plan prices assistants quoted were fully faithful to the official page.
2. **What only you offer, in plain words.** Exclusives, gift sets, samples, loyalty points and services such as in-store testing are reasons to choose a retailer over the brand’s own site; state them on product pages, not only in banners.
3. **Answers to routine and comparison questions.** Expert guides on routines, hair types and fragrance families give assistants something to cite. Our article on [the product content AI shopping assistants prefer](https://underneath.agency/resources/product-content-ai-shopping-assistants-prefer) explains what kind of copy helps.
4. **Reviews that hold up.** Encourage detailed, verified reviews on product pages; assistants summarize public reviews.
5. **Brand-partner consistency.** Make sure your product names and descriptions match the brand’s, so an assistant recognizes it is the same product you sell.
6. **Measurement by question and surface.** Track category, comparison and where-to-buy questions separately across ChatGPT, Gemini, Perplexity, Copilot and Google’s AI features, and connect them to orders, basket size and loyalty sign-ups. If a new launch you carry is missing from answers, [why ChatGPT misses new products](https://underneath.agency/resources/why-chatgpt-misses-new-products) walks through the usual causes.

## Which beauty retail questions can’t the evidence answer yet?

It does not show how often AI answers send beauty shoppers to one retailer rather than another.

The beauty surveys measure use and influence, not which store a shopper chose. Ulta’s “double the conversion” comes from the company itself, and Amazon’s Rufus figure is Amazon’s claim to Glossy. The Amazon beauty sales estimate comes from an agency’s data, not Amazon’s filings. Adobe’s figures cover all US retail. Glossy’s comparison of assistants was one reporter’s test, not a study. OpenAI and Google document what they consider, but not how they weigh one retailer against another for the same product. Checkout inside AI tools is also in flux: Reuters says OpenAI ended Instant Checkout in March 2026, while FashionUnited says it was scaled back.

## Where should a beauty retailer start?

Start by checking which store assistants name for your best-selling categories, and why.

A useful first step is an audit of the category, comparison, routine and where-to-buy questions your shoppers ask, across the main assistants, with the merchant each answer names and the price it quotes. That shows where you lose the sale to a brand site or a marketplace, and which fixes would move the most orders. If you would like us to run that audit with you and plan the work, tied to orders, basket size and loyalty sign-ups, [get in touch with our team](https://underneath.agency/contact). On our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page you can see how that audit leads into feed accuracy, clearer loyalty and exclusive details, and routine guides an assistant can cite.

## Frequently asked questions

### Do beauty shoppers buy what AI recommends?

Many say it influences them. In Criteo’s survey, 57% of beauty shoppers who use AI assistants said recommended products influence what they ultimately buy.

### Should a beauty retailer sell inside ChatGPT or Gemini?

It depends on the trade-off. In-chat checkout removes friction, but retailers such as Ulta prefer purchases on their own sites, where they keep the customer relationship and data.

### Why would an assistant send a shopper to the brand’s site instead of ours?

OpenAI lists “whether they are the maker or primary seller” among merchant ranking factors. Retailers need to make price, stock, perks and services clear to compete.

### Does Amazon have an advantage in AI shopping for beauty?

It has its own assistant, Rufus, and Amazon says Rufus users spent 80% more. Other retailers are partnering with ChatGPT and Google to compete.

## Sources

- NielsenIQ (2026), [The AI beauty advisor era has arrived: is your product content ready?](https://nielseniq.com/global/en/insights/analysis/2026/the-ai-beauty-advisor-era-has-arrived-is-your-product-content-ready/)
- Criteo (2026), [New Health & Beauty research: 5 shopper trends every marketer should know](https://www.criteo.com/blog/5-beauty-shopper-trends-2026/)
- Google Cloud and Ulta Beauty (2026), [Ulta Beauty and Google Introduce Gemini-Enabled Shopping Experiences](https://www.googlecloudpresscorner.com/2026-04-22-Ulta-Beauty-and-Google-Introduce-Gemini-Enabled-Shopping-Experiences-That-Streamline-Beauty-Discovery-and-Purchase)
- Reuters, via Zawya (2026), [Retailers tap AI shopping traffic but fight to keep customer data](https://www.zawya.com/en/business/americas/retailers-tap-ai-shopping-traffic-but-fight-to-keep-customer-data-424957)
- FashionUnited (September 29, 2026), [US consumers would let AI agents buy clothes, but not without a say](https://fashionunited.com/news/retail/us-consumers-would-let-ai-agents-buy-clothes-but-not-without-a-say/2026092974919)
- Glossy (2026), [Beauty Briefing: What to know about Sephora’s and Ulta’s AI partnerships](https://www.glossy.co/beauty/beauty-briefing-what-to-know-about-sephoras-and-ultas-ai-partnerships/)
- Glossy (2026), [Is agentic shopping the next big thing in beauty? Sephora and Ulta are betting yes](https://www.glossy.co/beauty/is-agentic-shopping-the-next-big-thing-in-beauty-sephora-and-ulta-are-betting-yes/)
- Cosmetics Business (2026), [Sephora bets on artificial intelligence with ChatGPT app launch](https://cosmeticsbusiness.com/sephora-bets-on-ai-with-chatgpt-app-launch)
- Circana (2026), [US Prestige and Mass Beauty Retail Deliver a Positive Performance in 2025](https://www.circana.com/post/us-prestige-and-mass-beauty-retail-deliver-a-positive-performance-in-2025-circana-reports)
- RetailBoss, summarizing McKinsey & Company (2025), [Beauty E-Commerce Set to Capture 31% of Global Retail Sales by 2030](https://retailboss.co/beauty-e-commerce-capture-31-percent-global-retail-sales-2030)
- Ulta Beauty (2026), [Second Quarter Fiscal 2026 Results](https://www.ulta.com/investor/news-events/press-releases/detail/234/ulta-beauty-announces-second-quarter-fiscal-2026-results-and)
- OpenAI (2026), [Shopping with ChatGPT Search](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search)
- TechCrunch (2026), [AI traffic to US retailers rose 393% in Q1, and it’s boosting their revenue too](https://techcrunch.com/2026/04/16/ai-traffic-to-us-retailers-rose-393-in-q1-and-its-boosting-their-revenue-too/)
- Underneath (2026), [AI pricing accuracy study](https://underneath.agency/research/ai-pricing-accuracy-study)

---

This is the Markdown twin of https://underneath.agency/resources/beauty-retailers-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Which pages should we target to show up in AI recommendations?"
description: "Ranked “best of” lists on other sites are the page type AI engines cite most when naming brands; your own site is a small share of citations."
canonical: "https://underneath.agency/resources/best-of-lists-ai-recommendations"
published: 2026-10-11
updated: 2026-10-11
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Which pages should we target to show up in AI recommendations?

Start with ranked “best of” lists on other people’s sites: they were the single most-cited content format in a study of 102 brands across five AI engines. A brand’s own website was a small share of citations, and AI search engines leaned heavily on third-party reviews and publishers. Being on the right list is not a guarantee of being named, but it puts you in the evidence the answer is built from.

## The short version

1. Ranked “best of” lists were 35.7% of the content pages AI engines cited across 102 tracked brands, more than any other format ([Kumar](https://arxiv.org/abs/2606.20065)).
2. Only 2.9% of citations in that study pointed at the brand’s own website; 75.2% pointed at other companies’ sites.
3. When a web-enabled GPT model was asked to rank cars in the US, 81.9% of the sources it used were independent publishers and review sites, not brand sites ([Chen and colleagues](https://arxiv.org/abs/2509.08919)).
4. In [our study of cited lists](https://underneath.agency/research/self-promoting-best-lists-study), answers that cited an independent list named its top pick 43.4% of the time; other answers to the same question named that pick 33.4% of the time.
5. In [our brand study](https://underneath.agency/research/brand-entity-ai-recommendations-study), coverage on independent sites was the strongest predictor of being recommended that we measured.

## Which kind of page do AI engines cite most when they recommend brands?

Ranked “best of” lists, the roundups titled “best CRM tools” or “top running shoes”. The largest study we found tracked 102 brands across ChatGPT, Gemini, Perplexity, Claude and Grok from March to May 2026, collecting 149,912 source citations ([Kumar](https://arxiv.org/abs/2606.20065)).

About 59% of cited pages were content, such as articles, guides and videos. The rest were homepages, product pages and other landing pages. Within content, the ranked list led.

| Content format | Share of cited content |
|---|---|
| Ranked “best of” list | 35.7% |
| General article | 31.0% |
| How-to guide | 9.7% |
| Comparison | 4.5% |

Counted across all citations, lists were about 21% of everything cited. The author argues that one list including your brand can be reused across many different questions. That is his reading of the citation shares, not a tested effect.

The author also co-founded the tracking platform the data comes from, so treat this as a vendor’s study.

## Why do other people’s pages matter more than your own site?

Because AI engines mostly cite third parties when recommending. In the same study, only 2.9% of citations pointed at the tracked brand’s own website. Another 75.2% pointed at other companies in the same space, such as competitors, peers and vendors.

Independent research shows the same tilt. University of Toronto researchers ran 1,000 ranking-style questions, such as “Rank the best smartphones from 1 to 10”, through Google and a web-enabled GPT model ([Chen and colleagues](https://arxiv.org/abs/2509.08919)). For US car questions, Google’s results were 45.1% independent publishers; the AI model’s sources were 81.9% independent publishers and review sites.

Engines differ in how far they lean. For questions about well-known brands, 93.5% of ChatGPT’s sources were independent publishers, against 63.4% for Gemini. Independent publishers were the majority of sources for every engine in that study; [your own pages carry less weight](https://underneath.agency/resources/do-ai-engines-cite-your-own-website) in this kind of answer.

## Does being on a list get you named in the answer?

It often does, but not always. In [our study of cited lists](https://underneath.agency/research/self-promoting-best-lists-study), we collected every page cited by six AI surfaces in the US on 26 September 2026 that was titled as a ranked list. There were 1,574 such pages.

When an answer cited an independent list, it named that list’s top pick 43.4% of the time. Other answers to the same question that did not cite the list named the same pick 33.4% of the time. The gap fits lists shaping answers, but our study measured association, not cause.

[Our brand study](https://underneath.agency/research/brand-entity-ai-recommendations-study) points the same way. Each tenfold increase in independent sites naming a brand in the pages an assistant cited went with 4.7 times the odds of that brand being recommended. It was the strongest predictor we measured, stronger than having a Wikipedia article.

## Should you publish your own “best of” list?

The evidence does not show a payoff. In our study, 24.2% of cited numbered lists with an identifiable publisher ranked their own publisher first. So AI engines do cite such lists.

But these lists were a small part of what answers cited: 1.1% of all citations. And answers named a self-ranking publisher at a rate we could not tell apart from an independent list’s top pick, 47.1% against 43.4%. Our study did not test whether readers trust such lists less.

## Which lists are worth pursuing?

The ones AI assistants already go looking for. In [our study of the searches assistants run](https://underneath.agency/research/ai-hidden-searches-study), ChatGPT ran a search aimed at a named publication, ranking or award in 43.8% of its answers.

When an assistant’s search named a source, such as NerdWallet or Avvo, the answer cited that source 44.0% of the time. In comparable answers that did not name it, the figure was 8.1%. That is an association, but it suggests the recognized rankings in your category are where to start.

Size matters too. Kumar found that on their first tracking run, global household-name brands appeared in 73% of category answers that did not name them, against 11% for small brands. His advice for small brands is to build presence in mainstream press, Wikipedia and YouTube before fine-tuning for each engine. Whether paid media can stand in for that presence is covered in [whether ad spend helps AI recommendations](https://underneath.agency/resources/does-ad-spend-help-ai-brand-recommendations).

## What should you do about it?

Treat inclusion in respected ranked lists as a core channel, alongside your own site.

1. Ask an AI assistant your category’s buying questions and note which lists it cites. Those lists, and the publishers behind them, are your target set.
2. Find the rankings and awards your category’s buyers trust. Assistants often search for these by name.
3. Give list editors what they need to judge you fairly: clear pricing, specifications, and access for independent testing.
4. Pursue several independent lists rather than one. Coverage across many independent sites was the strongest signal in our brand study. Once an AI agent reads about you directly, your own site matters more; see [off-site mentions versus a readable site](https://underneath.agency/resources/off-site-mentions-vs-site-readability-ai-agents).
5. Do not count on ranking yourself first on your own site. We found no evidence it helps.
6. Re-check every few weeks. Cited lists change between engines and over time.

If you want help building that coverage and tracking it, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

No study has tested whether getting onto a list causes AI engines to recommend a brand more.

- The main list figures come from one vendor’s tracking data on 102 brands, sampled for convenience and skewed toward software, fintech and consumer brands.
- Our list study comes from one collection date in the US. Its repeated runs were minutes apart, not days.
- Nobody has measured how long a list stays cited, or what happens when a list is updated and drops a brand.
- The research does not show whether readers trust AI answers that cite vendor-written lists less.
- Engines change their sources after updates, so today’s mix of cited page types may shift.

## Frequently asked questions

### What type of page do AI search engines cite most for brand recommendations?

Ranked “best of” lists. In a study of 102 brands across five AI engines, they were 35.7% of cited content pages, ahead of general articles at 31.0%.

### Does my own website get cited when AI recommends my category?

Rarely, in current data. Only 2.9% of citations pointed at the tracked brand’s own website in a study of 149,912 citations.

### Do AI engines cite brands’ own “best of” lists that rank themselves first?

Yes, sometimes. In our study, 24.2% of cited numbered lists with an identifiable publisher ranked that publisher first, but they were only 1.1% of all citations.

### Is PR coverage worth it for AI search?

The evidence leans that way, though it is observational. In our brand study, coverage on more independent sites went with 4.7 times the odds of being recommended for each tenfold increase.

## Sources

- Kumar (2026), [Generative Engine Optimization at Scale: Measuring Brand Visibility Across AI Search Engines](https://arxiv.org/abs/2606.20065), arXiv:2606.20065.
- Chen, Wang, Chen and Koudas (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919), arXiv:2509.08919.
- Underneath (2026), [How many “best of” lists cited by AI rank their own brand first?](https://underneath.agency/research/self-promoting-best-lists-study)
- Underneath (2026), [Do Wikipedia and schema make AI assistants recommend a brand?](https://underneath.agency/research/brand-entity-ai-recommendations-study)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)

---

This is the Markdown twin of https://underneath.agency/resources/best-of-lists-ai-recommendations. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How do beverage brands get recommended by AI assistants?"
description: "By being the clear answer to need-state questions: 27% of US consumers already took AI drink advice, and shoppers check ingredients, price and reviews."
canonical: "https://underneath.agency/resources/beverage-brands-ai-recommendations"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How does a beverage brand get its drinks recommended by AI assistants?

By being the clearest, best-supported answer to the need-state and occasion questions shoppers now ask: a drink for energy without a crash, a soda with less sugar, a bottle for Saturday’s dinner. More than a quarter of US consumers have already used AI to pick a beverage, so the brands named in those answers get the trial purchase, and the repeat habit that follows.

## The short version

1. In [EY’s 2026 Consumer Beverage Survey](https://www.ey.com/en_us/newsroom/2026/03/ey-consumer-beverage-survey) of 2,512 adults, 27% of US consumers had used AI-based beverage recommendations in the past year, and 45% of Brazilian consumers had.
2. Shoppers choose drinks by what is in them: 58% of US consumers pay attention to ingredients, 66% choose lower-sugar, lower-calorie drinks and 52% will pay more for drinks that support their health goals (EY).
3. The prize is large and shifting toward function: shoppers spent nearly $295 billion on US retail beverages in 2025, according to [FMI’s Power of Beverage 2026](https://www.supermarketnews.com/beverages/function-is-becoming-primary-beverage-purchase-driver-report) report built on Circana data, and energy drinks grew 15%.
4. A won shopper is a habit, not a single can: [NIQ](https://develop.nielseniq.com/wp-content/uploads/sites/4/2026/06/NIQ-Perspective-Functional-Beverages-June-2026.pdf) found functional-beverage spend per buyer rose 12.0% in a year, against 5.4% for all beverages, with millennials spending $153 a year each.
5. For wine, beer and spirits, a 2025 survey commissioned by the alcohol platform [DRINKS](https://aijourn.com/consumer-appetite-for-ai-driven-drink-recommendations-is-growing/) found 31% of US adults over 21 had already used AI to help choose alcohol.

## Who buys beverages now, and what is one new drinker worth?

One shopper buys on need, taste and price, and a new drinker’s value lies in a year of repeat purchases.

Beverages remain one of food retail’s strongest categories. FMI’s report, written with Circana sales data and a May 2026 survey of 2,003 US grocery shoppers, puts 2025 retail beverage spending at nearly $295 billion, up 3% on 2024. [Food Business News’ summary](https://www.foodbusinessnews.net/articles/30845-functional-beverages-are-having-a-moment) of the same report says non-alcoholic beverages grew 5% in dollar sales, led by energy drinks.

What shoppers say defines value is telling for anyone hoping to be recommended. Sixty-one percent say low price defines good value in a beverage, followed by taste (46%), trusted quality (40%) and a brand they trust (39%).

The purchase is cheap and frequent, so a single sale is small. The value sits in repetition. NIQ’s June 2026 analysis of functional beverages, drinks sold for energy, hydration, gut health, mood or focus, says “Habit is replacing trial.” It found a functional-beverage buyer spent $119 a year on average, millennials $153 and Gen X $136. Millennials drive 40% of the category’s dollars. For a brand, the commercial question is how many new people try the drink once, and how many of them make it part of a routine.

Big companies pay heavily for brands that have built that routine. In May 2025 [PepsiCo closed its acquisition of poppi](https://www.pepsico.com/newsroom/press-releases/2025/pepsico-completes-acquisition-of-poppi-accelerating-strategic-portfolio-transformation), the prebiotic soda, for $1.95 billion. PepsiCo credited poppi’s community- and culture-first approach, including viral TikTok campaigns and influencer partnerships. Our reading: discovery, not distribution alone, built that brand.

## Where do AI assistants already sit in a drink shopper’s decisions?

Early, at the moment a shopper names a need or an occasion and asks what to buy.

The EY survey, fielded online in late 2025 and weighted to the US and Brazilian census, gives the clearest picture so far:

- 27% of US consumers and 45% of Brazilian consumers used AI-based beverage recommendations in the past year, and 70% of Brazilians said they are very likely to use them next year.
- US consumers also find functional drinks through online grocery recommendations (19%), fitness and health apps (17%) and loyalty apps (16%).
- 80% of Gen Z and 75% of millennials drink functional beverages at least every two weeks, against 65% overall.

Alcohol shoppers are moving in the same direction. In the DRINKS survey of 1,000 US consumers, run by Dynata in 2025, 71% said they would be interested in AI help if liquor stores and online retailers offered it, 44% trusted AI to recommend a bottle of wine or liquor, and 39% would let it pick drinks for a party. Only 30% called human expert advice “very important” to alcohol purchases. DRINKS sells AI tools to alcohol retailers, so it has an interest in that result.

The assistants have also tried to become places to buy. In December 2025 [Instacart launched an app inside ChatGPT](https://www.supermarketnews.com/grocery-technology/instacart-launches-end-to-end-shopping-app-on-chatgpt) that built a grocery cart from a conversation and let the user pay through OpenAI’s Instant Checkout without leaving ChatGPT. That launch was documented by both companies. We follow that path from meal plan to cart in [our food ecommerce guide](https://underneath.agency/resources/food-ecommerce-sales-from-ai-search). OpenAI scaled the feature back in March 2026, according to [FashionUnited](https://fashionunited.com/news/retail/us-consumers-would-let-ai-agents-buy-clothes-but-not-without-a-say/2026092974919), though its help page still describes it for some eligible merchants. Across retail generally, [Adobe’s data reported by TechCrunch](https://techcrunch.com/2026/04/16/ai-traffic-to-us-retailers-rose-393-in-q1-and-its-boosting-their-revenue-too/) showed AI traffic to US retail sites up 393% year over year in the first quarter of 2026. Adobe did not break out beverages.

## What do drink shoppers ask AI assistants?

They ask about a need, an occasion or an ingredient, and rarely about a brand by name at first.

The prompts below are illustrative, written by us to show the shape of beverage questions. They are not captured from real users.

| Stage | Illustrative prompt |
|---|---|
| Need state | “Best electrolyte drink with no added sugar for long runs” |
| Alternative | “What’s a healthier soda that still tastes like soda?” |
| Comparison | “Olipop or poppi: which has less sugar?” |
| Ingredient check | “Energy drinks without sucralose or artificial colors” |
| Occasion | “Non-alcoholic drinks for a baby shower that feel festive” |
| Pairing and budget | “A red wine under $25 that goes with grilled salmon” |
| Where to buy | “Where can I buy this sparkling tea near me?” |

These questions match how the category is now sold. FMI says shoppers increasingly buy on functional benefits and specific occasions, and it urges retailers to organize the aisle around need states. NIQ’s advice to brands is to focus “less on claim variety and more on use occasions tied to consumer need states.” Our inference: an assistant answering a need-state question is doing exactly that sorting, out loud, for one shopper. Sportswear brands meet the same activity-based questions, as [what shoppers ask AI about training gear](https://underneath.agency/resources/sportswear-brands-ai-search) shows.

Each question can also turn into several searches. Google says AI Overviews and AI Mode [may use a “query fan-out” technique](https://developers.google.com/search/docs/appearance/ai-features), and that [AI Mode draws on shopping data for billions of products](https://blog.google/products/search/ai-mode-search/). Both are documented by Google. In [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study), ChatGPT looked for reviews in 46.2% of its answers to buyer questions.

## How does a mention in an AI answer become sales?

Through a trial purchase at a store, then repeat purchases if the drink earns a place in a routine.

The path for a beverage differs from a software purchase or a booked appointment:

1. **Named for a need.** The assistant puts a few drinks on a short list for “energy without a crash” or “alcohol-free for dinner.”
2. **Checked on the facts.** The shopper compares sugar, caffeine, ingredients and price, the attributes EY and FMI say drive the choice.
3. **Bought where it is convenient.** Most drinks are bought at grocery, convenience, club or online retailers; from late 2025 that list also included a cart built inside ChatGPT through Instacart. Some brands also sell bundles on their own sites. Turning those into repeat orders is covered in [winning subscribers through AI search](https://underneath.agency/resources/subscription-ecommerce-customers-ai-search).
4. **Repeated.** If the drink works for the shopper, the purchase becomes a weekly habit, which is where NIQ’s $119 to $153 a year per buyer comes from.

The first purchase may not even happen on your website, so the result shows up as retail velocity, the rate at which a product sells at each store, rather than as a web visit. When shoppers do click through, Adobe found AI visitors to US retail sites converted 42% better than other visitors in March 2026.

For wine, the direct channel is under pressure, which raises the value of every recommendation. [Sovos ShipCompliant’s 2026 report](https://sovos.com/press-releases/2026-direct-to-consumer-wine-shipping-report-reveals-record-declines-as-market-downturn-deepens/) found winery shipments to consumers fell 15% in volume and 6% in value in 2025, a loss of 967,000 cases. Napa wineries held up best, adding 1% to the value of their shipments. Our inference: when fewer people find wineries by visiting them, a recommendation for a specific bottle carries more weight.

## Why does an assistant name one drink and not another?

Mostly because of product facts, reviews and third-party coverage it can find; the exact rules are not public.

What is documented: [OpenAI says](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search) ChatGPT’s product results are not ads and are selected by ChatGPT based on the query and context. It names structured metadata from first-party and third-party providers, such as price and product description, plus other third-party content and review summaries drawn from public websites. Merchants on Shopify are already connected through Shopify Catalog, and others can apply to send a direct product feed.

What has been observed in studies (none of them on beverages):

- **Facts beat fame, until products look the same.** In tests of three AI systems on skincare products, [Chu and Hou](https://arxiv.org/abs/2606.17443) found rating, price and reviews explained 82.4% of how products were ranked and brand name only 1.2%. When every product had identical specs, the famous brand won all 670 valid trials. A challenger drink needs a clear, checkable difference.
- **Each assistant reads different sources.** For the same consumer shopping questions, [Uberti-Bona Marin and colleagues](https://arxiv.org/abs/2609.18729) found ChatGPT and Gemini shared only 5.4% of the websites they displayed, and editorial and review sites made up 56.7% of ChatGPT’s sources.
- **Lists change between asks.** In [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), only 25.2% of the brands ChatGPT named appeared in all five repeats of the same question.
- **Independent mentions matter.** In [our brand entity study](https://underneath.agency/research/brand-entity-ai-recommendations-study), each tenfold increase in the number of independent websites naming a brand went with 4.7 times the odds of being recommended.

Our inference for beverages: the trust factors are accurate nutrition and ingredient facts, retailer ratings and review volume, taste tests and “best of” lists from known publications, creator reviews, and wide retail availability. FMI found 45% of Gen Z say a drink becomes more appealing when they learn about it from a creator, so creator coverage may matter twice, to people and to the assistants that read about them. We explain the research on product facts in [what drives AI product recommendations](https://underneath.agency/resources/what-drives-ai-product-recommendations).

## What does a drink brand lose when it is missing from AI answers?

It loses trial at the moment a shopper picks a new drink, where the category’s growth now comes from.

Direct evidence of lost sales does not exist yet, so we label the reasoning:

- **The growth is in need states.** NIQ found energy claims added $2.7 billion in dollar growth, and FMI reported double-digit sales growth for drinks offering mood support, immune health and digestive health. These are exactly the questions shoppers bring to assistants.
- **Habits lock in.** NIQ’s point that habit is replacing trial cuts both ways. We infer that a shopper who settles on a rival’s drink after an AI answer may not try another for months.
- **Some product pages cannot be read.** Adobe found about 34% of retail product pages could not be properly accessed by AI. A drink whose details sit only in images or behind age gates may simply be skipped.

Treat any precise estimate of lost revenue from AI absence as a guess. No beverage company has published one.

## How does GEO work for a beverage brand?

Generative engine optimization (GEO) makes your drinks easy for AI assistants to find, describe accurately and support with outside evidence.

For a beverage company, that work usually covers:

1. **Clean product facts everywhere.** One consistent name per product and flavor, with sugar, calories, caffeine, key ingredients and pack sizes stated in text on your site, your retailer listings and your product feeds. Adobe’s finding on unreadable product pages makes this the first fix. If assistants already get your facts wrong, see [how to fix wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers).
2. **Pages for need states and occasions.** Plain pages that answer “drink for afternoon focus,” “alcohol-free options for dinner parties” or “what to pair with this wine,” with honest comparisons to the alternatives people consider.
3. **Third-party coverage.** Taste tests, roundups in food and drink publications, creator reviews on YouTube and TikTok, and genuine community discussion. For small brands, this is the main lever; see [how a small brand gets recommended by AI](https://underneath.agency/resources/how-small-brands-get-recommended-by-ai). Gift guides play a similar part for [toy brands answering gift questions](https://underneath.agency/resources/toy-brands-product-discovery-ai-search).
4. **Retail reputation.** Ratings and reviews on Amazon, Walmart, Target and grocery sites, plus clear where-to-buy information, because the purchase usually happens there.
5. **Product feeds and catalogs.** Keep Shopify Catalog, merchant feeds and any OpenAI product feed complete and current, since OpenAI documents that it reads that metadata.
6. **Responsible claims.** State functional benefits only as your evidence and labeling allow. In tests, [manipulative product copy was flagged and demoted](https://underneath.agency/resources/can-you-game-ai-shopping-rankings), and an exaggerated health claim is a legal and reputational risk, not a shortcut.
7. **Measurement.** Ask a fixed set of need-state and occasion questions in ChatGPT, Gemini and Google’s AI features many times, track how often your drinks appear and which sources are cited, and watch retail sales in the accounts those answers point to.

None of this guarantees a recommendation. It raises the odds that when an assistant looks for the best answer, it finds yours, stated accurately and backed by others.

## What can’t the evidence tell beverage brands yet?

It shows shoppers asking AI about drinks, not how much any brand sells because of it.

- **Self-reported surveys.** EY, FMI and DRINKS asked people what they did and would do. Behavior can differ, and the DRINKS survey was commissioned by a company that sells AI tools to alcohol retailers.
- **No beverage-specific studies of AI picks.** The ranking research we cite tested skincare and general consumer products such as electronics. We infer it applies to drinks because the purchase is similarly attribute-driven, but no one has tested beverages directly.
- **Store sales are hard to trace.** A recommendation that leads to a purchase in a supermarket aisle leaves no click. We found no beverage company that publishes sales attributed to AI answers. The broader question is weighed in [does AI visibility drive business results](https://underneath.agency/resources/does-ai-visibility-drive-business-results).
- **Alcohol is a special case.** Age rules and state-by-state shipping laws shape where alcohol can be bought, and we found no public data on how assistants handle alcohol shopping across markets.

## Where should a beverage brand start?

Start by asking the need-state and occasion questions your future drinkers ask, and see whose drinks are named.

That first check usually shows three things: which of your products appear and for which needs, which rival drinks and sources appear instead, and whether your sugar, caffeine, ingredient and price facts are quoted correctly. From there, the work is to fix the facts, earn the coverage and make your retail listings ready for the shoppers an assistant sends.

When growth depends on trial and repeat purchase, [send us a note](https://underneath.agency/contact) and we will show where your drinks appear in AI answers, why rivals are named instead, and which changes are most likely to put your products on more shoppers’ short lists. What that work covers, from clean nutrition facts and need-state pages to retail reviews and product feeds, is set out on our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page.

## Frequently asked questions

### Do people really ask ChatGPT which drink to buy?

Many do. In EY’s 2026 survey, 27% of US consumers had used AI-based beverage recommendations in the past year, and in a 2025 DRINKS survey, 31% of US adults over 21 had used AI to help choose alcohol.

### Does a famous beverage brand automatically win AI recommendations?

Not automatically. In one controlled study, the famous brand won every test when products were identical, but rating, price and reviews explained 82.4% of rankings once products differed. A clear, checkable difference gives challengers a chance.

### Can a drink be bought inside ChatGPT?

Groceries could. From December 2025, Instacart’s app in ChatGPT built a cart from a conversation and offered checkout without leaving ChatGPT, so drinks sold through Instacart’s retail partners could be bought there. OpenAI scaled in-chat checkout back in March 2026, according to FashionUnited, though its help page still describes it for some eligible merchants.

### Should functional drink brands make health claims to get recommended?

Only claims your evidence and labels support. Studies of AI shopping systems found manipulative copy was often flagged and demoted, and inflated claims create legal and reputational risk. This article does not give health advice.

## Sources

- EY (2026-03-09), [EY Consumer Beverage Survey: health-led choices, generational changes and digital discovery are redefining beverage expectations](https://www.ey.com/en_us/newsroom/2026/03/ey-consumer-beverage-survey)
- Supermarket News (2026-07-31), [Function is becoming primary beverage purchase driver: report](https://www.supermarketnews.com/beverages/function-is-becoming-primary-beverage-purchase-driver-report)
- Food Business News, Caleb Wilson (2026-08-13), [Functional beverages are having a moment](https://www.foodbusinessnews.net/articles/30845-functional-beverages-are-having-a-moment)
- NIQ (2026-06), [NIQ Perspective: Functional Beverages](https://develop.nielseniq.com/wp-content/uploads/sites/4/2026/06/NIQ-Perspective-Functional-Beverages-June-2026.pdf)
- DRINKS via The AI Journal (2025-04-22), [Consumer appetite for AI-driven drink recommendations is growing](https://aijourn.com/consumer-appetite-for-ai-driven-drink-recommendations-is-growing/)
- PepsiCo (2025-05-19), [PepsiCo completes acquisition of poppi, accelerating strategic portfolio transformation](https://www.pepsico.com/newsroom/press-releases/2025/pepsico-completes-acquisition-of-poppi-accelerating-strategic-portfolio-transformation)
- Supermarket News, Mark Hamstra (2025-12-08), [Instacart launches end-to-end shopping app on ChatGPT](https://www.supermarketnews.com/grocery-technology/instacart-launches-end-to-end-shopping-app-on-chatgpt)
- TechCrunch (2026-04-16), [AI traffic to US retailers rose 393% in Q1, and it’s boosting their revenue too](https://techcrunch.com/2026/04/16/ai-traffic-to-us-retailers-rose-393-in-q1-and-its-boosting-their-revenue-too/)
- Sovos ShipCompliant (2026), [2026 Direct-to-Consumer Wine Shipping Report reveals record declines as market downturn deepens](https://sovos.com/press-releases/2026-direct-to-consumer-wine-shipping-report-reveals-record-declines-as-market-downturn-deepens/)
- FashionUnited (September 29, 2026), [US consumers would let AI agents buy clothes, but not without a say](https://fashionunited.com/news/retail/us-consumers-would-let-ai-agents-buy-clothes-but-not-without-a-say/2026092974919)
- OpenAI (2026), [Shopping with ChatGPT Search](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- Google (2025-03-05), [Expanding AI Overviews and introducing AI Mode](https://blog.google/products/search/ai-mode-search/)
- Chu and Hou (2026), [Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems](https://arxiv.org/abs/2606.17443), arXiv:2606.17443.
- Uberti-Bona Marin and colleagues (2026), ["If I Had to Buy Just ONE: Galaxy S26 Ultra": Auditing AI-Generated Product Recommendations](https://arxiv.org/abs/2609.18729), arXiv:2609.18729.
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)
- Underneath (2026), [Do Wikipedia and schema make AI assistants recommend a brand?](https://underneath.agency/research/brand-entity-ai-recommendations-study)

---

This is the Markdown twin of https://underneath.agency/resources/beverage-brands-ai-recommendations. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Can AI search help biotechs find partners and investors?"
description: "Partly: AI tools now help investors and scientists screen biotech deals, so assets with public, cited science are easier to find before partnering talks."
canonical: "https://underneath.agency/resources/biotech-partnering-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Can AI search help a biotech company find partners, licensees and investors?

It can help a biotech get onto more scouting lists, but only when its science, pipeline and partnering status are public, consistent and backed by sources an assistant trusts. Investors and researchers already use AI tools to find and summarize deals and studies, while pharma licensing hit a 10-year high in 2025. No study yet shows how often pharma dealmakers use public AI assistants to scout assets, so treat this as a reasoned bet, not a proven channel.

## The short version

1. Licensing is where much of a biotech’s money now comes from: [J.P. Morgan’s Q4 2025 report](https://www.jpmorgan.com/content/dam/jpmorgan/documents/cb/insights/outlook/jpm-biopharma-deck-q4-2025.pdf), using DealForma data, counts $250.2 billion in announced biopharma licensing value across 516 deals in 2025, the highest since 2016, with upfront cash steady at 7% of deal value.
2. Venture money is scarcer: [IQVIA](https://www.iqvia.com/locations/emea/blogs/2026/01/biopharma-m-and-a-outlook-for-2026) puts 2025 US and European biotech venture funding at $24 billion across 410 rounds, down 14%, with only 11 US biotech IPOs.
3. Buyers are hungry: over $230 billion of industry revenue faces loss of exclusivity by 2030, IQVIA estimates, and emerging companies hold 70% of clinical-stage assets, most of them unpartnered.
4. The people who screen biotechs already lean on AI. In a [PitchBook and Web Summit survey](https://websummit.com/wp-media/2025/11/PitchBook_Web-Summit-Survey-Results-Press-Release_FINAL-Web-Summit.docx.pdf) of 116 investors, the top uses were summarizing due diligence materials (34%) and identifying relevant deals (26%); in [Elsevier’s survey](https://www.elsevier.com/about/press-releases/elseviers-global-survey-of-3-000-researchers-reveals-less-than-half-have) of 3,000 researchers, 58% use AI tools at work, up from 37% in 2024.
5. When AI answers cite sources, independent ones dominate: in a study of ChatGPT’s health answers, [Jacques and colleagues](https://arxiv.org/abs/2601.17109) found medical institutions supplied 30.6% of 615 cited sources and peer-reviewed journals 5.7%.

## Who does a biotech actually need to reach, and what is a deal worth?

Pharma dealmakers, investors and scientific collaborators. One license can bring a nine-figure upfront payment.

A biotech has three audiences, and each controls a different kind of money:

- **Pharma business development and licensing teams.** They decide whether to license an asset or buy the company. J.P. Morgan counts 41 licensing deals in 2025 with upfront payments above $100 million. IQVIA lists the largest: Pfizer paid a record $1.25 billion upfront for ex-China rights to 3SBio’s PD-1/VEGF bispecific antibody, and GSK paid $500 million upfront in a Hengrui alliance worth up to $12 billion.
- **Investors.** Venture, crossover and later public investors fund the years before a deal. Money is tighter and more concentrated: Massachusetts biotech companies raised $6.85 billion of venture capital in 2025, down 13% and the smallest total since 2019, according to a MassBio report [covered by Banker & Tradesman](https://bankerandtradesman.com/report-biotech-uncertainty-instability-bogged-down-2025/).
- **Research collaborators.** Academic labs, hospitals and other biotechs bring data, patients for trials and validation. They also cite your work, which matters for the other two audiences.

Platform biotechs add a fourth: pharma companies that buy discovery services. The landmark example is AI itself; IQVIA notes that Lilly and NVIDIA committed up to $1 billion over five years to a joint AI drug discovery lab.

The demand side is strong. IQVIA estimates big pharma’s deal capacity at $1.3 trillion and calls the patent cliff a source of “mounting pressure” to replenish revenue. Competition for that money is global: 40% of assets big pharma in-licensed in 2025 had a Chinese licensor, IQVIA reports, and the value of deals for China-originated assets rose from $45 billion to $105 billion. A US or European biotech is no longer competing only with its neighbors for a buyer’s attention.

## Where do AI assistants already show up in biotech dealmaking?

In investor screening and scientific literature research, documented in surveys; pharma scouting use of public assistants is unmeasured.

**Investors.** The PitchBook and Web Summit survey found 26% of investors use AI to identify relevant deals. S&P Global Market Intelligence’s [2025 outlook survey](https://press.spglobal.com/2025-04-01-S-P-Global-Market-Intelligences-Annual-Private-Equity-and-Venture-Capital-Outlook-Indicates-Optimism-Amid-Macroeconomic-Caution) of more than 100 private equity, venture and limited partner respondents found 31% saw due diligence as the most useful place for generative AI and 22% picked deal sourcing. Neither survey is limited to life sciences investors.

**Scientists.** Elsevier found that 61% of researchers using AI tools use them to find and summarize the latest research, and 51% use them for literature reviews. A collaborator looking for a delivery platform or a validated target may start with an AI summary of the field before reading papers.

**Pharma dealmakers.** Specialist scouting tools for business development teams are on the market, and data providers such as [CB Insights](https://www.cbinsights.com/research/report/book-of-scouting-reports-ai-drug-discovery) sell AI drug discovery scouting reports. We found no survey that measures how often licensing teams type questions into ChatGPT, Gemini or Perplexity. Our inference is that public assistants are used for early orientation (who works on a target, which modalities are crowded) while decisions still rest on paid databases, confidential data rooms and meetings.

That last step remains physical. [BioSpectrum Asia’s preview of BIO 2025](https://biospectrumasia.com/news/26/26201/bio-international-convention-2025-opens-in-boston-bringing-together-20000-global-biotech-leaders.html) reported 20,000+ registrants, 10,000+ partnering delegates and 60,000+ partnering meetings expected in Boston. AI search does not replace those meetings. It can influence who asks for one.

## What might a dealmaker or investor ask an AI assistant about your field?

Questions about targets, modalities, stages and who is still unpartnered. We wrote these examples ourselves; none were collected from real buyers.

| Audience | Illustrative question |
|---|---|
| Pharma licensing team | “Which biotechs have Phase 2 oral GLP-1 programs that are not yet partnered?” |
| Pharma search and evaluation | “Who is developing PD-1/VEGF bispecific antibodies outside China?” |
| Venture investor | “Which seed-stage radiopharmaceutical startups have published preclinical data?” |
| Crossover investor | “What readouts are expected in in vivo CAR-T in the next 12 months?” |
| Academic collaborator | “Which companies have lipid nanoparticle platforms that reach tissues beyond the liver?” |
| Platform customer | “Which AI drug discovery companies have taken a molecule into the clinic?” |

A single question rarely stays single. Google says its AI Overviews and AI Mode [may use a “query fan-out” technique](https://developers.google.com/search/docs/appearance/ai-features), running related searches on subtopics, and OpenAI says [ChatGPT search rewrites a prompt](https://help.openai.com/en/articles/9237897-chatgpt-search) into one or more targeted queries. A question about a target can quietly become separate searches for trials, publications and recent deals. When [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study) logged ChatGPT’s searches, 43.8% of its answers included a search aimed at a named publication, ranking or award. In biotech, the named authorities are journals, trial registries and the trade press.

## How does being named in an AI answer lead to a license or a funding round?

By adding your asset to a scouting list early, before the non-confidential deck and the partnering meeting request.

As we read the evidence, the path looks like this:

1. An analyst, investor or scientist asks about a target, modality or indication.
2. The answer names a handful of companies and cites the pages it used: publications, trial registry entries, press coverage, company pipeline pages.
3. The reader checks the cited science and the company’s own pipeline page, then requests a non-confidential overview or a meeting at the next partnering event.
4. Confidential diligence follows, then a term sheet: an upfront payment plus milestones for a license, a round for an investor, or a joint grant or data-sharing agreement for a collaborator.

The money sits at step 4, but the selection happens at steps 1 to 3. J.P. Morgan’s data show why the economics favor partnering now: licensing headline values rose mainly through milestone packages, while venture rounds concentrated in later-stage companies. For a preclinical or Phase 1 company, being found by a licensor may be the most realistic route to cash.

None of this shows up cleanly in analytics. A pharma scout who met your asset in an AI answer may contact you months later through a banker or a conference app. We discuss the attribution problem in [how to prove GEO caused sales](https://underneath.agency/resources/prove-geo-caused-sales).

## What makes an AI assistant cite one biotech and not another?

The companies do not publish their rules; studies point to independent, institutional and recent sources, which suits science.

**Documented by the platforms.** Google and OpenAI both say their AI answers search the web and show links to sources. Neither explains why one company working on a target is named and another is not.

**Observed in studies.**

- In consumer health questions, ChatGPT drew most citations from institutions such as medical centers, government agencies and Wikipedia, Jacques and colleagues found; peer-reviewed journals supplied 5.7% of the 615 cited sources. That study covers patient questions, not dealmaking, but it shows how heavily assistants lean on institutional names in health.
- For US software questions, AI search drew 72.7% of its sources from earned media, such as independent reviews and publications, [Chen and colleagues](https://arxiv.org/abs/2509.08919) report. Biotech has no equivalent study.
- In [our study of brand entities](https://underneath.agency/research/brand-entity-ai-recommendations-study), independent coverage was the strongest predictor we measured: each tenfold increase in the number of independent sites naming a brand in the cited pages went with 4.7 times the odds of being recommended.
- [Our freshness study](https://underneath.agency/research/ai-source-freshness-study) found the four assistants cited pages first published about half as long ago as Google’s top 10 for the same questions (ratio 0.50). A pipeline page last updated two readouts ago works against you.

**What scientists say they trust.** In Elsevier’s survey, 74% of researchers called peer-reviewed research trustworthy, and 59% said an AI tool that automatically cites references would increase their confidence in using it. Biotech buyers think the same way, we infer: an AI answer that cites your paper or trial record is more useful to them than one that cites your homepage.

**Our inference for biotech.** The proof a licensing team checks is largely public or could be: publications, posters, trial registrations, patents, press releases with data, and coverage in trade outlets. An assistant can find and cite those same documents. A company whose science lives only in a confidential deck gives both the human scout and the assistant very little to work with.

## What does a biotech lose when AI answers leave it out?

Mostly early meetings and inbound interest, though no study has priced that loss for biotech.

- **A crowded, global field.** With 40% of big pharma in-licensed assets now coming from Chinese licensors, a scout comparing options for a target sees more candidates than before. An asset absent from the first summary must be found some other way.
- **Fewer funding routes.** With US and European biotech venture funding at $24 billion and IPOs scarce, partnering is a funding source, not only an exit. Missing early scouting conversations can mean a longer, costlier runway, we infer.
- **Answers move.** A single good answer is not a fixed position. Assistants disagree with each other and with themselves from run to run, as our [guide to why AI answers about your brand change](https://underneath.agency/resources/why-ai-answers-about-your-brand-change) explains.
- **Wrong facts travel.** An outdated stage, a discontinued program or a confused company name can be repeated in answers. Our guide to [fixing wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) covers corrections.

## What does GEO work look like for a biotech?

Making your science, pipeline and partnering status easy to find, cite and repeat accurately, without promises of placement.

1. **One consistent entity.** Use the same company name, modality, targets, indications and stage on your website, registry entries, partnering-event profiles, investor materials and databases. Small companies with similar names are easy to confuse.
2. **A public, dated pipeline page.** State each program’s target, modality, indication, stage and whether it is available for partnering, with the date of the last update and links to data.
3. **Citable science.** Publish papers and posters, keep trial registrations current, and post data releases as plain web pages rather than only as slide PDFs. These are the sources researchers already say they trust.
4. **Independent coverage.** Trade press articles, conference presentations and partner announcements give assistants third-party confirmation. The general method is in [how brands build authority for AI search](https://underneath.agency/resources/how-brands-build-authority-for-ai-search); notability for an encyclopedia entry is discussed in [why Wikipedia matters for AI search](https://underneath.agency/resources/why-wikipedia-matters-for-ai-search).
5. **Plan around newness.** Young companies and new programs are often missing from assistants’ answers, as we explain in [why ChatGPT misses new products](https://underneath.agency/resources/why-chatgpt-misses-new-products).
6. **Measure several assistants.** Ask the questions a scout, investor or collaborator would ask in ChatGPT, Gemini, Perplexity, Claude, Copilot and Google’s AI features, repeatedly, and record which companies and sources come back.

Stay inside the rules. Overstated efficacy claims, selective data and planted coverage are scientific, legal and reputational risks for a company that will face diligence; see [legitimate GEO versus manipulation](https://underneath.agency/resources/legitimate-geo-vs-manipulation). Many of the startup lessons in [how AI startups win customers from AI search](https://underneath.agency/resources/ai-startups-customers-from-ai-search) also apply to young biotechs.

## Which parts of the AI scouting picture are still unproven?

The link between AI visibility and signed biotech deals is unmeasured; the surveys describe general AI use, not assistant-led scouting.

- **No pharma scouting data.** We found no survey of how often business development teams use public assistants to find assets.
- **The investor surveys are broad.** PitchBook’s sample is 116 investors attending one technology conference, and S&P’s covers private markets in general, not life sciences funds.
- **Health citation research covers patients.** The ChatGPT health study looked at consumer questions; dealmaker questions may draw on different sources.
- **Deal totals differ by database.** J.P. Morgan reports $250.2 billion of licensing value and IQVIA $232 billion for 2025, because they count deals differently. Both agree it was a decade high.
- **Nothing ties an AI mention to a term sheet.** The wider evidence is reviewed in [does AI visibility drive business results](https://underneath.agency/resources/does-ai-visibility-drive-business-results).

## Where should a biotech start if it wants more partnering conversations?

Ask the questions your future licensees and investors would ask, see which companies are named, then fill the public gaps.

List the targets, modalities and indications you work on, and write the questions a pharma scout, a fund analyst and an academic collaborator would ask about each. Ask each question in several assistants, then ask it again on another day. Record which companies appear, which papers and registries are cited, and whether your stage and partnering status come back correctly. In biotech, the gap is usually science kept in a confidential deck instead of on citable public pages.

If you want help, [talk to us about a partnering visibility review](https://underneath.agency/contact). It shows which of your programs surface when dealmakers and investors ask about your field, which sources the assistants rely on, and which missing public evidence is most likely costing you scouting calls, partnering meetings and term sheets. The [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page describes how we help move science out of the confidential deck and onto dated pipeline pages, citable data releases and independent coverage.

## Frequently asked questions

### Do pharma business development teams use ChatGPT to find assets?

No published survey measures it. Investors do use AI: 26% of investors in one survey use it to identify relevant deals.

### Should a biotech publish partnering status on its website?

Usually yes, for programs you want to out-license. A clear, dated statement helps human scouts and gives assistants something accurate to repeat.

### Do scientific publications help AI visibility?

Plausibly. Researchers trust peer-reviewed work, and AI health answers cite institutions and journals, but no study has tested biotech dealmaker questions.

### Can a preclinical company be found by AI assistants?

It can, if its data are public and covered by independent sources. Very new companies are often missing, so expect months, not weeks.

### Does GEO replace partnering events?

No. Events like BIO’s convention, with 60,000+ partnering meetings expected in 2025, remain where deals start. GEO can affect who requests a meeting.

## Sources

- J.P. Morgan (2025-12), [Q4 2025 Biopharma Licensing and Venture Report](https://www.jpmorgan.com/content/dam/jpmorgan/documents/cb/insights/outlook/jpm-biopharma-deck-q4-2025.pdf)
- IQVIA (2026-01-19), [Biopharma M&A: Outlook for 2026](https://www.iqvia.com/locations/emea/blogs/2026/01/biopharma-m-and-a-outlook-for-2026)
- Banker & Tradesman (2026), [Report: Biotech uncertainty, instability bogged down 2025](https://bankerandtradesman.com/report-biotech-uncertainty-instability-bogged-down-2025/)
- BioSpectrum Asia (2025-06), [Bio International Convention 2025 opens in Boston](https://biospectrumasia.com/news/26/26201/bio-international-convention-2025-opens-in-boston-bringing-together-20000-global-biotech-leaders.html)
- PitchBook and Web Summit (2025-11-11), [New PitchBook and Web Summit Investor Survey Finds AI Playing a Growing Role in Investment Decisions](https://websummit.com/wp-media/2025/11/PitchBook_Web-Summit-Survey-Results-Press-Release_FINAL-Web-Summit.docx.pdf)
- S&P Global Market Intelligence (2025-04-01), [Annual Private Equity and Venture Capital Outlook](https://press.spglobal.com/2025-04-01-S-P-Global-Market-Intelligences-Annual-Private-Equity-and-Venture-Capital-Outlook-Indicates-Optimism-Amid-Macroeconomic-Caution)
- Elsevier (2025), [Global survey of 3,000 researchers](https://www.elsevier.com/about/press-releases/elseviers-global-survey-of-3-000-researchers-reveals-less-than-half-have)
- CB Insights (2025), [Book of scouting reports: AI drug discovery](https://www.cbinsights.com/research/report/book-of-scouting-reports-ai-drug-discovery)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- OpenAI Help Center (2025), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Jacques and colleagues (2026), [Authority Signals in AI Cited Health Sources: A Framework for Evaluating Source Credibility in ChatGPT Responses](https://arxiv.org/abs/2601.17109), arXiv:2601.17109.
- Chen, Wang, Chen and Koudas (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919), arXiv:2509.08919.
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [Do Wikipedia and schema make AI assistants recommend a brand?](https://underneath.agency/research/brand-entity-ai-recommendations-study)
- Underneath (2026), [How fresh are the pages AI engines cite?](https://underneath.agency/research/ai-source-freshness-study)

---

This is the Markdown twin of https://underneath.agency/resources/biotech-partnering-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How building materials companies win specifications in AI search"
description: "Increasingly, yes: product research is moving online and into AI, so specifiers can only name materials whose performance data AI can find and check."
canonical: "https://underneath.agency/resources/building-materials-specification-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Will architects and contractors find our products when they ask AI what to specify?

They can, if the performance, compliance and sustainability facts behind your products are published where assistants and specifiers can verify them. Architects still use AI sparingly for product research, but they call it one of their least efficient tasks, and AI referrals to product research platforms are growing fast.

This guide is for manufacturers of building materials and building products: envelope, roofing, insulation, structural, concrete, interiors, finishes and fixtures. These companies win business when an architect or engineer writes a product into a specification, when a contractor buys it at bid time, and when a distributor keeps it on the shelf. The goal of AI visibility here is more specifications and better inquiries, not website traffic.

## The short version

1. AI use among architects is early: only 6% regularly used AI tools in their practice, according to the American Institute of Architects’ Journey to Specification research, as reported by [Architect magazine](https://www.architectmagazine.com/?p=332599) and [Dezeen](https://www.dezeen.com/2025/03/14/ai-architecture-study-american-architect/), although 53.1% had experimented with it.
2. Product research is where the opportunity sits: architects in that study named product research and updating product lists among their most inefficient tasks, yet fewer than 10% of firms had tried AI on those tasks.
3. Specifiers look for proof on origin and sustainability: in the [2026 edition of the AIA study](https://aiamaine.org/aiamainenews/2026/3/13/new-research-shows-architects-eager-for-greater-influence-in-building-product-innovation), 72% of architects preferred US-manufactured products, and the share that proactively recommend sustainability options rose from 58% in 2020 to 79% in 2025.
4. AI referrals are rising on product research platforms: [ArchiPro](https://business.archipro.co.nz/articles/archipro-building-and-interior-product-research-report), which serves New Zealand and Australia, recorded 69,192 sessions from AI platforms in the second quarter of 2026, up 662% on a year earlier.
5. Every specification feeds a very large build market: [US Census Bureau](https://www.census.gov/construction/c30/pdf/release.pdf) figures show construction spending running at an annual rate of $2,203.1 billion in August 2026.

## Who decides which building materials get used, and what is one specification worth?

Architects and engineers write the specification, contractors buy against it, and distributors supply it.

A building product usually has to win more than once. First, a specifier, such as an architect, specification writer, interior designer or engineer, names it as the basis of design. Then a contractor or installer prices it at bid time and may propose an “or equal” alternative. Finally, a distributor or dealer supplies it, and on residential projects the homeowner often has a say. Each of those people researches products, and each can now ask an assistant.

The work they choose materials for is substantial. In August 2026, the Census Bureau estimated annual rates of $882.3 billion for residential construction, $773.0 billion for private nonresidential construction and $547.8 billion for public construction.

Price pressure is real at the contractor stage. In the [AGC and Sage 2026 Construction Hiring and Business Outlook](https://www.agc.org/sites/default/files/users/user21902/2026%20Construction%20Hiring%20and%20Business%20Outlook%20Report_Final.pdf), 53% of contractors named materials costs among their biggest concerns for 2026. Makers of drones, robots and sensors meet the same buyers; see [how jobsite technology reaches contractor shortlists](https://underneath.agency/resources/construction-technology-demand-ai-search).

No public figure shows what one specification is worth, because order values depend on the product and project. Our inference: the value is in what follows. A product written into a specification is priced by every bidding contractor, and one that earns a place in a firm’s master specification can recur across many later projects. A product that is never considered is never priced at all.

## How much are specifiers, contractors and homeowners using AI today?

Lightly among architects, more among contractors and homeowners, and growing fastest in early product discovery.

Among architects, the Journey to Specification study, run by AIA with Deltek and ConstructConnect and answered by more than 500 AIA-registered architects, found only 8% of firms had built AI into their processes. Most architects who did use AI were using chatbots such as ChatGPT, grammar tools and image generators. A majority were optimistic about AI for complex problems (84%), but nearly all (90%) had concerns, led by inaccuracy.

Contractors have moved faster. In the [2026 Houzz State of AI in Construction and Design Report](https://www.constructionowners.com/press-release/houzz-survey-finds-ai-adoption-soars-among-construction-and-design-pros-while-homeowners-rely-on-the-experts), 52% of construction firms used AI for everyday business tasks. The same habit shapes [how contractors pick construction software](https://underneath.agency/resources/construction-software-customers-ai-search). Homeowners are experimenting too: 22% of renovating homeowners had used AI tools for their projects, and among them 50% used it to compare options.

Engineers who specify structural, mechanical and electrical products follow a similar pattern. In the [2026 State of Marketing to Engineers](https://advertising.globalspec.com/state-of-marketing-to-engineers-research-report) from TREW Marketing and GlobalSpec, 69% of technical buyers used generative AI during the purchasing process. The same survey found 76% routinely researched in online technical publications, slightly more than the 74% who used supplier and vendor websites.

The clearest product-discovery signal comes from ArchiPro, a platform where architects, builders and homeowners research building products in New Zealand and Australia. Its AI referrals generated 174,729 pageviews in the quarter, up 745% on a year earlier. ArchiPro notes that AI traffic is still relatively small. It is not US data, but it shows the direction of travel.

## What do specifiers and builders ask AI about materials?

Questions that combine performance, code compliance, sustainability, origin and the project type.

We wrote the following examples to illustrate the pattern. They were not collected from real users, and we have not tested what any assistant answers to them.

| Need | Illustrative question |
|---|---|
| Basis of design | “Continuous insulation options for a steel-stud wall in a four-story building that meets the fire test for exterior walls” |
| Code evidence | “Fiber cement cladding with a current ICC-ES evaluation report” |
| Sustainability | “Lower-carbon concrete mixes with product-specific environmental declarations available near Denver” |
| Origin | “Roof membranes made in the US for a Buy America project” |
| Substitution | “What is an approved equal to the specified acoustic ceiling panel for a school?” |
| Residential | “Composite or PVC decking for a coastal home with salt spray” |

Every one of these questions turns on documents a manufacturer controls: test reports, evaluation reports, listings, environmental declarations, origin statements, warranties and installation instructions. If those exist only inside gated downloads or scanned PDFs, an assistant may not be able to read them, and a specifier in a hurry may move on.

## How does an AI answer turn into a specification, a bid and an order?

Through early research, a basis-of-design choice, bidding, a submittal, then an order through distribution.

1. **Researched.** During design, an architect or engineer looks for products that meet performance, code and sustainability goals, using manufacturer sites, research platforms, representatives, peers and, increasingly, AI.
2. **Specified.** The product is named in the project specification, sometimes with listed alternatives.
3. **Bid.** Contractors price it. Some propose substitutions, which the design team accepts or rejects.
4. **Submitted and approved.** The contractor submits product data for approval before ordering.
5. **Ordered and installed.** The order goes through a distributor or dealer.
6. **Repeated.** A product that performs may be kept in the office’s master specification for future projects. That is our inference about how value compounds, not a measured rate.

The inquiry also tends to land on the manufacturer’s own site. In ArchiPro’s quarter, visitors who clicked through to suppliers’ websites generated 11,014 inquiries there, against 2,512 sent through the platform itself. Our inference: AI visibility mostly shows its value in steps one and two, and the evidence of it may arrive as an inquiry or a specification that no report attributes to AI.

## What makes an assistant name one product and not another?

Verifiable, consistent facts in sources it trusts; the platforms describe their search process, not their choices.

**Documented by the platforms.** [OpenAI says](https://help.openai.com/en/articles/9237897-chatgpt-search) ChatGPT search typically rewrites a question into one or more targeted queries sent to search partners, and that sites must allow OAI-SearchBot to be eligible. [Google says](https://developers.google.com/search/docs/appearance/ai-features) AI Overviews and AI Mode may use a “query fan-out” across subtopics, and that a page must be indexed and eligible to show a snippet to be a supporting link. Neither explains how products are chosen.

**Observed in our studies.** In [our study of assistants’ hidden searches](https://underneath.agency/research/ai-hidden-searches-study), ChatGPT searched for a named publication, ranking or award in 43.8% of its answers. In [our study of business facts in AI answers](https://underneath.agency/research/ai-business-facts-accuracy-study), phone numbers differed from the business’s profile 30.6% of the time when that number was missing from the business’s own website, against 1.6% when it was there. The study measured local businesses, not products, but the lesson carries: when your own sources disagree, answers drift.

**Observed in specifier research.** The 2026 AIA study found 73% of architects highly value being involved in product development by manufacturers, while only 24% want to take part directly. Specifiers want manufacturers who understand their problems, and evidence of that, such as published project collaborations, is something an assistant can find.

**Our inference for building materials.** The trust factors are the documents specifiers already check, published so they can be read and matched: performance data with the test standard and lab named, current code evaluation reports and listings, environmental declarations and origin statements, installation and warranty terms, completed projects naming the architect, and trade-press or award coverage. ArchiPro makes a similar point, arguing that products connected to completed projects, professionals and technical documentation give AI richer context than isolated product pages.

## What does a manufacturer lose when AI answers leave its products out?

It loses consideration at the design stage, where the specification, and often the order, is decided.

We found no measurement of specifications lost to AI, so we label our reasoning:

- **Early exclusion is hard to reverse.** If a product is not considered at design, the contractor prices whatever was specified, and a substitution request has to argue against it. That is our inference about how specification works.
- **Preferences are shifting toward proof.** With 72% of architects preferring US-made products and 79% proactively recommending sustainable options, missing origin or environmental facts can remove a product from a shortlist even when it qualifies.
- **Wrong facts are worse than none.** An assistant that repeats an outdated fire rating, a withdrawn evaluation report or a discontinued color can rule a product out. Our guide to [fixing wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) explains where to start.
- **The loss leaves no trace.** A specifier who never sees your product sends no inquiry. Our guide to [why analytics miss AI visibility](https://underneath.agency/resources/why-analytics-miss-ai-visibility) explains why such losses rarely show up in reports.

## How does GEO work for a building materials company?

Generative engine optimization (GEO) makes your products easy for assistants to find, describe accurately and support with independent evidence.

For a building materials manufacturer, the work usually includes:

1. **Readable product data.** Performance values, test standards, sizes, colors and limitations on product pages in plain text, not only inside data-sheet PDFs.
2. **Open technical files.** Specification text, CAD details and BIM objects available without a sign-up wall.
3. **Current approvals.** Evaluation reports, listings and certifications with their numbers, scope and dates, matching the issuing body’s own record.
4. **Sustainability and origin facts.** Environmental declarations, recycled content and where each product is made, stated per product rather than per company.
5. **Consistent channel data.** The same names, specifications and approvals on distributor, dealer and research-platform listings as on your site, so assistants do not meet conflicting facts.
6. **Built proof.** Project pages that name the building, the architect and the products used, published with permission.
7. **Independent coverage.** Trade publications, continuing-education courses, awards and association programs, the named sources assistants search for. Our guide on [how brands build authority for AI search](https://underneath.agency/resources/how-brands-build-authority-for-ai-search) covers the approach.
8. **Measurement.** Ask a fixed set of performance, code, sustainability and substitution questions across ChatGPT, Gemini, Perplexity, Copilot and Google’s AI features, repeat them, and track which products and sources appear. Compare the trend with inquiries and specification wins.

No manufacturer can make an assistant write its product into an answer. What this work can do is make your product the easiest one for a specifier, or an assistant, to check. Manufacturers that also sell made-to-order components can compare notes with our guide for [contract manufacturers](https://underneath.agency/resources/contract-manufacturers-rfqs-from-ai-search), and our guide on [product content AI shopping assistants prefer](https://underneath.agency/resources/product-content-ai-shopping-assistants-prefer) is useful for retail and residential lines.

## What can’t the specification data tell us yet?

It shows research habits and AI curiosity, not how many specifications AI answers have changed.

- **No attribution.** None of the sources links an AI answer to a specification, bid or order.
- **The AI finding is dated.** The 6% figure comes from an AIA study reported in March 2025; adoption may have moved since.
- **Interested parties.** ConstructConnect and Deltek sell to manufacturers and architects; ArchiPro sells listings to suppliers; Houzz runs a platform for building professionals; GlobalSpec sells advertising.
- **Different markets.** ArchiPro’s data covers New Zealand and Australia, not the US.
- **Our studies did not test material questions.** Applying their findings to specification is our inference.

## Where should a building materials company start?

Start by asking assistants the questions a specifier asks before naming a product, and see what they say about yours.

That first check usually shows whether your products appear for their performance, code and sustainability questions, whether the facts given match your current data, which platforms, publications and distributor listings the answers rely on, and which competing products appear instead. The work then is to publish your product facts in readable, consistent form and earn the independent coverage that confirms them.

If more of your revenue should start with a basis-of-design listing, [arrange a review of your products’ AI visibility with us](https://underneath.agency/contact). We test the performance, code and substitution questions specifiers ask in your category, trace which data sheets, listings and publications the answers draw on, and set out the product facts to publish so more specifications and qualified inquiries follow. Our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page explains how that work runs for a manufacturer, from readable data sheets and current approvals to matching distributor listings.

## Frequently asked questions

### Do architects use ChatGPT to find building products?

Some do, but regular use was low: 6% of architects in the AIA study used AI regularly, mostly chatbots, grammar tools and image generators. Product research is one of the tasks they find least efficient, which is where AI use could grow.

### Are BIM objects and specification text still worth producing?

Yes. Specifiers use them in design, and published, ungated files are also evidence an assistant can find. Keep them current and consistent with your product pages.

### Should we put our data sheets behind a form?

Gating can collect leads, but it can also hide your facts from assistants and from specifiers in a hurry. One middle path is to keep core technical data open and gate only deeper tools.

### Does AI visibility matter if we sell through distributors?

Yes, because the choice of product is often made before the distributor is involved. Keep distributor listings consistent with your own data so answers do not contradict you.

## Sources

- Architect Magazine (2025), [AI in Architecture: Revolution or Risk?](https://www.architectmagazine.com/?p=332599)
- Dezeen (2025-03-14), [Only six per cent of architects regularly using AI says AIA study](https://www.dezeen.com/2025/03/14/ai-architecture-study-american-architect/)
- American Institute of Architects via AIA Maine (2026-02-25), [New research shows architects eager for greater influence in building product innovation](https://aiamaine.org/aiamainenews/2026/3/13/new-research-shows-architects-eager-for-greater-influence-in-building-product-innovation)
- ArchiPro (2026), [ArchiPro Building and Interior Product Research Report](https://business.archipro.co.nz/articles/archipro-building-and-interior-product-research-report)
- Houzz via Construction Owners Club (2026-09-03), [Houzz Survey Finds AI Adoption Soars Among Construction and Design Pros](https://www.constructionowners.com/press-release/houzz-survey-finds-ai-adoption-soars-among-construction-and-design-pros-while-homeowners-rely-on-the-experts)
- Associated General Contractors of America and Sage (2026), [2026 Construction Hiring and Business Outlook](https://www.agc.org/sites/default/files/users/user21902/2026%20Construction%20Hiring%20and%20Business%20Outlook%20Report_Final.pdf)
- TREW Marketing and GlobalSpec (2026), [State of Marketing to Engineers research report](https://advertising.globalspec.com/state-of-marketing-to-engineers-research-report)
- US Census Bureau (2026-10-01), [Monthly Construction Spending, August 2026](https://www.census.gov/construction/c30/pdf/release.pdf)
- OpenAI Help Center (2026), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Google Search Central (2026), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [Do AI answers match a business’s Google profile?](https://underneath.agency/research/ai-business-facts-accuracy-study)

---

This is the Markdown twin of https://underneath.agency/resources/building-materials-specification-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How advisory firms get found when owners ask AI for help"
description: "Owners now ask AI about exits, cash and growth before calling anyone. Advisory firms get named by being specific, credentialed and independently covered."
canonical: "https://underneath.agency/resources/business-advisory-firms-clients-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How does a business advisory firm get found when owners ask AI for help?

By being the firm whose specialty, credentials and track record are easy for an AI assistant to find and confirm when an owner describes a problem. Owners of private companies already put growth, cash and exit questions to ChatGPT, and most of them are years behind on the planning an advisor would do. Firms that make their kind of help easy to understand and verify are the ones an answer can name.

## The short version

1. Owners already consult AI: in Wakefield Research’s 2025 survey for [TD Bank](https://stories.td.com/volumes/default/Wakefield-Research-Analysis-of-Results-for-TD-Bank-4.10.25.pdf), 30% of small business owners said they go to generative AI such as ChatGPT for advice, and 42% plan to use AI and other tools to improve their outlook, against 27% who plan to hire a financial advisor.
2. The exit wave is large and unprepared: the Exit Planning Institute’s research, as summarized by [T. Rowe Price](https://www.troweprice.com/financial-intermediary/us/en/insights/articles/2025/q4/adding-value-with-exit-strategy-planning-for-small-business-owners.html), found 75% of owners want to exit within a decade, representing $14 trillion of business wealth, yet its [2025 generational report](https://exit-planning-institute.org/hubfs/25SOOR-Generational.pdf) says only 13% of respondents have a formal exit plan.
3. Sellers come to market late: in the IBBA and M&A Source [Market Pulse survey for Q2 2026](https://www.webull.com/news/15462423712777216), between 60% and 90% of sellers had done less than one year of exit planning, or none.
4. Fractional leadership is growing from a small base: [Lightcast](https://lightcast.io/blog/rise-of-fractional-leadership) counted at least 34,000 US workers with “fractional” in their job title in 2025, up 265% since 2019, and 97% of 2026 fractional job postings came from small and medium companies.
5. AI answers lean on outside evidence: in our studies, ChatGPT ran 3.7 searches per buyer question, and the number of independent sites naming a brand was the strongest predictor of being recommended that we measured.

This article is about how advisory firms appear in AI answers. Nothing here is financial, tax, legal or valuation advice. If you run a CPA or tax firm, [our article on accounting firms in AI search](https://underneath.agency/resources/accounting-firms-clients-ai-search) covers that market; wealth managers who advise owners should read [how wealth management firms win clients when prospects check them with AI](https://underneath.agency/resources/wealth-management-firms-clients-ai-search).

## Who hires a business advisor, and what is an owner client worth?

The owner-CEO of a private company at a turning point. Growth stalls, cash tightens, a partner leaves or retirement nears.

That buyer is common and often older. The SBA’s Office of Advocacy counts [36.2 million small businesses](https://advocacy.sba.gov/?p=30239) in the United States, and T. Rowe Price, citing the Exit Planning Institute, reports that 51% of privately held businesses are owned by baby boomers. Few of them plan alone: 56% reported having a trusted group of advisors, typically a financial advisor, a CPA, an attorney, family members and other owners. Each of those advisers is now checked in AI answers too, as [how professional services firms win clients](https://underneath.agency/resources/professional-services-firms-clients-ai-search) explains.

What the firm sells depends on the moment, and so does what a client is worth:

- **Exit and succession planning.** A multi-year engagement. T. Rowe Price notes that an effective exit strategy typically takes three to five years to execute.
- **Sell-side advice and brokerage.** Paid largely at the sale. In Q2 2026 lower middle market deals averaged 11 to 12 months from engagement to close, and 87% of deals over $5 million attracted at least three offers, according to the IBBA and M&A Source survey of 255 business brokers and M&A advisors.
- **Fractional executives.** A monthly retainer that can run for years. Lightcast found finance accounts for 46% of fractional job postings in the United States, with business management at 12%. Finance specialists are covered in [winning CFO and deal work through AI](https://underneath.agency/resources/financial-consulting-firms-clients-ai-search).
- **Growth and turnaround advice.** Project work that often turns into an ongoing advisory seat.

Each of these is a long relationship with a high-trust buyer. A single well-matched owner can be worth years of fees, which is why where that owner first hears a firm’s name matters.

## Are business owners already asking AI for business advice?

Yes. A meaningful share of owners already use AI for advice, alongside their bank and their peers.

The clearest evidence is the TD Bank survey, conducted by Wakefield Research in March 2025 among owners of companies with 100 employees or fewer and revenue of $100,000 or more. When owners look for financial advice, the top source is a banking or financial partner (46%), followed by other small business owners (45%), and then generative AI (30%). Using AI, apps and websites was the top step owners planned to take to improve their outlook (42%), well ahead of hiring a financial advisor (27%).

The habit is broad. OpenAI’s analysis of 1.5 million conversations found that [49% of messages are “Asking”](https://openai.com/index/how-people-are-using-chatgpt/), a category it says shows people value ChatGPT most as an advisor. Age still matters, and it cuts both ways for this market. [Pew Research Center](https://www.pewresearch.org/short-reads/2025/06/25/34-of-us-adults-have-used-chatgpt-about-double-the-share-in-2023/) found 58% of adults under 30 had used ChatGPT, against 25% of those 50 to 64 and 10% of those 65 and older.

Our reading: older owners nearing an exit are less likely to type the question themselves, but their successors, younger co-owners and the CFO preparing the company for sale are more likely to. Younger owners also plan earlier. In the [Exit Planning Institute’s generational findings](https://blog.exit-planning-institute.org/key-owner-readiness-generational-soor), about 52% of millennial respondents named exit planning a top priority, while T. Rowe Price’s summary notes only 14% of baby boomers did.

## What do owners ask AI before they call an advisor?

Problem questions come first, then requests for names and reputation checks. We wrote the prompts below as examples of typical phrasing; none comes from a real owner’s chat history.

| Moment | Illustrative prompt |
|---|---|
| Exit, years out | “How do I make my HVAC company worth more before I sell it in five years?” |
| Choosing the kind of help | “Exit planning advisor vs business broker vs M&A advisor: which do I need for a $6M company?” |
| Succession | “How do I hand my family business to my daughter without hurting her siblings?” |
| Cash and finance | “Do I need a fractional CFO or a better bookkeeper at $8M revenue?” |
| Growth stall | “Advisors who help owner-run manufacturers in Ohio scale past $20M” |
| Checking a name | “Is [firm name] legit? What do business owners say about them?” |

Two things set this market apart. The owner often does not yet know which kind of advisor they need, so the answer to the “which do I need” question frames the whole search. And the advice is personal and confidential, so owners use AI to prepare privately before they talk to anyone, including the CPA or banker who would normally make the referral.

The need for that help is plain in the data. In the TD Bank survey, 34% of owners plan to retire in the next 10 years, 46% have no succession plan, and 23% don’t even know where to start. The Exit Planning Institute adds that only 5% of baby boomer respondents have a dedicated exit planning team, and only 27% of them have had a formal valuation, even though more than half plan to exit within five years.

## How does an AI answer turn into a discovery call?

Through education first. The answer explains the kind of help needed and names options, and the owner checks one before calling.

The path for an advisory firm usually looks like this: a problem question → an answer that defines the options (exit planner, broker, fractional CFO, a free public program) → a follow-up asking for firms that serve that kind of company → the firm’s page on that problem, its credentials and its people → a confidential discovery call → an engagement, often a retainer that can lead to a success fee at sale.

Two features of this market shape that path. First, outside help is now the norm for exits. In the Exit Planning Institute’s research, 68% of owners had sought outside advice on their exit in 2023, up from 38% in 2013. Second, the free alternatives are large. America’s SBDC says there are [nearly 1,000 local centers](https://americassbdc.org/about-us/) providing no-cost business consulting, and the [SBA lists these centers](https://www.sba.gov/local-assistance/resource-partners/small-business-development-centers-sbdc) among its resource partners. A reasonable expectation is that a general “where can I get help” question surfaces those public programs, while a specific question about a $10 million company preparing for sale is where private firms can be named.

Referral still matters at the end of the path. In a 2025 TaxDome survey of 350 US businesses reported by [CPA Practice Advisor](https://www.cpapracticeadvisor.com/2025/08/19/survey-of-smbs-shows-how-they-choose-and-evaluate-their-accountant-firm/167532/), 57% found their current accountant through a peer referral. We infer advisory firms see something similar: the referral supplies a name, and an AI assistant is one of the places the owner now goes to check it.

## What decides which advisory firms an AI assistant names?

Whatever evidence the assistant finds when it searches. That means independent coverage, credible directories and rankings, and plain pages about the firm’s specialty.

Here is what is documented by the platforms, what studies have observed, and what we infer.

- **Documented by Google.** [Google’s guidance on AI features](https://developers.google.com/search/docs/appearance/ai-features) says AI Overviews and AI Mode may use a “query fan-out” technique, issuing multiple related searches across subtopics to find supporting pages.
- **Documented by OpenAI.** [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search) typically rewrites a question into one or more targeted queries that it sends to search partners, and may use general location to make results local.
- **Observed in our studies.** In [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study), ChatGPT ran a mean of 3.7 searches per buyer question, and in 43.8% of its answers it searched for a named publication, ranking or award. In [our brand-entity study](https://underneath.agency/research/brand-entity-ai-recommendations-study), each tenfold increase in the number of independent sites naming a brand went with 4.7 times the odds of being recommended.
- **Observed: self-ranking lists are cited.** Of 269 AI-cited “best X” lists with an identifiable publisher, [24.2% ranked their own publisher first](https://underneath.agency/research/self-promoting-best-lists-study). Advisors publishing “top exit planners” lists that put themselves at number one is a pattern buyers and assistants both meet.
- **Observed: reputation checks rely on reviews.** In [our “is it legit?” study](https://underneath.agency/research/is-it-legit-ai-reputation-study), 88.0% of answers cited a review or complaint platform.

What we infer for advisory firms: credentials that can be checked in public, such as the Certified Exit Planning Advisor (CEPA) designation, the IBBA’s Certified Business Intermediary (CBI) or M&A Source’s Master Intermediary (M&AMI), give an assistant something solid to repeat. So do bylined articles in regional business journals, association talks and podcasts that name the firm next to the kind of owner it serves. A firm whose website says only “we help businesses grow” gives an assistant nothing specific to match to a question.

## What does it cost an advisory firm to be missing from AI answers?

The cost is the owner who prepares with AI, learns the options from it, and calls someone else. Nobody has measured it in dollars yet.

Timing makes this market unforgiving. The IBBA survey found retirement was the leading reason owners went to market across every deal segment, peaking at 72% of sales between $1 million and $2 million, and most came to market with less than a year of planning. An owner who decides to act tends to act within months, and the firm the owner hears about while researching is the one in the conversation when that happens.

The absence risk is our inference, not a measured loss. What is measured is how unstable visibility is. When [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study) put the same question to ChatGPT five times, just 25.2% of the brands it named turned up in every run. A firm that appears once may not appear the next time, and a firm with thin evidence appears less often. A two-partner advisory boutique is not shut out by that pattern; our article on [whether AI assistants favor big brands over smaller competitors](https://underneath.agency/resources/do-ai-assistants-favor-big-brands) looks at where smaller firms still get named.

The market for owner advice is also getting more crowded. Lightcast counted 677 US job postings for fractional roles in 2025, more than double the year before, and nearly 1,000 in the first seven months of 2026. More firms and independent advisors are publishing about the same owner problems, which makes specific, verifiable proof more valuable. Technology advisers, such as fractional CIOs, face a similar market; see [how IT consulting firms win clients](https://underneath.agency/resources/it-consulting-firms-clients-ai-search).

## How does GEO work for a business advisory firm?

Generative engine optimization (GEO) makes a firm’s specialty, credentials and proof easy for AI assistants to find, match and repeat.

1. **Problem pages in the owner’s words.** One page for each situation the firm handles well: preparing a company for sale in three to five years, family succession, a first fractional CFO, a growth stall at a specific revenue size. Name the industries and regions served and explain how an engagement works.
2. **The “which kind of help” page.** A plain comparison of exit planner, broker, M&A advisor and fractional executive, including when a free SBDC adviser is the better fit. Honest comparisons answer the question owners actually ask first.
3. **Checkable credentials and people.** Partner bios with designations, licenses, prior operating roles and industries served, matched across the firm’s site, LinkedIn, association directories and designation registries.
4. **Independent coverage.** Bylines in regional business journals, talks at trade associations in the industries the firm serves, podcast appearances and peer-group sessions. Our studies suggest independent coverage carries more weight than anything the firm says about itself; [how brands build authority for AI search](https://underneath.agency/resources/how-brands-build-authority-for-ai-search) explains the approach.
5. **Reviews and client stories within confidentiality.** Many owners cannot be named. Anonymized case notes with industry, size and outcome type, plus reviews where clients are willing, give assistants evidence without breaching trust.
6. **Consistent facts.** The same firm name, address, partners and services everywhere. If an answer already gets something wrong, [here is how to correct it](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers).
7. **Regular testing.** Ask the problem, comparison and reputation questions across ChatGPT, Gemini, Perplexity, Claude and Google AI Mode each quarter, several times each, and track which firms and sources appear.

Advisors who are also registered investment advisers, CPAs or attorneys must keep every public statement within their profession’s rules on advertising and testimonials. Those rules work in a firm’s favor, because a sober statement about a designation, a deal type or an industry served is exactly what an assistant can repeat without distortion. GEO cannot guarantee that any assistant names a firm; it makes the evidence about the firm accurate, specific and easy to find.

## Which questions does the current research leave open for advisory firms?

It shows how owners seek advice and how assistants search. It does not show how often a mention becomes an engagement.

- **No engagement data.** We found no public data linking an advisory firm’s AI visibility to discovery calls or signed clients.
- **No independent study of advisory answers.** We know of no independent, published study of which advisory firms AI assistants name. Our own studies cover other industries, and the patterns may not transfer.
- **Surveys from interested parties.** The owner-readiness figures come from the Exit Planning Institute, whose members sell exit planning, and the advice figures from a bank survey; the fractional counts come from job-title and posting data.
- **AI and valuations.** In the IBBA’s Q1 2026 survey, [67% of advisors reported no material impact of AI on valuations](https://aijourn.com/the-market-pulse-survey-q1-2026-reports-the-latest-trends-in-business-sales-up-to-50m/). How AI changes the deal itself is still unclear.

## Where should a business advisory firm start?

Pick the two or three owner situations you most want to win. Then see what AI assistants tell an owner in each.

Write down how an owner would describe each situation, including company size, industry, place and timing. Ask each question in ChatGPT, Gemini, Perplexity and Google AI Mode several times. Note which kinds of help the answer recommends, which firms and directories it names, which pages it cites, and what it says when asked about your firm by name. The gap between those answers and the owners you serve best is the plan.

If your partners want more discovery calls from owners who are a good fit, before a competitor or a free program frames the decision for them, [ask us how assistants currently present your firm to owners](https://underneath.agency/contact). We will test the owner questions that matter to your practice, show the sources assistants rely on, and plan the pages, listings and coverage that connect your firm to those owners. For a partner-level view of the engagement, the [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page walks through each step, from owner problem pages to checkable credentials and independent coverage.

## Frequently asked questions

### Do business owners really use ChatGPT for business advice?

Many do. In a 2025 TD Bank survey, 30% of small business owners said they go to generative AI such as ChatGPT for advice, behind their bank (46%) and other owners (45%).

### Will AI replace exit planners and M&A advisors?

Not on current evidence. In 2023, 68% of owners had sought outside advice on their exit, up from 38% in 2013, and 67% of deal advisors saw no material AI effect on valuations.

### Can a small advisory firm appear in AI answers next to large firms?

It can, when the question is specific. Assistants search for evidence, and independent sites naming a firm were the strongest predictor of recommendation in our study, ahead of size alone.

### Should we publish a “best advisors” list that ranks our own firm first?

We would not. AI engines do cite such lists, but self-ranking is easy for owners to spot, and independent coverage is the stronger and safer signal.

## Sources

- TD Bank and Wakefield Research (2025-04), [TD Bank Financial Preparedness Survey: analysis of results](https://stories.td.com/volumes/default/Wakefield-Research-Analysis-of-Results-for-TD-Bank-4.10.25.pdf)
- T. Rowe Price (2025), [Adding value with exit strategy planning for small business owners](https://www.troweprice.com/financial-intermediary/us/en/insights/articles/2025/q4/adding-value-with-exit-strategy-planning-for-small-business-owners.html)
- Exit Planning Institute (2025), [2025 State of Owner Readiness: Generational Report](https://exit-planning-institute.org/hubfs/25SOOR-Generational.pdf)
- Exit Planning Institute (2025-08-18), [Key Owner Readiness Findings Across Generations](https://blog.exit-planning-institute.org/key-owner-readiness-generational-soor)
- IBBA and M&A Source via PR Newswire (2026-08-25), [The Market Pulse Survey Q2 2026 Reports the Latest Trends in Business Sales up to $50M](https://www.webull.com/news/15462423712777216)
- IBBA and M&A Source via PR Newswire (2026-06-30), [The Market Pulse Survey Q1 2026 Reports the Latest Trends in Business Sales up to $50M](https://aijourn.com/the-market-pulse-survey-q1-2026-reports-the-latest-trends-in-business-sales-up-to-50m/)
- Lightcast (2026-08-20), [The Rise of Fractional Leadership](https://lightcast.io/blog/rise-of-fractional-leadership)
- U.S. SBA Office of Advocacy (2025), [New Advocacy Report Shows the Number of Small Businesses in the U.S. Exceeds 36 million](https://advocacy.sba.gov/?p=30239)
- U.S. Small Business Administration (n.d.), [Small Business Development Centers](https://www.sba.gov/local-assistance/resource-partners/small-business-development-centers-sbdc)
- America’s SBDC (n.d.), [About us](https://americassbdc.org/about-us/)
- CPA Practice Advisor (2025-08-19), [Survey of SMBs Shows How They Choose and Evaluate Their Accounting Firm](https://www.cpapracticeadvisor.com/2025/08/19/survey-of-smbs-shows-how-they-choose-and-evaluate-their-accountant-firm/167532/)
- OpenAI (2025-09-15), [How people are using ChatGPT](https://openai.com/index/how-people-are-using-chatgpt/)
- Pew Research Center (2025-06-25), [34% of U.S. adults have used ChatGPT, about double the share in 2023](https://www.pewresearch.org/short-reads/2025/06/25/34-of-us-adults-have-used-chatgpt-about-double-the-share-in-2023/)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- OpenAI Help Center (2026), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [Do Wikipedia and schema make AI assistants recommend a brand?](https://underneath.agency/research/brand-entity-ai-recommendations-study)
- Underneath (2026), [How many “best of” lists cited by AI rank their own brand first?](https://underneath.agency/research/self-promoting-best-lists-study)
- Underneath (2026), [“Is this brand legit?” How AI assistants build a reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study)
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)

---

This is the Markdown twin of https://underneath.agency/resources/business-advisory-firms-clients-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Could AI search engines filter out manipulative GEO content?"
description: "In lab tests, yes: one defense cut manipulative GEO’s success from 50.32% to 6.20% while keeping most useful sources. No live engine is known to use it yet."
canonical: "https://underneath.agency/resources/can-ai-search-filter-manipulative-geo"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Could AI search engines filter out manipulative GEO content?

In lab tests, yes: a two-stage defense cut the success of manipulative rewriting from 50.32% to 6.20% while keeping most useful sources. Simpler filters did far worse, and no study shows a live AI search engine running such a defense today. The practical lesson is that gains built on gaming an engine’s preferences, rather than on better content, are the gains most likely to disappear.

## The short version

1. A research defense called GEO Defender cut manipulative rewriting’s success rate from 50.32% to 6.20% and kept 94.12% of the honest sources answers relied on, in a lab test across five AI systems ([Li and colleagues](https://arxiv.org/abs/2609.02964)).
2. A simple filter for odd-looking text flagged only 15 of 616 manipulated documents, because optimized pages read like normal, well-written pages (Li and colleagues).
3. Telling an AI to “be skeptical” can backfire: in a test of 12 AI systems, it raised the rate at which they recommended fake products by 10.5 percentage points ([Luo and Chen](https://arxiv.org/abs/2606.13610)).
4. A detector for optimized pages was right 93.95% of the time on a test set of 562 pages, but its authors warn that subtle and human-made optimization is harder to catch ([Chu and colleagues](https://arxiv.org/abs/2608.16824)).

## What counts as manipulative GEO?

It is rewriting a page purely to win AI citations and sway answers, regardless of whether the page gets better. GEO, generative engine optimization, means adapting content so AI search engines select and cite it. [Li and colleagues](https://arxiv.org/abs/2609.02964) argue that optimization and manipulation “differ in intent rather than mechanism.”

Their definition of a GEO attack is narrow. It is a rewrite that pushes one document into the small set of sources an AI answer draws on.

The rewritten page “need not contain false statements or embedded instructions.” It simply fits the engine’s tastes more closely than its rivals do.

That is what makes filtering hard. Statistics, citations and an authoritative tone are what GEO adds.

As the authors put it, those same features “equally characterize high-quality benign content.” A filter that targets them risks removing good pages too.

## Can a defense really block it without hurting answers?

In one controlled test, mostly yes. Li and colleagues built a benchmark from 100 search questions. For each, they took ten real web pages and replaced one with a version rewritten by one of seven published GEO methods.

Without any defense, the AI answer cited the manipulated page about half the time: 50.32% on average across five AI systems, including GPT-5.5 and Claude Opus 4.8. With their two-stage defense, that fell to 6.20%. The defense still used 94.12% of the honest sources, compared with answers built from the untouched page set.

The defense works in two places. First, it re-orders the candidate pages so manipulated ones are less likely to make the shortlist.

Second, it gives the AI a written set of lessons about which sources to trust, learned from earlier successes and failures. On GPT-5.5, the re-ordering step alone cut attack success from 45.45% to 10.06%.

Two details matter for the real world. The defense was built from just 36 examples using three GEO methods, and it still worked on four methods it had never seen. And a judging AI found answer quality almost unchanged, with an average score drop of 0.01 on its comparison scale.

## Why don’t simple filters work?

Because manipulated pages look like good pages, and blunt warnings make AI systems discard good evidence or behave worse. The test compared three cruder approaches against the full defense:

| Approach | What happened in the tests |
|---|---|
| Filter out unusual-looking text | Flagged 15 of 616 manipulated documents, and 252 honest ones |
| Add a fixed “watch for manipulation” instruction | Modest protection; on GPT-5.5, use of honest sources fell from 89.66% to 71.67% |
| Tell the AI to be skeptical of unfamiliar brands | Fake-product recommendations rose 10.5 points across 12 systems |
| Keep only brands most sources agree on | Caught the fakes but dropped 52% to 79% of legitimate recommendations |

The first two rows come from [Li and colleagues](https://arxiv.org/abs/2609.02964). The odd-text filter failed because GEO rewriting “does not necessarily produce” strange-looking text. The safety instruction mostly worked by making the AI use fewer sources of any kind.

The last two rows come from [Luo and Chen](https://arxiv.org/abs/2606.13610), who planted fake product reviews in real search results. Their skepticism prompt backfired most on commercial systems, raising fake recommendations by 24 points on average. Ranking pages by publisher type, with editorial sites first and open forums last, helped every system but removed only 17% of fake recommendations. Whether a rival could use such tactics against you is covered in [how competitors game AI recommendations](https://underneath.agency/resources/can-competitors-game-ai-recommendations).

## Can engines detect optimized pages before they reach an answer?

Partly, and detection is improving fast. [Chu and colleagues](https://arxiv.org/abs/2608.16824) built a detector for pages rewritten by eight families of GEO methods. On a separate test set of 562 pages, it was right 93.95% of the time.

When they ran it on real Google and Gemini search results for 1,000 real user questions, it flagged 898 of 10,095 pages, or 8.90%. The authors stress these are estimates. Live pages have no ground truth, and “GEO is not inherently malicious,” so a flag does not mean a page is false.

They also name a weakness. Subtle edits and optimization done by people, rather than by AI tools, were harder to detect. A detector trained on today’s methods may miss tomorrow’s. Some self-promotion is plain to see yet still cited, as [our study of self-ranking “best of” lists](https://underneath.agency/research/self-promoting-best-lists-study) found.

## What happens to manipulative tactics when a defense is in place?

They stop paying, and optimizers drift back toward plain, factual writing. [Bagga and colleagues](https://arxiv.org/abs/2511.20867) tested this in a simulated shopping engine. They added one sentence to the ranking instructions asking it to demote manipulative product descriptions.

They then let an automated optimizer try to find wording that ranked well without being flagged. Starting from aggressive styles, the flag rate fell by 60 percentage points on average. [Product copy built on superlatives](https://underneath.agency/resources/can-you-game-ai-shopping-rankings) started out flagged 100% of the time and ended below 8%.

The winning wording was not cleverer manipulation. The final prompts converged on “careful, fact-grounded prose,” with phrases like “avoids exaggeration or manipulation.” Under even a simple defense, honest content was the best-ranking strategy.

Older tricks are already treated differently. Li and colleagues cite a 2026 study finding that classic black-hat SEO is removed by the retrieval systems of AI-enhanced search engines. Tactics aimed at the answer-writing step, they report, still worked there.

## What should you do about it?

Build visibility that would survive a defense like this, and audit anything that would not.

1. **Ask what each tactic adds for a reader.** If a rewrite only mimics what engines like, such as invented statistics or borrowed authority, assume a future filter will target it. We look at that risk in [when AI search optimization backfires](https://underneath.agency/resources/can-geo-backfire-on-your-brand).
2. **Prefer cooperative optimization.** One research method that rewrote pages to match engine preferences while keeping answers accurate improved visibility measures by 35.99% on average ([Wu and colleagues](https://arxiv.org/abs/2510.11438)). The same authors found adversarial methods “always degrade engine utility.”
3. **Keep facts checkable.** Defenses in these papers judge sources partly on how trustworthy the publisher is. Clear authorship, dates and sources help.
4. **Don’t rely on a gap that engines can close.** Engines change their pipelines without notice. Measure your AI visibility regularly so you see shifts early.

For help building durable visibility, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

Whether any live AI search engine uses defenses like these. The evidence has clear limits:

- GEO Defender was tested on a research pipeline with a small re-ranking system and 616 test cases, not on ChatGPT, Gemini or Google’s AI features in production.
- The shopping test used simulated rankers with a researcher-added instruction, not a deployed shopping assistant.
- The fake-review test used frozen search results, mainly in Chinese, from one snapshot in April 2026.
- No study yet measures how often real brands lose AI citations because a filter removed their pages.
- All of these are preprints, and none has been repeated by an independent team.

## Frequently asked questions

### Do ChatGPT or Google filter out GEO-optimized pages?

No published study shows it either way. The strongest evidence is a lab defense that cut manipulative rewriting’s success from 50.32% to 6.20%, not a test of a live engine.

### Will normal GEO work be penalized if engines add defenses?

Probably not, if it improves the page. In the lab test, the defense kept 94.12% of honest sources in use, and a separate study found fact-grounded writing ranked best under a simple defense.

### Can engines tell AI-rewritten content from GEO-optimized content?

That is the hard part. One detector reached 93.95% accuracy on a test set, but its authors found subtle and human-made optimization harder to catch.

### Why not just tell the AI to ignore manipulative sources?

Because it can backfire. In one test of 12 AI systems, a skepticism instruction raised fake-product recommendations by 10.5 percentage points on average.

## Sources

- Li, Shao, Lin, Guan, Zhou and Shi (2026), [When Optimization Becomes Manipulation: Defending Generative Search against Malicious Generative Engine Optimization](https://arxiv.org/abs/2609.02964), arXiv:2609.02964.
- Luo and Chen (2026), [One Polluted Page Is Enough: Evaluating Web Content Pollution in LLM Recommenders](https://arxiv.org/abs/2606.13610), arXiv:2606.13610.
- Chu, Leng, Li, Shen, Shen and Zhang (2026), [GEO-Flag: Detecting and Measuring GEO-Optimized Web Content](https://arxiv.org/abs/2608.16824), arXiv:2608.16824.
- Bagga, Farias, Korkotashvili, Peng and Wu (2025), [E-GEO: A Testbed for Generative Engine Optimization in E-Commerce](https://arxiv.org/abs/2511.20867), arXiv:2511.20867.
- Wu, Zhong, Kim and Xiong (2025), [What Generative Search Engines Like and How to Optimize Web Content Cooperatively](https://arxiv.org/abs/2510.11438), arXiv:2510.11438.

---

This is the Markdown twin of https://underneath.agency/resources/can-ai-search-filter-manipulative-geo. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Can a competitor game AI into recommending them first?"
description: "In lab tests, yes: hidden text on a product page pushed a product to the top of AI recommendations. Stronger, defended AI engines resist far better."
canonical: "https://underneath.agency/resources/can-competitors-game-ai-recommendations"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Can a competitor game AI assistants into recommending their product first?

In controlled tests, yes: researchers have pushed a chosen product to the top of AI recommendations by adding crafted text to its web page. On a live AI search engine the effect was smaller but still large. The strongest current AI models, and engines that add simple defenses, resist crude tricks much better, and real-world success rates appear low so far.

## The short version

1. In a 2024 test with ten made-up coffee machines, [Kumar and Lakkaraju](https://arxiv.org/abs/2404.07981) used added text to make a $199 machine the AI almost never showed its top pick in most cases.
2. On Perplexity’s online model, [Pfrommer and colleagues](https://arxiv.org/abs/2406.03589) raised promoted products by almost 3 positions on average with text planted on product pages.
3. A scan of 1.2 billion web addresses by [Khodayari and colleagues](https://arxiv.org/abs/2604.27202) found 1,521 hidden instructions aimed at reputation. AI models obeyed such instructions at most 8% of the time.
4. In [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), ChatGPT’s first pick changed at least once for 80.0% of questions over five runs, so one answer cannot tell you whether a rival is gaming it.
5. Engines can defend themselves: one defense tested by [Li and colleagues](https://arxiv.org/abs/2609.02964) cut attack success from 50.32% to 6.20%.

## How have researchers made an AI recommend one product first?

By adding a carefully crafted string of text to the product’s own information page. [Kumar and Lakkaraju](https://arxiv.org/abs/2404.07981) built a catalog of ten fictitious coffee machines and asked an open-source AI model for affordable options. They then used an automated search to find a text string that, when added to one product’s description, made the AI list that product first.

The first target, ColdBrew Master, cost $199 and the AI almost never recommended it for an “affordable” request. With the added text, it became the top recommendation in most cases. The second target, QuickBrew Express ($89), usually ranked second. The added text often lifted it to first place.

The trick was fragile at first. When the order of the product list was shuffled, the first version of the text helped in about 40% of the evaluations and changed nothing in about 60%. For the second product, a version tuned to one fixed order was as likely to hurt as to help. Tuning the text across shuffled orders made it far more reliable.

This was a lab setting: one open-source model, Llama-2, and a made-up catalog. It was not tested on ChatGPT, Gemini or other live products.

## Does it work on real AI search engines?

Partly: tests on a live engine showed real but smaller gains than in the lab. [Pfrommer and colleagues](https://arxiv.org/abs/2406.03589) collected 1147 webpages from manufacturers’ own sites across 50 product categories. For each category, they tried to promote the product the AI ranked lowest by inserting instructions into its page.

They then hosted manipulated pages online and asked Perplexity’s Sonar Large Online model to read and compare them. The planted text was repeated 15 times through each page. The authors note it could be made invisible to human visitors with ordinary web tricks. Promoted products rose by almost 3 positions on average, and more than half the gap to the top spot.

There are limits. The researchers handed the engine the page addresses directly, so the test skipped the step where the engine finds pages on its own. The full Perplexity product was only tested informally, in 2024.

A later study, summarized in [a 2026 survey by Martinez](https://arxiv.org/abs/2607.14035), went further. Using pages the researchers controlled on production AI search engines, manipulation raised the recommendation rate of a fictitious camera from 34.0 to 59.4%.

## Are smarter AI models harder to fool?

Not automatically, though the newest models and simple defenses do resist crude tricks well. In the Pfrommer tests, GPT-4 Turbo was more vulnerable than the older GPT-3.5, which led the authors to conclude that better capability does not bring built-in protection.

Other results are more reassuring for honest brands:

| Test | What happened |
|---|---|
| Hidden instructions found on real web pages, given to 13 AI models in a page-summary task ([Khodayari and colleagues](https://arxiv.org/abs/2604.27202)) | Small models obeyed up to 8.0% of the time on plain text; closed-source models 0.6% overall |
| Fourteen manipulative rewrites of shopping listings, with a one-line warning added to the AI’s instructions ([Bagga and colleagues](https://arxiv.org/abs/2511.20867)) | GPT-5 and Claude gave them little ranking gain, whether or not they flagged them |
| A two-stage engine defense across seven attacks ([Li and colleagues](https://arxiv.org/abs/2609.02964)) | Attack success fell from 50.32% to 6.20% on average |

These are all controlled tests. They show that defenses exist, not that every engine uses them. The shopping tests are covered in [whether sellers can game AI shopping rankings](https://underneath.agency/resources/can-you-game-ai-shopping-rankings).

## Is anyone actually doing this today?

Yes, on a small scale, and often out of sight. [Khodayari and colleagues](https://arxiv.org/abs/2604.27202) analyzed 1.2 billion web addresses and confirmed 15.3K hidden AI instructions on 11.7K pages. Of these, 1,521 were reputation manipulation across 139 sites, mostly product or content promotion, forced citations and demands for positive reviews. About 87% of all the injections they found were not visible to human readers. We cover that scan in [websites hiding instructions for AI](https://underneath.agency/resources/websites-hiding-instructions-for-ai-search).

Openly self-serving content is far more common than hidden code. In [our self-ranking lists study](https://underneath.agency/research/self-promoting-best-lists-study), 24.2% of the “best X” lists with an identifiable publisher cited by six AI surfaces ranked their own publisher first. Yet answers named the top pick of these lists at about the same rate as independent lists: 47.1% against 43.4% of answer and list pairs. Ranking yourself first did not show a measurable advantage in that sample.

## Would you notice if a rival gamed the answers?

Probably not from a single check, because AI rankings already move around on their own. [Pfrommer and colleagues](https://arxiv.org/abs/2406.03589) point out that a reader cannot tell from the output whether the AI was deceived, since nobody knows what the “correct” order should be.

Natural variation makes this harder. In [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), we asked ChatGPT, Gemini and Perplexity the same 20 buyer questions five times on 26 September 2026. ChatGPT named the same first brand in two runs only 53.8% of the time, and its first pick changed at least once for 80.0% of questions. A competitor jumping to first place in one answer may be noise.

## What should you do about it?

Monitor patterns over time, and compete on evidence rather than tricks. Practical steps:

1. Track your category’s AI answers repeatedly, across several assistants and wordings, so you can see a sustained shift rather than one-off noise.
2. When a rival suddenly dominates, read the pages the answers cite. Look for hidden text, invented statistics or fake reviews.
3. Report clear manipulation to the AI platform and, where relevant, to consumer protection bodies.
4. Do not copy the tactic. Engines are building [defenses that demote manipulated pages](https://underneath.agency/resources/can-ai-search-filter-manipulative-geo), and planted instructions can damage your brand if discovered.
5. Make your own pages easy to verify: clear specifications, real reviews and independent coverage give an AI something solid to cite.

If you want help building visibility that holds up as engines tighten their defenses, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The research shows the vulnerability exists but has not measured how often rivals succeed on live assistants. Specifically:

- The strongest results come from open-source models, made-up catalogs or pages handed directly to the engine.
- Tests on live products, such as Perplexity, date from 2024; the products have changed since.
- No study we reviewed tracks a real competitor gaming ChatGPT, Gemini or Google’s AI features over time.
- The in-the-wild scan finds hidden instructions but cannot say how many changed real answers.
- Defenses were tested in research settings; we do not know which ones each engine runs.

## Frequently asked questions

### Can a company pay or trick ChatGPT into recommending it first?

Lab studies show that crafted text on a product page can move an AI’s recommendation, but tests on current commercial assistants are limited. In one 2026 study, closed-source models obeyed hidden page instructions only 0.6% of the time.

### How would a competitor manipulate AI recommendations?

The documented methods are crafted text strings, hidden instructions in a page’s code, and planted content such as fake reviews. In one scan, about 87% of hidden AI instructions were invisible to human readers.

### How can I tell if a competitor is gaming AI answers?

Look for a sustained pattern across many runs and assistants, then inspect the cited pages. In our study, ChatGPT’s first pick changed at least once for 80.0% of questions, so single answers are unreliable evidence.

### Should my brand use these tactics too?

No. Engines are adding defenses, one of which cut attack success from 50.32% to 6.20% in testing, and hidden manipulation carries legal and reputational risk.

## Sources

- Kumar, A. and Lakkaraju, H. (2024), [Manipulating Large Language Models to Increase Product Visibility](https://arxiv.org/abs/2404.07981), arXiv:2404.07981.
- Pfrommer, S., Bai, Y., Gautam, T. and Sojoudi, S. (2024), [Ranking Manipulation for Conversational Search Engines](https://arxiv.org/abs/2406.03589), arXiv:2406.03589.
- Khodayari, S., Zhang, X., Acharya, B. and Pellegrino, G. (2026), [Indirect Prompt Injection in the Wild: An Empirical Study of Prevalence, Techniques, and Objectives](https://arxiv.org/abs/2604.27202), arXiv:2604.27202.
- Bagga, P. S., Farias, V. F., Korkotashvili, T., Peng, T. and Wu, Y. (2025), [E-GEO: A Testbed for Generative Engine Optimization in E-Commerce](https://arxiv.org/abs/2511.20867), arXiv:2511.20867.
- Li, H., Shao, Y., Lin, X., Guan, Z., Zhou, M. and Shi, J. (2026), [When Optimization Becomes Manipulation: Defending Generative Search against Malicious Generative Engine Optimization](https://arxiv.org/abs/2609.02964), arXiv:2609.02964.
- Martinez, O. (2026), [Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026)](https://arxiv.org/abs/2607.14035), arXiv:2607.14035.
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)
- Underneath (2026), [How many “best of” lists cited by AI rank their own brand first?](https://underneath.agency/research/self-promoting-best-lists-study)

---

This is the Markdown twin of https://underneath.agency/resources/can-competitors-game-ai-recommendations. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Can optimizing for AI search backfire on your brand? | Underneath"
description: "Yes. In controlled tests, common rewriting tricks often lowered rankings and hype got flagged, but careful optimization did not harm answer quality."
canonical: "https://underneath.agency/resources/can-geo-backfire-on-your-brand"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Can optimizing for AI search backfire on your brand?

Yes, it can, and in controlled tests it often did. Popular rewriting tricks lowered pages in AI answers more often than many teams expect, and hype or manipulation was flagged or demoted when engines had even simple defenses. Careful work that adds real substance did not show the same downside.

## The short version

1. In a benchmark of AI answer engines, adding statistics to pages lowered their rank in 19 of 24 test settings ([Puerto and colleagues](https://arxiv.org/abs/2506.11097)).
2. In a shopping test, product descriptions rewritten with superlatives were flagged as questionable 100% of the time by every AI ranker that had a simple warning built in ([Bagga and colleagues](https://arxiv.org/abs/2511.20867)).
3. A research defense against manipulative rewriting cut its success rate from 50.32% to 6.20% ([Li and colleagues](https://arxiv.org/abs/2609.02964)).
4. Some defenses also catch honest brands: filters that caught at least 90% of planted fake brands also threw out roughly two thirds of genuine recommendations ([Luo and colleagues](https://arxiv.org/abs/2606.13610)).
5. Optimization does not have to make answers worse: a careful rewriting method raised visibility by 35.99% on average while keeping answer quality ([Wu and colleagues](https://arxiv.org/abs/2510.11438)).

## Can GEO tactics lower your visibility instead of raising it?

Yes. In the largest independent test of common rewriting tricks, many made pages less visible, not more.

The C-SEO Bench study, by [Puerto and colleagues](https://arxiv.org/abs/2506.11097), rewrote pages using popular generative engine optimization (GEO) methods and measured where AI answers then cited them. Out of 54 combinations of method and topic, only 3 showed a reliable improvement. Adding statistics, a widely repeated tip, lowered a page’s rank in 19 of 24 test settings. For product recommendations on Anthropic’s Claude Haiku 3.5, 26 out of 30 cases moved pages down.

The effects were also unpredictable. For one rewriting method on retail products, pages moved up in 26.2% of cases and down in 12.8%. In 61.0% of cases, nothing changed. A 2026 survey of 45 studies by [Martinez](https://arxiv.org/abs/2607.14035) describes a test that kept a simulated search step in place. There, rewriting only the page body cut how often pages reached the top 10 after re-sorting by 16%. Our related guide covers [how GEO content tactics hold up in tests](https://underneath.agency/resources/do-geo-content-tactics-work).

## Do AI engines penalize content that looks manipulative?

Sometimes, and simple warnings were enough in tests. Hype and invented proof stood out to the stronger AI rankers.

[Bagga and colleagues](https://arxiv.org/abs/2511.20867) built a shopping test with Amazon product listings and five AI rankers. They added one sentence to each ranker’s instructions, asking it to down-rank and flag misleading descriptions. Descriptions rewritten with superlatives were then flagged 100% of the time by every ranker. On GPT-5, those descriptions fell an average of 4.14 places. When the researchers let software rewrite copy to dodge the warning, it drifted to careful, fact-based prose. Their reading is that rank gains under the warning came from genuine content improvement, not manipulation.

Defenses built by researchers go further. [Li and colleagues](https://arxiv.org/abs/2609.02964) tested a two-stage filter on five AI systems against seven manipulative rewriting attacks. It reduced the average attack success from 50.32% to 6.20%. These are lab defenses; no study we found confirms what ChatGPT or Google deploy today. We weigh the odds in [whether AI search can filter manipulative GEO](https://underneath.agency/resources/can-ai-search-filter-manipulative-geo).

Live engines do not screen out every promotional format today. In [our study of self-ranking “best of” lists](https://underneath.agency/research/self-promoting-best-lists-study), 24.2% of cited numbered lists with an identifiable publisher ranked that publisher first. Being cited now is no guarantee for later, since the lab defenses above target exactly this kind of self-interest.

## Can honest content get caught by these defenses?

Yes. Defenses that block manipulation also remove some legitimate sources, so honest brands carry part of the cost.

The research defense above kept 94.12% of genuine evidence in use, which means a slice of honest material was dropped. [Luo and colleagues](https://arxiv.org/abs/2606.13610) tested 12 AI systems on fake products planted in web pages. Two consensus filters, which keep only brands that other evidence backs up, caught the fake in at least 90% of cases. But they also discarded roughly two thirds of the real brands people would have been recommended.

The [Martinez survey](https://arxiv.org/abs/2607.14035) warns that an overly strict anti-promotion filter “may penalize small publishers that legitimately describe their products.” If your content reads like an advertisement, it may sit closer to the line than you think.

## Can aggressive optimization hurt how AI describes your brand?

Yes. Manipulative rewriting made AI answers worse in tests, and weakly sourced pages are common among optimized content.

[Wu and colleagues](https://arxiv.org/abs/2510.11438) compared careful rewriting with “hijack” and “poisoning” attacks. The attacks raised visibility but always lowered answer quality and reliability. That is a poor trade for a brand whose name appears next to a misleading answer.

[Chu and colleagues](https://arxiv.org/abs/2608.16824) scanned 10,095 pages returned by Google Search and Gemini for 1,000 real questions. On pages their detector judged to be optimized for AI, 69.34% of the sources those pages cited were rated low on how easily they could be checked. Separately, [Khodayari](https://arxiv.org/abs/2604.27202) found hidden instructions to AI systems on live websites. AI systems obeyed them at most 8% of the time in tests; we cover the reputational risk in [websites hiding instructions for AI search](https://underneath.agency/resources/websites-hiding-instructions-for-ai-search).

## Does optimizing for AI search necessarily make answers less diverse?

No study shows that it does. The evidence points to gains canceling out, and to defenses as a separate risk to variety.

Visibility in an answer is shared, so one page’s gain is another’s loss. Puerto’s team found that the advantage of a rewriting method shrank as more sites adopted it, “eventually converging toward zero at full adoption.” That suggests a crowded race, not fewer sources.

There are two real pressures toward sameness. First, in Bagga’s shopping test, automated rewriting converged on one shared style; the authors list as an open question whether product descriptions would converge if everyone used it. Second, strict defenses can thin out legitimate choices, as Luo’s filters did. For now, Google’s AI Overviews draw on a wide range of sites. In a US study by [Xu and colleagues](https://arxiv.org/abs/2605.14021), 56.2% of the sites they cited appeared only once in 40 days. Whether widespread optimization narrows that range has not been measured.

## What should you do about it?

Treat GEO as a product-quality decision with brand risk, not as a bag of tricks.

1. Ask any team or vendor for evidence that a tactic raised citations in a test with a comparison group, not a single before-and-after.
2. Drop hype. Superlatives, invented awards and emotional pressure were the patterns AI rankers flagged most readily.
3. Add substance that a reader can check: real figures with sources, clear product facts and plain answers. Whether polishing the prose alone helps is covered in [readable writing and AI visibility](https://underneath.agency/resources/does-readable-writing-help-ai-visibility).
4. Never hide text or instructions for AI systems, and audit your site for any that agencies or plugins added.
5. Keep investing in being findable in ordinary search. Puerto’s team found that a page’s position in the results handed to the AI mattered more than any rewrite.
6. Measure before and after any rewrite, on the same questions, across more than one AI assistant.

For how to tell legitimate work from manipulation, see [what separates legitimate GEO from manipulation](https://underneath.agency/resources/legitimate-geo-vs-manipulation). If you want help building that kind of program, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

Most of this evidence comes from lab tests, not from watching real brands over months.

- No study we found tracked a real brand that was demoted by ChatGPT, Gemini or Google for optimized content.
- The defenses tested are research prototypes; what engines deploy is not public.
- Diversity under widespread optimization has been modeled and discussed, not measured on live engines.
- The shopping and benchmark tests used fixed sets of pages, which can differ from the open web.
- Results varied by AI model, so a finding on one engine may not hold on another.

## Frequently asked questions

### Can GEO hurt my Google rankings?

The research has not measured that directly. The Martinez survey reports a simulated search test where rewriting page bodies for AI made pages less likely to reach the top results.

### Will ChatGPT penalize my content for being optimized?

There is no public evidence that it does today. In lab tests, a single warning in an AI ranker’s instructions was enough to flag superlative-heavy product copy 100% of the time.

### Is adding statistics to content risky?

Adding invented or decorative statistics can be. In C-SEO Bench, adding statistics lowered rank in 19 of 24 settings, so add numbers only when they inform the reader.

### If everyone optimizes for AI, does it stop working?

Gains from a shared trick shrink as more sites adopt it. Puerto’s team found the advantage of popular rewriting methods fell toward zero at full adoption.

## Sources

- Puerto and colleagues (2025), [C-SEO Bench: Does Conversational SEO Work?](https://arxiv.org/abs/2506.11097), arXiv:2506.11097.
- Bagga and colleagues (2025), [E-GEO: A Testbed for Generative Engine Optimization in E-Commerce](https://arxiv.org/abs/2511.20867), arXiv:2511.20867.
- Li and colleagues (2026), [When Optimization Becomes Manipulation: Defending Generative Search against Malicious Generative Engine Optimization](https://arxiv.org/abs/2609.02964), arXiv:2609.02964.
- Luo and colleagues (2026), [One Polluted Page Is Enough: Evaluating Web Content Pollution in LLM Recommenders](https://arxiv.org/abs/2606.13610), arXiv:2606.13610.
- Wu and colleagues (2025), [What Generative Search Engines Like and How to Optimize Web Content Cooperatively](https://arxiv.org/abs/2510.11438), arXiv:2510.11438.
- Martinez (2026), [Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026)](https://arxiv.org/abs/2607.14035), arXiv:2607.14035.
- Chu and colleagues (2026), [GEO-Flag: Detecting and Measuring GEO-Optimized Web Content](https://arxiv.org/abs/2608.16824), arXiv:2608.16824.
- Khodayari (2026), [Indirect Prompt Injection in the Wild: An Empirical Study of Prevalence, Techniques, and Objectives](https://arxiv.org/abs/2604.27202), arXiv:2604.27202.
- Xu and colleagues (2026), [Measuring Google AI Overviews: Activation, Source Quality, Claim Fidelity, and Publisher Impact](https://arxiv.org/abs/2605.14021), arXiv:2605.14021.
- Underneath (2026), [How many “best of” lists cited by AI rank their own brand first?](https://underneath.agency/research/self-promoting-best-lists-study)

---

This is the Markdown twin of https://underneath.agency/resources/can-geo-backfire-on-your-brand. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Can GEO be used to push false information into AI answers?"
description: "Yes. In tests, one planted page made AI systems recommend a fake product in up to 27% of cases, and fabricated evidence was treated as real."
canonical: "https://underneath.agency/resources/can-geo-push-false-information-into-ai-answers"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Can GEO be used to push false information into AI answers?

Yes: researchers have shown that a few planted pages can make AI systems recommend products that do not exist, and that invented evidence is often taken at face value. GEO itself is not dishonest, but the same techniques that make good content easy to cite can make weak or false claims look well supported. For brands, the risk runs both ways: false claims about you, and false claims that crowd out your accurate ones.

## The short version

1. In a test of 12 AI systems on 225 real products, one fake-review page at the top of the search results led to a fake product being recommended in up to 27% of cases; with the top three results polluted, up to 73.8% ([Luo and Chen](https://arxiv.org/abs/2606.13610)).
2. When fooled, AI systems added invented social proof, such as claims of community popularity, 1.5 to 11 times more often than when they resisted (Luo and Chen).
3. Fabricated clinical-trial claims beat a famous skincare brand in 73.3% of head-to-head tests across three commercial AI systems ([Chu and Hou](https://arxiv.org/abs/2606.17443)).
4. An audit of real Google and Gemini results estimated that 8.90% of pages showed signs of GEO, and 69.34% of the citations on those pages pointed to sources that were hard to verify ([Chu and colleagues](https://arxiv.org/abs/2608.16824)).
5. AI answers repeat their sources’ weaknesses: in [our Reddit study](https://underneath.agency/research/ai-reddit-citations-study), 20.8% of sentences citing a Reddit thread were not supported by it.

## How could GEO make a false claim look well supported?

By building a chain of pages that appear to confirm each other. [Chu and colleagues](https://arxiv.org/abs/2608.16824) describe the basic move. An operator publishes a false claim on a site that is easy to create or edit, then cites that page from another site.

The link works and the cited page does support the claim. But, as they put it, “the source itself was strategically created to support it.” An AI answer that summarizes both pages can present the claim as backed by evidence.

Generative search makes this easier to miss. A traditional results page shows competing sources side by side.

An AI answer blends them into one response, and checking where a claim came from takes extra clicks. The authors warn GEO “can be misused to make weak or false information appear well supported,” while stressing that GEO is “not inherently malicious.”

## Has anyone shown it working on real AI systems?

Yes, in a controlled test built on real search results, and in at least one public exposé. On March 15, 2026, China Central Television reported a black market of GEO operators. According to [Luo and Chen](https://arxiv.org/abs/2606.13610), they seeded fake reviews so a fake brand appeared in mainstream Chinese AI assistants’ recommendations “within hours.”

Luo and Chen then measured the effect without polluting the real web. They took real search results for 225 products in 15 categories and swapped the real brand name in some pages for an invented one. Twelve commercial and open AI systems then made recommendations.

| Pages polluted | Fake product recommended |
|---|---|
| One page, in the top search slot | Up to 27% of cases for the most exposed systems |
| One page, in slots two to ten | Nearly inert |
| Top three pages | 13.3% to 73.8%, depending on the system |

Placement mattered most. Once fooled, the AI put the fake brand in first place 57% of the time. And the top slot is often the easiest to fill. In their data, a user-posted page such as a forum or Q&A thread held it in over half of all searches.

Every system tested was vulnerable. Larger and commercial systems were not reliably safer.

Everyday categories such as dining, personal services and supplements were the most exposed. Phones and home appliances, where AI systems know the real brands well, were the least.

## Do AI systems add their own false support?

Often, and that makes the false claim more convincing. In Luo and Chen’s test, fooled answers dressed the fake brand up rather than just naming it. They used social-proof phrases such as “frequently recommended” in technical communities, which did not appear in the planted pages.

In their analysis of the open systems, fooled answers used such phrases 1.5 to 11 times more often than answers that resisted. The authors call this a form of confabulation: confident detail with nothing behind it.

Invented evidence also works as input. [Chu and Hou](https://arxiv.org/abs/2606.17443) found three commercial AI systems largely ignored pushy sales copy.

Yet they treated a fabricated clinical citation “as if it were real evidence.” Authority claims of that kind beat a famous brand in 73.3% of tests. Our guide to [gaming AI shopping rankings with product copy](https://underneath.agency/resources/can-you-game-ai-shopping-rankings) covers this test in more detail.

## How common are optimized pages with weak sourcing?

Roughly one page in eleven showed GEO signals in one real-world audit, with weak sourcing common on those pages. [Chu and colleagues](https://arxiv.org/abs/2608.16824) ran a detector on pages returned by Google Search and by Gemini with Google grounding. They covered 1,000 real user questions and 10,095 pages.

The detector flagged 8.90% of pages. Among pages that declared a 2026 modification date, the share was 16.36%. Across the citations on flagged pages, 69.34% got a low verifiability rating, meaning the cited source had little editorial accountability or could not be reached.

The two channels differed. Low-verifiability citations made up 74.15% on flagged pages surfaced by Gemini, against 45.88% for Google Search. Platforms differed too: no flagged pages among 613 Wikipedia pages, but 20.37% of Amazon pages.

These are estimates. The detector can make mistakes, and only 19.57% of pages declared a modification date at all. A low verifiability rating means weak accountability, not that a claim is false. We look at the wider audit in [how much web content is optimized for AI](https://underneath.agency/resources/how-much-web-content-is-optimized-for-ai-search).

## Why does this matter for your brand?

Because AI answers pass on what their sources say, including errors. Our own studies show how loosely answers can track their evidence:

- In [our Reddit study](https://underneath.agency/research/ai-reddit-citations-study), 20.8% of sentences citing a Reddit thread were not supported by the post or its top comments.
- In [our reputation study](https://underneath.agency/research/is-it-legit-ai-reputation-study), 71.4% of claims that a problem was common rested on evidence that did not show how common it was.
- In [our pricing study](https://underneath.agency/research/ai-pricing-accuracy-study), most prices that differed from a vendor’s pricing page were not invented: for 39 of 64, the same figure appeared elsewhere on the vendor’s own site.

The pattern is consistent. AI answers tend to repeat what is on the pages they find.

If those pages are wrong, or planted, the answer can be too. A competitor, critic or fraudster can target your category with the same techniques. Some cited pages are machine-written too, as [AI-generated pages in AI citations](https://underneath.agency/resources/are-ai-search-engines-citing-ai-content) shows.

## What should you do about it?

Make accurate information about you easy to find and confirm, and watch for false claims early.

1. **Monitor AI answers about your brand and category.** Ask the main assistants the questions buyers ask, regularly, and note claims you cannot trace to a real source.
2. **Publish checkable facts on your own site.** Clear, dated pages on pricing, specifications and policies give AI systems an accurate source to cite. Remove stale figures that could be quoted as current.
3. **Make sure credible third parties have it right.** In one test, ranking editorial sources above open forums reduced fake recommendations, and AI systems resisted fakes best where they already knew the real brands.
4. **Watch open forums and review sites.** Planted content worked best in the top search slot, which user-posted pages often held.
5. **Never fabricate evidence yourself.** Invented studies or reviews are what researchers flag as manipulation, and the tactic has already drawn national-television exposure and regulatory action in China.

For help strengthening accurate AI visibility, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

How often false information actually spreads through live AI answers, and who is doing it. Key gaps:

- The fake-product test used frozen search results, mainly in Chinese, from April 2026, and rewrote pages locally rather than on the live web.
- The prevalence audit estimates pages with GEO signals, not intent, and does not check whether claims on those pages are false.
- No study has tracked a real disinformation campaign from planted pages to AI answers over time.
- No published study measures how quickly live engines remove false claims once they are reported.
- Most of this research is in preprints that have not yet been repeated by other teams.

## Frequently asked questions

### Can someone make ChatGPT say false things about my company?

Research shows it is possible in principle. In a controlled test of 12 AI systems, one planted page in the top search slot led to a fake product being recommended in up to 27% of cases.

### Is GEO the same as spreading misinformation?

No. Researchers stress that GEO “is not inherently malicious,” but the same techniques can make weak claims look well supported. One audit estimated 8.90% of pages in Google and Gemini results showed GEO signals.

### How do AI search engines decide whether a source is trustworthy?

Published evidence is limited. In one test, re-ordering search results to put editorial sources first and open forums last cut fake recommendations for all six open systems tested, but only partly.

### What should I do if an AI answer repeats a false claim about my brand?

Find the source it cites and correct the record there and on your own site. In our pricing study, most differing figures traced back to a real page, so fixing the source is usually the lever.

## Sources

- Luo and Chen (2026), [One Polluted Page Is Enough: Evaluating Web Content Pollution in LLM Recommenders](https://arxiv.org/abs/2606.13610), arXiv:2606.13610.
- Chu, Leng, Li, Shen, Shen and Zhang (2026), [GEO-Flag: Detecting and Measuring GEO-Optimized Web Content](https://arxiv.org/abs/2608.16824), arXiv:2608.16824.
- Chu and Hou (2026), [Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems](https://arxiv.org/abs/2606.17443), arXiv:2606.17443.
- Underneath (2026), [When does a Reddit thread become evidence in Google’s AI?](https://underneath.agency/research/ai-reddit-citations-study)
- Underneath (2026), [“Is this brand legit?” How AI assistants build a reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study)
- Underneath (2026), [How faithfully do AI assistants quote software prices?](https://underneath.agency/research/ai-pricing-accuracy-study)

---

This is the Markdown twin of https://underneath.agency/resources/can-geo-push-false-information-into-ai-answers. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Can smaller websites get cited by AI search engines?"
description: "Yes. AI engines cite a long tail of little-known sites, but being cited is not being recommended, and well-known brands still win most answers."
canonical: "https://underneath.agency/resources/can-small-websites-get-cited-by-ai"
published: 2026-10-08
updated: 2026-10-10
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Can smaller, lower-traffic websites get cited by AI search engines?

Yes. Most AI engines cite a long tail of little-known sites, and some lean toward less popular domains than Google does. But being cited is not the same as being recommended. Well-known brands still appear in far more AI answers than small ones.

## The short version

1. Copilot and Gemini cited domains with traffic rankings over 22,000 places lower on average than Google and Bing, in [a 2025 study of six AI engines](https://arxiv.org/abs/2512.09483); ChatGPT was the exception.
2. In Google’s AI Overviews, 56.2% of cited sites were cited only once in 40 days, [Xu and colleagues found](https://arxiv.org/abs/2605.14021).
3. The YouTube videos AI Overviews cited from outside page one had a median of 2,776 views, in [our YouTube study](https://underneath.agency/research/ai-overview-youtube-videos-study).
4. In a vendor dataset, niche brands appeared in just 11% of relevant AI answers, against 73% for household names, [Kumar reports](https://arxiv.org/abs/2606.20065).

## Can smaller websites get cited by AI search engines?

Yes, and often. Many cited sites receive only one or two citations. [Zhang and colleagues](https://arxiv.org/abs/2512.09483) compared six AI engines with Google and Bing on 55,936 queries in July and August 2025. Except for ChatGPT, the AI engines tended to cite less popular domains than traditional search.

The gap was largest for two engines. Copilot and Gemini cited domains whose average traffic rank was over 22,000 places lower than Google’s and Bing’s. The authors also found that 37% of all cited domains appeared only in AI answers, never in traditional results.

The traffic ranking used here is a research list of website popularity. It measures whole domains, not individual pages.

## Do AI engines favor big, popular websites?

The evidence is mixed and depends on the engine. A [global study by Aral and colleagues](https://arxiv.org/abs/2602.13415) ran 24,000 search queries in 243 countries in 2024 and 2025. It found Google’s AI Overviews linked to the long tail of the web significantly less than traditional search, and to the top 1,000 sites more.

[Grossman and colleagues](https://arxiv.org/abs/2604.27790) found the opposite pattern for the first source shown. In traditional Google results, the top source came from one of the 1,000 most popular sites for 52.7% of queries. For AI Overviews the figure was 40.0%, and for Gemini 32.6%.

These studies used different query sets, popularity lists and years. The honest summary is that no engine reliably shuts out smaller sites, and none reliably favors them. For the kinds of sites engines lean on, see [which websites AI search relies on most](https://underneath.agency/resources/which-websites-do-ai-search-engines-cite).

## How spread out are AI citations?

Very spread out, with a short list of repeat winners. In the Xu study of Google AI Overviews in spring 2026, 56.2% of cited sites were cited exactly once over 40 days. On Google’s first page of results the share was 52.5%.

The top of the list is less dominant in AI Overviews than on Google’s results page. The five most-cited sites took 20.0% of AI Overview citations, against 47.0% of first-page links.

News is the exception. [Yang](https://arxiv.org/abs/2507.05301) found that for OpenAI models in 2025, the 20 most-cited news outlets took 67.3% of citations in the news sample. A small publisher competing on news faces a far more concentrated field.

## Is being cited the same as being recommended?

No. A small site can be cited as evidence while its brand is left out of the answer. [Kumar](https://arxiv.org/abs/2606.20065), whose company sells AI visibility tracking, measured more than 100 brands between March and May 2026. Household names appeared in 73% of relevant unbranded answers on the first run, mid-market brands in 44% and niche brands in 11%.

Other studies show the same pull toward big names. [Chen and colleagues](https://arxiv.org/abs/2509.08919) asked ChatGPT and Perplexity 50 unbranded cola questions in 2025. Niche brands were 12.3% of ChatGPT’s brand mentions and 5.8% of Perplexity’s.

New companies face the steepest climb. [Sharma](https://arxiv.org/abs/2601.00912) tested 112 Product Hunt startups. A version of ChatGPT without web search recognized them 99.4% of the time when asked by name. It surfaced them in only 3.32% of discovery questions such as “What are the best AI tools launched this year?”.

## What helps a smaller site get cited?

Depth, relevance and visible proof, more than size. [Zhu and Chang](https://arxiv.org/abs/2603.20062) compared 14 Tokyo hotel websites against Gemini’s citations. One small independent [hotel with no schema markup](https://underneath.agency/resources/deep-content-without-technical-seo-gemini) won direct citations through a 33-question FAQ and a detailed sightseeing guide. A polished boutique with shallow content did not. The audit is small and exploratory.

Our own data shows small creators getting in. In [our YouTube study](https://underneath.agency/research/ai-overview-youtube-videos-study), videos cited from outside page one had a median of 2,776 views, against 34,100 for cited videos Google also showed. Within the same search, a bigger channel lowered the chance of citation.

For local businesses, reviews matter. In [our ChatGPT local picks study](https://underneath.agency/research/chatgpt-local-picks-google-profile-study), businesses with more reviews than the local median were 19.5 points more likely to be listed.

Small advantages can matter in tests too. In [a simulated skincare study by Chu and Hou](https://arxiv.org/abs/2606.17443) using three AI models, well-known brands won 100% of recommendations when products were otherwise identical. That lead disappeared when a rival had a rating advantage of less than 0.1 stars.

## Can optimization close the gap with bigger sites?

Possibly, in simulations. The original GEO study by [Aggarwal and colleagues](https://arxiv.org/abs/2311.09735) tested content rewrites on a simulated engine. Adding citations to sources raised visibility by 115.1% for websites ranked fifth in search, while the top-ranked website lost 30.3% on average.

Live evidence is weaker. In Sharma’s startup study, scores for AI-oriented page optimization showed no link to discovery. For Perplexity, [traditional signals such as links from other sites](https://underneath.agency/resources/is-seo-still-important-for-ai-search) did.

## What should you do about it?

Compete where size matters least: specific questions, deep answers and proof from others.

1. Target narrow, specific questions where large sites have thin coverage.
2. Answer them in depth on your own pages, with concrete facts, FAQs and guides.
3. Publish short, useful videos. Google’s AI cites small channels.
4. If you serve a local market, build review volume on your Google profile.
5. Earn links and mentions from other sites, which help discovery on engines that search the web.
6. Measure brand mentions, not only citations, because the two diverge for smaller brands.

Our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) works on this kind of focused visibility.

## What does the research not tell us yet?

No study we reviewed followed small sites over time to see what moved them from cited to recommended.

- Studies disagree on whether AI engines favor popular sites, and they use different popularity lists.
- The brand-size figures come from a vendor dataset and small tests, not independent large samples.
- The optimization gains come from simulated engines, not live ones.
- Several samples are narrow: Product Hunt startups, Tokyo hotels or cola brands.
- In [our brand entity study](https://underneath.agency/research/brand-entity-ai-recommendations-study), 21 of 110 options all four assistants named had no Wikipedia article, and 11 were local businesses. Why such businesses break through is still unexplained.

## Frequently asked questions

### Do AI search engines only cite big websites?

No. In one study of AI Overviews, 56.2% of cited sites were cited only once in 40 days, which points to a long tail of smaller sources.

### Does ChatGPT prefer popular websites?

More than most AI engines. In a 2025 study of six engines, ChatGPT was the only one that did not cite less popular domains than Google and Bing.

### Can a new startup show up in ChatGPT answers?

Rarely at first. One study found a version of ChatGPT without web search surfaced Product Hunt startups in 3.32% of discovery questions.

### Is YouTube a way for small brands to get into AI Overviews?

It can be. In our study, cited videos from outside page one had a median of 2,776 views, far fewer than the videos Google ranked.

## Sources

- Zhang et al. (2025), [Source Coverage and Citation Bias in LLM-based vs. Traditional Search Engines](https://arxiv.org/abs/2512.09483), arXiv:2512.09483.
- Aral et al. (2026), [The Rise of AI Search: Implications for Information Markets and Human Judgement at Scale](https://arxiv.org/abs/2602.13415), arXiv:2602.13415.
- Grossman et al. (2026), [How Generative AI Disrupts Search: An Empirical Study of Google Search, Gemini, and AI Overviews](https://arxiv.org/abs/2604.27790), arXiv:2604.27790.
- Xu et al. (2026), [Measuring Google AI Overviews: Activation, Source Quality, Claim Fidelity, and Publisher Impact](https://arxiv.org/abs/2605.14021), arXiv:2605.14021.
- Yang (2025), [News Source Citing Patterns in AI Search Systems](https://arxiv.org/abs/2507.05301), arXiv:2507.05301.
- Kumar (2026), [Generative Engine Optimization at Scale: Measuring Brand Visibility Across AI Search Engines](https://arxiv.org/abs/2606.20065), arXiv:2606.20065.
- Chen et al. (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919), arXiv:2509.08919.
- Sharma (2026), [The Discovery Gap: How Product Hunt Startups Vanish in LLM Organic Discovery Queries](https://arxiv.org/abs/2601.00912), arXiv:2601.00912.
- Zhu and Chang (2026), [The End of Rented Discovery: How AI Search Redistributes Power Between Hotels and Intermediaries](https://arxiv.org/abs/2603.20062), arXiv:2603.20062.
- Chu and Hou (2026), [Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems](https://arxiv.org/abs/2606.17443), arXiv:2606.17443.
- Aggarwal et al. (2024), [GEO: Generative Engine Optimization](https://arxiv.org/abs/2311.09735), arXiv:2311.09735.
- Underneath (2026), [The YouTube videos Google’s AI cites are small](https://underneath.agency/research/ai-overview-youtube-videos-study)
- Underneath (2026), [Which Google Maps businesses does ChatGPT recommend?](https://underneath.agency/research/chatgpt-local-picks-google-profile-study)
- Underneath (2026), [Do Wikipedia and schema make AI assistants recommend a brand?](https://underneath.agency/research/brand-entity-ai-recommendations-study)

---

This is the Markdown twin of https://underneath.agency/resources/can-small-websites-get-cited-by-ai. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Can sellers game AI shopping rankings with manipulative copy?"
description: "Sometimes, on weaker AI systems. In tests, GPT-5 and Claude flagged and demoted manipulative product copy, and any gain vanished once rivals copied it."
canonical: "https://underneath.agency/resources/can-you-game-ai-shopping-rankings"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Can sellers game AI shopping rankings with manipulative product copy?

Sometimes, but mostly on weaker AI systems, and the gains are fragile. In a test of 14 manipulation tactics across five AI shopping rankers, GPT-5 and Claude flagged almost all of them and pushed the products down the list. Fabricated evidence did move some systems, which is exactly why it is a legal and reputational risk rather than a strategy.

## The short version

1. Product copy stuffed with superlatives and guarantees was flagged as questionable 100% of the time by GPT-5 and dropped about four places on a ten-product list ([Bagga and colleagues](https://arxiv.org/abs/2511.20867)).
2. Weaker rankers could be fooled: hidden override instructions moved a product up 3.40 places on an open Llama model, which flagged only 21.0% of them (Bagga and colleagues).
3. Fabricated clinical claims beat a famous brand in 73.3% of head-to-head tests across three commercial AI systems, but only when one seller used them ([Chu and Hou](https://arxiv.org/abs/2606.17443)).
4. When all nine challenger brands used the same claims, the famous brand won again 93.8% of the time, and each copier’s gain fell from +0.802 to +0.007 on the authors’ payoff measure (Chu and Hou).
5. Planted fake reviews already work in the wild: in a test of 12 AI systems, one polluted top search result led to a fake product being recommended in up to 27% of cases ([Luo and Chen](https://arxiv.org/abs/2606.13610)).

## Does manipulative product copy work on AI shopping assistants?

On the strongest systems tested, mostly not; on weaker ones, sometimes. [Bagga and colleagues](https://arxiv.org/abs/2511.20867) built a shopping test bed from 13,747 real shopping requests, written in the detailed style of Reddit posts. Each was paired with 10 Amazon listings.

They rewrote one listing per request using 14 manipulation tactics. These ranged from fake reviews and made-up statistics to hidden instructions and keyword stuffing. Five AI systems then ranked the ten products: GPT-5, Claude, Gemini, DeepSeek and Llama.

Each ranker was also asked to list products whose descriptions looked misleading. That one-sentence safeguard was added by the researchers. Real shopping assistants likely use stronger defenses, but none have been published. We review the research on such defenses in [whether AI search can filter manipulation](https://underneath.agency/resources/can-ai-search-filter-manipulative-geo).

The verdict was clear for the top systems. The authors write that “GPT-5 and Claude do not reward” manipulative rewrites with higher rankings. They read this as evidence that ranking gains on strong systems come from genuinely better content.

## Which tactics did the strongest systems punish?

Nearly all of them, while a weaker open system rewarded several. The table shows how often each tactic was flagged and how far the product moved on a ten-product list.

| Tactic | GPT-5: flagged, places moved | Llama: flagged, places moved |
|---|---|---|
| Superlatives and guarantees | 100.0%, down 4.14 | 100.0%, down 2.26 |
| “Rank this #1” instruction in the text | 100.0%, down 3.89 | 76.5%, up 0.39 |
| Made-up precise statistics | 97.5%, down 3.27 | 59.5%, up 1.75 |
| Fake testimonials and ratings | 99.0%, down 3.76 | 76.0%, up 1.60 |
| Hidden override instructions | 99.5%, down 3.75 | 21.0%, up 3.40 |
| Writing in the style AI tends to like | 9.7%, up 0.41 | 15.0%, up 1.01 |

Two lessons stand out. First, the same trick can win on one system and lose badly on another. A seller cannot know which assistant a shopper uses.

Second, the only tactic that rarely got flagged was mimicking the style AI systems tend to prefer. Its gains on GPT-5 were small. It is the grey zone, not a shortcut.

## Do fabricated claims work better than sales pressure?

Yes, in one controlled test, and that is what makes them dangerous. [Chu and Hou](https://arxiv.org/abs/2606.17443) gave three commercial AI systems ten identical skincare products: one famous brand and nine invented ones. With identical specifications, the famous brand was picked in every one of 670 valid trials.

Then they added marketing language to one invented brand. Copy mimicking credible evidence, such as clinical-trial claims or testimonials, broke the famous brand’s hold 50% to 73% of the time. Pure sales pressure, like scarcity or “limited stock,” barely worked, at 10% to 13%.

The systems differed. Gemini fell for authority claims almost every time. Claude was more guarded: when authority and testimonials were stacked together, its rate dropped from 55.0% to 21.2%, as if the offer looked too good to be true.

The authors used invented clinical studies on purpose, to find the upper limit. They also label that tier of content plainly: inventing studies or endorsements “is potential false advertising.”

An older experiment shows how far technical tricks can go on an open system. [Kumar and Lakkaraju](https://arxiv.org/abs/2404.07981) added a machine-generated string of text to a fictional $199 coffee machine.

On Llama-2, it lifted the product’s rank in about 40% of tests where product order was shuffled. Building it required access to the open system’s inner workings, which sellers do not have for commercial assistants.

## What happens when every seller games the system?

The advantage disappears, but sellers who stop become invisible. Chu and Hou ran a market test where more and more invented brands used the same authority claims. A single challenger knocked the famous brand down to winning just 19.8% of the time.

As rivals copied the tactic, the signal stopped standing out. With all nine challengers using it, the famous brand won 93.8% of the time again. The first mover’s gain of +0.802 on the authors’ payoff measure shrank to +0.007.

Worse, across 4,745 trials, invented brands that did not use the claims received zero recommendations. The authors call it a prisoner’s dilemma: everyone feels forced to join, and nobody ends up ahead. They suggest platforms may need to check authority claims, because each brand’s copy looks legitimate on its own. We look at the cost of standing still in [what happens to brands that skip GEO](https://underneath.agency/resources/what-happens-if-you-skip-geo).

## What are the risks beyond being flagged?

Legal exposure, public exposure and a shrinking payoff. Fake reviews are already a live business.

[Luo and Chen](https://arxiv.org/abs/2606.13610) describe a March 15, 2026 report by China Central Television. It exposed operators who seeded fake reviews so a fake brand appeared in leading Chinese AI assistants’ recommendations “within hours.”

The researchers then tested 12 AI systems on 225 real products. One fake-review page at the top of the search results led to the fake product being recommended in up to 27% of cases. With the top three results polluted, the worst system recommended it 73.8% of the time.

That shows the tactic can work. It also shows why it is risky: it was exposed on national television, and the authors note Chinese regulators launched enforcement in response. Our guide to [pushing false information into AI answers](https://underneath.agency/resources/can-geo-push-false-information-into-ai-answers) covers the wider risk.

Reputation is also judged by AI from reviews. In [our AI reputation study](https://underneath.agency/research/is-it-legit-ai-reputation-study), 88.0% of answers about whether a brand was legitimate cited a review or complaint platform. Claims resting only on those platforms were negative 56.5% of the time.

## What should you do about it?

Compete on real, specific product information, and check that nothing on your listings could read as fabricated.

1. **Lead with verifiable facts.** In Chu and Hou’s test, a small real edge, such as a 7.3% lower price, was enough for an unknown brand to win half the time.
2. **Cite real evidence only.** Genuine certifications and published trials are legitimate. Invented studies or endorsements are the tier the researchers call potential false advertising.
3. **Strip hype that adds nothing.** Superlatives and guarantees were flagged by every ranker in the shopping test.
4. **Audit agency and marketplace copy.** Ask anyone writing listings for you to confirm they do not add hidden text, fake reviews or invented figures.
5. **Test across assistants.** A gain on one system can be a penalty on another.

For help improving product content that holds up across AI assistants, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

How live shopping assistants such as ChatGPT shopping or Amazon’s assistant respond to these tactics. The limits are real:

- The shopping test used simulated rankers with a researcher-added safeguard over a fixed list of ten Amazon products, in English.
- The fabricated-claims test used skincare products and invented brands; the authors did not repeat the marketing experiments in other categories.
- The fake-review test used frozen search results, mostly in Chinese, from April 2026.
- No study tracks real sellers over time to see whether manipulation gains last or trigger penalties.

## Frequently asked questions

### Can fake reviews get my product recommended by AI?

In tests, sometimes, which is why it is a serious risk. One fake-review page at the top of search results led AI systems to recommend a fake product in up to 27% of cases. The practice was exposed by Chinese state television.

### Do AI shopping assistants penalize hype?

The strongest ones tested did. Copy built on superlatives and guarantees was flagged 100% of the time by GPT-5 and dropped about four places on a ten-product list.

### Does mentioning clinical studies help AI rank my product?

Real evidence is a legitimate signal; invented evidence is not. Fabricated clinical claims beat a famous brand in 73.3% of tests, but the researchers call invented studies potential false advertising.

### If competitors game AI rankings, should we copy them?

The research suggests not. When every challenger used the same claims, the famous brand won again 93.8% of the time and each copier’s gain nearly vanished.

## Sources

- Bagga, Farias, Korkotashvili, Peng and Wu (2025), [E-GEO: A Testbed for Generative Engine Optimization in E-Commerce](https://arxiv.org/abs/2511.20867), arXiv:2511.20867.
- Chu and Hou (2026), [Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems](https://arxiv.org/abs/2606.17443), arXiv:2606.17443.
- Kumar and Lakkaraju (2024), [Manipulating Large Language Models to Increase Product Visibility](https://arxiv.org/abs/2404.07981), arXiv:2404.07981.
- Luo and Chen (2026), [One Polluted Page Is Enough: Evaluating Web Content Pollution in LLM Recommenders](https://arxiv.org/abs/2606.13610), arXiv:2606.13610.
- Underneath (2026), [“Is this brand legit?” How AI assistants build a reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study)

---

This is the Markdown twin of https://underneath.agency/resources/can-you-game-ai-shopping-rankings. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "What do commercial health sites cited by ChatGPT have in common?"
description: "In a 615-source ChatGPT audit, cited commercial health sites mostly showed medical review (71.1%), schema (86.8%) and pages over 1,500 words (68.4%)."
canonical: "https://underneath.agency/resources/chatgpt-cited-commercial-health-sites"
published: 2026-10-11
updated: 2026-10-11
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# What do commercial health sites cited by ChatGPT have in common?

They stack visible trust signals that hospitals and government sites skip: a stated medical review, schema markup and long, comprehensive pages. In a January 2026 audit of ChatGPT’s health answers, commercial publishers showed all three far more often than institutional sources did. The audit describes cited pages only, so it shows a pattern, not a recipe that guarantees citation.

## The short version

1. In an audit of 615 cited sources by Jacques and colleagues, ChatGPT drew 75.7% of its health citations from institutions such as hospitals, government agencies and Wikipedia.
2. Commercial health publishers earned 12.4% of citations, and they stated a medical review on 71.1% of their cited pages.
3. Of those commercial pages, 86.8% carried schema markup, the machine-readable labels that describe a page to software, and 68.4% ran past 1,500 words.
4. Government health pages, by contrast, stated a medical review only 13.4% of the time, and only 31.1% of them were that long.

## Which kinds of health sources does ChatGPT cite?

Mostly institutions with built-in authority; commercial publishers take about one citation in eight.

[Jacques and colleagues](https://arxiv.org/abs/2601.17109) took a random sample of 100 questions from HealthSearchQA, a set of 3,173 consumer health questions that Google Research built from real search suggestions. One researcher entered each question into ChatGPT 5.2 Pro, in a fresh account and a separate chat. Every question asked ChatGPT to include sources with links. All answers were collected on January 11, 2026.

The team then coded every cited page. Of 615 usable sources, 75.7% came from organizations with inherent institutional authority. Medical institutions led, followed by government resources, Wikipedia, professional associations and journals.

The rest split almost evenly. Commercial health information platforms took 12.4% and small professional or practice websites took 11.9%. The commercial group was led by names such as Healthline, WebMD and Medical News Today, with 76 citations spread over 28 organizations.

## What do the cited commercial sites have in common?

Three signals appear again and again: a stated medical review, schema markup and long, comprehensive content.

The authors describe these as “compensatory credibility signals”. A commercial publisher has no hospital or government name behind it, so its pages show their vetting openly instead.

| Signal on cited pages | Commercial health platforms | Government resources | Medical institutions |
|---|---|---|---|
| States a medical review | 71.1% | 13.4% | 49.5% |
| Comprehensive content (over 1,500 words) | 68.4% | 31.1% | 31.9% |

The study reports that cited commercial platforms implemented schema markup in 86.8% of cases. Government sources and journals sat below the sample’s typical rate on schema, at 61.3% and 42.9%.

One signal the commercial group did not lead on was freshness. Only 18.4% of cited commercial pages showed a date from 2024 to 2026. Small practice websites took a different route: 61.6% of their cited pages carried a recent date.

## How do commercial sites compare with the whole cited set?

They beat the cited set as a whole on review, schema and length, but not on freshness.

Across all 615 cited sources, 70.4% did not state any medical review at all. Schema markup appeared on 74.3% of all cited pages, below the commercial rate. Comprehensive length was common too, but the commercial group still ran above the overall share.

The cited set was also strong on classic web authority. Sources had a median Domain Authority of 89, a 0 to 100 score from the SEO tool Moz that estimates the strength of a site’s links. Commercial platforms that were cited had a median score of 73, well above the small practice sites.

In other words, the commercial publishers that ChatGPT cited were large, established sites that also invested in visible quality markers. A new site that copies the markers would not copy the years of links and reputation behind them.

## Does this mean those signals earn the citation?

No. The audit only looked at pages ChatGPT cited, so it cannot say what the signals cause.

The study has no comparison group of commercial health pages that ChatGPT passed over. If uncited commercial pages also state a medical review and use schema, the signals would not explain anything. The authors present their findings as a baseline to monitor over time, not as a ranking formula.

Two more cautions matter. The coding was largely automated: pages were scraped and labeled by an AI model working from the researchers’ codebook, with missing fields checked by hand. And the split between commercial publishers and small practices partly used a Domain Authority cutoff of 45 when the type was unclear.

## What do other studies say about schema, length and review?

They suggest schema alone does little, while substance and evidence on the page matter more.

Our [study of pages cited by AI Overviews](https://underneath.agency/research/ai-overview-cited-pages-study) took 486 US searches and compared cited pages with the uncited pages ranking beside them. Once pages on the same search were compared, no schema type held up as a reliable predictor of citation. Position in Google came first: 41.7% of pages ranked 1 to 3 were cited. Our study also notes a controlled test by Ahrefs in which adding schema did not raise citations. Our guide to [on-page signals linked to AI citations](https://underneath.agency/resources/on-page-signals-linked-to-ai-citations) covers the rest of that evidence.

Length looks more promising, but only with structure. [Zhang and colleagues](https://arxiv.org/abs/2604.25707), three independent researchers, analyzed ChatGPT, Google and Perplexity citations. On their measure of how much a cited page shaped the answer, the average rose with length, from very short pages to pages longer than 3,000 words. The pages that did best were also better organized and closer to the question, so length worked as a container for evidence, not as a target. Whether reorganizing a page alone helps is covered in [content structure and AI citations](https://underneath.agency/resources/does-content-structure-increase-ai-citations).

Finally, evidence beats tone. In a 252,000-trial controlled test by [Vishwakarma and colleagues](https://arxiv.org/abs/2605.25517) at Sprinklr, a software vendor, AI models preferred pages that backed claims with evidence. That is close to what a stated medical review promises a reader.

## What should you do about it?

If you publish health or other expert content without institutional backing, show your vetting and depth plainly.

1. State who reviewed each page, with their qualifications and the review date, where a qualified reviewer really did the work. A named author on its own is not required; see [whether cited pages need author bylines](https://underneath.agency/resources/do-ai-cited-pages-need-author-bylines).
2. Write pages that answer the full question a patient or buyer has, with sections, definitions, figures and references, rather than short, thin posts.
3. Add accurate schema markup as basic hygiene, but do not expect it to earn citations on its own.
4. Keep important pages dated and current; the cited commercial pages lagged here, and small practices did not.
5. Keep building search ranking and links, since the cited sources were overwhelmingly high-authority sites.
6. Track which of your pages ChatGPT actually cites for your core questions, and compare them with pages it skips.

If you want a structured way to run that tracking, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The audit is a careful first map, but it is one engine, one day and one subject.

- It covers ChatGPT 5.2 Pro on January 11, 2026, with a prompt that asked for sources. Other assistants and ordinary prompts may cite differently.
- It used 100 of 3,173 questions; the authors say a different sample might change the results.
- It describes cited pages only, with no uncited pages to compare against.
- Health is a special case. Results may not carry over to finance, software or local services.
- No study here changed a real page and measured its citations before and after.

## Frequently asked questions

### Does a medical review statement help a page get cited by ChatGPT?

It is common on cited commercial health pages, but no study has shown it causes citation. In the ChatGPT audit, 71.1% of cited commercial pages stated a medical review, against 13.4% of cited government pages.

### Does schema markup help health sites appear in AI answers?

There is no reliable evidence that schema alone does. Most cited commercial health pages used it, but our AI Overview study found no schema type that held up once pages on the same search were compared.

### How long should health content be to be cited by AI?

Length went with citation only when it carried real substance. Most cited commercial health pages ran past 1,500 words, and a separate study found the most-used pages were long and well structured, not merely long.

### Which health websites does ChatGPT cite most?

Institutions dominate. In the audit, Wikipedia, Mayo Clinic and Cleveland Clinic led, and 75.7% of citations went to institutional sources.

## Sources

- Jacques, Datuowei, Jones, Basch, Vanderpool, Udeozo and Chapa (2026), [Authority Signals in AI Cited Health Sources: A Framework for Evaluating Source Credibility in ChatGPT Responses](https://arxiv.org/abs/2601.17109), arXiv:2601.17109.
- Zhang, He and Yao (2026), [From Citation Selection to Citation Absorption: A Measurement Framework for Generative Engine Optimization Across AI Search Platforms](https://arxiv.org/abs/2604.25707), arXiv:2604.25707.
- Vishwakarma, Kumar and Jamidar (2026), [What Gets Cited: Competitive GEO in AI Answer Engines](https://arxiv.org/abs/2605.25517), arXiv:2605.25517.
- Underneath (2026), [What pages cited by AI Overviews have in common: 3,096 pages](https://underneath.agency/research/ai-overview-cited-pages-study)

---

This is the Markdown twin of https://underneath.agency/resources/chatgpt-cited-commercial-health-sites. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Will optimizing content for ChatGPT hurt our Google rankings?"
description: "Not if you protect pages already earning Google clicks. In one field test, a guarded rewrite lifted ChatGPT referrals with no Google loss beyond the site trend."
canonical: "https://underneath.agency/resources/chatgpt-optimization-google-rankings"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Will optimizing content for ChatGPT hurt our Google rankings?

Not necessarily: the one field test published so far found no measurable Google harm when pages already earning Google clicks were left alone. The evidence is thin, comes from a single website, and does not prove the two channels never conflict. What it does show is how to run the work so you would notice if they did.

## The short version

1. In a five-month field test on one website, Google clicks to rewritten pages fell about 25% while clicks across the whole site fell about 20%, a decline the authors read as the general trend ([Watanabe and Nakayashiki](https://arxiv.org/abs/2606.04362)).
2. The same site’s ChatGPT referrals grew 5.7 times overall, but pages nobody touched grew 3.5 times, so most of the growth was ChatGPT itself, not the optimization (Watanabe and Nakayashiki).
3. ChatGPT and Google pick different pages: only 8.3% of the pages ChatGPT cited ranked in Google’s top 10 for the question ([our AI citations and Google rankings study](https://underneath.agency/research/ai-citations-google-rankings-study)).
4. In a lab test of 252,000 trials, formatting-only edits had little effect on which source AI assistants cited, while topic match, price, dates and list position mattered most ([Vishwakarma and colleagues](https://arxiv.org/abs/2605.25517)).

## Did optimizing for ChatGPT hurt Google traffic in the field test?

No measurable harm was found, but the test protected the pages that mattered most to Google. [Watanabe and Nakayashiki](https://arxiv.org/abs/2606.04362) work at Glasp and studied their own site, glasp.co, so this is a self-reported case. In January 2026 they rewrote titles as questions and turned page summaries into short standalone answers across hundreds of thousands of pages.

The safeguard is the important part. They called it an “SEO Guard”: any page earning meaningful Google clicks was locked and never rewritten. Pages with neither Google nor AI interest were taken down, and only the rest were edited. How to read that AI interest from your own records is covered in [what AI bots request in server logs](https://underneath.agency/resources/ai-bot-server-logs-content-demand).

Google clicks to the rewritten section fell about 25% from the second half of 2025 to 2026. Site-wide, Google clicks fell about 20% over the same window, and the rewritten pages’ impressions stayed in their normal range. The authors read the decline as the general trend, not damage from the rewrite, and conclude that “AEO and SEO need not be in tension” (AEO, answer engine optimization, is their term for this work).

## Why does a ChatGPT gain look bigger than it really is?

Because ChatGPT’s own growth lifts every site, so raw before-and-after numbers overstate what the optimization did. On the same site, total ChatGPT referrals grew 5.7 times on monthly figures. Pages that received no optimization at all grew 3.5 times over the same months.

The rewritten pages grew 6.1 times. Comparing them with the untouched pages, the authors put the real effect at about 1.8 to 2.3 times, and call even that “suggestive” rather than proven. A stricter check could not rule out that a jump of that size happened by chance.

This matters for the Google question too. If you only watch one channel, a rising ChatGPT line can hide a falling Google line, or the reverse. Compare optimized pages with similar pages you left alone, in both channels.

## Do ChatGPT and Google even reward the same pages?

Mostly not, which is why work for one does not automatically help or hurt the other. In [our study of 80 US buyer questions](https://underneath.agency/research/ai-citations-google-rankings-study), only 8.3% of the pages ChatGPT cited ranked in Google’s top 10 for the question. Claude was the most Google-like assistant, at 25.7%.

The gap is wide. For ChatGPT, 68.5% of cited pages were not in Google’s top 100 for the question or for any of the searches it ran behind the scenes. Google rankings describe less than a third of what ChatGPT cites. The reverse question is covered in [whether ChatGPT citations boost Google rankings](https://underneath.agency/resources/do-chatgpt-citations-boost-google-rankings).

Google’s own AI surfaces differ from each other as well. On searches that showed both, [the AI Overview, the AI summary at the top of Google’s results, cited 29.7% of top-10 pages and AI Mode cited 16.3%](https://underneath.agency/research/ai-mode-vs-ai-overviews-study). So “optimizing for AI” and “ranking on Google” overlap only in part.

## Which changes help AI assistants without touching what Google rewards?

Changes to substance, not layout, are where the lab evidence points. [Vishwakarma and colleagues](https://arxiv.org/abs/2605.25517), researchers at the software vendor Sprinklr, ran 252,000 trials in which six AI assistants chose between two versions of a product review. The test was simulated: each assistant saw exactly two sources, with brand names removed.

Topic match, a stated price, a recent date and being listed first won citations across all six assistants. Formatting-only edits, such as breaking dense text into sections, “had no impact.” That is useful for the Google question: the edits that worked are content edits you would usually want on the page anyway.

A second lab study points the same way. [Wu and colleagues](https://arxiv.org/abs/2510.11438) at Carnegie Mellon built a tool that rewrites pages to suit AI engines and reported an average improvement of 35.99% in their visibility measures. They found this did not lower the quality of the AI engines’ answers. Neither study measured Google rankings.

Dates are a low-risk change worth testing. In [our study of 3,096 top-10 pages](https://underneath.agency/research/ai-overview-cited-pages-study), a machine-readable date was the one page feature linked to AI Overview citation that held up, at 7.9 points. And [pages published in the last 90 days](https://underneath.agency/research/ai-source-freshness-study) made up 17.4% to 22.6% of each AI assistant’s dated citations, against 6.9% of Google’s top 10. Keeping pages current is a change AI assistants appear to notice more than Google’s top 10 does.

## What should you do about it?

Treat Google-earning pages as protected and run AI-focused changes as a measured experiment. In practice:

1. List the pages that earn meaningful Google clicks today. Exclude them from rewrites, or change them only after a small test.
2. Start AI-focused work on pages that earn little from Google: they carry the least risk.
3. Prefer substance over layout: a clear direct answer, stated prices and specifications, a visible date, and claims backed by evidence.
4. Keep a comparison group of similar pages you do not change. Judge both ChatGPT referrals and Google clicks against that group, not against last quarter.
5. Watch Google clicks and impressions for the edited pages monthly. A drop that outpaces the rest of the site is your signal to stop and review. Deciding [who owns AI search visibility](https://underneath.agency/resources/who-should-own-ai-search-visibility) makes that review someone’s job.

If you want help setting up that kind of guarded program, see [our generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The research has not tested the Google risk properly, so treat today’s answer as provisional.

- The only field evidence is one website, one AI engine and about five months. The authors also say their changes were a bundle, so no single tactic can be credited or blamed.
- The safeguard itself is the main reason for the result. Nobody has published what happens when pages that earn Google clicks are rewritten for AI.
- The lab studies measured AI citations only. They say nothing about Google rankings.
- Our studies compare which pages each system cites or ranks on given days. They do not show what happens to rankings when a page changes.

## Frequently asked questions

### Can SEO and GEO work against each other?

They can in principle, but the one field test found no measurable conflict. On the site studied, Google clicks to rewritten pages fell about 25% against about 20% site-wide, which the authors treat as the general trend.

### Should we rewrite our best-ranking pages for ChatGPT?

Not as a first step. The only safe result published so far came from locking pages that earned meaningful Google clicks and rewriting everything else.

### Does formatting content for AI assistants hurt SEO?

The evidence does not say, but formatting alone barely moved AI citations in a lab test. In 252,000 simulated trials, structure-only edits had little effect, while topic match, prices and dates mattered.

### Why is our ChatGPT traffic growing when we have not changed anything?

ChatGPT’s own growth lifts most sites. On one site, pages with no optimization at all grew 3.5 times in ChatGPT referrals over about five months.

## Sources

- Watanabe and Nakayashiki (2026), [Disentangling Answer Engine Optimization from Platform Growth: A Log-Based Natural Experiment on ChatGPT Referral Traffic](https://arxiv.org/abs/2606.04362), arXiv:2606.04362.
- Vishwakarma, Kumar and Jamidar (2026), [What Gets Cited: Competitive GEO in AI Answer Engines](https://arxiv.org/abs/2605.25517), arXiv:2605.25517.
- Wu, Zhong, Kim and Xiong (2025), [What Generative Search Engines Like and How to Optimize Web Content Cooperatively](https://arxiv.org/abs/2510.11438), arXiv:2510.11438.
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude cite pages that rank?](https://underneath.agency/research/ai-citations-google-rankings-study)
- Underneath (2026), [AI Mode vs AI Overviews: how different are the sources?](https://underneath.agency/research/ai-mode-vs-ai-overviews-study)
- Underneath (2026), [What pages cited by AI Overviews have in common: 3,096 pages](https://underneath.agency/research/ai-overview-cited-pages-study)
- Underneath (2026), [How fresh are the pages AI engines cite?](https://underneath.agency/research/ai-source-freshness-study)

---

This is the Markdown twin of https://underneath.agency/resources/chatgpt-optimization-google-rankings. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Is ChatGPT Referral Growth Proof That Your GEO Is Working?"
description: "Not from growth alone. In one site’s logs, ChatGPT referrals grew 5.7 times, but untouched pages grew 3.5 times. Credit needs a comparison group."
canonical: "https://underneath.agency/resources/chatgpt-referral-growth-and-geo"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Our ChatGPT referral traffic is growing; can we credit our GEO work?

Not from the growth figure alone, because ChatGPT’s own growth lifts traffic to sites that did nothing. In the best-documented case, total ChatGPT referrals grew 5.7 times, but pages that received no optimization grew 3.5 times over the same months. To credit your work, compare the pages you changed with similar pages you left alone, and expect a far smaller, less certain effect than the headline.

## The short version

1. On one large website, ChatGPT referrals grew 5.7 times in five months, while pages that received no optimization grew 3.5 times (Watanabe and Nakayashiki, who work for the site’s owner).
2. After netting out that platform growth, the optimization’s effect was about 1.8 to 2.3 times, and a stricter check could not rule out chance.
3. A reanalysis of the same data by other researchers found the estimate ranged from 1.183 to 2.404 depending only on which week was treated as the start (Kato and colleagues).
4. Google’s own design choices move clicks too: in a test with 1,100 US users, Google’s AI Mode cut click-through to websites by 18.8 percentage points (Wang and colleagues).
5. A review of 45 studies concluded that claims about GEO return on investment outstrip the academic evidence (Martinez).

## Does rising ChatGPT referral traffic prove your GEO work is paying off?

No: ChatGPT’s growth lifts referral traffic to almost every site, whether or not it was optimized. [Watanabe and Nakayashiki](https://arxiv.org/abs/2606.04362) studied glasp.co, a website with hundreds of thousands of question-and-answer pages. In January 2026 the company optimized one section of the site for AI answers and left the rest alone.

On monthly totals, ChatGPT referrals to the whole site grew 5.7 times by May 2026. But the untouched pages grew 3.5 times over the same window, with no optimization at all. The authors call this the platform tailwind, and it was most of the growth.

The authors work at the company that owns the site, and this is one domain studied through ChatGPT alone. Still, it is the clearest public evidence on the question, and its lesson is general: a before-and-after figure on its own cannot separate your work from the market. The same caution applies to visibility scores; see [whether GEO work raised your AI visibility](https://underneath.agency/resources/did-geo-work-raise-ai-visibility).

## How much of the growth was the optimization in the best-documented case?

About 1.8 to 2.3 times, against headline figures of 5.7 times or more. Read from the lowest day to the single best day, the daily series rose 38.2 times. That is the kind of number that ends up in case studies, and the authors call it a fragile basis for any estimate.

The optimized pages grew 6.1 times while the untouched pages grew 3.5 times. Measured week by week against the untouched pages, the optimized section showed a jump of 1.82 times when the work began. The authors put the defensible effect at about 1.8 to 2.3 times.

Even that is not settled. The optimized pages were already gaining share before the work started, by about 11% a month. When the researchers tested fake start dates in the months before, jumps nearly as large appeared by chance; the stricter check gave 0.16, short of the usual bar for proof. The authors suspect “many published AEO multiples are similarly inflated by conflating treatment with tailwind.”

## What else can move AI referral numbers besides your work?

Analytics changes, timing choices and platform design changes can all shift the figures. In the glasp.co data, the share of ChatGPT visits counted as engaged rose from 0.486 in January to 0.896 in May. Most of that jump came from a mid-March change in how bot traffic was filtered, not from visitors behaving differently.

[Kato and colleagues](https://arxiv.org/abs/2609.11915) reanalyzed the published data. Moving the assumed start of the work from 16 December 2025 to 6 January 2026 changed the estimated effect from 1.183 to 2.404 times. Allowing for the mid-March measurement change in their fullest models gave 1.337 to 1.455 times.

The platforms also change underneath you. In a randomized test with 1,100 US participants, [Wang and colleagues](https://arxiv.org/abs/2608.18352) found that hiding Google’s AI Overviews raised click-through to websites by 8.8 percentage points. Switching people to AI Mode cut it by 18.8 points. We cover what that could mean for your traffic in [our article on AI Mode as Google’s default](https://underneath.agency/resources/ai-mode-default-traffic-loss).

## What does a fair test of GEO look like?

Compare pages you changed with similar pages you did not, over the same weeks. Untouched pages on the same site share the same platform growth, brand and analytics settings. The difference between the two groups is what your work can claim.

Researchers use the same logic for platform effects. [Khosravi and Yoganarasimhan](https://arxiv.org/abs/2602.18455) compared English Wikipedia articles, exposed to AI Overviews by default, with the same articles in German and French. They estimate AI Overviews cut English Wikipedia’s search traffic by 5.45% against the German comparison.

The strongest design is random assignment. Watanabe and Nakayashiki say holding back a random subset of eligible pages would have turned their estimate into a clean experiment; they did not run one. Their comparison did catch one useful result: organic Google clicks to the optimized pages fell about 25%, close to the 20% fall across the whole site, so the work did not visibly harm search traffic. Whether the reverse holds, with ChatGPT citations lifting Google, is covered in [whether ChatGPT citations boost Google rankings](https://underneath.agency/resources/do-chatgpt-citations-boost-google-rankings).

## Should referral traffic be your main measure of GEO?

No: people can read an AI answer without clicking, so referrals capture only part of the effect. Kato and colleagues point out that referral sessions measure a different event from being seen. They measured mentions directly instead, asking two ChatGPT models 56 questions; Glasp appeared in 33.8% of one model’s answers.

Referrals also reflect only the engines your analytics can see. On glasp.co, AI engines other than ChatGPT sent less than 1% of AI-referred visits in the study window. A brand can gain visibility in answers that send few or no visitors. Server logs add another view, showing [which pages AI bots request](https://underneath.agency/resources/ai-bot-server-logs-content-demand).

The wider evidence on traffic is thin. [Martinez](https://arxiv.org/abs/2607.14035) reviewed 45 studies and found no technique with a stable, long-term effect on discovery or on what users do. One industry study reported a 20% traffic lift against a control group, but without enough detail to judge it. The survey’s conclusion: “claims about GEO return on investment clearly outstrip the academic evidence.”

## What should you do about it?

Before crediting GEO for referral growth, set up a comparison that can tell your work apart from ChatGPT’s growth.

1. Keep a control group. Leave a similar set of pages untouched, ideally chosen at random, and compare growth between the two.
2. Report the gap between groups, not the raw multiple. If both grew, only the difference belongs to your work.
3. Fix the start date in advance and note any analytics or bot-filter changes, since both can move the estimate.
4. Track AI answers directly as well as clicks: whether you are mentioned, how often, and on which engines. What predicts those mentions is examined in [our study of Wikipedia, schema and recommendations](https://underneath.agency/research/brand-entity-ai-recommendations-study).
5. Watch Google’s organic traffic at the same time, to make sure [AI-focused changes do not cost you search visits](https://underneath.agency/resources/chatgpt-optimization-google-rankings).
6. Treat early results as suggestive until they hold for several months.

If you want help designing a measurement plan like this, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

There is still no randomized, multi-site study showing how much GEO work raises AI referral traffic.

- The key evidence is one website, one engine and five months, published by the site’s own staff.
- Its optimized and untouched pages differed in content and purpose, so they were not a perfect match.
- The work was a bundle of changes, so no single tactic can be credited.
- Nobody has yet linked AI mentions or referrals to sales in a published controlled study.

## Frequently asked questions

### How do I measure the ROI of GEO?

Compare the pages you optimized with similar untouched pages over the same period. On one site, untouched pages grew 3.5 times on their own, so only growth beyond that belongs to the work.

### Why is our ChatGPT traffic growing when we have not changed anything?

ChatGPT itself is growing. In one study, pages that received no optimization still saw ChatGPT referrals grow 3.5 times in about five months.

### Are published GEO case studies with big traffic multiples reliable?

Treat them with caution. In the one carefully analyzed case, a 38.2-times peak and a 5.7-times monthly rise shrank to an effect of about 1.8 to 2.3 times once platform growth was removed.

### Can a GEO change hurt our Google rankings?

In the one case studied, no measurable harm was found. Organic clicks to the optimized pages fell about 25%, in line with a 20% fall across the whole site.

## Sources

- Keisuke Watanabe and Kazuki Nakayashiki (2026), [Disentangling Answer Engine Optimization from Platform Growth: A Log-Based Natural Experiment on ChatGPT Referral Traffic](https://arxiv.org/abs/2606.04362), arXiv:2606.04362.
- Masahiro Kato, Daiki Honma and Taka Kato (2026), [Generative Marketing Mix Modeling: A Causal Inference Framework Linking GEO and GEM to Business Impact](https://arxiv.org/abs/2609.11915), arXiv:2609.11915.
- Olivier Martinez (2026), [Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026)](https://arxiv.org/abs/2607.14035), arXiv:2607.14035.
- Stephanie T. Wang, Jeffrey Gleason, Yakov Bart, Christo Wilson and Danaé Metaxa (2026), [AI in Search Reduces Publisher Referrals Without Improving User Experience: Experimental Evidence](https://arxiv.org/abs/2608.18352), arXiv:2608.18352.
- Mehrzad Khosravi and Hema Yoganarasimhan (2026), [Impact of AI Search Summaries on Website Traffic: Evidence from Google AI Overviews and Wikipedia](https://arxiv.org/abs/2602.18455), arXiv:2602.18455.

---

This is the Markdown twin of https://underneath.agency/resources/chatgpt-referral-growth-and-geo. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Where ChatGPT fits in the customer journey vs Google"
description: "Later than Google: in a 2026 US and British panel, search usually opened the journey, while people more often browsed before using an AI assistant than after."
canonical: "https://underneath.agency/resources/chatgpt-vs-google-customer-journey"
published: 2026-10-08
updated: 2026-10-10
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Where does ChatGPT fit in the customer journey compared with Google?

ChatGPT usually sits later in the journey than Google, not at its start. In a 2026 panel of US and British web users, people tended to read web pages after a Google search but before opening an AI assistant. That suggests what buyers read on the open web may travel with them into the AI conversation, though no study has yet proven that link.

## The short version

1. Web pages followed a search in 60.2% of search sessions but came before the assistant in 45.5% of assistant sessions (Iannelli and Ai, 2026).
2. Comparing the same people, search and assistants differed by 20.6 percentage points in this before-and-after pattern, and by 20.3 a month later.
3. Just over half of assistant sessions (51.9%) showed web activity before the first prompt, and a search came before it in 33.8%.
4. Assistant sessions ran longer than search sessions, a median of 22 minutes against 9, and switched between the assistant and the web 5.2 times on average.
5. ChatGPT also reads the web for itself: in our study it ran 3.7 searches per buyer question before answering.

## Does ChatGPT replace Google at the start of the journey?

Mostly not: search tends to open the journey, while AI assistants sit deeper inside it.

That is the central finding of [Iannelli and Ai](https://arxiv.org/abs/2607.04282), who linked an opt-in panel’s AI assistant activity to the same people’s searches and page visits. The panel covered the United States and Great Britain in February 2026. The assistants were standalone ones such as ChatGPT and Gemini, not Google’s AI Overviews.

Two cautions come first. The authors work at Scrunch AI, a commercial company, so this is vendor research. The paper was also under review and not yet peer reviewed when we read it.

The table compares where web page visits fell, counting each person equally.

| Session centered on | Pages visited before | Pages visited after |
|---|---|---|
| A Google or other search | 46.7% | 60.2% |
| An AI assistant | 45.5% | 38.9% |

Around search, page visits mostly came afterward: a query, then the results it led to. Around assistants the order flipped. In the authors’ words, “search tends to open the observed journey, while conversational assistants sit deeper inside it.” That is one reason the evidence does not show [ChatGPT replacing Google search](https://underneath.agency/resources/is-chatgpt-replacing-google).

The gap between the two patterns, measured within the same people, was 20.6 percentage points. It held at 20.3 points when the team repeated the analysis on March 2026 data.

## What do people do before they open an AI assistant?

Often they have already been browsing; just over half of assistant sessions had web activity before the first prompt.

Counting each person equally, 51.9% of sessions showed some web activity before the first prompt. A search engine visit came before the first prompt in 33.8% of sessions and after the last response in 24.4%. Search appeared somewhere in 46.0% of assistant sessions.

Sessions with web activity only before the assistant (18.3%) outnumbered those with web activity only after it (10.5%). The popular picture of “ask the AI first, then click out” describes roughly one session in ten. It fits other evidence that [people click fewer websites with AI answers](https://underneath.agency/resources/ai-answer-engines-fewer-website-clicks).

The authors are careful about what timing can show. Their data held no page content and no prompt text. So it cannot prove that the pages someone read were why they opened ChatGPT, or that both belonged to one task.

## How is an AI assistant session different from a search session?

It lasts longer and mixes more activity: a median of 22 minutes, against 9 for a search session. Inside assistant sessions, prompts and responses made up 59.7% of recorded steps, searches 8.8% and page visits 31.5%. People switched between the assistant and the web 5.2 times on average, against 4.0 switches in search sessions.

A search session, in plain terms, is a quick hop from a query to a few pages. An assistant session is a longer working conversation that dips in and out of the web. The authors describe AI assistants as a possible “synthesis layer deeper in the journey,” while search “remains an access layer.” They stress these are broad patterns, not labels proven for each session.

## Does the pattern hold when people research purchases?

Yes, within the site categories the researchers could check, shopping included. They grouped sessions by the kind of sites people visited and found the same reversal in every group. It measured 19.8 points for shopping sites, 21.1 for news and 26.8 for reference sites.

This is a rough check, not a study of buying journeys. Only 22.5% of assistant sessions visited sites that could be put in a category.

Research on Google points the same way for AI answers in general. [Grossman and colleagues](https://arxiv.org/abs/2604.27790) ran 11,500 queries through Google in December 2025.

An AI Overview, the AI summary at the top of Google’s results, appeared on only 17.4% of plain product searches. It appeared on 88.2% of product comparison searches and 92% of product questions.

Those two sets were written by AI from real shopping searches, so treat them as a test, not real traffic. The authors conclude that AI search plays a larger role in the research stage than at the final purchase.

## Does what buyers read earlier shape what the assistant tells them?

Possibly, but no study has shown it yet.

What we do know is that the assistant reads the web itself. In [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study), ChatGPT ran a mean of 3.7 searches per buyer question before answering. None of the 509 searches we collected repeated the user’s question word for word.

The pages it cites are often not the ones Google ranks highly. In [our rankings study](https://underneath.agency/research/ai-citations-google-rankings-study), 8.3% of the pages ChatGPT cited ranked in Google’s top 10 for the question as typed. So a buyer can meet your brand on Google first, then meet a different set of sources inside ChatGPT.

## What should you do about it?

Treat AI assistants as a research step in the middle of the journey, not as its front door. Practical steps:

1. Keep investing in search. Search still opens most observed journeys, so your Google presence shapes what buyers see before they ask an assistant.
2. Publish the material buyers compare. Comparison and question searches drew AI answers far more often than plain product searches, so pages that answer “how does X compare with Y?” matter.
3. Check what assistants say about you. They run their own searches and cite pages that often do not rank on Google, so ask them the questions your buyers ask.
4. Measure the whole journey. A buyer who read your page last week may [never click through from ChatGPT](https://underneath.agency/resources/ai-assistant-sessions-without-website-visits), so watch branded searches and direct visits, not only AI referrals.

If you want help mapping where AI assistants enter your buyers’ research, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The research shows the order of events, not their cause or the buyer’s intent. The main gaps:

- Whether earlier reading shapes the prompt. The panel study kept no page content or prompt text, so the link is a reasonable guess, not a finding.
- Business buyers. The panel was an opt-in group of web users that the authors call AI-forward, not a sample of B2B purchasers.
- Mobile apps. Assistant use in native mobile apps and through APIs was not observed.
- Time and place. The data covers two months, February and March 2026, in two countries.
- Google’s own AI. AI Overviews and AI Mode were outside the panel study, because they sit inside the results page.

## Frequently asked questions

### Do people use ChatGPT before or after Google?

More often after browsing than before it, in the one large study that tracked both. In a 2026 US and British panel, page visits came before the first assistant prompt in 45.5% of sessions and after the last response in 38.9%.

### Is ChatGPT replacing Google search?

The data does not show replacement. Search appeared somewhere in 46.0% of assistant sessions, and the authors describe coexistence while saying their design cannot settle whether AI reduces search overall.

### Where in the buying process do AI answers show up most?

At the research stage. In one 2025 test, Google showed AI Overviews on 88.2% of product comparison searches but only 17.4% of plain product searches.

### Does ChatGPT find the same pages as Google?

Often not. In our study, 8.3% of the pages ChatGPT cited ranked in Google’s top 10 for the same question.

## Sources

- Iannelli, M. and Ai, A. (2026), [The New Shape of Search: How Conversational AI Recomposes Information Seeking](https://arxiv.org/abs/2607.04282), arXiv:2607.04282.
- Grossman, R., Liu, S., Chen, M. K., Smith, M., Borcea, C. and Chen, Y. (2026), [How Generative AI Disrupts Search: An Empirical Study of Google Search, Gemini, and AI Overviews](https://arxiv.org/abs/2604.27790), arXiv:2604.27790.
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude cite pages that rank?](https://underneath.agency/research/ai-citations-google-rankings-study)

---

This is the Markdown twin of https://underneath.agency/resources/chatgpt-vs-google-customer-journey. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How chemical companies win B2B buyers through AI search"
description: "Chemical suppliers reach B2B buyers in AI search when grades, specifications, safety data sheets and regulatory status are public and easy to verify."
canonical: "https://underneath.agency/resources/chemical-suppliers-b2b-buyers-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# When a formulator asks AI who supplies a chemical, does our company come up?

It can, when an assistant can find and check your grades, specifications, applications, regulatory status and where you ship. Chemical buyers are technical, cautious and slow to switch, so AI search rarely closes a sale. What it can do is decide which suppliers a formulator or buyer requests samples and documents from first.

This guide is for chemical producers and distributors that sell to industrial customers: specialty and basic chemicals, ingredients, additives, coatings raw materials and water treatment products. It covers legitimate commercial discovery only. It is not safety, regulatory or legal advice.

## The short version

1. Chemicals lead industrial sourcing: on Thomasnet, [chemicals were the most sourced product category of 2025](https://distributionstrategy.com/2026/01/industrial-sourcing-behavior-shifts-in-2025-signaling-strategic-imperatives-for-distributors/), led by specialty formulations, coatings and cleaning compounds, and buyers increasingly evaluated alternative and secondary suppliers.
2. AI is now routine inside chemical companies: in a [Journal of Business Chemistry survey](https://www.businesschemistry.org/article/5126-2/) of 124 German chemical and pharmaceutical companies, active AI use rose from 34% in 2020 to 76% in 2025, with generative AI tools such as ChatGPT part of daily work.
3. Industrial customers dominate demand: more than 80% of basic and specialty chemicals are consumed by the industrial sector, the American Chemistry Council said in [SCI’s 2026 outlook](https://soci.org/chemistry-and-industry/cni-data/2026/1/chemistry-in-2026-navigating-the-year-of-uncertainty), and US output rose only 0.7% in 2025, so growth often means winning share from other suppliers.
4. Buyers want technically capable partners: specialty chemical manufacturers in [L.E.K. Consulting’s 2026 survey](https://www.lek.com/insights/industrials/four-trends-shaping-us-specialty-chemicals-2026) favored end-market-specific and regional distributors that offer technical knowledge and formulation support.
5. Regulation is part of every search: the [TSCA Inventory](https://www.epa.gov/tsca-inventory/about-tsca-chemical-substance-inventory) lists more than 86,000 chemicals, and the EU’s [REACH rules](https://environment.ec.europa.eu/topics/chemicals/reach-regulation_en) require registration of substances above 1 tonne per year per company, so buyers ask about regulatory status as often as about price.

## Who buys industrial chemicals, and how do they decide?

Formulators, engineers and procurement teams at manufacturers, who qualify a supplier slowly and then buy for years.

Several people at the customer judge a new chemical supplier. An R&D chemist or formulator chooses the ingredient and grade. A process or plant engineer cares about handling, compatibility and consistency. A procurement or category manager negotiates price, supply security and terms, and a regulatory or environmental, health and safety specialist checks documents. Across business purchases generally, [Forrester](https://www.forrester.com/press-newsroom/forrester-2026-the-state-of-business-buying/) finds procurement professionals are decision-makers in 53% of buying cycles.

The decision follows a familiar path. A buyer identifies a need, such as a new product, a reformulation, a supply problem or a tariff change. They research candidate suppliers, request a technical data sheet and a safety data sheet (SDS), order samples, run lab and plant trials, and only then add a supplier to the approved list. Once a material is specified into a formulation, switching means requalifying, so a supplier that wins a spec can keep the volume for years. Building product makers run a similar race, covered in our guide to [how architects find materials through AI](https://underneath.agency/resources/building-materials-specification-ai-search).

The market is under pressure, which raises the stakes of each new account. The ACC expected growth in only nine of the 20 key chemistry end-use industries it tracks in 2025, rising to 12 in 2026. Global share has shifted too: according to the European chemical industry group Cefic, quoted by SCI, China now accounts for 46% of global chemical sales. Distributors feel the same squeeze: [Brenntag](https://www.brenntag.com/en-de/media/news/brenntag-reports-fullyear-2025-financial-results.html), the large chemical distributor, reported 2025 sales of EUR 15.2 billion, down 3.7%. In a flat market, the supplier that gets into more early comparisons gets more chances to win.

There is no public benchmark for what a typical chemical account is worth. Our inference: because qualification is slow and switching is costly, one spec-in can be worth far more than the first order suggests.

## Where does AI search already sit in chemical buying?

In early research and comparison, used widely by technical buyers but checked against documents and supplier sites.

Evidence on chemical buyers specifically is thin, so we separate what is measured from what is inferred.

- **Chemical companies use AI at work.** The German survey found 91% of companies rated digitalization relevant or very relevant in 2025, and active AI use more than doubled to 76%. It describes generative AI tools, such as ChatGPT and company assistants, as increasingly embedded in daily operations. It does not measure supplier searches.
- **Technical buyers use generative AI in purchasing.** The [2026 State of Marketing to Engineers research](https://advertising.globalspec.com/state-of-marketing-to-engineers-research-report) by GlobalSpec and TREW Marketing, which surveyed engineers across industries, found 69% of technical buyers use generative AI during the purchasing process, but rate their trust in its answers at 4.7 out of 10.
- **Independent sources are gaining weight.** In the same survey, online technical publications (76%) edged out supplier and vendor websites (74%) as the place technical buyers routinely research purchases.
- **Specialty chemical firms are going digital.** L.E.K. reports significant increases since 2022 in the use of AI, generative AI and lead generation software among specialty chemical companies.
- **Buyers still want people.** Distributor ChemPoint’s [2025 industry survey](https://www.chempoint.com/en-ca/insights/2025-chempoint-chemical-industry-survey-report) of over 150 chemical industry professionals frames its findings around why AI has not replaced the human touch in B2B buying.

Our reading: chemists and buyers use AI to find candidate suppliers and grades faster, then verify everything against technical data sheets, SDSs and trusted publications before they contact anyone.

## Which supplier questions do chemical buyers put to AI?

Questions that join an application, a grade, a regulatory status, a pack size and a region.

We wrote the examples below for illustration; they are not records of real buyer prompts.

| What the buyer needs | Example question (our wording) |
|---|---|
| Reformulation | “Alternatives to fluorinated surfactants for an industrial floor coating, with US supply” |
| Food and personal care grades | “Food-grade sodium citrate distributors in Texas, tote quantities, kosher certified” |
| Sustainable inputs | “Bio-based plasticizers compatible with flexible PVC that are registered under REACH” |
| Documentation | “Where can I get the technical data sheet and SDS for this grade?” |
| Local supply | “Water treatment chemical distributors in the Midwest with bulk delivery” |
| Comparison | “Compare these two distributors on technical support, lead time and minimum order” |

Each question works as a filter. If your product pages do not state the grade, the typical specifications, the applications, the certifications, the pack sizes and the regions you serve, an assistant has nothing to match. If your documents sit behind a login with no public summary, a buyer may never learn you carry the grade.

Reformulation questions deserve special attention. L.E.K. found roughly half of specialty chemical companies have reformulated or phased out products containing chemicals of concern in the past three years, with a similar share expecting to continue. Every reformulation starts with a search for an alternative ingredient and a supplier who can document it.

## How does an AI answer turn into a sample request and a contract?

Through a candidate list, a document check, samples, qualification trials and, finally, a place on the approved supplier list.

1. **Candidate list.** A formulator or buyer asks an assistant, a directory such as Thomasnet or a search engine for suppliers of an ingredient for an application.
2. **Document check.** They look for the technical data sheet, the SDS, regulatory status and certifications on supplier websites and trusted publications.
3. **Sample request.** The supplier receives a request for samples, usually alongside other suppliers. This is the first moment the supplier sees the opportunity.
4. **Qualification.** Lab work, plant trials, quality audits and customer approvals follow, often over months.
5. **Spec-in and supply.** The approved material goes into production, with repeat orders as long as quality, supply and price hold.

An assistant’s answer reaches only the candidate list and the document check. Product performance, documentation quality, technical service and price decide the rest. A supplier that is missing from the candidate list never gets to prove any of those. Component makers face the same early cut when [engineers ask AI for a part](https://underneath.agency/resources/industrial-manufacturers-ai-search).

## What makes an assistant name one chemical supplier over another?

Information it can find and corroborate; the platforms document how they search, not how they rank suppliers.

**Documented by the platforms.** [OpenAI says](https://help.openai.com/en/articles/9237897-chatgpt-search) ChatGPT search typically rewrites a question into one or more targeted queries that it sends to search providers, and that sites should allow OpenAI’s search crawler, OAI-SearchBot, to be eligible for inclusion. [Google says](https://blog.google/products/search/ai-mode-search/) AI Mode issues multiple related searches across subtopics and data sources, a technique it calls “query fan-out.”

**Observed in our studies**, which covered buyer questions in several industries but not chemicals:

- ChatGPT ran a mean of 3.7 searches per buyer question before answering ([hidden searches study](https://underneath.agency/research/ai-hidden-searches-study)).
- Each tenfold increase in independent sites naming a brand in the cited pages went with 4.7 times the odds of being recommended ([brand entity study](https://underneath.agency/research/brand-entity-ai-recommendations-study)).
- Of 269 numbered “best” lists cited by AI surfaces, 24.2% ranked their own publisher first ([self-promoting lists study](https://underneath.agency/research/self-promoting-best-lists-study)), so a distributor’s own “top suppliers” page is unlikely to be the only list an assistant reads.
- Only 25.2% of the brands ChatGPT named for a question appeared in all five repeats ([consistency study](https://underneath.agency/research/ai-recommendation-consistency-study)).

**Our inference for chemicals.** The trust factors are the ones a regulatory specialist already checks: product identity with CAS numbers, grade and typical properties, applications, regulatory status by market, certifications such as ISO 9001, kosher, halal or food-safety schemes with the issuing body named, packaging sizes, plants and warehouses, and up-to-date SDS access. Coverage in trade publications, technical papers, association listings and manufacturer line cards gives an assistant independent confirmation.

## Why do safety data and regulatory facts matter so much in AI answers?

Because buyers filter suppliers by compliance first, and an assistant can only repeat what is published.

Chemical buying runs on documents. In the US, OSHA’s [Hazard Communication standard](https://www.osha.gov/hazcom) gives safety data sheets a specified 16-section format, and the agency’s [SDS guide](https://www.osha.gov/sites/default/files/publications/OSHA3514.pdf) explains what each section contains. TSCA requires EPA to keep the Inventory of chemical substances made or imported in the United States. In the EU, REACH entered into force in 2007, and a company that receives a consumer inquiry about a substance of very high concern in an article must reply within 45 days.

For AI visibility, the lesson is practical, not legal. Buyers ask regulatory questions, and assistants answer from whatever pages they find. A supplier that states, accurately and with dates, which grades are TSCA-listed or REACH-registered, and where the current SDS can be requested, gives the assistant a reliable source. A supplier that says nothing leaves the answer to distributors, directories or outdated pages. And a supplier that overstates a safety or sustainability attribute risks having that claim repeated to buyers and regulators, which is why every such statement should pass your regulatory and legal review first.

Nothing here changes existing duties. Suppliers of regulated or dual-use substances screen customers and restrict sales as the law requires, and AI visibility work should never make restricted products easier to obtain.

## What does a chemical supplier lose when assistants leave it out?

Sample requests it never hears about, during a period when buyers are actively testing alternative suppliers and ingredients.

We found no public measurement of chemical sales lost to AI absence, so this is our reasoning, with the evidence that supports it:

- **Buyers are diversifying.** Thomas’ 2025 data shows procurement teams evaluating alternative and secondary suppliers to reduce risk. A supplier absent from the first list is absent from that evaluation. Packaging buyers keep second suppliers too, as our guide to [winning packaging quote requests through AI](https://underneath.agency/resources/packaging-suppliers-quote-requests-ai-search) explains.
- **Research happens before contact.** In the engineering survey, 62% of the buying process happened online before technical buyers engaged with sales. Suppliers learn about an opportunity only at the sample request.
- **Wrong facts disqualify.** An assistant that says you do not carry a grade, do not ship to a region or lack a certification removes you quietly. In [our study of business facts](https://underneath.agency/research/ai-business-facts-accuracy-study), 18.9% of AI answers about local businesses stated at least one fact that differed from the business’s Google profile. That study did not cover chemicals, but it shows how easily details drift. If an assistant already misstates one of your grades or certifications, our guide to [fixing wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) shows how to find the distributor page or directory listing it came from.

## How does GEO work for a chemical company?

Generative engine optimization (GEO) makes your products, documents and credentials easy for AI assistants to find, describe and verify.

For a chemical producer or distributor, the work usually includes:

1. **Product and grade pages in text.** Name, CAS number, grade, typical properties, applications, compatible systems, pack sizes and regions served, written on the page rather than only in a PDF.
2. **Document access that machines can see.** A public description of each technical data sheet and SDS, with a clear way to request the current version, even when the document itself sits behind a form.
3. **Regulatory status by market.** Plain, dated statements of TSCA, REACH and other listings, reviewed by your regulatory team.
4. **Consistent identity across channels.** The same product names, grades and facts on your site, manufacturer line cards, distributor pages, Thomasnet and association directories. Our guide on [how brands build authority for AI search](https://underneath.agency/resources/how-brands-build-authority-for-ai-search) explains why matching product facts and outside coverage reinforce each other.
5. **Independent technical coverage.** Trade press articles, conference papers, application notes co-published with customers, and association memberships.
6. **Readable, crawlable pages.** Make sure product pages load without logins and that search crawlers can reach them. Our guide on [what happens when AI agents can’t read your site](https://underneath.agency/resources/when-ai-agents-cant-read-your-site) covers common blockers.
7. **Measurement.** Ask a fixed set of application, grade, regulatory and regional questions across ChatGPT, Gemini, Perplexity, Copilot and Google’s AI features, repeatedly, and compare the answers with the sample requests you actually receive.

No provider can guarantee that an assistant will name your company. The aim is to make your company the easiest supplier for an assistant, and then a chemist, to verify.

## What can’t the current data tell a chemical supplier?

It shows AI spreading through chemical companies, not how often AI answers create sample requests or contracts.

- **No chemical-specific attribution.** We found no public data linking AI answers to chemical sample requests, qualifications or sales.
- **Usage surveys are broad.** The German survey covers AI use at work across chemical and pharmaceutical companies, not supplier searches. The engineering survey covers many industries.
- **Some results are gated.** ChemPoint’s survey numbers are not public, so we did not use them.
- **Our studies are cross-industry.** Applying them to chemicals is our inference.

Toll and custom manufacturers face a similar buyer path; our guide on [how contract manufacturers win RFQs through AI search](https://underneath.agency/resources/contract-manufacturers-rfqs-from-ai-search) covers it.

## Where should a chemical company begin?

Begin by asking assistants the application, grade and regulatory questions your best customers would ask.

That check shows whether your company is named for the products and markets that matter, whether grades, documents and regulatory status are described correctly, which distributors, directories and publications the answers rely on, and which suppliers appear in your place. The work then is to publish the facts buyers need in a form assistants can read, and to earn the independent coverage that confirms them.

If your growth depends on new sample requests turning into qualified accounts, [ask us to review how AI search describes your products](https://underneath.agency/contact). We put your application, grade and regulatory questions to the main assistants, show which distributors and competing producers are named for them, and point to the product, document and coverage gaps that most limit qualified sample requests. How the follow-on work is run, from grade pages and dated regulatory statements to distributor consistency, is described on our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page.

## Frequently asked questions

### Do chemists and buyers really use ChatGPT to find suppliers?

Many use generative AI at work: 76% of German chemical and pharmaceutical companies reported active AI use in 2025, and 69% of technical buyers across industries used it in purchasing. Most still verify suppliers through documents and trusted publications.

### Should we publish safety data sheets openly?

That is a decision for your regulatory and legal teams. At minimum, describe each product and grade in text and make it clear how a buyer can get the current SDS, so an assistant can point buyers to you.

### Does this matter for distributors as well as producers?

Yes. Distributors are often the supplier a buyer contacts, and L.E.K. found manufacturers favoring distributors with technical knowledge and regional reach. Distributor pages need the same grade, document and regulatory facts as the producer’s.

### Can GEO help sell restricted or controlled chemicals?

No. GEO is about accurate commercial information for legitimate industrial buyers. Customer screening and legal restrictions apply as before.

## Sources

- Distribution Strategy Group (2026-01-30), [Industrial Sourcing Behavior Shifts in 2025 (Thomas 2025 Annual Sourcing Activity Report)](https://distributionstrategy.com/2026/01/industrial-sourcing-behavior-shifts-in-2025-signaling-strategic-imperatives-for-distributors/)
- Daubenfeld, Hasselbach and Just, Journal of Business Chemistry (2025-10), [Artificial Intelligence in the German Chemical and Pharmaceutical Industry: A Comparative Analysis of Empirical Survey Results from 2020 and 2025](https://www.businesschemistry.org/article/5126-2/)
- SCI Chemistry & Industry (2026-01), [Chemistry in 2026: Navigating the year of uncertainty](https://soci.org/chemistry-and-industry/cni-data/2026/1/chemistry-in-2026-navigating-the-year-of-uncertainty)
- L.E.K. Consulting (2026-05-29), [Four Trends Shaping US Specialty Chemicals in 2026](https://www.lek.com/insights/industrials/four-trends-shaping-us-specialty-chemicals-2026)
- Brenntag (2026), [Brenntag reports full-year 2025 financial results](https://www.brenntag.com/en-de/media/news/brenntag-reports-fullyear-2025-financial-results.html)
- ChemPoint (2025), [2025 ChemPoint Chemical Industry Survey Report](https://www.chempoint.com/en-ca/insights/2025-chempoint-chemical-industry-survey-report)
- GlobalSpec and TREW Marketing (2026), [State of Marketing to Engineers research report](https://advertising.globalspec.com/state-of-marketing-to-engineers-research-report)
- Forrester (2026-01-21), [Forrester’s 2026 Buyer Insights: GenAI Is Upending B2B Buying](https://www.forrester.com/press-newsroom/forrester-2026-the-state-of-business-buying/)
- US Environmental Protection Agency, [About the TSCA Chemical Substance Inventory](https://www.epa.gov/tsca-inventory/about-tsca-chemical-substance-inventory)
- European Commission, [REACH Regulation](https://environment.ec.europa.eu/topics/chemicals/reach-regulation_en)
- Occupational Safety and Health Administration, [Hazard Communication](https://www.osha.gov/hazcom)
- Occupational Safety and Health Administration, [Hazard Communication Standard: Safety Data Sheets](https://www.osha.gov/sites/default/files/publications/OSHA3514.pdf)
- OpenAI Help Center (2026), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Google (2025-03-05), [Expanding AI Overviews and introducing AI Mode](https://blog.google/products/search/ai-mode-search/)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [Do Wikipedia and schema make AI assistants recommend a brand?](https://underneath.agency/research/brand-entity-ai-recommendations-study)
- Underneath (2026), [How many “best of” lists cited by AI rank their own brand first?](https://underneath.agency/research/self-promoting-best-lists-study)
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)
- Underneath (2026), [Do AI answers match a business’s Google profile?](https://underneath.agency/research/ai-business-facts-accuracy-study)

---

This is the Markdown twin of https://underneath.agency/resources/chemical-suppliers-b2b-buyers-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Chinese vs Western AI Models: Does Your Brand Show Up?"
description: "Likely yes. One study found Chinese AI models named brands far more often than Western ones, though its authors have a stake. Language clearly matters."
canonical: "https://underneath.agency/resources/chinese-vs-western-ai-brand-visibility"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Does my brand show up differently in Chinese AI models than in Western ones?

Probably, but the evidence is thinner than the headlines suggest. The only direct comparison found that Chinese AI models named brands far more often than Western models when asked identical English questions, and named Chinese brands most of all. Its authors work for a firm that shares a name with their case-study brand, so the firmer ground is independent research showing that the language of a question changes which brands and sources AI assistants reach for.

## The short version

1. In one study of 1,909 English-language questions, three Chinese AI models named the brand asked about in 88.9% of answers and three Western models in 58.3% (Huang and colleagues, whose firm shares a name with the brand they featured).
2. The same study found Chinese brands named in 31.2% of Western answers and 96.8% of Chinese ones, while Western brands moved only 9.2 points.
3. Across 66 European brands, asking in a brand’s home language instead of English raised local champions’ recommendation share by 0.80 on a 0-to-1 scale, against 0.15 for global brands (Żatuchin, who runs an AI visibility company).
4. University of Toronto researchers who translated 100 shopping questions into five languages found ChatGPT switched to different websites in each language, while Claude kept citing English-language ones.
5. In a fake-brand test of 12 Chinese and Western models, every one could be fooled, so neither group is safer from manipulation.

## Do Chinese AI models mention brands more often than Western ones?

In the only direct test we found, yes: Chinese models named brands far more often. [Huang and colleagues](https://arxiv.org/abs/2601.00869) asked three Western models (GPT-4o, Claude and Gemini) and three Chinese ones (Qwen3, DeepSeek and Doubao) about 30 collaboration and productivity software brands. They kept only questions written entirely in English.

Across those 1,909 question-and-model pairs, the Chinese models mentioned the brand in 88.9% of answers and the Western models in 58.3%, a gap of 30.6 percentage points. The gap depended on where the brand came from.

| Brand origin | Western models | Chinese models |
|---|---|---|
| Western brands | 72.1% | 81.3% |
| Chinese brands | 31.2% | 96.8% |
| Global or mixed brands | 71.6% | 88.6% |

The gap also depended on the kind of question. For “best tools for X” questions it was 46.9 points; for “What is X?” questions it was 12.5. When the Chinese models did name a brand, they described it more warmly: +0.71 against +0.42 on a scale running from −1 (negative) to +1 (positive).

## How much weight should you put on that finding?

Some, but not much on its own, because the study has a clear conflict of interest and a narrow scope. The three authors work at OmniEdge (Zhizibianjie) AI Consulting in Shenzhen. Their case-study brand, Zhizibianjie, was named in 65.6% of 32 questions to Chinese models and in none of 32 to Western ones. The paper declares no conflict, but the shared name is plain to see.

The scope is narrow too. It covers one software category, six models and one collection window, and the paper’s own dates do not agree with each other. To keep every question in English, the authors removed 891 question-and-model pairs (31.8%), mostly about Chinese brands whose names are written in Chinese characters. The Chinese-brand figures therefore rest on a filtered subset.

The authors also say their design cannot show cause and effect. Their 18-month plan for building visibility in other markets is advice, not a tested result. We would not base a budget on this paper alone.

## Is it the AI model or the language of the question?

Both matter, and language alone changes which websites and brands an assistant reaches for. [Chen and colleagues at the University of Toronto](https://arxiv.org/abs/2509.08919) translated 100 English shopping questions into Chinese, Japanese, German, French and Spanish. They put them to ChatGPT, Perplexity, Gemini and Claude.

Under non-English questions, the assistants [cited more websites written in that language](https://underneath.agency/resources/do-ai-engines-cite-local-language-sources). ChatGPT switched to an almost entirely different set of sites in each language. Claude kept citing much the same English-language authorities whatever the language. Chinese questions were strongly localized, and changing the language moved results more than rewording a question did.

[Żatuchin](https://arxiv.org/abs/2606.23165) ran 35,640 answers from GPT-5.4, Gemini 3.1 Pro and Perplexity about 66 brands in 12 European languages. For local champions, asking in the home language instead of English raised the share of answers naming them by 0.80 on a 0-to-1 scale; for global brands, by 0.15. The author is affiliated with Rankfor.AI, an AI visibility company, so treat the paper as vendor research.

Even within English-speaking markets, the country changes the answer. In [our country study](https://underneath.agency/research/ai-recommendations-by-country-study), two ChatGPT answers from the same country shared 0.594 of their brands; two from different countries shared 0.429.

## Which brands are most affected?

Brands that are strong in one home market but thin in English-language sources are affected most. In the Huang study, the gap between Chinese and Western models was 9.2 points for Western brands and 65.6 points for Chinese brands. In the Żatuchin study, local champions were rarely named in English buyer questions and named in almost every home-language one.

The Toronto team saw the same pattern by category. Brand lists stayed more alike across languages where a few global brands dominate, such as cameras and laptops, and differed more in localized categories such as home appliances.

Local publications matter at the margin. Żatuchin found Wikipedia was the most-cited website in 11 of the 12 languages. In Lithuanian, the national business daily vz.lt edged it out, with 4.38% of citations.

## Are Chinese AI assistants easier to manipulate?

Not on the available evidence: in one test, Chinese and Western models were similarly easy to fool. [Luo and Chen](https://arxiv.org/abs/2606.13610) describe a March 2026 broadcast on China Central Television. It exposed paid operators seeding fake reviews online to push a fake brand into the top picks of mainstream Chinese AI assistants within hours.

Their own test rewrote real products into fake ones in the web pages an assistant reads. Across 12 commercial and open models, including GPT-5.4, Claude, Gemini, Qwen and DeepSeek, a single polluted page fooled models up to 27% of the time. Replacing the top three pages raised this to 73.8%.

Their main test used Chinese pages. When they repeated it with English pages, 8 of 12 models landed within 10 points of their Chinese results. Models resisted best in categories whose real brands they knew well, which again points to how much an assistant already knows about you.

## What should you do about it?

Treat each AI ecosystem and each language as a separate market, and measure each one directly. Our guide to [planning GEO market by market](https://underneath.agency/resources/geo-strategy-for-global-brands-across-languages) sets out the wider approach for global brands.

1. Check your visibility in the languages your buyers use, not only in English. [An English-only visibility check](https://underneath.agency/resources/english-only-ai-visibility-audits) can understate a locally strong brand.
2. If you sell in China, test Chinese assistants such as Qwen, DeepSeek and Doubao separately. Do not assume results from ChatGPT or Gemini carry over.
3. Earn coverage in respected local-language publications in each market. Translating your own site is a start; the Toronto researchers found that answers in other languages lean on local sources.
4. Watch “best X for Y” questions most closely, since that is where the Chinese-versus-Western gap was widest.
5. Monitor your category for fake or misleading pages, which can sway assistants in both ecosystems.

If you want help setting up multilingual tracking, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

No one has independently repeated a Chinese-versus-Western brand comparison across several product categories.

- The only direct comparison covers one software category and comes from authors with a stake in the result.
- It used English questions only. How Chinese assistants treat Western brands in Chinese-language questions is untested in the papers we reviewed.
- The language studies compared languages and Western-built assistants, not Chinese and Western models side by side.
- Whether bilingual content causes a brand to appear more often has not been tested; the evidence is correlation.
- AI models change often, so every figure here describes one moment in time.

## Frequently asked questions

### Do DeepSeek and Qwen recommend different brands than ChatGPT?

In one study, yes. Chinese models named the brand asked about in 88.9% of English answers against 58.3% for Western models, but the authors have a conflict of interest and studied one software category.

### Is an English-language AI visibility check enough for a global brand?

No, not if you sell in several languages. Across 66 European brands, switching to the home language raised local champions’ share of recommendations by 0.80 on a 0-to-1 scale, against 0.15 for global brands.

### Will translating my website get my brand into Chinese AI answers?

Nobody has tested that. The Toronto study suggests that answers in other languages draw on local-language sources, so earned coverage in local media probably matters more than translation alone.

### Are Chinese AI models more positive about brands?

One study says yes. When they named a brand, Chinese models scored +0.71 on a −1 to +1 tone scale against +0.42 for Western models, from the same conflicted study.

## Sources

- Huang Junyao, Situ Ruimin and Ye Renqin (2025), [Cultural Encoding in Large Language Models: The Existence Gap in AI-Mediated Brand Discovery](https://arxiv.org/abs/2601.00869), arXiv:2601.00869.
- Mahe Chen, Xiaoxuan Wang, Kaiwen Chen and Nick Koudas (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919), arXiv:2509.08919.
- Dmitrij Żatuchin (2026), [The Language Blind Spot: How Query Language and Brand Recognition Tier Shape AI-Constructed Brand Reputation Across Twelve European Languages](https://arxiv.org/abs/2606.23165), arXiv:2606.23165.
- Minghao Luo and Liang Chen (2026), [One Polluted Page Is Enough: Evaluating Web Content Pollution in LLM Recommenders](https://arxiv.org/abs/2606.13610), arXiv:2606.13610.
- Underneath (2026), [Same question, four countries: do AI recommendations change?](https://underneath.agency/research/ai-recommendations-by-country-study)

---

This is the Markdown twin of https://underneath.agency/resources/chinese-vs-western-ai-brand-visibility. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How cloud providers win enterprise deals when buyers ask AI"
description: "By being the cloud AI answers name for real architecture, cost and sovereignty questions, with proof architects can check before they commit."
canonical: "https://underneath.agency/resources/cloud-infrastructure-enterprise-demand-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How can a cloud provider win enterprise deals when buyers ask AI first?

By being the provider AI answers name when architects ask about workloads, costs, sovereignty and migration, and by making the proof behind that answer easy to check. Cloud is a concentrated market where three hyperscalers hold most of the spend, so an AI answer that reaches for the default names is a real risk for everyone else. It is also an opening: buyers now ask much more specific questions than “which cloud is best,” and specific questions have specific answers.

## The short version

1. Enterprise spending on cloud infrastructure services reached $143 billion in the second quarter of 2026, and Amazon, Microsoft and Google held 28%, 20% and 15% of it, according to [Synergy Research Group](https://www.srgresearch.com/articles/q2-cloud-market-passes-143-billion-highest-growth-rate-in-eight-years).
2. Customers are large and committed: in [Flexera’s 2026 survey](https://vmblog.com/news/flexera-finds-cloud-value-is-rising-while-ai-waste-grows/), 76% of large enterprises spent over $5 million a month on cloud, and [Canalys](https://channeldrive.in/market-research/hyperscaler-cloud-marketplace-sales-to-hit-us85-billion-by-2028-canalys/) counts more than $360 billion of multi-year commitments to the top three providers.
3. Those commitments steer software buying too: Canalys expects sales through hyperscaler marketplaces to reach $85 billion by 2028, up from $16 billion in 2023.
4. There is room for challengers: [Gartner](https://www.gartner.com/en/newsroom/press-releases/2026-02-09-gartner-says-worldwide-sovereign-cloud-iaas-spending-will-total-us-dollars-80-billion-in-2026) expects sovereign cloud spending of $80 billion in 2026 to shift 20% of current workloads from global to local providers.
5. A new kind of buyer is arriving: at Vercel, fewer than 3% of deployments were triggered by AI coding agents at the start of 2026, against more than half by June, [eCommerceNews reported](https://ecommercenews.com.au/story/vercel-unveils-agent-tools-as-deployments-turn-ai-led).

## Who signs for cloud infrastructure, and how much is one account worth?

A platform team, finance and procurement buy it together, and one customer can be worth millions a month.

The market is big and still accelerating. Synergy estimates that quarterly cloud infrastructure revenue grew 43% year on year in the second quarter of 2026, the highest rate in eight years, with the top three holding 67% of the public cloud segment. [Gartner](https://www.gartner.com/en/newsroom/press-releases/2024-11-19-gartner-forecasts-worldwide-public-cloud-end-user-spending-to-total-723-billion-dollars-in-2025) forecast $723.4 billion of public cloud spending in 2025, growth of 21.5%.

A cloud contract is rarely signed off by one person. In Flexera’s 2026 survey, 71% of organizations ran a cloud center of excellence and 63% had a FinOps team that manages cloud costs. In practice that means:

- **Architects and platform engineers** decide what is technically possible and what they will support.
- **FinOps and finance** decide what is affordable; 85% of Flexera’s respondents called managing cloud spend a key challenge.
- **Security and compliance** decide what is allowed, which increasingly includes where data may live.
- **Procurement** turns the decision into a multi-year commitment.

Most enterprises already run more than one cloud. Flexera found AWS (83%) and Azure (79%) in active use for enterprise workloads, and 73% of organizations operating hybrid environments. Gartner predicts 90% of organizations will adopt a hybrid approach through 2027. So the commercial question is less “which cloud” than “which cloud for this workload.”

What a customer is worth shows up in public filings. DigitalOcean, a challenger that serves developers and growing companies, reported $1,125 million of annual run-rate revenue in its [second quarter of 2026 results](https://www.webull.com/news/15344369267385344). Its customers spending over $1 million a year grew 73%, it signed its first nine-figure annual commitments, and its weighted average contract life rose from 1.6 years to over 3 years. A won workload grows, and the contract around it gets longer.

Cloud software vendors, the second group this article covers, sell into the same budgets. Canalys reports that enterprises increasingly spend their existing cloud commitments on third-party software bought through hyperscaler marketplaces, and that CrowdStrike and Snowflake were among the first to claim $1 billion of cumulative marketplace sales. Data platforms such as Snowflake have their own buying path, covered in [how data platforms get shortlisted](https://underneath.agency/resources/data-infrastructure-enterprise-demand-ai-search).

## Where does AI already sit in a cloud buying decision?

At the research stage, in Google’s results, and increasingly inside the coding tools that provision infrastructure.

No public survey isolates cloud buyers. The closest evidence covers business software: in [G2’s August 2025 survey](https://learn.g2.com/ai-search-surging-for-b2b-buyers) of more than 1,000 software buyers, 87% said AI chatbots were changing how they research. G2 sells visibility on its review platform, so it has an interest in that finding.

Google’s own results already answer most technology searches with AI. The B2B software and technology keywords in [our study of 1,248 US searches](https://underneath.agency/research/ai-overviews-frequency-study), the group closest to cloud platforms, showed an AI Overview on 96.0% of searches, more often than any of the other seven industries.

The newer development is that software, not only people, now makes infrastructure choices. At Vercel, a hosting platform, the share of deployments triggered by AI coding agents rose from fewer than 3% to more than half in about six months, according to eCommerceNews’s report of the company’s figures. When an agent sets up a project, the database, hosting and services it reaches for become the default. Database makers already feel this, as our guide to [how AI helps pick the database](https://underneath.agency/resources/database-companies-developers-ai-search) explains. A reasonable expectation is that the defaults agents pick will matter more for developer-led cloud products each quarter. No published study yet shows how agents choose.

## What do architects and FinOps teams ask AI assistants about cloud?

Mostly workload questions about architecture, cost, location and migration, rarely a general “which cloud” question. We wrote the sample prompts below to show the shape of those workload questions; they are not logged queries from real cloud buyers.

| Stage | Illustrative prompt |
|---|---|
| Architecture | “What is the best way to run GPU inference for a mid-sized fintech without a hyperscaler contract?” |
| Alternatives | “Alternatives to AWS for a SaaS company that wants simpler pricing and EU hosting” |
| Cost | “Is it cheaper to run Kubernetes on Azure, Google Cloud or a smaller provider for 200 nodes?” |
| Sovereignty | “Which cloud providers offer sovereign cloud regions in Germany that meet government requirements?” |
| Migration | “How hard is it to move from Heroku to a container platform, and which ones do teams choose?” |
| Marketplace | “Can I buy this observability tool through AWS Marketplace and count it against our commitment?” |

The answer to each of these is rarely written from one page. For a GPU or Kubernetes question, Google says AI Overviews and AI Mode [may use a “query fan-out” technique](https://developers.google.com/search/docs/appearance/ai-features), firing off multiple related searches across subtopics and data sources. OpenAI’s help pages say [ChatGPT search rewrites a question](https://help.openai.com/en/articles/9237897-chatgpt-search) into one or more targeted queries and passes them to its search partners. An architecture question, we infer, sets off searches for documentation, pricing, benchmarks and practitioner discussion all at once.

## How does an AI answer turn into cloud consumption and commitments?

Through two paths: a developer signup that grows into a commitment, and an enterprise shortlist that ends in a contract.

**The developer path.** An engineer or a coding agent asks how to run something, gets an answer that names a platform, signs up and starts a small workload. If it works, usage grows. DigitalOcean’s numbers show where that leads: revenue from customers spending $1 million or more a year now makes up 23% of its total and grew 214%. The AI answer sits at the very first step, where the platform is chosen.

**The enterprise path.** An architect asks a workload question, the answer names a few providers or tools, the team runs a proof of concept, and procurement folds the winner into a multi-year agreement. For cloud software vendors there is a third step: the purchase is often routed through a hyperscaler marketplace so it counts against an existing commitment. Canalys expects more than 50% of marketplace sales to flow through channel partners by 2027, so a partner may be the one asking the assistant.

On both paths the revenue rarely shows up as a click from an AI answer. It arrives later as a signup, a sales conversation or a marketplace order, which is why attribution is hard; we cover this in [what lost clicks mean for pipeline](https://underneath.agency/resources/ai-answers-pipeline-revenue).

## Why does an AI assistant name one cloud provider and skip another?

No platform documents how it picks providers; studies show strong defaults toward famous names that specific evidence can overcome.

**Documented by the platforms.** Both Google and OpenAI say their AI answers search the web and link what they used. Neither says how it decides which hyperscaler, neocloud or cloud tool ends up in the answer.

**Observed in studies.** Assistants lean toward names they know. In controlled tests by [Chu and Hou](https://arxiv.org/abs/2606.17443), three AI models always picked a real brand over invented ones when products were otherwise identical, but a rating edge of just 0.075 stars was enough for an unknown name to win half the time. In [our brand study](https://underneath.agency/research/brand-entity-ai-recommendations-study) of four assistants, a brand named on ten times as many independent sites had 4.7 times the odds of being recommended. [Our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study) caught ChatGPT searching for a named publication, ranking or award in 43.8% of its answers and for reviews in 46.2%. What that pull toward famous names means for challengers to the big three is the subject of [do AI assistants favor big brands](https://underneath.agency/resources/do-ai-assistants-favor-big-brands).

**What cloud buyers check.** Architects verify claims in public: documentation, pricing calculators, status pages, compliance attestations, region lists, benchmarks and practitioner forums. Flexera found 53% of cloud leaders named security and compliance as their top challenge for cloud-based AI work. Security vendors meet that scrutiny head on, as our guide to [cloud security shortlists in AI answers](https://underneath.agency/resources/cloud-security-enterprise-buyers-ai-search) shows.

**Our inference.** For a challenger, the advantage that wins an evaluation (a lower price for a workload, a region, a certification, a simpler pricing model) is the same thing an assistant needs to see stated plainly and confirmed elsewhere. A reasonable expectation is that a specific, verifiable difference helps a challenger more than general claims to be “developer-friendly.” Nobody has yet tested this for cloud providers.

## What does a cloud provider lose when AI answers leave it out?

A missed workload can mean a missed multi-year commitment, though no study has measured the loss directly.

- **Commitments lock in choices.** With more than $360 billion of multi-year commitments to the top three, Canalys notes that customers increasingly spend that money on software bought through those providers’ marketplaces. A vendor not considered when a workload is placed may wait years for the next opening, we infer.
- **Sovereignty is a new opening, with a deadline.** Gartner forecasts sovereign cloud spending to grow 35.6% in 2026, with Europe projected at 83% growth. Buyers asking where they can host regulated workloads are forming new shortlists now.
- **Neoclouds are moving fast.** Synergy counts nine neocloud companies, providers built for AI workloads, among the top 40 cloud providers. The challenger field is getting more crowded, not less.
- **Agents pick defaults.** If more deployments start in AI coding tools, a platform that agents do not reach for loses the first step of the developer path, our inference from Vercel’s figures.

## What does GEO look like for a cloud provider or cloud software vendor?

It makes your workload-level advantages easy for AI assistants to find, confirm and repeat. No provider, large or small, can be promised a slot in the answer.

1. **One precise identity.** State which workloads you are best for, where you host, and how you price, the same way on your site, documentation, marketplace listings, analyst profiles and partner pages. Mixed messages produce vague answers; see [how to fix wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers).
2. **Readable documentation and pricing.** Keep pricing, region lists, compliance attestations and limits on plain, crawlable pages rather than behind sign-in or in scripts an assistant cannot render. We explain the rendering problem in [when AI agents can’t read your site](https://underneath.agency/resources/when-ai-agents-cant-read-your-site).
3. **Independent proof.** Benchmarks run by others, analyst coverage, migration case studies with named customers, and conference talks by practitioners give assistants something independent to cite. Our guide to [how brands build authority for AI search](https://underneath.agency/resources/how-brands-build-authority-for-ai-search) covers how that third-party proof builds up.
4. **Honest comparison content.** Workload-by-workload comparisons and migration guides answer the questions buyers actually ask; ranked lists on other sites still carry more weight, as we show in [best-of lists](https://underneath.agency/resources/best-of-lists-ai-recommendations).
5. **Developer and agent paths.** Clear quick-starts, templates and setup instructions are what a coding agent or a developer follows, so they deserve the same care as the homepage.
6. **Workload-by-region tracking.** Ask each workload and region question repeatedly in ChatGPT, Google AI Overviews and AI Mode, Gemini, Perplexity, Copilot and Claude, and log the answers; [how many prompts to track](https://underneath.agency/resources/how-many-prompts-to-track-ai-visibility) covers how many to run.

## What remains unmeasured about AI answers and cloud workload decisions?

Nobody has measured whether appearing in AI answers moves a workload, or a commitment, toward one provider.

- **No cloud-specific buyer survey on AI use.** The surveys we found cover software buyers in general, and the platforms that publish them have commercial interests. Surveys of software buyers more broadly, and their limits, are weighed in [how B2B SaaS companies generate revenue from AI search](https://underneath.agency/resources/b2b-saas-revenue-from-ai-search).
- **Agent choices are a black box.** Vercel’s figures show agents deploying, not how they choose a provider, and one platform’s data may not generalize.
- **Market figures move quickly.** Synergy’s quarterly shares and Gartner’s forecasts are estimates and are revised often.
- **The tie to consumption revenue is unproven.** Evidence linking AI visibility to business results is the thinnest of all; [does AI visibility drive business results](https://underneath.agency/resources/does-ai-visibility-drive-business-results) sets out what exists.

## How can a cloud provider see which workloads AI answers send elsewhere?

Run the workload questions behind your recent wins and see which providers the answers name, and why.

Pick the five or six workloads where you win deals today, and the regions and compliance needs that come with them. Put those questions to ChatGPT, Gemini, Copilot, Perplexity and Google’s AI features several times each, noting who is named, which sources are cited, and whether your pricing, regions and certifications come out right. The gaps usually point to missing independent proof or unreadable documentation, not missing blog posts.

For an outside view, [talk to us about a workload-level audit](https://underneath.agency/contact). We will map where AI answers place you on the questions that lead to proofs of concept and multi-year commitments, and which gaps most likely cost you enterprise deals. Our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page sets out how the work after that audit is organized, workload by workload, from readable pricing and region pages to independent proof.

## Frequently asked questions

### Can a smaller cloud provider be named next to AWS, Azure and Google?

Yes, for specific workloads. In controlled tests a small, clear advantage beat a famous name, while generic questions favor the leaders.

### Do hyperscaler marketplace listings help AI visibility?

Untested. They matter commercially, with marketplace sales forecast to reach $85 billion by 2028, and a clear listing gives assistants one more consistent description.

### Should cloud pricing pages be public?

For the facts buyers compare, yes. Assistants can only repeat prices they can read, and FinOps teams check costs before any proof of concept.

### Does sovereign cloud demand change what buyers ask AI?

Likely. Gartner expects sovereign cloud spending to reach $80 billion in 2026, which makes location and compliance questions more common.

### Do AI coding agents count as buyers?

Increasingly they choose the first platform. At Vercel, agents triggered more than half of deployments by mid-2026.

## Sources

- Synergy Research Group (2026-07-30), [Q2 Cloud Market Passes $143 Billion; Highest Growth Rate in Eight Years](https://www.srgresearch.com/articles/q2-cloud-market-passes-143-billion-highest-growth-rate-in-eight-years)
- Gartner (2024-11-19), [Gartner Forecasts Worldwide Public Cloud End-User Spending to Total $723 Billion in 2025](https://www.gartner.com/en/newsroom/press-releases/2024-11-19-gartner-forecasts-worldwide-public-cloud-end-user-spending-to-total-723-billion-dollars-in-2025)
- Gartner (2026-02-09), [Gartner Says Worldwide Sovereign Cloud IaaS Spending Will Total $80 Billion in 2026](https://www.gartner.com/en/newsroom/press-releases/2026-02-09-gartner-says-worldwide-sovereign-cloud-iaas-spending-will-total-us-dollars-80-billion-in-2026)
- Flexera, via VMblog (2026-03-18), [Flexera Finds Cloud Value is Rising While AI Waste Grows](https://vmblog.com/news/flexera-finds-cloud-value-is-rising-while-ai-waste-grows/)
- Canalys, via Channel Drive (2024), [Hyperscaler cloud marketplace sales to hit US$85 billion by 2028](https://channeldrive.in/market-research/hyperscaler-cloud-marketplace-sales-to-hit-us85-billion-by-2028-canalys/)
- DigitalOcean, via Webull (2026-08-05), [DigitalOcean Announces Second Quarter 2026 Financial Results](https://www.webull.com/news/15344369267385344)
- eCommerceNews Australia (2026-06), [Vercel unveils agent tools as deployments turn AI-led](https://ecommercenews.com.au/story/vercel-unveils-agent-tools-as-deployments-turn-ai-led)
- G2 (2025-08), [How AI Chat is Rewriting B2B Software Buying](https://learn.g2.com/ai-search-surging-for-b2b-buyers)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- OpenAI Help Center (2025), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Chu and Hou (2026), [Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems](https://arxiv.org/abs/2606.17443), arXiv:2606.17443.
- Underneath (2026), [When does Google show an AI Overview? 1,248 US searches](https://underneath.agency/research/ai-overviews-frequency-study)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [Do Wikipedia and schema make AI assistants recommend a brand?](https://underneath.agency/research/brand-entity-ai-recommendations-study)

---

This is the Markdown twin of https://underneath.agency/resources/cloud-infrastructure-enterprise-demand-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How cloud security vendors reach enterprise buyers in AI search"
description: "Cloud security vendors win AI shortlists with proof assistants can check: clear category fit, cloud coverage, marketplace listings and third-party evidence."
canonical: "https://underneath.agency/resources/cloud-security-enterprise-buyers-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How do cloud security vendors get onto enterprise shortlists when buyers ask AI first?

By making it easy for an AI answer, and the buyer who checks it, to place you in the right category with verifiable proof. Cloud security is the fastest-growing part of the security market, and buyers are consolidating onto fewer platforms. That makes each AI-assisted shortlist worth more and harder to rejoin once it is set.

## The short version

1. Cloud security is the fastest-growing slice of security spending: Gartner forecasts it at 28.8% growth in 2026, from $13.0 billion in 2025 to $17.1 billion, with posture management alone growing 33.4%, according to an analysis of the forecast by [Louis Columbus](https://softwarestrategiesblog.com/2026/04/01/top-10-fastest-growing-security-categories-gartner-2026-forecast/).
2. Buyers are moving to platforms: in [Fortinet’s 2025 State of Cloud Security Report](https://www.voicendata.com/research/fortinet-report-highlights-cloud-security-strategies-for-2025-8643178), 67% were implementing posture management, 62% were adopting cloud-native application protection platforms (CNAPP), and 97% preferred one unified platform.
3. A large share of cloud software is now bought through the cloud providers’ own stores: [Canalys](https://channeldrive.in/market-research/hyperscaler-cloud-marketplace-sales-to-hit-us85-billion-by-2028-canalys/) projects marketplace sales of $85 billion by 2028, up from $16 billion in 2023.
4. Google search already answers this category with AI: in [our study of 1,248 searches](https://underneath.agency/research/ai-overviews-frequency-study), 96.0% of B2B software and technology keywords showed an AI Overview, the highest of eight industries.
5. Category leadership is valued at a premium: Google closed its [$32 billion purchase of Wiz](https://www.clearygottlieb.com/news-and-insights/news-listing/google-completes-32-billion-acquisition-of-wiz) in March 2026, about two years after Wiz reported [$500 million in annual recurring revenue](https://techcrunch.com/2024/10/23/wiz-hopes-to-hit-1b-in-arr-in-2025-before-an-ipo-after-turning-down-googles-23b).

## Who signs for cloud security, and how much is a won account worth?

A CISO signs, but cloud platform, DevOps and compliance teams shape the choice; won accounts expand for years.

Cloud security covers several product types that buyers increasingly see as one decision: posture management (CSPM), which finds misconfigurations; workload protection (CWPP), which protects running servers, containers and serverless code; cloud access security brokers (CASB); and CNAPP, the platforms that bundle them. Gartner’s 1Q26 forecast, as summarized by Columbus, puts workload protection at $6.0 billion in 2025 and posture management at $4.7 billion, the two largest cloud categories.

The buyer’s environment explains why the evaluation is technical and crowded with stakeholders. In the Fortinet survey, produced by Cybersecurity Insiders:

- over 78% used two or more cloud providers, and 54% ran hybrid models that mix on-premises and public cloud;
- 61% named security and compliance as the main barrier to cloud adoption;
- 76% reported a shortage of cloud security expertise;
- 64% lacked confidence in their ability to handle real-time threat detection.

So the buyer is short of people, running several clouds, and answerable to auditors. Those constraints become the questions they ask: which tool covers AWS, Azure and Google Cloud equally, which needs no agents, which maps findings to the frameworks their auditors use.

The stakes for the buyer are concrete. IBM’s [Cost of a Data Breach Report 2025](https://www.bakerdonelson.com/webfiles/Publications/20250822_Cost-of-a-Data-Breach-Report-2025.pdf) found that 30% of breaches involved data spread across multiple environments, which cost an average of $5.05 million and took 276 days to identify and contain, the longest of any storage location.

For the vendor, a won customer is a platform relationship. Consolidation means one contract can grow from posture management into workload, identity and data modules. The market has priced that potential highly: Wiz said it had reached $500 million in annual recurring revenue in October 2024, and Google paid $32 billion for the company, a deal that closed on March 11, 2026.

## Where does AI search already sit in a cloud security evaluation?

Early, at research and shortlisting, though no public study isolates cloud security buyers.

The nearest data comes from technology buyers as a whole. [TrustRadius’s 2026 B2B Buying Disconnect Report](https://www.demandgenreport.com/industry-news/news-brief/trustradius-ai-has-changed-how-buyers-research-but-not-what-they-trust/53853/), a survey of nearly 2,500 technology buyers and vendors, found:

- 63% of buyers used AI during their purchase journey;
- 94% of those buyers fact-checked its answers at least some of the time;
- 83% shortlisted three or fewer products;
- 74% used reviews to inform their decision, while analyst reports were used by only 13%.

Short shortlists are the key number for cloud security vendors. If an AI answer helps decide which three platforms get a proof of value, a fourth vendor may never be tested. TrustRadius earns money selling review visibility to software vendors, so weigh its figures with that in mind.

Google’s results point the same way. In [our study of when Google shows an AI Overview](https://underneath.agency/research/ai-overviews-frequency-study), B2B software and technology keywords had the highest rate of the eight industries we tested: 96.0% as observed and 83.0% after adjusting for the kind of searches each industry had. Longer, more specific wording raised the rate for the same topic, from 59.4% as written to 87.5% in a long form. Cloud security questions tend to be long and specific, so we infer that many of them will meet an AI answer before any vendor’s page.

## What do cloud architects type into an assistant when weighing CNAPP or CSPM tools?

Questions about category, cloud coverage, deployment, compliance and consolidation. We wrote the example prompts below to reflect a multi-cloud security team’s concerns; none were taken from real buyer sessions.

| Buying need | Illustrative prompt |
|---|---|
| Category | “Do we need a CSPM tool or a full CNAPP for a 200-account AWS estate?” |
| Coverage | “Which cloud security platforms cover AWS, Azure and Google Cloud with the same depth?” |
| Deployment | “Agentless or agent-based workload protection for Kubernetes: what are the trade-offs?” |
| Alternatives | “What are the alternatives to Wiz now that Google owns it?” |
| Comparison | “Microsoft Defender for Cloud or Palo Alto Networks Cortex Cloud for a mostly Azure company?” |
| Compliance | “Which CNAPP vendors are FedRAMP authorized?” |
| Buying route | “Which cloud security tools can we buy through AWS Marketplace with our committed spend?” |

Each of these questions can trigger several searches. Google says that behind AI Overviews and AI Mode it [may use a “query fan-out” technique](https://developers.google.com/search/docs/appearance/ai-features), issuing “multiple related searches across subtopics and data sources.” OpenAI says [ChatGPT search typically rewrites](https://help.openai.com/en/articles/9237897-chatgpt-search) a question “into one or more targeted queries” sent to search providers. A vendor whose coverage, deployment model and certifications are spelled out on public pages gives those searches something to find, we infer.

## How does an AI answer lead to a signed cloud security platform deal?

Through a short path: AI answer, shortlist, proof of value, then a marketplace purchase that expands across clouds and modules.

1. A cloud architect or security leader asks an assistant a category or comparison question.
2. The answer names a few platforms and links sources; the buyer checks them, since 94% of AI-using technology buyers fact-check.
3. Two or three vendors connect to a test cloud account for a proof of value. Agentless products are built to connect quickly, which shortens this step.
4. The winner is bought, increasingly through a cloud provider’s marketplace.
5. The contract grows as the customer adds accounts, clouds and modules.

Step 4 is particular to cloud security. Canalys reports that enterprise customers have committed to spend over $360 billion on the top three cloud providers’ services on a multi-year basis, and that buyers are using marketplaces to “burn down a portion of their cloud credits on third-party software.” Canalys names CrowdStrike among the first vendors to publicly claim $1 billion of total sales through marketplaces. For a buyer, a marketplace listing means no new budget line and faster procurement. For a vendor, it shortens the gap between an AI-assisted shortlist and a signed order. The providers themselves compete for those commitments, as our guide to [how cloud providers win enterprise deals](https://underneath.agency/resources/cloud-infrastructure-enterprise-demand-ai-search) describes.

Step 5 is why each shortlist matters so much. With 97% of Fortinet’s respondents preferring a unified platform, the first module bought is often the foothold for the rest, we infer.

## Why do assistants name some cloud security platforms and skip others?

The platforms don’t say how; studies of brands across assistants point to independent coverage and prominence.

**Documented by the platforms.** Both Google and OpenAI describe AI answers that look things up on the web and link the pages behind them. Neither says how one posture management or workload protection vendor gets chosen over another.

**Observed in studies.** In [our brand study](https://underneath.agency/research/brand-entity-ai-recommendations-study) across four assistants, coverage on independent sites was the strongest predictor we measured: each tenfold increase in the number of independent sites naming a brand went with 4.7 times the odds of being recommended. In a vendor dataset analyzed by [Kumar](https://arxiv.org/abs/2606.20065), global household-name brands appeared in 73% of unbranded answers on the first day and the least-known tier in 11%. The big-brand head start is examined in [do AI assistants favor big brands](https://underneath.agency/resources/do-ai-assistants-favor-big-brands).

That matters in cloud security, where the largest platform vendors, the cloud providers’ own tools and a few heavily funded specialists dominate coverage. A smaller vendor competes on specifics, not on fame. Endpoint vendors face a similar contest, covered in [how EDR vendors win enterprise deals](https://underneath.agency/resources/endpoint-security-enterprise-customers-ai-search).

**Our inference on what a cloud buyer checks.** Most of what a multi-cloud security team verifies sits on public pages that an assistant can read as well:

- which clouds, regions and services are supported, and with what depth;
- deployment model: agentless scanning, agents, or both;
- authorizations and attestations such as FedRAMP, SOC 2, ISO 27001 and Cloud Security Alliance STAR entries;
- marketplace listings and cloud-provider partner designations;
- original research on cloud attacks and misconfigurations that journalists and practitioners cite;
- reviews from practitioners who run the product in production.

Vendors that keep these cloud facts public, current and identical across their site, marketplace listings and review profiles plausibly give assistants better material to repeat. That idea has not been tested for cloud security.

## What does a cloud security vendor lose when AI answers leave it out?

Mostly platform decisions, which lock in for a contract term; no study has measured the dollar cost.

- **Shortlists are small.** With 83% of technology buyers shortlisting three or fewer products, a vendor left out of the first answer competes for one of very few remaining places.
- **Consolidation raises the stakes.** When buyers want one platform, losing the first evaluation can mean losing every module that would have followed, until the next renewal, we infer.
- **Marketplace buying speeds up the winner.** Committed cloud spend lets the chosen vendor close quickly, which leaves less time for a missing vendor to enter late.
- **Wrong facts are a cost too.** If an assistant says your product needs agents when it does not, or omits a cloud you support, buyers may rule you out. Our guide to [fixing wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) covers what to do.

## What does GEO look like for a CNAPP or CSPM vendor?

Publishing your category fit, cloud coverage and compliance proof where assistants can find and repeat it. Nothing guarantees a recommendation.

1. **One clear category statement.** Say plainly whether you are a CSPM, CWPP, CNAPP or a point tool, which clouds you cover and for whom. Use the same wording on your site, documentation, marketplace listings, review profiles and analyst briefings.
2. **Public technical facts.** Publish supported services, deployment options, data residency and integration details as readable web pages, not only in gated PDFs or behind a demo.
3. **Compliance pages that answer the question.** Map controls to the frameworks buyers name, and show authorization status and attestation dates where buyers can verify them.
4. **Marketplace and partner presence.** Keep cloud marketplace listings complete and consistent with your site, since buyers use them and assistants may find them.
5. **Independent coverage.** Publish original cloud threat research, brief analysts, and earn coverage in security publications. See [how brands build authority for AI search](https://underneath.agency/resources/how-brands-build-authority-for-ai-search) and [which pages to target](https://underneath.agency/resources/best-of-lists-ai-recommendations).
6. **Honest comparison content.** Explain how your approach differs from the cloud providers’ native tools and from larger platforms; our article on [whether comparison pages help B2B citations](https://underneath.agency/resources/do-comparison-pages-help-b2b-ai-citations) sets expectations.
7. **Measurement across assistants.** Put real CSPM, CNAPP and compliance questions to ChatGPT, Google AI Overviews and AI Mode, Gemini, Perplexity, Copilot and Claude, repeating each one, and log which platforms are named.

For the wider security picture, including how CISOs judge trust, see our article on [cybersecurity software and AI search](https://underneath.agency/resources/cybersecurity-software-revenue-from-ai-search). Identity vendors have their own guide on [how AI answers name identity platforms](https://underneath.agency/resources/iam-enterprise-demand-ai-search).

## What remains unproven about AI search in cloud security buying?

Nobody has measured how often cloud security buyers consult assistants, or whether AI visibility lifts revenue.

- **Cloud security buyers are not studied on their own.** TrustRadius’s figures describe technology buyers as a whole. We found no public, vendor-neutral survey of how cloud security buyers use AI assistants.
- **Survey sponsors have interests.** Fortinet sells cloud security products and TrustRadius sells review visibility. The Gartner figures reach us through a secondary analysis, not a Gartner press release.
- **The choice of names is a black box.** Only the platforms know how an assistant picks which cloud security vendors to list, and the list shifts between runs and between assistants.
- **Revenue links are unproven.** Whether appearing in AI answers changes win rates or deal size has not been measured publicly; see [does AI visibility drive business results](https://underneath.agency/resources/does-ai-visibility-drive-business-results).

## How can a cloud security vendor check whether AI answers put it into proofs of value?

Check how AI answers describe your coverage, deployment and compliance on the questions your buyers ask.

List the category, coverage, comparison, compliance and buying-route questions your buyers ask. Repeat each one in every major assistant, since the named platforms shift from run to run. Record whether you are named, which sources are cited, and whether the facts about your clouds, agents and authorizations are right. Where you go unnamed or described wrongly, the usual cause is a technical detail that is not public, or too little independent coverage.

For a second pair of eyes, [ask us to audit your place on cloud security shortlists](https://underneath.agency/contact). We will show which enterprise shortlists AI answers put you on or leave you off, and which gaps most likely cost you proof-of-value invitations and platform deals. The steps that usually follow, such as a clear category statement, public facts on which clouds you cover and consistent marketplace listings, are laid out on our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page.

## Frequently asked questions

### Do cloud security buyers really use AI assistants to choose vendors?

Technology buyers do: 63% used AI in their purchase journey in TrustRadius’s 2026 survey. No public study isolates cloud security buyers.

### Does a cloud marketplace listing help AI visibility?

Untested. It clearly helps buying: Canalys projects $85 billion of marketplace sales by 2028. A consistent listing also gives assistants one more accurate description, we infer.

### Should we write pages comparing ourselves with the cloud providers’ native tools?

Yes, as long as the comparison is fair and names specific differences in coverage and deployment. Buyers often weigh native tools first, and independent sources will still carry more weight than your own page.

### Can a smaller cloud security vendor compete with the big platforms in AI answers?

It can, but it starts behind: well-known brands appeared in 73% of unbranded answers in one dataset, the least-known in 11%. Independent coverage is the strongest lever we have measured.

## Sources

- Louis Columbus, Software Strategies Blog (2026-04-01), [Gartner’s $246.2B Security Forecast shows 10 categories growing 2x to 3x the market](https://softwarestrategiesblog.com/2026/04/01/top-10-fastest-growing-security-categories-gartner-2026-forecast/), summarizing Gartner’s Information Security Market Current Outlook, 1Q26
- Fortinet and Cybersecurity Insiders, via Voice&Data (2025-01-22), [Fortinet report highlights cloud security strategies for 2025](https://www.voicendata.com/research/fortinet-report-highlights-cloud-security-strategies-for-2025-8643178)
- Canalys, via Channel Drive (2024-08-16), [Hyperscaler cloud marketplace sales to hit US$85 billion by 2028](https://channeldrive.in/market-research/hyperscaler-cloud-marketplace-sales-to-hit-us85-billion-by-2028-canalys/)
- Cleary Gottlieb (2026-03-11), [Google Completes $32 Billion Acquisition of Wiz](https://www.clearygottlieb.com/news-and-insights/news-listing/google-completes-32-billion-acquisition-of-wiz)
- TechCrunch, Julie Bort (2024-10-23), [Wiz hopes to hit $1B in ARR in 2025 before an IPO](https://techcrunch.com/2024/10/23/wiz-hopes-to-hit-1b-in-arr-in-2025-before-an-ipo-after-turning-down-googles-23b)
- IBM and Ponemon Institute (2025-07), [Cost of a Data Breach Report 2025](https://www.bakerdonelson.com/webfiles/Publications/20250822_Cost-of-a-Data-Breach-Report-2025.pdf)
- Demand Gen Report, James Hickey (2026-07-30), [TrustRadius: AI Has Changed How Buyers Research, But Not What They Trust](https://www.demandgenreport.com/industry-news/news-brief/trustradius-ai-has-changed-how-buyers-research-but-not-what-they-trust/53853/)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- OpenAI Help Center (2025), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Kumar (2026), [Generative Engine Optimization at Scale: Measuring Brand Visibility Across AI Search Engines](https://arxiv.org/abs/2606.20065), arXiv:2606.20065
- Underneath (2026), [When does Google show an AI Overview? 1,248 US searches](https://underneath.agency/research/ai-overviews-frequency-study)
- Underneath (2026), [Do Wikipedia and schema make AI assistants recommend a brand?](https://underneath.agency/research/brand-entity-ai-recommendations-study)

---

This is the Markdown twin of https://underneath.agency/resources/cloud-security-enterprise-buyers-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How CNC machine shops win RFQs through AI search"
description: "Yes, if assistants can verify your machines, materials, tolerances, certifications and location; instant-quote platforms are already built for that."
canonical: "https://underneath.agency/resources/cnc-machine-shops-rfqs-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Can AI search bring our CNC shop the RFQs that now go to instant-quote platforms?

It can, but only for jobs where an assistant can confirm that your shop holds the tolerance, runs the material, carries the certification and sits in the right place. The instant-quote platforms already publish those facts in a form machines can read; most job shops still keep them in a brochure or in the owner’s head.

This guide is for owners and sales leads of precision CNC machining shops: milling, turning, 5-axis and Swiss work for aerospace, defense, medical, semiconductor and industrial customers. Our broader article on how contract manufacturers win RFQs through AI search covers reshoring and general supplier discovery. This one stays with the machine shop and its specific rival: the online platform that quotes a part in minutes.

## The short version

1. The platforms are growing fast: in the second quarter of 2026, [Xometry](https://www.streetinsider.com/Press+Releases/Xometry+Reports+Record+Second+Quarter+2026+Results/26861407.html) grew marketplace revenue 45% to $215 million and active buyers 20% to 89,557, and [Protolabs](https://www.webull.com/news/15321326548034560) grew CNC machining revenue 13.6%.
2. Machine shops are investing, but lagging: the [AMT’s order report](https://www.automation.com/article/583.4-million-usd-new-machinery-orders-us-economic-strengths) put US metalworking machinery orders at $2.77 billion in the first five months of 2026, up 31.9%, while orders from contract machine shops ran nearly 10% below their recent average in May.
3. Buyers find sourcing painful: in the 2026 [Fictiv and MISUMI survey](https://www.digitalengineering247.com/article/fictiv-releases-annual-state-of-manufacturing-report/Additive-Manufacturing) of more than 300 manufacturing leaders, 81% said supplier sourcing is too time-consuming and costly, up from 73% in 2025.
4. Engineers are starting to use AI, cautiously: in the 2026 [State of Marketing to Engineers](https://advertising.globalspec.com/state-of-marketing-to-engineers-research-report) survey, 21% routinely used generative AI platforms in purchase research, up 8 points, and they rated their trust in AI answers 4.7 out of 10.
5. Defense work now has a new gate: the CMMC cybersecurity rule took effect on November 10, 2025, and from November 10, 2026 applicable defense contracts can require a third-party Level 2 assessment.

## Who buys CNC machining, and what is one account worth?

Engineers and sourcing managers at manufacturers, and a good account is worth years of repeat part orders.

The person who starts the search is usually a mechanical or manufacturing engineer with a drawing and a deadline. A buyer or sourcing manager then owns the supplier relationship, the quote comparison and the purchase order. In aerospace, defense and medical work, a quality engineer also checks the shop’s certifications before the first article ships. Chemical buyers run a similar document check, described in [how formulators find chemical suppliers](https://underneath.agency/resources/chemical-suppliers-b2b-buyers-ai-search).

The industry is investing. According to the US Manufacturing Technology Orders report from AMT, the Association For Manufacturing Technology, metalworking machinery orders reached $583.4 million in May 2026, 47.8% more than a year earlier. AMT calls contract machine shops “the largest customer of manufacturing technology by industry,” but notes their orders have been falling behind total orders since December 2025, with much of the growth coming from aerospace. The builders selling those machines face their own AI shortlist, covered in [how machinery makers get shortlisted](https://underneath.agency/resources/machinery-companies-buyers-ai-search).

There is no public benchmark for what an average job shop account is worth. The platforms show the shape of the value instead. At Xometry, accounts spending at least $50,000 over twelve months grew 23%, from 1,653 to 2,039, in the year to June 2026. Our inference: a shop’s revenue rests on a few dozen repeat accounts, so one well-matched RFQ from a new program can matter more than a month of website traffic.

## Where do instant-quote platforms and AI search meet the buyer?

At the start of a job, when the engineer wants a fast price or a short list of capable shops.

Instant-quote platforms answer the engineer’s first question, price and lead time, without a phone call. Xometry describes itself as an “AI-native marketplace” and, in its second-quarter 2026 report, said it launched a process recommender that suggests the best of 20 manufacturing techniques when a buyer uploads a part. It also scores each job against a supplier’s “machine characteristics, quality history and on-time shipping record.” In other words, the platform is doing the matching a buyer used to do by phone.

Buyers like that convenience. In the Fictiv and MISUMI survey, 97% of leaders called digital manufacturing platforms essential for production, up from 86% in 2024. Fictiv is itself a manufacturing platform, so read that figure as a signal from an interested party, not a neutral count.

General AI assistants are a second, newer route, and the engineers who send RFQs are only part way there. In the TREW Marketing and GlobalSpec survey of more than 1,000 engineers and technical buyers, 69% use generative AI somewhere in the purchasing process, but the share who make it a routine part of purchase research is 21%, up 8 points. About 31% routinely use industry directory websites, and 77% still use a search engine more often than an AI platform. Engineers do 62% of the buying process online before contacting a vendor, and for 26% the first contact with sales is driven by pricing or inventory questions.

Two points stand out for a job shop. Search engines still dominate, so the pages that rank for capability searches are also what many assistants read. And the first reason engineers contact a vendor is price and availability, which is exactly what instant-quote platforms supply without contact.

Directories are adding AI too. In January 2026, Thomas launched an AI search that lets buyers run multi-attribute searches in natural language; [Digital Commerce 360 reported](https://www.digitalcommerce360.com/2026/01/19/thomas-ai-search-performance-based-ads-industrial-sourcing/) that in testing it drove more than 15% more supplier evaluations than the old search.

## What do engineers ask when they look for a machine shop?

Questions that stack a process, a material, a tolerance, a certification, a quantity and a place.

The sample questions below are ours, built to show how a drawing turns into a search. They are not captured from real buyers or from AI logs.

| What the buyer needs | Example question (illustrative) |
|---|---|
| Aerospace capability | “AS9100 machine shops with 5-axis capacity for aluminum housings near Wichita” |
| Defense eligibility | “ITAR-registered CNC shops that can machine titanium brackets, 50 pieces” |
| Medical precision | “Swiss turning shops holding half a thousandth on 316 stainless pins, ISO 13485” |
| Cybersecurity status | “CMMC Level 2 machine shops in Ohio for defense parts” |
| Platform or local shop | “Should I use an online quoting service or a local machine shop for 200 production parts?” |
| Speed | “Machine shops that can deliver prototype parts in five days” |

Each filter removes most of the market. Our own research suggests the details matter to the answer, not just to the buyer. In [our prompt phrasing study](https://underneath.agency/research/ai-prompt-phrasing-study), adding “on a tight budget” to a question kept the original first brand only 15.3% of the time, against 68.0% when the same question was simply asked again. That study covered consumer and business questions, not machining, so applying it here is our inference: a tolerance, material or certification in the question will likely reshuffle which shops are named.

Location matters in a documented way. [OpenAI says](https://help.openai.com/en/articles/9237897-chatgpt-search) ChatGPT search rewrites a question into targeted searches and may use the general location from a user’s IP address to localize them. In [our study of four assistants](https://underneath.agency/research/ai-assistants-brand-agreement-study), recommendations for questions that named a place overlapped between assistants far less than national ones, 0.160 against 0.390 on a scale where 1 is identical. Local shop answers are less settled, which means a well-documented regional shop has room to appear.

## How does a shop named by an assistant end up with the purchase order?

Through a shortlist, a capability check, an RFQ, a first article and then repeat orders.

1. **Named or not.** An engineer asks an assistant, a directory or a platform who can make the part. A shop that is not named is not on the list.
2. **Checked.** The engineer opens the shop’s site to confirm machines, materials, tolerances and certifications. Thin pages end the visit here.
3. **Asked to quote.** The RFQ goes to two or three shops, often alongside an instant quote from a platform as a price reference.
4. **First article.** The winning shop ships first-article parts with inspection reports. Regulated programs add a quality audit.
5. **Repeat orders.** A shop that passes keeps the part number for revisions and production. This is where the value of the first mention shows up.

Showing up in an assistant’s answer gets a shop through the shortlist and the capability check, and no further. Price, lead time, quality and responsiveness still win the job. What visibility changes is the mix of RFQs: a shop described clearly for its hardest work is more likely to be asked about that work, not just the commodity parts where platforms win on speed. For the wider supplier search, including reshoring, read [how contract manufacturers win RFQs](https://underneath.agency/resources/contract-manufacturers-rfqs-from-ai-search).

## What makes an assistant name one machine shop over another?

Facts it can find and check outside your own claims; the platforms document how they search, not how they pick.

Documented by the platforms: ChatGPT search sends rewritten queries to search providers, and a site must allow OpenAI’s search crawler, OAI-SearchBot, to be eligible for inclusion. [Google explains](https://blog.google/products/search/ai-mode-search/) that AI Mode splits a question into subtopics and runs many related searches, a method it calls “query fan-out,” so a single machining question can become separate searches for material, tolerance and location. Neither company publishes how it chooses a supplier.

Observed in studies:

- **Facts drift where sources disagree.** In [our business facts study](https://underneath.agency/research/ai-business-facts-accuracy-study), 18.9% of AI answers about local businesses stated at least one address, phone, website or hours fact that differed from the Google profile. A shop with an old address on one directory and a new one on its site invites that mistake.
- **Engineers verify.** With trust in AI answers at 4.7 out of 10, the TREW and GlobalSpec data suggest an assistant’s mention starts a check, not a purchase.
- **Familiarity breaks ties.** In the same survey, 70% said they were likely to choose the better-known brand when two solutions look technically similar.

Our inference for machine shops: the trust signals are the ones a quality engineer already asks for, published where software can read them. That means certificates with scope and registrar, a real equipment list, materials and tolerance ranges, part size limits, quantities, inspection equipment, and the industries you serve. For defense work, it also means stating your export-control and cybersecurity status accurately. Under the [International Traffic in Arms Regulations](https://www.ecfr.gov/api/renderer/v1/content/enhanced/current/title-22?part=122), anyone in the US manufacturing defense articles must register with the State Department even if it never exports, and the first-tier fee is $3,000 a year. Under the CMMC rule, explained by [Squire Patton Boggs](https://www.squirepattonboggs.com/insights/publications/the-cmmc-dfars-final-rule-goes-live-ready-or-not-here-it-comes/), primes must flow requirements down to subcontractors that handle controlled defense information, and status is recorded in the government’s Supplier Performance Risk System. Claiming a status you do not hold is a compliance problem, not a marketing one.

## What does a machine shop lose when assistants leave it out?

It loses the RFQs that never arrive, and it cedes the easy first quote to the platforms.

We have not found any measurement of RFQs lost to AI absence, so the reasoning below is labeled:

- **The platforms are taking share.** Xometry’s chief executive said the company aims “to rapidly penetrate the vast, fragmented offline market.” Its buyer count and its large accounts both grew faster than the machine shop order trend AMT reports. We infer that part of that growth comes from work that once went to independent shops by phone and referral.
- **Buyers want less effort.** With 81% of manufacturing leaders saying sourcing takes too much time and money, the supplier that is easiest to verify has an advantage before quoting starts. [Packaging suppliers chasing quote requests](https://underneath.agency/resources/packaging-suppliers-quote-requests-ai-search) face the same test.
- **Absence is silent.** An engineer who builds a shortlist from an assistant and three websites never calls the shop that was missing. No CRM records the loss. To see how fewer search clicks show up in pipeline, read [what lost clicks to AI answers mean for revenue](https://underneath.agency/resources/ai-answers-pipeline-revenue).

There is also a dependence question. Shops that take overflow work as platform suppliers gain volume but not the customer relationship. Our article on [whether AI search reduces dependence on marketplaces](https://underneath.agency/resources/ai-search-marketplace-dependence) looks at that trade-off in another industry; we infer the same tension applies to machining.

## How does GEO work for a CNC machine shop?

Generative engine optimization (GEO) makes your shop easy for AI assistants to find, describe correctly and verify.

For a machine shop, the work usually covers:

1. **Capability pages in plain text.** A page for each process (milling, turning, 5-axis, Swiss, EDM, grinding) listing machines, axis counts, travel limits, materials, typical tolerances, quantities, secondary operations and lead-time ranges.
2. **Certifications with proof.** ISO 9001, AS9100, ISO 13485, ITAR registration and CMMC status stated exactly, with scope and dates, and matching what registrars and customers can look up.
3. **One consistent identity.** The same company name, address, phone and capability list on your site, Thomasnet, LinkedIn, Google Business Profile, association directories and customer supplier portals. Our guide to [fixing wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) shows how errors spread.
4. **Independent coverage.** Features in trade publications, a profile in association directories, published case studies with customer permission, and talks at shows such as IMTS. Engineers rank online technical publications as their top research source, and assistants read the same pages. A 30-person shop can still be named; our piece on [how small brands get recommended by AI](https://underneath.agency/resources/how-small-brands-get-recommended-by-ai) explains what makes that happen.
5. **An honest answer to the platform question.** A page that explains when a local shop beats an online quote, for example on tight tolerances, complex setups, inspection documentation or ongoing engineering support, and when it does not.
6. **Crawl access.** Allow the search crawlers the assistants document, and keep capability pages outside logins and PDF-only downloads.
7. **Measurement.** Ask a fixed set of process, material, tolerance, certification and location questions across ChatGPT, Gemini, Perplexity, Copilot and Google’s AI features, repeatedly, and log who is named and which pages are cited. Then compare with RFQs received.

No one can promise that an assistant will name your shop. The aim is to make your shop the easiest one to confirm when it is a real fit.

## What can’t the data tell a machine shop owner yet?

How many RFQs AI answers create, and how assistants treat machining questions specifically.

- **Usage is not attribution.** The engineer survey measures research habits. We found no public data linking AI mentions to RFQs or purchase orders.
- **Interested sources.** TREW sells marketing to technical firms, GlobalSpec sells advertising, Fictiv and Xometry run manufacturing platforms, and Protolabs competes with job shops. Their numbers are useful but not neutral.
- **No machining-specific answer studies.** Our studies tested buyer questions across several industries, not machine shop searches. Applying their findings to CNC work is our inference.
- **Account values are private.** Platform reports give thresholds such as $50,000 a year, not typical job shop contract sizes.

## Where should a CNC machine shop start?

Ask assistants for shops that can do your best work, and check whether you are named and described correctly.

Pick the ten jobs you most want more of, phrase them the way an engineer would, with material, tolerance, certification and region, and run them across the main assistants. Note whether you appear, what they say about your capabilities, and which shops and platforms show up instead. That usually shows which facts are missing from your site and which outside sources the answers lean on.

If you would rather quote fewer commodity parts and more of the tight-tolerance work your machines were bought for, [ask us to check how assistants describe your shop](https://underneath.agency/contact). We run the process, material and certification questions your buyers ask, show which shops and platforms are named in your place, and list the capability facts that are missing or wrong. To see how fixing those facts is handled, from process pages to certification proof and directory cleanup, read about our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## Frequently asked questions

### Will instant-quote platforms replace local machine shops?

The data does not show that. Platforms are growing fast, and Xometry also routes work to partner shops. Complex, tight-tolerance and documented regulated work still depends on a shop’s engineering and quality systems.

### Should our shop list itself on Xometry or Thomasnet?

That is a business decision about margin and customer ownership. For AI visibility, complete and consistent directory profiles help because they are outside sources that confirm your capabilities and location.

### Does our ITAR registration or CMMC status help in AI answers?

We have no evidence that it changes rankings. It matters because defense buyers filter on it. State it accurately and only if it is current.

### Do engineers trust what ChatGPT says about suppliers?

Only partly. They rated trust in AI answers 4.7 out of 10 in the 2026 survey, so most will check your website before sending an RFQ.

## Sources

- Xometry via StreetInsider (2026-08-04), [Xometry Reports Record Second Quarter 2026 Results](https://www.streetinsider.com/Press+Releases/Xometry+Reports+Record+Second+Quarter+2026+Results/26861407.html)
- Protolabs via Business Wire (2026-07-31), [Protolabs Reports Financial Results for the Second Quarter of 2026](https://www.webull.com/news/15321326548034560)
- AMT via Automation.com (2026-07-13), [$583.4 Million in New Machinery Orders Highlight US Economic Strengths](https://www.automation.com/article/583.4-million-usd-new-machinery-orders-us-economic-strengths)
- Digital Engineering 24/7 (2026), [Fictiv Releases Annual State of Manufacturing Report](https://www.digitalengineering247.com/article/fictiv-releases-annual-state-of-manufacturing-report/Additive-Manufacturing)
- TREW Marketing and GlobalSpec (2026), [State of Marketing to Engineers research report](https://advertising.globalspec.com/state-of-marketing-to-engineers-research-report)
- Digital Commerce 360 (2026-01-19), [Thomas adds AI search and performance-based ads for industrial sourcing](https://www.digitalcommerce360.com/2026/01/19/thomas-ai-search-performance-based-ads-industrial-sourcing/)
- Electronic Code of Federal Regulations (2026), [22 CFR Part 122: Registration of Manufacturers and Exporters](https://www.ecfr.gov/api/renderer/v1/content/enhanced/current/title-22?part=122)
- Squire Patton Boggs (2025-09), [The CMMC DFARS Final Rule Goes Live: Ready or Not, Here It Comes](https://www.squirepattonboggs.com/insights/publications/the-cmmc-dfars-final-rule-goes-live-ready-or-not-here-it-comes/)
- OpenAI Help Center (2026), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Google (2025-03-05), [Expanding AI Overviews and introducing AI Mode](https://blog.google/products/search/ai-mode-search/)
- Underneath (2026), [Does rewording a question change AI brand recommendations?](https://underneath.agency/research/ai-prompt-phrasing-study)
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude agree on brands?](https://underneath.agency/research/ai-assistants-brand-agreement-study)
- Underneath (2026), [Do AI answers match a business’s Google profile?](https://underneath.agency/research/ai-business-facts-accuracy-study)

---

This is the Markdown twin of https://underneath.agency/resources/cnc-machine-shops-rfqs-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How collaboration software wins teams that ask AI first"
description: "By being named when teams ask AI which tool to try, then turning those free signups into seat growth and company-wide contracts that renew."
canonical: "https://underneath.agency/resources/collaboration-software-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How do collaboration software companies win teams that ask AI first?

By being the tool an AI assistant names when a team lead asks what to try, then turning that free signup into seats, paid plans and, eventually, a company-wide contract. The category already has its first public warning: monday.com told investors in 2025 that Google’s AI features were shrinking the small-business traffic its self-serve funnel runs on. Canva, on the other side, says ChatGPT has become one of its larger sources of visits.

## The short version

1. Bank of America’s analysis of [monday.com’s web traffic](https://www.itiger.com/news/2561490983) found search-driven visits fell 23.5% year over year in the second quarter of 2025 and 25.3% in July, which the analyst linked to Google’s wider rollout of AI Overviews.
2. A year later, monday.com’s [August 2026 guidance](https://s29.q4cdn.com/881027206/files/doc_financials/2026/q2/v2/MNDY-USQ_Transcript_2026-08-10.pdf) still assumed no rebound in top-of-funnel activity, while its customers paying over $100,000 a year grew 37% and those over $500,000 grew 68%.
3. [Canva](https://techcrunch.com/2026/02/18/canva-gets-to-4b-in-revenue-as-llm-referral-traffic-rises/) says it was one of the top 10 domains ChatGPT sends visitors to, and that traffic from AI referrals is “in double-digit percentages” (without saying of what), as its business accounts doubled to $500 million in annual recurring revenue.
4. In [Capterra’s survey](https://cioinfluence.com/machine-learning/ai-and-security-drive-project-management-software-investment-capterra-survey-finds/) of 2,545 project management professionals, 55% said adding AI features prompted their most recent purchase and 71% ranked security as critical.
5. In [our study of Google’s AI Overviews](https://underneath.agency/research/ai-reddit-citations-study), 35.4% of B2B software answers cited Reddit, with r/projectmanagers among the most cited communities.

## Who chooses a team’s work tool today, and what does a growing account bring in?

A team lead usually picks the first tool; IT and procurement later turn it into a contract.

Collaboration software is bought where people actually work, and work is still spread out. [Gallup’s 2025 data](https://www.gallup.com/workplace/694361/hybrid-work-retreat-barely.aspx) shows the share of remote-capable US employees working hybrid slipped from 55% to 51%, while fully remote and fully on-site work each rose two points. In tech, remote-capable staff are as likely to be fully remote (47%) as hybrid (45%). Even fully on-site employees increasingly say their team is spread across locations: 27% in 2025, up from 13% in 2023.

That keeps demand steady but the field crowded. According to [Ramp’s spend data](https://ramp.com/vendors/slack), 54% of organizations with a vendor in the communication category use Slack. And buyers are trimming as much as adding: [Zylo](https://zylo.com/blog/saas-benchmarks) reports that 53% of all SaaS licenses are underutilized, and lists meeting, productivity and project management tools among the applications most often duplicated. So the typical buying moment is either a team looking for something better than what it has, or IT asking which tools to consolidate onto. Support leaders re-shopping for AI agents follow a similar path, covered in [how support software gets chosen through AI](https://underneath.agency/resources/customer-support-software-ai-search).

What a customer is worth depends on how far it grows after the first seat. The public companies in the category report this directly:

- [Asana](https://seekingalpha.com/pr/20420470) ended fiscal 2026 with 25,928 “Core” customers spending $5,000 or more a year, and 817 spending $100,000 or more.
- [Atlassian](https://s206.q4cdn.com/270053503/files/doc_financials/2026/q1/TEAM-Q1-2026-Earnings-Release.pdf) serves over 300,000 customers, of which 53,017 had more than $10,000 in annual cloud revenue, up 13% year over year.
- [monday.com](https://ppc.land/monday-com-reports-google-ai-search-changes-impacting-traffic-acquisition) added 258 net new customers paying over $50,000 a year and a record 144 paying over $100,000 in one quarter of 2025.

The pattern is the same everywhere: a small team starts on a free or cheap plan, and the money arrives when the tool spreads to departments and a central buyer signs a larger agreement.

## Where does AI search already sit in that journey?

At the very top, where teams look for a tool to try, and the evidence is company-reported.

The clearest case is monday.com. On its [Q2 2025 earnings call](https://s29.q4cdn.com/881027206/files/doc_financials/2025/q2/MNDY-USQ_Transcript_2025-08-11.pdf), co-CEO Roy Mann said high-intent buyers still click through, and that “the drop that we see is just on volume because they are experimenting with AI on top.” The company described the effect as small and concentrated in the small-business segment. Bank of America’s analyst was less relaxed. Working from web traffic estimates, he noted that less than 30% of monday.com’s signups come from Google, and that extending July’s traffic trend implied a 5.2% decline in self-serve gross new annual recurring revenue in 2026. By August 2026 the company was planning its year without any recovery in performance marketing, and leaning on enterprise expansion instead.

Canva shows the other side of the same shift. Its co-founder told TechCrunch that the company sees ChatGPT and similar assistants as top-of-the-funnel acquisition platforms, and that it is allocating resources to appear in their answers alongside its search work. Canva said users had more than 26 million conversations with its app inside ChatGPT by October 2025.

Cross-category buyer surveys lean in the same direction. In [G2’s March 2026 survey](https://company.g2.com/news/g2-research-the-answer-economy) of software buyers across categories, 51% said they now start research with an AI chatbot more often than with Google. G2 sells to software vendors, so treat that as a direction rather than a precise share for collaboration tools.

## What do team leads and IT ask assistants about work tools?

Questions about fit for a team, replacing an incumbent, consolidating tools, security and AI features.

These sample prompts are our own, written to mirror how team leads and IT buyers phrase a search for a work tool; none was observed in real use:

- Fit for a team: “Best project management tool for a 15-person hybrid marketing team that works across three time zones.”
- Replacing an incumbent: “Alternatives to monday.com with a free plan that doesn’t cap seats.”
- Living with a suite: “Slack or Microsoft Teams for a company that already runs on Google Workspace?”
- Consolidation: “Can one tool replace our wiki, task tracker and video messaging?”
- Security and control: “Which whiteboard tools support single sign-on, audit logs and EU data residency?”
- AI features: “Which work management tools have AI agents that can update tasks for us?”

Vendors already write for the alternatives question. [ClickUp’s guide](https://clickup.com/blog/monday-alternatives/) to monday.com alternatives, for example, puts ClickUp first in its comparison table. Pages like that are common in software, and assistants can cite them; see [our article on comparison pages](https://underneath.agency/resources/do-comparison-pages-help-b2b-ai-citations) for what the research says about their value.

The AI-features question matters more here than in most categories. Capterra found that wanting AI features was the top trigger for a [new project management tool purchase](https://underneath.agency/resources/project-management-software-customers-ai-search), ahead of cost or user growth. A reasonable expectation is that buyers ask assistants which tools have the AI capabilities they want, which means your AI features need to be described in plain, public, checkable terms.

## How does an AI recommendation turn into revenue for a collaboration tool?

Through a free signup that spreads: one team, then more seats, then a central contract.

**Signup.** For product-led collaboration tools, the AI answer’s job is to send a team lead to the signup page. That click is the part analytics can see, and it is where monday.com felt the drop and Canva felt the gain.

**Seat growth.** Collaboration tools grow by invitation. A reasonable expectation, which we infer from how these products spread, is that one AI-sourced team can be worth far more than its first plan once colleagues join.

**Enterprise contract.** The largest revenue arrives when IT or procurement standardizes. Canva’s business accounts, defined as companies with more than 25 seats, grew 100% to $500 million in annual recurring revenue. monday.com’s growth in 2026 came mainly from upmarket customers. At that stage, buyers are more likely to ask an assistant to compare shortlisted tools on security, administration and integrations, and those answers rarely show up as referral traffic. We infer they surface instead as direct visits, branded searches and demo requests.

**Renewal.** The contract then has to survive renewals and consolidation reviews. Asana reported dollar-based net retention of 96% for customers spending $100,000 or more, which means retention, not just acquisition, sets the value of each new account.

## What gets a work management tool named in an AI answer?

Platforms document little; studies point to independent coverage, community discussion and clear, checkable facts.

**What Google says.** AI Mode, according to Google, relies on a [“query fan-out” technique](https://blog.google/products/search/ai-mode-search/): it issues multiple related searches across subtopics and data sources, then combines the results. A question about the best tool for a hybrid team can therefore pull in pricing pages, review sites and community threads at once.

**Observed in our studies.** In our Reddit study, 17.9% of all AI Overviews in our sample cited at least one Reddit thread, and more than a third did in B2B software. The most cited communities included r/projectmanagers (9 citations) and r/projectmanagement, where practitioners discuss the tools they use. In [our brand entity study](https://underneath.agency/research/brand-entity-ai-recommendations-study), independent coverage was the strongest predictor we measured: each tenfold increase in independent sites naming a brand in the cited pages went with 4.7 times the odds of being recommended. The same study found big suite brands such as Google Workspace, Slack or Zoom named in passing in answers to a project management question, rather than put forward as the pick. We infer that being famous in an adjacent category does not make a product the answer to a specific team’s question.

**Observed in other research.** In a test on skincare products, [researchers found](https://arxiv.org/abs/2606.17443) that well-known brands were recommended 100% of the time when all products had identical specifications, but that dominance disappeared once a competitor had less than a 0.1-star rating advantage. That is a different category, so treat it only as a signal: a clearly stated advantage may be how a challenger gets named next to Microsoft, Google or Atlassian.

**Trust factors specific to collaboration software.** Security comes first: in Capterra’s survey, 71% of buyers ranked it as critical and 39% made purchases because of security needs. Integrations and administration controls follow, because the tool has to fit the suite a company already runs. And [Capterra’s wider buyer research](https://capterra.com/resources/tech-trends-successful-buyer-purchase-journey/) found successful buyers were at least 50% more likely to factor in product comparison sites and expert recommendations when building their shortlist. A reasonable expectation is that public trust pages, integration directories and review profiles carry weight when an assistant compares tools.

## What does it cost a collaboration tool to be left out?

Mostly small-business volume first, then a weaker pipeline into enterprise expansion.

monday.com is the documented example. Its executives called the impact minor and limited to the small-business end, while the analyst modeled a measurable hit to self-serve growth and the company stopped planning for a recovery. Both views agree on where the cost lands: on the free and low-tier signups that feed seat growth. For a product-led company, those signups are the raw material for tomorrow’s enterprise accounts, so a smaller top of the funnel can show up in large-customer growth a year or two later. That last step is our inference; no company has published it.

The loss is also hard to see from inside. When an assistant answers “which tool should my team use?” with three names and yours is not among them, nothing appears in analytics. How fewer search clicks play out further down the funnel is the subject of [our article on AI answers and pipeline](https://underneath.agency/resources/ai-answers-pipeline-revenue).

## Where does GEO effort go for a collaboration product?

Into what assistants can find and trust about your tool, from positioning to pricing; no placement is guaranteed.

Six areas usually make up generative engine optimization (GEO) for a work management or messaging product:

1. **One clear story everywhere.** Describe who the product is for, what it replaces and what it connects to in the same words on your site, review profiles, help center and app marketplace listings (Slack, Microsoft Teams, Google Workspace and Atlassian all run app directories). See [why well-known brands still miss AI recommendations](https://underneath.agency/resources/why-well-known-brands-miss-ai-recommendations).
2. **Independent coverage.** Earn mentions in the comparison sites, newsletters, podcasts and “best of” lists your buyers read. Our article on [which lists matter](https://underneath.agency/resources/best-of-lists-ai-recommendations) covers how to pick them.
3. **Community presence, done honestly.** Practitioner communities like the project management subreddits are where tool choices get discussed and, in our data, cited. Answer questions openly as the vendor; do not plant posts.
4. **Public, specific security and admin pages.** Single sign-on, audit logs, data residency, compliance reports and admin controls should be readable without a login, because that is what the enterprise comparison turns on.
5. **Accurate pricing.** Per-seat prices with annual and monthly billing are easy for answers to garble. In [our pricing study](https://underneath.agency/research/ai-pricing-accuracy-study), Asana’s Starter plan was quoted at $10.99 a month, which is the annual-billing rate; billed monthly it was $13.49. Keep one current pricing page and retire old figures elsewhere.
6. **Measurement by buying stage.** Track a fixed set of team-fit, alternatives, consolidation and security questions across ChatGPT, Gemini, Perplexity, Copilot and Google, and connect the results to signups, seat growth and enterprise pipeline. For how to set that up, see [how many prompts to track](https://underneath.agency/resources/how-many-prompts-to-track-ai-visibility).

## What can’t we yet say about AI search and collaboration revenue?

No one has shown how much AI visibility adds to a collaboration company’s revenue, rather than accompanying it.

The monday.com figures are an analyst’s estimates from third-party traffic data, and the company disputed how much they mattered. Canva’s referral claims are self-reported to a journalist and not broken down. The Capterra and G2 surveys come from review platforms that sell to software vendors. Our Reddit, brand entity and pricing studies record what AI answers cite, not what team leads do next. We also cannot yet tell how much of monday.com’s small-business slowdown came from AI Overviews rather than Google’s ranking updates or wider demand. Treat the channel as real and growing, with the size of its effect on contracts still unmeasured.

## How should a collaboration company check whether assistants are sending it teams?

Find out which AI answers name you when team leads ask for a tool like yours.

A practical first step is an audit of your team-fit, alternatives, consolidation and security questions across the main assistants and Google’s AI features, set against your signup, seat-growth and enterprise pipeline data, so the gaps are ranked by what they are worth. We can help rank those gaps by seat and contract value and plan the fixes: [ask us to review where assistants place you for team-fit questions](https://underneath.agency/contact). The ongoing work, from public security and admin pages to tracking answers against signups and seat growth, is outlined on our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page.

## Frequently asked questions

### Do AI Overviews really reduce signups for collaboration tools?

In the one public case, monday.com reported softer small-business acquisition and an analyst linked a fall of about a quarter in its search-driven visits to AI Overviews. The company said high-intent buyers still click through. No study has yet measured the effect across the category.

### Should product-led collaboration companies care about AI visibility if most revenue is enterprise?

Yes. Enterprise accounts in this category often start as a single team on a free plan, so the AI answer that sends that team to you can be the first step of a large contract. The enterprise comparison stage also happens partly inside assistants.

### Does publishing an “alternatives to” page help?

It can, if it is fair and specific. Assistants look for pages that answer alternatives questions, and vendors in this category publish many of them. Pages that only praise your own product are a weak bet.

### How can a collaboration software company tell if AI is sending it customers?

Combine AI referral signups in analytics with a “how did you hear about us?” question at signup, and a regular check of how often assistants name you for your buyers’ questions. Then follow those accounts as they add seats.

## Sources

- Benzinga via Tiger Brokers (2025), [Monday.com Stock Slides As SEO Traffic Declines Raise Growth Fears](https://www.itiger.com/news/2561490983)
- monday.com (2025), [Q2 2025 earnings call transcript](https://s29.q4cdn.com/881027206/files/doc_financials/2025/q2/MNDY-USQ_Transcript_2025-08-11.pdf)
- monday.com (2026), [Q2 2026 earnings call transcript](https://s29.q4cdn.com/881027206/files/doc_financials/2026/q2/v2/MNDY-USQ_Transcript_2026-08-10.pdf)
- PPC Land (2025), [Monday.com reports Google AI search changes impacting traffic acquisition](https://ppc.land/monday-com-reports-google-ai-search-changes-impacting-traffic-acquisition)
- TechCrunch (2026), [Canva gets to $4B in revenue as LLM referral traffic rises](https://techcrunch.com/2026/02/18/canva-gets-to-4b-in-revenue-as-llm-referral-traffic-rises/)
- Asana (2026), [Asana Announces Fourth Quarter and Fiscal Year 2026 Results](https://seekingalpha.com/pr/20420470)
- Atlassian (2025), [Q1 Fiscal Year 2026 Earnings Release](https://s206.q4cdn.com/270053503/files/doc_financials/2026/q1/TEAM-Q1-2026-Earnings-Release.pdf)
- Capterra via CIO Influence (2025), [AI and Security Drive Project Management Software Investment, Capterra Survey Finds](https://cioinfluence.com/machine-learning/ai-and-security-drive-project-management-software-investment-capterra-survey-finds/)
- Capterra (2024), [Capterra’s 2025 Tech Trends Report](https://capterra.com/resources/tech-trends-successful-buyer-purchase-journey/)
- Gallup (2025), [Hybrid Work in Retreat? Barely.](https://www.gallup.com/workplace/694361/hybrid-work-retreat-barely.aspx)
- Ramp (2026), [Slack vendor data](https://ramp.com/vendors/slack)
- Zylo (2026), [SaaS benchmarks](https://zylo.com/blog/saas-benchmarks)
- G2 (2026), [In the Answer Economy, Don’t Win the Click — Win the Answer](https://company.g2.com/news/g2-research-the-answer-economy)
- ClickUp (2026), [Best monday.com Alternatives & Competitors](https://clickup.com/blog/monday-alternatives/)
- Google (2025), [Expanding AI Overviews and introducing AI Mode](https://blog.google/products/search/ai-mode-search/)
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems (2026), [Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems](https://arxiv.org/abs/2606.17443)
- Underneath (2026), [Reddit citations in AI answers](https://underneath.agency/research/ai-reddit-citations-study), [brand entity and AI recommendations](https://underneath.agency/research/brand-entity-ai-recommendations-study) and [software pricing accuracy](https://underneath.agency/research/ai-pricing-accuracy-study)

---

This is the Markdown twin of https://underneath.agency/resources/collaboration-software-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Can you combine AI engines into one visibility score?"
description: "Only with care. Score each engine separately first, then combine with stated weights; pooling raw counts lets the most talkative engine dominate."
canonical: "https://underneath.agency/resources/combine-ai-engines-visibility-score"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Can we combine ChatGPT, Gemini and Perplexity into one AI visibility score?

You can, but only after scoring each engine on its own, and the single number should never replace the per-engine view. The engines cite very different numbers of sources, mostly cite different pages and can move in opposite directions for the same brand. Add raw counts together and the score mostly reflects whichever engine cites the most.

## The short version

1. Gemini produced between 19.9 and 50.1 citations per answer depending on the topic, against 4.2 to 6.3 for ChatGPT search, so pooled raw counts are dominated by Gemini ([Sielinski, 2026](https://arxiv.org/abs/2607.10341)).
2. In a test of 15 commercial prompts across four AI engines, 96.4% of cited pages appeared on only one engine ([Tannenbaum, 2026](https://arxiv.org/abs/2609.22655)).
3. Across 80 buyer questions, ChatGPT, Gemini, Perplexity and Claude shared about a third of their recommended options and agreed on the top pick for 10.0% of questions ([our brand agreement study](https://underneath.agency/research/ai-assistants-brand-agreement-study)).
4. For one brand over five weeks, Perplexity’s mention rate climbed from 37% to 62% while ChatGPT’s slid from 45% to 20%, a split a single score would flatten ([Kumar, 2026](https://arxiv.org/abs/2606.20065)).

## Why not just add up results across all the engines?

Because the engines produce very different volumes, so the biggest one swamps the rest. [Sielinski](https://arxiv.org/abs/2607.10341) measured citations across ten topics. Gemini produced between 19.9 and 50.1 citations per answer, Perplexity between 14.5 and 42.2, and ChatGPT search between 4.2 and 6.3.

His advice is direct. Adding raw citation counts across platforms weights each one by its volume, so a dense platform overshadows sparse ones. The fix is to turn each engine’s results into a share first, then combine the shares, so each engine counts equally regardless of volume. The same applies to the error ranges: work them out per engine, then combine.

Volume differences also show up in other studies. In a June 2026 test, ChatGPT cited 17.8 distinct pages per prompt on average and Microsoft Copilot 4.5 ([Tannenbaum, 2026](https://arxiv.org/abs/2609.22655)). A page competing for a slot in a four-source answer faces a different contest from one in an 18-source answer. Our guide to [how many sources each engine cites](https://underneath.agency/resources/how-many-sources-ai-search-engines-cite) compares the counts across studies.

## Do the engines even measure the same thing?

Not really; they mostly cite different pages and often recommend different brands. [Tannenbaum](https://arxiv.org/abs/2609.22655) put 15 commercial prompts to ChatGPT, Copilot, Google and Perplexity on 6 June 2026. Of the 528 distinct pages cited, 96.4% appeared on only one engine. In 84.9% of engine pairs, the two engines shared no cited page at all for the same prompt. Other studies of [whether AI engines cite the same sources](https://underneath.agency/resources/do-ai-search-engines-cite-the-same-sources) point the same way.

No single engine was a good stand-in for the others. The broadest one, ChatGPT, captured 42.6% of all the pages the four engines cited between them; Copilot captured 11.4%. The author is affiliated with Aiso Boost Ltd., and the prompts were about AI visibility software, an unusual topic, so the exact figures may not transfer. Still, it is one reason [tracking ChatGPT alone falls short](https://underneath.agency/resources/is-tracking-chatgpt-enough).

Recommendations diverge too. In our [brand agreement study](https://underneath.agency/research/ai-assistants-brand-agreement-study), two assistants’ recommended options for the same question overlapped by 0.327 on a scale where 1 means identical. A [Dutch audit of product questions](https://arxiv.org/abs/2609.18729) found the ChatGPT and Gemini apps shared only 5.4% of their cited websites for the same question.

## What can a single combined score hide?

It can hide engines moving in opposite directions. Ranqo, which sells AI visibility tracking, followed one brand, ParallelDots, for five weeks on the same prompts ([Kumar, 2026](https://arxiv.org/abs/2606.20065)). Perplexity’s mention rate climbed from 37% to 62% while ChatGPT’s slid from 45% to 20%. An average of the two would look almost flat while both engines changed sharply.

The author treats this as one illustrative brand, not proof. But it shows the risk: a blended number can say “stable” when the right response is to work on one engine and protect gains on another.

A combined score can also hide differences between models from the same company. [Kato and colleagues](https://arxiv.org/abs/2609.11915) tracked one brand, Glasp, across 2,240 answers from two OpenAI models. Overall, the newer model named it in 378 of 1,120 answers and the older one in 311 of 1,120. Yet the older model led on the English questions, and rates varied widely by use case. The authors conclude that one overall rate is inadequate.

## If we do combine engines, how should we weight them?

By a stated, defensible rule, such as each engine’s share of your buyers, and never by raw volume. Any weighting is a choice that changes the answer. Ranqo’s own cross-engine figure, for example, weights ChatGPT at 0.30, Gemini, Perplexity and Claude at 0.20 each and Grok at 0.10.

[Martinez](https://arxiv.org/abs/2609.06811), in a survey of GEO measurement, argues that a single score conceals the choices behind it when several weightings are reasonable. His advice is to report the range of scores across those weightings. When comparing brands or periods, apply the same weights to both, rather than comparing scores built on different mixes.

In practice, that means three numbers per engine and a clearly labeled blend. If your buyers mostly use ChatGPT, weight it heavily and say so. If you do not know where your buyers are, show the blend under two or three plausible weightings.

## Do different engines need different measurement plans?

Yes, because each engine varies in its own way. [Sielinski](https://arxiv.org/abs/2603.08924) found each engine had a typical level of repeat consistency regardless of answer length: repeated runs of a question shared about 0.30 of their sources on Gemini, 0.40 on ChatGPT search and 0.50 on Perplexity.

Engines also differ in how concentrated their sources are. A Swiss study scored citation concentration at 0.782 for Google AI Mode and 0.671 for Perplexity, on a scale where 1 means one site takes everything ([Schulte, Bleeker and Kaufmann, 2026](https://arxiv.org/abs/2604.07585)). Its advice is to set a separate baseline for each engine rather than one threshold for all.

They differ in whether they show links at all. In Ranqo’s study of CRM software questions, Perplexity included source links in 95% of answers and Claude in 10%. A citation-based score is therefore partly a measure of each product’s design, not just of your brand.

## What should you do about it?

Keep each engine as its own line, and treat any combined score as a summary on top. Steps:

1. Measure each engine separately, with enough repeated runs for that engine, and report its own rate and range.
2. Convert each engine’s results to shares before combining, so a high-volume engine does not dominate.
3. Choose and publish the weights, ideally based on where your buyers actually ask. Show how the blend changes under other reasonable weights.
4. Watch for engines moving in opposite directions. A flat blend with diverging engines needs action, not reassurance.
5. Keep mentions, citations and first-place picks as separate measures; they behave differently on each engine.
6. Read our [brand agreement study](https://underneath.agency/research/ai-assistants-brand-agreement-study) for how far the main assistants disagree on the same questions.

If you want a dashboard built this way, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The case against raw pooling is strong; the right weights are still an open question. Gaps:

- No study has measured what share of real buyers in a given category use each engine, which is what sensible weights need. AI companies do not publish it.
- The clearest evidence on citation volume covers only Gemini, ChatGPT search and Perplexity, on consumer topics.
- The page-overlap figures come from 15 prompts on one topic in one week, from a single author affiliated with a private firm.
- The opposite-direction example is one brand tracked by a vendor. How often engines diverge for the same brand has not been measured across many brands.
- No study yet compares how well different blended scores predict business results such as traffic or sales.

## Frequently asked questions

### Should AI visibility be reported per engine or as one number?

Per engine first. Engines cite different volumes and mostly different pages, so a single number is only useful as a labeled summary on top of the per-engine results.

### Why does one AI engine dominate my combined visibility score?

Probably because raw counts were added up. Gemini produced up to 50.1 citations per answer in one study, against at most 6.3 for ChatGPT search, so adding counts mostly measures Gemini.

### How should we weight ChatGPT, Gemini and Perplexity in one score?

Use a stated rule tied to where your buyers ask, apply it consistently, and show how the score changes under other reasonable weights.

### Do ChatGPT and Gemini cite the same sources?

Rarely. In one audit of product questions, the two apps shared only 5.4% of their cited websites for the same question.

## Sources

- Sielinski (2026), [From Stochastic to Stable: Rank Stability and Structural Sufficiency in AI Visibility Measurement](https://arxiv.org/abs/2607.10341), arXiv:2607.10341.
- Sielinski (2026), [Quantifying Uncertainty in AI Visibility: A Statistical Framework for Generative Search Measurement](https://arxiv.org/abs/2603.08924), arXiv:2603.08924.
- Tannenbaum (2026), [Scoring With the Engine: Retrieval Exposure, Cross-Engine Divergence, and the Limits of Engine-Agnostic GEO Scores](https://arxiv.org/abs/2609.22655), arXiv:2609.22655.
- Kumar (2026), [Generative Engine Optimization at Scale: Measuring Brand Visibility Across AI Search Engines](https://arxiv.org/abs/2606.20065), arXiv:2606.20065.
- Kato, Honma and Kato (2026), [Generative Marketing Mix Modeling: A Causal Inference Framework Linking GEO and GEM to Business Impact](https://arxiv.org/abs/2609.11915), arXiv:2609.11915.
- Martinez (2026), [Measuring GEO Visibility: Prompt Corpora Define the Answer Market](https://arxiv.org/abs/2609.06811), arXiv:2609.06811.
- Uberti-Bona Marin and colleagues (2026), ["If I Had to Buy Just ONE: Galaxy S26 Ultra": Auditing AI-Generated Product Recommendations](https://arxiv.org/abs/2609.18729), arXiv:2609.18729.
- Schulte, Bleeker and Kaufmann (2026), [Don’t Measure Once: Measuring Visibility in AI Search (GEO)](https://arxiv.org/abs/2604.07585), arXiv:2604.07585.
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude agree on brands?](https://underneath.agency/research/ai-assistants-brand-agreement-study)

---

This is the Markdown twin of https://underneath.agency/resources/combine-ai-engines-visibility-score. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How construction software companies win customers in AI search"
description: "Often they do: contractors use AI at work, and the shortlist for estimating, bidding and project software forms before your demo request arrives."
canonical: "https://underneath.agency/resources/construction-software-customers-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Do contractors find our construction software when they ask AI what to use?

Many now start there, and whether your product appears depends on what assistants can verify about it. Contractors have folded AI into office work, and the shortlist for estimating, bidding, project management and field apps is often set before a vendor hears about the deal.

This guide is for companies that sell construction management software to contractors: general contractors, specialty trades, home builders and remodelers. Jobsite hardware, such as drones, robots and sensors, follows a different path that we cover separately. Here the prize is a subscription that grows from one module to a company-wide platform.

## The short version

1. Contractors are adopting AI fast: in the [AGC and Sage 2026 Construction Hiring and Business Outlook](https://www.agc.org/sites/default/files/users/user21902/2026%20Construction%20Hiring%20and%20Business%20Outlook%20Report_Final.pdf), 61% of 951 responding firms used AI or planned to invest more in it, up from 44% a year earlier, and 23% already used it for estimating.
2. Buyers want fewer, broader tools: [Capterra’s analysis](https://www.capterra.com/resources/ai-in-construction-management-software-on-time-budget/) of more than 8,200 construction software buyer conversations found estimating requested by 72% and project management by 68%, with three out of four buyers who stated a preference wanting both in one platform.
3. Most buyers are small: 55% of those Capterra conversations came from businesses with 10 or fewer employees, in an industry the [Associated General Contractors of America](https://www.agc.org/learn/construction-data) counts at more than 919,000 establishments.
4. Accounts grow after the first sale: [Procore](https://www.boardroomalpha.com/sec/pcor-8-k-2026-02-12-0001628280-26-007662/pcor-q425x8xkxexx991.htm) ended 2025 with 17,850 organic customers, 2,710 of them above $100,000 in annual recurring revenue, and 78% of its recurring revenue came from customers using four or more products.
5. General AI tools dominate daily use: in [Mastt’s 2026 survey](https://www.mastt.com/research/ai-in-construction-project-management-2026) of construction project managers, ChatGPT reached 66.7% of respondents, while construction-specific AI stayed under 10%.

## Who buys construction software, and how much is one contractor account worth?

Owners and operations leaders at contractors buy it, and a good account grows for years through modules and renewals.

At a small contractor, the owner usually decides, often after an office manager or estimator has found the options. At a larger firm, the buying group widens: operations and preconstruction leaders, the chief estimator, project managers, the controller or chief financial officer when accounting is involved, and IT. Specialty contractors weigh field time tracking and change orders; home builders and remodelers weigh client communication and selections; general contractors weigh bidding, document control and subcontractor coordination.

The market they serve is large and fragmented. The [US Census Bureau](https://www.census.gov/construction/c30/pdf/release.pdf) put construction spending at an annual rate of $2,203.1 billion in August 2026, of which residential work was $882.3 billion. AGC counts an industry of 8.0 million employees.

Public company data shows what a customer can become:

- **Procore** reported revenue of $1,323 million for 2025, 115 customers above $1,000,000 in annual recurring revenue, a gross revenue retention rate of 95% and a net revenue retention rate of 106%. By our arithmetic, that is roughly $74,000 of revenue per organic customer on average, though a few large accounts pull the average up.
- **JobTread**, which sells mainly to residential builders and remodelers, says it [reached 10,000 companies in February 2026](https://jobtread.com/blog/jobtread-vs-buildertrend-construction-management-software-comparison).

Our inference: a contractor won today is worth far more than its first-year contract, because satisfied accounts add products, projects and users. So the shortlist a contractor writes before its first demo carries more weight than its size suggests.

## How far has AI reached into contractors’ software decisions?

Far into daily work, and into early research; no study yet measures how often it picks the software.

The AGC and Sage survey found 45% of firms used AI for office and administrative work and 20% for design or preconstruction. Among residential and design firms, the [2026 Houzz State of AI in Construction and Design Report](https://www.constructionowners.com/press-release/houzz-survey-finds-ai-adoption-soars-among-construction-and-design-pros-while-homeowners-rely-on-the-experts) found 52% of construction firms using AI for everyday business tasks, with sales and marketing the leading use, at 64%. In Mastt’s smaller, global sample of 108 project management professionals, Microsoft Copilot followed ChatGPT at 50.9%.

The habit of asking a general assistant is already there. What remains unmeasured is how often a contractor asks it which software to buy. The nearest evidence is cross-industry: in Gartner Digital Markets’ [2025 software buying report](https://cdn-static.bizzabo.com/bizzabo.file.upload/yVkIqwmWRDKbEhqeE3aQ_Gartner-Digital-Markets_2025-Report_Making-The-List.pdf), which surveyed 3,500 software buyers, small enterprises used ChatGPT or other generative AI tools most often (31%) when building a shortlist. Buyers carried an average of 4.4 options on their initial list. Since 55% of construction software buyers in Capterra’s data are businesses with 10 or fewer employees, we infer this small-business pattern is relevant here.

Vendors are moving the same way. JobTread’s AI Connector gives contractors access to their job data through AI platforms such as Claude and ChatGPT. In August 2026, Capterra surveyed 10 construction software providers and found three had agentic AI live for customers.

## What do contractors ask when they shop for software?

They ask by trade, company size, workflow and the accounting system they already run.

These questions are our own illustrations of the pattern, not prompts recorded from contractors or answers captured from an assistant.

| Need | Illustrative question |
|---|---|
| Category, by trade and size | “Best project management software for a 25-person electrical contractor” |
| Alternatives | “Cheaper alternatives to Procore for a general contractor doing $40 million a year” |
| Head-to-head | “JobTread or Buildertrend for a remodeler with three crews?” |
| Estimating | “Estimating and takeoff software that exports to Sage 300 CRE” |
| Bidding | “Bid management tools for a mechanical subcontractor that receives 50 invitations to bid a month” |
| Field | “Daily log app that works offline on remote sites” |
| Compliance | “Construction project software with FedRAMP authorization for federal work” |

Each question carries facts a vendor can either state or leave out: the trades served, company size, accounting integrations, offline support, pricing model, security authorizations. Capterra’s finding that most buyers want estimating and project management together also suggests buyers will ask whether one platform covers both, so say plainly what yours covers and what it connects to. Product makers meet the same kind of filtering when contractors ask what to specify, as our guide to [how specifiers find building materials through AI](https://underneath.agency/resources/building-materials-specification-ai-search) explains.

## How does an AI shortlist turn into a trial, a rollout and a renewal?

AI helps form the list; demos, trials and a first project decide the sale; expansion decides the value.

1. **Listed.** A contractor asks an assistant, peers, review sites and search which tools fit. Gartner Digital Markets found buyers narrow about 4.4 options to an average of 3.5 on the shortlist.
2. **Demonstrated or trialed.** Across business buying, [Forrester](https://www.forrester.com/press-newsroom/forrester-2026-the-state-of-business-buying/) reports that more than 60% of buyers now use a trial before buying.
3. **Proven on live projects.** The software runs on a few new jobs before the company moves its whole backlog.
4. **Rolled out.** Estimating, project management, financials and field modules spread across the firm.
5. **Renewed and expanded.** At Procore, 52% of recurring revenue came from customers using six or more products.

An assistant can influence step one. Everything after depends on the product, onboarding and support. A vendor missing from step one does not get to show any of it.

## Why does an assistant name one construction platform over another?

Because it can find independent, consistent evidence for it; the platforms do not publish how they choose.

**What the platforms document.** [OpenAI says](https://help.openai.com/en/articles/9237897-chatgpt-search) ChatGPT search typically rewrites a question into one or more targeted queries for search partners, and that a site must allow OAI-SearchBot to be eligible for inclusion. [Google says](https://developers.google.com/search/docs/appearance/ai-features) its AI Overviews and AI Mode may run a “query fan-out” of related searches, and that pages need no special markup beyond being indexed and eligible for a snippet.

**What our studies observed for business software.** In [our AI Overviews frequency study](https://underneath.agency/research/ai-overviews-frequency-study), 96.0% of B2B software and technology keywords showed an AI Overview, the highest of the industries we tested. In [our YouTube study](https://underneath.agency/research/ai-overview-youtube-videos-study), the AI Overview cited a YouTube video on 91.0% of B2B software searches. In [our Reddit study](https://underneath.agency/research/ai-reddit-citations-study), 35.4% of B2B software AI Overviews cited Reddit. And in [our study of self-ranking lists](https://underneath.agency/research/self-promoting-best-lists-study), 24.2% of the numbered “best” lists that AI engines cited ranked their own publisher first.

**What the buyers say.** In the Gartner Digital Markets data, customer reviews, industry experts and search were the most influential sources when shortlisting.

**Our inference for construction software.** The evidence an assistant can find is the evidence a contractor already checks: reviews on software marketplaces, video walkthroughs, honest comparisons, peer discussion in contractor communities, trade-press coverage, named customer stories by trade, and clear statements about integrations, pricing model and security. Procore, for example, lists a FedRAMP Moderate authorization and a 2026 TrustRadius Buyer’s Choice award in its results release, both facts a federal contractor or a reviewer can verify.

## What happens to a software vendor that AI answers skip?

It misses buyers who are consolidating tools, and it rarely sees the loss in its own data.

We found no direct measurement of lost trials, so we label the reasoning:

- **Consolidation favors the named.** When three out of four buyers want estimating and project management in one platform, a buyer who asks for one integrated system may never meet a strong point solution that the assistant did not mention. That is our inference.
- **Competition is tight on both sides.** Contractors themselves are under pressure: 44% of AGC respondents named increased competition for projects as a concern for 2026. Software that promises faster estimates and bids speaks to that worry, but only if the buyer hears of it.
- **Rivals publish the comparisons.** JobTread publishes its own comparison with Buildertrend, and a smaller rival, BuildTools, publishes [one comparing itself with JobTread](https://www.buildtools.com/blog/2026-03-31-jobtread-vs-buildtools-modern-builder-software). When only your competitors describe your product, theirs is the account assistants can find.
- **The loss is silent.** A contractor who never shortlists you never books a demo. Our guide to [why analytics miss AI visibility](https://underneath.agency/resources/why-analytics-miss-ai-visibility) explains the gap.

## What does GEO look like for a construction software company?

Generative engine optimization (GEO) means making your product easy for assistants to find, describe accurately and support with outside evidence.

For construction software, it usually covers:

1. **Pages by trade and company size.** Separate, plain-text pages for the trades and segments you serve, with the workflows each one uses.
2. **Integration facts.** The accounting, scheduling, design and payroll systems you connect to, how, and with what limits, matched on partners’ listings. Jobsite hardware makers state which platforms they connect to as well, as [our guide to jobsite technology](https://underneath.agency/resources/construction-technology-demand-ai-search) shows.
3. **Pricing model clarity.** Even without publishing prices, state how you charge (per user, per project, by construction volume or flat) and what is included.
4. **Security and compliance.** Authorizations, audits and data-hosting facts, kept current and dated.
5. **Review platforms.** Complete, accurate profiles and a steady flow of reviews from real customers on the marketplaces contractors use.
6. **Fair comparisons.** Comparison and alternative pages that describe rivals accurately. Our review of [whether comparison pages help B2B brands get cited](https://underneath.agency/resources/do-comparison-pages-help-b2b-ai-citations) covers what they can and cannot do.
7. **Video with transcripts.** Product walkthroughs for each workflow, with titles and transcripts that name the trade and the task.
8. **Independent coverage.** Trade publications, association events, awards and contractor communities, the sources assistants and buyers both trust. Our research on [which pages AI engines cite when they recommend brands](https://underneath.agency/resources/best-of-lists-ai-recommendations) shows how much weight outside lists carry.
9. **Measurement.** Ask a fixed set of category, alternative, integration and trade questions across ChatGPT, Gemini, Perplexity, Copilot and Google’s AI features, repeat them over time, and compare the trend with demo and trial requests.

No construction software vendor can make an assistant recommend it. What this work does is remove the reasons an assistant, or a contractor, might leave you out. Our broader guide to [B2B SaaS revenue from AI search](https://underneath.agency/resources/b2b-saas-revenue-from-ai-search) covers the general software path, and the guide to [project management software](https://underneath.agency/resources/project-management-software-customers-ai-search) covers the horizontal category construction tools compete with.

## Which gaps remain in what we know about construction software and AI?

We know contractors use AI at work, not how often an AI answer decides their software purchase.

- **No attribution data.** None of the sources here links an AI answer to a construction software trial or contract.
- **Interested sources.** Capterra and Gartner Digital Markets earn money from software vendors, Mastt and Sage sell construction software, Houzz runs a platform for building professionals, and the JobTread and BuildTools comparisons are marketing.
- **Small or partial samples.** Mastt surveyed 108 people; Capterra’s provider survey covered 10 vendors; AGC’s respondents are its members.
- **Our studies used general software queries.** Applying their findings to construction-specific questions is our inference.
- **Account values are mostly private.** Procore is public; most construction software companies are not.

## How should a construction software company begin?

Begin by asking assistants the questions your last ten new customers asked, and record who gets named.

That first pass usually shows whether you appear for your trades and company sizes, whether your integrations and pricing model are described correctly, which review sites, videos and discussions the answers draw on, and which rivals are named instead. From there the work is to state the facts assistants need and build the outside evidence that confirms them.

If your next stage of growth relies on contractors booking demos and starting trials, [we can review how AI answers describe your product](https://underneath.agency/contact). That review runs the trade, company-size and integration questions contractors ask, shows which rival platforms and review sites shape the answers, and ranks the fixes most likely to add qualified trials. Pages by trade, integration facts and review profiles are the core of our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization), which explains how that work is scoped and measured.

## Frequently asked questions

### Do contractors use ChatGPT to choose construction software?

Many use it at work: ChatGPT reached 66.7% of respondents in Mastt’s survey. We found no study measuring how often it decides the software purchase itself, so treat it as one input among reviews, peers and demos.

### Does a big brand like Procore always win AI answers?

Not necessarily. Large brands have more independent coverage, but questions that name a trade, company size or accounting system narrow the field, and smaller products that state those facts clearly can be named for them.

### Do review sites matter for AI visibility?

Buyers say reviews are among their most influential sources, and review platforms are outside sources assistants can find. Keep profiles complete, current and consistent with your own site.

### Should a construction software company put its prices on its website?

That is a sales decision. At minimum, explain how pricing works and what is included, so neither buyers nor assistants have to guess.

## Sources

- Associated General Contractors of America and Sage (2026), [2026 Construction Hiring and Business Outlook](https://www.agc.org/sites/default/files/users/user21902/2026%20Construction%20Hiring%20and%20Business%20Outlook%20Report_Final.pdf)
- Associated General Contractors of America (2026), [Construction Data](https://www.agc.org/learn/construction-data)
- Capterra (2026-09), [Does AI in Construction Management Software Actually Keep Projects on Time and Budget?](https://www.capterra.com/resources/ai-in-construction-management-software-on-time-budget/)
- Procore, Form 8-K exhibit 99.1 (2026-02-12), [Procore Announces Fourth Quarter and Full Year 2025 Financial Results](https://www.boardroomalpha.com/sec/pcor-8-k-2026-02-12-0001628280-26-007662/pcor-q425x8xkxexx991.htm)
- JobTread (2026), [JobTread vs. Buildertrend: Construction Management Software Comparison](https://jobtread.com/blog/jobtread-vs-buildertrend-construction-management-software-comparison)
- BuildTools (2026-03-31), [JobTread vs. BuildTools](https://www.buildtools.com/blog/2026-03-31-jobtread-vs-buildtools-modern-builder-software)
- Mastt (2026-07-23), [State of AI in Construction Project Management 2026](https://www.mastt.com/research/ai-in-construction-project-management-2026)
- Houzz via Construction Owners Club (2026-09-03), [Houzz Survey Finds AI Adoption Soars Among Construction and Design Pros](https://www.constructionowners.com/press-release/houzz-survey-finds-ai-adoption-soars-among-construction-and-design-pros-while-homeowners-rely-on-the-experts)
- US Census Bureau (2026-10-01), [Monthly Construction Spending, August 2026](https://www.census.gov/construction/c30/pdf/release.pdf)
- Gartner Digital Markets (2025), [Making the List: 2025 software buying report](https://cdn-static.bizzabo.com/bizzabo.file.upload/yVkIqwmWRDKbEhqeE3aQ_Gartner-Digital-Markets_2025-Report_Making-The-List.pdf)
- Forrester (2026-01-21), [Forrester’s 2026 Buyer Insights: GenAI Is Upending B2B Buying](https://www.forrester.com/press-newsroom/forrester-2026-the-state-of-business-buying/)
- OpenAI Help Center (2026), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Google Search Central (2026), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- Underneath (2026), [When does Google show an AI Overview? 1,248 US searches](https://underneath.agency/research/ai-overviews-frequency-study)
- Underneath (2026), [The YouTube videos Google’s AI cites are small](https://underneath.agency/research/ai-overview-youtube-videos-study)
- Underneath (2026), [When does a Reddit thread become evidence in Google’s AI?](https://underneath.agency/research/ai-reddit-citations-study)
- Underneath (2026), [How many “best of” lists cited by AI rank their own brand first?](https://underneath.agency/research/self-promoting-best-lists-study)

---

This is the Markdown twin of https://underneath.agency/resources/construction-software-customers-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How construction technology companies win demand in AI search"
description: "It can, when AI can verify your accuracy, compliance and service record, because contractors now shortlist drones, robots and sensors before calling."
canonical: "https://underneath.agency/resources/construction-technology-demand-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Can AI search put our jobsite technology on a contractor’s shortlist?

Yes, if an assistant can find and confirm what your drone, robot, scanner, sensor or telematics product does on a real jobsite. Contractors are adopting field technology quickly and using AI for everyday work, and much of their vendor research now happens before a sales call.

This guide is for companies that sell construction hardware and field technology: robotics, drones, reality capture, equipment telematics, jobsite sensors, and prefab or modular systems. Construction management software is a different sale with a different path, covered in our companion guide on construction software. Here the prize is a pilot on one project that becomes a standard across a contractor’s jobs.

## The short version

1. Robotics went mainstream fast: in the [BuiltWorlds 2026 Equipment and Robotics Benchmarking Report](https://builtworlds.com/news/builtworlds-report-finds-jobsite-robotics-adoption-doubles-year-year/), 79% of surveyed contractors used jobsite robotics to some degree, up from 29% in 2025.
2. Reality capture is routine: [BuiltWorlds found](https://builtworlds.com/insights/2026-field-solutions-reality-capture-tech-specialty-report/) nearly 58% of organizations use it on most or every project, with drones used by 100% of those that capture.
3. Drones face a compliance shake-up: 36.1% of members in an [Associated Builders and Contractors survey](https://compactequip.com/business/abc-releases-drone-usage-survey-report/) fly drones, DJI makes the primary drone for 97.1% of them, and 53.5% name regulatory and compliance issues as the biggest hurdle.
4. Contractors already use AI at work: in the [2026 Houzz State of AI in Construction and Design Report](https://www.constructionowners.com/press-release/houzz-survey-finds-ai-adoption-soars-among-construction-and-design-pros-while-homeowners-rely-on-the-experts), 52% of construction firms used AI for everyday business tasks, up 20 percentage points in a year.
5. Jobsite technology sells into a huge market: August 2026 construction spending ran at an annual rate of $2,203.1 billion, [according to the US Census Bureau](https://www.census.gov/construction/c30/pdf/release.pdf).

## Who buys construction technology, and what is a customer worth?

Contractors buy it as a small team, and one won account can mean a fleet rollout plus years of service.

A jobsite technology purchase passes through several hands. At a general or specialty contractor, a virtual design and construction (VDC) or innovation manager usually finds and tests new tools. An operations vice president or equipment manager approves the rollout. Superintendents decide whether crews keep using it. Safety directors weigh in on anything that flies, lifts or moves near workers. On large programs, the owner or developer may require a technology, such as weekly scans for progress payments.

People, not gadgets, are the starting point. When [Trimble surveyed](https://www.trimble.com/en/blog/trimble/article/future-construction-technology-trends-contractor-survey) 1,800 civil, general and specialty contractors, workforce skills, hiring and retention topped their concerns for 2026, and if cost were no object they would invest first in AI and precise positioning. A product pitched as an answer to a labor problem fits how these buyers frame their needs.

The budget behind those decisions is large. The Census Bureau reported $773.0 billion in annual private nonresidential construction and $547.8 billion in public construction in August 2026, the work where most of this equipment is used.

There is no public benchmark for what one contractor account is worth to a field-technology company. Public data shows the shape instead:

- **Entry tickets can be small.** In the ABC survey, 62% of drone-flying contractors had invested between $1,000 and $10,000 in drones. Procore’s Kris Lengieza told [Construction Executive](https://constructionexec.com/article/contech-for-christmas-what-technologies-contractors-are-seeking-most-at-years-end/) that a quadruped robot that once cost about $100,000 can now be bought for about $10,000.
- **The value is in what follows the hardware.** [Caterpillar said](https://www.digitalcommerce360.com/2026/01/30/caterpillar-cat-ai-assistant-sales-q4-2025/) it had more than 1.6 million connected assets in 2025 and is working to grow services revenue from $24 billion toward $30 billion by 2030.

Our inference: for most field-technology companies, the first unit sold matters less than the rollout across a contractor’s projects and the data, software and service revenue attached to every unit after it.

## Where do AI assistants already sit in a contractor’s technology research?

In everyday office work and early research, while the final check still happens with peers, trade media and demos.

Construction has moved from curiosity to daily use. In the Houzz survey of more than 600 US construction and design professionals, 80% of construction firms that use AI do so daily, and sales and marketing was the leading use, at 64%. Houzz’s sample leans toward residential firms, so treat it as a signal of habit, not of commercial buying.

On larger projects, [Mastt’s 2026 survey of 108 construction project management professionals](https://www.mastt.com/research/ai-in-construction-project-management-2026) found ChatGPT was the most used AI tool, at 66.7% of respondents, followed by Microsoft Copilot at 50.9%. Mastt sells project software, and its sample is small and global.

How those habits reach purchases is clearer among technical buyers in general. In the [2026 State of Marketing to Engineers](https://advertising.globalspec.com/state-of-marketing-to-engineers-research-report) research from TREW Marketing and GlobalSpec, 69% of technical buyers used generative AI during purchasing, yet rated its answers only 4.7 out of 10 for trust. That pattern fits construction: AI starts the list, and people confirm it.

Equipment publisher [KHL](https://www.khl.com/news/construction-equipment-buyers-3-insights-suppliers-cant-ignore/8127333.article) found construction equipment decisions now usually happen within 1–3 months, and that by first contact buyers have often researched suppliers, read industry content, talked to peers and drawn up a shortlist. Manufacturers are building for this too: Caterpillar introduced a Cat AI Assistant that its chief executive said will help customers buy, maintain, manage and operate its equipment.

## Which questions do contractors ask about jobsite technology?

Questions about accuracy, compliance, compatibility and support on a specific type of project.

We wrote the examples below ourselves to show the shape of these questions. None was captured from a real contractor or a live AI answer.

| Situation | Illustrative question |
|---|---|
| Drone compliance | “Which drones can we fly on a federally funded highway job now that foreign drones are on the FCC Covered List?” |
| Reality capture | “360 camera or terrestrial LiDAR for weekly progress capture on a 300,000 square foot hospital?” |
| Robotics | “Robotic layout versus a total station crew for mechanical, electrical and plumbing layout on a high-rise” |
| Telematics | “Telematics that can read a mixed fleet of Caterpillar, Deere and Komatsu machines in one dashboard” |
| Sensors | “Wireless concrete maturity sensors that state transportation departments accept” |
| Prefab | “Bathroom pod manufacturers that can supply a 200-room hotel in Texas” |

Each question is a filter, and the filters are concrete: a tolerance, a regulation, a machine brand, an agency approval, a project type. If your website, dealer pages and trade coverage never state that your scanner meets a given accuracy, that your drone platform is compliant for federal work or that your telematics reads a competitor’s machines, an assistant has nothing to match. Building product makers face the same filters, as our guide to [how architects choose materials with AI](https://underneath.agency/resources/building-materials-specification-ai-search) shows.

## How does an AI answer become a pilot and then a fleet rollout?

Through a shortlist, a demo or site walk, a pilot on one job, then standardization and recurring service.

1. **Shortlisted.** An innovation manager asks an assistant, colleagues and trade press which products fit a problem, such as layout labor or progress reporting for an owner.
2. **Demonstrated.** The vendor shows the product on a real site, often with a dealer or reseller.
3. **Piloted.** One project tests it. Piloting is the fastest-growing stage of robotics adoption: BuiltWorlds found 32% of contractors had piloted or trialed a robot on at least one jobsite, up from 12% a year earlier.
4. **Standardized.** If the pilot proves accuracy, labor savings or safety, the contractor rolls it out to other projects and writes it into its methods. In the BuiltWorlds survey, improved accuracy was the most cited benefit of robotics, at 75%, ahead of reducing manual effort at 63% and safety at 56%.
5. **Renewed.** Data plans, software, calibration, parts and service bring revenue for the life of the fleet.

AI visibility can affect step one, and steps two and three only happen for products that make the shortlist. Proof on the jobsite decides the rest.

## Why does an assistant recommend one robot, drone or scanner over another?

Facts it can find and confirm in independent sources; the platforms document how they search, not how they choose.

**Documented by the platforms.** [OpenAI says](https://help.openai.com/en/articles/9237897-chatgpt-search) ChatGPT search typically rewrites a question into one or more targeted queries for its search partners, and that a site must allow its OAI-SearchBot crawler to be eligible. [Google says](https://developers.google.com/search/docs/appearance/ai-features) AI Overviews and AI Mode may use a “query fan-out” technique and that there are no additional requirements beyond being indexed and eligible for a snippet. Neither publishes how products are picked.

**Observed in our studies.** In [our study of the hidden searches assistants run](https://underneath.agency/research/ai-hidden-searches-study), ChatGPT ran a mean of 3.7 searches per buyer question, and in 43.8% of its answers it searched for a named publication, ranking or award. Industry awards matter here: FieldAI, the most highly rated robotics company in the BuiltWorlds survey, went on to win the 2026 Contractor’s Choice Award. [Our brand entity study](https://underneath.agency/research/brand-entity-ai-recommendations-study) found that a brand named on ten times as many independent sites had 4.7 times the odds of being recommended.

**Observed in buyer surveys.** Recognition tips close decisions: in the engineers survey, 70% of technical buyers were likely to choose the better-known brand when two solutions are technically similar. KHL found after-sales service was the most important purchasing factor for equipment, ahead of price and technology.

**Our inference for field technology.** The trust factors an assistant can check are the ones a superintendent already asks about: stated accuracy with test conditions, compliance status, the machines and software you integrate with, where service and parts come from, how data is stored and who owns it, and named contractors using the product on named projects.

## What does a construction technology company lose when AI leaves it out?

It loses shortlists formed during a reshuffle, when many contractors are choosing new tools for the first time.

We found no study measuring pilots lost to AI absence, so here is the reasoning, labeled:

- **Many contractors are still first-time buyers.** Adoption is rising quickly, and pilots grew fastest. A product that is absent when a firm writes its first shortlist may have to displace an incumbent later. That is our inference.
- **Regulation is forcing re-shopping.** On December 22, 2025, the FCC added foreign-made drones and critical components to its Covered List, which blocks new equipment authorizations, according to [DroneLife](https://dronelife.com/2025/12/27/fcc-adds-foreign-made-drones-and-components-to-covered-list-what-it-means-for-operators-and-manufacturers/). With DJI the primary drone for 97.1% of the ABC contractors who fly, many will ask what to buy next. If an assistant cannot find your compliance status, or repeats an outdated one, you may be filtered out.
- **The miss is invisible.** A contractor who never hears of you never fills in a demo form. Our guide to [why analytics miss AI visibility](https://underneath.agency/resources/why-analytics-miss-ai-visibility) explains why this rarely shows up in web reports.

## How does GEO work for a construction technology company?

Generative engine optimization (GEO) makes your product easy for AI assistants to find, describe correctly and confirm.

For field technology, the work usually covers:

1. **Specification pages in plain text.** Accuracy, range, battery life, payload, operating conditions and test methods, written as text rather than only in spec-sheet PDFs or videos.
2. **Compliance pages kept current.** Regulatory status, certifications and the date each was confirmed. When rules change, update the page the same week.
3. **Integration facts.** Which machine brands, file formats and construction software platforms you work with, stated plainly and matched on partners’ own pages. How those platforms reach contractors is covered in [our companion guide on construction software](https://underneath.agency/resources/construction-software-customers-ai-search).
4. **Named proof.** Case studies with the contractor, project type and measured result, published with permission. Our guide to [building authority for AI search](https://underneath.agency/resources/how-brands-build-authority-for-ai-search) covers why outside confirmation counts.
5. **Independent coverage.** Benchmarking reports, trade publications, association surveys and industry awards, the named sources assistants search for.
6. **Service and support facts.** Dealer and service locations, response times and training options, consistent across your site, dealer sites and directories.
7. **Video with text.** Equipment-in-action video with transcripts and clear titles, so the footage can be read as well as watched.
8. **Measurement.** Ask a fixed set of category, compliance and comparison questions across ChatGPT, Gemini, Perplexity, Copilot and Google’s AI features, repeatedly, and record which products and sources are named. Then compare the trend with demo requests and pilots.

Nothing here forces an assistant to recommend a robot, drone or scanner. The work makes yours the easiest one to verify when a contractor, or an assistant working for one, checks. Equipment makers with dealer networks face similar questions; our guide for [medical equipment companies](https://underneath.agency/resources/medical-equipment-leads-ai-search) shows a parallel path in another regulated field, and the guide for [contract manufacturers](https://underneath.agency/resources/contract-manufacturers-rfqs-from-ai-search) covers capability-led discovery.

## Which questions can’t the construction data answer yet?

It shows contractors adopting field technology and using AI, not how often AI answers decide which product they pilot.

- **Usage is not attribution.** None of the surveys here links an AI answer to a demo, pilot or purchase.
- **Many sources sell something.** BuiltWorlds runs a paid membership, Houzz and Mastt sell software, Trimble and Procore sell construction technology, and GlobalSpec sells advertising.
- **Samples are narrow.** Mastt surveyed 108 people worldwide; the Houzz sample leans residential; ABC’s respondents are its own members.
- **Our studies did not test construction questions.** Applying their findings to jobsite technology is our inference.
- **Account values are private.** Public figures show entry costs and one manufacturer’s services revenue, not typical contract values.

## Where should a construction technology company start?

Start by asking assistants the questions your best customers asked before their first pilot, and see who gets named.

That first check usually shows whether your product appears for your category and project types, whether your accuracy and compliance facts are described correctly, which publications and awards the answers lean on, and which rival products appear instead. The work then is to publish facts an assistant can verify and earn the independent coverage that confirms them.

If your growth depends on turning demos into pilots and pilots into fleet rollouts, [book a review of how assistants talk about your product](https://underneath.agency/contact). We ask the category, compliance and comparison questions contractors use, show which rival products, benchmarks and awards the answers lean on, and identify what would most likely bring more qualified demo requests. Read how our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) turns those findings into plain specification pages, current compliance status and named jobsite proof.

## Frequently asked questions

### Do contractors really use ChatGPT to research equipment?

Many use AI at work: 52% of construction firms in the Houzz survey, and ChatGPT led the tools in Mastt’s survey. We found no construction-specific data on how often AI starts an equipment purchase.

### Does the FCC Covered List mean contractors must replace their drones?

Not by itself. DroneLife reports that previously authorized fleets remain legal to fly and that the listing mainly freezes new models. Separate rules can apply on federally funded work, so contractors check each project, and vendors should state their status clearly.

### Can a startup compete with Caterpillar or DJI in AI answers?

It can be named for narrow needs, such as a specific accuracy, project type or compliance requirement, where fewer products fit. Recognition still helps larger brands, so independent proof matters more for smaller ones.

### Should we publish our pricing?

Assistants and buyers both look for prices, but the decision depends on your sales model. At minimum, state what is included, how the product is sold or leased, and what service costs cover.

## Sources

- BuiltWorlds (2026), [BuiltWorlds Report Finds Jobsite Robotics Adoption Doubles Year Over Year](https://builtworlds.com/news/builtworlds-report-finds-jobsite-robotics-adoption-doubles-year-year/)
- BuiltWorlds (2026-06-23), [2026 Field Solutions: Reality Capture Tech Specialty Report](https://builtworlds.com/insights/2026-field-solutions-reality-capture-tech-specialty-report/)
- Compact Equipment (2026-08-25), [ABC Releases Drone Usage Survey Report](https://compactequip.com/business/abc-releases-drone-usage-survey-report/)
- DroneLife (2025-12-27), [FCC Adds Foreign-Made Drones and Components to Covered List](https://dronelife.com/2025/12/27/fcc-adds-foreign-made-drones-and-components-to-covered-list-what-it-means-for-operators-and-manufacturers/)
- Houzz via Construction Owners Club (2026-09-03), [Houzz Survey Finds AI Adoption Soars Among Construction and Design Pros](https://www.constructionowners.com/press-release/houzz-survey-finds-ai-adoption-soars-among-construction-and-design-pros-while-homeowners-rely-on-the-experts)
- Mastt (2026-07-23), [State of AI in Construction Project Management 2026](https://www.mastt.com/research/ai-in-construction-project-management-2026)
- US Census Bureau (2026-10-01), [Monthly Construction Spending, August 2026](https://www.census.gov/construction/c30/pdf/release.pdf)
- Digital Commerce 360 (2026-01-30), [Caterpillar debuts “Cat AI assistant,” grows sales in Q4 2025](https://www.digitalcommerce360.com/2026/01/30/caterpillar-cat-ai-assistant-sales-q4-2025/)
- Construction Executive (2025-11-20), [Contech for Christmas: What Technologies Contractors Are Seeking Most at Year’s End](https://constructionexec.com/article/contech-for-christmas-what-technologies-contractors-are-seeking-most-at-years-end/)
- KHL (2026-07-22), [Construction equipment buyers: 3 insights suppliers can’t ignore](https://www.khl.com/news/construction-equipment-buyers-3-insights-suppliers-cant-ignore/8127333.article)
- TREW Marketing and GlobalSpec (2026), [State of Marketing to Engineers research report](https://advertising.globalspec.com/state-of-marketing-to-engineers-research-report)
- Trimble (2026-01-20), [Navigating the future of construction: key insights from contractor survey](https://www.trimble.com/en/blog/trimble/article/future-construction-technology-trends-contractor-survey)
- OpenAI Help Center (2026), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Google Search Central (2026), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [Do Wikipedia and schema make AI assistants recommend a brand?](https://underneath.agency/research/brand-entity-ai-recommendations-study)

---

This is the Markdown twin of https://underneath.agency/resources/construction-technology-demand-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How specialist consulting firms win clients through AI search"
description: "By being the firm AI can tie to a specific problem, industry and place, and confirm in outside sources, because clients now research consultants before calling."
canonical: "https://underneath.agency/resources/consulting-firms-clients-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How does a specialist consulting firm get on the shortlist when clients ask AI for advice?

By being the firm an assistant can match to a specific problem, industry and place, and then confirm in sources it trusts. Business buyers now research providers before they contact anyone, and a growing share start that research in an AI chatbot, so the list of firms worth a call is often drawn up without you.

This article is for founders and managing partners of consulting firms, especially boutique and specialist firms built around one function, one industry or one kind of problem. It covers how companies find and shortlist consultants and the alternatives they now weigh. Management consulting reputation and rankings, and technology consulting, are separate questions.

## The short version

1. Specialists are growing: [IE University’s 2026 consulting industry report](https://static.ie.edu/T%26C/Industry%20Report_Consulting_2026.pdf) puts the global management consulting market at $374 billion to $492 billion in 2026 and names boutique and specialist firms as the fastest-growing segment.
2. Research happens first: in [Responsive’s survey](https://www.digitalcommerce360.com/2026/01/05/ai-reshapes-b2b-buying-rfps/) of 350 business buyers at organizations that issue at least 10 requests for proposal a year, 90% researched before first contact with a vendor.
3. AI is in that research: about one third of those buyers named web search, peer recommendations and AI chatbots as their main ways of finding new vendors, and [48% of US buyers](https://www.responsive.io/news/buyer-intelligence-2025) said they use generative AI for vendor discovery.
4. Expertise beats price at the decision: 52% of the same buyers said industry expertise carries significant weight, ahead of price at 49%.
5. Firms know they must sell better: in [Deltek’s 2026 Clarity research](https://www.consultancy.uk/news/45503/professional-services-firms-look-to-boost-client-satisfaction-to-boost-financial-results), consulting leaders ranked more effective sales and client service (27%) and new markets (24%) as their top priorities.

## Who hires a specialist consulting firm, and what is a client worth?

Functional and business-unit leaders with a specific problem, and the value of a client usually lies in follow-on work.

The market is large and splitting in two. IE’s 2026 report puts North America at approximately 37% of global consulting revenue, operations consulting at approximately 30% and financial services clients at approximately 24%. It describes a market where the largest firms keep complex transformation work at the top, boutique and specialist firms gain share at the other end, and mid-tier generalists face the greatest pressure.

A consulting engagement is seldom bought by one executive alone. [Forrester’s 2026 State of Business Buying](https://www.forrester.com/press-newsroom/forrester-2026-the-state-of-business-buying/) found the typical purchase involves 13 internal stakeholders and nine external influencers, and procurement was a decision-maker in 53% of buying cycles. For a boutique, that means the executive who wants help is often joined by a finance lead and a procurement team that ask for alternatives.

The alternatives are widening. [Lightcast found](https://lightcast.io/blog/rise-of-fractional-leadership) at least 34,000 US workers with “fractional” in their job title in 2025, up 265% from 2019, and 97% of fractional job postings in 2026 came from medium and small companies. Finance made up 46% of those postings and business management 12%. A company that once hired a finance or operations consultancy may now compare it with a part-time executive. Finance consultancies facing that choice have their own guide on [winning CFO and deal work through AI](https://underneath.agency/resources/financial-consulting-firms-clients-ai-search).

What one client is worth varies too widely for a public benchmark, and we found none for boutique firms. Our inference from how consulting is sold: a first engagement is often a scoped project, and the value comes from the follow-on work, referrals and the case it creates for the next client in the same niche.

## How do companies find and shortlist consultants now?

They research alone first, build a long list, cut it to three or fewer, and only then talk to firms.

The clearest recent evidence comes from Responsive’s 2025 survey of buyers involved in strategic vendor selection, reported by Digital Commerce 360. It covers business purchases of all kinds, not consulting alone:

| Finding | Share |
|---|---|
| Research before first contact with a vendor | 90% |
| Use generative AI for vendor discovery (US buyers) | 48% |
| Use generative AI for vendor discovery (other regions) | 14% |
| Use generative AI for discovery at companies with 2,000+ employees | 42% |
| Start with a preferred vendor in mind | 61% |
| Open to switching from that preference | 45% |
| Name the response to a request for proposal as the most important factor | 81% |

Buyers in that survey typically began with five to eight vendors and narrowed to three or fewer before formal evaluation. The narrowing happens in research, where referrals, web search and now AI assistants all play a part.

[Gartner’s survey](https://www.gartner.com/en/newsroom/press-releases/2026-05-20-gartner-survey-finds-sixty-nine-percent-of-b-two-b-buyers-turn-to-sales-reps-to-validate-ai-generated-insights) of 645 business buyers, published in 2026, points the same way: they used an average of seven information sources in a recent purchase, and 45% used generative AI, mainly to gather information on vendors and products. And the marketplaces are following buyers into the assistants. In 2026 Clutch, a directory and review site for business service providers, [launched an app inside ChatGPT](https://www.demandgenreport.com/?p=52857) so buyers can compare providers, reviews and pricing signals without leaving the conversation.

Referrals still matter. Our inference is that AI does not replace the colleague who recommends a consultant; it is where the buyer checks that recommendation and finds the two or three other firms to compare it with. The same pattern runs across law, accounting and design, as our guide to [how professional services firms win clients](https://underneath.agency/resources/professional-services-firms-clients-ai-search) explains.

## Which questions do clients ask AI about consultants?

Questions that pair a problem with an industry, a company size or a place, plus questions about alternatives.

These prompts are our own examples of the shapes such questions take. We did not record them from real buyers, and they do not show what any assistant answers.

| Need | Example prompt |
|---|---|
| Niche match | “Pricing strategy consultants for mid-sized industrial distributors” |
| Industry experience | “Consultants who have run ERP selection for food manufacturers” |
| Place | “Supply chain consulting firms in the Midwest for a mid-sized manufacturer” |
| Size of firm | “Boutique alternative to a large firm for a post-merger integration” |
| Alternative model | “Fractional CFO or a finance transformation consultancy?” |
| Do it ourselves | “Can we build a market-entry plan with AI instead of hiring consultants?” |
| Validation | “Is this firm any good at operational turnarounds?” |

The last two rows matter most for a boutique. A buyer may ask an assistant whether to hire anyone at all, and a firm whose method and results are not described anywhere public gives the assistant nothing to weigh against doing the work in-house.

## How does an AI mention turn into a consulting engagement?

Through a longlist, a check of your people and cases, a conversation, a proposal, then follow-on work.

1. **Longlist.** The buyer asks colleagues and an assistant for firms that fit the problem. The answer is a starting list, not a decision.
2. **Check.** The buyer reads your site, your consultants’ profiles and any outside coverage. In the Responsive survey, industry expertise (52%) outweighed price (49%) and product fit (46%) in final decisions.
3. **Conversation.** A short list, usually three or fewer, is invited to talk. Gartner found 69% of business buyers prefer to validate AI-generated insights with sales reps, which for a consulting firm usually means a partner.
4. **Proposal.** A scoped proposal or a formal response to a request for proposal, often reviewed by procurement.
5. **Engagement and follow-on.** The first project leads to extensions, new workstreams and referrals in the same niche.

AI visibility mostly affects step 1 and part of step 2. The quality of your people and proposal decides the rest. But a firm left off the first list cannot be compared at all, and 45% of buyers in the Responsive survey said they were open to switching from the firm they had in mind.

## What decides whether an assistant names a boutique firm?

Specific, consistent facts and independent sources that confirm them; the assistants document how they search, not how they choose.

What is documented: [OpenAI says](https://help.openai.com/en/articles/9237897-chatgpt-search) ChatGPT search rewrites a question into one or more targeted queries sent to search providers, and that a site must allow its crawler, OAI-SearchBot, to be eligible for inclusion. OpenAI does not publish how firms are chosen for an answer.

What our studies observed, across buyer questions in several industries rather than consulting alone:

- **Independent coverage was the strongest signal we measured.** Across [our brand entity study](https://underneath.agency/research/brand-entity-ai-recommendations-study), each tenfold increase in independent sites naming a brand in the cited pages was linked to 4.7 times the odds of being recommended.
- **Named sources get cited.** In [our study of hidden searches](https://underneath.agency/research/ai-hidden-searches-study), when an assistant’s search named a source, the answer cited that source 44.0% of the time, against 8.1% when no search named it.
- **Some “best firm” lists rank their own author first.** [Our study of self-ranking lists](https://underneath.agency/research/self-promoting-best-lists-study) found that 24.2% of AI-cited numbered “best X” lists with an identifiable publisher put that publisher first. Consulting has many such lists.
- **The firms named change between runs.** In [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), only 25.2% of the brands ChatGPT named for a question came back in all five repeats, so a boutique named once may be missing the next time.

Rankings and directories are some of those outside sources. [Consultancy.uk’s 2026 ranking](https://www.consultancy.uk/news/44246/top-consulting-firms-in-the-uk-2026-the-rankings) assessed more than 1,400 firms using more than 800 consultant surveys and 650 client surveys, sorted into more than 60 areas of expertise and industries. Our inference for boutiques: being one of a few firms listed for a narrow area is more achievable, and more useful, than appearing low on a general list. How rankings and reputation work for strategy firms is covered in [our guide for management consulting firms](https://underneath.agency/resources/management-consulting-firms-clients-ai-search).

## What does a boutique lose when AI leaves it out?

Invitations to compete that it never knew existed; nobody has yet measured the loss directly.

We found no study of consulting engagements lost to AI absence, so we label the reasoning:

- **The list is cut early.** If buyers go from five to eight firms to three or fewer before talking to anyone, a firm missing from the research stage is out before the conversation. We infer that this loss shows up as proposals never requested, which no pipeline report records.
- **Referrals get checked.** A client who hears your name from a colleague may ask an assistant about you. If the answer is thin, wrong or names a rival, the referral weakens.
- **Alternatives get named instead.** With fractional executives growing and large firms offering specialist teams, an assistant asked about a problem can recommend a model, not just a firm. A boutique that does not explain why its model fits is easy to skip.
- **Wrong facts travel.** An outdated service line, a departed partner or an old office can follow you into answers. Our guide to [fixing wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) shows how to trace an error to the page behind it.

## How does GEO work for a specialist consulting firm?

Generative engine optimization (GEO) makes your firm easy for AI assistants to find, describe correctly and confirm in outside sources.

For a boutique or specialist firm, the work usually covers:

1. **A sharp niche, written down.** Pages that state the problems you solve, for which industries and company sizes, and where, in plain sentences rather than slogans. Our guide on [how HR consultancies win clients through AI](https://underneath.agency/resources/hr-consulting-firms-clients-ai-search) shows this for one function.
2. **Named people with real experience.** Consultant bios with sectors, past roles and published work, consistent with their LinkedIn profiles.
3. **Evidence of results.** Case studies with the client’s permission, or anonymized with the industry, size, problem and measured outcome stated.
4. **Consistent profiles.** The same firm name, service lines, locations and leaders on your site, directories such as Clutch and Consultancy.org, association listings and partner sites.
5. **Independent coverage.** Articles in trade publications your clients read, conference talks, podcasts, association roles and research others cite. Our guide on [how small brands get recommended by AI](https://underneath.agency/resources/how-small-brands-get-recommended-by-ai) explains why outside sources carry weight.
6. **Honest comparison content.** When a boutique fits better than a large firm, a fractional executive or doing it in-house, and when it does not.
7. **Crawl access and measurement.** Allow the crawlers the assistants document, then ask a fixed set of niche, alternative and validation questions across ChatGPT, Gemini, Perplexity, Copilot and Google’s AI features, repeatedly, and record who is named and which sources are cited.

A boutique cannot buy its way into an answer, and anyone who promises it a mention is overselling. Neighboring professional services face the same pattern, as our articles on [accounting firms](https://underneath.agency/resources/accounting-firms-clients-ai-search) and [IT service firms](https://underneath.agency/resources/it-service-firms-leads-ai-search) show. Design-led firms meet it too, covered in [how clients find engineering firms through AI](https://underneath.agency/resources/engineering-firms-project-inquiries-ai-search).

## Which questions about consulting buyers and AI does the research leave open?

It shows buyers researching with AI, not how many consulting engagements begin with an AI answer.

- **Cross-industry surveys.** The Responsive, Gartner and Forrester figures cover business buying in general. Whether consulting buyers behave the same way is our inference.
- **Interested parties.** Responsive sells software for answering requests for proposal, Deltek sells software to project-based firms, and Clutch and Consultancy.uk run directories and rankings.
- **No attribution data.** We found no public study linking AI answers to proposals requested or fees won in consulting.
- **Our studies are not consulting-specific.** They covered buyer questions in several sectors, and assistants change how they search over time.

## Where should a specialist consulting firm start?

Start by asking assistants the questions your best clients asked before they hired you.

Ask them in the client’s words: the problem, the industry, the company size, the place and the alternatives they considered. Then look at whether your firm is named, whether your niche and people are described correctly, which sources the answers rely on, and which firms or models are recommended instead. The work that follows is to state your niche where assistants read it and to earn the outside coverage that confirms it.

If your firm grows through a handful of new client relationships a year, [ask us to look at how assistants describe your practice](https://underneath.agency/contact). We research your niche the way your clients would, show which firms, fractional options and rankings the answers favor, and list the changes most likely to bring more requests for proposals. A boutique weighing the longer effort can see on our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page how niche pages, consultant profiles and outside coverage are built up.

## Frequently asked questions

### Do clients really use ChatGPT to find consultants?

Many business buyers now use generative AI in vendor research: 48% of US buyers in Responsive’s survey used it for vendor discovery. Referrals and web search remain just as common, so AI is one input among several.

### Should a boutique firm try to appear on “best consulting firms” lists?

Independent rankings and directories that assess firms by area of expertise are useful outside sources. Be wary of lists published by a competitor that ranks itself first; our study found that pattern in 24.2% of AI-cited numbered lists.

### Can a small firm compete with large consultancies in AI answers?

It can be named where the question is narrow: a specific problem, industry, size or place. Large firms still benefit from recognition, so the niche has to be stated clearly and confirmed by outside sources.

### Does this replace referrals and networking?

No. Referrals still start many engagements. AI changes how a referred firm is checked and which other firms the client compares it with.

## Sources

- IE University Talent & Careers (2026), [Industry Report: Consulting 2026](https://static.ie.edu/T%26C/Industry%20Report_Consulting_2026.pdf)
- Digital Commerce 360 (2026-01-05), [AI reshapes B2B buying and RFPs](https://www.digitalcommerce360.com/2026/01/05/ai-reshapes-b2b-buying-rfps/)
- Responsive (2025-10-15), [GenAI Overtakes Search for a Quarter of B2B Buyers](https://www.responsive.io/news/buyer-intelligence-2025)
- Consultancy.uk (2026-09), [Professional services firms look to boost client satisfaction to boost financial results](https://www.consultancy.uk/news/45503/professional-services-firms-look-to-boost-client-satisfaction-to-boost-financial-results)
- Consultancy.uk (2026-05-27), [Top Consulting Firms in the UK 2026: The Rankings](https://www.consultancy.uk/news/44246/top-consulting-firms-in-the-uk-2026-the-rankings)
- Lightcast (2026-08-20), [The Rise of Fractional Leadership](https://lightcast.io/blog/rise-of-fractional-leadership)
- Forrester (2026-01-21), [Forrester’s 2026 Buyer Insights: GenAI Is Upending B2B Buying](https://www.forrester.com/press-newsroom/forrester-2026-the-state-of-business-buying/)
- Gartner (2026-05-20), [Gartner survey finds 69% of B2B buyers turn to sales reps to validate AI-generated insights](https://www.gartner.com/en/newsroom/press-releases/2026-05-20-gartner-survey-finds-sixty-nine-percent-of-b-two-b-buyers-turn-to-sales-reps-to-validate-ai-generated-insights)
- Demand Gen Report (2026-05-12), [Clutch launches first B2B services marketplace app on ChatGPT](https://www.demandgenreport.com/?p=52857)
- OpenAI Help Center (2026), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Underneath (2026), [Do Wikipedia and schema make AI assistants recommend a brand?](https://underneath.agency/research/brand-entity-ai-recommendations-study)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [How many “best of” lists cited by AI rank their own brand first?](https://underneath.agency/research/self-promoting-best-lists-study)
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)

---

This is the Markdown twin of https://underneath.agency/resources/consulting-firms-clients-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Can AI assistants win new subscribers for a consumer app?"
description: "Yes. One in ten US adults first heard of their latest app from an AI assistant, but the app store listing and the first hour of use still close the sale."
canonical: "https://underneath.agency/resources/consumer-app-subscribers-from-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Can AI assistants win new subscribers for a consumer app?

Yes, and they already do: in a 2026 survey, one in ten US adults said an AI assistant was where they first heard of the last app they downloaded, as often as Google. But an AI recommendation only starts the sale. The app store page, the trial and the first hour of use still decide whether that person becomes a paying subscriber, so consumer software companies need to be named by AI assistants and ready for the scrutiny that follows.

## The short version

1. In [AppTweak’s](https://www.apptweak.com/en/aso-blog/ai-app-discovery-survey) August 2026 survey of 1,000 US adults, 58% had used an AI assistant to find or choose an app, and 10% first learned about their most recent download from one, level with Google (10.7%) and social media (10.4%).
2. An AI recommendation is rarely the last step: 59% of people who received one checked reviews, ratings or screenshots in the store before deciding, and only 16% downloaded on the recommendation alone.
3. Competition for attention is rising fast: [RevenueCat](https://www.revenuecat.com/state-of-subscription-apps) counted 14,700+ new subscription apps launched in January 2026 alone, up from about 2,000 a month in January 2022.
4. Each subscriber is worth little at first, so acquisition has to be cheap: in the same RevenueCat data, the median North American app turned 2.6% of downloads into payers within 35 days, and a payer brought in $32 in the first year.
5. New products are hard for assistants to find unprompted: in a [study of 112 Product Hunt launches](https://arxiv.org/abs/2601.00912), ChatGPT recognized them 99.4% of the time when asked by name but surfaced them in only 3.32% of open discovery questions.

## Who chooses a consumer subscription today, and what is a subscriber worth?

An individual decides alone, often within minutes, and a typical subscriber is worth tens of dollars a year.

Consumer subscription software, or B2C SaaS, covers apps such as language tutors, fitness and meditation apps, photo editors, VPNs, note-taking tools and AI assistants themselves. There is no buying committee and no sales call. One person searches, compares, downloads, tries and pays, usually on a phone. Note-taking and task apps also sell to work users, a contest covered in [how productivity tools win users from AI](https://underneath.agency/resources/productivity-software-ai-search-growth).

App spending is large and still climbing. [Sensor Tower’s State of Mobile 2026](https://sensortower.com/press/press-release-boosted-by-gen-ai-services-consumers-spent-more-money-in-apps-than-games-for-first-time) found that in-app purchases reached $167 billion globally in 2025, and that spending on non-game apps passed games for the first time, helped by generative AI services.

Per subscriber, though, the numbers are small. RevenueCat’s 2026 report, built on more than 115,000 apps, gives these medians:

| Measure (median app) | Value |
|---|---|
| Downloads that became payers within 35 days, North America | 2.6% |
| Revenue per payer in the first year, North America | $32 |
| Revenue per payer in the first year, global | $23 |
| Download-to-paid within 35 days, hard paywall vs. freemium | 10.7% vs. 2.1% |

Retention is the second half of the value. RevenueCat’s summary of the report says about 72% of annual subscribers cancelled in their first year in 2026, up from about 56% in 2025, and the first month accounts for 35% of annual cancellations. A subscriber found cheaply and kept for years is the whole business model. Brands that ship products on subscription face the same math, as our guide to [winning subscribers through AI search](https://underneath.agency/resources/subscription-ecommerce-customers-ai-search) shows.

## Where do assistants fit between a need and an app store search?

At the start of the search and in comparisons, alongside the app stores rather than in place of them.

The app stores are still the main door. [Apple](https://ads.apple.com/app-store) says 70% of App Store visitors use search to discover apps, and almost 65% of downloads happen directly after a search. That is documented by the platform that runs the store.

AI assistants now sit just before that door. In AppTweak’s survey, 34% used an AI assistant at some point before their most recent app download, and 53% would consider an AI assistant the next time they look for an app, second only to the app stores at 62.8%. Among people who had used AI while choosing an app, the most common reason was to find an app for a specific need, followed by learning about an app they had heard of and comparing apps they were already considering. AppTweak sells app marketing software, so it has an interest in the result, and the answers are self-reported.

Broader surveys of US consumers show the same drift toward AI:

- [Menlo Ventures’ 2026 State of Consumer AI](https://menlovc.com/wp-content/uploads/2026/09/menlo_ventures_consumer_ai_report-2026.pdf), a survey of 5,067 US adults, found 25% of Americans use AI every day, up from 19% a year earlier. Among shoppers who use AI, 43% used it to narrow a shortlist and 37% made a final purchase decision based on its recommendation.
- [Goodwater Capital’s 2026 consumer survey](https://www.goodwatercap.com/insights/2026-consumer-survey/) found 19% of Americans now begin a search with an AI assistant; among daily AI users, 37% start with AI against 32% with a search engine.

The assistants are also becoming places where apps live. In October 2025 [OpenAI launched apps inside ChatGPT](https://openai.com/index/introducing-apps-in-chatgpt/), starting with partners including Canva, Coursera, Spotify and Zillow, and said developers could reach “over 800 million ChatGPT users.” OpenAI documents that ChatGPT “can also suggest apps when they’re relevant to the conversation.”

## What do consumers ask AI assistants before they subscribe?

Questions about a need, a shortlist, a comparison, a price and whether the app is safe to trust.

The examples below are illustrative, written by us to show the buying stages. They are not captured from real users.

| Stage | Illustrative prompt |
|---|---|
| Need | “What’s a good app to learn Spanish in 15 minutes a day?” |
| Shortlist | “Best budgeting apps for a couple who both get paid biweekly” |
| Alternatives | “Cheaper alternatives to Headspace that still have sleep stories” |
| Comparison | “Is Duolingo or Babbel better for real conversations?” |
| Price | “How much is NordVPN per month, and is the two-year plan worth it?” |
| Trust | “Is this photo editing app safe, and is it easy to cancel?” |

These questions rarely become a single search. A request like “cheaper alternatives to Headspace” can fan out: Google says AI Overviews and AI Mode [may use a “query fan-out” technique](https://developers.google.com/search/docs/appearance/ai-features), running several related searches before answering. In [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study) of 80 buyer questions, ChatGPT looked for reviews in 46.2% of its answers. Our inference: for consumer apps, those reviews include store ratings, review sites and community threads.

Price questions carry a specific risk for subscriptions. In [our pricing study](https://underneath.agency/research/ai-pricing-accuracy-study) of 45 software and subscription products, 61.9% of the plan prices that four assistants quoted were fully faithful to the official page. Another 3.8% had the right amount but dropped a condition, usually by presenting an annual-billing price as the monthly price. One answer gave NordVPN’s Basic monthly price for its Complete plan.

## How does an AI recommendation become a paying subscriber?

Through four steps: named in the answer, checked in the store, downloaded, then converted during a short trial.

1. **Named.** The assistant puts the app on a short list for a need the person described.
2. **Checked.** In AppTweak’s survey, after an AI recommendation, 35.6% looked at reviews, ratings or screenshots in the store, and 10.6% searched for opinions on Google, Reddit or YouTube. Only 9.6% downloaded right away.
3. **Downloaded.** Asked to choose between an app an AI recommended and a different app at the top of store search, 17.3% picked the AI recommendation and 7.7% the store’s top result. Most (69.8%) said they would check reviews first. AppTweak notes this is a stated preference, not observed behavior.
4. **Converted.** The trial then has to work quickly. RevenueCat found 55.4% of all 3-day trial cancellations happen on the day the trial starts. Trials of 17 to 32 days converted at a median of 42.5%, against 25.5% for trials of under four days.

The money arrives at step four, but the first three steps decide who reaches it. With a $32 first-year payer and a 2.6% download-to-paid rate, a paid download has to cost very little to pay back. We infer that an unpaid recommendation from an assistant is valuable mainly because it lowers that cost, not because it replaces the trial.

## Why does an assistant suggest one app and not its rival?

Mostly evidence about you on other sites, found through search; nobody outside the platforms knows the exact rules.

What is documented: OpenAI says ChatGPT can suggest an app from its own app directory when it fits the conversation, and Google says its AI features can run several related searches. Neither publishes how it ranks consumer apps in an answer.

What has been observed:

- **Recognition is not recommendation.** In Sharma’s Product Hunt study, ChatGPT recognized new products 99.4% of the time when asked by name and Perplexity 94.3%, but discovery questions surfaced them only 3.32% and 8.29% of the time. Perplexity turned up the launches with more referring domains and more Reddit presence more often. It is one author’s test of a small ChatGPT model through its developer interface, so app makers should read it as a signal rather than a rule.
- **The list changes between asks.** When [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study) repeated each question five times, only 25.2% of the brands ChatGPT named showed up every time.
- **People use several assistants.** Menlo found the average AI user now uses 3.0 general AI assistants, up from 2.2 a year earlier. In AppTweak’s survey, 39.2% of US adults had used ChatGPT and 30.9% Gemini when choosing an app.

Our inference for consumer subscriptions: the trust factors are store ratings and review volume, independent “best app” roundups from known publications, community threads, accurate pricing and cancellation terms, and clear privacy information. Big brands start with an advantage; we cover that in [do AI assistants favor big brands](https://underneath.agency/resources/do-ai-assistants-favor-big-brands).

## What does an app lose when assistants recommend its rivals instead?

A missing app loses the cheapest new subscribers in a market where new launches keep multiplying.

The evidence is indirect, so we label the reasoning:

- **The field is getting crowded.** RevenueCat counted 14,700+ new subscription apps launched in January 2026, yet apps launched before 2020 still earn 69% of subscription revenue. Established names hold the advantage, and a new app that assistants do not surface has to buy every download instead.
- **The winners pull away.** RevenueCat found the median app grew monthly recurring revenue by 5.3% a year, while the top tenth grew by 306% or more. We infer that an extra low-cost channel matters most for apps in the middle.
- **AI categories face the sharpest test.** In RevenueCat’s report, AI-powered apps made 41% more revenue per payer but saw subscribers churn 30% faster. When an assistant can do the job itself, Menlo warns, an app’s competition becomes “the general AI assistant the consumer already has open.” Our guide to [how AI software companies stand out](https://underneath.agency/resources/ai-software-companies-in-ai-search) covers that crowded field.

What no one has measured is how many subscribers a given app loses by being absent from AI answers. Treat any precise figure as a guess.

## How does GEO work for a consumer subscription business?

Generative engine optimization (GEO) makes your app easy for AI assistants to find, describe correctly and back with outside evidence.

For a consumer subscription business, that usually means:

1. **Clear identity.** One consistent name, category and description across the website, both app stores, and profiles, so assistants do not confuse you with a similar app. If an assistant already describes you wrongly, see [how to fix wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers).
2. **Answer-ready pages.** Plain pages for each core need (“learn Spanish on a commute”), honest comparison pages, and a pricing page that states the monthly and annual price, the trial length and how to cancel.
3. **Third-party coverage.** Earned reviews in publications that write “best app for…” roundups, creator reviews on YouTube, and genuine participation in the communities where your users talk.
4. **Store reputation.** Ratings, review replies and screenshots matter twice: assistants may draw on them, and 59% of people check the store after an AI recommendation.
5. **Presence across assistants.** Check ChatGPT, Gemini and Google’s AI features at least. In March 2026, [Statcounter](https://gs.statcounter.com/press/google-gemini-overtakes-perplexity-to-become-second-largest-source-of-ai-chatbot-referrals-to-websites) measured ChatGPT at 78.16% of AI chatbot referrals to websites and Gemini at 8.65%, with Gemini’s share growing fast.
6. **Measurement.** Ask a fixed set of need-based questions many times and track how often you are named, then add a “Where did you hear about us?” question to onboarding, because [analytics miss most AI visibility](https://underneath.agency/resources/why-analytics-miss-ai-visibility).

Nothing here forces an assistant to name your app. It raises the odds that, when it looks, it finds strong and accurate evidence.

## What can’t the evidence tell a consumer app yet?

It shows people use AI to find apps, not the subscription revenue any one app gains.

- **Self-reported surveys.** AppTweak, Menlo and Goodwater asked people what they did. Those answers can differ from behavior, and AppTweak’s sample came from an online research panel.
- **Vendor interests.** AppTweak sells app marketing tools and RevenueCat sells subscription software. Their data are large but not neutral.
- **No public conversion data from AI referrals for consumer apps.** We found no consumer subscription company that publishes trial or paid rates for visitors arriving from AI assistants. For broader evidence, see [does showing up in AI answers drive business results](https://underneath.agency/resources/does-ai-visibility-drive-business-results).
- **App picks are unexplained.** Only the platforms know how an assistant settles on which apps to suggest, and the picks shift between runs.

## How can a subscription app see whether AI assistants send it trial starts?

Ask the questions your future subscribers ask, in the assistants they use, and note which apps get named.

That first check usually shows three things: whether you appear for your core needs, which competitors and sources appear instead, and whether your prices and trial terms are quoted correctly. From there, the work is to fix the facts, earn the third-party evidence and make the store page ready for the people an assistant sends.

If your growth depends on cheaper downloads and better trial conversion, [ask us to review where your app stands in AI answers](https://underneath.agency/contact). We will map where your app appears in AI answers, why competitors appear instead, and which changes are most likely to bring more qualified people into your trial. What happens after that review, from accurate pricing and cancellation pages to earned app roundups, is set out on our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page.

## Frequently asked questions

### Do people really use ChatGPT to find apps?

Many do. In AppTweak’s August 2026 survey of 1,000 US adults, 58% had used an AI assistant at least occasionally to find or choose an app, and 21% did so half the time or more.

### Does an AI recommendation replace app store optimization?

No. After an AI recommendation, 59% of people went to the store to check reviews, ratings or screenshots before deciding, so the store listing still has to close the download.

### Which AI assistants matter most for consumer apps?

ChatGPT and Gemini, on current evidence. In AppTweak’s survey, 39.2% of US adults had used ChatGPT and 30.9% Gemini when choosing an app; Claude, Copilot and Perplexity were in single digits.

### Will an assistant suggest an app that has only just launched?

Rarely at first. Across 112 Product Hunt launches in one study, ChatGPT knew them almost every time when asked by name but brought them up in only 3.32% of open discovery questions.

## Sources

- AppTweak, Micah Motta (2026-09-14), [The state of AI app discovery: what 1,000 U.S. adults told us about how they find apps](https://www.apptweak.com/en/aso-blog/ai-app-discovery-survey)
- RevenueCat (2026), [State of Subscription Apps 2026](https://www.revenuecat.com/state-of-subscription-apps)
- RevenueCat (2026), [The State of Subscription Apps in 10 minutes: lessons, trends, and benchmarks for 2026](https://www.revenuecat.com/blog/growth/subscription-app-trends-benchmarks-2026)
- Sensor Tower (2026-01-21), [Boosted by Gen AI services, consumers spent more money in apps than games for first time](https://sensortower.com/press/press-release-boosted-by-gen-ai-services-consumers-spent-more-money-in-apps-than-games-for-first-time)
- Apple (2026), [Apple Ads: App Store](https://ads.apple.com/app-store)
- Menlo Ventures (2026-09-15), [2026: The State of Consumer AI](https://menlovc.com/wp-content/uploads/2026/09/menlo_ventures_consumer_ai_report-2026.pdf)
- Goodwater Capital (2026-06-08), [Goodwater’s 2026 US Consumer Survey](https://www.goodwatercap.com/insights/2026-consumer-survey/)
- OpenAI (2025-10-06), [Introducing apps in ChatGPT and the new Apps SDK](https://openai.com/index/introducing-apps-in-chatgpt/)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- Statcounter (2026-04-02), [Google Gemini overtakes Perplexity to become second largest source of AI chatbot referrals to websites](https://gs.statcounter.com/press/google-gemini-overtakes-perplexity-to-become-second-largest-source-of-ai-chatbot-referrals-to-websites)
- Sharma (2025), [The Discovery Gap: How Product Hunt Startups Vanish in LLM Organic Discovery Queries](https://arxiv.org/abs/2601.00912), arXiv:2601.00912.
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [How faithfully do AI assistants quote software prices?](https://underneath.agency/research/ai-pricing-accuracy-study)
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)

---

This is the Markdown twin of https://underneath.agency/resources/consumer-app-subscribers-from-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How electronics brands win shoppers who compare specs in AI"
description: "By being the model AI names when shoppers compare specs and prices: AI answers lean on independent lab reviews, retailer data and accurate prices."
canonical: "https://underneath.agency/resources/consumer-electronics-sales-from-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How do consumer electronics brands and retailers win shoppers who compare specs in AI?

By making sure your models are the ones an AI assistant names when a shopper describes a need, a budget and a few specs, and that the price and details it quotes are right. In electronics that answer is built mostly from independent reviews and retailer listings, not from your own site, so the work is as much about earned coverage and clean product data as about your pages. The prize is the unit sale, and increasingly the trade-up to a better model.

## The short version

1. Electronics is the biggest online category: [Adobe](https://news.adobe.com/news/2026/01/adobe-holiday-shopping-season) counted $59.8 billion spent online on electronics in the 2025 US holiday season, up 8.2%, and named electronics among the categories where shoppers used AI most.
2. The market grows by trading up, not by volume: the [Consumer Technology Association](https://www.ces.tech/press-releases/cta-despite-tariffs-and-economic-headwinds-us-consumer-tech-revenue-to-hit-565-billion-in-2026) projects $565 billion in US consumer tech revenue in 2026, up 3.7%, while unit shipments grow just 0.7%.
3. AI shoppers in this category buy: [Adobe found](https://blog.adobe.com/en/publish/2025/03/17/adobe-analytics-traffic-to-us-retail-websites-from-generative-ai-sources-jumps-1200-percent) that conversion from AI traffic was highest in electronics and jewelry, because shoppers use AI to narrow options by screen size, resolution and price.
4. AI answers about electronics are built on reviewers: in one study, 92.1% of the sources US AI search used for consumer electronics rankings were independent “earned” sites such as TechRadar, Tom’s Guide and RTINGS ([Chen and colleagues](https://arxiv.org/abs/2509.08919)).
5. Retailers are building their own assistants: Amazon says its Rufus assistant was used by 300 million customers in 2025 and drove nearly $12 billion in incremental annualized sales ([Modern Retail](https://www.modernretail.co/technology/amazon-says-its-ai-shopping-assistant-is-gaining-traction-with-rufus-users-up-115/)).

## Who buys consumer electronics today, and what is one sale worth?

A deliberate, deal-aware shopper buys, and each sale is worth more when that shopper trades up.

The electronics buyer is rarely impulsive. A TV, a laptop, a camera or a pair of noise-canceling headphones is a planned purchase, researched on spec sheets and reviews and often timed to a sale. Best Buy’s chief executive described the customer on the retailer’s March earnings call as one who is “still spending, but is value-focused and attracted to sales moments,” according to [PYMNTS](https://www.pymnts.com/earnings/2026/best-buy-bets-ai-consumers-avoid-big-ticket-items/). Adobe’s holiday data shows how much price drives timing: discounts on electronics peaked at 30.9% off listed price and on televisions at 24.3%.

What a sale is worth depends on which model the shopper lands on. The industry is not growing by selling more boxes: CTA forecasts unit shipments up just 0.7% in 2026, while revenue rises 3.7% to $565 billion. Growth comes from premium features and from shoppers moving up a tier. Adobe saw exactly that last holiday season: the share of units sold for the most expensive goods rose 56% in electronics compared with the rest of the year. For a maker, being named as the better choice in a comparison is worth the price gap between two models on every unit. For a retailer, it is the margin on the item plus any protection plan, installation or accessory sold with it. Categories with repeat orders change that math, because one recommendation can start years of purchases, as our guide to [pet products recommended by AI](https://underneath.agency/resources/pet-product-brands-ai-search) shows.

Retailer scale shows what is at stake. Best Buy reported second-quarter revenue of $9.78 billion and comparable sales up 4.1%, according to [HomePage News](https://www.homepagenews.com/retail-articles/best-buy-advances-in-agentic-ai-after-solid-q2-earnings-comps/).

## Where do AI assistants already sit in an electronics purchase?

At the research and narrowing stage, and increasingly at price tracking and checkout.

The evidence comes from retailers, platforms and analytics firms, not from guesses:

- **Research and narrowing.** In Adobe’s survey of 5,000 US consumers, 87% of those who had used AI for shopping said they were more likely to use it for larger or more complex purchases, which describes most electronics. Adobe’s traffic data adds that AI visitors convert best in electronics and jewelry, and gives televisions as the example: shoppers use AI to narrow by screen size, resolution and price.
- **Buyer’s guides.** [OpenAI documents](https://openai.com/index/chatgpt-shopping-research/) that its shopping research feature in ChatGPT “performs especially well in detail-heavy categories like electronics,” and that it looks across the internet for “price, availability, reviews, specs, and images.” One of its own example requests is a gaming laptop “under $1000 with a screen that’s over 15 inches.”
- **Price watching.** [Google documents](https://blog.google/products/shopping/agentic-checkout-holiday-ai-shopping/) that shoppers can track an item’s price and, with eligible merchants, have Google buy it once the price falls within budget. Its AI Mode draws on a Shopping Graph of more than 50 billion product listings, 2 billion of which are updated every hour.
- **Retailer assistants.** Amazon says customers who use Rufus are 60% more likely to complete a purchase. Best Buy launched Ask Blue, a conversational shopping and support assistant that, in the words of incoming chief executive Jason Bonfig, “combines product knowledge, support, resources, customer reviews, availability and pricing.”

Traffic is growing fast from a small base. Adobe measured a 693.4% rise in visits to US retail sites from AI tools in the 2025 holiday season and named electronics among the categories where these services were used most. Toys were on that list too, as our guide to [how toy brands get found through AI](https://underneath.agency/resources/toy-brands-product-discovery-ai-search) explains.

## Which questions do electronics shoppers ask AI assistants?

Spec-and-budget questions, head-to-head model comparisons, “is it worth it” questions and timing questions.

We wrote the prompts below to illustrate the kinds of questions electronics shoppers ask; they are not observed data:

- Spec and budget: “Best 65-inch TV under $1,000 for a bright living room.”
- Head to head: “Sony or Bose noise-canceling headphones for long flights?”
- Use case: “Which laptop for 4K video editing with at least 32 GB of memory?”
- Worth the upgrade: “Is an OLED TV worth the extra money over a mini-LED for movies?”
- Compatibility: “Will this soundbar work with my TV’s eARC port?”
- Timing and price: “Will this TV be cheaper on Black Friday?” or “Is this a good price for this camera?”

Two things make these questions different from most industries. They are full of numbers, so an answer that gets a spec or price wrong can push the shopper toward a rival model. And they often name two or three models, so the assistant is choosing between specific products, not just brands. A real study example shows the stakes: asked which phone has the best camera, Google’s AI Overview answered that “the Oppo Find X9 Ultra is widely rated as the best overall camera phone,” while ChatGPT said “my pick right now is … iPhone 17 Pro Max” ([Uberti-Bona Marin and colleagues](https://arxiv.org/abs/2609.18729)). A phone choice often ties a household to one ecosystem, the subject of [how device brands win ecosystem decisions](https://underneath.agency/resources/consumer-tech-brands-ai-search).

## How does an AI answer turn into an electronics sale?

Through a shortlist of models, a click to a retailer, and a purchase timed to price.

For a **maker**, the path is: a shopper describes a need; the assistant names a few models with reasons; the shopper checks a review or two and a retailer page; then buys, often on a promotion. The sale may land at Best Buy, Amazon or the brand’s own store, so a maker’s AI visibility shows up as retail sell-through, not as traffic to its site. That is our inference from how electronics is sold, not a measured split.

For a **retailer**, there are three routes. The assistant can cite its product page; in one study, Perplexity’s electronics answers drew on BestBuy.com as a dominant source (Chen and colleagues). Its catalog can sit inside the assistant: Best Buy said it teamed with OpenAI to bring its product catalog into ChatGPT and is working with Google on “agentic shopping” so customers can buy through the Gemini app and Google Search. And its own assistant can answer the question on its site, as Rufus and Ask Blue do.

The trade-up is where the money concentrates. When an answer explains why a step-up model is worth it for the shopper’s use, it can move the sale up a tier. We infer that this is where AI answers matter most for revenue in electronics, given that growth in the category comes from premium models rather than more units.

## Which sources decide whether a TV, laptop or headphone is named?

Mostly independent reviewers and retailer listings, as studies observe; the platforms document only part of their methods.

**Documented by the platforms.** OpenAI says shopping research is trained to “read trusted sites, cite reliable sources,” that results are “based on publicly available retail sites,” and that it avoids “low-quality or spammy sites.” Google says AI Mode shopping responses bring together price, reviews and inventory information from its Shopping Graph. Neither publishes how it chooses one model over another.

**Observed in studies.**

- *Reviewers dominate.* [Chen and colleagues](https://arxiv.org/abs/2509.08919) found that for consumer electronics in the US, AI search drew 92.1% of its sources from earned sites, while Google’s results leaned more on brand content at 32.9%. In a second test, Claude’s most-used domains were TechRadar, Tom’s Guide and RTINGS; Perplexity mixed in more brand and retail sources, at about 31.6%, with YouTube and BestBuy.com prominent. Lab-testing reviewers such as RTINGS and Wirecutter are therefore part of the shelf an AI shops from.
- *Reviews are fairly recent.* In the same study, the reviews Claude cited for electronics had a mean age of about 117 days and a median of 62. Our own [freshness study](https://underneath.agency/research/ai-source-freshness-study) found that pages published in the last 90 days made up 17.4% to 22.6% of each assistant’s dated citations, against 6.9% of Google’s top 10.
- *Hard facts beat hype.* Controlled tests summarized in [our article on what drives AI product recommendations](https://underneath.agency/resources/what-drives-ai-product-recommendations) found that ratings, prices and reviews outweigh brand name when assistants can see them.
- *Electronics is harder to fake.* In tests of planted fake pages, [Luo and Chen](https://arxiv.org/abs/2606.13610) found technical products such as phones and PCs among the least exposed categories: with English search results, the planted fake brand fooled the models in 43% of smartphone tests on average, against 87% for San Francisco restaurants.

**Trust factors specific to electronics.** We infer from the sources above that assistants weigh measured performance (brightness, battery life, noise canceling), consistent model names across retailers, current price and stock, review volume and ratings, and recency of reviews. Model-year naming is a real risk: a 2025 and a 2026 version of the same TV can differ in price and panel, and an answer that mixes them misleads the shopper.

## What does an electronics company lose when its models are missing or misdescribed?

Unit sales and trade-ups at the moment of choice, and margin when prices are quoted wrong.

The cost of absence follows from how few models an answer names. Uberti-Bona Marin and colleagues found ChatGPT expressing a first-person product preference in 79% of product-recommending responses, against 7% for Gemini and 2% for AI Overviews. A shopper told “my pick is” one phone is unlikely to research a model the assistant never mentioned. We infer that, in a category growing through premium models, losing that moment costs both the unit and the higher-priced unit.

Misdescription is the second cost. OpenAI itself warns that shopping research “might make mistakes about product details like price and availability.” In software, [our pricing study](https://underneath.agency/research/ai-pricing-accuracy-study) found that only 61.9% of plan prices quoted by four assistants were fully faithful to the official page; no comparable measurement exists for electronics, but with deep discounts and model-year changes we expect price errors to matter at least as much. The wider cost of doing nothing is laid out in [what happens if you skip GEO](https://underneath.agency/resources/what-happens-if-you-skip-geo).

## How does GEO work for an electronics maker or retailer?

By making your models easy for AI systems to find, compare and trust, without any promise of placement.

Generative engine optimization, GEO, for electronics covers:

- **Product and entity clarity.** One canonical name per model and model year, the same on your site, retailer listings and spec sheets, so an assistant does not mix two products.
- **Complete, comparable specs.** Measured, plainly stated specs (brightness, ports, battery life, weight, warranty) on pages that can be read without heavy scripts. Our guide to [product content AI shopping assistants prefer](https://underneath.agency/resources/product-content-ai-shopping-assistants-prefer) covers what tests reward and what backfires.
- **Independent review coverage.** Getting review units to the lab testers, editorial reviewers and YouTube channels that AI answers draw on, and making corrections when their data is out of date. This is digital PR aimed at the earned layer, not paid placement; our note on [which pages to target](https://underneath.agency/resources/best-of-lists-ai-recommendations) helps pick the reviewers and roundups first.
- **Retailer data.** Accurate titles, specs, prices and stock in the feeds behind Google’s Shopping Graph, ChatGPT shopping and retailer assistants such as Rufus and Ask Blue. OpenAI offers merchants an allowlisting process to appear in shopping research. Our guide to [getting products chosen by AI shopping assistants](https://underneath.agency/resources/ecommerce-brands-ai-shopping-assistants) covers keeping listings consistent across channels.
- **Ratings and reviews.** Volume and recency of customer reviews on retailer sites, where assistants can see them.
- **Launch timing.** New models need fresh coverage quickly; [our article on new products](https://underneath.agency/resources/why-chatgpt-misses-new-products) explains why assistants often miss them.
- **Measurement across engines.** The Uberti-Bona Marin study found that ChatGPT and Gemini shared only 5.4% of displayed domains for the same question, and another study found brand overlap across repeated runs of only 48% for consumer electronics ([Schulte and colleagues](https://arxiv.org/abs/2604.07585)). Visibility must be measured over many runs and engines, not one screenshot.

## What can’t anyone yet measure about AI and electronics sales?

How much of an electronics brand’s sales AI answers cause, and how stable each assistant’s choices are.

- No public data splits electronics sales by AI influence. Adobe measures traffic and conversion; Amazon’s Rufus figures are Amazon’s own estimate of incremental sales.
- The source studies use ranking-style questions; real shoppers ask messier, constraint-heavy questions, and results may differ.
- Platforms change fast. OpenAI’s shopping research, Google’s agentic checkout and Best Buy’s ChatGPT catalog are all less than a year old.
- We have no measurement of how often assistants quote wrong electronics prices or mix model years; our price figure comes from software.
- Advertising is arriving inside assistants, and how paid placements will sit next to organic answers is not yet documented for electronics.

## Where should a consumer electronics company start?

Start by checking which models AI assistants name for your top spec-and-price questions, and why.

Pick the 30 to 50 questions that drive your unit sales and trade-ups, from “best TV under $1,000” to head-to-head comparisons with your closest rival, and test them across ChatGPT, Gemini, Google AI Mode, Perplexity, Copilot and retailer assistants, several times each. Note which models are named, which reviewers and retailers are cited, and whether prices and specs are right. That shows whether the gap is review coverage, product data or naming. To have us do it with you, [reach our team here](https://underneath.agency/contact): we map where your models appear for the questions that decide electronics purchases and build a plan to improve how your products are found and described, so more shortlists, and more trade-ups, include you. Our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page describes how that plan is carried out, covering model naming, retailer feeds and coverage from the reviewers assistants cite.

## Frequently asked questions

### Do AI assistants recommend electronics from brand websites?

Rarely as their main source. In one study, 92.1% of US AI search sources for consumer electronics were independent earned sites; Perplexity drew more on brand and retail pages than ChatGPT or Claude ([Chen and colleagues](https://arxiv.org/abs/2509.08919)).

### Does AI shopping traffic actually buy electronics?

Adobe reports that conversion from AI traffic is highest in electronics and jewelry, and Amazon says Rufus users are 60% more likely to complete a purchase. Neither figure is a controlled measurement of AI’s effect on a brand’s sales.

### Can we pay to be recommended?

Not in organic answers. OpenAI says shopping research results are organic and based on publicly available retail sites. Paid placements in assistants are a separate, newer channel.

### How quickly do new models appear in AI answers?

It varies. Assistants that search the web cite recent pages more than Google’s top results do, but new launches often go missing until reviews exist; see [why ChatGPT misses new products](https://underneath.agency/resources/why-chatgpt-misses-new-products).

## Sources

- Adobe (2026-01-07), [Holiday Shopping Season Drove a Record $257.8 Billion Online with Consumers Embracing Generative AI Tools](https://news.adobe.com/news/2026/01/adobe-holiday-shopping-season)
- Adobe (2025-03-17), [Traffic to U.S. retail websites from generative AI sources jumps 1,200 percent](https://blog.adobe.com/en/publish/2025/03/17/adobe-analytics-traffic-to-us-retail-websites-from-generative-ai-sources-jumps-1200-percent)
- Consumer Technology Association (2026-01-04), [U.S. Consumer Tech Revenue to Hit $565 Billion in 2026](https://www.ces.tech/press-releases/cta-despite-tariffs-and-economic-headwinds-us-consumer-tech-revenue-to-hit-565-billion-in-2026)
- OpenAI (2025-11-24), [Introducing shopping research in ChatGPT](https://openai.com/index/chatgpt-shopping-research/)
- Google (2025-11-13), [Let AI do the hard parts of your holiday shopping](https://blog.google/products/shopping/agentic-checkout-holiday-ai-shopping/)
- Modern Retail (2026-04-30), [Amazon says its AI shopping assistant is gaining traction](https://www.modernretail.co/technology/amazon-says-its-ai-shopping-assistant-is-gaining-traction-with-rufus-users-up-115/)
- PYMNTS (2026-03-03), [Best Buy Bets on AI as Consumers Avoid Big-Ticket Items](https://www.pymnts.com/earnings/2026/best-buy-bets-ai-consumers-avoid-big-ticket-items/)
- HomePage News (2026-08-27), [Best Buy Advances in Agentic AI After Solid Q2 Earnings, Comps](https://www.homepagenews.com/retail-articles/best-buy-advances-in-agentic-ai-after-solid-q2-earnings-comps/)
- Chen and colleagues (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919)
- Uberti-Bona Marin and colleagues (2026), ["If I Had to Buy Just ONE: Galaxy S26 Ultra": Auditing AI-Generated Product Recommendations](https://arxiv.org/abs/2609.18729)
- Luo and Chen (2026), [One Polluted Page Is Enough: Evaluating Web Content Pollution in LLM Recommenders](https://arxiv.org/abs/2606.13610)
- Schulte and colleagues (2026), [Don’t Measure Once: Measuring Visibility in AI Search (GEO)](https://arxiv.org/abs/2604.07585)
- Underneath (2026), [AI source freshness study](https://underneath.agency/research/ai-source-freshness-study)
- Underneath (2026), [AI pricing accuracy study](https://underneath.agency/research/ai-pricing-accuracy-study)

---

This is the Markdown twin of https://underneath.agency/resources/consumer-electronics-sales-from-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How consumer fintech apps win customers from AI answers"
description: "By being named, and described accurately, when people ask AI which app to use and whether it is safe. In money questions, trust facts decide the answer."
canonical: "https://underneath.agency/resources/consumer-fintech-apps-customers-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How do consumer fintech apps win customers when people ask AI about money?

By being named in the answer when someone asks which budgeting, payment or credit-building app to use, and by being described accurately when they ask whether it is safe. Many people already take money questions to AI assistants, and in finance those answers lean on established, well-ranked sources. For a consumer fintech app, AI visibility is won with clear, verifiable facts about fees, protections and the bank behind the account.

## The short version

1. Money questions have moved into AI: in an [Intuit Credit Karma survey](https://stacker.com/stories/personal-finance-investing/rise-fin-ai-why-americans-are-trusting-generative-ai-their), 66% of Americans who had used generative AI said they had used it to seek financial advice, and 34% of those asked about budgeting and expense management.
2. Younger customers lead: [Experian](https://cuinsight.com/gen-z-millennials-are-using-ai-for-personal-finance-advice-report-finds) found 67% of Gen Z and 62% of millennials it surveyed use AI to help with their personal finances.
3. Google shows AI answers on most finance searches: in [our study](https://underneath.agency/research/ai-overviews-frequency-study) of US keywords, financial services and insurance searches triggered an AI Overview 88.0% of the time.
4. A customer is worth a lot once the app becomes primary: [Chime](https://stocks.observer-reporter.com/observerreporter/article/bizwire-2026-8-5-chime-reports-second-quarter-2026-financial-results) reported average revenue per active member of $260 a year across 10.4 million active members.
5. Safety facts are regulated: the [FDIC’s rule](https://www.federalregister.gov/documents/2024/01/18/2023-28629/fdic-official-signs-and-advertising-requirements-false-advertising-misrepresentation-of-insured) on deposit insurance misrepresentation, with compliance required by January 1, 2025, warns that fintech arrangements can confuse consumers about whether their money is insured.

We write here about how money apps get found and described by AI assistants. Nothing in this guide is financial, legal or regulatory advice, so check fee, insurance and lending wording with your compliance team.

## Who chooses a consumer fintech app, and what is a customer worth?

Individuals choose them, usually at a moment of change; one who makes the app primary brings years of revenue.

The buyer is a person, not a procurement team. They are often younger, often comparing two or three apps on a phone, and often prompted by an event: a new job, a credit card denial, a budget that stopped working, or an app that shut down.

The economics reward depth, not downloads:

| App | Latest public figures | What drives revenue |
|---|---|---|
| Chime | 10.4 million active members, average revenue per active member of $260, quarterly revenue of $670 million | Card spending, instant transfers and pay advances once members deposit their pay |
| Cash App | 59 million monthly transacting actives, 9.4 million primary banking actives, quarterly gross profit of $1.97 billion | Banking, borrowing and card use by people who move their paycheck in |

Chime’s pay-advance product, MyPay, originated $4.5 billion in the second quarter of 2026. Block, which owns Cash App, reported in its [results coverage](https://siliconangle.com/2026/08/05/block-shares-slip-despite-second-quarter-beat-raised-2026-guidance/) that Cash App’s monthly transacting actives grew only 3%, while its primary banking actives grew 17%. The lesson for growth teams, we infer: the valuable customer is the one who [trusts the app with their paycheck](https://underneath.agency/resources/neobanks-account-openings-ai-search), and trust is exactly what people ask AI about.

Demand is moving between categories too. [Sensor Tower](https://sensortower.com/press/press-release-boosted-by-gen-ai-services-consumers-spent-more-money-in-apps-than-games-for-first-time) reports that [credit and lending apps](https://underneath.agency/resources/lending-platforms-borrowers-ai-search) saw downloads climb 18% in 2025, while investing and crypto apps declined.

## Where does AI already sit in people’s money decisions?

At the start: people ask AI to explain, plan and compare before they download anything.

- **Advice and planning.** In Credit Karma’s survey of 1,019 adults, finance was the second most common use for generative AI (41%), behind health and wellness. Topics included financial goal setting (35%), budgeting (34%) and optimizing savings (33%). Credit Karma is an Intuit company with its own stake in AI-assisted finance, so read its figures as a vendor survey.
- **Credit and saving.** Experian’s respondents said AI tools had helped with saving and budgeting (60%) and credit score improvement (48%). Savers also ask where to keep their money, the subject of [how digital banks win depositors](https://underneath.agency/resources/digital-banks-depositors-ai-search).
- **Everyday tasks.** In [Menlo Ventures’ 2026 consumer survey](https://menlovc.com/wp-content/uploads/2026/09/menlo_ventures_consumer_ai_report-2026.pdf), 23% of people who pay bills use AI to help. Only 20% had given AI agents access to financial accounts, far fewer than for email.
- **Inside the assistant.** Intuit signed a [contract worth more than $100 million](https://techcrunch.com/2025/11/18/intuit-signs-100m-deal-with-openai-to-bring-its-apps-to-chatgpt) with OpenAI so that Credit Karma and TurboTax work inside ChatGPT, letting users review credit options there.

People do not trust AI blindly. In Credit Karma’s survey, 80% of those who acted on AI financial advice said they still research and validate it, and 52% said they had made a poor financial decision or mistake based on it. That second check is where an app’s reviews, disclosures and third-party coverage do their work.

## What do people ask AI before choosing a money app?

Questions about fit, fees, safety and alternatives. We wrote the examples below to illustrate common questions; they are not observed data.

| Need | Illustrative prompt |
|---|---|
| Budgeting | “Best budgeting app for a couple who want to share some accounts but not all” |
| Switching | “What are good Mint alternatives that still connect to my credit union?” |
| Credit building | “How can I build credit with no credit history, and which apps help?” |
| Payments | “Cash App vs Venmo: which charges less for instant transfers?” |
| Safety | “Is my money in this app FDIC insured, and what happens if the app goes under?” |
| Reputation | “Is this cash advance app legit? What do users complain about?” |

Switching moments matter most. When Intuit announced it was shutting down Mint, which [reportedly had 3.6 million monthly active users in 2021](https://betakit.com/as-intuit-winds-down-mint-financial-planning-app-monarch-money-comes-to-canada/), rivals such as Monarch moved quickly to win those users. Today, a large share of “what should I use instead” questions goes to assistants first, we infer, and the apps named in those answers start with an advantage.

## How does an AI answer turn into a funded account?

Through a short chain: answer, safety check, download, account link, then the step that matters most, regular deposits.

1. **The answer names the app.** A person asks for options, and the assistant lists two or three.
2. **The safety check.** The person asks a follow-up about fees, deposit insurance or complaints, or checks reviews.
3. **Download and link.** They install the app and link a bank account.
4. **Primary use.** They set up direct deposit or start paying through the app, where most revenue begins.

Each step can fail on an inaccurate answer. If an assistant misstates a fee, a deposit insurance arrangement or a credit-building feature, the person may drop out before downloading, and the app never learns why. We suggest tracking AI visibility against funded or active accounts, and asking new customers during sign-up where they first heard of the app.

## What makes an assistant trust a fintech app enough to name it?

In finance, mostly sources that already rank and trusted third parties; platforms document little about how they choose.

**Documented.** Google says its AI features may use [“query fan-out”](https://developers.google.com/search/docs/appearance/ai-features), running several related searches across subtopics before answering. Neither Google nor OpenAI publishes how it chooses between consumer finance apps.

**Observed in our studies.**

- **Finance answers lean on pages that already rank.** In [our citation study](https://underneath.agency/research/ai-overview-citations-study), 34.8% of finance and insurance AI Overview citations were page-one Google results, and nerdwallet.com appeared in 10.0% of all AI Overviews in the sample.
- **Comparison sites are a direct route.** In [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study), when ChatGPT’s own search named a source such as NerdWallet, the answer cited that source 44.0% of the time, against 8.1% when it did not. Investing apps meet the same pattern, covered in [how investing platforms get chosen by AI](https://underneath.agency/resources/investment-platforms-customers-ai-search).
- **Reputation questions go to review platforms.** In [our reputation study](https://underneath.agency/research/is-it-legit-ai-reputation-study), 88.0% of answers to “is this brand legit?” cited a review or complaint platform.
- **Communities matter less in finance.** In [our Reddit study](https://underneath.agency/research/ai-reddit-citations-study), only 5.7% of AI Overviews for financial services searches cited Reddit.

**Our inference.** For a consumer app, the deciding facts are regulatory and verifiable: which bank holds the money, whether deposit insurance applies and on what conditions, what each fee is, and how complaints are handled. The FDIC’s rule says fintech growth “has also blurred the distinction” between banks and non-banks “in the eyes of many consumers.” Chime’s own release states plainly: “Chime is not FDIC-insured,” and names its partner banks. An assistant that finds facts this clear can repeat them; one that finds vague or conflicting claims may hedge, or name a rival with clearer pages.

## What does GEO look like for a consumer fintech app?

Generative engine optimization (GEO) makes your app’s fit, fees and safety easy for assistants to find, check and repeat.

For a consumer fintech app, the work usually covers:

1. **Exact safety facts.** A public page stating the partner bank, how deposit insurance works for your accounts and its conditions, data security, and what happens to funds if something goes wrong, written with your compliance team.
2. **Plain fee and feature pages.** Every fee, limit and timing in text, not only in app screens, so comparisons in AI answers are correct.
3. **Comparison-site presence.** Accurate listings and reviews on the finance comparison and editorial sites that assistants search and cite.
4. **Reputation work.** Monitor Trustpilot, the BBB and app store reviews, answer complaints, and fix the causes; assistants summarize what those platforms say.
5. **Switching-moment content.** Clear, honest pages for people leaving another app or starting over with credit, without promising outcomes.
6. **Search foundations.** Because finance answers lean on page-one results, keep key pages ranking well in Google and Bing.

When an assistant gets a fact wrong, follow [how to fix wrong information about your brand in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers). Marketing claims in finance carry regulatory weight, so GEO content should go through the same review as any other advertising. Mortgage lenders face the same duty with rate claims, covered in [how mortgage lenders win borrowers through AI](https://underneath.agency/resources/mortgage-platforms-leads-ai-search). No one can guarantee an assistant will recommend an app; GEO makes the evidence it finds accurate and verifiable. Subscription economics for consumer apps in general are covered in [can AI assistants win new subscribers for a consumer app?](https://underneath.agency/resources/consumer-app-subscribers-from-ai-search).

## What does the evidence not tell fintech apps yet?

It shows people ask AI about money; it does not show how assistants pick apps or how many sign-ups follow.

- **Surveys come from interested parties.** The Credit Karma and Experian figures come from companies that sell financial products; independent, app-specific surveys are scarce.
- **No public link to sign-ups.** We found no public data connecting AI visibility to downloads or funded accounts for a consumer fintech app. What is known in general is summarized in [does AI visibility drive business results?](https://underneath.agency/resources/does-ai-visibility-drive-business-results).
- **Our studies are snapshots.** Our citation and frequency findings come from US searches in September 2026, and answers change between assistants and over time.
- **Assistant apps are new.** How often people use finance apps inside ChatGPT, rather than ordinary answers, is not yet public.

## Where should a consumer fintech app start?

Start by asking assistants the questions your future customers ask, and check every fee and safety fact in the answers.

A useful first check covers your core use cases (budgeting, payments, credit building), the switching questions in your category, and the safety and “is it legit” questions about your app. It shows whether you are named, which rivals and comparison sites appear instead, and whether your fees, partner bank and protections are described correctly.

If your growth depends on people choosing your app and trusting it with their paycheck, [ask us for a review of your app in AI answers](https://underneath.agency/contact). We will compare how assistants describe you and your competitors, list the facts they get wrong or cannot find, and plan the content, coverage and reputation work, reviewed with your compliance team, to close those gaps. You can see how that work is structured, from partner-bank and fee pages to comparison-site listings, on our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page.

## Frequently asked questions

### Do people really ask ChatGPT which budgeting app to use?

Many ask AI about money. In Credit Karma’s survey, 66% of generative AI users had sought financial advice, and 34% of those asked about budgeting. No public data counts app recommendations specifically.

### Why do AI answers about finance apps cite sites like NerdWallet?

Finance answers lean on established, well-ranked sources. NerdWallet appeared in 10.0% of the AI Overviews in our sample, and ChatGPT cited named sources far more often when its own search mentioned them.

### What is the biggest AI risk for a consumer fintech app?

A wrong safety or fee fact. The FDIC has warned that fintech arrangements can confuse consumers about deposit insurance, so state your partner bank and coverage conditions plainly.

### Can GEO tell people our app is the best choice for them?

No. GEO makes accurate information about your app easy to find. It does not give personal financial advice, and claims should pass the same compliance review as your advertising.

## Sources

- Intuit Credit Karma, via Stacker (2026-06-29), [The rise of fin-AI: Why Americans are trusting generative AI with their wallets](https://stacker.com/stories/personal-finance-investing/rise-fin-ai-why-americans-are-trusting-generative-ai-their)
- CUInsight (2025), [Gen Z, millennials are using AI for personal finance advice, report finds](https://cuinsight.com/gen-z-millennials-are-using-ai-for-personal-finance-advice-report-finds)
- Chime, via Business Wire (2026-08-05), [Chime Reports Second Quarter 2026 Financial Results](https://stocks.observer-reporter.com/observerreporter/article/bizwire-2026-8-5-chime-reports-second-quarter-2026-financial-results)
- SiliconANGLE (2026-08-05), [Block shares slip despite second-quarter beat and raised 2026 guidance](https://siliconangle.com/2026/08/05/block-shares-slip-despite-second-quarter-beat-raised-2026-guidance/)
- Sensor Tower (2026-01-21), [Boosted by gen AI services, consumers spent more money in apps than games for first time](https://sensortower.com/press/press-release-boosted-by-gen-ai-services-consumers-spent-more-money-in-apps-than-games-for-first-time)
- Menlo Ventures (2026-09-15), [2026: The State of Consumer AI](https://menlovc.com/wp-content/uploads/2026/09/menlo_ventures_consumer_ai_report-2026.pdf)
- TechCrunch (2025-11-18), [Intuit signs $100M+ deal with OpenAI to bring its apps to ChatGPT](https://techcrunch.com/2025/11/18/intuit-signs-100m-deal-with-openai-to-bring-its-apps-to-chatgpt)
- BetaKit (2023), [As Intuit winds down Mint, financial planning app Monarch Money comes to Canada](https://betakit.com/as-intuit-winds-down-mint-financial-planning-app-monarch-money-comes-to-canada/)
- Federal Register, FDIC (2024-01-18), [FDIC Official Signs and Advertising Requirements, False Advertising, Misrepresentation of Insured Status](https://www.federalregister.gov/documents/2024/01/18/2023-28629/fdic-official-signs-and-advertising-requirements-false-advertising-misrepresentation-of-insured)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- Underneath (2026), [When does Google show an AI Overview?](https://underneath.agency/research/ai-overviews-frequency-study)
- Underneath (2026), [Do AI Overviews cite the pages that rank?](https://underneath.agency/research/ai-overview-citations-study)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [“Is this brand legit?” How AI assistants build a reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study)
- Underneath (2026), [When does a Reddit thread become evidence in Google’s AI?](https://underneath.agency/research/ai-reddit-citations-study)

---

This is the Markdown twin of https://underneath.agency/resources/consumer-fintech-apps-customers-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Will AI shortlist our devices when buyers choose a brand?"
description: "By being named, with correct compatibility and launch facts, when households ask AI which device or ecosystem to buy into, a choice that lasts years."
canonical: "https://underneath.agency/resources/consumer-tech-brands-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Will AI assistants put our devices on the shortlist when buyers choose a brand and ecosystem?

Only if your brand is described clearly and consistently across the independent sources AI assistants draw on, including what your devices work with and what your latest models change. For a consumer technology brand, the moment that matters is the rare one when a household picks an ecosystem or considers switching, because loyalty after that choice is very high. AI assistants are increasingly present at that moment, and they tend to favor brands that are already widely written about.

## The short version

1. The choice is sticky: in a 2026 SellCell survey reported by [MacRumors](https://www.macrumors.com/2026/04/16/iphone-loyalty-survey/), 96.4% of iPhone users and 86.4% of Android users planned to stay with their platform, and 17.4% of loyal iPhone users said they were invested in the Apple ecosystem.
2. Households keep spending inside that choice: [Deloitte](https://www.deloitte.com/us/en/about/press-room/connectivity-mobile-trends-survey.html) found surveyed households spending about $896 on connected devices in 2025, up from $764 in 2024, and consumers who see their providers as both innovative and responsible spend 62% more a year on devices.
3. AI already advises on this: product recommendations are among the top five uses regular gen AI users report for chatbots ([Deloitte Connected Consumer](https://www.deloitte.com/us/en/insights/industry/telecommunications/connectivity-mobile-trends-survey.html)), and 53% of surveyed consumers now use or experiment with gen AI.
4. Big brands start ahead: in one vendor’s tracking data, global household-name brands appeared in 73% of unbranded AI answers on day one, against 11% for niche brands ([Kumar](https://arxiv.org/abs/2606.20065)).
5. New launches are easy to miss: in a test of 112 recently launched Product Hunt products, mostly software, ChatGPT without web search surfaced them in only 3.32% of discovery answers and Perplexity in 8.29% ([Sharma](https://arxiv.org/abs/2601.00912)).

## Who chooses a consumer tech brand, and what is that customer worth?

Households choose, usually once per ecosystem, and the customer is worth years of devices and services afterward.

Consumer technology brands sell more than a device. A phone, a watch, a smart speaker or a video doorbell pulls the buyer into an ecosystem of apps, accessories, subscriptions and future devices that work best together. The [Consumer Technology Association](https://www.ces.tech/press-releases/cta-despite-tariffs-and-economic-headwinds-us-consumer-tech-revenue-to-hit-565-billion-in-2026) projects US consumer spending on software and services to rise 4.2% in 2026 to nearly $194 billion, and describes consumers “anchoring to subscription services.” The device is the entry point to that revenue.

Loyalty data shows how much rests on the first choice. In the SellCell survey of 5,000 US smartphone users, 96.4% of iPhone users said they would stick with an iPhone, and of the few planning to leave, 69.7% said they would choose Samsung. [CIRP](https://cirpapple.substack.com/p/did-iphone-loyalty-reach-its-limit), which tracks actual purchases, reported that the share of iPhone owners buying another iPhone peaked at 94% and was 89% in the twelve months to May 2025, while Samsung’s loyalty was growing from a lower base. For every brand outside the leader’s ecosystem, the opportunities are the first purchase, the switch and the adjacent device, such as a smart lock or earbuds bought to work with what the household already owns.

The smart home is now mainstream and cautious. Parks Associates research reported by the [Fiber Broadband Association](https://fiberbroadband.org/2025/01/29/the-smart-home-in-2025-outlook-and-opportunities) found that about 45% of US internet households own at least one smart home device and the average household has about 17 connected devices. Self-described innovators fell from about 60% of smart home device owners in 2018 to 11%; most buyers now wait until devices are “several generations later and proven.” That buyer researches, and increasingly asks an assistant.

## Where do AI assistants enter brand consideration for devices?

In general chatbots, inside shopping features that remember the buyer, and in the voice assistants built into devices.

- **General chatbots.** Deloitte’s 2025 survey of about 3,500 US consumers found product recommendations among the top five chatbot uses of regular gen AI users, and most chatbot users said the help they received was as good as a human’s.
- **Shopping features with memory.** [OpenAI documents](https://openai.com/index/chatgpt-shopping-research/) that its shopping research feature builds on “ChatGPT’s understanding of you from past conversations and your ChatGPT memory,” giving the example that if ChatGPT knows you are into gaming, “it can factor that in when helping you find a new laptop.” A reasonable expectation is that the devices a buyer already owns will shape what an assistant suggests next.
- **Assistants inside devices.** [Google documents](https://blog.google/products/google-nest/gemini-for-home/) that Gemini for Home will replace Google Assistant on existing speakers and displays and offers “expert advice” through Gemini Live, with “I’m thinking about buying a new car…” as an example of the in-depth questions it handles. Several of the companies whose assistants answer these questions also sell devices; we have no evidence that this affects which brands are recommended, but brand leaders should know who runs the assistant.

The source mix changes with the stage of the question. [Chen and colleagues](https://arxiv.org/abs/2509.08919) observed that for comparison questions such as “Garmin vs Apple Watch,” independent earned sources dominated across systems, while purchase-ready questions such as “Best price for Samsung Galaxy S24 unlocked” raised the share of brand sources. Your own site matters most late in the journey; independent coverage matters most when the shortlist forms.

## Which questions decide whether a device brand is considered?

Ecosystem fit, switching, alternatives, compatibility and launch questions decide it, more than raw specs.

We wrote these prompts to illustrate the questions households ask when choosing a brand; they are not observed data:

- Ecosystem: “Which smart lock works best with Apple Home and Matter?” or “Google Home or Alexa for a family of four in a rental?”
- Switching: “Should I switch from iPhone to a Galaxy if I use a Mac at work?”
- Alternatives: “Video doorbells like Ring that don’t need a subscription.”
- Compatibility: “Will these earbuds support spatial audio on my Android phone?”
- Launch: “Is the new watch worth upgrading from the one two generations back?”
- Use case: “Best smart home starter kit for someone who isn’t technical.”

These questions are about brands and ecosystems more than single models, which is the difference from [a spec-by-spec electronics comparison](https://underneath.agency/resources/consumer-electronics-sales-from-ai-search). Our article on [what drives AI product recommendations](https://underneath.agency/resources/what-drives-ai-product-recommendations) covers how product data drives model-level picks.

## How does AI visibility turn into device sales and ecosystem revenue?

By placing your brand on the shortlist at a first purchase or a switch, which then compounds through repeat purchases.

The path we infer from the loyalty data runs like this: a household reaches a decision point, such as a new home, a broken device or a launch; it asks an assistant which brand or ecosystem fits; the assistant names a few and explains why; the household checks reviews and compatibility, then buys. If the answer named your brand with the right compatibility facts, you win a device and often the next several. If it did not, the household may stay outside your ecosystem for years, given loyalty rates near 90% or more.

Trust affects the size of the prize. Deloitte found that among companies seen as innovative, those also seen as excelling at data responsibility enjoy 25% higher annual device spending from their customers. Privacy, security and update policies are part of how assistants describe a brand, so they are part of brand consideration in AI answers, not only on the packaging. That last point is our inference.

## What decides whether an assistant names your brand?

Platforms document little; studies point to independent coverage, clear positioning and a consistent brand identity.

**Documented by the platforms.** OpenAI says shopping research is trained to “read trusted sites, cite reliable sources,” and avoids “low-quality or spammy sites.” Google says Gemini for Home combines Gemini Live with Google Search for in-the-moment help. Neither explains how brands are chosen.

**Observed in studies.**

- *Independent coverage.* In [our brand entity study](https://underneath.agency/research/brand-entity-ai-recommendations-study), each tenfold increase in the number of independent sites naming a brand went with 4.7 times the odds of being recommended. Our article on [whether AI assistants favor big brands](https://underneath.agency/resources/do-ai-assistants-favor-big-brands) covers the same pattern.
- *Positioning in buyers’ words.* In tests by [Malthouse and colleagues](https://arxiv.org/abs/2609.16304), a well-known tool brand that never appeared for a plain category request appeared in 35.4% of lists when shoppers described their needs, and in 81.3% when prompts used its own positioning language. Brands whose positioning matches how buyers describe their use case are easier to retrieve.
- *A split identity.* In our brand entity study, 10.2% of the options assistants recommended were named differently by different assistants, and 447 were recorded with a parent brand. Consumer tech brands with parent brands, sub-brands and model families (think Galaxy, Pixel, Echo or Nest) are exposed to this.
- *Stable presence.* [Uberti-Bona Marin and colleagues](https://arxiv.org/abs/2609.18729) found that whether any brand or product appeared in an answer changed in only 11% to 13% of repeated queries, even when the specific products named varied. Being in or out tends to persist.

**Trust factors specific to consumer tech.** We infer from these sources that assistants weigh independent reviews, clear compatibility and standards support (such as Matter, which the [Connectivity Standards Alliance](https://csa-iot.org/all-solutions/matter/) says is “rolling out to millions of smart home devices”), software update and support records, privacy reputation, and subscription requirements, all of which buyers ask about directly.

## Why do launches and compatibility facts go wrong in AI answers?

Because assistants often answer from older information, and compatibility facts change with every update.

Launch timing is the main risk. In [Sharma’s test](https://arxiv.org/abs/2601.00912) of products newly launched on Product Hunt, mostly software rather than devices, ChatGPT without web search found them in 3.32% of discovery answers; Perplexity, which searches, reached 8.29%. Assistants that do search lean toward recent pages: our [freshness study](https://underneath.agency/research/ai-source-freshness-study) found that pages from the last 90 days were 17.4% to 22.6% of each assistant’s dated citations, against 6.9% of Google’s top 10. A launch with fresh independent coverage has a chance; one with only a press release may not. Our article on [why ChatGPT misses new products](https://underneath.agency/resources/why-chatgpt-misses-new-products) explains this in more depth.

Compatibility is the second risk. Support for a platform, a standard or a feature can arrive in a firmware update months after launch, and older reviews will still say it is missing. Deloitte found that one-third of gen AI users had encountered incorrect or misleading information, and 53% mostly or always verify outputs. A buyer who checks and finds your compatibility misdescribed may drop you from the shortlist. Our guide to [fixing wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) covers how to correct it.

## How does GEO work for a consumer technology brand?

By making your brand, ecosystem and current models easy for AI systems to identify, verify and explain, without promising placement.

Generative engine optimization, GEO, for a consumer tech brand covers:

- **Brand and entity clarity.** One consistent way of naming the parent brand, sub-brands and model generations across your site, retailers, app stores and press, so assistants do not split or merge your products.
- **Compatibility stated plainly.** Public, current pages that list what each device works with (platforms, standards, phones, voice assistants), what requires a subscription, and what changed in each update.
- **Launch coverage.** Review units, briefings and comparison pages against the previous generation, timed so independent coverage exists when buyers start asking. Fresh pages are cited more by assistants that search.
- **Use-case positioning.** Pages and coverage that describe who the device is for in buyers’ own words, so needs-based questions retrieve your brand.
- **Independent authority.** Digital PR aimed at the reviewers, publications and creators assistants cite, as set out in [how brands build authority for AI search](https://underneath.agency/resources/how-brands-build-authority-for-ai-search).
- **Reputation and trust signals.** Clear privacy, security and update-support commitments, and customer reviews on major retailers and app stores.
- **Measurement by stage.** Separate tracking for informational, comparison and purchase-ready questions across ChatGPT, Gemini, Google AI Mode, Perplexity, Copilot and assistants built into devices.

## What is still unknown about AI and device ecosystem choices?

How much AI answers change ecosystem choices, and whether device-making platforms favor their own products.

- No public study measures how often an AI answer causes a household to switch ecosystems; the loyalty figures come from surveys and purchase tracking.
- We found no evidence on whether assistants run by companies that also sell devices favor their own brands, in either direction.
- Most brand studies test single-product categories, not ecosystems; how assistants reason about multi-device compatibility is not measured.
- Memory and personalization are new; how much a buyer’s past conversations change brand suggestions has not been measured publicly.
- Study results date quickly as models and features change.

## Where should a consumer technology brand start?

Start by testing the ecosystem, switching, compatibility and launch questions that decide brand consideration in your categories.

List the 30 to 50 questions that open or close your ecosystem to a new household, ask them several times in each major assistant and in the voice assistants on devices, and record whether your brand appears, how your compatibility and latest models are described, and which sources are cited. Fix factual errors first, then close the coverage gaps. If you would like a partner for it, [start a conversation with us](https://underneath.agency/contact): we map how AI assistants describe your brand and ecosystem at the decision points that set years of device and service revenue, and build a plan to make your brand easier to find, verify and choose. That plan, covering consistent sub-brand naming, plain compatibility pages and launch coverage, is explained step by step on our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page.

## Frequently asked questions

### Do AI assistants simply recommend the biggest device brands?

They start with a large advantage: in one vendor’s data, global brands appeared in 73% of unbranded answers on day one against 11% for niche brands. But independent coverage and clear positioning also matter, so smaller brands are not locked out.

### Should we publish comparison pages against rival ecosystems?

Factual, current comparison and compatibility pages help assistants answer purchase-ready questions, where brand sources carry more weight. For comparison-stage questions, independent coverage matters more.

### How long after a launch do new devices show up in AI answers?

There is no fixed delay. Assistants that search the web cite recent pages more often, but new products are often missed until independent reviews exist.

### Can a firmware update fix how AI describes our compatibility?

Only once the sources assistants read are updated. Publish the change clearly and ask reviewers to update their pages.

## Sources

- MacRumors (2026-04-16), [iPhone Loyalty Hits 96.4% as Android Users Four Times More Likely to Switch](https://www.macrumors.com/2026/04/16/iphone-loyalty-survey/)
- CIRP (2025-05-21), [Did iPhone Loyalty Reach its Limit While Samsung Loyalty Grows?](https://cirpapple.substack.com/p/did-iphone-loyalty-reach-its-limit)
- Deloitte (2025), [Connected Consumer survey press release](https://www.deloitte.com/us/en/about/press-room/connectivity-mobile-trends-survey.html)
- Deloitte (2025), [2025 Connected Consumer survey](https://www.deloitte.com/us/en/insights/industry/telecommunications/connectivity-mobile-trends-survey.html)
- Consumer Technology Association (2026-01-04), [U.S. Consumer Tech Revenue to Hit $565 Billion in 2026](https://www.ces.tech/press-releases/cta-despite-tariffs-and-economic-headwinds-us-consumer-tech-revenue-to-hit-565-billion-in-2026)
- Fiber Broadband Association (2025-01-29), [The smart home in 2025: outlook and opportunities](https://fiberbroadband.org/2025/01/29/the-smart-home-in-2025-outlook-and-opportunities)
- OpenAI (2025-11-24), [Introducing shopping research in ChatGPT](https://openai.com/index/chatgpt-shopping-research/)
- Google (2025-08-20), [Gemini for Home: Your household’s new, more helpful assistant](https://blog.google/products/google-nest/gemini-for-home/)
- Connectivity Standards Alliance (n.d.), [Matter](https://csa-iot.org/all-solutions/matter/)
- Kumar (2026), [Generative Engine Optimization at Scale: Measuring Brand Visibility Across AI Search Engines](https://arxiv.org/abs/2606.20065)
- Malthouse and colleagues (2026), [Evaluating Brand Retrieval and Ranking in Large Language Model Recommendations](https://arxiv.org/abs/2609.16304)
- Chen and colleagues (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919)
- Uberti-Bona Marin and colleagues (2026), ["If I Had to Buy Just ONE: Galaxy S26 Ultra": Auditing AI-Generated Product Recommendations](https://arxiv.org/abs/2609.18729)
- Sharma (2025), [The Discovery Gap: How Product Hunt Startups Vanish in LLM Organic Discovery Queries](https://arxiv.org/abs/2601.00912)
- Underneath (2026), [Brand entity and AI recommendations study](https://underneath.agency/research/brand-entity-ai-recommendations-study)
- Underneath (2026), [AI source freshness study](https://underneath.agency/research/ai-source-freshness-study)

---

This is the Markdown twin of https://underneath.agency/resources/consumer-tech-brands-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Is my content business at risk from Google’s AI summaries?"
description: "It depends on your content. Factual pages lose visits to Google’s AI summaries; opinion and experience content gained, until AI Mode eroded the gain."
canonical: "https://underneath.agency/resources/content-business-risk-from-ai-search"
published: 2026-10-08
updated: 2026-10-10
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Is my content business at risk from Google’s AI summaries?

It depends on what your content is. Research on Wikipedia and Reddit suggests that factual content, which a summary can restate, loses visits, while opinion, advice and experience content can gain. Google’s conversational AI Mode erodes even those gains, so no content business should assume it is safe.

## The short version

1. English Wikipedia, a factual reference site, lost about 5% of its search traffic after AI Overviews became Google’s US default in 2024 (Khosravi and Yoganarasimhan).
2. Reddit communities that AI Overviews can cite gained 12.0% more daily comments, and experience-based communities gained 2.3 to 2.8 times more than fact-based ones (Zhang, Cui and Zhang).
3. After AI Mode rolled out globally in August 2025, that extra gain for experience-based communities fell 59% in commenters and disappeared in comments.
4. 50.63% of the pages AI Overviews cite carry display ads, so many of the sites supplying the answers depend on the clicks the answers absorb (Xu and colleagues, 2026).
5. Forcing every search into AI Mode cut the share of people clicking through to news sites by 12.5 points and to Reddit by 21.2 points, in a 2026 US experiment.

## Which kinds of content are most at risk?

Factual content that a summary can restate in full is most at risk of losing visits.

[Zhang, Cui and Zhang](https://arxiv.org/abs/2605.16428) split online content into two kinds. Fact-based content, such as reference entries or a quick technical answer, can be judged from a summary alone. Experience-based content, such as opinions, advice and personal stories, has to be read in full to be worth anything.

For the first kind, an AI summary does the job and replaces the visit. For the second, the summary works more like a trailer that sends people to the full discussion. In the authors’ words, factual-content platforms “face substitution risk, as AI can directly satisfy the information need.”

| Content type | Examples | What an AI summary tends to do |
|---|---|---|
| Fact-based | Definitions, reference facts, quick how-to answers | Answers the question, so fewer people click |
| Experience-based | Reviews, advice threads, personal experience, opinion | Points people to the full content, so more may visit |

Most content mixes both. Of the 105,012 Reddit communities the study classified, an AI model labeled 56.6% fact-based and 43.4% experience-based.

## What happened to a factual reference site?

Wikipedia lost about 5% of its English search traffic once AI Overviews became the US default.

[Khosravi and Yoganarasimhan](https://arxiv.org/abs/2602.18455) compared English Wikipedia articles with the same articles in German and French through 2024. English search traffic fell by about 5.45% relative to German and 4.82% relative to French. That is roughly 100.27 million fewer search visits a month. Wider evidence on [whether AI Overviews reduce clicks](https://underneath.agency/resources/do-ai-overviews-reduce-clicks) points the same way.

That figure covers only the first page a visitor lands on. Article-to-article links make up about 40% of Wikipedia’s traffic, so a lost search visit may also cost the pages a reader would have clicked next. Wikipedia runs no ads. For an ad-supported publisher losing the same traffic, the authors estimate roughly $0.901M to $3.090M of ad revenue a month.

Press reports collected by [Xu, Iqbal and Montgomery](https://arxiv.org/abs/2605.14021) describe harder cases, such as a travel blog that shut down after its traffic fell 90%. These are single stories, not measured studies, but they show the tail of the risk.

## What happened to discussion and experience content?

It gained: Reddit communities that AI Overviews could cite saw more people join their discussions.

Google keeps adult communities out of AI Overviews but still lists them in normal results. That gave Zhang and colleagues a clean comparison across 105,012 Reddit communities from January 2024 to July 2025. After AI Overviews expanded internationally in August 2024, citable communities gained 12.0% more daily comments and 12.4% more daily commenters.

The gain came mostly from new people taking part, not existing members posting more. It was 2.3 times larger for comments, and 2.8 times larger for commenters, in experience-based communities than in fact-based ones. Smaller communities gained most in relative terms: in the second-smallest size group, comments rose 56% and commenters 77%.

Reddit is a special case. Khosravi and Yoganarasimhan point out that Reddit is cited unusually often in AI Overviews following its data-sharing agreement with Google. [Our study of Reddit citations in Google’s AI](https://underneath.agency/research/ai-reddit-citations-study) looks at which threads get cited. A business blog with the same kind of content may not see the same lift.

## Does AI Mode change the picture?

Yes: Google’s conversational AI Mode erased much of the gain experience content had won.

AI Mode is Google’s chat-style search, where people can ask follow-up questions instead of clicking through. Zhang and colleagues extended their data through December 2025, after AI Mode rolled out globally in August 2025. The extra gain for experience-based communities fell 59% in commenters, and for comments it disappeared entirely. A summary that can answer follow-up questions starts to replace the discussion itself.

A field experiment by [Wang and colleagues](https://arxiv.org/abs/2608.18352) points the same way. When 1,100 US volunteers had every Google search sent to AI Mode for a week in March 2026, the share clicking through to news sites fell 12.5 points, to Reddit 21.2 points and to Wikipedia 9.9 points. Our article on [AI Mode as Google’s default](https://underneath.agency/resources/ai-mode-default-traffic-loss) covers that experiment in detail.

## Are big publishers or small sites more exposed?

The research disagrees, so treat any confident answer with caution.

[Grossman and colleagues](https://arxiv.org/abs/2604.27790) found that AI Overviews lean less on the most popular websites than Google’s normal results. Counting only the first source for each search, sites in the 1,000 most popular domains held the top spot for 52.7% of searches in normal results and 40.0% in AI Overviews. The authors conclude that generative search may help niche sites at the expense of large ones.

[Aral, Li and Zuo](https://arxiv.org/abs/2602.13415), using a different set of searches across many countries, found the opposite. In their data, AI search pointed to the most-visited websites more often, and to the long tail of smaller sites less often, than normal search.

What both agree on is exposure. Xu and colleagues found that 50.63% of the pages AI Overviews cite run display ads. A large share of the content that feeds these summaries belongs to businesses that earn money from the visits the summaries can absorb.

## What should you do about it?

Sort your content by whether a summary can replace it, and move investment toward what it cannot.

1. Audit your traffic pages. Mark which ones mainly state facts or answer a quick question, and which offer experience, judgment, tools or discussion.
2. Treat the factual pages as the most exposed. On Wikipedia’s evidence, [being cited did not stop the loss](https://underneath.agency/resources/ai-overview-citations-vs-lost-clicks).
3. Build content that has to be consumed in full: firsthand reviews, original data, community discussion and tools people use rather than read.
4. Watch AI Mode closely. It removed much of the gain experience content had won from AI Overviews.
5. Reduce dependence on ad-supported page views where you can, since the answer now often sits above your page.

If you want help assessing which of your pages are exposed, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The research covers two very large platforms, not a typical content business.

- Wikipedia is a nonprofit reference site; Reddit is a forum with a data deal with Google. Neither looks like a niche publisher or a company blog.
- The Reddit study measures comments and commenters, not page views or revenue.
- Content types were labeled by an AI model, which adds some error.
- Video, audio and other formats are untested; the Reddit authors flag this themselves.
- Studies disagree on whether large or small sites lose more.
- The data after AI Mode’s launch covers only a few months, so the long-run effect is unknown.

## Frequently asked questions

### Will Google’s AI Overviews kill my blog traffic?

Not necessarily, but factual pages are at risk. English Wikipedia lost about 5% of its search traffic after AI Overviews launched, while opinion and advice communities on Reddit gained engagement.

### Does AI search help forums and online communities?

It did at first. Reddit communities that AI Overviews could cite gained 12.0% more daily comments, but much of that gain faded after AI Mode rolled out.

### Which topics see the most AI Overviews?

In one 2026 study of general-knowledge questions, business (84%), history (81%) and education (74%) questions produced AI Overviews most often, while sports, technology and travel (32% to 34%) did least, according to [Huang and colleagues](https://arxiv.org/abs/2603.16138).

### Are news publishers losing clicks to Google’s AI?

The evidence says yes for AI Mode. In a 2026 US experiment, forcing every search into AI Mode cut the share of people clicking through to news sites by 12.5 percentage points.

## Sources

- Zhang, Cui and Zhang (2026), [The Impact of AI Search on the Online Content Ecosystem: Evidence from Google and Reddit](https://arxiv.org/abs/2605.16428), arXiv:2605.16428.
- Khosravi and Yoganarasimhan (2026), [Impact of AI Search Summaries on Website Traffic: Evidence from Google AI Overviews and Wikipedia](https://arxiv.org/abs/2602.18455), arXiv:2602.18455.
- Xu, Iqbal and Montgomery (2026), [Measuring Google AI Overviews: Activation, Source Quality, Claim Fidelity, and Publisher Impact](https://arxiv.org/abs/2605.14021), arXiv:2605.14021.
- Wang, Gleason, Bart, Wilson and Metaxa (2026), [AI in Search Reduces Publisher Referrals Without Improving User Experience: Experimental Evidence](https://arxiv.org/abs/2608.18352), arXiv:2608.18352.
- Grossman, Liu, Chen, Smith, Borcea and Chen (2026), [How Generative AI Disrupts Search: An Empirical Study of Google Search, Gemini, and AI Overviews](https://arxiv.org/abs/2604.27790), arXiv:2604.27790.
- Aral, Li and Zuo (2026), [The Rise of AI Search: Implications for Information Markets and Human Judgement at Scale](https://arxiv.org/abs/2602.13415), arXiv:2602.13415.
- Huang and colleagues (2026), [Answer Bubbles: Information Exposure in AI-Mediated Search](https://arxiv.org/abs/2603.16138), arXiv:2603.16138.

---

This is the Markdown twin of https://underneath.agency/resources/content-business-risk-from-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How contract manufacturers win RFQs through AI search"
description: "It can, when your capabilities, certifications and location are written where AI can verify them, because engineers now use AI to build supplier longlists."
canonical: "https://underneath.agency/resources/contract-manufacturers-rfqs-from-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Will AI search put our plant on the supplier shortlist before the RFQ goes out?

It can, if an assistant can find and verify what you make, to what standard, in what volumes and where. Most technical buyers now use generative AI somewhere in a purchase, and most of their evaluation happens before they contact a supplier, so the shortlist is often drawn up before your sales team knows there is a project.

This article is for contract and custom manufacturers: machine shops, fabricators, molders and private-label producers that win work through requests for quotation (RFQs). The prize here is not website traffic. It is a place on a shortlist that can lead to years of production orders.

## The short version

1. Demand for outside production is strong: on Thomasnet, [contract and private-label manufacturing was the top category buyers sourced in 2025](https://distributionstrategy.com/2026/01/industrial-sourcing-behavior-shifts-in-2025-signaling-strategic-imperatives-for-distributors/), across more than 1.5 million monthly sourcing sessions.
2. Engineers use AI but check it: in the [2026 State of Marketing to Engineers](https://advertising.globalspec.com/state-of-marketing-to-engineers-research-report) research by TREW Marketing and GlobalSpec, 69% of technical buyers used generative AI in purchasing, they rated its answers 4.7 out of 10 for trust, and 62% of the buying journey happened online before contacting a vendor.
3. Supplier switching is creating openings: 32% of contract manufacturers were quoting reshoring projects in the [2026 Reshoring Initiative survey](https://www.designnews.com/industry/us-reshoring-accelerates-despite-policy-uncertainties-and-skilled-worker-shortages), double the 16% in 2025, and in [Kearney’s survey](https://www.kearney.com/documents/d/asset-library-291362522/2026-reshoring-index-pdf) three in four respondents had moved sourcing from China to other low-cost countries.
4. Value sits in a few large accounts: [Xometry](https://www.manilatimes.net/2026/02/24/tmt-newswire/globenewswire/xometry-reports-record-fourth-quarter-and-strong-full-year-2025-results/2283753) had 81,821 active buyers at the end of 2025, and 1,760 accounts spending at least $50,000 a year.
5. Brand recognition decides close calls: 70% of technical buyers said they were likely to choose the better-known brand when two solutions are technically similar.

## Who buys from contract manufacturers, and what is a customer worth?

Engineers and sourcing managers at manufacturers, and one won account can mean years of repeat production orders.

The buyer is usually a design or manufacturing engineer who needs parts made, working with a sourcing or procurement manager who owns the supplier relationship. Thomas’ 2025 sourcing report found two-thirds of active buyers on its platform represent small and midsize businesses, with a significant share under age 45.

The supplier side is fragmented. According to the [National Association of Manufacturers](https://nam.org/mfgdata/facts-about-manufacturing-expanded/), there were 239,265 manufacturing firms in the US in 2022, and around three-quarters had fewer than 20 employees. Thomas says it connects buyers with more than 500,000 suppliers. Most of these firms have no marketing department, which is exactly why discovery matters: a buyer cannot shortlist a shop it has never heard of.

The money moving through this market is large. The [US Census Bureau](https://www.census.gov/manufacturing/m3/adv/pdf/durgd.pdf) reported new orders for manufactured durable goods of $338.6 billion in August 2026 alone.

What one customer is worth varies by process and program, and there is no public benchmark for an average contract value. Public data shows the shape instead. At Xometry, the marketplace for custom manufacturing that also owns Thomas, active buyers grew 20% to 81,821 in 2025, while accounts spending at least $50,000 over twelve months grew 18% to 1,760. Our inference: a small share of accounts carries much of the value, so one well-qualified RFQ can matter more than thousands of website visits.

## Where does AI already sit in supplier discovery?

In early research and longlisting, used by a majority of technical buyers but trusted only after checking.

The best industry-specific evidence is the 2026 State of Marketing to Engineers survey from TREW Marketing and GlobalSpec, with Elektor:

| Technical buyers, 2026 (TREW Marketing and GlobalSpec) | Share |
|---|---|
| Use generative AI during the purchasing process | 69% |
| Never use it to evaluate vendors or research options (42% in 2025) | 31% |
| Use it often (6% in 2025) | 11% |
| Notice AI summaries at the top of search results | 75% |
| Of those, read the summary and then continue to the results | 69% |
| Of those, find the summary usually enough | 6% |
| Routinely research in online technical publications | 76% |
| Routinely research on supplier and vendor websites | 74% |

Two findings matter most for a supplier. Average trust in generative AI answers was 4.7 out of 10, so engineers use AI to start a search and then verify on supplier websites and in trade publications. And the online share of the journey keeps rising: buyers 35 and under complete 66% of the process online before their first vendor conversation. The same habit shapes how [engineers choose components with AI](https://underneath.agency/resources/industrial-manufacturers-ai-search).

Industrial platforms are moving the same way. In January 2026, Thomas launched an AI search that lets buyers run detailed, multi-attribute searches in natural language. [Digital Commerce 360 reported](https://www.digitalcommerce360.com/2026/01/19/thomas-ai-search-performance-based-ads-industrial-sourcing/) that in testing it drove more than 15% more supplier evaluations than the old search.

Procurement teams are earlier in adoption. In [EFESO’s survey of European procurement](https://www.efeso.com/wp-content/uploads/2026/04/EFESO-The-State-of-Generative-AI-in-European-Procurement.pdf), 79% of respondents had been exposed to AI, but only 31% actively used generative AI tools for work. Across business buying generally, [Forrester](https://www.forrester.com/press-newsroom/forrester-2026-the-state-of-business-buying/) describes generative AI searches as the starting point for B2B buyers, followed by validation from trusted people and sources.

## Which supplier questions do buyers ask AI?

Questions that combine a process, a material, a standard, a volume and a place, the same filters an RFQ uses.

The prompts below are illustrative, written by us to show the shape of supplier questions. They are not captured from real buyers.

| Need | Illustrative prompt |
|---|---|
| Capability and certification | “AS9100 certified machine shops for 5-axis titanium parts, low volume” |
| Regulated production | “ISO 13485 injection molders for medical device housings in the Midwest” |
| Reshoring | “US alternative to an overseas supplier for stamped steel brackets, about 50,000 a year” |
| Private label | “Contract manufacturers for private-label skin care with low minimum orders” |
| Speed | “Sheet metal fabricators that can turn prototypes in under two weeks” |
| Comparison | “Compare these three molders on capacity, certifications and industries served” |

Each prompt is a filter. If your site, directory profiles and trade coverage never say that you hold a given certification, run a given process or serve a given region, an assistant has nothing to match. The Thomas data shows which filters are active: buyers in aerospace and defense sought [precision CNC machining](https://underneath.agency/resources/cnc-machine-shops-rfqs-ai-search) and fabrication, food and beverage buyers sought contract manufacturing for supplements and vitamins, and the sharpest growth came in military equipment, automation, robots and prototyping services.

## How does an AI answer become an RFQ and a contract?

Through a longlist, a website check, a quote competition and qualification, then repeat production orders.

1. **Longlisted.** An engineer asks an assistant or an industrial platform for suppliers that match the job. The AI answer is a starting list, not a decision.
2. **Checked.** The engineer visits supplier websites and trade publications to confirm capabilities, certifications, equipment and industries served. TREW and GlobalSpec found buyers reach out when they need what they cannot find online: pricing and availability.
3. **Asked to quote.** The supplier receives an RFQ, usually alongside rivals. In the Reshoring Initiative’s survey, [reported by Floor Covering News](https://www.fcnews.net/2026/09/reshoring-survey-finds-u-s-manufacturing-momentum-building/), contract manufacturers competed against imports on an average of 38% of quotes, and when they lost to imports, price was the main factor 94% of the time.
4. **Qualified and awarded.** Samples, first-article inspection and audits follow before production starts.
5. **Repeated.** A qualified supplier tends to keep receiving orders for the life of the program. That is where the value of the first mention is realized.

AI visibility affects the first two steps. Price, quality and delivery still decide the last three. A shop that is never longlisted never gets the chance to compete on them. Equipment makers chasing [a place on engineers’ bid lists](https://underneath.agency/resources/industrial-equipment-leads-ai-search) face the same first cut.

## What decides whether an assistant names your plant?

Facts it can find and confirm in outside sources; the platforms describe how they search, not how they choose.

What is documented by the platforms: [OpenAI says](https://help.openai.com/en/articles/9237897-chatgpt-search) ChatGPT search rewrites a question into one or more targeted queries sent to search providers, and that a site must allow OpenAI’s search crawler, OAI-SearchBot, to be eligible for inclusion. [Google says](https://blog.google/products/search/ai-mode-search/) AI Mode uses a “query fan-out” technique, running multiple related searches across subtopics. Neither publishes how suppliers are chosen.

What has been observed:

- **Assistants run several searches per question.** In [our study of hidden searches](https://underneath.agency/research/ai-hidden-searches-study), ChatGPT ran a mean of 3.7 searches per buyer question before answering.
- **Independent coverage matters most of what we measured.** In [our brand entity study](https://underneath.agency/research/brand-entity-ai-recommendations-study), each tenfold increase in independent sites naming a brand in the cited pages went with 4.7 times the odds of being recommended.
- **Answers vary.** In [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), only 25.2% of the brands ChatGPT named for a question appeared in all five repeats.
- **Familiarity tips decisions.** In the engineers survey, 53% said brand familiarity influenced their most recent purchase.

Our inference for contract manufacturing: the trust factors are the ones a sourcing manager already checks, stated where machines can read them. That means certifications with the registrar named, process and equipment lists, materials, tolerances, volume ranges, industries served, plant locations, export-control registrations where they apply, and coverage in trade publications and industry directories. Chemical producers face a similar checklist, set out in [how chemical suppliers reach B2B buyers](https://underneath.agency/resources/chemical-suppliers-b2b-buyers-ai-search).

## What does a contract manufacturer lose when AI leaves it out?

It loses RFQs it never sees, at a time when many buyers are actively looking for new suppliers.

We have no direct measurement of RFQs lost to AI absence, so we label the reasoning:

- **Buyers are moving work.** In the Reshoring Initiative survey, 36% of original equipment manufacturers (OEMs) had reshored or were actively reshoring in 2026, and 79% of contract manufacturers said at least some customers had discussed reshoring with them. Kearney found most companies moved sourcing to other low-cost countries rather than home: only 20% of its survey respondents had looked at domestic manufacturing. Either way, supplier searches are happening.
- **The search happens before contact.** With 62% of the journey online, a supplier missing from the longlist is not in the comparison at all. We infer the loss shows up as RFQs that never arrive, which no CRM records.
- **Wrong facts are costly.** An assistant that says you lack a certification you hold, or places you in the wrong state, filters you out. Correcting a missing certification or a wrong plant location starts with the page the assistant read; our guide to [fixing wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) walks through it.

## How does GEO work for a contract manufacturer?

Generative engine optimization (GEO) makes your plant easy for AI assistants to find, describe accurately and verify.

For a contract manufacturer, the work usually covers:

1. **Capability pages in plain text.** One page per process with materials, tolerances, part sizes, volumes, secondary operations and lead-time ranges, written as text, not only in brochures, PDFs or photos.
2. **Certifications you can prove.** Each certification named with its scope, registrar and expiry, matching the registrar’s own listing.
3. **One consistent identity.** The same company name, plant addresses, capabilities and certifications on your site, Thomasnet, LinkedIn, association directories and customer supplier portals.
4. **Independent coverage.** Trade publication features, association memberships, case studies published with customer permission, conference talks and supplier awards. For why outside sources matter, see [how small brands get recommended by AI](https://underneath.agency/resources/how-small-brands-get-recommended-by-ai).
5. **Comparison-ready content.** Clear answers to how you compare with offshore sourcing on lead time and total cost, without naming rivals unfairly. Our review of [whether comparison pages help B2B brands get cited](https://underneath.agency/resources/do-comparison-pages-help-b2b-ai-citations) covers the limits.
6. **Crawl access.** Make sure your site allows the search crawlers the assistants document, and that key pages load without logins.
7. **Measurement.** Ask a fixed set of capability, certification and location questions across ChatGPT, Gemini, Perplexity, Copilot and Google’s AI features, repeatedly, and track who is named and which sources are cited. Then match those findings to RFQs received.

None of this guarantees a mention, and no honest provider can promise one. It makes your plant the easiest one for an assistant, and then an engineer, to verify. Hardware companies face the same verification test with their own buyers; makers of jobsite drones, robots and sensors are one example, covered in [how construction technology reaches contractor shortlists](https://underneath.agency/resources/construction-technology-demand-ai-search).

## Which questions can’t the data answer for manufacturers yet?

It shows engineers using AI in research, not how many RFQs or contracts AI answers create.

- **Usage is not attribution.** The TREW and GlobalSpec survey measures how buyers research, not which supplier choices came from AI. We found no public data linking AI answers to RFQ volume.
- **Surveys come from interested parties.** TREW sells marketing to engineering firms, GlobalSpec sells advertising, Thomas and Xometry run marketplaces, and the Reshoring Initiative promotes reshoring.
- **No manufacturing-specific ranking studies.** Our studies covered buyer questions across several industries, not supplier searches. Applying them to contract manufacturing is our inference.
- **Deal values are private.** Public figures show account counts and spending thresholds, not typical contract sizes for a job shop.

## Where should a contract manufacturer start?

Start by asking assistants the supplier questions your best customers would ask, and see whether you appear.

That first check usually shows whether your plant is named for your core processes and certifications, whether your locations and capabilities are described correctly, which directories and publications the answers rely on, and which rival shops are listed instead. The work then is to state your capabilities where assistants read them and earn the outside coverage that confirms them.

If your growth depends on a handful of new production accounts a year, [ask us to review your supplier visibility](https://underneath.agency/contact). We will show where your plant appears when buyers ask AI for suppliers, why other shops are named instead, and which changes are most likely to bring more qualified RFQs. Those changes, such as capability pages per process, provable certifications and one consistent plant identity, are what our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page covers.

## Frequently asked questions

### Do engineers really use ChatGPT to find suppliers?

Many use generative AI somewhere in the purchase: 69% in the 2026 TREW and GlobalSpec survey. Most still verify on supplier websites and in trade publications before contacting anyone.

### Does being on Thomasnet still matter?

Yes. Thomas reports more than 1.5 million monthly sourcing sessions and now offers natural-language AI search. Directory profiles are also outside sources an assistant can find, so keep them complete and consistent.

### Can a small job shop compete with large contract manufacturers in AI answers?

It can be named for specific combinations of process, material, certification and location, where fewer suppliers match. Large firms still benefit from recognition, and 70% of technical buyers lean toward the better-known brand.

### Will AI search replace our sales team?

No. Buyers contact suppliers for pricing, availability and complex requirements. AI changes who gets contacted, not the need for quoting and engineering conversations.

## Sources

- Distribution Strategy Group (2026-01-30), [Industrial Sourcing Behavior Shifts in 2025 (Thomas 2025 Annual Sourcing Activity Report)](https://distributionstrategy.com/2026/01/industrial-sourcing-behavior-shifts-in-2025-signaling-strategic-imperatives-for-distributors/)
- TREW Marketing and GlobalSpec (2026), [State of Marketing to Engineers research report](https://advertising.globalspec.com/state-of-marketing-to-engineers-research-report)
- Digital Commerce 360 (2026-01-19), [Thomas rolls out AI search and performance-based ads to speed industrial sourcing](https://www.digitalcommerce360.com/2026/01/19/thomas-ai-search-performance-based-ads-industrial-sourcing/)
- Xometry via GlobeNewswire (2026-02-24), [Xometry Reports Record Fourth Quarter and Strong Full Year 2025 Results](https://www.manilatimes.net/2026/02/24/tmt-newswire/globenewswire/xometry-reports-record-fourth-quarter-and-strong-full-year-2025-results/2283753)
- National Association of Manufacturers (2026), [Facts About Manufacturing](https://nam.org/mfgdata/facts-about-manufacturing-expanded/)
- US Census Bureau (2026-09-25), [Advance Report on Durable Goods Manufacturers’ Shipments, Inventories and Orders, August 2026](https://www.census.gov/manufacturing/m3/adv/pdf/durgd.pdf)
- Design News (2026), [US Reshoring Accelerates Despite Policy Uncertainties and Skilled Worker Shortages](https://www.designnews.com/industry/us-reshoring-accelerates-despite-policy-uncertainties-and-skilled-worker-shortages)
- Floor Covering News (2026-09), [Reshoring survey finds US manufacturing momentum building](https://www.fcnews.net/2026/09/reshoring-survey-finds-u-s-manufacturing-momentum-building/)
- Kearney (2026), [2026 Reshoring Index](https://www.kearney.com/documents/d/asset-library-291362522/2026-reshoring-index-pdf)
- EFESO (2026), [The State of Generative AI in European Procurement](https://www.efeso.com/wp-content/uploads/2026/04/EFESO-The-State-of-Generative-AI-in-European-Procurement.pdf)
- Forrester (2026-01-21), [Forrester’s 2026 Buyer Insights: GenAI Is Upending B2B Buying](https://www.forrester.com/press-newsroom/forrester-2026-the-state-of-business-buying/)
- OpenAI Help Center (2026), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Google (2025-03-05), [Expanding AI Overviews and introducing AI Mode](https://blog.google/products/search/ai-mode-search/)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [Do Wikipedia and schema make AI assistants recommend a brand?](https://underneath.agency/research/brand-entity-ai-recommendations-study)
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)

---

This is the Markdown twin of https://underneath.agency/resources/contract-manufacturers-rfqs-from-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How do makeup brands get recommended when shoppers ask AI?"
description: "By giving AI assistants shade, undertone and finish facts they can match to a shopper’s question, backed by reviews and independent coverage."
canonical: "https://underneath.agency/resources/cosmetics-brands-product-discovery-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How do makeup brands get recommended when shoppers ask AI?

By making every shade, undertone, finish and use case easy for AI assistants to read, compare and confirm from sources they trust. Makeup shoppers already ask AI which foundation matches their skin, and the answer names a handful of products. A brand whose shade facts are missing or inconsistent is easy to leave out, and a brand that is named can win a shopper who was not loyal to anyone.

## The short version

1. Makeup is the largest prestige beauty category: [Circana](https://www.circana.com/post/us-prestige-and-mass-beauty-retail-deliver-a-positive-performance-in-2025-circana-reports) reports US prestige beauty sales of $36 billion in 2025, with makeup up 4%, and mass-market beauty at $72.7 billion.
2. AI is already in the beauty aisle: in [Criteo’s June 2026 survey](https://www.criteo.com/blog/5-beauty-shopper-trends-2026/) across six countries, 38% of beauty and personal care shoppers use AI assistants, and 57% of those say the products recommended in AI answers influence what they buy.
3. Shade is where online makeup sales break: in The Benchmarking Company’s March 2026 survey of more than 4,300 American women, [reported by CosmeticOBS](https://client.cosmeticobs.com/en/articles/patterns-54/in-search-of-the-right-shade-9287), 41% had trouble finding their complexion shade, rising to 56% among women with darker skin.
4. Shoppers are open to switching: Criteo found only 34% rarely switch from their preferred beauty brands, and up to 53% of fragrance, skincare and makeup purchases came from a brand the shopper had not bought in the previous 12 months.
5. Retailers have moved: [Ulta Beauty](https://www.googlecloudpresscorner.com/2026-04-22-Ulta-Beauty-and-Google-Introduce-Gemini-Enabled-Shopping-Experiences-That-Streamline-Beauty-Discovery-and-Purchase) made its products shoppable inside Google’s AI Mode and the Gemini app in April 2026, a month after [Sephora](https://www.glossy.co/beauty/is-agentic-shopping-the-next-big-thing-in-beauty-sephora-and-ulta-are-betting-yes/) launched its app inside ChatGPT.

## How do people shop for makeup now, and why is each new customer worth fighting for?

They research hard, compare many products and switch brands easily, so every purchase is a chance to win someone new.

A makeup purchase looks small, but the decision behind it is not. Criteo’s commerce data shows US shoppers browse an average of 19 makeup products before buying. Reviews and ratings are the top purchase influence, with nearly half of shoppers relying on them, followed by friends and family, in-store testing and search engines. More than half of cosmetics and perfume purchases (57%) involve an online touchpoint.

Loyalty is thin. Because only about a third of shoppers rarely switch, and up to 53% of purchases go to a brand the shopper had not bought in a year, makeup brands win and lose customers constantly. The prize for being chosen is larger than one order: a foundation or concealer that matches becomes a repeat purchase, and a shopper who finds her shade tends to stop searching. We infer that the first match is where much of a complexion customer’s value is decided, though no public source we found puts a dollar figure on it.

The category is growing, but growth is uneven. Circana says all prestige makeup segments grew in dollars, with unit softness in face and eye makeup, and that lip was the fastest-growing segment in mass retail. Value, social media and “skinification,” hybrid products that blend color and skincare benefits, drove much of the growth. Each of those trends creates new questions that shoppers now put to AI. Jewelry shows the same effect, where the shift to lab-grown stones now has its own group of shopper questions, as our guide to [jewelry brands recommended by AI](https://underneath.agency/resources/jewelry-brands-ai-product-recommendations) shows.

## Where do AI assistants already sit in a makeup purchase?

At the start, as a shopping advisor that narrows dozens of options to a few named products.

The evidence is recent and consistent. Criteo’s figures above come from a survey of 4,595 shoppers in the US, UK, France, Germany, Japan and South Korea. [NielsenIQ](https://nielseniq.com/global/en/insights/analysis/2026/the-ai-beauty-advisor-era-has-arrived-is-your-product-content-ready/) reports that 49% of shoppers have already received beauty product recommendations from AI tools such as ChatGPT, Claude, Gemini and Copilot, and that 84% are more likely to buy when key product attributes can be easily compared. NielsenIQ’s point about the shelf is the one executives should remember: a physical shelf might show 50 or more products, while an AI answer may show only one or two.

Retail data points the same way, though it covers all of retail, not beauty alone. [Adobe’s analysis](https://techcrunch.com/2026/04/16/ai-traffic-to-us-retailers-rose-393-in-q1-and-its-boosting-their-revenue-too/) of more than 1 trillion visits to US retail sites found AI traffic rose 393% in the first quarter of 2026 compared with a year earlier, and that AI visitors converted 42% better than other traffic in March 2026.

[The biggest beauty retailers](https://underneath.agency/resources/beauty-retailers-ai-search) are building for this. Ulta, with more than 1,500 stores and more than 46 million loyalty members, says shoppers can now receive Ulta product recommendations, compare options and complete checkout for eligible purchases inside AI Mode and Gemini. Its chief technology officer said the company sees “a clear shift in how guests are discovering and shopping for beauty, with AI playing a much bigger role in that journey.” Sephora’s ChatGPT app launched in March 2026. Glossy’s reporters, testing the assistants, noted that ChatGPT gave long answers covering “the various finishes and the various undertones.” That is a description of one test, not a measurement, but it shows the vocabulary these answers run on.

## Which questions do makeup shoppers ask AI?

Questions about shade, undertone, finish, skin type, wear time, price and dupes, usually in one sentence.

We wrote the prompts below to illustrate the kinds of questions makeup shoppers type; they are not observed queries:

- Shade and undertone: “Which foundation shades suit medium skin with an olive undertone, in a satin finish?”
- Finish and skin type: “Best dewy foundation for dry, mature skin under $40.”
- Use case: “Long-wear concealer for dark circles on deep skin that won’t crease.”
- Wear: “Lipstick that doesn’t transfer onto a mask or a coffee cup.”
- Values: “Cruelty-free mascara for sensitive eyes that holds a curl.”
- Value and dupes: “Cheaper alternative to a prestige brow gel with the same hold.”

NielsenIQ lists similar real-world question shapes, such as a “good cruelty-free foundation for combination skin.” Google noted in 2022 that [foundation is the most-searched category within makeup](https://blog.google/products/shopping/more-ways-to-shop-in-ar). These questions differ from [skincare questions](https://underneath.agency/resources/skincare-brands-ai-search), which turn on ingredients and claims, and from retailer questions about where to buy. A makeup question is usually about how a product will look on one person’s face.

## How does an AI answer turn into a makeup sale?

Through a named product and shade, a click to a retailer or brand site, a first order, then repeat purchases.

**The answer.** The assistant names a few products, sometimes with a suggested shade. If your shade names, undertone labels and finish descriptions are clear, a reasonable expectation is that the assistant can match them to the shopper’s words; if they are vague, it has less to work with.

**The click or the checkout.** For most makeup brands, the sale lands at a retailer: Sephora, Ulta, Amazon or a mass chain. With Ulta’s catalog shoppable inside Google’s AI surfaces, and OpenAI documenting that [ChatGPT may show an Instant Checkout option](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search) for some eligible products and merchants, part of that sale may never touch a website. ChatGPT’s option is in flux, though: OpenAI scaled it back in March 2026, according to [FashionUnited](https://fashionunited.com/news/retail/us-consumers-would-let-ai-agents-buy-clothes-but-not-without-a-say/2026092974919), even as the help page still lists it. In-assistant buying makes sales harder to see in analytics, which is one reason brands should measure what AI says directly.

**The shade check.** This is where makeup differs from most categories. Google reported in 2022 that more than 60% of online beauty shoppers had decided not to buy a beauty item online because they did not know what color or shade to choose, and 41% had returned an item because it was the wrong shade. Perfect Corp, a try-on vendor, reports that Benefit Cosmetics saw a [113% increase in conversion rate](https://www.perfectcorp.com/business/blog/makeup/ecommerce-conversion-rate-optimization) and a 20% increase in add-to-cart actions after adding eyebrow virtual try-on. Those are vendor-reported results, but they show that confidence about shade is what turns interest into an order. Clothing has the same problem with fit, covered in [how apparel brands win customers through AI](https://underneath.agency/resources/apparel-brands-customers-ai-search).

**The repeat.** A matched shade is replenished. Our inference is that AI visibility for a complexion product pays back over several purchases, not one.

## What decides whether an AI assistant names your makeup product?

Platforms document product data, price and reviews; studies add that a small, visible quality edge can beat a famous name.

**Documented by the platform.** OpenAI says that when choosing products, ChatGPT considers structured metadata from first-party and third-party providers, such as price and product description, along with other third-party content and reviews. It can display review summaries built from public websites and labels such as “Most popular” that it generates itself. If the same foundation is sold by Sephora, Ulta and the brand itself, ChatGPT orders those sellers by availability, price, quality and whether the seller is the maker or primary seller. On Google’s side, AI Mode pulls from [shopping data for billions of products](https://blog.google/products/search/ai-mode-search/), firing off several related searches at once and then merging what they return. OpenAI’s documented “Try on” button covers clothes and accessories, not makeup, so shade confidence in AI answers still depends on words and data.

**Observed in a study.** In a [laboratory study of skincare recommendations](https://arxiv.org/abs/2606.17443) across three AI models, well-known brands were recommended 100% of the time when all products had the same specifications, but that dominance disappeared when a competitor had a rating advantage of less than 0.1 stars. The same study found that invented clinical claims also shifted picks, which is a reason for platforms to verify, not a tactic: we cover the risks in [can you game AI shopping rankings](https://underneath.agency/resources/can-you-game-ai-shopping-rankings). For makeup, we infer the practical reading is that a challenger brand needs a clear, verifiable difference, for example a wider shade range or better ratings from shoppers with a given skin tone.

**Trust factors specific to makeup.** Shade range and how inclusive it really is matter more here than in almost any category. In the Benchmarking Company survey, 39% of women had stopped using a brand because its range was not extensive enough, rising to 61% among women with darker skin. When a product failed, the reasons were the wrong undertone (66%), a color too light (52%) and a finish that looked different in photos than in real life (51%). Reviews that mention skin tone, undertone and wear time, retailer listings that agree with the brand site, and independent “best foundation for” coverage are the evidence an assistant can find. Our [guide to which pages to target](https://underneath.agency/resources/best-of-lists-ai-recommendations) explains why third-party lists matter.

## What does a makeup brand lose when AI leaves it out?

New customers at the moment they were ready to switch, and often the repeat orders that follow a shade match.

The evidence on lost sales is indirect, so we are careful here. What is documented: shoppers are open to new brands, more than a third use AI while shopping for beauty in Criteo’s data, and those who do mostly say the answers influence what they buy. Adobe found about 34% of retail product pages cannot be properly accessed by AI, across all retail, which suggests many brands start with a gap they have not measured. What we infer: when an answer names two or three foundations for “olive undertone, satin finish,” the brands not named lose a shopper who might have stayed for years. And when an answer names your product but suggests a shade from outdated or wrong data, the cost shows up as a return or a bad review instead of a missed sale. [What happens if you skip GEO](https://underneath.agency/resources/what-happens-if-you-skip-geo) covers the broader effect on search traffic.

## How does GEO work for a cosmetics brand?

It makes your shade, finish and use-case facts easy for AI to find, trust and repeat.

For a makeup brand, generative engine optimization (GEO) means getting AI answers to name your products and quote their shades and finishes correctly. It cannot buy a recommendation, and no one can promise a placement. For makeup, it usually covers seven things:

1. **Shade-level product facts.** Every shade with its depth, undertone, finish, coverage, skin type and wear claims, written the same way on your site, in product feeds and in retailer listings. NielsenIQ argues that AI systems rely on “structured, complete, and easily interpretable product information”; our article on [what product content AI shopping assistants prefer](https://underneath.agency/resources/product-content-ai-shopping-assistants-prefer) covers what tests show.
2. **Answers to shade questions.** Plain pages that answer the questions shoppers ask: how your undertone labels work, which shade suits which skin, how the finish wears on oily or dry skin, which shades replace a discontinued one.
3. **Reviews with context.** Encouraging honest reviews on retailer sites that mention skin tone, undertone and wear, since reviews are the top purchase influence and assistants summarize them. Never fake ones: [fake reviews and AI recommendations](https://underneath.agency/resources/fake-reviews-fake-brands-ai-recommendations) explains the risk.
4. **Independent coverage.** Beauty editors, creators’ written reviews and “best of” roundups that test your products across skin tones. Digital PR that earns this coverage gives assistants a source other than your own claims. Our [study of what makes assistants recommend a brand](https://underneath.agency/research/brand-entity-ai-recommendations-study) tests how much outside coverage matters.
5. **Consistent facts across retailers.** Shade names, prices and availability that match on Sephora, Ulta, Amazon and your site, so an assistant does not see conflicting information.
6. **Readable pages.** Product pages that AI systems can actually access and read, not shade data hidden in images or scripts.
7. **Measurement across AI surfaces.** Tracking a fixed set of shade, finish and dupe questions across ChatGPT, Google AI Mode, Gemini, Perplexity and retailer assistants, repeated over time, because answers vary from run to run.

## Which makeup questions does the research leave open?

How often AI answers name the right shade, and how much makeup revenue comes through AI today.

We found no public data on beauty sales completed inside AI assistants, on how accurate AI shade suggestions are, or on whether AI-referred makeup shoppers return products less often. The most specific survey figures come from Criteo and NielsenIQ, both companies that sell services to brands, and Criteo’s spans six countries rather than the US alone. Adobe’s conversion and traffic figures cover all US retail. The skincare study ran in a laboratory, with invented products, so it shows how models can respond, not what they do on live shopping pages. Google’s shade figures date from 2022. And the retailer launches at Ulta and Sephora are months old, with no published results yet.

## Where should a makeup brand start?

With your hero complexion products: check what AI assistants say about their shades, finishes and where to buy them.

Pick the 20 to 30 questions your shoppers really ask about your foundations, concealers and lip products, by shade and undertone. Run them across the main AI assistants and retailer assistants, and record which brands are named, which shades are suggested, which retailer gets the click and where the facts are wrong. That shows whether you have a visibility problem, an accuracy problem or both. Should you want a hand, [our team can run it with you](https://underneath.agency/contact): we will map how AI assistants describe your range today and set out the product-data, review and coverage work most likely to win more first purchases and fewer wrong-shade returns. Shade-level product facts, reviews with skin-tone context and matching retailer listings are each part of our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization), which explains how the work runs.

## Frequently asked questions

### Do AI assistants offer makeup virtual try-on?

Not in ChatGPT’s documented shopping features. OpenAI’s help page describes a “Try on” button for clothes and accessories. Makeup try-on lives on brand and retailer sites and in tools from vendors such as Perfect Corp, so in AI answers your shade descriptions do the work.

### Should a makeup brand focus on its own site or on retailers?

Both. Most makeup sales happen at retailers, and assistants read retailer listings and reviews as well as brand sites. The facts should match everywhere.

### Is AI visibility different for skincare and makeup?

Yes. Skincare questions turn on ingredients and claims, which carry regulatory limits. Makeup questions turn on shade, undertone, finish and wear, which are matters of accurate product description.

### Can a smaller makeup brand compete with the big names in AI answers?

In tests, yes, when it has a clear and checkable advantage. The evidence is gathered in [how small brands get recommended by AI](https://underneath.agency/resources/how-small-brands-get-recommended-by-ai) and [whether AI assistants favor big brands](https://underneath.agency/resources/do-ai-assistants-favor-big-brands).

## Sources

- Circana (February 10, 2026), [US Prestige and Mass Beauty Retail Deliver a Positive Performance in 2025](https://www.circana.com/post/us-prestige-and-mass-beauty-retail-deliver-a-positive-performance-in-2025-circana-reports)
- Criteo (2026), [New Health & Beauty research: 5 shopper trends every marketer should know](https://www.criteo.com/blog/5-beauty-shopper-trends-2026/)
- NielsenIQ (2026), [The AI beauty advisor era has arrived: Is your product content ready?](https://nielseniq.com/global/en/insights/analysis/2026/the-ai-beauty-advisor-era-has-arrived-is-your-product-content-ready/)
- CosmeticOBS, reporting The Benchmarking Company survey (2026), [In search of the right shade](https://client.cosmeticobs.com/en/articles/patterns-54/in-search-of-the-right-shade-9287)
- Google (November 17, 2022), [Use new AR features to shop for beauty products and shoes](https://blog.google/products/shopping/more-ways-to-shop-in-ar)
- Google Cloud and Ulta Beauty (April 22, 2026), [Ulta Beauty and Google Introduce Gemini-Enabled Shopping Experiences That Streamline Beauty Discovery and Purchase](https://www.googlecloudpresscorner.com/2026-04-22-Ulta-Beauty-and-Google-Introduce-Gemini-Enabled-Shopping-Experiences-That-Streamline-Beauty-Discovery-and-Purchase)
- Glossy (2026), [Is agentic shopping the next big thing in beauty? Sephora and Ulta are betting yes](https://www.glossy.co/beauty/is-agentic-shopping-the-next-big-thing-in-beauty-sephora-and-ulta-are-betting-yes/)
- TechCrunch, reporting Adobe data (April 16, 2026), [AI traffic to US retailers rose 393% in Q1, and it’s boosting their revenue too](https://techcrunch.com/2026/04/16/ai-traffic-to-us-retailers-rose-393-in-q1-and-its-boosting-their-revenue-too/)
- OpenAI (2026), [Shopping with ChatGPT search](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search)
- FashionUnited (September 29, 2026), [US consumers would let AI agents buy clothes, but not without a say](https://fashionunited.com/news/retail/us-consumers-would-let-ai-agents-buy-clothes-but-not-without-a-say/2026092974919)
- Google (March 5, 2025), [Expanding AI Overviews and introducing AI Mode](https://blog.google/products/search/ai-mode-search/)
- Perfect Corp (n.d.), [Ecommerce conversion rate optimization with makeup virtual try-on](https://www.perfectcorp.com/business/blog/makeup/ecommerce-conversion-rate-optimization)
- Chu and Hou (2026), [Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems](https://arxiv.org/abs/2606.17443)

---

This is the Markdown twin of https://underneath.agency/resources/cosmetics-brands-product-discovery-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Why do AI assistants keep naming the same CRMs?"
description: "Because the leaders dominate the sources AI search reads. Challenger CRMs get named by owning specific needs: an industry, a team size, a budget."
canonical: "https://underneath.agency/resources/crm-software-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Why do AI assistants keep naming the same CRMs, and how do we join them?

Because the category leaders are everywhere in the sources AI search reads, from their own help pages to Reddit threads and review videos, so generic “best CRM” questions return a familiar handful of names. Challengers get named when the question gets specific, by industry, team size, budget or integration, and when their own facts are easy to verify. For a CRM company, the path to more demos runs through those specific questions, not through the generic one.

## The short version

1. The market is concentrated: Salesforce, citing IDC, says it held a 20.7% share of the CRM market in 2024 and has been ranked first for the 12th consecutive year; HubSpot reports 306,000 customers and $3.13+ billion in 2025 revenue.
2. The leaders are also among the most-cited sources: in our study of 481 Google AI Overviews across eight industries, hubspot.com was cited in 29 (6.0%) and salesforce.com in 28 (5.8%).
3. Generic CRM questions are hard to move: in a monitoring study by Rankshift that we summarize in our phrasing research, HubSpot appeared at 98% on all seven near-synonym CRM prompts, with 1,176 runs each.
4. Specific questions change the list: in our phrasing study, adding a budget to business software questions cut the overlap with the original brand list to 0.363, against 0.708 when the same question was simply asked again.
5. Practitioner sources matter in this category: AI Overviews for B2B software searches cited Reddit in 35.4% of cases in our study, with r/crmsoftware among the three most-cited communities.

## Who chooses a CRM, and how long does a won account last?

Sales, marketing and service leaders buy it with IT, and the winner usually stays the system of record for years.

The category is led by a few large platforms. In a [Salesforce summary of IDC data](https://www.salesforce.com/news/stories/idc-crm-market-share-ranking-2025/), Salesforce led all CRM vendors in 2024 with a 20.7% share, its 12th consecutive year at number one. Its [second quarter of fiscal 2027](https://www.salesforce.com/news/press-releases/2026/08/26/fy27-q2-earnings/) brought revenue of $11.3 billion and a current remaining performance obligation, contracted revenue due within a year, of $33.5 billion. [HubSpot](https://www.hubspot.com/our-story) reports 306,000 customers and [$3.13+ billion in 2025 revenue](https://www.hubspot.com/company-news). [Microsoft](https://www.microsoft.com/en-us/investor/earnings/fy-2026-q4/press-release-webcast) reported Dynamics 365 revenue up 13% in the quarter to June 30, 2026.

Below them sits a long tail of challengers, each with a large base of its own. [Zoho](https://www.zoho.com/crm/) says its CRM is trusted by 300K+ businesses; [Pipedrive](https://www.pipedrive.com/en/about) reports 100,000+ companies using its CRM. Then come industry CRMs, for real estate, life sciences, legal, nonprofits and financial advisors, which often compete on fit rather than breadth.

A CRM is rarely replaced. Data migration, integrations and user training make switching painful, so a customer won is usually a multi-year subscription that expands by seats and add-ons. That makes the first shortlist unusually valuable. Marketing software buyers shortlist in a similar way, as our guide to [how marketing software wins AI-first buyers](https://underneath.agency/resources/marketing-software-ai-search-growth) shows.

## At which points in a CRM purchase do assistants show up?

At the start, where buyers ask for a shortlist, and at the comparison stage, where they check price and fit.

Software buyers in general lean on third-party sources to narrow their options. Gartner Digital Markets’ [2025 survey of 3,500 software buyers](https://cdn-static.bizzabo.com/bizzabo.file.upload/yVkIqwmWRDKbEhqeE3aQ_Gartner-Digital-Markets_2025-Report_Making-The-List.pdf) across categories found 41% use customer reviews to pare down their shortlist and 62% say the product trial is their top factor in the final decision. It also found 59% regret at least one software purchase made in the past 18 months, often replacing it with another vendor’s product. AI assistants now compress the review-reading step into one answer.

CRM is also one of the categories where Google’s AI answers are most common. Our [AI Overview frequency study](https://underneath.agency/research/ai-overviews-frequency-study) used “crm software” as one of its four B2B software seed topics. In [our citations study](https://underneath.agency/research/ai-overview-citations-study), brands’ own sites were among the most-cited sources in software, naming salesforce.com, hubspot.com, zapier.com and zoho.com. We have not found a published survey of how many CRM buyers start in ChatGPT or Gemini specifically, so we do not quote one; [our B2B SaaS article](https://underneath.agency/resources/b2b-saas-revenue-from-ai-search) covers the cross-category surveys.

## Which questions do CRM buyers ask AI assistants?

Mostly questions with a constraint attached: industry, team size, budget, sales motion or the tools the CRM must connect to.

These example prompts are ours, written to show how CRM buyers attach constraints to a question; we did not observe them:

- **Generic:** “What is the best CRM for a small business?”
- **Industry:** “Best CRM for a commercial real estate brokerage with 25 agents.”
- **Budget:** “Cheapest CRM with email sequences and a mobile app for a five-person sales team.”
- **Alternatives:** “Salesforce alternatives that are easier to administer without a full-time admin.”
- **Comparison:** “HubSpot vs Pipedrive vs Zoho for an outbound B2B team.”
- **Integration:** “Which CRM integrates best with Microsoft 365 and our NetSuite ERP?”

The first question is the one the leaders own. The rest are where answers open up. Our [prompt phrasing study](https://underneath.agency/research/ai-prompt-phrasing-study) found that for business software questions, a plain rerun of the same question kept 0.708 of the brands, adding context kept 0.526, and adding a budget kept only 0.363. When the buyer’s situation changes the question, the shortlist changes with it. Our [country study](https://underneath.agency/research/ai-recommendations-by-country-study) points the other way for geography: on global products such as CRM software, the country gap in ChatGPT’s brand lists was only 0.034.

## How does an AI answer turn into CRM pipeline?

Through the shortlist and the trial or demo: the answer names you, the buyer tests you, and the contract renews.

**Shortlist.** A buyer who asks for “the best CRM for a 25-agent brokerage” gets three to five names and often stops there. Being one of them is the whole game at this stage.

**Validation.** Buyers do not take the answer on trust. [Gartner reported in May 2026](https://www.gartner.com/en/newsroom/press-releases/2026-05-20-gartner-survey-finds-sixty-nine-percent-of-b-two-b-buyers-turn-to-sales-reps-to-validate-ai-generated-insights) that 69% of B2B buyers turn to sales reps to validate AI-generated insights. That cross-industry survey of 645 buyers also found 45% said they used generative AI during a recent purchase, mainly to gather information on vendors and products. A reasonable expectation is that your sales team will be asked to confirm or correct what an assistant said about your pricing, integrations and limits.

**Trial or demo.** Self-serve CRMs see this as signups; sales-led CRMs see it as demo requests. With a trial the top deciding factor for 62% of software buyers, the answer only opens the door; the product has to win the trial. HR platforms face the same gate, covered in [whether AI search decides HR software demos](https://underneath.agency/resources/hr-software-ai-search).

**Contract and expansion.** The value is in what follows: seats added as the team grows, then marketing, service and AI add-ons. That is why we infer that a CRM vendor should weigh AI visibility by the lifetime value of the segment asking, not by the number of mentions.

## Why do assistants put some CRMs on the list and leave others off?

Platforms document only how they search; studies show leaders dominate generic questions and practitioner sources matter.

**Documented by the platforms.** According to [Google](https://developers.google.com/search/docs/appearance/ai-features), AI Overviews and AI Mode may use “query fan-out”, running multiple related searches across subtopics and data sources, with no technical requirements beyond those of normal search. A question about a real estate CRM can therefore pull in pages on pricing, integrations and reviews at the same time.

**Observed in studies.** Leaders are hard to dislodge on generic prompts: Rankshift, a monitoring vendor whose study we summarize in our phrasing research, found HubSpot at 98% on every one of seven near-synonym CRM prompts, run 1,176 times each. Our [four-assistant study](https://underneath.agency/research/ai-assistants-brand-agreement-study), which used “What is the best CRM for a small business?” as an example question, found B2B software had the highest agreement between assistants, an overlap of 0.543. In Google’s answers, [our Reddit study](https://underneath.agency/research/ai-reddit-citations-study) found AI Overviews for B2B software cited Reddit in 35.4% of cases, and r/crmsoftware was cited 7 times, among the three most-cited communities. A YouTube channel named CRM Central was among the most-cited channels in our [YouTube study](https://underneath.agency/research/ai-overview-youtube-videos-study).

**Our inference about CRM-specific trust factors.** We expect assistants to rely on what CRM buyers check: review platforms, practitioner discussions, integration marketplaces, implementation partner content, industry associations for vertical CRMs, and clear public pricing. No platform lists any of them as something that moves a CRM up an answer. For why leaders tend to win close calls, see [our article on whether AI assistants favor big brands](https://underneath.agency/resources/do-ai-assistants-favor-big-brands).

## What does a CRM vendor lose when the answer skips it?

A seat-based contract that would have grown for years, lost before your sales team ever saw the lead.

Because most CRMs are chosen from a short list and rarely replaced, a vendor absent from the answer for its best-fit segment loses deals it never knew existed. The cost is highest for challengers in their home segment: if a brokerage asks for a real estate CRM and the answer lists only horizontal leaders, the specialist built for that buyer is invisible at the moment that matters. Purchase regret is common across software, so a buyer steered to a poor fit may come back to the market later, but the window can be years.

## What does GEO involve for a CRM vendor chasing specific segments?

Generative engine optimization (GEO) here means making your segment fit easy for AI search to find and check. A place on the shortlist is never assured.

1. **Pick the questions you can win.** Map the industries, team sizes, budgets and integrations where you are the best answer, and focus there rather than on “best CRM”.
2. **Publish fit pages.** Write a plain page for each industry and use case you serve, with real customer examples, limits and integrations.
3. **Make pricing readable.** State per-user prices, billing terms and what each plan includes in text, not only in images or calculators.
4. **Earn practitioner coverage.** Reddit threads, YouTube walkthroughs, partner blogs and independent reviews are cited sources in this category; see [our article on Reddit and AI Overviews](https://underneath.agency/resources/does-reddit-shape-google-ai-overviews).
5. **Be fair in comparisons.** Publish honest alternatives and comparison pages; see [our summary of comparison-page research](https://underneath.agency/resources/do-comparison-pages-help-b2b-ai-citations).
6. **Track specific prompts.** Monitor a fixed set of constrained buyer questions across ChatGPT, Gemini, Perplexity, Copilot and Google, and compare it with trials and demo requests. The questions you track decide what you see; see [how question phrasing changes AI sources](https://underneath.agency/resources/does-question-phrasing-change-ai-sources).

## What hasn’t been measured yet about AI answers and CRM sales?

No published study links CRM visibility in AI answers to demos, trials or revenue.

- The market-share figure comes from Salesforce’s own summary of IDC data, not from IDC directly.
- The Rankshift figure is a vendor’s monitoring result, reported in our study, not independently checked by us.
- Our studies are US, one-day snapshots across several categories; none isolates CRM.
- How much a single Reddit thread or review video changes a CRM’s chance of being named has not been measured.

## Which buyer questions should a CRM vendor test first?

The industry, team-size and budget questions your best-fit customers ask, checked against what the main assistants answer.

List those questions, plus the integration ones that precede your best deals, run them across the assistants, note who is named and what the answers cite, and rank the gaps by the contract value of each segment. We can run that ranking with you and tie it to seats, demos and renewals; [talk to us about a CRM visibility review](https://underneath.agency/contact). Our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page covers what comes next for a challenger CRM: fit pages by industry, readable per-user pricing and tracking of the constrained questions you can win.

## Frequently asked questions

### Why does ChatGPT recommend HubSpot and Salesforce so often?

They are the largest vendors and appear across the sources AI search reads, including their own sites, which Google’s AI Overviews cited often in our study. On generic CRM questions, one monitoring study found HubSpot named 98% of the time.

### Can a challenger CRM get onto AI shortlists?

Yes, for specific needs. Our phrasing research found that adding a budget or context changes most of the brands named, so challengers have room in constrained questions where they are the best fit. See [how small brands get recommended by AI](https://underneath.agency/resources/how-small-brands-get-recommended-by-ai).

### Do industry-specific CRMs have an advantage in AI search?

They can, when the buyer names the industry and the vendor’s pages state its fit clearly. The research has not measured this for CRM specifically.

### Does Reddit matter for CRM visibility in AI answers?

On Google, yes: AI Overviews for B2B software cited Reddit in 35.4% of cases in our study, and r/crmsoftware was among the most-cited communities.

## Sources

- Salesforce (2025), [Salesforce Named #1 CRM Provider by IDC Market Share for 2025](https://www.salesforce.com/news/stories/idc-crm-market-share-ranking-2025/)
- Salesforce (2026), [Salesforce Delivers Record Second Quarter Fiscal 2027 Results](https://www.salesforce.com/news/press-releases/2026/08/26/fy27-q2-earnings/)
- HubSpot (2026), [Our Story](https://www.hubspot.com/our-story) and [Company News](https://www.hubspot.com/company-news)
- Microsoft (2026), [FY26 Fourth Quarter Earnings Press Release](https://www.microsoft.com/en-us/investor/earnings/fy-2026-q4/press-release-webcast)
- Zoho (2026), [Zoho CRM](https://www.zoho.com/crm/)
- Pipedrive (2026), [About Pipedrive](https://www.pipedrive.com/en/about)
- Gartner Digital Markets (2025), [Making the List: How Software Buyers Pare Down Their Options](https://cdn-static.bizzabo.com/bizzabo.file.upload/yVkIqwmWRDKbEhqeE3aQ_Gartner-Digital-Markets_2025-Report_Making-The-List.pdf)
- Gartner (2026), [Gartner Survey Finds 69% of B2B Buyers Turn to Sales Reps to Validate AI-Generated Insights](https://www.gartner.com/en/newsroom/press-releases/2026-05-20-gartner-survey-finds-sixty-nine-percent-of-b-two-b-buyers-turn-to-sales-reps-to-validate-ai-generated-insights)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- Underneath (2026), [AI Overview citations](https://underneath.agency/research/ai-overview-citations-study), [prompt phrasing](https://underneath.agency/research/ai-prompt-phrasing-study), [Reddit citations](https://underneath.agency/research/ai-reddit-citations-study), [YouTube videos](https://underneath.agency/research/ai-overview-youtube-videos-study), [four-assistant agreement](https://underneath.agency/research/ai-assistants-brand-agreement-study), [country differences](https://underneath.agency/research/ai-recommendations-by-country-study) and [AI Overview frequency](https://underneath.agency/research/ai-overviews-frequency-study)

---

This is the Markdown twin of https://underneath.agency/resources/crm-software-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How support software gets chosen when buyers ask AI"
description: "By being named while support leaders re-shop for AI agents, with clear outcome pricing and proof of resolution rates that buyers and AI answers can check."
canonical: "https://underneath.agency/resources/customer-support-software-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How do customer support software companies get chosen when buyers ask AI?

By being named, accurately, at the moment support leaders are re-shopping their helpdesk for AI agents, and then proving in a trial that the resolution rates hold. Almost every support team is re-evaluating its tools at once, and pricing is moving from seats to outcomes, so buyers have more questions than ever and fewer settled answers. That makes this category unusually open to vendors who are easy for AI assistants to describe correctly.

## The short version

1. [Gartner’s survey](https://www.gartner.com/en/newsroom/press-releases/2024-12-09-gartner-survey-reveals-85-percent-of-customer-service-leaders-will-explore-or-pilot-customer-facing-conversational-genai-in-2025) of 187 support leaders found 85% would explore or pilot customer-facing conversational AI in 2025, and [Gartner predicts](https://www.gartner.com/en/newsroom/press-releases/2025-03-05-gartner-predicts-agentic-ai-will-autonomously-resolve-80-percent-of-common-customer-service-issues-without-human-intervention-by-20290) AI agents will resolve 80% of common service issues without a person by 2029.
2. Among enterprise technology leaders interviewed by [Andreessen Horowitz](https://a16z.com/ai-enterprise-2025/), over 90% said they were testing third-party apps for customer support, which means they are shopping.
3. Pricing is changing under buyers’ feet: [Intercom](https://fin.ai/pricing) charges $0.99 per outcome for its AI agent, and [Zendesk](https://www.zendesk.com/pricing/) bills AI agents per automated resolution on top of per-agent seats. In [our pricing study](https://underneath.agency/research/ai-pricing-accuracy-study), Zendesk tiers that no longer exist showed up in AI answers.
4. The money is upmarket: at [Freshworks](https://s21.q4cdn.com/987526491/files/doc_financials/2026/q2/Freshworks_Q2_2026_Earnings_Call_Transcript.pdf), customers paying over $100,000 a year grew 25% and make up about 40% of total recurring revenue, while its helpdesk business grew 3%, a sign that share in this market is won by displacing incumbents.
5. [HubSpot](https://transcripts.platformaeronaut.com/transcripts/HUBS-2Q25-transcript), which sells a helpdesk as well as a CRM, told investors that clicks from AI assistants convert much better than organic search because buyers arrive deep in their research.

## Who signs for a helpdesk, and how much is a support account worth?

A head of support or customer experience leads the choice; IT, finance and security sign off on larger deals.

The buyer owns a cost center under pressure and a customer experience under scrutiny. [Zendesk’s CX Trends research](https://cxtrends.zendesk.com/) reports that 85% of customer experience leaders say customers will drop brands over unresolved issues, even on first contact, and that 74% of consumers now expect customer service to be available around the clock because of AI. Those two numbers explain the purchase: leaders need more resolutions without proportionally more agents.

The value of a customer depends on team size and, increasingly, on how much automated work it buys. Freshworks reported that customers contributing more than $50,000 a year now represent over 55% of its total recurring revenue. Its helpdesk business ended the second quarter of 2026 at $400 million in annual recurring revenue, growing 3% year over year, and the company expects low-single-digit growth for the year. In a market growing that slowly, new revenue mostly comes from winning accounts away from someone else.

Pricing is in transition. Zendesk describes its pricing as primarily seat-based, per agent per month, with AI agents billed by “automated resolutions,” meaning requests resolved without escalation to a person. Intercom prices its Fin AI agent at $0.99 per outcome, and $9.99 for each prospect it qualifies and routes to sales. For a buyer, that turns a familiar question (“how many seats do we need?”) into a harder one: how many conversations will the AI actually resolve, and what will that cost?

## Why is almost every support team shopping at once?

Because AI agents changed what a helpdesk is expected to do, and leaders are under pressure to adopt them.

Gartner’s survey found 44% of support leaders exploring a customer-facing voicebot, 11% piloting one and 5% already deployed. Gartner also expects the shift to cut operating costs by 30% by 2029. Those are forecasts, but they describe the brief that support leaders are receiving from their own executives.

Vendors are answering with resolution claims, all self-reported and measured differently:

- Intercom says Fin averages a 76% resolution rate across 12,000+ customers and handles 2 million resolutions a week.
- HubSpot said its customer agent had 4,000 customers in mid-2025, with resolution rates averaging 55%.
- Freshworks said its AI agent’s deflection rates average 50%, reaching as high as 80% for mature deployments.

Each vendor counts a “resolution” its own way, so buyers ask someone to compare them. Increasingly that someone is an AI assistant, followed by a trial.

## At what point do support leaders bring AI assistants into a helpdesk search?

At the comparison stage, and the published signals come from vendors who sell support software.

HubSpot’s chief executive told investors in August 2025 that organic search traffic is declining as AI Overviews give answers, that the company is now cited by AI assistants more than any other CRM, and that clicks from AI assistants convert much better than organic search because buyers are already deep in their research. The same call noted that over 20,000 HubSpot customers had used its ChatGPT and Claude connectors. HubSpot sells to the same mid-market teams that buy helpdesks, so its experience is the closest public signal for this category, though it reports it for its business as a whole.

Outside helpdesks, [G2’s 2026 survey](https://company.g2.com/news/g2-research-the-answer-economy) of software buyers in general found 51% of buyers now start research with an AI chatbot more often than with Google. G2 sells to software vendors, so read it as a direction rather than a measurement for helpdesks specifically. Tools sold to the same teams are covered in [how collaboration software wins AI-first teams](https://underneath.agency/resources/collaboration-software-ai-search).

## Which questions do support leaders ask AI assistants?

Questions about switching, AI resolution, total cost, integrations and fit for their team size.

We wrote the prompts below to show how a head of support might phrase a helpdesk search; none was taken from a real buyer:

- Switching: “Best Zendesk alternatives for a 40-agent B2B support team that wants AI agents included.”
- AI performance: “Which helpdesk AI agents actually resolve billing questions without a human?”
- Cost: “How much would Intercom Fin cost us at 30,000 conversations a month compared with Zendesk?”
- Integrations: “Which support platforms integrate natively with Salesforce, Jira and Shopify?”
- Fit: “Best helpdesk for an ecommerce brand with 5 agents and seasonal spikes.”
- Risk: “Which customer service platforms are HIPAA compliant and let us keep data in the EU?”

Cost questions are where answers are most likely to go wrong, because outcome pricing is new and older pages still carry older prices. In our pricing study, 39 of the 64 prices that differed from the vendor’s page were found on another page of the vendor’s own site, including Zendesk’s article on its 2023 pricing update.

## How does a helpdesk named in an AI answer become a signed contract?

The assistant puts you on the shortlist; a trial proves the resolution rate; the contract follows.

**Shortlist.** A support leader asks for alternatives or a comparison and gets a short list with reasons. If your product is not on it, the evaluation proceeds without you.

**Trial.** This category lets buyers test before talking to sales. Intercom offers a 14-day trial with no credit card and unlimited Fin outcomes during it; Zendesk also runs a 14-day trial. A buyer can point the AI agent at its own help center and see how many questions it resolves, so a good trial converts the claim into evidence.

**Demo and procurement.** Larger teams add security review, integration checks and a migration plan. Freshworks’ figures show where the value sits: customers paying over $100,000 a year grew 25% year over year. Those deals rarely show up as a single AI referral; we infer they appear as direct visits, branded searches and demo requests from a buyer who already knows your name. Sales software vendors face the same pattern, covered in [turning AI answers into sales pipeline](https://underneath.agency/resources/sales-software-pipeline-from-ai-search).

**Expansion.** Outcome pricing means revenue grows as the AI resolves more conversations, so the first deal is the start of usage-based growth rather than a fixed seat count.

## Why do assistants name some helpdesks and not others?

The assistants explain little; the research favors accurate pages, independent comparisons and current, structured information.

**What Google has said.** Google’s description of AI Mode includes a [“query fan-out” technique](https://blog.google/products/search/ai-mode-search/) that issues multiple related searches across subtopics, then combines the results. A question about Zendesk alternatives can pull in pricing pages, review sites, comparison articles and help-center pages at once.

**Observed in our studies.** In [our study of self-ranking “best of” lists](https://underneath.agency/research/self-promoting-best-lists-study), 24.2% of AI-cited numbered lists with an identifiable publisher ranked that publisher first, and Zendesk was among the most frequent publishers of such lists, with 3. We found no statistically detectable difference in how often answers named the top pick of a self-ranking list versus an independent one, so these lists neither clearly help nor clearly hurt. In [our study of agent-readable websites](https://underneath.agency/research/agent-readable-web-study), only 3.2% of top websites returned a plain text version when an AI agent asked for one; zendesk.com and intercom.com were among the early adopters.

**Observed in other research.** A study of [B2B software citations](https://arxiv.org/abs/2509.10762) using 70 industry-targeted prompts found the page qualities most strongly associated with being cited were metadata and freshness, clear page structure and structured data. That is a correlation, not proof that adding them causes citations.

**Trust factors specific to support software.** Buyers weigh how a vendor measures resolution, which integrations work out of the box, how migration is handled, and whether security and compliance claims can be verified. A reasonable expectation is that assistants answering those questions lean on whatever is public: your pricing page, trust center, integration directory, help center and independent reviews.

## What does a helpdesk vendor lose when an assistant leaves it off the list?

The switch you never heard about, in a market where switching is how share moves.

When a helpdesk market grows in the low single digits, as Freshworks expects for its own helpdesk business, the main source of new revenue is accounts leaving an incumbent. Those moments are now frequent, because AI agents are prompting re-evaluations across the market. We infer that an AI answer that omits you, or quotes your old pricing, costs you the shortlist at exactly the time buyers are most willing to move. Vendors with wrong prices in AI answers face a second cost: a buyer who sees an outdated figure may rule you out before visiting your site. We look at how falling search clicks affect pipeline more generally in [our article on AI answers and pipeline](https://underneath.agency/resources/ai-answers-pipeline-revenue).

## What should GEO cover for a helpdesk or AI agent vendor?

Your pricing, resolution claims and integrations, made easy for assistants to describe correctly, with no guaranteed recommendation.

For a helpdesk vendor, generative engine optimization (GEO) tends to come down to six pieces of work:

1. **Pricing that explains itself.** Publish one current page that states seat prices, outcome prices and what counts as a resolution, and retire or redirect old pricing articles. When an assistant still quotes a retired tier, follow the steps in [how to fix wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers).
2. **Resolution claims with a method.** State how you measure resolution or deflection, on what kind of traffic, and link to customer examples. Unexplained percentages are hard for buyers, and assistants, to compare.
3. **Honest comparison and migration pages.** Buyers ask switching questions, so publish fair “from Zendesk” or “from Freshdesk” migration guides and comparisons. Whether such pages earn citations is weighed in our [comparison pages](https://underneath.agency/resources/do-comparison-pages-help-b2b-ai-citations) article.
4. **Independent coverage and reviews.** Earn places in the comparison sites, analyst notes, newsletters and communities support leaders read, and keep review profiles current. Our article on [which “best of” lists matter](https://underneath.agency/resources/best-of-lists-ai-recommendations) helps pick targets.
5. **A help center built to be read.** Support vendors know this discipline better than anyone: Gartner found 61% of support leaders have a backlog of knowledge articles to edit. Apply the same standard to your own documentation, integration pages and trust center, with clear structure and visible update dates.
6. **Measurement by buying stage.** Track switching, AI-performance, cost, integration and compliance questions across ChatGPT, Gemini, Perplexity, Copilot and Google, and connect the results to trials, demos and pipeline.

## What remains unmeasured about AI search in helpdesk buying?

Nobody has published how AI visibility shifts win rates or contract value for helpdesk vendors.

Most of the numbers here are vendor-reported, including every resolution rate, and they use different definitions. The Gartner figures are a survey and a forecast. HubSpot’s comments describe its business as a whole, not only its helpdesk. G2 and Andreessen Horowitz describe software buying broadly, not this category. Our own studies record what AI answers cite and say about vendors such as Zendesk and Intercom, not which helpdesk a buyer then signs with. Treat AI assistants as an established part of the support software shortlist, with the size of their effect on revenue still unmeasured.

## How can a helpdesk vendor see which AI answers are costing it trials?

Compare what assistants say about your pricing, resolution rates and alternatives with what your trial data shows.

Run the switching, cost, AI-performance and integration questions your buyers ask through the main assistants and Google’s AI features, then match the answers against your trial and demo records to see which wrong or missing answers cost you evaluations, seats or outcome-priced volume. To have us do that work alongside your team, [talk to us about a helpdesk visibility audit](https://underneath.agency/contact). The follow-on work, including pricing pages that explain outcome charges and resolution claims stated with their method, is described on our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page.

## Frequently asked questions

### Do AI assistants get helpdesk pricing right?

Often not fully. In our study of software prices quoted by four assistants, many of the differences came from older pages on the vendor’s own site, and outdated Zendesk tiers were among them. Outcome-based pricing adds a new way to be misquoted, so it pays to check.

### Should we publish our AI agent’s resolution rate?

Yes, with the method. Buyers compare resolution claims across vendors, and a number without a definition is hard to trust or compare. Say what counts as resolved, on which kinds of conversations, and show examples.

### Are “best helpdesk software” lists on our own site worth it?

They are common in this category and AI engines cite them. Our study found about one in four cited numbered lists ranked their own publisher first, with no detectable difference in how often answers named that pick. A fair comparison page is a safer investment.

### How do we know if AI assistants are sending us trials?

Watch AI referrals into your trial and demo pages, ask new signups how they heard about you, and recheck how often assistants name your helpdesk for buyers’ questions. Then trace those accounts to the seats and resolution volume they end up buying.

## Sources

- Gartner (2024), [Gartner Survey Reveals 85% of Customer Service Leaders Will Explore or Pilot Customer-Facing Conversational GenAI in 2025](https://www.gartner.com/en/newsroom/press-releases/2024-12-09-gartner-survey-reveals-85-percent-of-customer-service-leaders-will-explore-or-pilot-customer-facing-conversational-genai-in-2025)
- Gartner (2025), [Gartner Predicts Agentic AI Will Autonomously Resolve 80% of Common Customer Service Issues Without Human Intervention by 2029](https://www.gartner.com/en/newsroom/press-releases/2025-03-05-gartner-predicts-agentic-ai-will-autonomously-resolve-80-percent-of-common-customer-service-issues-without-human-intervention-by-20290)
- Andreessen Horowitz (2025), [How 100 Enterprise CIOs Are Building and Buying Gen AI in 2025](https://a16z.com/ai-enterprise-2025/)
- Freshworks (2026), [Q2 2026 earnings call transcript](https://s21.q4cdn.com/987526491/files/doc_financials/2026/q2/Freshworks_Q2_2026_Earnings_Call_Transcript.pdf)
- HubSpot (2025), [Q2 2025 earnings call transcript](https://transcripts.platformaeronaut.com/transcripts/HUBS-2Q25-transcript)
- Intercom (2026), [Fin pricing](https://fin.ai/pricing) and [Fin](https://fin.ai/)
- Intercom (2026), [Intercom pricing](https://www.intercom.com/pricing)
- Zendesk (2026), [Zendesk pricing](https://www.zendesk.com/pricing/)
- Zendesk (2026), [CX Trends 2026](https://cxtrends.zendesk.com/)
- G2 (2026), [In the Answer Economy, Don’t Win the Click — Win the Answer](https://company.g2.com/news/g2-research-the-answer-economy)
- Google (2025), [Expanding AI Overviews and introducing AI Mode](https://blog.google/products/search/ai-mode-search/)
- AI Answer Engine Citation Behavior: Bringing the GEO-16 Framework in B2B SaaS (2025), [AI Answer Engine Citation Behavior An Empirical Analysis of the GEO16 Framework](https://arxiv.org/abs/2509.10762)
- Underneath (2026), [software pricing accuracy](https://underneath.agency/research/ai-pricing-accuracy-study), [self-ranking “best of” lists](https://underneath.agency/research/self-promoting-best-lists-study) and [agent-readable websites](https://underneath.agency/research/agent-readable-web-study)

---

This is the Markdown twin of https://underneath.agency/resources/customer-support-software-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How cybersecurity software firms can earn revenue from AI search"
description: "By getting named in the AI answers that now build security shortlists, backed by third-party proof that CISOs and AI assistants can both check."
canonical: "https://underneath.agency/resources/cybersecurity-software-revenue-from-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How can cybersecurity software companies generate revenue from AI search?

By earning a place in the AI answers where security buyers now form their first shortlist, because the shortlist is where enterprise deals are won or lost. More software buyers now start vendor research with an AI chatbot, as the surveys covered below show. In security, whether an assistant names you depends less on what you say about yourself than on proof that others publish and buyers can verify.

## The short version

1. The security market is crowded and committee-driven: Momentum Cyber’s market map, as cited by [Expert Insights](https://expertinsights.com/industry-perspectives/how-senior-cisos-actually-vet-security-vendors), tracks more than 1,500 cybersecurity companies, and ActualTech Media found decision committees of 11 or more people 60% of the time at companies with more than 5,000 employees. No public study yet measures how many CISOs use a chatbot to narrow that field.
2. Security buyers distrust what vendors tell them: in a [Sophos-commissioned survey](https://assets.sophos.com/X24WTUEQ/at/bhmx4pwtc5vz7scc97jjb4c9/sophos-the-cybersecurity-trust-reality-in-2026.pdf) of 5,000 IT and security leaders, only 5% fully trusted their security vendors and 79% found new vendors hard to assess.
3. The same survey ranked verifiable evidence, such as certifications, third-party assessments and a public trust center, as the top driver of trust in a security vendor.
4. For US software questions, AI search drew 72.7% of its sources from independent “earned” sites, against 45.4% for Google, according to [Chen and colleagues](https://arxiv.org/abs/2509.08919).
5. The prize is large but slow to win: [Gartner](https://www.crnasia.com/india/news-network/news/gartner-forecasts-worldwide-end-user-spending-on-information-security-to-total-213-bn-in-2025) forecasts $121,154 million of security software spending in 2026, and most security purchases run on a 6–12-month cycle.

## Who buys cybersecurity software, and what is a customer worth?

A security-led committee buys it slowly and carefully, and a won enterprise customer is worth a lot.

Security spending is big and still rising. Gartner forecasts worldwide end-user spending on information security of $213 billion in 2025, rising 12.5% to $240 billion in 2026. Security software is the largest slice and, Gartner says, the fastest growing one: from $105,940 million in 2025 to $121,154 million in 2026.

Each buyer’s budget is tighter than those totals suggest. [IANS Research and Artico Search](https://www.iansresearch.com/resources/press-releases/detail/ians-research-and-artico-search-release-security-budget-benchmark-report), surveying 587 CISOs, found average security budget growth slowed to 4% in 2025, down from 8% in 2024. A new vendor is usually asking for money that was planned for something else.

The buying group is large, and it is not only the CISO:

- [Forrester’s 2026 study of business buying](https://www.forrester.com/press-newsroom/forrester-2026-the-state-of-business-buying/) found the typical decision involves 13 internal stakeholders and nine external influencers, and procurement staff are decision-makers in 53% of buying cycles.
- When a purchase includes generative AI features, which most security products now claim, Forrester found the [buying group doubles](https://forrester.com/blogs/state-of-business-buying-2026), from seven members to 14.
- In [ActualTech Media’s survey](https://www.actualtechmedia.com/blog/cybersecurity-buyers-report-part-3/) of 327 US security professionals, run in late 2023, most prospects were on a 6–12-month purchase cycle. In companies with more than 5,000 employees, the decision committee had 11 or more people 60% of the time.
- In the Sophos survey, only 1% of organizations said senior leadership plays no role in security purchases.

A listed security vendor’s results show what one enterprise logo can be worth. [SentinelOne](https://www.01net.it/sentinelone-announces-fourth-quarter-and-fiscal-year-2026-financial-results/) reported 1,667 customers paying $100,000 or more a year as of January 31, 2026, within $1,119.1 million of annualized recurring revenue. Subscription contracts like these are built to renew, so one enterprise win can pay out for years.

The buyer’s stakes are high too, which explains the caution. [IBM’s 2025 breach report](https://newsroom.ibm.com/2025-07-30-ibm-report-13-of-organizations-reported-breaches-of-ai-models-or-applications,-97-of-which-reported-lacking-proper-ai-access-controls) put the average cost of a US data breach at a record $10.22 million, against $4.44 million worldwide. A CISO who picks the wrong tool carries that risk personally.

## How early do AI assistants enter a security purchase?

At the start of research and, increasingly, at the shortlist, though no public study isolates security buyers alone.

The best recent evidence covers B2B software in general. In March 2026, [G2](https://company.g2.com/news/g2-research-the-answer-economy), which runs a software review platform, surveyed 1,076 software buyers: 51% started research with an AI chatbot more often than with Google, up from 29% in G2’s 2025 report. 71% relied on AI chatbots at some point, and 41% used Deep Research tools for structured software evaluations.

In [G2’s 2025 survey](https://research.g2.com/cmos-2025-buyer-behavior-report-research-g2) of 1,100 decision-makers, AI chatbots were already the source that most influenced vendor shortlists:

| Source that most influenced the shortlist | Share of buyers |
|---|---|
| Generative AI chatbots | 17.1% |
| Software review sites | 15.1% |
| Vendor websites | 12.8% |
| Market research firms | 10.6% |

Enterprise buyers, at companies with 1,000 to 5,000 employees, named review sites (56%) and AI search (55%) as their top two research sources. That pairing matters for security, where peer validation has always carried weight.

Two cautions. G2 runs a software review marketplace and sells vendors visibility on it, so it has an interest in these findings. And Forrester reports that answer engines “often deliver incomplete or unreliable information, creating mistrust,” so buyers check AI answers against trusted sources. In security, that checking step is where deals are decided.

The market also gives buyers good reason to ask a machine for help. [Expert Insights](https://expertinsights.com/industry-perspectives/how-senior-cisos-actually-vet-security-vendors) notes that Momentum Cyber’s market map tracks more than 1,500 cybersecurity companies. Narrowing that field to a handful is exactly the job buyers now hand to assistants, we infer. Zero trust access is one such field, covered in [how network security vendors win zero trust buyers](https://underneath.agency/resources/network-security-zero-trust-customers-ai-search).

## What do CISOs and security architects ask AI assistants?

Questions about categories, alternatives, head-to-head comparisons, compliance and proof. We wrote the security prompts in this table ourselves to illustrate the stages; none was captured from a real buyer.

| Buying stage | Illustrative prompt |
|---|---|
| Category | “What are the best endpoint detection and response platforms for a mid-sized hospital network?” |
| Alternatives | “What are the main alternatives to CrowdStrike Falcon for a manufacturer with a small security team?” |
| Comparison | “SentinelOne or Microsoft Defender for Endpoint for a company already on Microsoft 365?” |
| Compliance | “Which SIEM vendors are FedRAMP authorized?” or “Which data loss prevention tools help with HIPAA audits?” |
| Proof | “How did these vendors do in the latest MITRE ATT&CK Evaluations?” |
| Due diligence | “Has this vendor had a breach, and how did it handle disclosure?” |
| Fit | “Does this tool integrate with Splunk and ServiceNow?” |

A single security question can trigger a batch of web searches behind the scenes. Google documents that AI Overviews and AI Mode [may use a “query fan-out” technique](https://developers.google.com/search/docs/appearance/ai-features), “issuing multiple related searches across subtopics and data sources.” OpenAI documents that [ChatGPT search rewrites a question](https://help.openai.com/en/articles/9237897-chatgpt-search) into one or more targeted queries that it sends to search partners.

Our own research shows which kinds of pages ChatGPT goes looking for in those searches. In [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study) of 80 buyer questions across eight industries, ChatGPT searched for reviews in 46.2% of its answers and for a named publication, ranking or award in 43.8%. In security, we infer, those named sources would include analyst reports, independent test results and peer-review platforms. Endpoint vendors lean on such results, as our guide to [how EDR vendors win enterprise deals](https://underneath.agency/resources/endpoint-security-enterprise-customers-ai-search) explains.

## How does a mention in an AI answer become a security contract?

Via the shortlist: a named security vendor gets evaluated, and evaluation leads to a proof of concept, then a contract.

The path, as we understand it from the evidence above, has five steps:

1. A CISO, architect or IT lead asks an assistant a category, alternatives or compliance question.
2. The assistant replies with a few security vendors and links to a handful of supporting pages.
3. The buyer checks those vendors against evidence: review sites, certifications, test results, peers.
4. The vendors that survive are invited into a request for proposal or a proof of concept.
5. One wins a multi-year contract that renews and expands.

The surveys suggest the AI answer weighs heavily at step 2. In G2’s 2026 data, 85% of buyers thought more highly of a vendor when AI included it in an answer, and 69% chose a different vendor than they had planned because it was part of a chatbot’s recommendation. Step 3 is where security differs from other software: the buyer and the board want proof, not claims.

A CISO who first met you in ChatGPT seldom arrives through a link your analytics can trace. A buyer who met you in an AI answer may arrive months later through a reseller, a peer introduction or a direct demo request. That makes attribution hard, a problem we cover in [what lost clicks mean for pipeline](https://underneath.agency/resources/ai-answers-pipeline-revenue) and [how to prove GEO caused sales](https://underneath.agency/resources/prove-geo-caused-sales).

## Why does an assistant name one security vendor and not another?

The platforms do not say; studies point to independent sources, and security buyers reward proof they can check.

**What the platforms disclose.** Google and OpenAI say their AI answers search the web and link to the pages they draw on. Neither explains how it picks the security products it recommends.

**What researchers have measured.** In the analysis of US software questions by [Chen and colleagues](https://arxiv.org/abs/2509.08919), AI search drew 72.7% of its sources from earned sites such as reviews and independent publications, against 45.4% for Google, which leaned more on vendor-owned pages. That matters for due-diligence questions: [our reputation study](https://underneath.agency/research/is-it-legit-ai-reputation-study) found 88.0% of answers to “Is this brand legit?” cited a review or complaint platform. In [our Reddit study](https://underneath.agency/research/ai-reddit-citations-study), r/cybersecurity was among the communities Google’s AI cited most.

B2B software is also where assistants agree most. In [our four-assistant study](https://underneath.agency/research/ai-assistants-brand-agreement-study), the overlap between two assistants’ recommendations was highest for B2B software, at 0.543 on a scale from 0 to 1, against 0.227 for home and local services. Agreement is still far from complete.

**What security buyers trust.** The Sophos survey, run by the research firm Vanson Bourne, ranked what most raises trust in a security vendor:

| Rank for boards | Rank for security teams | Driver of trust |
|---|---|---|
| 1 | 1 | Verifiable evidence: bug bounty, trust center, advisories, third-party assessments, certifications |
| 2 | 3 | Transparent, timely communication during incidents |
| 3 | 4 | Expert commentary after major incidents |
| 5 | 5 | Performance in analyst reports |
| 7 | 7 | Performance in independent tests such as MITRE or SE Labs |

The barriers were about information: 47% said vendor information was not factual or detailed enough, 41% met conflicting information, and 38% struggled to find what they needed. Sophos is itself a security vendor, so treat the survey as one input.

**Our inference.** The evidence a CISO uses to verify a vendor is mostly public: certifications, test results, advisories, reviews, analyst mentions. An assistant can reach the same pages, provided they are not locked behind a form. So a security vendor that keeps its proof public, current and consistent should, we expect, give the CISO and the assistant more to go on. Nobody has tested that expectation on security vendors yet.

## What is at stake for a security vendor that assistants skip?

Mostly places in RFPs and proofs of concept, not clicks; no study has priced that loss in dollars.

- **The first list is hard to join later.** If AI answers shape the shortlist, a vendor left out may never be evaluated. G2 puts it bluntly: if AI leaves you out, “the buyer may not even know you exist.” G2 is not neutral on this point.
- **Long cycles raise the price of a miss.** With most purchases on a 6–12-month cycle and budgets set annually, missing one evaluation can mean waiting until the winner’s contract comes up for renewal, we infer. A SIEM migration is one such window, covered in [how SIEM vendors reach enterprise shortlists](https://underneath.agency/resources/siem-enterprise-pipeline-ai-search).
- **Shortlist presence flickers.** Across five runs of the same question in [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), just 25.2% of the brands ChatGPT named showed up every time. A security vendor that appears in some runs and not others drops out of evaluations without ever learning they happened.
- **Distrust drives switching.** In the Sophos survey, 45% said a lack of trust makes them more likely to switch vendors. An AI answer that describes your product vaguely or wrongly is, we infer, a renewal risk as well as a pipeline one.

## How does GEO work for a cybersecurity software company?

It makes your proof easy for AI assistants to find, read and repeat, in the sources they already use. It cannot make any assistant recommend a particular security product.

1. **One clear identity.** Describe what you protect, for whom, and how it deploys, the same way on your site, review profiles, analyst briefings and company databases. Inconsistent descriptions confuse buyers and assistants alike; see [how to fix wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers).
2. **A public, readable trust center.** Put certifications such as SOC 2 and ISO 27001, FedRAMP status, vulnerability advisories and incident history on plain web pages, not only in gated PDFs. That answers the 38% of buyers who could not find what they needed.
3. **Independent validation and coverage.** Take part in independent tests, brief analysts, and publish threat research that security journalists cite. Boards ranked expert commentary after major incidents third among trust drivers; this is where digital PR earns its keep. Our guide on [building authority for AI search](https://underneath.agency/resources/how-brands-build-authority-for-ai-search) covers that PR work in more depth.
4. **Reviews and peer communities.** Encourage detailed customer reviews on platforms buyers use, and take part honestly in practitioner communities. Fake reviews and planted posts are a risk, as we explain in [legitimate GEO versus manipulation](https://underneath.agency/resources/legitimate-geo-vs-manipulation).
5. **Content for real buyer questions.** Publish honest comparison and alternatives pages, compliance mapping pages, integration pages and deployment detail. Ranked lists on other sites matter more than your own pages, as we show in [which pages to target](https://underneath.agency/resources/best-of-lists-ai-recommendations) and [whether comparison pages help](https://underneath.agency/resources/do-comparison-pages-help-b2b-ai-citations).
6. **Tracking that mirrors how CISOs search.** Put the same security questions to ChatGPT, Google AI Overviews and AI Mode, Gemini, Perplexity, Copilot and Claude, several times each; see [how many prompts to track](https://underneath.agency/resources/how-many-prompts-to-track-ai-visibility).

## Which gaps remain in the evidence on AI and security purchases?

Whether being named by assistants wins security contracts: nobody has published a study of that link.

- **No public, vendor-neutral study of CISOs’ AI use.** Figures circulate online about how many security leaders use AI assistants, some attributed to analyst firms, but we could not trace them to a public source, so we do not use them.
- **The surveys have interests and limits.** G2’s data cover all B2B software and G2 sells review visibility. Sophos sells security products. ActualTech Media’s cycle data predate the latest shift to AI search.
- **Revenue effects have the thinnest support.** When [Martinez](https://arxiv.org/abs/2607.14035) reviewed 45 studies of AI search optimization, traffic and conversions had the least evidence behind them, a gap we discuss in [does AI visibility drive business results](https://underneath.agency/resources/does-ai-visibility-drive-business-results).
- **Selection is a black box.** Outside the platforms, nobody knows exactly how assistants pick vendors, and answers change between runs and between assistants.

## How can a security vendor see whether AI answers are feeding its RFPs?

Check which vendors assistants name for the questions CISOs ask before an RFP, then fix the proof gaps.

List the questions your buyers ask at each stage: category, alternatives, comparisons, compliance and due diligence. Put each one to every major assistant more than once, since security answers shift between runs. Note which vendors appear, which pages are cited, and how your certifications, test results and incident history are described. In security, what is missing is usually outside proof such as test results or certifications, not more blog posts.

To have that check run against your enterprise pipeline, [ask us for a security shortlist review](https://underneath.agency/contact). It shows which CISO questions name you, which name rivals instead, and which missing public proof most likely costs you proofs of concept and renewals. Closing those gaps, from an ungated trust center to independent test results and analyst coverage, is the work our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page sets out.

## Frequently asked questions

### Do CISOs really use ChatGPT to research security vendors?

Software buyers in general increasingly do, as the G2 survey described above shows. No public study isolates CISOs, so their share is unknown.

### Will a Gartner Magic Quadrant spot get a security vendor named by ChatGPT?

Plausibly, but untested. Boards and security teams ranked analyst reports fifth among trust drivers, and ChatGPT searched for named rankings in 43.8% of answers in our study.

### Should we ungate our trust center and technical documents?

Usually yes, for the facts buyers verify. 38% of IT and security leaders in the Sophos survey struggled to find the vendor information they needed.

### Are “alternatives to” pages worth writing for security buyers?

Yes, if honest and specific, but expect third-party sources to carry more weight: AI search drew 72.7% of US software sources from earned sites.

### How many quarters before GEO work reaches a signed security contract?

Expect it to be slow. Most security purchases run on 6–12-month cycles, so a new shortlist place may take quarters to become signed revenue.

## Sources

- G2, Tim Sanders (2026-04-15), [In the Answer Economy, Don’t Win the Click — Win the Answer](https://company.g2.com/news/g2-research-the-answer-economy)
- G2, Sydney Sloan (2025), [Proving Value in the Age of AI: 2025 Buyer Behavior Report highlights](https://research.g2.com/cmos-2025-buyer-behavior-report-research-g2)
- Sophos and Vanson Bourne (2026-03), [The Cybersecurity Trust Reality in 2026](https://assets.sophos.com/X24WTUEQ/at/bhmx4pwtc5vz7scc97jjb4c9/sophos-the-cybersecurity-trust-reality-in-2026.pdf)
- Gartner, via CRN Asia (2025-07-29), [Gartner forecasts worldwide end-user spending on information security to total $213 bn in 2025](https://www.crnasia.com/india/news-network/news/gartner-forecasts-worldwide-end-user-spending-on-information-security-to-total-213-bn-in-2025)
- IANS Research and Artico Search (2025-08-05), [2025 Security Budget Benchmark Report](https://www.iansresearch.com/resources/press-releases/detail/ians-research-and-artico-search-release-security-budget-benchmark-report)
- Forrester (2026-01-21), [Forrester’s 2026 Buyer Insights: GenAI Is Upending B2B Buying](https://www.forrester.com/press-newsroom/forrester-2026-the-state-of-business-buying/)
- Forrester (2026-01), [The State Of Business Buying: Risk-Averse Buyers Demand Proof, Not Promises](https://forrester.com/blogs/state-of-business-buying-2026)
- ActualTech Media (2024-01-10), [Cybersecurity Buyers Report Part 3](https://www.actualtechmedia.com/blog/cybersecurity-buyers-report-part-3/)
- SentinelOne, via 01net (2026-03-12), [SentinelOne Announces Fourth Quarter and Fiscal Year 2026 Financial Results](https://www.01net.it/sentinelone-announces-fourth-quarter-and-fiscal-year-2026-financial-results/)
- IBM (2025-07-30), [IBM Report: 13% of Organizations Reported Breaches of AI Models or Applications](https://newsroom.ibm.com/2025-07-30-ibm-report-13-of-organizations-reported-breaches-of-ai-models-or-applications,-97-of-which-reported-lacking-proper-ai-access-controls)
- Expert Insights, Mirren McDade (2026-08-11), [Senior CISOs On How They Make Their Security Investments](https://expertinsights.com/industry-perspectives/how-senior-cisos-actually-vet-security-vendors)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- OpenAI Help Center (2025), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Chen, Wang, Chen and Koudas (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919), arXiv:2509.08919.
- Martinez (2026), [Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026)](https://arxiv.org/abs/2607.14035), arXiv:2607.14035.
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [“Is this brand legit?” How AI assistants build a reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study)
- Underneath (2026), [When does a Reddit thread become evidence in Google’s AI?](https://underneath.agency/research/ai-reddit-citations-study)
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude agree on brands?](https://underneath.agency/research/ai-assistants-brand-agreement-study)
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)

---

This is the Markdown twin of https://underneath.agency/resources/cybersecurity-software-revenue-from-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How do data platforms get shortlisted when data leaders ask AI?"
description: "By being named, accurately, in the AI answers data leaders use to frame architecture choices, then winning the proof of concept on the buyer’s own data."
canonical: "https://underneath.agency/resources/data-infrastructure-enterprise-demand-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How do data platforms get shortlisted when data leaders ask AI?

By being named, and described accurately, in the AI answers data leaders use to frame architecture choices: warehouse or lakehouse, batch or streaming, build or buy. The shortlist that follows is short, and the contract that follows a won proof of concept tends to grow for years. In a market reshaped by $8 billion and $11 billion acquisitions, buyers also ask AI who still does what, and stale answers cost deals.

## The short version

1. The prize compounds: [Snowflake](https://www.01net.it/snowflake-reports-financial-results-for-the-second-quarter-of-fiscal-2027/) reported 828 customers with more than $1 million of trailing product revenue and a net revenue retention rate of 126%, meaning existing customers spent 26% more than a year earlier.
2. Enterprise reach is the battleground: [Databricks](https://www.databricks.com/company/newsroom/press-releases/databricks-raising-series-k-investment-100-billion-valuation) says over 60% of the Fortune 500 use its platform, and [IBM](https://newsroom.ibm.com/2025-12-08-ibm-to-acquire-confluent-to-create-smart-data-platform-for-enterprise-generative-ai) says Confluent serves more than 40% of the Fortune 500.
3. The vendor map is being redrawn: IBM agreed to buy Confluent for $11 billion, [Salesforce](https://www.salesforce.com/news/press-releases/2025/05/27/salesforce-signs-definitive-agreement-to-acquire-informatica/) agreed to buy Informatica for about $8 billion, and [Fivetran and dbt Labs](https://www.fivetran.com/press/fivetran-dbt-labs-complete-merger-to-create-the-data-infrastructure-for-trusted-ai-agents) merged in 2026.
4. Data practitioners already work inside AI tools: in [dbt Labs’ 2025 survey](https://www.getdbt.com/resources/reports/state-of-analytics-engineering-2025) of 459 data professionals, 80% used AI in their daily work, up from 30% a year earlier.
5. The buyer’s pain is specific: poor data quality was the top challenge, cited by over 56% of the same respondents, and data quality tools were the second-largest area of planned new investment.

## Who chooses a data platform, and how much can one customer spend?

A chief data officer or head of data platform leads, with architects, security, finance and often a systems integrator.

Data infrastructure covers the pipes and platforms that move, store, transform and serve data: ingestion and integration tools, warehouses and lakehouses, streaming platforms, transformation, quality and governance. The buyer is usually a data leader with an architecture problem, not a developer picking a tool for one app. That is the main difference from choosing a database for one application: here the decision is the shape of the whole data estate. The cloud underneath is a separate choice, covered in [how cloud providers win enterprise deals through AI](https://underneath.agency/resources/cloud-infrastructure-enterprise-demand-ai-search).

The money is in expansion. Most platforms charge for consumption, so revenue grows as the customer moves more workloads onto them. Snowflake’s second quarter of fiscal 2027 shows the pattern:

- Product revenue of $1.49 billion, up 37% year on year.
- 828 customers with trailing 12-month product revenue above $1 million, up 27%.
- 829 Forbes Global 2000 customers, and $9.00 billion of remaining performance obligations.
- 692 net new customers in the quarter.

Databricks, which is private, says more than 15,000 customers use its platform and valued itself at over $100 billion in its 2025 funding round. Confluent, the streaming company, had more than 6,500 clients when IBM agreed to buy it. Fivetran and dbt Labs say their combined business supports more than 100,000 data teams.

Budgets are moving in the right direction for sellers. The dbt Labs survey found that 30% of respondents reported budget increases for their data teams, against 9% the year before, while 40% reported headcount increases.

## How far into data platform decisions has AI already reached?

In the daily work of the people who shape the decision, and in buyers’ research before vendor calls.

The people who evaluate data platforms use AI constantly. In the dbt Labs survey, 80% of data practitioners used AI in their daily workflow, and 70% used it for analytics development, mainly through general-purpose assistants such as ChatGPT, Claude and Gemini. The same tools they use to write code are a short step from the tools they ask about architecture. The same practitioners weigh in on reporting tools, covered in [how analytics and BI tools get shortlisted](https://underneath.agency/resources/analytics-bi-software-ai-search).

Technology buyers in general use AI in research and check it. In [TrustRadius’s 2026 survey](https://www.demandgenreport.com/industry-news/news-brief/trustradius-ai-has-changed-how-buyers-research-but-not-what-they-trust/53853/) of nearly 2,500 technology buyers and vendors, 63% of buyers used AI during their purchase journey, 94% of those fact-checked its answers at least some of the time, and 83% shortlisted three or fewer products. Because TrustRadius sells review visibility, its numbers are best read as one input among several. The finding matters for data infrastructure because a shortlist of three is a small door for any platform.

## Which questions do data leaders ask AI assistants?

Architecture, comparison, migration, cost and vendor-risk questions. These prompts are our own illustrations of how data leaders phrase architecture questions; we did not collect them from real buyers.

| Decision | Illustrative prompt |
|---|---|
| Architecture | “Should a 2,000-person insurer build a lakehouse on open table formats or stay with a cloud data warehouse?” |
| Streaming | “Do we need Kafka for real-time fraud scoring, or can our warehouse handle it?” |
| Integration | “Which data integration tools have reliable connectors for SAP and Salesforce with change data capture?” |
| Cost | “How do Snowflake and Databricks compare on cost for 50 TB of mostly SQL analytics?” |
| AI readiness | “What data stack do we need before we put AI agents on our customer data?” |
| Vendor risk | “What changes for Informatica customers after the Salesforce acquisition?” |
| Governance | “Which catalogs support column-level lineage across dbt, Spark and our warehouse?” |

Two features of these questions favor some vendors over others, we infer. First, they are framed around architectures, not products, so an answer usually describes a pattern and then names the vendors that fit it. A vendor that is not tied to the pattern in public writing is easy to leave out. Second, the vendor map changes fast: an assistant answering from older training data may still describe a merged or acquired company as independent, as we explain in [do AI assistants answer from training data](https://underneath.agency/resources/do-ai-assistants-answer-from-training-data).

## How does an AI answer turn into consumption revenue for a data platform?

Through the architecture decision: the AI answer frames the pattern, the shortlist follows, the proof of concept decides.

1. A data leader or architect asks an assistant how to solve a problem: real-time data, AI readiness, a migration.
2. The answer describes an architecture and names a few vendors that fit it.
3. The team checks those vendors with peers, analysts, partners and documentation.
4. Two or three run a proof of concept on the buyer’s own data, often with a systems integrator.
5. The winner signs a consumption commitment and then grows as workloads move.

The AI answer matters most at step 2, where the frame is set. A vendor framed as “the streaming option” or “the cheaper warehouse” enters the evaluation with that label. Step 5 is where the value is: a net revenue retention rate of 126%, as Snowflake reported, means each year’s customers keep spending more. Losing step 2, by this logic, costs a stream of revenue rather than one deal. That is our inference from the published figures, not a measured link.

Tracing a consumption contract back to an AI answer is hard. A buyer who met you in an AI answer may arrive through a partner, a cloud marketplace or an inbound demo request months later. Our article on [what lost clicks mean for pipeline](https://underneath.agency/resources/ai-answers-pipeline-revenue) looks at how to account for that delay.

## Why does an assistant name one warehouse, lakehouse or pipeline vendor and not another?

No assistant publishes its vendor-selection rules; studies point to independent coverage, and data teams test claims on their own data.

**Documented by the platforms.** Google documents that AI Overviews and AI Mode [may run several related searches](https://developers.google.com/search/docs/appearance/ai-features) for one question, and OpenAI documents that [ChatGPT search rewrites questions](https://help.openai.com/en/articles/9237897-chatgpt-search) into targeted queries. Neither explains how a warehouse or integration vendor earns a place in the answer.

**Observed in studies.** In [our four-assistant study](https://underneath.agency/research/ai-assistants-brand-agreement-study), assistants agreed on brands more for B2B software than for any other industry, with an overlap of 0.543 on a scale from 0 to 1. That still leaves plenty of disagreement: a vendor named by one assistant may be missing from another. Of the signals in [our brand study](https://underneath.agency/research/brand-entity-ai-recommendations-study), independent coverage mattered most, with each tenfold rise in independent sites naming a brand going with 4.7 times the odds of a recommendation.

**What data buyers check.** Data leaders verify with tests on their own data, reference customers, partner opinions and documentation. TrustRadius found analyst reports were used by only 13% of buyers to make purchase decisions, a 63% fall since 2022, while 74% of buyers used reviews.

**Our inference.** Data platforms are judged on specifics: connector coverage, supported formats, latency, governance features, pricing units and where data stays. Those are facts an assistant can repeat only if they are public and consistent. A reasonable expectation is that vendors with clear, current, independently confirmed specifics are described more accurately and named more often for the right architectures. That idea has not yet been tested on warehouse, streaming or integration vendors.

## What slips away when a data platform is absent from architecture answers?

Lost architecture decisions, which are rarely reopened, though no study has measured the cost directly.

- **Platforms are sticky.** Once pipelines, models and governance are built on a platform, switching means rebuilding them. A vendor left out of an architecture decision may wait years for the next one, we infer.
- **Expansion goes to the incumbent.** With consumption pricing and retention above 100%, the vendor that wins the first workload usually wins the next ones too.
- **Consolidation reshuffles the map.** After the Confluent, Informatica and dbt deals, buyers will ask who now owns what and what changes for them. An assistant that answers with outdated information can send buyers elsewhere.
- **AI projects raise the stakes.** The dbt Labs survey found 45% of respondents planned to increase investment in AI tooling, the largest category. Data platforms that are described as “ready for AI” in answers stand to gain from that spending, our inference. Data science platforms chase it too; see [how data science platforms win enterprise buyers](https://underneath.agency/resources/data-science-platforms-ai-search).

## What does GEO look like for a warehouse, streaming or integration vendor?

It connects your platform to the data architectures buyers ask about, through sources assistants already rely on. Whether an assistant then names you stays outside anyone’s control.

1. **One clear position per architecture.** Say plainly which patterns you fit (streaming, lakehouse, integration, quality) and which you do not, the same way on your site, docs, partner pages and marketplace listings.
2. **Public specifics.** Publish connector lists, supported formats, limits, security attestations and pricing units on readable pages. If those pages are hard for machines to load, read [when AI agents can’t read your site](https://underneath.agency/resources/when-ai-agents-cant-read-your-site).
3. **Independent validation.** Reference architectures written by cloud providers and integrators, customer talks at practitioner conferences, and independent benchmarks with published methods give assistants sources other than you. For data vendors, this is the practical side of [how brands build authority for AI search](https://underneath.agency/resources/how-brands-build-authority-for-ai-search).
4. **Up-to-date corporate facts.** After a merger, rename or acquisition, update every profile, partner listing and knowledge source quickly. When an answer still shows a pre-merger owner or an old product name, [how to fix wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) sets out the steps.
5. **Architecture content, not only product pages.** Write honest guides to the decisions buyers face, including when a competitor’s approach fits better. Roundups of warehouses, pipelines and catalogs published by others also count, as our review of [best-of lists](https://underneath.agency/resources/best-of-lists-ai-recommendations) shows.
6. **Tracking by architecture question.** Record how ChatGPT, Gemini, Perplexity, Claude, Copilot and Google’s AI features answer your lakehouse, streaming, migration and vendor-risk questions, asking each one repeatedly.

## What is still unknown about AI’s role in data platform decisions?

No public study shows whether AI answers change which data platform wins the architecture decision.

- **No data-leader survey on AI research.** The dbt Labs survey measures AI use at work, not in vendor research; the TrustRadius figures cover technology buyers in general. The broader software-buying picture is in [how B2B SaaS companies generate revenue from AI search](https://underneath.agency/resources/b2b-saas-revenue-from-ai-search).
- **Company figures are self-reported.** Customer counts and Fortune 500 shares come from the companies’ own releases.
- **How assistants handle mergers is untested.** We found no study of how quickly AI answers reflect acquisitions.
- **Revenue effects are the least proven part.** The cross-industry evidence is reviewed in [does AI visibility drive business results](https://underneath.agency/resources/does-ai-visibility-drive-business-results).

## Which architecture questions should a data platform test before its next proof of concept?

The warehouse, streaming, migration and vendor-risk questions behind your proofs of concept, and how answers frame you.

List the decisions that lead to your proofs of concept: the architectures, migrations, integrations and vendor-risk questions. Ask each one in several assistants, more than once, because a single answer is a thin sample. Record whether you are named, how you are labeled, which sources are cited, and whether your connectors, pricing and ownership are described correctly. In data infrastructure, the gaps usually trace back to thin independent proof or to ownership and product facts that predate a merger.

For a second pair of eyes, [talk to us about a data platform visibility review](https://underneath.agency/contact). We will map how assistants label your platform in the architecture choices that come before proofs of concept and consumption commitments, and flag the gaps most likely to be costing you enterprise pipeline. Tying your platform to the right architectures, publishing connector and pricing specifics and keeping post-merger facts current are covered on our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page.

## Frequently asked questions

### Do data engineers use AI to research vendors?

They use AI daily for work: 80% of data practitioners in dbt Labs’ survey. No public study measures their vendor research separately.

### Does an acquisition hurt a data vendor’s AI visibility?

It can confuse answers. Assistants that answer from older training data may describe the company as it was before the deal.

### Are analyst reports still worth the effort?

Less than before for buyers directly: TrustRadius found 13% of buyers used them to decide, down 63% since 2022.

### Should we publish our connector list and pricing units?

Yes. Integration and cost are among the first questions data leaders ask, and assistants can only repeat facts they can read.

### How long does it take for GEO to show in data platform pipeline?

Expect quarters, not weeks. Architecture decisions are slow, and proofs of concept add months before a commitment is signed.

## Sources

- Snowflake, via 01net (2026-09-02), [Snowflake Reports Financial Results for the Second Quarter of Fiscal 2027](https://www.01net.it/snowflake-reports-financial-results-for-the-second-quarter-of-fiscal-2027/)
- Databricks (2025-08-19), [Databricks is raising a Series K Investment at >$100 billion valuation](https://www.databricks.com/company/newsroom/press-releases/databricks-raising-series-k-investment-100-billion-valuation)
- IBM (2025-12-08), [IBM to Acquire Confluent to Create Smart Data Platform for Enterprise Generative AI](https://newsroom.ibm.com/2025-12-08-ibm-to-acquire-confluent-to-create-smart-data-platform-for-enterprise-generative-ai)
- Salesforce (2025-05-27), [Salesforce Signs Definitive Agreement to Acquire Informatica](https://www.salesforce.com/news/press-releases/2025/05/27/salesforce-signs-definitive-agreement-to-acquire-informatica/)
- Fivetran (2026-06-01), [Fivetran + dbt Labs Complete Merger to Create the Data Infrastructure for Trusted AI Agents](https://www.fivetran.com/press/fivetran-dbt-labs-complete-merger-to-create-the-data-infrastructure-for-trusted-ai-agents)
- dbt Labs (2025), [2025 State of Analytics Engineering Report](https://www.getdbt.com/resources/reports/state-of-analytics-engineering-2025)
- TrustRadius, via Demand Gen Report (2026), [TrustRadius: AI Has Changed How Buyers Research, But Not What They Trust](https://www.demandgenreport.com/industry-news/news-brief/trustradius-ai-has-changed-how-buyers-research-but-not-what-they-trust/53853/)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- OpenAI Help Center (2025), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude agree on brands?](https://underneath.agency/research/ai-assistants-brand-agreement-study)
- Underneath (2026), [Do Wikipedia and schema make AI assistants recommend a brand?](https://underneath.agency/research/brand-entity-ai-recommendations-study)

---

This is the Markdown twin of https://underneath.agency/resources/data-infrastructure-enterprise-demand-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How data science platforms win buyers who research with AI"
description: "By being named and described accurately in the AI answers that frame platform comparisons, governance checks and cost estimates, then winning the pilot."
canonical: "https://underneath.agency/resources/data-science-platforms-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How do data science platforms win enterprise buyers who research with AI?

By being named, and described accurately, in the AI answers that now frame platform comparisons, governance checks and cost estimates, then winning the proof of concept on the buyer’s own data. A data science platform is a large, sticky purchase that grows with every workload a customer adds. The evaluation is technical and slow, but the shortlist forms early, increasingly in assistants and Google’s AI answers that draw heavily on practitioner communities and independent comparisons.

## The short version

1. The category is large and still accelerating: Databricks announced in August 2026 that it had crossed a $7 billion revenue run-rate with more than 80% year-over-year growth, and [Sacra](https://sacra.com/c/databricks/) reports its net dollar retention remained above 140%.
2. Incumbents are hard to dislodge: according to [Menlo Ventures](https://menlovc.com/perspective/2025-the-state-of-generative-ai-in-the-enterprise/), incumbents hold 56% of the AI infrastructure market as builders keep working on the data platforms they have trusted for years.
3. Governance is the gate: [Gartner](https://www.gartner.com/en/newsroom/press-releases/2025-02-26-lack-of-ai-ready-data-puts-ai-projects-at-risk) predicts that through 2026, organizations will abandon 60% of AI projects unsupported by AI-ready data, and [Anaconda’s survey](https://www.anaconda.com/state-of-data-science-report-2024) of more than 3,000 practitioners found 42% of organizations cite security as their main AI challenge.
4. Google’s AI answers are almost always present for this kind of search: in [our study](https://underneath.agency/research/ai-overviews-frequency-study), B2B software and technology keywords showed an AI Overview on 96.0% of searches.
5. Practitioner communities feed those answers: Google’s AI Overviews cited Reddit on [35.4% of B2B software searches](https://underneath.agency/research/ai-reddit-citations-study) we tested, mostly threads where practitioners discuss their own work.

## Who chooses a data science platform, and how much can one account spend?

A platform or data leader buys it for many users, after practitioners have tested it, and revenue grows with usage.

The buying group is mixed. A chief data officer, chief information officer or head of data platform signs the contract; data scientists, machine learning engineers and data engineers decide whether it works; security, finance and procurement decide whether it can be bought. Each group checks different things: practitioners check notebooks, languages and libraries, platform teams check scale and integration, and security checks governance and access control.

AI has raised the stakes. In [IBM’s survey of 2,000 CEOs](https://newsroom.ibm.com/2025-05-06-ibm-study-ceos-double-down-on-ai-while-navigating-enterprise-hurdles), 68% said an integrated enterprise-wide data architecture is critical for cross-functional collaboration and 72% saw their proprietary data as key to unlocking the value of generative AI. Yet 50% said rapid investment had left them with disconnected technology. That is the problem a data science platform claims to solve, and it is why governance and integration questions now sit next to features in every evaluation.

Account value on a data science platform rises and falls with consumption. Most leading platforms charge for compute, storage and processing rather than seats, so revenue grows as customers move more workloads onto them. Databricks, the clearest public example, reported reaching a $5.4 billion revenue run-rate in early 2026, with more than $1.4 billion from its AI products, before crossing $7 billion by August. Sacra’s estimate of net dollar retention above 140% means the average existing customer spent far more than it did a year earlier. Winning the first workload is the start of a multi-year expansion.

Buyers are also consolidating. In [G2’s 2026 buyer behavior report](https://sell.g2.com/2026-buyer-behavior-report), a survey of software buyers across categories, 84% had consolidated at least three best-of-breed tools into all-in-one platforms in the past year. For data science, we infer that means fewer, larger platform decisions, each worth more to the winner. The warehouses and pipelines underneath face similar choices; see [how data platforms get shortlisted by AI](https://underneath.agency/resources/data-infrastructure-enterprise-demand-ai-search).

## Where do data scientists and platform leaders run into AI answers?

In the practitioner’s daily work, in early research, and in the Google answers that shape comparisons.

Practitioners use AI constantly. Anaconda found 87% of practitioners were increasing AI adoption, in work such as data cleaning, task automation and predictive modeling. In the [2025 Stack Overflow Developer Survey](https://stackoverflow.co/company/press/archive/stack-overflow-2025-developer-survey/), 84% of respondents used or planned to use AI tools in their development process. A data scientist who asks an assistant how to deploy a model or track experiments is, in effect, asking which platform to use. Reporting tools meet the same kind of question, as our guide to [analytics and BI shortlists in AI](https://underneath.agency/resources/analytics-bi-software-ai-search) shows.

Google’s AI answers sit on top of the comparison searches. With AI Overviews on 96.0% of B2B software and technology searches in our sample, a platform evaluator searching “feature store options” or a vendor comparison will usually read Google’s summary before any vendor page. According to Google, AI Mode relies on a [“query fan-out” technique](https://blog.google/products/search/ai-mode-search/) that runs multiple related searches across subtopics and merges the results, so one evaluator’s question can pull in pricing, governance and review pages together.

Machines are also reading the documentation. Mintlify, a documentation platform, reports that agents made 37.9% of requests to the “dev infrastructure and data” documentation sites it hosts in August 2026, up from 15.5% in February, according to its [own traffic data](https://mintlify.com/data.md/).

## Which questions do data science platform buyers ask AI assistants?

Comparison, fit, governance, cost and migration questions, asked by practitioners and platform leaders in different words.

Practitioners and platform leaders word these questions differently; the examples below are ours, not observed queries:

- Comparison: “Databricks vs Snowflake for machine learning workloads at a 200-person analytics team.”
- Fit: “Best data science platform for a regulated bank that must keep data on-premises.”
- Alternatives: “Open-source alternatives to a managed notebook platform for a small team.”
- Governance: “Which platforms offer column-level lineage, access control and audit logs for models?”
- Cost: “How much does a data science platform cost per year for 50 data scientists?”
- Migration: “How hard is it to move our notebooks and pipelines from one platform to another?”

Comparison and fit questions decide the shortlist. Governance and cost questions decide who survives the security review and the finance review. Cost questions are especially risky for consumption-priced platforms: in [our pricing study](https://underneath.agency/research/ai-pricing-accuracy-study), only 61.9% of plan prices four assistants quoted for 45 software products were fully faithful to the official page, and usage-based pricing is harder to summarize than a seat price.

## How does an AI mention become a consumption contract for a data science platform?

Through the proof of concept: the platform is named, practitioners test it, a pilot runs, and usage grows.

**Practitioner adoption.** Platforms often land through the tools practitioners already use. GitHub’s [Octoverse 2025](https://github.blog/news-insights/octoverse/octoverse-a-new-developer-joins-github-every-second-as-ai-leads-typescript-to-1/) counts 2.4 million repositories using notebooks, up 75% in a year, and reports that Python remains dominant in AI and data science, with 2.6 million contributors. A practitioner who meets a platform in an assistant’s answer about a notebook or deployment problem may try it before anyone opens a procurement ticket.

**Shortlist and proof of concept.** Platform leaders then compare a few candidates and run a pilot on their own data. We infer that an assistant’s answer to a comparison or governance question shapes which candidates get a pilot slot.

**Contract and expansion.** Consumption contracts grow as more teams and workloads move onto the platform. That is where Sacra’s net dollar retention estimate of above 140% for Databricks comes from: the first workload brings in the next.

## Why do assistants name some data science platforms more readily than others?

No platform documents it; studies point to independent sources, practitioner discussion and established brands.

**Documented by the platform.** Apart from Google’s description of query fan-out, no AI assistant documents how it chooses which software vendors to name.

**What studies show.** [Chen and colleagues](https://arxiv.org/abs/2509.08919) found independent “earned” sites supplied 72.7% of AI search’s sources across the US categories they tested, compared with 45.4% of Google’s. In our Reddit research, half of Google’s Reddit citations came from communities of practitioners answering questions about their own work. Software answers are also unusually stable: in [our four-assistant study](https://underneath.agency/research/ai-assistants-brand-agreement-study), B2B software had the highest agreement between assistants of eight industries (an overlap score of 0.543 on a 0-to-1 scale), and in [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study) its recommended brands were the most stable across repeated runs (0.708).

We infer that this stability favors established platforms, which matches Menlo’s observation that incumbents hold most of the AI infrastructure market. Menlo, which invests in Databricks, notes that newer infrastructure companies are also growing fast, so the door is not closed. Firms that build machine learning solutions for clients face a different question, covered in [how machine learning companies get found](https://underneath.agency/resources/machine-learning-companies-ai-search).

**Trust factors specific to data science platforms.** A reasonable expectation is that assistants and evaluators look for the same evidence:

- Public, detailed governance and security documentation: lineage, access control, audit, certifications and data residency.
- Honest, reproducible performance and cost comparisons, with methods and dates.
- Integration lists for the clouds, data stores, languages and libraries buyers already run.
- Practitioner discussion and independent reviews that describe real deployments, including limitations.

## What does it cost a data science platform to be left out?

A multi-year consumption contract, because platform choices are made rarely and changed even more rarely.

When buyers consolidate tools into fewer platforms, each decision carries more revenue and happens less often. A platform that is missing from the comparison answer in the year a buyer consolidates may wait years for the next chance. The same stability that protects incumbents works against challengers: if assistants name the same few platforms run after run, a newer platform has to give them a specific reason to name it, such as a clear fit for a regulated industry, a deployment model or a price point.

Misdescription is a cost too. A platform that is named but described with an outdated price, a missing certification or a retired feature can be cut at the governance or finance review. Our article on [fixing wrong brand information](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) covers how to trace and correct those errors.

## How does GEO work for a data science platform?

It makes your platform easy to find, compare and verify in AI answers; it cannot promise a shortlist place.

Six workstreams make up generative engine optimization (GEO) for a data science platform:

1. **Consistent product facts.** Describe capabilities, deployment options, supported clouds and certifications the same way on your site, documentation, cloud marketplaces and review profiles.
2. **Governance and security in public.** Publish ungated pages on lineage, access control, audit, residency and compliance, since these answer the questions that decide the security review.
3. **Transparent cost information.** Explain how consumption pricing works, with worked examples for typical team sizes and workloads, so assistants and finance teams have accurate figures to quote.
4. **Fair comparison and migration content.** Publish honest comparisons and migration guides that name trade-offs. Whether such pages earn AI citations is covered in [our article on comparison pages](https://underneath.agency/resources/do-comparison-pages-help-b2b-ai-citations).
5. **Practitioner presence.** Support the communities where data scientists and engineers discuss tools, such as GitHub, Stack Overflow, Reddit and conference talks, with accurate answers and real examples rather than promotion. Our article on [Reddit and Google’s AI Overviews](https://underneath.agency/resources/does-reddit-shape-google-ai-overviews) explains what those citations do and do not mean.
6. **Measurement by buyer and stage.** Track practitioner, comparison, governance and cost questions separately across ChatGPT, Gemini, Perplexity, Copilot and Google’s AI features, repeated over time, and compare with trials, proofs of concept and expansion.

## Where is the evidence thin on AI search and data science platform choices?

Two things remain unmeasured: evaluators’ use of assistants, and whether AI visibility changes which platform wins.

No public survey isolates data science platform buyers. The buyer evidence here comes from practitioner and CEO surveys, a cross-category software survey and our own studies of AI answers for B2B software in general. Databricks’ figures are company announcements and Sacra’s estimates, and Menlo Ventures invests in several of the companies it describes. No published study yet follows a platform evaluation from an AI answer to a proof of concept and a consumption contract. Treat any claimed link between AI visibility and platform revenue as something to test.

## How can a data science platform see whether AI answers win it proof-of-concept slots?

Check how assistants compare and describe your platform when evaluators ask about fit, governance and cost.

The first piece of work is an audit of comparison, fit, governance and cost questions across the main assistants and Google’s AI features, checking both whether you are named and whether the facts are right, and matching the gaps against your pipeline of trials and proofs of concept. We can [run that platform audit with your team](https://underneath.agency/contact) and point to where better public facts could earn more evaluation slots. For how those facts get published and checked over time, including governance pages and worked consumption-cost examples, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## Frequently asked questions

### Do AI assistants compare data science platforms accurately?

Not always. In our pricing study, about six in ten software prices quoted by assistants were fully faithful to the official page, and consumption pricing is harder to quote. Check the governance and cost facts assistants give about you, not just whether you are named.

### Does Reddit matter for data science platform visibility in AI search?

It matters on Google. Its AI Overviews cited Reddit on about a third of B2B software searches we tested, mostly practitioner communities. How much those threads shape each answer is less clear.

### Can a newer data science platform compete with incumbents in AI answers?

Yes, but it needs a specific, verifiable reason to be named, such as a deployment model, a regulated-industry fit or a cost advantage, documented on independent sources and not only on its own site.

### How should a platform measure the impact of AI search?

Track a fixed set of buyer questions by stage, ask new trial users and evaluation teams how they found you, and follow those accounts through proofs of concept to consumption and expansion.

## Sources

- Sacra (2026), [Databricks revenue, valuation and funding](https://sacra.com/c/databricks/), including Databricks’ August 2026 announcement
- Menlo Ventures (2025), [2025: The State of Generative AI in the Enterprise](https://menlovc.com/perspective/2025-the-state-of-generative-ai-in-the-enterprise/)
- Gartner (2025), [Lack of AI-Ready Data Puts AI Projects at Risk](https://www.gartner.com/en/newsroom/press-releases/2025-02-26-lack-of-ai-ready-data-puts-ai-projects-at-risk)
- Anaconda (2024), [State of Data Science 2024 Report](https://www.anaconda.com/state-of-data-science-report-2024)
- IBM (2025), [IBM Study: CEOs Double Down on AI While Navigating Enterprise Hurdles](https://newsroom.ibm.com/2025-05-06-ibm-study-ceos-double-down-on-ai-while-navigating-enterprise-hurdles)
- G2 (2026), [2026 Buyer Behavior Report](https://sell.g2.com/2026-buyer-behavior-report)
- Stack Overflow (2025), [Stack Overflow’s 2025 Developer Survey Reveals Trust in AI at an All Time Low](https://stackoverflow.co/company/press/archive/stack-overflow-2025-developer-survey/)
- GitHub (2025), [Octoverse: A new developer joins GitHub every second as AI leads TypeScript to #1](https://github.blog/news-insights/octoverse/octoverse-a-new-developer-joins-github-every-second-as-ai-leads-typescript-to-1/)
- Google (2025), [Expanding AI Overviews and introducing AI Mode](https://blog.google/products/search/ai-mode-search/)
- Mintlify (2026), [Data: agents vs human traffic](https://mintlify.com/data.md/)
- Chen and colleagues (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919)
- Underneath (2026), [When does Google show an AI Overview?](https://underneath.agency/research/ai-overviews-frequency-study), [Reddit citations](https://underneath.agency/research/ai-reddit-citations-study), [software pricing accuracy](https://underneath.agency/research/ai-pricing-accuracy-study), [four-assistant agreement](https://underneath.agency/research/ai-assistants-brand-agreement-study) and [recommendation consistency](https://underneath.agency/research/ai-recommendation-consistency-study)

---

This is the Markdown twin of https://underneath.agency/resources/data-science-platforms-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How database companies get chosen when AI helps pick the stack"
description: "By being the database AI coding tools and assistants reach for by default, then backing it with docs, benchmarks and proof enterprise architects can check."
canonical: "https://underneath.agency/resources/database-companies-developers-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How does a database company get chosen when AI helps pick the stack?

By becoming the database that AI coding tools set up by default and that AI assistants name for specific workloads, then giving enterprise architects proof they can check. The first database decision is increasingly made inside an AI tool, before anyone compares vendors. The large contracts still come later, through an architecture review that rewards evidence over popularity.

## The short version

1. At [Supabase](https://supabase.com/blog/supabase-series-f), more than 60% of new databases are now launched by some sort of AI tool, and database launches grew 600% in a year.
2. At Neon, over 80% of databases were created automatically by AI agents rather than people, [Databricks said](https://www.databricks.com/company/newsroom/press-releases/databricks-agrees-acquire-neon-help-developers-deliver-ai-systems) when it agreed to buy the company in 2025.
3. Developers lean on AI but check it: in [Stack Overflow’s 2025 survey](https://stackoverflow.co/company/press/archive/stack-overflow-2025-developer-survey/) of more than 49,000 developers, 84% used or planned to use AI tools, yet 46% did not trust the accuracy of their output.
4. The enterprise prize is large: [MongoDB](https://webassets.mongodb.com/_com_assets/cms/mtj3tpvjctfbi5x1q-MongoDB-Inc-Announces-Second-Quarter-Fiscal-2027-Financial-Results.pdf) reported 2,999 customers paying $100,000 or more a year, and says about 75% of the Fortune 100 rely on it.
5. PostgreSQL is the default many tools reach for: for the third year in a row it led Stack Overflow’s ranking of databases developers want to use (47%) and want to keep using (66%).

## Who chooses a database, and what is a customer worth?

Developers usually make the first choice; architects and procurement turn it into an enterprise standard and a contract.

A database rarely starts with a purchase order. A developer picks one for a prototype, the prototype goes into production, and the choice hardens as data piles up. Moving data later is expensive and risky, so the first pick tends to stay. That is why database companies spend so much effort on free tiers, documentation and developer communities. DevOps tools grow the same way, as our guide to [how DevOps tools get recommended to engineers](https://underneath.agency/resources/devops-platforms-ai-search) shows.

The enterprise stage looks different. An architecture or platform team reviews the database for security, compliance, support, cost at scale and fit with the rest of the stack. Procurement then signs a multi-year subscription or a cloud commitment. Security tools sold to developers face a similar double review, covered in [how application security vendors win demand](https://underneath.agency/resources/application-security-demand-ai-search).

Public filings show what that second stage is worth. MongoDB reported revenue of $771.8 million for its second quarter of fiscal 2027, up 30%, from more than 70,600 customers. Its remaining performance obligations, contracted revenue not yet recognized, reached $1,519.2 million, up 91% year on year. Its count of customers paying $100,000 or more a year has climbed every quarter in the table it publishes.

Private companies show what the developer stage is worth. Supabase raised $500 million at a $10 billion pre-money valuation in June 2026, with nearly 10 million developers building on it. Large data platforms are buying their way into the same market: Databricks agreed to acquire Neon, and Snowflake announced the acquisition of [Crunchy Data](https://www.snowflake.com/en/news/press-releases/snowflake-acquires-crunchy-data-to-bring-enterprise-ready-postgres-offering-to-the-ai-data-cloud/), whose press release put PostgreSQL use at 49% of all developers and called it “a massive $350 billion market opportunity.” That figure is Snowflake’s own estimate.

## How has AI changed the way databases get picked?

AI coding tools now create many new databases, so an agent’s default choice has become a sales channel.

The clearest evidence comes from the vendors themselves:

- **Neon.** Databricks said internal telemetry showed over 80% of databases provisioned on Neon were created by AI agents rather than humans.
- **Supabase.** Its CEO wrote in June 2026 that more than 60% of new databases are launched by an AI tool, and that growth accelerated as Claude Code and Codex expanded the number of people who can build.
- **Integrations.** Supabase became an [official ChatGPT app](https://supabase.com/blog/supabase-is-now-an-official-chatgpt-app) with 29 tools for running queries and managing projects. MongoDB launched a hosted service that connects coding agents such as Claude Code and Codex to live Atlas data.

These are company-reported figures from firms with an interest in the trend, and they describe developer-led products more than enterprise estates. Still, the direction is plain: when a person asks an AI tool to build an app, the tool usually picks the database too. The same goes for other parts of the stack, as our guide to [how developer tools win AI-first users](https://underneath.agency/resources/developer-tools-ai-search) explains.

Developers themselves use AI heavily but cautiously. In Stack Overflow’s 2025 survey, 84% used or planned to use AI tools, up from 76% in 2024. Only 31% used AI agents at the time, and 35% visited Stack Overflow after running into problems with AI answers. Human communities remain where developers go to check.

## Which questions do developers and architects ask AI about databases?

Workload, comparison, compatibility, cost and migration questions. We wrote the sample prompts below to show how developers and architects tend to phrase database questions; none were taken from real query logs.

| Who asks | Illustrative prompt |
|---|---|
| Developer, new app | “What database should I use for a multi-tenant SaaS app with user logins and file storage?” |
| Developer, AI feature | “Can I do vector search in Postgres, or do I need a separate vector database?” |
| Comparison | “MongoDB vs PostgreSQL with JSON columns for a product catalog that changes often” |
| Architect, scale | “Which distributed SQL databases handle writes across three regions with strong consistency?” |
| Cost | “How much does a managed Postgres database cost at 2 TB with read replicas?” |
| Migration | “How hard is it to move from Oracle to PostgreSQL, and which tools help?” |
| Risk | “Which databases changed their open source license, and what does that mean for us?” |

Two kinds of answer are at stake. A developer’s question often ends in a setup command, so being named means being installed. An architect’s question ends in a shortlist for a proof of concept. Pricing answers are fragile for usage-priced products: in [our pricing study](https://underneath.agency/research/ai-pricing-accuracy-study), only 61.9% of plan prices quoted by four assistants for 45 software products were fully faithful to the vendor’s page.

## How does an agent’s database pick turn into usage revenue and contracts?

Through a long funnel: an AI tool’s default becomes a free project, then production use, then a contract.

1. A developer or a coding agent needs a database and asks for one, or the agent picks one itself.
2. The tool names or installs a database, usually one it has seen in documentation, templates and examples.
3. The project grows on a free or usage-priced tier.
4. If the app succeeds, usage and the bill grow with it.
5. The company standardizes, and procurement signs a larger commitment with support, security and compliance terms.

MongoDB’s customer table shows the shape of the later steps: more than 70,600 customers, of which 2,999 pay $100,000 or more a year. AI visibility, on this reading, matters most at steps 1 and 2, where volume is decided, and again when architects research alternatives for a large workload. That last point is our inference, not a measured link. The clouds those workloads run on face their own shortlist, covered in [how cloud platforms win enterprise buyers](https://underneath.agency/resources/cloud-infrastructure-enterprise-demand-ai-search).

The revenue rarely arrives with a referral tag. A database chosen by an agent in March may show up as an enterprise deal a year later. We cover the attribution problem in [what lost clicks mean for pipeline](https://underneath.agency/resources/ai-answers-pipeline-revenue).

## What decides whether an AI assistant recommends a database?

No platform documents how it picks databases; studies and developer behavior point to public, checkable sources.

**Documented by the platforms.** Google says its AI features may issue several related searches for one question, and OpenAI says ChatGPT search rewrites a question into targeted queries. Neither says how a database gets recommended. Coding agents have no published selection rules either.

**Observed in studies.** Software is where assistants agree most. In [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), B2B software had the most stable brand lists across repeated runs, with a mean overlap of 0.708. Stable lists are good news for established names and hard for newcomers to break into. Developer communities also matter as sources: in [our Reddit study](https://underneath.agency/research/ai-reddit-citations-study), 35.4% of Google’s AI Overviews for B2B software and technology searches cited at least one Reddit thread.

**What developers trust.** Developers check what AI tells them. Stack Overflow found 46% distrusted the accuracy of AI output, up from 31% a year earlier. Documentation, community threads and benchmarks they can rerun carry more weight than claims.

**Our inference.** An assistant or agent can only recommend a database it has seen used. Clear setup guides, framework templates, code examples, honest comparison pages and independent benchmarks are the raw material it works from. A reasonable expectation is that databases with that material in public, current and consistent places get named more often. No study has yet tested this for databases. Training data also lags: an assistant answering from memory may describe last year’s product, as we explain in [why ChatGPT misses new products](https://underneath.agency/resources/why-chatgpt-misses-new-products).

## What does it cost a database company to be missing?

Mostly lost defaults, which compound, though no study has put a number on the loss.

- **Defaults compound.** If more than 60% of new databases on a platform are launched by AI tools, a database that agents do not reach for misses a growing share of new projects at the moment of choice, we infer.
- **Switching is rare.** Once data and code depend on a database, moving is costly. A project lost at the prototype stage is often lost for its lifetime, our reading of how database adoption works.
- **Popularity is self-reinforcing.** PostgreSQL topped Stack Overflow’s “want to use” and “want to keep using” lists for three years running. When assistants and agents favor what is common, the common choice gets more common.
- **Wrong answers cost credibility.** A developer who gets a wrong limit, price or feature from an AI answer and then hits it in production blames the product, not the assistant.

## How does GEO work for a database company?

It makes your database easy for assistants and agents to understand, set up and compare honestly. No one can promise that an agent will install it.

1. **One clear identity per workload.** Say plainly what your database is best for, and what it is not for, in the same words across docs, your site, cloud marketplace listings and community answers.
2. **Agent-ready documentation.** Keep quick-starts, connection strings, limits and pricing on readable pages; publish framework templates and integrations where coding tools look. Docs that only render with scripts are a common blind spot, as [when AI agents can’t read your site](https://underneath.agency/resources/when-ai-agents-cant-read-your-site) explains.
3. **Benchmarks others can rerun.** Publish methods and scripts, and encourage independent tests. Developers trust results they can reproduce, and independent write-ups give assistants a source other than you.
4. **Honest comparisons and migration guides.** Explain trade-offs against the databases buyers already use. Our research summary on [whether comparison pages help B2B citations](https://underneath.agency/resources/do-comparison-pages-help-b2b-ai-citations) covers what such pages can and cannot do.
5. **Real community presence.** Answer questions on developer forums and Reddit as identified staff, and fix the misunderstandings you find. Why those threads matter is covered in [how Reddit shapes Google’s AI Overviews](https://underneath.agency/resources/does-reddit-shape-google-ai-overviews).
6. **Measurement in assistants and agents.** Track which databases ChatGPT, Gemini, Claude, Perplexity, Copilot and Google’s AI features name for your workloads, and what coding agents install when asked to build typical apps. For sample sizes, read [how many prompts to track](https://underneath.agency/resources/how-many-prompts-to-track-ai-visibility).

## Which questions about databases and coding agents has nobody answered yet?

No public study shows how coding agents pick a database, or whether AI visibility moves enterprise database revenue.

- **Agent figures come from vendors.** The Neon and Supabase numbers are company-reported and cover their own platforms.
- **No study of agent selection.** We found no published test of which databases coding agents pick, or why.
- **We could not use one standard source.** DB-Engines publishes a widely watched popularity ranking, but we could not save its page for verification, so we do not quote its scores.
- **The link to subscription and consumption revenue is thin.** The evidence is reviewed in [does AI visibility drive business results](https://underneath.agency/resources/does-ai-visibility-drive-business-results) and [the B2B SaaS view of AI search revenue](https://underneath.agency/resources/b2b-saas-revenue-from-ai-search).

## How can a database company find out what agents install and assistants recommend?

Have assistants and coding agents build the apps your customers build, and log which database each one picks.

Pick a dozen typical projects and architecture questions for your strongest workloads. Give them to ChatGPT, Gemini, Claude and Copilot and to two or three coding agents, repeating each run several times. Note which database is named or installed, which sources are cited, and whether your limits, prices and features are described correctly. The gaps usually point to missing templates, unreadable docs or a lack of independent benchmarks.

We can run those tests with you: [ask us to audit what agents and assistants say about your database](https://underneath.agency/contact). We will show where AI answers and coding agents place you on the questions that lead to new projects and enterprise evaluations, and which gaps most likely cost you developer adoption and paid usage. Agent-ready documentation, rerunnable benchmarks and honest migration guides make up most of our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) for a database company, and that page explains how each is run.

## Frequently asked questions

### Do AI coding agents really choose the database?

Often. Neon reported over 80% of its databases were created by AI agents, and Supabase more than 60% of new ones by AI tools.

### Should we be Postgres-compatible to be recommended?

Not necessarily, but PostgreSQL is the common default: 66% of developers who used it wanted to keep using it in Stack Overflow’s survey.

### Do benchmarks help AI visibility?

Untested, but independent, reproducible benchmarks give assistants a source beyond your own site, and developers distrust unverified claims.

### Will enterprise architects trust an AI recommendation?

Not on its own. 46% of developers distrust AI output, so expect architects to verify every claim in a proof of concept.

### How do we know what agents install today?

Test it. Ask several coding agents to build typical apps and record the database each one sets up, repeating runs over time.

## Sources

- Supabase, Paul Copplestone (2026-06-04), [Supabase Series F](https://supabase.com/blog/supabase-series-f)
- Supabase, Greg Richardson (2026-05-08), [Supabase Is Now an Official ChatGPT App](https://supabase.com/blog/supabase-is-now-an-official-chatgpt-app)
- Databricks (2025-05-14), [Databricks Agrees to Acquire Neon to Deliver Serverless Postgres for Developers + AI Agents](https://www.databricks.com/company/newsroom/press-releases/databricks-agrees-acquire-neon-help-developers-deliver-ai-systems)
- Snowflake (2025-06-02), [Snowflake Acquires Crunchy Data to Bring Enterprise Ready Postgres Offering to the AI Data Cloud](https://www.snowflake.com/en/news/press-releases/snowflake-acquires-crunchy-data-to-bring-enterprise-ready-postgres-offering-to-the-ai-data-cloud/)
- MongoDB (2026-09-01), [MongoDB, Inc. Announces Second Quarter Fiscal 2027 Financial Results](https://webassets.mongodb.com/_com_assets/cms/mtj3tpvjctfbi5x1q-MongoDB-Inc-Announces-Second-Quarter-Fiscal-2027-Financial-Results.pdf)
- Stack Overflow (2025-07-29), [Stack Overflow’s 2025 Developer Survey](https://stackoverflow.co/company/press/archive/stack-overflow-2025-developer-survey/)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- OpenAI Help Center (2025), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Underneath (2026), [How faithfully do AI assistants quote software prices?](https://underneath.agency/research/ai-pricing-accuracy-study)
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)
- Underneath (2026), [When does a Reddit thread become evidence in Google’s AI?](https://underneath.agency/research/ai-reddit-citations-study)

---

This is the Markdown twin of https://underneath.agency/resources/database-companies-developers-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Can deep content without technical SEO get cited by Gemini?"
description: "Sometimes. In a 14-hotel Gemini audit, an independent hotel with no schema but a 33-question FAQ was cited directly; a polished rival was not."
canonical: "https://underneath.agency/resources/deep-content-without-technical-seo-gemini"
published: 2026-10-11
updated: 2026-10-11
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Can a site with deep content but no technical SEO get cited by Gemini?

Yes, in at least one documented case. In a 2026 audit of Gemini’s hotel answers in Tokyo, a small independent hotel with no schema markup and no technical SEO was cited directly. Its site answered travelers’ questions in depth. The audit is small and exploratory, so treat it as a strong hint, not a rule.

## The short version

1. In an audit of 1,357 Gemini citations for 156 Tokyo hotel questions, researchers at Blossom AI compared 14 hotel websites (Zhu and Chang, March 2026).
2. Hotels that Gemini cited directly scored 8.6 out of 15 on content depth; hotels it skipped scored 3.4, and every hotel scoring 6 or more was cited.
3. Kadoya Hotel, with no schema markup and no technical SEO, scored 13 out of 15 thanks to a 33-question FAQ and a 13-attraction sightseeing guide, and was cited.
4. Hotel K5, with a 9.6 out of 10 guest rating and full technical setup, scored 5 and was never cited directly; Gemini described it using booking sites instead.

## What did the Gemini hotel audit find?

Cited hotels had much deeper question-answering content than skipped ones, regardless of technical polish.

[Zhu and Chang](https://arxiv.org/abs/2603.20062) are two researchers at Blossom AI, a San Francisco company. In March 2026 they put 156 hotel questions to Gemini 2.5 Flash with Google Search grounding switched on. Grounding means Gemini runs Google searches and cites the pages it uses. The questions covered nine Tokyo districts in English and Japanese and produced 1,357 citations.

They then audited the websites of seven hotels Gemini cited directly and seven it did not. Each site was scored from 0 to 3 on five kinds of content: an FAQ, an area guide, a blog, access information and distinctive content. The key test was depth, meaning whether each feature was rich enough to rank in Google for real traveler questions.

| Group (14 hotels) | Average depth score |
|---|---|
| Cited directly by Gemini | 8.6 out of 15 |
| Not cited | 3.4 out of 15 |

The split was clean. Every hotel scoring 6 or above was cited, and every hotel below 6 was not.

## Why did the hotel without schema beat the one with it?

Because Kadoya’s pages answered real questions in depth, while K5’s features were present but thin.

Hotel K5 is a design boutique hotel with a 9.6 out of 10 rating on Booking.com. It has bilingual content, Schema.org markup, a neighborhood page and an FAQ. Each of those features is brief, so it scored only 5 out of 15. Gemini named K5 in three answers but drew its facts from Expedia, Hotels.com and an editorial site, not from K5’s own website.

Kadoya Hotel is an independent with no schema markup and no technical SEO. It scored 13 out of 15. Its 33-question FAQ, 13-attraction sightseeing guide and regular blog create more than 20 indexed pages that answer traveler questions directly.

The authors draw a two-step lesson. First, a hotel’s site needs content deep enough to rank in Google for relevant questions. Second, that content has to answer the question better than competing booking or editorial pages. The authors conclude that brand prestige and technical sophistication did not determine direct citation; content depth did.

## How often does Gemini cite a business’s own website?

Not often. Booking sites and other intermediaries took most citations, even when hotels were named.

[Online travel agencies such as Booking.com and Expedia](https://underneath.agency/resources/ai-search-marketplace-dependence) took 55.3% of all 1,357 citations. Hotels’ own websites took 8.2% of English citations and 11.0% of Japanese ones. Japanese hotel sites tend to carry deeper neighborhood, transit and area content, which the authors link to their higher direct citation.

[How the question was asked](https://underneath.agency/resources/does-question-phrasing-change-ai-sources) mattered too. Questions about atmosphere or experience drew 55.9% of citations from sources other than booking sites, against 30.8% for plain booking-style questions. Hotels’ own sites held a steady share in both, which suggests deep pages help across many kinds of questions.

Gemini is more open to brand-owned pages than some other assistants. A comparison of AI engines by [Chen and colleagues](https://arxiv.org/abs/2509.08919) at the University of Toronto found Gemini drew 21.2% of its sources for niche brands from brand-owned sites. ChatGPT drew just 4.9%.

## Does technical SEO matter at all for AI citations?

Ranking still matters; schema on its own shows little sign of mattering.

The hotel authors are clear that Kadoya’s pages still had to rank in Google to be found. Gemini’s grounding runs on Google Search, so a page that never appears in search never reaches the second step. Kadoya’s 20-plus indexed pages, each answering traveler questions directly, cleared that first step without any technical SEO.

Our own data supports both halves. In our [study of 3,096 pages](https://underneath.agency/research/ai-overview-cited-pages-study) ranking in Google’s top 10, AI Overviews cited 41.7% of pages in positions 1 to 3. That is about twice the rate at 7 to 10. Once pages on the same search were compared, no schema type held up as a reliable predictor of citation.

Rankings for the question as typed are only part of the story. In our [four-assistant study](https://underneath.agency/research/ai-citations-google-rankings-study), 16.6% of Gemini’s citations ranked in Google’s top 10 for the buyer’s question. And 49.8% of Gemini’s citations sat outside Google’s top 100 for the question and for the related searches the study tracked. Pages that cover many specific questions give an assistant more ways to find them.

## Is depth the same as length or an FAQ page?

No. Depth means covering real questions with specific answers; format and word count alone do little.

K5 had an FAQ and a neighborhood page and still failed, because both were brief. A large study of ChatGPT, Google and Perplexity citations by [Zhang and colleagues](https://arxiv.org/abs/2604.25707), three independent researchers, found the same. [Pages in a question-and-answer format](https://underneath.agency/resources/do-faq-pages-help-ai-citations) had 5.74% less influence on answers, on the authors’ measure, than other pages. The pages that shaped answers most held definitions, figures, comparisons and how-to steps.

A 2026 controlled test by [Vishwakarma and colleagues](https://arxiv.org/abs/2605.25517) at Sprinklr, a software vendor, points the same way. Across six AI models, pages with fuller specifications and more comprehensive analysis were cited first more often, while formatting-only changes had no consistent effect.

## What should you do about it?

Invest first in pages that answer your buyers’ specific questions in depth, and treat technical markup as a supporting task.

1. List the concrete questions buyers ask before they choose you, such as access, timing, prices, comparisons and local details.
2. Answer each in depth on your own site, with specifics no booking or review site has.
3. Build depth, not features. A long, organized FAQ beats a short one; one rich guide beats several thin pages.
4. Publish in the languages your buyers search in. Japanese searches, where hotel sites carry deeper content, produced more direct hotel citations than English ones.
5. Keep basic technical health so pages can be crawled and ranked, but do not expect schema alone to earn citations.
6. Check which sources Gemini cites when it mentions you. If it names you but cites intermediaries, your own pages are not answering the question well enough.

If you want help running that check across assistants, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The hotel finding is exploratory, and its authors say plainly that it shows correlation, not causation.

- The depth audit covered only 14 hotels in one city, scored by the authors on their own scale.
- Better-known hotels may both invest in content and be cited for brand reasons; the study cannot separate the two.
- It tested one engine, Gemini 2.5 Flash, in March 2026. ChatGPT, Perplexity and Claude use different search systems.
- Each question was asked once; a repeat of 20 questions showed citations vary from run to run.
- The authors work for a company and propose, but did not run, a test that adds content to a hotel site and tracks citations.

## Frequently asked questions

### Does Gemini need schema markup to cite a website?

No. In the Tokyo hotel audit, Kadoya Hotel had no schema markup and was cited directly, while a hotel with full markup but thin content was not.

### What kind of content does Gemini cite from business websites?

Deep pages that answer specific customer questions. Cited hotel sites scored 8.6 out of 15 on question-answering depth, against 3.4 for hotels Gemini skipped.

### Can a small business compete with big brands in AI search?

Sometimes, on specific questions. An independent hotel with a 33-question FAQ was cited directly, while a highly rated boutique rival reached Gemini answers only through booking and editorial sites.

### Is an FAQ page enough to get cited by AI?

No. Hotel K5 had an FAQ and was not cited because it was brief, and a separate study found question-and-answer formatting alone went with 5.74% less influence on answers.

## Sources

- Zhu and Chang (2026), [The End of Rented Discovery: How AI Search Redistributes Power Between Hotels and Intermediaries](https://arxiv.org/abs/2603.20062), arXiv:2603.20062.
- Chen, Wang, Chen and Koudas (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919), arXiv:2509.08919.
- Zhang, He and Yao (2026), [From Citation Selection to Citation Absorption: A Measurement Framework for Generative Engine Optimization Across AI Search Platforms](https://arxiv.org/abs/2604.25707), arXiv:2604.25707.
- Vishwakarma, Kumar and Jamidar (2026), [What Gets Cited: Competitive GEO in AI Answer Engines](https://arxiv.org/abs/2605.25517), arXiv:2605.25517.
- Underneath (2026), [What pages cited by AI Overviews have in common: 3,096 pages](https://underneath.agency/research/ai-overview-cited-pages-study)
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude cite pages that rank?](https://underneath.agency/research/ai-citations-google-rankings-study)

---

This is the Markdown twin of https://underneath.agency/resources/deep-content-without-technical-seo-gemini. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Can AI search put your dental tech on a dentist’s shortlist?"
description: "Dentists already use AI at work. To make their shortlist, a dental technology company needs clear, checkable facts on cost, clearance and fit."
canonical: "https://underneath.agency/resources/dental-technology-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Can AI search put your dental technology on a dentist’s shortlist?

Yes, if an assistant can find plain, checkable answers about what your product costs, what it is cleared to do and which practice systems it works with. Dentists already use AI inside their practices, and the shortlist for a scanner, imaging unit or practice software increasingly starts with a question typed by an owner dentist or a DSO team. No study yet counts how often those questions name a vendor, so this guide separates what is measured from what we infer.

## The short version

1. Dentists are past the curiosity stage: 43.3% use AI for at least one task in their practice and 26.4% more plan to, according to the [ADA Health Policy Institute](https://www.ada.org/resources/research/health-policy-institute/dental-practice-research/dentists-ai-usage-and-attitudes) in July 2026.
2. Owners are buying under a squeeze: since January 2021, prices for dental equipment and supplies rose 23% while dental reimbursement rose 19%, in the [ADA’s Q2 2026 economic report](https://www.ada.org/-/media/project/ada-organization/ada/ada-org/files/resources/research/hpi/state_us_dental_economy_q22026.pdf).
3. Cost is the main barrier: in an [ADA panel on intraoral scanners](https://adanews.ada.org/ada-news/2021/september/ace-panel-report-finds-about-half-of-dentists-use-intraoral-scanners/), 66% of dentists without one named the investment as the reason, while 34% were considering a purchase.
4. The buyer is consolidating: in 2023, 26.0% of dentists less than 10 years out of school worked in groups with at least 10 locations, and 20.7% of orthodontists were in DSOs, according to [HPI’s practice data](https://www.ada.org/-/media/project/ada-organization/ada/ada-org/files/resources/research/hpi/hpi_evolving_dental_practice_model_2023.pdf?rev=897fa6b028054de1970f70334a3c47aa&hash=8A609BED936215ADDF0762956FE063B4).
5. Assistants check proof before they name anyone: in [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study), ChatGPT looked for reviews in 46.2% of answers and for prices in 23.8%.

## Who buys dental technology, and what is one customer worth?

Owner dentists, DSO leadership and specialists buy it, and a single sale often opens years of software, service and upgrades.

The market is large and still spread across many small buyers. HPI counted 203,631 practicing dentists in the US in 2023, and its [economic impact estimate](https://www.ada.org/resources/research/health-policy-institute/dental-practice-research/economic-impact-of-dental-offices) puts dental office revenue at $478 billion a year, or $2.4 million per dentist. Four buyer groups matter most:

- **Owner dentists in solo and small group practices.** They sign the lease, sit through the demo and live with the result. Their questions are practical: will this pay for itself, and will my team use it?
- **DSOs.** Dental support organizations run billing, marketing and purchasing for many offices. HPI found 13.6% of general dentists affiliated with a DSO in 2023, and nearly one-quarter of dentists in Colorado, Nevada and Georgia. A DSO decision can cover dozens or hundreds of offices.
- **Specialists.** Orthodontists (20.7% DSO-affiliated) and oral surgeons (15.7%) buy imaging, scanning and planning tools that general practices may never need.
- **Labs.** Dental labs buy scanners, design software and printers, and often influence which systems their client practices choose.

Deal sizes vary with the product. In a 2022 [ADA News guide](https://adanews.ada.org/ada-news/2022/june/digital-dentistry-what-to-know-about-a-few-popular-technologies/), a cone-beam CT (CBCT) unit cost $50,000-$100,000 to buy, an intraoral scanner $5,000-$23,000 and a 3D printer $300-$20,000. Prices have moved since, but the shape holds: capital purchases in the tens of thousands, then software, service and materials for years. The category is substantial even for one company. [Align Technology](https://orthodonticproductsonline.com/industry-news/company-news/align-technology-q4-2025-financial-results-record-revenue/), maker of the iTero scanner, reported $789.6 million in imaging systems and CAD/CAM services revenue for 2025.

## How far has AI reached into dental practices and their buying?

Inside the practice, far: two in five dentists use AI tools. For vendor research, nobody has measured it yet.

HPI’s July 2026 survey is the clearest picture of where dentists stand. 22.8% use AI for imaging and diagnostics, 13.6% for insurance verification and 10.1% for business analytics. The line is clinical judgment: less than 5% use AI for treatment recommendations, and four out of five (82.6%) do not plan to. Around one-third (32.4%) say they are overworked, which is why administrative tools draw the most interest.

That matters for vendors in two ways. First, many of your buyers are now comfortable typing a question into an assistant. Second, they are skeptical of AI that oversteps, so an answer that overstates your product’s clinical role may hurt you with exactly the dentist you want.

What no survey measures is how often a dentist asks ChatGPT, Gemini or Perplexity which scanner or software to buy. Across all business purchases, [Gartner’s 2026 survey](https://www.gartner.com/en/newsroom/press-releases/2026-05-20-gartner-survey-finds-sixty-nine-percent-of-b-two-b-buyers-turn-to-sales-reps-to-validate-ai-generated-insights) of 645 buyers found 45% had used generative AI in a recent purchase, mainly to gather information on vendors and products. Dentists buying for their own practice are a narrower group, so treat that as a direction, not a dental figure.

## Which questions do dentists and DSO teams ask before they buy?

Questions about payback, fit with existing systems, clinical scope and price. We wrote the prompts below in the voice of real buyers; they are illustrations, not logged queries.

| Buyer | Illustrative prompt |
|---|---|
| General dentist, first scanner | “Best intraoral scanner for a general practice doing mostly single crowns, and what does it cost?” |
| Implant-focused practice | “Is a CBCT worth it if I place about five implants a month?” |
| Owner under cost pressure | “Cheapest way to go digital for impressions without replacing my practice software” |
| DSO clinical director | “Pearl vs Overjet for a 12-office group: which has more FDA clearances?” |
| Office manager | “Cloud practice management software that works with Dentrix imaging and insurance verification” |
| Orthodontist | “Which scanner works best with clear aligner workflows other than Invisalign?” |
| Lab owner | “Open-format scanners whose files my lab can use without extra fees” |

Wording changes answers. In [our phrasing study](https://underneath.agency/research/ai-prompt-phrasing-study), adding “on a tight budget” kept the original first brand only 15.3% of the time, against 68.0% when the same question was simply asked again. With cost the leading barrier, many dentists will add exactly that kind of qualifier, and a reasonable expectation is that they then see a different set of products.

## How does an AI answer turn into a sale for a dental technology company?

By sending a dentist to the demo already knowing your price range, your integrations and why you fit their practice.

The path we see in this industry runs through people as well as screens:

1. **Trigger.** An aging sensor, a new associate, a move into implants or aligners, rising lab bills, or a DSO acquisition that standardizes equipment.
2. **Research.** Continuing education courses, study clubs, peers, journals, trade shows and, increasingly, an assistant asked to compare options.
3. **Shortlist and demo.** A rep visit, often through a distributor, a booth at a dental meeting, or an online demo.
4. **Financing and install.** A lease or loan, training and integration with practice software and imaging.
5. **Expansion.** Software subscriptions, service plans, materials, upgrades and, for DSOs, office-by-office rollout.

The fifth step is where the value sits. Overjet, in its [own comparison page](https://www.overjet.com/blog/overjet-vs-pearl-dental-ai-software), says its software runs at more than 240 practices in one DSO and 147 in another. Those are the vendor’s claims, but they show why an enterprise buyer’s first impression matters: one evaluation can decide hundreds of offices.

Most of this happens without a trackable click. A dentist who reads an assistant’s comparison and then calls a distributor rep shows up in your pipeline as a rep referral, not as AI traffic. We explain the measurement problem in [how AI answers affect pipeline](https://underneath.agency/resources/ai-answers-pipeline-revenue).

## What decides whether an assistant names your scanner, imaging unit or software?

No platform documents how it picks dental products; studies show assistants look for reviews, prices and rankings first.

**Documented by the platforms.** Google says one question in AI Overviews or AI Mode [can be split into several related searches](https://developers.google.com/search/docs/appearance/ai-features), a technique it calls “query fan-out.” A question about scanners for crown work may fan out into searches about accuracy, cost and integrations. Google does not explain how a product ends up named.

**Observed in our studies.** In our hidden-searches study, ChatGPT ran a search aimed at a named publication, ranking or award in 43.8% of answers, besides the review and price searches above. In [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), only 25.2% of brands ChatGPT named appeared in all five runs of the same question, so one test tells you little. When [we asked assistants](https://underneath.agency/research/is-it-legit-ai-reputation-study) whether brands were legitimate, 88.0% of answers cited a review or complaint platform.

**What dental buyers check.** Regulatory scope and evidence. Dental AI shows how contested this is. Pearl [received FDA clearance](https://pk.dental-tribune.com/news/first-fda-clearance-granted-for-ai-platform-analyzing-both-2d-and-3d-dental-images/) in 2025 for AI analysis of both 2D and 3D images. Overjet’s comparison says it has ten FDA-cleared modules to Pearl’s seven. [Pearl’s comparison](https://hellopearl.com/blog/pearl-vs-overjet-which-is-best-for-your-practice-pearl) says it is the only dental AI independently validated, to 94% accuracy across more than 100,000 verified cases. Each vendor wrote the page that ranks itself first. An assistant asked “Pearl or Overjet?” has to reconcile both, and we infer it leans on whatever independent sources it can find.

**Our inference.** A reasonable expectation is that products with clearly stated clearances, published integration lists, honest cost ranges and independent reviews give assistants more to repeat accurately. No study has tested this for dental products.

## What does it cost a dental technology company to be left out?

Mostly demos and multi-office rollouts that never start, though no study has measured that loss in dentistry.

- **The first purchase locks in the rest.** A practice that buys a rival’s scanner tends to adopt that ecosystem’s software and workflows, and a DSO that standardizes on a rival does so across every office.
- **Price-sensitive buyers filter early.** With 66% of non-buyers in the ADA scanner panel citing cost, an assistant that cannot find your price range or financing options may leave you out of exactly the questions where cheaper entrants compete.
- **Wrong facts cost fit.** If an assistant says your software does not work with Open Dental or Eaglesoft when it does, the dentist who needs that integration moves on. Pearl’s page lists integrations including Open Dental, Dentrix, Eaglesoft, Planmeca Romexis and Sidexis, which is the kind of fact assistants can repeat. Our guide to [correcting wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) shows how to trace such an error back to its source.
- **Self-ranking lists fill the gap.** In [our self-ranking study](https://underneath.agency/research/self-promoting-best-lists-study), 24.2% of 269 cited numbered “best X” lists put their own publisher first. In a category where vendors publish their own comparisons, the absent company has no voice in the answer.

## How does GEO work for a dental technology company?

It makes your product’s scope, cost and fit easy to find and confirm across independent sources. Nobody can guarantee a recommendation.

1. **State clinical scope precisely.** For each product, say what it is cleared or indicated for and what it is not. Dentists distrust AI that oversteps clinical judgment, and assistants repeat what pages say.
2. **Publish integration facts.** One page per practice management and imaging system you support, with versions and limits. These answer the questions office managers and DSO IT teams ask.
3. **Explain cost and payback honestly.** Give price ranges, financing and the practice volume at which the purchase makes sense. Owners facing the fiscal squeeze ask this first, and assistants search for prices.
4. **Earn independent coverage.** Peer-reviewed studies, CE courses taught by clinicians who use your product, dental trade press and association coverage carry more weight than your own comparison page. See [how brands build authority for AI search](https://underneath.agency/resources/how-brands-build-authority-for-ai-search) and [whether comparison pages help B2B citations](https://underneath.agency/resources/do-comparison-pages-help-b2b-ai-citations).
5. **Keep reviews current** on the dental product review sites and software review platforms buyers read, and answer complaints in public.
6. **Align distributor catalogs.** Product names, specifications and prices should match across your site and every distributor listing, so an assistant does not meet conflicting facts.
7. **Name DSO results with permission.** Case studies that name the group, the number of offices and the measured change give enterprise buyers something checkable.
8. **Measure with real buyer wording.** Test owner, specialist, DSO and lab prompts, with budget and integration variations, across ChatGPT, Google AI Overviews and AI Mode, Gemini, Perplexity, Copilot and Claude, several runs each.

Health technology sellers facing hospital buyers have a different path; see [how healthcare software companies win health system deals](https://underneath.agency/resources/healthcare-software-ai-search). For imaging and scanner makers that also sell to hospitals, where buyers issue RFQs and check prices, contracts and service, see [how medical equipment sellers win RFQs](https://underneath.agency/resources/medical-equipment-leads-ai-search).

## What can’t today’s data tell a dental technology company?

How many dentists ask an assistant before buying equipment, and whether being named produces demos. No published study answers either.

- **No buying survey for this question.** HPI measures how dentists use AI in the practice, not how they research vendors.
- **The ADA cost ranges are from 2022** and the scanner panel from 2021; both describe direction, not today’s prices or adoption.
- **Vendor claims are vendor claims.** Pearl’s accuracy figure and Overjet’s clearance count come from pages each company wrote about the other.
- **Distributor influence is unmeasured.** Reps still shape many purchases, and no study shows how AI answers interact with that relationship.

## Where should a dental technology company start?

With the questions your best buyers ask before they book a demo, and a record of what assistants answer today.

List 20 to 30 prompts across owner dentists, specialists, DSO teams and labs, including cost, integration and comparison questions. Ask each major assistant several times. Note which products are named, which pages and review sites are cited, and whether your clearances, prices and integrations are described correctly. Companies selling software alongside hardware can also compare notes with [how B2B software companies earn revenue from AI search](https://underneath.agency/resources/b2b-saas-revenue-from-ai-search).

If you would rather have that done for you, [ask us for a dental AI visibility review](https://underneath.agency/contact). We will show which owner, DSO and specialist questions name your product, where assistants send dentists instead, and which missing or wrong facts are most likely costing you demo requests and rollouts. Stating clinical scope, integration lists and cost ranges where assistants and dentists can check them is the core of our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## Frequently asked questions

### Do dentists use ChatGPT to choose equipment?

Nobody has measured it. HPI found 43.3% of dentists use AI for at least one practice task, so many are comfortable with assistants, but buying research is unstudied.

### Should a dental technology company publish prices?

Publish ranges and financing at least. ChatGPT searched for prices in 23.8% of answers in our hidden-searches study, and 66% of dentists without a scanner cited cost.

### Do FDA clearances help AI visibility?

They give assistants a checkable fact. State each clearance and its scope plainly, and link to the public record, rather than relying on a comparison page you wrote.

### Do DSOs and solo practices need different content?

Yes. Owners ask about payback and ease of use; DSO teams ask about standardization, integrations, security and results across offices.

### How often should a dental technology company recheck what assistants say about its products?

Regularly, with several runs per question: only 25.2% of brands appeared in all five runs of the same question in our consistency study.

## Sources

- ADA Health Policy Institute (2026-07), [Dentists’ AI usage and attitudes](https://www.ada.org/resources/research/health-policy-institute/dental-practice-research/dentists-ai-usage-and-attitudes)
- ADA Health Policy Institute (2026), [The state of the US dental economy, Q2 2026](https://www.ada.org/-/media/project/ada-organization/ada/ada-org/files/resources/research/hpi/state_us_dental_economy_q22026.pdf)
- ADA Health Policy Institute (2024), [The evolving dental practice model: data update for 2023](https://www.ada.org/-/media/project/ada-organization/ada/ada-org/files/resources/research/hpi/hpi_evolving_dental_practice_model_2023.pdf?rev=897fa6b028054de1970f70334a3c47aa&hash=8A609BED936215ADDF0762956FE063B4)
- ADA Health Policy Institute, [Economic impact of dental offices in the United States](https://www.ada.org/resources/research/health-policy-institute/dental-practice-research/economic-impact-of-dental-offices)
- ADA News (2021-09-21), [ACE Panel report finds about half of dentists use intraoral scanners](https://adanews.ada.org/ada-news/2021/september/ace-panel-report-finds-about-half-of-dentists-use-intraoral-scanners/)
- ADA News (2022-06-10), [Digital dentistry: what to know about a few popular technologies](https://adanews.ada.org/ada-news/2022/june/digital-dentistry-what-to-know-about-a-few-popular-technologies/)
- Orthodontic Products (2026-02-05), [Align Technology reports record revenue for Q4 and fiscal 2025](https://orthodonticproductsonline.com/industry-news/company-news/align-technology-q4-2025-financial-results-record-revenue/)
- Dental Tribune (2025-05-29), [First FDA clearance granted for AI platform analyzing both 2D and 3D dental images](https://pk.dental-tribune.com/news/first-fda-clearance-granted-for-ai-platform-analyzing-both-2d-and-3d-dental-images/)
- Pearl (2025-10-02), [Pearl vs. Overjet: which dental AI is best for your practice?](https://hellopearl.com/blog/pearl-vs-overjet-which-is-best-for-your-practice-pearl)
- Overjet (2026-03-18), [Overjet vs Pearl AI: definitive comparison 2026](https://www.overjet.com/blog/overjet-vs-pearl-dental-ai-software)
- Gartner (2026-05-20), [Gartner survey finds 69% of B2B buyers turn to sales reps to validate AI-generated insights](https://www.gartner.com/en/newsroom/press-releases/2026-05-20-gartner-survey-finds-sixty-nine-percent-of-b-two-b-buyers-turn-to-sales-reps-to-validate-ai-generated-insights)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [Does rewording a question change AI brand recommendations?](https://underneath.agency/research/ai-prompt-phrasing-study)
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)
- Underneath (2026), [“Is this brand legit?” How AI assistants build a reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study)
- Underneath (2026), [How many “best of” lists cited by AI rank their own brand first?](https://underneath.agency/research/self-promoting-best-lists-study)

---

This is the Markdown twin of https://underneath.agency/resources/dental-technology-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How developer tool companies win users when developers ask AI"
description: "By being the tool AI coding agents and assistants pick for a developer’s stack, then turning that free signup into team plans and enterprise contracts."
canonical: "https://underneath.agency/resources/developer-tools-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How do developer tool companies win users when developers ask AI first?

By becoming the tool that AI assistants and coding agents choose when a developer asks “what should I use?”, then turning that free signup into a paid team and, later, an enterprise contract. For developer tools the AI answer is often not a recommendation a person reads. It is a package an agent installs. Being picked, or skipped, now happens inside the editor, before any developer visits your website.

## The short version

1. In the 2026 [JetBrains Developer Ecosystem research](https://blog.jetbrains.com/research/2026/08/ai-coding-agent-adoption-2026/), 90% of professional developers used AI coding agents at work at least weekly and 68% used them daily.
2. In [Amplifying’s test](https://amplifying.ai/research/claude-code-picks/report) of 2,430 open-ended prompts such as “i need a database, what should i use,” Claude Code chose GitHub Actions for CI/CD in 94% of answers and Stripe for payments in 91%, while traditional cloud providers received zero primary picks for deployment.
3. Agents now read documentation more than people do. Across docs sites hosted by [Mintlify](https://mintlify.com/data.md/), agents made 61.87% of requests in August 2026, and agent readership grew 7.7x from February to August while human readership grew 1.2x.
4. Agents are already creating accounts. [Databricks](https://www.databricks.com/company/newsroom/press-releases/databricks-agrees-acquire-neon-help-developers-deliver-ai-systems) said over 80 percent of the databases provisioned on Neon were created automatically by AI agents rather than by humans.
5. Signups already carry the trace of AI answers: the developer platform [Vercel](https://vercel.com/blog/how-were-adapting-seo-for-llms-and-ai-search) says ChatGPT refers around 10% of new signups to it, up from 1% six months earlier.

## Who buys developer tools, and what is a new user worth?

Developers choose them, usually for free, and the company pays later as usage spreads across teams.

The buying journey runs bottom-up. In the [2025 Stack Overflow Developer Survey](https://survey.stackoverflow.co/2025/work), 48% of developers said they endorsed or influenced the purchase of new technology in their organization in the past year, and a fifth of those influenced a substantial addition to the company’s tech stack. Asked what makes them endorse a tool, developers ranked an easy-to-use API first, a robust and complete API second and a reputation for quality third. “AI integration or AI Agent capabilities” came ninth of ten.

The audience is huge and growing fast. GitHub’s [Octoverse 2025 report](https://github.blog/news-insights/octoverse/octoverse-a-new-developer-joins-github-every-second-as-ai-leads-typescript-to-1/) counts more than 180 million developers on the platform, with more than 36 million joining in a single year. Most of those people will never sign a contract themselves. Their choices, repeated across a company, decide which tool the company ends up paying for.

What a new user is worth depends on how far that spread goes. Free-to-paid conversion is thin: the research firm [Sacra estimates](https://sacra.com/research/supabase-170m-year-growing-221-yoy/) that Supabase’s paying customers are about 2.5% of its registered users, at roughly $700 a year each. The value sits in expansion. [GitLab’s fiscal 2026 results](https://s204.q4cdn.com/984476563/files/doc_financials/2026/q4/Gitlab-4Q26-Earnings-Press-Release.pdf) report 1,456 customers paying more than $100,000 a year, up 18% year over year, and a dollar-based net retention rate of 118%: existing customers spent 18% more than a year before. A single developer’s first install can be the start of a six-figure account, which is why the first pick matters so much.

## Where do AI assistants already sit in the developer’s workflow?

Inside the editor and the terminal, where agents now write code, choose packages and read documentation.

Use is close to universal. In the Stack Overflow survey, 84% of respondents used or planned to use AI tools in their development process, and 51% of professional developers used them daily. JetBrains reports that around 39% of professional developers worldwide used Claude Code at work in May to July 2026, up from 18% in January, and 47% in the United States. OpenAI’s Codex grew from 3% to 16% adoption over the same period. The makers of those coding tools run their own version of this race, covered in [how AI coding tools win developers](https://underneath.agency/resources/ai-coding-tools-developers-ai-search). These are not search engines in the usual sense, but they answer the same question a search used to: which tool should I use for this?

The old discovery channels still matter, but they are shifting. Technical documentation remained the top resource for learning to code, used by 68% of respondents in the Stack Overflow survey, and Stack Overflow itself was the most used community platform at 84%. Yet the number of new questions posted there has fallen to levels [not seen since 2009](https://blog.pragmaticengineer.com/are-llms-making-stackoverflow-irrelevant/), according to data reported by The Pragmatic Engineer. Developers who once asked a forum now ask an assistant.

The reader of your documentation has changed too. Mintlify, a documentation platform, reports that 70% of the docs sites it hosts serve more machine requests than human page loads. For the “dev infrastructure and data” sites in its August 2026 leaderboard, the agent share was 37.9%, up from 15.5% in February. These are a vendor’s own measurements of its own customers, but the direction is consistent across every industry group it reports.

Developers do not trust the output blindly. In the Stack Overflow survey, 46% of developers distrusted the accuracy of AI tools against 33% who trusted it, and 35% said they visit Stack Overflow after running into issues with AI responses. A reasonable expectation is that developers accept an agent’s first pick for low-stakes choices and check it for anything that touches production, security or cost.

## Which questions lead developers to a tool?

Two kinds: questions a developer asks a chat assistant, and requests that leave the choice to a coding agent.

The second kind is new and specific to this industry. Amplifying’s benchmark used real, open-ended prompts with no tool named, such as “how do i deploy this?”, “add user authentication” and “what testing framework works best with this stack.” The agent did not just suggest a tool. Amplifying recorded what it installed, configured and committed. API vendors meet the same test when an agent writes the integration; see [how API companies win agent-chosen integrations](https://underneath.agency/resources/api-companies-ai-search).

The first kind looks more like classic research. We wrote the chat prompts below as examples of how a developer frames a tooling question; they are not logged queries:

- Category within a stack: “Best feature flag service for a Next.js app on Vercel.”
- Alternatives: “Open-source alternatives to Datadog for a small team.”
- Comparison: “Supabase vs Firebase for a SaaS app with row-level security.”
- Price and limits: “Which error-tracking tools have a free tier above 10,000 events a month?”
- Fit and risk: “Which CI services support self-hosted runners and SOC 2 reports?”

Questions that name a framework or hosting stack decide which tools enter the running. Comparison, pricing and compliance questions decide who gets past a team lead. For broader B2B software buyer questions, see [our article on SaaS shortlists](https://underneath.agency/resources/b2b-saas-revenue-from-ai-search).

## How does an AI pick turn into revenue for a developer tool?

Through usage: the agent installs you, the project grows on you, the team pays, and the company expands.

**The pick.** When an agent picks a tool, the tool ships with the code. Databricks put it plainly when it bought Neon: “four out of every five databases on their platform are spun up by code, not humans.” Sacra estimates that AI coding agents now create more than 60% of new Supabase databases and that database launches there are up 600% year over year, calling coding agents its fastest-growing driver of usage. These are a buyer’s statement and an analyst’s estimates, but both point the same way. Database vendors can go deeper with [how a database gets chosen for the stack](https://underneath.agency/resources/database-companies-developers-ai-search).

**The signup.** Some of the path is visible as referrals. Vercel saw ChatGPT’s share of its new signups reach around 10%, up from 4.8% the previous month. Much of the path is not: an agent that installs a package or provisions a database leaves no referral in analytics, and we infer it shows up instead as unexplained growth in free accounts.

**The team plan and the contract.** A project that grows on a free tier eventually hits limits, needs collaboration or needs a security review. GitLab’s 118% net retention shows how much of a developer tool’s revenue comes from accounts that grow after the first sale. The AI pick is worth far more than the first plan it produces, because it decides which tool is already inside the codebase when the company starts paying.

## What decides whether an AI assistant or coding agent picks your tool?

Platforms say little; studies show that the project’s existing stack and the tool’s established position drive most picks.

**Documented by the platform.** No AI coding agent vendor publishes how its agent chooses third-party tools. Anything stronger than the observations below is inference.

**Observed in studies.** In Amplifying’s Claude Code study, context mattered more than phrasing. The same request produced Vercel for a Next.js project and Railway for a Python project, and picks stayed stable across five phrasings of the same request (76% average stability). Models agreed with each other 90% of the time within one language ecosystem. An agent can know a library well and still not install it: Redux was mentioned 23 times but never chosen as the primary recommendation, and AWS Amplify was mentioned 24 times but never recommended.

Different agents pick differently. In a second Amplifying study of [1,470 responses from Claude Code and OpenAI Codex](https://amplifying.ai/research/codex-vs-claude-code-picks/report), the two agents agreed on the top tool in 7 of 12 categories, and 6 of those 7 agreements were to build it themselves. Codex picked Statsig for feature flags 27% of the time; Claude never did, though it mentioned Statsig in 28% of those answers. Amplifying calls this a pick-rate gap, not evidence of steering. Statsig is owned by OpenAI.

**Trust factors specific to developer tools.** We infer that what helps an agent pick a tool is what helps a developer trust it: a clear, complete API, plenty of working examples in the frameworks people use, accurate documentation an agent can read, and a public record of reliability. Few sites make that easy for machines. In our study of 5,902 top websites, only [11.5% published a valid llms.txt file](https://underneath.agency/research/llms-txt-adoption-study), and just [3.2% returned Markdown](https://underneath.agency/research/agent-readable-web-study) when an agent asked for it. For what happens when an agent cannot read a site, see [our article on unreadable sites](https://underneath.agency/resources/when-ai-agents-cant-read-your-site).

## What does it cost a developer tool to be skipped?

Often the whole project, because the tool an agent installs on day one rarely gets replaced later.

Some categories are already close to locked. In Amplifying’s study, shadcn/ui took 90% of primary picks for UI components, Stripe 91% for payments and GitHub Actions 94% for CI/CD. CI/CD, deployment and monitoring vendors can see [how DevOps platforms get recommended](https://underneath.agency/resources/devops-platforms-ai-search).

The biggest competitor may be no tool. Claude Code built a custom solution in 252 of 2,073 identifiable picks, 12% of the total, making “build it yourself” its single most common recommendation. Custom code won 68.8% of picks for feature flags and 47.7% for authentication, ahead of LaunchDarkly and NextAuth.js. For vendors of those tools, the loss is invisible: no comparison page lost, no sales call missed, just code that never needed you.

The loss also hides in your data. A developer who never visits your site because the agent chose a rival leaves no trace in your analytics, a gap we describe in [why analytics miss AI visibility](https://underneath.agency/resources/why-analytics-miss-ai-visibility).

## What does GEO mean for an API, SDK or developer platform?

It makes your tool easy for assistants and coding agents to find, understand and install; it cannot guarantee a pick.

For a developer tool, generative engine optimization (GEO) comes down to six pieces of work:

1. **Documentation agents can read.** Serve clean text or Markdown, publish an llms.txt, keep reference pages complete and current, and consider an MCP server so agents can search your docs directly. Mintlify reports that MCP tool calls across its docs sites grew 6.0x from January to August 2026.
2. **Examples in the stacks that matter.** Since the project’s stack shapes the pick, publish working quickstarts, templates and integration guides for the frameworks your buyers use. We infer that a tool with a clean example for a stack is more likely to be chosen for that stack.
3. **Community presence.** Keep answers on GitHub issues, Stack Overflow and developer forums accurate and current. These are the places developers check when they do not trust an AI answer.
4. **Independent coverage.** Earn mentions in tutorials, newsletters, conference talks and comparison write-ups by people who use your tool. Assistants that search the web lean on these, and our article on [how brands build authority](https://underneath.agency/resources/how-brands-build-authority-for-ai-search) covers the evidence.
5. **Honest comparisons and clear pricing.** Publish fair comparison pages and a single, current pricing page with free-tier limits stated plainly. Assistants misquote software pricing often enough to matter: in [our pricing study](https://underneath.agency/research/ai-pricing-accuracy-study), only 61.9% of plan prices quoted by four assistants for 45 software products matched the official page in full.
6. **Measurement across agents and stacks.** Test fixed “what should I use” requests in several agents and project types, repeated over time, and compare with signup sources and what new users say when asked how they found you. Checking one assistant is not enough; see [is tracking ChatGPT enough](https://underneath.agency/resources/is-tracking-chatgpt-enough).

## What is still unmeasured about how coding agents choose developer tools?

How agents pick tools, and whether better documentation changes those picks, is still largely unmeasured.

The agent-choice evidence comes mainly from one research group, Amplifying, testing JavaScript and Python projects; its own caveats say it cannot separate a tool’s quality from how often it appears in training data. Mintlify’s traffic figures describe its own customers, measured by its own method. Sacra’s Supabase figures are estimates. No published study yet shows that improving a tool’s documentation or examples raises its pick rate in coding agents, or ties AI picks to paid conversions and enterprise contracts. Any GEO plan for a developer tool should treat those links as hypotheses to test.

## How can a developer tool company see whether agents are steering paid teams its way?

Test which tools coding agents and assistants pick for the jobs your product does, in real project stacks.

A useful first step is an audit that runs your category’s real “what should I use” requests across the main agents and assistants, in the stacks your users work in, and checks the results against your signup and expansion data. That shows where you are picked, where a rival or custom code wins, and which gaps matter most for paid teams. Our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page sets out how that work runs for a developer tool, from agent-readable docs and stack-specific examples to repeated pick tests. To run that test with outside help, tied to free-to-paid conversion and expansion revenue, [talk to us about an agent pick-rate audit](https://underneath.agency/contact).

## Frequently asked questions

### Do AI coding agents recommend specific developer tools?

Yes. In Amplifying’s tests, Claude Code and Codex installed specific tools in response to open-ended requests, with near-unanimous picks in some categories, such as Vercel for Next.js deployment. They also often built custom code instead of using any tool.

### Does an llms.txt file get a developer tool picked more often?

No study has shown that. An llms.txt and Markdown versions of docs make pages easier for agents to read, and agents read documentation heavily, but the link from those files to pick rates is unmeasured.

### Is Stack Overflow still worth investing in for developer marketing?

Yes, as a trust check rather than a discovery channel. New questions have fallen sharply, but in the 2025 survey it was still the most used community platform, and developers go there when AI answers fail them.

### How can a developer tool company see whether agents pick it?

Run the same open-ended requests in several agents and project types, repeat them over weeks, and record the primary pick, alternatives and mentions. Then compare those rates with self-reported attribution at signup.

## Sources

- JetBrains Research (2026), [AI coding agent adoption in 2026](https://blog.jetbrains.com/research/2026/08/ai-coding-agent-adoption-2026/)
- Amplifying (2026), [What Claude Code Actually Chooses](https://amplifying.ai/research/claude-code-picks/report)
- Amplifying (2026), [Codex vs Claude Code: what AI coding agents pick](https://amplifying.ai/research/codex-vs-claude-code-picks/report)
- Mintlify (2026), [Data: agents vs human traffic](https://mintlify.com/data.md/)
- Databricks (2025), [Databricks Agrees to Acquire Neon to Help Developers Deliver AI Systems](https://www.databricks.com/company/newsroom/press-releases/databricks-agrees-acquire-neon-help-developers-deliver-ai-systems)
- Vercel (2025), [How we’re adapting SEO for LLMs and AI search](https://vercel.com/blog/how-were-adapting-seo-for-llms-and-ai-search)
- Stack Overflow (2025), [2025 Developer Survey: Work](https://survey.stackoverflow.co/2025/work), [AI](https://survey.stackoverflow.co/2025/ai) and [press release](https://stackoverflow.co/company/press/archive/stack-overflow-2025-developer-survey/)
- GitHub (2025), [Octoverse: A new developer joins GitHub every second as AI leads TypeScript to #1](https://github.blog/news-insights/octoverse/octoverse-a-new-developer-joins-github-every-second-as-ai-leads-typescript-to-1/)
- Sacra (2026), [Supabase at $170M/year growing 221% YoY](https://sacra.com/research/supabase-170m-year-growing-221-yoy/)
- GitLab (2026), [GitLab Reports Fourth Quarter and Full Fiscal Year 2026 Financial Results](https://s204.q4cdn.com/984476563/files/doc_financials/2026/q4/Gitlab-4Q26-Earnings-Press-Release.pdf)
- The Pragmatic Engineer (2025), [Are LLMs making StackOverflow irrelevant?](https://blog.pragmaticengineer.com/are-llms-making-stackoverflow-irrelevant/)
- Underneath (2026), [llms.txt adoption](https://underneath.agency/research/llms-txt-adoption-study), [Markdown for AI agents](https://underneath.agency/research/agent-readable-web-study) and [software pricing accuracy](https://underneath.agency/research/ai-pricing-accuracy-study)

---

This is the Markdown twin of https://underneath.agency/resources/developer-tools-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How DevOps companies get recommended when engineers ask AI"
description: "By being the tool AI answers and coding agents find in docs, GitHub and engineer communities, then converting team adoption into platform deals."
canonical: "https://underneath.agency/resources/devops-platforms-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How do DevOps companies get recommended when engineers ask AI which tool to use?

By being easy for AI assistants and coding agents to find, read and trust: complete public documentation, working examples on GitHub, honest comparisons, and a reputation in the communities engineers already consult. DevOps tools are usually adopted by engineers first and bought by the enterprise later, so an AI answer that names you can start a free trial that becomes a platform contract. Engineers also distrust AI output more than most buyers do, which makes verifiable technical proof the deciding factor.

## The short version

1. Engineers drive the purchase: 48% of developers endorsed or influenced a new technology purchase in the past year, in [Stack Overflow’s 2025 survey](https://survey.stackoverflow.co/2025/work) of more than 49,000 developers.
2. They already work through AI: 84% use or plan to use AI tools in development, and ChatGPT (82%) is the most used assistant, [Stack Overflow reports](https://survey.stackoverflow.co/2025/ai). Yet 46% distrust the accuracy of AI output.
3. Platform buying is now the norm: 90% of organizations have adopted at least one internal platform, according to Google Cloud’s [2025 DORA report](https://cloud.google.com/blog/products/ai-machine-learning/announcing-the-2025-dora-report) of nearly 5,000 technology professionals.
4. Small teams become large contracts: [GitLab](https://s204.q4cdn.com/984476563/files/doc_financials/2026/q4/Gitlab-4Q26-Earnings-Press-Release.pdf) had 1,456 customers paying more than $100,000 a year and a dollar-based net retention rate of 118% at the end of fiscal 2026.
5. AI already sends signups to developer platforms: [Vercel](https://vercel.com/blog/how-were-adapting-seo-for-llms-and-ai-search) said in 2025 that ChatGPT was bringing in around 10% of its new signups, against 1% six months before.

## Who picks a DevOps tool, who pays for it, and how large can the account grow?

Engineers choose it, platform teams standardize it, and procurement signs; a won account often grows for years.

The buying journey usually runs bottom-up. In the Stack Overflow survey, nearly half of developers endorsed or influenced a technology purchase last year, and one in five of those influenced a substantial addition to their company’s tech stack. The tools they pick (continuous integration, deployment, infrastructure as code, observability, incident response) then meet a platform team that decides what the whole company will standardize on. Security tools travel a similar path with a second buyer; see [how AppSec vendors win demand](https://underneath.agency/resources/application-security-demand-ai-search).

That platform layer is now common. DORA found 90% of organizations have adopted at least one platform, and that a high-quality internal platform is tied to an organization’s ability to unlock the value of AI. The [CNCF and SlashData](https://www.cncf.io/announcements/2026/03/24/cncf-and-slashdata-report-finds-platform-engineering-tools-maturing-as-organizations-prepare-for-ai-driven-infrastructure/) found 28% of organizations have a dedicated platform engineering team. Separately, the [CNCF reports](https://www.cncf.io/announcements/2026/03/24/cncf-and-slashdata-report-finds-cloud-native-community-reaches-nearly-20-million-developers/) that the share of developers working without formalized DevOps or platform practices fell from 20% to 12%, and that the cloud native developer community reached 19.9 million.

GitLab’s results show how team adoption turns into enterprise revenue. At the end of fiscal 2026, GitLab passed $1 billion in annual recurring revenue and reported:

| Customer size | Customers | Growth over the year |
|---|---|---|
| More than $5,000 of annual recurring revenue | 10,682 | 8% |
| More than $100,000 | 1,456 | 18% |
| More than $1 million | 155 | 26% |

The largest accounts grew fastest, and a dollar-based net retention rate of 118% means existing customers spent more each year. For a DevOps vendor, the first team that adopts the tool is the start of a long expansion, not the end of a sale.

## Where in an engineer’s day do AI assistants already show up?

Everywhere in the daily workflow, and increasingly in documentation reading, where agents now read more than people do.

- **Assistants are standard.** Stack Overflow found 51% of professional developers use AI tools daily. DORA found 90% of technology professionals use AI at work. The assistants themselves compete for those engineers, as [how AI coding tools win developers](https://underneath.agency/resources/ai-coding-tools-developers-ai-search) explains.
- **AI is a learning channel.** 44% of developers used AI tools to learn to code, up from 37%. Technical documentation remained the top learning resource, at 68%.
- **Agents read docs.** Mintlify, which hosts documentation sites, reports on its [live data page](https://www.mintlify.com/data) that agent readership grew 7.7 times from February to August 2026, while human readership grew 1.2 times. Mintlify sells documentation tools, so treat this as vendor data.
- **Signups already arrive from ChatGPT.** At Vercel, a deployment platform, ChatGPT referred around 10% of new signups, up from 4.8% the previous month, the company reported.

DevOps work is also where engineers are most cautious about AI. Stack Overflow found 76% of developers do not plan to use AI for deployment and monitoring, the most resisted task in the survey. We infer that engineers will use AI to shortlist tools but verify the answer against documentation, peers and a trial before trusting a tool with production.

## Which questions do engineers ask AI about DevOps tools?

Questions about fit, migration, alternatives, standards, pricing and security. We wrote the example prompts below to show how a platform engineer or SRE might phrase these questions; they are not logged from real users.

| Stage | Illustrative prompt |
|---|---|
| Category | “What is the best CI/CD setup for a monorepo with 200 engineers on Kubernetes?” |
| Alternatives | “What are the main alternatives to Jenkins for a team moving to the cloud?” |
| Comparison | “GitHub Actions or GitLab CI for a company that self-hosts its runners?” |
| Standards | “Which observability tools accept OpenTelemetry data without vendor-specific agents?” |
| Migration | “How hard is it to move our Terraform state to another infrastructure-as-code tool?” |
| Pricing | “Which incident management tools charge per responder rather than per seat?” |
| Compliance | “Which deployment platforms have SOC 2 Type II and support FedRAMP environments?” |

A coding agent may ask similar questions on a developer’s behalf while it works, reading docs directly, as Mintlify’s data suggest. For API vendors, [how agents choose an integration](https://underneath.agency/resources/api-companies-ai-search) covers that case in depth.

Standards questions matter more than they used to. The CNCF announced in May 2026 that OpenTelemetry had graduated, describing it as allowing organizations to [change analysis tools without rewriting code](https://www.cncf.io/announcements/2026/05/21/cloud-native-computing-foundation-announces-opentelemetrys-graduation-solidifying-status-as-the-de-facto-observability-standard/). When switching is easier, the shortlist is reopened more often.

## How does a tool named in an AI answer end up as an enterprise DevOps contract?

Through a product-led path: answer, documentation, free use, team adoption, platform standard, enterprise contract.

1. An engineer asks an assistant, or an agent searches, for a tool to solve a specific problem.
2. The answer names a few tools and links documentation, GitHub repositories or community threads.
3. The engineer reads the docs and tries the free tier or open source version.
4. The team adopts it; the platform team evaluates it for company-wide use.
5. Procurement and security review the vendor, and an enterprise contract follows, then expansion.

The AI answer matters most at steps 1 to 3, where no salesperson is involved. Vercel’s numbers show that step can convert at scale. The value lands much later: GitLab’s 118% net retention shows how accounts grow once a tool becomes part of the platform. That gap in time makes attribution hard; we cover it in [how to prove GEO caused sales](https://underneath.agency/resources/prove-geo-caused-sales).

## Why does an assistant recommend one CI/CD or observability tool over another?

The assistants don’t say; research points to independent and community sources, and engineers reward proof they can run themselves.

**What the search providers publish.** According to Google, AI Overviews and AI Mode [may use a “query fan-out” technique](https://developers.google.com/search/docs/appearance/ai-features), so one question about replacing Jenkins can spawn multiple related searches. OpenAI writes that [ChatGPT search typically rewrites a question](https://help.openai.com/en/articles/9237897-chatgpt-search) into targeted queries for its search partners. Neither company explains how a particular pipeline or monitoring tool gets picked.

**What studies have found.** Independent sources dominate: on the US software questions examined by [Chen and colleagues](https://arxiv.org/abs/2509.08919), AI search drew 72.7% of its sources from earned sites, against 45.4% for Google. Engineering forums count too: [our Reddit study](https://underneath.agency/research/ai-reddit-citations-study) found 35.4% of Google AI Overviews on B2B software searches cited a Reddit thread, and specialist communities supplied 58.7% of the Reddit threads shown in software searches. And ChatGPT went hunting for reviews or ratings in 46.2% of its answers in [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study).

**What engineers trust.** These factors are specific to the field:

- **Documentation they can run.** Documentation is the top learning resource for 68% of developers. Complete, current, example-rich docs are what both an engineer and an agent read before choosing.
- **Peer recommendation.** In the CNCF radar, 91% of developers said they would recommend GitHub Actions to peers. Engineers repeat what peers endorse in forums, talks and repositories.
- **Open standards and portability.** Support for standards such as OpenTelemetry signals that a buyer is not locked in.
- **Skepticism of claims.** 46% of developers distrust AI accuracy and 75% said not trusting AI’s answers is the top reason they would still ask a person for help. Marketing claims that cannot be tested are discounted twice.

**Our inference.** A DevOps vendor’s strongest evidence is technical and public: docs, code, benchmarks, changelogs, status pages and community answers. We would expect a tool with deep, accurate public docs and real community use to give AI answers more to work with, though nobody has tested that on DevOps vendors.

## What happens to a DevOps vendor that AI answers skip?

It loses trials it never sees, because engineers try whatever the answer names and move on.

- **Silent losses.** Product-led funnels start without a form fill. If an assistant names three tools and not yours, the trial never happens and nothing records the loss.
- **Agents that can’t read you.** If documentation is hard to parse, an agent may fall back on other sources or other tools. We describe this in [what happens when AI agents can’t read your site](https://underneath.agency/resources/when-ai-agents-cant-read-your-site).
- **A single mention can vanish.** [Our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study) repeated each question to ChatGPT five times, and only 25.2% of the brands ChatGPT named were there every time, a warning for any pipeline tool counting on one good answer.
- **Expansion depends on adoption.** Fewer teams adopting today means fewer accounts that grow into six- and seven-figure contracts later, we infer.

## What does GEO look like when your buyers read docs before marketing?

It makes your docs, code and benchmarks easy for engineers and agents to find, with no guarantee of a recommendation.

1. **Treat documentation as the front door.** Keep it public, complete, versioned and fast. Give clear install steps, limits and pricing. Offering clean text versions can help agents; in [our llms.txt study](https://underneath.agency/research/llms-txt-adoption-study), 11.5% of top sites published a valid file, but no study yet shows that AI answers use it. See [whether agents recommend readable websites](https://underneath.agency/resources/do-ai-agents-recommend-readable-websites).
2. **Ship working examples.** Maintain example repositories, templates and integration guides on GitHub for the stacks buyers use. Data tools face the same stack question; see [how a database gets picked for the stack](https://underneath.agency/resources/database-companies-developers-ai-search).
3. **Write honest comparison and migration guides.** Engineers ask “X or Y” and “how do I move from Z.” Answer with real trade-offs; see [whether comparison pages help](https://underneath.agency/resources/do-comparison-pages-help-b2b-ai-citations).
4. **Earn community and third-party proof.** Conference talks, foundation projects, independent benchmarks and technical publications carry more weight than your own pages. Take part in communities honestly; planted posts backfire, as we explain in [legitimate GEO versus manipulation](https://underneath.agency/resources/legitimate-geo-vs-manipulation).
5. **Keep facts consistent.** Pricing, limits and compliance status should match across docs, pricing pages and marketplaces; see [how to fix wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers).
6. **Watch both assistants and agents.** Ask your tooling questions repeatedly in ChatGPT, Google AI Overviews and AI Mode, Gemini, Perplexity, Copilot and Claude, and check your docs logs for agent visits; [how many prompts to track](https://underneath.agency/resources/how-many-prompts-to-track-ai-visibility) sets the sample size. The wider software approach, without the developer-first twist, is in [B2B SaaS revenue from AI search](https://underneath.agency/resources/b2b-saas-revenue-from-ai-search).

## What is still unknown about AI search and DevOps tool adoption?

No one has measured whether AI mentions lift a DevOps vendor’s pipeline or its net retention.

- **Few DevOps-specific studies.** Most AI search research covers software in general, not CI/CD or observability. Engineering software has the same gap, with no public survey of how engineers choose CAD, simulation or PLM tools, as our guide to [CAD, simulation and PLM software](https://underneath.agency/resources/engineering-software-ai-search) notes.
- **Vendor data.** Mintlify, Vercel and GitLab report their own numbers, and Stack Overflow sells advertising to tool vendors.
- **Agents are new.** How coding agents choose tools, and how stable those choices are, is barely studied.
- **Revenue effects remain unproven for tools engineers adopt.** [Does AI visibility drive business results](https://underneath.agency/resources/does-ai-visibility-drive-business-results) weighs the general evidence.

## How can a DevOps vendor find out whether AI answers are losing it trials?

Ask the questions engineers ask, read what assistants and coding agents return, then fix the docs behind it.

Start from what a platform engineer or SRE would ask at each stage: category, alternatives, comparisons, migration, standards and pricing. Put each question to ChatGPT, Gemini, Claude and the other assistants several times, and point a coding agent at your docs to see what it can extract. Note which tools come back, which sources are cited, and whether your rate limits, pricing tiers and integrations are stated correctly.

We can do this with your team: [talk to us about a docs and AI visibility review](https://underneath.agency/contact). It shows where AI answers put you in front of engineers, and which gaps in your documentation and third-party proof are most likely costing you free-tier signups, team adoption and platform contracts. Our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page describes how the ongoing work runs for engineer-led products, from documentation and example repositories to repeated checks in assistants and coding agents.

## Frequently asked questions

### Do engineers trust AI tool recommendations?

Not fully. 46% of developers distrust AI accuracy, so they check docs and peers before adopting a tool.

### Is documentation more important than marketing pages for AI visibility?

Probably, for DevOps. Documentation is developers’ top learning resource (68%), and agent reads of docs are growing fast on Mintlify-hosted sites.

### Does Reddit matter for DevOps tools?

For Google’s AI Overviews, clearly: 35.4% of those on B2B software searches in our study cited a Reddit thread.

### How long before AI visibility shows up in revenue?

Trials can start within days; enterprise contracts follow team adoption, which can take quarters.

### Should we publish an llms.txt file?

It is cheap, but unproven: 11.5% of top sites publish one, and no study shows AI answers use it.

## Sources

- Stack Overflow (2025-07), [2025 Developer Survey: Work](https://survey.stackoverflow.co/2025/work)
- Stack Overflow (2025-07), [2025 Developer Survey: AI](https://survey.stackoverflow.co/2025/ai)
- Stack Overflow (2025-07-29), [Stack Overflow’s 2025 Developer Survey Reveals Trust in AI at an All Time Low](https://stackoverflow.co/company/press/archive/stack-overflow-2025-developer-survey/)
- Google Cloud (2025-09), [Announcing the 2025 DORA Report](https://cloud.google.com/blog/products/ai-machine-learning/announcing-the-2025-dora-report)
- GitLab (2026-03), [GitLab Reports Fourth Quarter and Full Fiscal Year 2026 Financial Results](https://s204.q4cdn.com/984476563/files/doc_financials/2026/q4/Gitlab-4Q26-Earnings-Press-Release.pdf)
- CNCF and SlashData (2026-03-24), [Platform engineering tools maturing as organizations prepare for AI-driven infrastructure](https://www.cncf.io/announcements/2026/03/24/cncf-and-slashdata-report-finds-platform-engineering-tools-maturing-as-organizations-prepare-for-ai-driven-infrastructure/)
- CNCF and SlashData (2026-03-24), [Cloud native community reaches nearly 20 million developers](https://www.cncf.io/announcements/2026/03/24/cncf-and-slashdata-report-finds-cloud-native-community-reaches-nearly-20-million-developers/)
- CNCF (2026-05-21), [CNCF Announces OpenTelemetry’s Graduation](https://www.cncf.io/announcements/2026/05/21/cloud-native-computing-foundation-announces-opentelemetrys-graduation-solidifying-status-as-the-de-facto-observability-standard/)
- Vercel (2025-06-10), [How we’re adapting SEO for LLMs and AI search](https://vercel.com/blog/how-were-adapting-seo-for-llms-and-ai-search)
- Mintlify (2026), [Agent traffic data](https://www.mintlify.com/data)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- OpenAI Help Center (2025), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Chen, Wang, Chen and Koudas (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919), arXiv:2509.08919.
- Underneath (2026), [When does a Reddit thread become evidence in Google’s AI?](https://underneath.agency/research/ai-reddit-citations-study)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)
- Underneath (2026), [How many websites have an llms.txt file? 2026 adoption data](https://underneath.agency/research/llms-txt-adoption-study)

---

This is the Markdown twin of https://underneath.agency/resources/devops-platforms-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How to tell whether GEO really improved your AI citations"
description: "Only with repeated measurement before and after, a margin of error, and an untreated comparison. A 3-point gain in AI citation share can be pure noise."
canonical: "https://underneath.agency/resources/did-geo-improve-ai-citations"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How do we know whether a GEO campaign actually improved our AI citations?

You know only if the gain beats the normal swing in AI answers, lasts for weeks, and is not matched where you changed nothing. A single before-and-after snapshot cannot tell you. In one study’s worked example, a rise in citation share from 8% to 11% on OpenAI’s search model would sit within ordinary noise.

## The short version

1. On OpenAI’s search model, a rise in citation share from 8% to 11% would fall within the usual margin of error, so it would prove nothing (Sielinski, 2026).
2. Across three engines and three topics, websites whose citation shares differed by less than 5 to 7 points usually could not be told apart (Sielinski, 2026).
3. With nothing changed, the top cited website on one topic swung by a factor of nearly 2 in nine days (Sielinski, 2026).
4. In one tracking platform’s data, brands that did nothing drifted down about 1.34 points per run on ChatGPT, so “before” is not a fixed line (Kumar, 2026).
5. A review of 45 studies found no technique with a proven lasting effect on organic AI visibility across engines (Martinez, 2026).

## Can a before-and-after comparison prove a GEO campaign worked?

Not on its own. AI answers change between runs, so two snapshots mostly measure noise.

Generative engine optimization, or GEO, means changing your content and presence so AI engines name and cite you more. The trouble is that the same question asked twice gives different sources. [Sielinski](https://arxiv.org/abs/2603.08924) sent the same queries to Gemini, Perplexity and OpenAI’s search model every day for nine days. He also checked that the cited pages themselves had not changed. The swings came mostly from the engines, not from edits to the web.

His conclusion is blunt. If your share of OpenAI’s citations rises from 8% to 11% after a campaign, the gain cannot be attributed to the campaign with confidence. The 3-point rise sits inside the typical margin of error for a website on that engine. Only larger effects, or effects confirmed by repeated measurement before and after, stand out. No real campaign was tested here, the topics were consumer products, and the author works for a company that sells AI visibility measurement.

## How big does a change need to be before it counts?

Bigger than most people expect. Gaps of a few points between two measurements are often within the noise.

Sielinski gives an example from running gear. With 200 queries, Tom’s Guide had about 9.5% of citations on OpenAI’s search model and Runner’s World about 6.0%. That looks like a clear lead. But the plausible ranges for the two sites overlapped almost entirely. Across his platforms and topics, sites that seemed to differ by less than 5 to 7 points usually could not be separated.

The same problem applies to brand mentions. In [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), a brand named in 3 of 5 runs could truly appear anywhere from 23.1% to 88.2% of the time. A move from 2 of 5 to 3 of 5 tells you almost nothing. The same caution applies to [a one-time AI visibility report](https://underneath.agency/resources/one-time-ai-visibility-report).

## Can a change in rankings prove progress?

Rarely. A cited website’s rank is itself uncertain, even among the most-cited sites.

In a second dataset of 10 topics, [Sielinski](https://arxiv.org/abs/2607.10341) measured how much each site’s rank could plausibly vary. Even among the ten most-cited websites, the typical range spanned 5.0 rank positions, and 18.7% of them had a range wider than ten positions. A site moving from rank 10 to rank 20 between two periods may have lost ground, or may simply be bouncing within that noise.

## Why does timing matter so much?

Because AI answers drift on their own, within hours and over weeks. A change you see may have happened without you.

Several studies show movement with no campaign at all:

- **Over days.** On OpenAI’s search model, the most-cited website for multivitamins swung by a factor of nearly 2 within nine days ([Sielinski](https://arxiv.org/abs/2603.08924)).
- **Day to day.** In four Swiss industries, about 65% of cited sources changed from one day to the next. Even a 14-day window was only enough to show direction, not to compare brands finely ([Schulte and colleagues](https://arxiv.org/abs/2604.07585)).
- **Within hours.** In our consistency study, a ChatGPT answer to the same question 4.4 hours later overlapped less with the earlier answers than they did with each other.
- **Slow decline.** On one tracking platform, brands that took no action lost about 1.34 points of visibility per run on ChatGPT ([Kumar](https://arxiv.org/abs/2606.20065)). The author co-founded the platform, and some of the drift may come from prompt changes.

The last point cuts both ways. If visibility was drifting down, holding steady after a campaign may itself be a gain. A simple before-and-after would miss it.

## What does a credible test of a GEO campaign look like?

One that compares changed pages or prompts with similar ones left alone, measured repeatedly over the same weeks.

[Martinez](https://arxiv.org/abs/2607.14035) reviewed 45 studies and recommends a design any team can borrow. Fix the measure and the threshold for success before you start. Measure an untreated baseline. Ask each question several times and in several wordings, across named engines and dates. Where you can, assign the change at random to some pages or topics and not others.

A control group matters because AI traffic is growing everywhere. In a field study by [Watanabe and Nakayashiki](https://arxiv.org/abs/2606.04362), pages that received no changes still grew their ChatGPT referrals 3.5 times over the same period. Without that comparison, the whole rise would have been credited to the optimization. We list the other rival explanations in [whether GEO work raised your visibility](https://underneath.agency/resources/did-geo-work-raise-ai-visibility).

It also helps to watch the right unit. On Kumar’s platform, 77.5% of brand, prompt and engine combinations were either always or never mentioned across runs. A prompt that moves from never mentioning you to always mentioning you is a stronger signal than a small shift in an average.

## What should you do about it?

Treat every GEO campaign as an experiment, and design the measurement before the work starts. In practice:

1. **Write down the success measure and the minimum gain** that would count, before any change goes live.
2. **Collect a baseline over several weeks,** with each prompt asked several times, so you know the normal swing.
3. **Hold back a control.** Leave comparable pages, products or prompt topics untouched, and track them alongside the ones you change.
4. **Measure after the change for as long as before,** using two-to-four-week averages rather than daily readings.
5. **Report ranges, not points.** “Up from 8% to 11%” means little without the margin of error around each figure.
6. **Check each engine separately.** A gain on Perplexity says nothing about ChatGPT. Our guide to [measuring share of citations](https://underneath.agency/resources/measuring-share-of-citations-in-ai-search) covers the setup.
7. **Be suspicious of any result that only one snapshot supports,** including a vendor’s or an agency’s.

If you want help setting up a test like this, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The research has not shown which GEO techniques reliably raise AI citations in live engines over time. Key gaps:

- **Real campaign tests.** The 8%-to-11% example is a worked illustration, not a tested campaign.
- **Lasting effects.** Martinez found no technique with a stable, long-term effect on organic visibility across engines.
- **The right threshold.** How big a gain must be depends on engine, topic and sample size, and no general table exists yet.
- **Competitor effects.** If rivals optimize too, your share can fall even when your own work succeeds. Few studies measure this.
- **Independent replication.** Several of the key studies come from companies that sell measurement or from the site being studied.

## Frequently asked questions

### How long should a GEO test run before we judge it?

At least several weeks before and after the change. In the Swiss study, even a 14-day window only showed direction, and the authors recommend rolling averages over two to four weeks.

### Is a 3-point increase in AI citation share meaningful?

Often not. On OpenAI’s search model, a rise from 8% to 11% fell within the typical margin of error for a single website.

### What is a control group in GEO?

It is a set of pages, products or prompts you deliberately leave unchanged and measure alongside the ones you optimize. In one field study, unchanged pages still grew their ChatGPT referrals 3.5 times, which shows why the comparison matters.

### Can a vendor dashboard prove our GEO worked?

Only if it reports repeated runs, ranges around each figure and an untreated comparison. Rankings alone are weak evidence: even top-ten websites had rank ranges spanning 5.0 positions.

## Sources

- Sielinski, R. (2026), [Quantifying Uncertainty in AI Visibility: A Statistical Framework for Generative Search Measurement](https://arxiv.org/abs/2603.08924), arXiv:2603.08924.
- Sielinski, R. (2026), [From Stochastic to Stable: Rank Stability and Structural Sufficiency in AI Visibility Measurement](https://arxiv.org/abs/2607.10341), arXiv:2607.10341.
- Schulte, J., Bleeker, M. and Kaufmann, P. (2026), [Don’t Measure Once: Measuring Visibility in AI Search (GEO)](https://arxiv.org/abs/2604.07585), arXiv:2604.07585.
- Kumar, P. (2026), [Generative Engine Optimization at Scale: Measuring Brand Visibility Across AI Search Engines](https://arxiv.org/abs/2606.20065), arXiv:2606.20065.
- Martinez, O. (2026), [Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026)](https://arxiv.org/abs/2607.14035), arXiv:2607.14035.
- Watanabe, K. and Nakayashiki, K. (2026), [Disentangling Answer Engine Optimization from Platform Growth: A Log-Based Natural Experiment on ChatGPT Referral Traffic](https://arxiv.org/abs/2606.04362), arXiv:2606.04362.
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)

---

This is the Markdown twin of https://underneath.agency/resources/did-geo-improve-ai-citations. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Our AI visibility went up: was it because of our GEO work?"
description: "Maybe. A rise can also come from chance, AI platform growth, engine updates, competitors or a changed prompt set, so you need a baseline and a control."
canonical: "https://underneath.agency/resources/did-geo-work-raise-ai-visibility"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Our AI visibility went up: was it because of our GEO work?

You cannot tell from the rise alone. Your own work is only one of at least four things that move an AI visibility score, alongside competitors’ changes, updates to the AI engines and changes to the set of questions being tracked. Ordinary run-to-run chance can produce swings of several points on top of that. Before you credit or blame a GEO program, you need a baseline measured the same way and a comparison group that did not get the work.

## The short version

1. A survey of GEO measurement lists at least four causes of a score change besides wording and scoring: your work, competitors’ work, engine drift and a change in the question set ([Martinez, 2026](https://arxiv.org/abs/2609.06811)).
2. On one website, ChatGPT referrals to optimized pages grew 6.1 times, but untouched pages on the same site grew 3.5 times; the part plausibly due to the work was about 1.8 to 2.3 times ([Watanabe and Nakayashiki, 2026](https://arxiv.org/abs/2606.04362)).
3. In one study, a rise in a site’s ChatGPT search citation share from 8% to 11% was within normal noise and could not be credited to any change ([Sielinski, 2026](https://arxiv.org/abs/2603.08924)).
4. A review of 45 GEO studies found no technique with a proven, lasting effect across AI platforms on being found organically ([Martinez, 2026](https://arxiv.org/abs/2607.14035)).

## What else could have raised our AI visibility?

At least four things besides your own work. [Martinez](https://arxiv.org/abs/2609.06811), in a critical survey of how GEO visibility is measured, lists them: your own intervention, competitors’ interventions, changes in the AI engine itself, and a change in the mix of questions being tracked. Changes in how questions are worded or answers are scored come on top.

The survey is explicit about the limit. A publisher sees a change in its score over time, not the effect of its own work. Comparing two dates, without recording what competitors did and which engine version answered, mixes all of these causes together. Even fixing the prompts does not remove the effect of competitors, because their pages change what the engine finds.

The last cause is easy to miss. If your tracking tool added questions, dropped some or reweighted them, the score can move even if no AI answer changed. Ask your vendor for the question list behind both readings.

## How much of a rise can be pure noise?

More than most dashboards admit: several percentage points, sometimes more. [Sielinski](https://arxiv.org/abs/2603.08924) asked three AI search engines the same 200 questions per topic every day for nine days. On ChatGPT search, a rise in a site’s citation share from 8% to 11% fell within the normal range of noise and could not be credited to an intervention.

Single readings can mislead badly. On one day, Gemini gave nationalgeographic.com a citation share of 0.032 for bird feeder questions; across the nine days its average was 0.005. Yahoo’s share for multivitamin questions on ChatGPT search ranged from 0.092 to 0.160 within the window, nearly a twofold swing for the top site. The author checked that the cited pages themselves had mostly not changed. That is why [a one-time visibility report](https://underneath.agency/resources/one-time-ai-visibility-report) needs repeated runs behind it.

Time alone moves answers too. In our [consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), the same ChatGPT question asked about 4.4 hours later overlapped with the earlier runs by 0.477, against 0.588 between runs minutes apart. Sielinski is affiliated with the company IQRush; the consistency study is our own.

## How much of the growth is just AI platforms growing?

Possibly most of it, if you measure AI referral traffic. [Watanabe and Nakayashiki](https://arxiv.org/abs/2606.04362) of Glasp studied their own site after a bundle of changes to one section in January 2026. The rest of the site was left alone and served as a comparison group.

Total ChatGPT referrals to the site grew 5.7 times on monthly figures. The optimized pages grew 6.1 times, but the untouched pages grew 3.5 times with no work at all. Measured from the lowest day to the single best day, the rise looked like 38.2 times, a figure the authors call a fragile basis for any claim.

Their best estimate of the effect of the work was a jump of about 1.8 to 2.3 times. Even that is only suggestive: a stricter check found jumps of similar size at made-up dates before the change. The authors suspect many published success stories are inflated the same way. This is one site, one engine and a bundle of changes, written up by the company that made them. For the traffic side in more detail, see [crediting ChatGPT referral growth](https://underneath.agency/resources/chatgpt-referral-growth-and-geo).

## Can competitors or engine updates move your score?

Yes, and without any change on your side. Citation share is a zero-sum measure: when one source gains share, another loses it. A [review of 45 GEO studies](https://arxiv.org/abs/2607.14035) notes that tested gains can shrink as more competitors adopt the same tactics, and that a before-and-after comparison without a matched baseline confuses the treatment with chance.

Engines also drift. Ranqo, which sells AI visibility tracking, followed brands that did little or nothing to change their visibility. On ChatGPT, their visibility for questions that did not name them slipped by about 1.34% per tracking run ([Kumar, 2026](https://arxiv.org/abs/2606.20065)). The author notes some of that may reflect changes to the question set, not the engine.

Short-term turnover is large. In a test of 15 commercial prompts, the exact pages an engine cited for the same prompt changed by 67.0% from one day to the next, on average across three engines ([Tannenbaum, 2026](https://arxiv.org/abs/2609.22655)). The author, who is affiliated with Aiso Boost Ltd., warns the figure may partly reflect changes in his own collection system.

## How can you tell whether your GEO work caused the change?

Compare against something that did not get the work, measured the same way at the same time. The studies above point to a few designs that are practical for a marketing team.

| Design | What it rules out | Source |
|---|---|---|
| Untouched pages or products on the same site as a comparison group | Platform growth and broad engine changes | Watanabe and Nakayashiki |
| Randomly holding back some pages from the changes | Most other causes, if the groups are large enough | Watanabe and Nakayashiki (proposed, not run) |
| A frozen question set measured before and after | Changes in the question mix | Martinez |
| Many runs before and after, reported with ranges | Run-to-run chance | Sielinski |
| Long enough windows | Day-to-day swings | Schulte and colleagues |

On window length, a Swiss study of four AI engines found per-brand estimates needed about 24 days of daily runs to become reasonably stable ([Schulte, Bleeker and Kaufmann, 2026](https://arxiv.org/abs/2604.07585)). Our guide to [testing whether GEO improved citations](https://underneath.agency/resources/did-geo-improve-ai-citations) shows how to set the threshold for a real gain.

## What should you do about it?

Set up the comparison before you start the work, not after the number moves. Steps:

1. Freeze your question list, assistants, locations and run counts before the program starts, and keep a version history of any change.
2. Measure a baseline of at least several weeks, with repeated runs, so you know your normal range.
3. Keep a comparison group: pages, products or topics you deliberately leave alone. Where you can, choose them at random.
4. Track the competitors named alongside you, and note known engine updates on your timeline.
5. Report the change as a range, and give the effect on treated items minus the change in the comparison group.
6. For traffic, compare AI referrals to treated pages with untreated pages on the same site, as in the Glasp study.

If you want a program measured this way from the start, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The research shows how a rise can mislead, but not yet how much GEO work reliably delivers. Gaps:

- The review of 45 studies found that content changes can affect answers once a page is already retrieved. It found no technique with a proven, lasting, cross-platform effect on being found or on clicks and sales.
- The best traffic evidence comes from one site, one engine and one bundle of changes, reported by the site’s owner.
- No study has yet split an observed change into its four causes with real data. The survey that names them offers a framework, not measurements.
- Most noise figures come from consumer product topics over days or weeks. Long-run engine drift through major model updates is barely measured.

## Frequently asked questions

### How do I know if GEO improved our AI visibility?

Compare the change on the pages or topics you worked on with a similar group you left alone, measured with the same questions and run counts. Without that comparison, platform growth and noise can look like success.

### Can AI visibility change without us doing anything?

Yes. On one site, ChatGPT referrals to pages nobody had touched grew 3.5 times from January to May 2026, which the authors put down to the growth of ChatGPT itself.

### How long should we measure before judging GEO results?

At least several weeks. One study found per-brand estimates took about 24 days of daily runs to become reasonably stable, and short windows mostly show noise.

### Can competitors lower our AI visibility?

Yes. Citation share is zero-sum within an answer, so competitors’ gains can lower your share even if your content and the engine did not change.

## Sources

- Martinez (2026), [Measuring GEO Visibility: Prompt Corpora Define the Answer Market](https://arxiv.org/abs/2609.06811), arXiv:2609.06811.
- Martinez (2026), [Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026)](https://arxiv.org/abs/2607.14035), arXiv:2607.14035.
- Watanabe and Nakayashiki (2026), [Disentangling Answer Engine Optimization from Platform Growth: A Log-Based Natural Experiment on ChatGPT Referral Traffic](https://arxiv.org/abs/2606.04362), arXiv:2606.04362.
- Sielinski (2026), [Quantifying Uncertainty in AI Visibility: A Statistical Framework for Generative Search Measurement](https://arxiv.org/abs/2603.08924), arXiv:2603.08924.
- Kumar (2026), [Generative Engine Optimization at Scale: Measuring Brand Visibility Across AI Search Engines](https://arxiv.org/abs/2606.20065), arXiv:2606.20065.
- Tannenbaum (2026), [Scoring With the Engine: Retrieval Exposure, Cross-Engine Divergence, and the Limits of Engine-Agnostic GEO Scores](https://arxiv.org/abs/2609.22655), arXiv:2609.22655.
- Schulte, Bleeker and Kaufmann (2026), [Don’t Measure Once: Measuring Visibility in AI Search (GEO)](https://arxiv.org/abs/2604.07585), arXiv:2604.07585.
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)

---

This is the Markdown twin of https://underneath.agency/resources/did-geo-work-raise-ai-visibility. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How digital banks win depositors when savers ask AI"
description: "By making rates, conditions and deposit insurance facts current and checkable, because savers now ask AI where to put money and AI often gets rates wrong."
canonical: "https://underneath.agency/resources/digital-banks-depositors-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How does a digital bank win depositors when savers ask AI where to put their money?

By being named for the right products and quoted correctly when savers ask which account pays the most and whether their money is safe. Savings is the money topic people ask AI about most often, and a deposit customer is worth years of low-cost funding. Rates change often and AI answers can lag them, so a digital bank wins in AI search by publishing current, plain and verifiable facts about rates, conditions, fees and deposit insurance.

## The short version

1. Savings is the top money question for AI: in a [J.D. Power survey reported by the ABA](https://bankingjournal.aba.com/2025/09/survey-consumers-increasingly-turn-to-ai-for-financial-advice/), savings strategies were the most popular money topic people asked AI about (45%), and 59% of respondents used AI for banking and financial services occasionally.
2. The rate gap drives the search: the [FDIC’s](https://www.fdic.gov/national-rates-and-rate-caps/national-rates-and-rate-caps-june-2026) national average savings rate was 0.38% in June 2026, while [Ally](https://mmx.prnewswire.com/media/MS1886857/Ally-2Q-2026-Press-Release.pdf) reported an average retail deposit yield of 3.12%.
3. Depositors are worth a lot: Ally holds $144 billion of retail deposits from 3.6 million deposit customers, and [SoFi](https://pulse2.com/sofi-sees-43-revenue-increase-for-q2-2026-to-record-1-2-billion-as-net-income-reaches-157-million/) says its deposits cost 156 basis points less than its warehouse funding, worth about $712.6 million a year.
4. Savers move: in a [Credit One Bank survey](https://keyt.com/?p=1447003) of 1,000 US savers in April 2026, 48% had moved savings to another bank in the previous 12 months.
5. AI gets rates wrong: when [Which?](https://www.which.co.uk/news/article/can-ai-really-find-you-the-best-savings-deal-aAy2E6f4Elky) tested ChatGPT in the UK, it quoted Marcus at 4.5% when the account paid 3.75%, and gave an out-of-date deposit protection limit.

This guide covers how digital banks are found and described in AI answers. It is not financial, legal or regulatory advice; check rate, fee and deposit insurance wording with your compliance team.

## Who chooses a digital bank, and what is a depositor worth?

Rate-aware savers, younger adults and people at life events choose them; each depositor brings years of low-cost funding.

By digital bank we mean a chartered bank that serves customers mainly online: Ally Bank, SoFi Bank, Marcus by Goldman Sachs, Charles Schwab Bank and the online savings brands of large banks. App brands that rely on a partner bank are a different case, covered in [our neighboring guide on neobanks](https://underneath.agency/resources/neobanks-account-openings-ai-search).

Who switches tells you who to reach. In the Credit One survey, roughly two-thirds of Gen Z respondents had moved some savings to another bank for a higher rate, against fewer than 3 in 10 baby boomers. Expecting parents were the most active rate shoppers: nearly 7 in 10 had moved savings in the past year. At Ally, millennials and younger customers remain the largest generation among new customers.

The value of a depositor shows in the banks’ results:

| Bank | Latest public figures | Why deposits matter |
|---|---|---|
| Ally | $144 billion of retail deposits, 92% FDIC insured; 3.6 million deposit customers; 63 thousand net new in the second quarter of 2026 | Deposits made up 87% of Ally’s funding |
| SoFi | $45.5 billion of deposits, up $5.3 billion in the quarter; 15.8 million members | Deposits replace costlier credit lines, and existing members opened 51% of new products |

The revenue path, we infer, is two-step: a depositor first lowers the bank’s funding cost, then becomes a customer for loans, cards or investing. SoFi says it aims to acquire members through one product and then sell them others, “potentially lowering customer acquisition costs and increasing lifetime value.” Banks that also offer investing can see [how investing apps get chosen in AI answers](https://underneath.agency/resources/investment-platforms-customers-ai-search).

## Where does AI enter a saver’s decision?

Early, and often: savings questions lead the money topics people take to AI assistants.

- **United States.** J.D. Power found that thirteen percent of respondents use AI for banking and financial services daily, and another 59% use it occasionally. ChatGPT was the most popular tool, particularly for people under 40; Google Gemini was preferred by people over 40.
- **United Kingdom, as a signal.** Which? cites a Lloyds Bank study finding 53% of UK adults use AI tools to help with savings. In [STRAT7’s April 2026 research](https://strat7.com/?p=48569), 55% of UK adults use AI platforms for financial questions.

Trust is the gap. In STRAT7’s study, ChatGPT scored 53% on trust for financial information, against 92% for bank websites. A reasonable expectation is that savers use AI to build a shortlist and explain options, then go to the bank’s own site to confirm the rate and the protection before they open an account. That makes the bank’s own pages part of the AI journey, not separate from it.

Many savers have already crossed the online line. In a 2024 [Bankrate survey reported by InvestmentNews](https://investmentnews.com/industry-news/news/two-thirds-of-us-savers-are-leaving-interest-on-the-table-251518), 51 percent of Americans had a savings or money market account with an online bank. Those who did not use one cited a preference for local branches (45 percent), satisfaction with their current institution (42 percent) and concerns over the security of their money (32 percent). Each of those objections is a question a saver can now put to an assistant.

## Which questions do savers ask AI before opening an account?

Questions about the best rate, the conditions behind it, safety and fees. We wrote the examples below to illustrate; they are not observed prompts.

| Need | Illustrative prompt |
|---|---|
| Best rate | “Which online savings accounts pay the highest rate right now with no minimum balance?” |
| Comparison | “Ally vs Marcus vs SoFi for an emergency fund” |
| Lump sum | “Where should I put $25,000 I won’t need for a year: savings account or CD?” |
| Safety | “Is my money safe in an online-only bank, and is it FDIC insured?” |
| Conditions | “Does this bank’s savings rate include a temporary bonus, and what happens after it ends?” |
| Fees | “Online checking accounts with no monthly fees and early direct deposit” |
| Switching | “My bank cut my savings rate. Is it worth moving my money?” |

The switching prompt matters most. In the Credit One survey, if a bank cut its rate by a full percentage point, fewer than 2 in 10 savers would switch immediately, but two-thirds would use the cut as a trigger to shop around. Just over half said eliminating monthly fees would prompt them to switch even at the same rate. Each rate move, we infer, sends a wave of savers back to search and AI to compare.

## How does a saver get from an AI answer to a funded deposit account?

Through a short chain that breaks easily: named, checked, opened, funded, kept.

1. **Named.** A saver asks for the best options and the answer lists a few banks, usually with rates.
2. **Checked.** The saver confirms the rate, the conditions and deposit insurance, often on the bank’s own site.
3. **Opened.** The account is opened online in minutes.
4. **Funded.** Money moves in. This is the step that creates value for the bank.
5. **Kept.** The depositor stays until a better offer appears, and may add loans or investing later.

An AI answer can break the chain at step two. If it quotes a rate higher than the one on offer, the saver arrives expecting a deal the bank does not offer, and may leave disappointed. If it quotes a rate that is too low or out of date, the bank may never be considered. In both cases the bank sees nothing in its analytics. We suggest tracking AI visibility against funded accounts, not clicks, and asking new customers during account opening where they first heard of the bank. The measurement problem is covered in [what to measure in AI visibility](https://underneath.agency/resources/what-to-measure-ai-visibility).

## Why do rates and deposit facts go wrong in AI answers?

Because rates change faster than many pages that AI answers draw on, and conditions are easy to drop.

**Observed by a consumer group.** On 10 March 2026, Which? asked the free version of ChatGPT about UK savings and checked the answers against Moneyfacts data. Most average and top rates were wrong; only regular saver rates were right. ChatGPT said Marcus offered 4.5%, when its deal paid 3.75%. It suggested a Chase account at 4.5% without saying the rate drops to 2.25% once a temporary bonus ends. It also gave the old UK deposit protection limit of £85,000, though the limit had risen to £120,000 on 1 December 2025, and answers varied when the question was repeated. OpenAI told Which? that for consumer products it recommends using ChatGPT’s built-in search tool, which shows sources. This was one journalist’s test in another country, but the failure modes apply to any bank that competes on rate.

**Observed in our studies.**

- **Prices are often quoted with a twist.** In [our pricing accuracy study](https://underneath.agency/research/ai-pricing-accuracy-study) of software plans, 61.9% of plan prices in AI answers were fully faithful to the vendor’s page and 7.6% differed for the same plan and term. Most differing figures came from another page on the vendor’s own site. A savings rate is a price, and the same risk applies, we infer, to old promotional pages.
- **Assistants lean to newer pages.** In [our freshness study](https://underneath.agency/research/ai-source-freshness-study), pages published in the last 90 days made up 17.4% to 22.6% of each assistant’s dated citations, against 6.9% of Google’s top 10.
- **Finance answers depend on the country.** In [our country study](https://underneath.agency/research/ai-recommendations-by-country-study), on questions such as insurance, lending and tax, the gap between same-country and different-country brand overlap was 0.180, against 0.034 for global products such as headphones. US savers should get US banks; that only works when a bank’s pages make its market and regulator plain.

## What decides whether a digital bank is named?

Independent evidence and clear, current facts; the platforms document how they search, not how they pick banks.

**Documented.** Google says AI Overviews and AI Mode [may use “query fan-out”](https://developers.google.com/search/docs/appearance/ai-features), issuing multiple related searches across subtopics. No assistant publishes how it ranks banks.

**Observed in our studies.** In [our frequency study](https://underneath.agency/research/ai-overviews-frequency-study), 88.0% of US financial services and insurance keywords triggered an AI Overview. Of everything [our brand entity study](https://underneath.agency/research/brand-entity-ai-recommendations-study) measured, coverage on independent sites predicted recommendations best: when the number of independent sites naming a brand in the cited pages rose tenfold, the brand had 4.7 times the odds of being recommended.

**Independent evidence banks already earn.** J.D. Power’s [2026 Direct Banking Satisfaction Study](https://www.01net.it/online-only-banking-providers-continue-to-win-over-customers-with-personalized-digital-experiences-but-some-struggle-with-customer-service-quality-jd-power-finds/) rated online bank high-yield savings at 689 out of 1,000, against 657 for neobanks. Marcus ranked highest among high-yield savings providers with 739, ahead of Ally at 728. Awards, ratings and editorial reviews of this kind are, we infer, exactly the third-party evidence an assistant can cite and a cautious saver trusts.

**Our inference on safety facts.** Deposit insurance is where a chartered digital bank has a clear story, and where precision matters. The FDIC’s official sign states that each depositor is insured up to at least $250,000. The FDIC has also moved the compliance date for showing its official digital sign on websites and apps to [January 1, 2027](https://www.fdic.gov/news/financial-institution-letters/2025/compliance-date-extension-sections-3284-and-3285-0), while it revises those rules. Banks that state the legal bank name behind each brand, the insurance facts and their conditions in plain text give assistants something accurate to repeat. Banks that bury them leave the answer to whatever page an assistant finds first.

## What does GEO look like for a digital bank?

Generative engine optimization (GEO) keeps your rates, conditions and protections easy for assistants to find, check and repeat.

For a digital bank, the work usually covers:

1. **Rate pages written for checking.** The current rate in text, the date it took effect, the conditions (minimums, bonus periods, caps) next to the number, and old promotional pages retired or updated.
2. **A clear entity.** The legal bank name behind each brand, its regulator and its FDIC insurance facts on every product page, consistent everywhere they appear.
3. **Comparison-site accuracy.** Correct, current listings on the rate tables and review sites that savers and assistants use. Many digital banks already pay these sites for referrals; the same listings now feed AI answers.
4. **Independent proof.** Satisfaction rankings, editorial reviews and press coverage that confirm service, not only rate.
5. **Life-event content.** Plain pages for savers with a lump sum, an emergency fund to build or a baby on the way, written without promising outcomes.
6. **Monitoring after every rate change.** Ask assistants your savers’ questions after each change and correct wrong facts at their source, following [how to fix wrong information about your brand in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers).

No one can guarantee that an assistant will name a bank. GEO makes sure that when it does, the rate, the conditions and the protections are right. For app-based finance more broadly, see [how consumer fintech apps win customers when people ask AI about money](https://underneath.agency/resources/consumer-fintech-apps-customers-ai-search). Vendors selling core, payments or fraud systems to banks face a different buyer; see [how banking technology vendors reach bank shortlists](https://underneath.agency/resources/banking-technology-enterprise-deals-ai-search).

## Which questions about AI and deposits remain open?

Several: how often savers open accounts from AI answers, and how fast answers catch up with rate changes.

- **No public link to funded accounts.** We found no public data connecting AI answers to account openings or balances at any bank. The cross-industry evidence, thin as it is, is gathered in [does AI visibility drive business results?](https://underneath.agency/resources/does-ai-visibility-drive-business-results).
- **Rate accuracy is measured once, abroad.** The Which? test was a single journalist’s check of one assistant in the UK. No comparable US test of savings rates has been published that we could find.
- **Survey sources have interests.** The Credit One survey was produced by a bank, and the UK figures come from a bank and a research agency. Treat them as signals.
- **Our studies are snapshots.** They cover US searches collected in September 2026; answers vary between assistants and over time.

## Where should a digital bank begin?

Begin by asking assistants the questions your savers ask, then check every rate, condition and insurance fact in the answers.

A useful first review covers best-rate questions for your products, head-to-head comparisons with the banks you compete with, safety questions about online banks, and the switching questions that follow a rate change. It shows whether you are named, which banks and comparison sites appear instead, and whether your rates and protections are stated correctly.

If your deposit growth depends on savers choosing you when they move money, [ask us to review how AI answers describe your bank](https://underneath.agency/contact). We will compare how assistants present you and your competitors, list the rates and facts they get wrong or miss, and plan the content, listings and coverage, reviewed with your compliance team, that keep those answers accurate after every rate change. Our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page shows how that work is organized for a bank, from rate-page upkeep and listing corrections to plain pages for savers.

## Frequently asked questions

### Do people really ask ChatGPT where to open a savings account?

Many ask AI about savings. In J.D. Power’s survey, savings strategies were the top money topic people asked AI about (45%), and 59% used AI for banking and financial services occasionally. No public data counts account openings from AI answers.

### Why would an AI answer show the wrong savings rate?

Rates change often, and answers can draw on old or promotional pages. In the Which? test, ChatGPT quoted a 4.5% rate for an account paying 3.75%.

### Does being a chartered bank help in AI answers?

It gives you a clear safety fact to state. Whether an assistant repeats it depends on whether your pages say it plainly and consistently.

### Can GEO make an assistant recommend our bank?

No. GEO makes accurate information about your bank easy to find and verify. It does not give personal financial advice, and its content should pass the same compliance review as your advertising.

## Sources

- ABA Banking Journal (2025-09), [Survey: Consumers increasingly turn to AI for financial advice](https://bankingjournal.aba.com/2025/09/survey-consumers-increasingly-turn-to-ai-for-financial-advice/)
- FDIC (2026-06), [National Rates and Rate Caps, June 2026](https://www.fdic.gov/national-rates-and-rate-caps/national-rates-and-rate-caps-june-2026)
- Ally Financial (2026-07-21), [Second quarter 2026 press release](https://mmx.prnewswire.com/media/MS1886857/Ally-2Q-2026-Press-Release.pdf)
- Pulse 2.0 (2026-07), [SoFi sees 43% revenue increase for Q2 2026](https://pulse2.com/sofi-sees-43-revenue-increase-for-q2-2026-to-record-1-2-billion-as-net-income-reaches-157-million/)
- Credit One Bank, via Stacker and KEYT (2026-07-15), [The 4% flashpoint: Why Americans are quietly dumping traditional bank accounts](https://keyt.com/?p=1447003)
- Which? (2026-03-26), [Can AI really find you the best savings deal?](https://www.which.co.uk/news/article/can-ai-really-find-you-the-best-savings-deal-aAy2E6f4Elky)
- STRAT7 (2026), [Who’s using AI for financial advice? New UK research](https://strat7.com/?p=48569)
- InvestmentNews (2024-03-27), [Two-thirds of US savers are leaving interest on the table](https://investmentnews.com/industry-news/news/two-thirds-of-us-savers-are-leaving-interest-on-the-table-251518)
- J.D. Power, via 01net (2026-04-29), [Online-only banking providers continue to win over customers with personalized digital experiences](https://www.01net.it/online-only-banking-providers-continue-to-win-over-customers-with-personalized-digital-experiences-but-some-struggle-with-customer-service-quality-jd-power-finds/)
- FDIC (2025-12), [Compliance date extension for sections 328.4 and 328.5 amendments](https://www.fdic.gov/news/financial-institution-letters/2025/compliance-date-extension-sections-3284-and-3285-0)
- Federal Register, FDIC (2024-01-18), [FDIC Official Signs and Advertising Requirements, False Advertising, Misrepresentation of Insured Status](https://www.federalregister.gov/documents/2024/01/18/2023-28629/fdic-official-signs-and-advertising-requirements-false-advertising-misrepresentation-of-insured)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- Underneath (2026), [When does Google show an AI Overview? 1,248 US searches](https://underneath.agency/research/ai-overviews-frequency-study)
- Underneath (2026), [How faithfully do AI assistants quote software prices?](https://underneath.agency/research/ai-pricing-accuracy-study)
- Underneath (2026), [How fresh are the pages AI engines cite?](https://underneath.agency/research/ai-source-freshness-study)
- Underneath (2026), [Same question, four countries: do AI recommendations change?](https://underneath.agency/research/ai-recommendations-by-country-study)
- Underneath (2026), [Do Wikipedia and schema make AI assistants recommend a brand?](https://underneath.agency/research/brand-entity-ai-recommendations-study)

---

This is the Markdown twin of https://underneath.agency/resources/digital-banks-depositors-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How do digital health apps win users when people ask AI first?"
description: "By giving assistants proof they can check: plain privacy terms, honest evidence and app store trust, now that 32% of US adults ask AI about health."
canonical: "https://underneath.agency/resources/digital-health-apps-users-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How do digital health apps win new users when people ask AI about their health first?

By being the app an assistant can describe accurately and a cautious person can verify: clear claims that stay inside the rules, plain privacy terms, published evidence and a strong app store reputation. One in three US adults now uses AI chatbots for health information, and most of them use general-purpose assistants rather than tools built by hospitals or insurers. For a consumer health app or a direct-to-consumer care program, that makes the assistant an early stop before the app store and the free trial.

This article is about how consumer digital health companies are found and chosen. It is not health advice, and nothing here says any app or program works.

## The short version

1. In [Rock Health’s survey of 8,000 US adults](https://fiercehealthcare.com/ai-and-machine-learning/ai-chatbot-use-health-information-16-2024-rock-health-survey), fielded in December 2025, 32% had used AI chatbots for health information, up from 16% a year earlier, and 64% of those users ask health questions weekly or more.
2. General assistants dominate: [74% of AI health users](https://hitconsultant.net/2026/03/23/rock-health-2025-survey-consumer-ai-adoption-chatgpt-healthcare/) turned to tools like ChatGPT, against 5% for chatbots offered by providers. The same people trust health apps far more than non-users do (55% vs. 25%).
3. The shelf is crowded: [IQVIA counts](https://drugstorenews.com/iqvia-releases-digital-health-trends-2024-report) 337,000 digital health apps, and US digital health startups raised [$7.4 billion across 244 deals](https://hitconsultant.net/2026/07/13/rock-health-h1-2026-digital-health-funding-report/) in the first half of 2026, with mental health and weight management drawing the most money.
4. Each new user is hard won: in [RevenueCat’s 2026 benchmarks](https://www.revenuecat.com/state-of-subscription-apps), the median health and fitness app converts 2.9% of downloads to paying users, and 68% of the category’s subscriptions are annual plans.
5. OpenAI now connects apps such as MyFitnessPal, Weight Watchers and Peloton inside [ChatGPT Health](https://openai.com/index/introducing-chatgpt-health/), which it says fields health questions from over 230 million people a week.

## Who are a digital health company’s customers, and what is one worth?

Mostly engaged adults who track their health and pay for an annual plan or a program, after a trial.

The consumer side of digital health covers wellness and fitness apps, mental health and sleep apps, weight management programs, condition management apps and wearable companions. Rock Health’s 2025 survey describes the people most likely to use AI for health as engaged trackers: 43% of AI users track sleep, against 28% of non-users, and they follow an average of four health metrics. They are the natural audience for health apps.

The commercial engine is a subscription or a paid program. RevenueCat’s 2026 report, which covers apps using its billing platform, gives the clearest public numbers for the health and fitness category:

| Health and fitness apps (RevenueCat 2026, medians) | Value |
|---|---|
| Downloads that become paying users | 2.9% |
| Trials that convert to paid | 37.7% |
| Trials started on the day of download | 82.1% |
| Share of subscriptions on annual plans | 68% |
| Realized lifetime value after one year, per payer | $35.64 |

Two things follow. First, the decision happens fast: most trials start the day the app is installed, so the research that matters happens before the download. Second, the category depends on annual commitment, which favors buyers who arrive already convinced. Our inference is that a user who was named a product by an assistant, checked it and then downloaded it is the kind of user that annual-plan economics need. No public data yet compares the lifetime value of users referred by AI with others.

Direct-to-consumer care programs, such as coaching and weight management programs, sell enrollment rather than an app subscription, at higher prices. We found no public benchmark for their customer value and will not invent one. Virtual care providers face their own questions on price and insurance; see [how telehealth companies get chosen](https://underneath.agency/resources/telehealth-patients-from-ai-search).

## Where do AI assistants sit in a health app user’s research?

At the start, for a growing third of adults, and mostly inside general assistants rather than health system tools.

Rock Health’s survey is the most direct evidence. Among US adults, ChatGPT was the most used chatbot for health information (23%), followed by Gemini (15%). [Fierce Healthcare’s summary](https://fiercehealthcare.com/ai-and-machine-learning/ai-chatbot-use-health-information-16-2024-rock-health-survey) notes that people use AI for questions to ask at appointments, for mental health needs and for “looking for specific providers and clinics.” After a chatbot answer, 42% searched for more information and 40% consulted a provider. Millennials (48%) and Gen Z (45%) lead adoption.

The platforms are building for this. When OpenAI launched ChatGPT Health in January 2026, it documented that people can connect Apple Health and wellness apps such as Function and MyFitnessPal, with Weight Watchers and Peloton among the listed connections. OpenAI says apps in Health must meet its privacy and security requirements and pass an additional security review, and that ChatGPT may suggest a connected app when helpful. It also states that Health “is not intended for diagnosis or treatment.” Those are statements documented by the platform; how often ChatGPT suggests a particular app is not public.

App discovery in general is shifting too. In [AppTweak’s survey](https://www.apptweak.com/en/aso-blog/ai-app-discovery-survey) of 1,000 US adults across all app categories, 10% first learned about their most recent download from an AI assistant, level with Google, and 59% still went to the store to check reviews, ratings or screenshots after an AI recommendation. We cover the wider app picture in [how consumer apps win subscribers from AI search](https://underneath.agency/resources/consumer-app-subscribers-from-ai-search).

## What do people ask AI assistants before they download a health app?

Which app fits a goal, how two programs compare, whether it protects data and whether it has evidence behind it.

We wrote these sample prompts ourselves to show how people shop for a health app or care program. None comes from a real user, and none asks for a diagnosis or treatment.

| Stage | Illustrative prompt |
|---|---|
| Category | “What’s a good sleep tracking app that works with my Apple Watch?” |
| Program choice | “Compare structured weight management programs that include coaching” |
| Alternatives | “Cheaper alternative to my meditation app with offline sessions” |
| Privacy | “Does this period tracking app share my data with advertisers?” |
| Evidence | “Which stress apps have published studies behind them?” |
| Cost and coverage | “Is this program covered by my employer or insurer?” |
| Legitimacy | “Is this app legit? What do reviews say about canceling?” |

Privacy questions deserve special weight in this industry. Rock Health found AI users are about twice as willing as non-users to share health data with health tech companies (23% vs. 11%), but the regulators are watching how apps handle it. The [Federal Trade Commission](https://www.ftc.gov/business-guidance/resources/complying-ftcs-health-breach-notification-rule-0) says its July 2024 amendments make clear that makers of health apps and connected devices must comply with the Health Breach Notification Rule, with civil penalties of up to $53,088 per violation. Our inference: an assistant asked whether an app protects data can only answer well if the company has said so clearly, in text, somewhere it can read.

## How does an AI answer turn into a paying user?

Through a short list, a store check, a same-day trial and an annual plan.

1. **Named for a goal.** The assistant lists a few apps or programs for “sleep tracking with a wearable” or “weight management with coaching.”
2. **Checked.** The person reads store ratings, privacy labels and reviews, or asks the assistant whether the company is legitimate. When [our brand reputation study](https://underneath.agency/research/is-it-legit-ai-reputation-study) asked assistants whether brands were legitimate, 88.0% of the answers cited a review or complaint platform.
3. **Downloaded and trialed.** In RevenueCat’s data, 82.1% of health and fitness trials start on the day of install.
4. **Converted.** About 37.7% of trials convert at the median, mostly to annual plans. That is when the recommendation becomes revenue.

For care programs, step three is an enrollment form or an eligibility check rather than an app trial, and the stakes of step two are higher because people are sharing more personal information.

There is also a newer path that skips the store. If an app is connected inside ChatGPT Health, OpenAI documents that the assistant can reference that app’s data or suggest the app during a conversation. We infer that being a well-described, trusted connected app may matter more over time, though no data yet shows how many users arrive this way.

## What decides which health app an assistant names?

Evidence the assistant can verify about the product and the company; the platforms do not publish their selection rules.

What is documented: OpenAI requires apps in ChatGPT Health to meet privacy and security requirements, and [Google says](https://developers.google.com/search/docs/appearance/ai-features) SEO best practices remain relevant for AI Overviews and AI Mode, with no additional requirements or special optimizations needed to appear. [Apple’s App Review Guidelines](https://developer.apple.com/app-store/review/guidelines/) add a gate that shapes what gets into the store at all: under guideline 1.4.1, apps must disclose data and methodology to support accuracy claims about health measurements, and medical apps may be reviewed with greater scrutiny.

What has been observed:

- **Health answers favor institutions.** In an audit of 615 sources cited by ChatGPT for consumer health questions, [Jacques and colleagues](https://arxiv.org/abs/2601.17109) found commercial health platforms earned 12.4% of citations. Those that were cited stated a medical review on 71.1% of pages. We explain the pattern in [what commercial health sites cited by ChatGPT have in common](https://underneath.agency/resources/chatgpt-cited-commercial-health-sites).
- **Lists move.** In [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), only 25.2% of the brands ChatGPT named appeared in all five repeats of a question, so a single check proves little.

What we infer: in this industry, the trust factors are unusually concrete. They include an accurate statement of regulatory status, published evidence that is described honestly, plain-language privacy terms, clinical or expert involvement that is named, a strong store rating and a clean record on cancellations and billing.

## Why is it hard for a health app to describe itself to an assistant?

Because the rules limit health claims, so the proof must be specific, accurate and plainly worded.

The [Food and Drug Administration revised its general wellness policy](https://cov.com/en/news-and-insights/insights/2026/01/fda-issues-revised-guidance-on-general-wellness-products) on January 6, 2026, superseding the 2019 version. As law firm Covington & Burling summarizes it, the update lets more non-invasive wearables that output measures such as blood pressure stay in the low-risk wellness category if their use is strictly wellness. But wellness products still may not claim clinical equivalence, clinical accuracy or medical grade. Covington notes the tension this creates for companies that want to show their product is validated without crossing into restricted claims. Wearable makers work under the same revised policy, as covered in [how wellness brands get recommended for sleep and recovery](https://underneath.agency/resources/wellness-brands-customers-ai-search).

That tension applies directly to AI answers. Anything an assistant says about an app, it first had to read somewhere. If an app’s site overstates its evidence, the company takes on legal risk; if the site says nothing specific, the assistant has nothing to repeat. IQVIA reports that more than 360 software-based digital therapies are now on the market, 140 of them prescription digital therapeutics, so many consumer apps sit beside products with formal clearances. Our inference: the winning description says exactly what the product is, whether it is cleared and for what, what the evidence shows and where it is published, and what the app does not do.

## How does generative engine optimization work for a digital health company?

Generative engine optimization (GEO) makes your app easy for assistants to find, describe accurately and verify through outside evidence.

For a consumer health company, the work usually covers:

1. **One consistent identity.** The same product name, category, platforms and plan details across your site, both app stores and any connected-app listings, so assistants do not confuse your app with a similarly named one.
2. **Regulatory status in plain words.** Whether the product is a general wellness product or cleared, with a link to the FDA listing where one exists. Never imply clearance you do not have.
3. **Evidence pages written for readers.** Summaries of published studies with links, who ran them and their limits, reviewed by named experts, without claims beyond what the evidence and your regulatory status allow.
4. **Privacy that answers the question.** A readable page on what data you collect, what you never share for advertising and how people delete it, matching your store privacy labels.
5. **Reputation where people check.** Strong ratings and fast responses in the app stores, clear cancellation terms and honest handling of complaints. Avoid any review tactic regulators would treat as deceptive; [fake reviews also distort AI answers](https://underneath.agency/resources/fake-reviews-fake-brands-ai-recommendations).
6. **Independent coverage.** Editorial reviews, clinician and dietitian mentions, research partnerships and employer or insurer listings that describe you accurately. Supplement makers lean on outside proof too, as [how supplement brands win AI customers](https://underneath.agency/resources/supplement-brands-customers-ai-search) shows.
7. **Measurement.** Ask a fixed set of illustrative category, comparison, privacy and legitimacy questions across ChatGPT, Gemini, Perplexity and Google’s AI features, repeatedly, and track who is named, what is said and which sources are cited. Our guide to [designing AI visibility tracking](https://underneath.agency/resources/how-to-design-ai-visibility-tracking) explains how.

No one can promise that an assistant will recommend a particular app or program. What this work does is give assistants an accurate, well-sourced account of your product, so the description they pass on to a would-be user is the right one.

## Which questions about health apps and AI remain unanswered?

How many downloads and enrollments AI answers create, and how assistants weigh clinical evidence, are still unknown.

- **Health questions are not app choices.** Rock Health and OpenAI measure health questions in general. We found no public data on how often people ask an assistant which health app or program to choose.
- **Surveys and vendors.** Rock Health is a venture fund that invests in digital health, AppTweak and RevenueCat sell to app makers, and surveys rely on what people report.
- **Evidence weighting is unmeasured.** No study we found tests whether published clinical studies make an assistant more likely to name an app. The Jacques audit looked at cited pages, not app recommendations.
- **Connected apps are new.** ChatGPT Health launched in January 2026; how often it suggests a connected app, and how it chooses among them, is not public.

## Where should a digital health company start?

Start by asking the questions your future users ask, then see which apps assistants name and what they say.

That first check usually shows whether your product is named for the goals you serve, whether assistants describe your regulatory status and privacy practices correctly, which reviews and articles shape the answer and which rivals appear instead. From there, the work is to publish proof that is accurate and compliant, strengthen your store reputation and keep every description consistent.

If your growth depends on trials that become annual subscribers or enrollments in a paid program, [talk to us about a review of your app’s AI visibility](https://underneath.agency/contact). We will map where your product appears in AI answers, why competitors are named instead, and which changes are most likely to bring more qualified people to your download and sign-up pages. Our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page lays out how that work runs for a health company, with claims kept inside your regulatory status and privacy terms written plainly.

## Frequently asked questions

### Do people use ChatGPT to find health apps?

Many use it for health questions, and some for app choices. Rock Health found 32% of US adults used AI chatbots for health information, and OpenAI lets people connect apps such as MyFitnessPal inside ChatGPT Health.

### Can a health app pay to be suggested in ChatGPT Health?

OpenAI does not describe paid placement in Health. It says connected apps need explicit permission, must meet its privacy and security requirements and pass an additional security review.

### Should a wellness app talk about clinical accuracy to get recommended?

No. FDA’s revised general wellness policy says wellness products should not claim clinical accuracy, clinical equivalence or medical grade. Describe validation honestly and within your regulatory status.

### Does privacy affect whether an assistant recommends an app?

We infer it does when people ask about privacy, because the assistant can only repeat what it finds. The FTC also enforces breach notification for health apps not covered by HIPAA.

## Sources

- Fierce Healthcare (2026-03-24), [AI chatbot use for health information up 16% from 2024: Rock Health survey](https://fiercehealthcare.com/ai-and-machine-learning/ai-chatbot-use-health-information-16-2024-rock-health-survey)
- HIT Consultant (2026-03-23), [Rock Health: 32% of Consumers Now Use AI for Health Information](https://hitconsultant.net/2026/03/23/rock-health-2025-survey-consumer-ai-adoption-chatgpt-healthcare/)
- HIT Consultant (2026-07-13), [Rock Health H1 2026 Digital Health Funding Recap](https://hitconsultant.net/2026/07/13/rock-health-h1-2026-digital-health-funding-report/)
- Drug Store News (2024), [IQVIA releases Digital Health Trends 2024 report](https://drugstorenews.com/iqvia-releases-digital-health-trends-2024-report)
- RevenueCat (2026), [State of Subscription Apps 2026](https://www.revenuecat.com/state-of-subscription-apps)
- OpenAI (2026-01), [Introducing ChatGPT Health](https://openai.com/index/introducing-chatgpt-health/)
- AppTweak (2026), [AI app discovery survey](https://www.apptweak.com/en/aso-blog/ai-app-discovery-survey)
- Federal Trade Commission (2024), [Complying with FTC’s Health Breach Notification Rule](https://www.ftc.gov/business-guidance/resources/complying-ftcs-health-breach-notification-rule-0)
- Covington & Burling (2026-01-08), [FDA Issues Revised Guidance on General Wellness Products](https://cov.com/en/news-and-insights/insights/2026/01/fda-issues-revised-guidance-on-general-wellness-products)
- Apple (2026), [App Review Guidelines](https://developer.apple.com/app-store/review/guidelines/)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- Jacques and colleagues (2026), [Authority Signals in AI Cited Health Sources: A Framework for Evaluating Source Credibility in ChatGPT Responses](https://arxiv.org/abs/2601.17109), arXiv:2601.17109.
- Underneath (2026), [“Is this brand legit?” How AI assistants build a reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study)
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)

---

This is the Markdown twin of https://underneath.agency/resources/digital-health-apps-users-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Do AI Agents Invent Facts About Your Business or Omit Them?"
description: "Mostly they leave facts out. Answering from the web instead of a business’s site, agents left out 45% of facts, up from 29%; wrong facts rose only 4% to 6%."
canonical: "https://underneath.agency/resources/do-ai-agents-invent-or-omit-business-facts"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Do AI agents make up facts about my business, or just leave them out?

Mostly they leave them out. In a large test of AI agents answering buyer questions about real businesses, outright wrong facts were rare, while facts the buyer asked for and never got were common. When a wrong fact does appear, it usually comes from somewhere on the web, often the business’s own pages.

## The short version

1. When AI agents built answers from the open web instead of the business’s own site, facts never mentioned rose from 29% to 45%, while facts stated wrongly rose only from 4% to 6% ([Finder and colleagues, 2026](https://arxiv.org/abs/2609.34951)).
2. Answers built off-site were 3.7 times more likely to contain none of the facts the buyer asked for (same study, written by the maker of an agent-readiness tool).
3. In Google’s AI Overviews, 2.66% of 98,020 claims were contradicted by a page the Overview itself cited ([Xu, Iqbal and Montgomery, 2026](https://arxiv.org/abs/2605.14021)).
4. In [our pricing study](https://underneath.agency/research/ai-pricing-accuracy-study), only 4 of 64 software prices that differed from the official page could not be found on any source we could fetch.

## Is a wrong fact or a missing fact the bigger risk?

A missing fact is the more common risk; outright wrong facts are comparatively rare.

The clearest evidence comes from [Finder and colleagues](https://arxiv.org/abs/2609.34951). They ran 37,927 agent journeys, each a buyer question about pricing, features or setup, across 1,056 real businesses and four AI agent setups. They then graded answers against facts captured from each business’s own site.

Fact by fact, the shift was toward silence, not invention:

| What happened to each asked fact | Answer built from the business’s site | Answer built from the wider web |
|---|---|---|
| Stated wrongly | 4% | 6% |
| Never mentioned | 29% | 45% |

The authors sum it up: poor readability “makes a fact unretrievable, not false.” Site-built answers got 48.3% of asked facts right, against 34.3% for web-built ones about the same business and question.

Read this with care. The authors work for ora, which sells the agent-readiness score the study uses, and they say so. Accuracy was graded on only 131 businesses. On that subset, the difference between the two groups of businesses was not statistically reliable; the gap appears when comparing answers about the same business.

## What happens when an agent cannot read your site?

It answers anyway, using other sources, and the answer gets thinner and more hedged.

In about 99% of journeys that hit a dead end on a business’s site, the agent still answered, [built from whatever it found elsewhere](https://underneath.agency/resources/when-ai-agents-cant-read-your-site). Only 56% of runs on hard-to-read sites ended grounded in the business’s own pages, against 78% on readable sites.

The answers also hedged more. Agents said they could not access the business 4.4 times more often when the site was hard to read. Training memory did not fill the gap either: it supplied only 7% to 10% of the finished answer whether or not the site was readable.

So the typical failure is not an agent confidently inventing a price. It is an agent saying it could not find the price, or giving a general answer with the specifics missing.

## How often do AI answers state something false outright?

Rarely in the large audits, though how often depends on the engine, the topic and how “false” is measured.

[Xu, Iqbal and Montgomery](https://arxiv.org/abs/2605.14021) checked 98,020 claims from Google’s AI Overviews (the AI summary at the top of Google’s results) against the pages each Overview cited, over 40 days in spring 2026. Only 2.66% were directly contradicted by a cited page. Another 6.98% were not mentioned by any cited page at all.

That second group matters for this question. A claim no cited page mentions could be made up, or could come from a page the Overview did not cite. The study cannot tell which. Even so, outright contradiction was the smallest failure.

Earlier answer engines did worse. In a 2024 audit, You.com and Perplexity each had around 30% of their statements unsupported by the sources they listed ([Narayanan Venkit and colleagues, 2024](https://arxiv.org/abs/2410.22349)). Unsupported is not the same as false, but it shows the risk was higher in older systems.

## When an AI gets a business fact wrong, where does it come from?

Usually from a real source, often one the business itself published, rather than from thin air.

Our own studies traced differing facts back to their likely sources:

- **Software prices.** In [our pricing study](https://underneath.agency/research/ai-pricing-accuracy-study), 64 quoted prices differed from the official pricing page. For 39 of them the same figure was on another page of the vendor’s own site, and for 21 more on a third-party page cited for the product. Only 4 were found nowhere.
- **Phone numbers.** In [our business facts study](https://underneath.agency/research/ai-business-facts-accuracy-study), 55 phone numbers differed from the business’s Google profile. 42 of them were on the business’s own website. Only 3 of the 633 numbers given could not be found on the profile, the website or any page the answer cited.

In both cases the wrong-looking fact was mostly an old or alternative fact still live on the web. An assistant that finds two official answers can quote either.

## Which facts go missing most often?

The research points to setup documentation, full price lists and opening hours.

- **Setup details.** In the agent study, 65% of asked setup facts went unmentioned whichever way the answer was built, because setup lives in documentation rather than on marketing pages.
- **Full price lists.** In our pricing study, answers named a price for 84.2% of the paid plans on the official page on average. The rest were simply left out.
- **Opening hours.** In our business facts study, 9.4% of answers gave Monday hours that differed from the Google profile, and 12.0% gave none.

Notice the pattern in that last case. For hours, leaving the fact out was about as common as stating a different one. Omission is the main risk, not the only one.

## What should you do about it?

Make your key facts easy for an agent to fetch, and remove the old versions that compete with them.

1. **Put prices, plans and core facts in plain text on pages agents can load.** Agents often cannot see content that only appears after a page’s scripts run, or pages behind bot blocks. See [whether agents recommend readable websites](https://underneath.agency/resources/do-ai-agents-recommend-readable-websites) for why this matters beyond accuracy.
2. **Retire or update old pages when facts change.** Most differing prices and phone numbers in our studies were still published somewhere by the business itself.
3. **Publish one version of each fact everywhere.** Use the same phone number, hours and plan names on your site, your Google profile and directories. How much those outside listings count once an agent is reading about you is covered in [off-site mentions versus site readability](https://underneath.agency/resources/off-site-mentions-vs-site-readability-ai-agents).
4. **Surface setup and documentation facts.** Agents rarely reached them, so link them clearly from pages agents do read.
5. **Ask assistants your buyers’ questions regularly.** Check what is missing, not just what is wrong.

If you want help making your site readable to AI agents, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The research shows omission dominates in the settings tested, but those settings are narrow.

- The main agent study comes from a company that sells agent-readiness scoring, uses its own score as the measure, and graded accuracy on only 131 businesses.
- That study did not match businesses on how much content they publish, so thinner sites may explain part of the gap.
- Its sample leans toward software and online commerce, in English. Local services and other languages are untested.
- The AI Overview audit cannot tell whether an unsupported claim was invented or taken from an uncited page.
- No study yet measures how often buyers act on a missing fact versus a wrong one, so the business cost of each is unknown.

## Frequently asked questions

### Do AI assistants make up prices?

Rarely, in our tests. Of 64 software prices that differed from the official page, only 4 could not be found on any source we could fetch; most were old prices still on the vendor’s own site.

### Why does ChatGPT leave out information about my company?

Often because it could not read your site. In one large agent study, only 56% of runs on hard-to-read sites ended grounded in the business’s own pages, against 78% on readable sites.

### Is a missing fact really less harmful than a wrong one?

Not necessarily, but it is more common. Answers built off-site were 3.7 times more likely to contain none of the facts the buyer asked for, which can cost a recommendation.

### How accurate are Google’s AI Overviews about facts?

In a 40-day audit, 2.66% of claims were contradicted by a page the Overview cited, and 6.98% were not mentioned by any cited page.

## Sources

- Finder, Elovic, Shalev and Yosef (2026), [AX is the New AEO](https://arxiv.org/abs/2609.34951), arXiv:2609.34951.
- Xu, Iqbal and Montgomery (2026), [Measuring Google AI Overviews: Activation, Source Quality, Claim Fidelity, and Publisher Impact](https://arxiv.org/abs/2605.14021), arXiv:2605.14021.
- Narayanan Venkit and colleagues (2024), [Search Engines in an AI Era: The False Promise of Factual and Verifiable Source-Cited Responses](https://arxiv.org/abs/2410.22349), arXiv:2410.22349.
- Underneath (2026), [How faithfully do AI assistants quote software prices?](https://underneath.agency/research/ai-pricing-accuracy-study)
- Underneath (2026), [Do AI answers match a business’s Google profile?](https://underneath.agency/research/ai-business-facts-accuracy-study)

---

This is the Markdown twin of https://underneath.agency/resources/do-ai-agents-invent-or-omit-business-facts. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Do AI agents recommend businesses with readable websites more?"
description: "In one large vendor study, AI agents clearly recommended businesses with agent-readable websites 20% of the time, against 11% for sites they struggled to read."
canonical: "https://underneath.agency/resources/do-ai-agents-recommend-readable-websites"
published: 2026-10-11
updated: 2026-10-11
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Do AI agents recommend businesses whose websites they can read more often?

Yes, in the largest study so far: AI agents clearly recommended businesses with agent-readable websites 20% of the time, against 11% for hard-to-read sites. When an agent cannot read a site, it answers anyway from other sources, and the answer is more likely to miss the facts the buyer asked for. The study was run by a company that sells agent-readiness scoring, so treat the size of the gap with care.

## The short version

1. Across 37,927 agent sessions on 1,056 real businesses, agent-ready businesses were clearly recommended 1.9 times as often, in [a study by ora research](https://arxiv.org/abs/2609.34951).
2. When the site was readable, 78% of the answer came from the business’s own pages; when it was not, 58% did.
3. Answers built from the business’s own site got 48.3% of the asked facts right, against 34.3% for answers built from elsewhere.
4. Few sites are set up for agents: only 3.2% of top websites served a Markdown version on request in [our study of 5,902 sites](https://underneath.agency/research/agent-readable-web-study).

## What did the study find?

Agent-ready businesses were clearly recommended almost twice as often. The gap held in every agent setup and every business category tested.

Finder and colleagues at ora research ran 37,927 agent sessions. An AI agent here means an assistant that searches the web, opens pages and reads them before answering. Each session was a buyer question about one business: its pricing, its features or how to set it up.

The 1,056 businesses were split into two groups by how readable their sites were to agents. The groups were matched on fame, on how much AI systems already knew about each brand, and on how often the brand was mentioned on other sites. Two AI judges from different companies rated every answer, and an answer counted as a clear recommendation only if both gave it the top score.

## What does “agent-ready” mean?

It means an agent can fetch your pages and read the content as text, without being blocked. The study used ora’s own scoring tool to measure it.

The score checks things such as how much content survives without JavaScript, the code that builds many pages in the browser after they load. It also checks structured data, Markdown versions of pages, whether pricing and documentation are reachable, and whether bot controls let in agents acting for a user. The authors write that across nearly 100,000 sites they have scored, fewer than 1% earn the tool’s top grade.

Our own measurements show how early this is. In our study of the top 10,000 websites, 45.8% of homepages carried JSON-LD structured data, a machine-readable description of the page. In [our crawler study](https://underneath.agency/research/ai-crawler-blocking-study), 15.2% of top sites blocked OpenAI’s GPTBot from the whole site.

## What happens when an agent cannot read your site?

It answers anyway, using other people’s pages instead of yours. In about 99% of sessions that hit a dead end on a site, the agent still answered.

For agent-ready businesses, 78% of the answer came from their own pages and 12% from web search. For the others, own-site content fell to 58% and web search rose to 25%. The share from the AI’s built-in knowledge stayed small either way, at only 7–10% of the answer.

The agents also searched more. Across the readiness scale, average web searches per session rose from 1.8 on the most readable sites to 4.5 on the least. In one example in the paper, an agent blocked twice by a software company’s site ended its answer on two competitors’ blogs.

The tone of answers changed too. When agents could not read a site, they were 4.4 times as likely to say they could not access the business. They were 3.0 times as likely to vouch for it using secondhand sources.

## Are answers less accurate without your website?

Yes, mostly because facts go missing, not because they are made up. Comparing answers about the same business, site-built answers were 41% more accurate.

Site-built answers got 48.3% of the asked facts right, against 34.3% for answers built from other sources. Answers built elsewhere were 3.7 times as likely to contain none of the facts the buyer asked for. Wrong facts rose only from 4% to 6%, while facts never mentioned rose from 29% to 45%. Our guide on [whether agents invent or omit facts](https://underneath.agency/resources/do-ai-agents-invent-or-omit-business-facts) looks at where the wrong ones come from.

| Fate of each asked fact | Built from the business’s site | Built from elsewhere |
|---|---|---|
| Stated wrong | 4% | 6% |
| Never mentioned | 29% | 45% |

Pricing showed the biggest gain from reading the site, and setup questions showed none.

Our own studies point the same way. In [our business facts study](https://underneath.agency/research/ai-business-facts-accuracy-study), AI answers gave a phone number that differed from the Google profile 30.6% of the time when the profile number was missing from the business’s own website. When it was there, the figure was 1.6%.

## Does the effect hold across agents and industries?

The direction held everywhere, but the size varied a lot. Baseline recommendation rates differed sevenfold between agent setups.

The lowest, Claude Code, clearly recommended 5% of the time; the highest, eve on GPT-5.4, 36%. Each recommended agent-ready businesses more. Across the twelve business categories, the lift ranged from 1.2 times in customer service to 2.5 times in IT infrastructure.

The study also tested whether [being talked about elsewhere](https://underneath.agency/resources/off-site-mentions-vs-site-readability-ai-agents) mattered once readability was known. Across all 1,056 businesses, neither of the two measures of off-site mentions predicted recommendation, while readability did.

That is one study’s result. In [our brand entity study](https://underneath.agency/research/brand-entity-ai-recommendations-study), independent coverage in the pages assistants cited was the strongest predictor of recommendation we measured.

## How much should you trust this study?

Treat it as strong early evidence with clear caveats. The authors sell the scoring tool they used, and they say so.

- **Vendor study.** ora built the readiness ranker and reports its scores “as the study’s instrument rather than as an independent measure.”
- **Content depth not matched.** A business that blocks agents may also publish less. The authors say the recommendation gap “cannot rule this out.”
- **AI judges.** Recommendations were rated by two AI systems, not by people, though the two agreed within one point on 88% of answers.
- **Accuracy on a subset.** Facts were graded for 131 businesses. Between the two groups as a whole, accuracy differed by under two points (51.7% against 49.8%), which was not a clear difference.
- **Scope.** Mostly software and commerce businesses, in English, asking about business facts.

## What should you do about it?

Make sure an agent that reaches your site can read it, then keep earning the visits through search. Practical steps:

1. Load your key pages with JavaScript turned off. If pricing or product facts disappear, agents probably cannot see them.
2. Check robots.txt and bot-protection settings for the AI crawlers and user-triggered agents you want to admit. In our crawler study, none of the 123 top-site pages ChatGPT cited that day was closed to its search crawler. Opting out of Google’s AI training is a separate call, weighed in [whether blocking Google-Extended hurts AI Overviews](https://underneath.agency/resources/does-blocking-google-extended-hurt-ai-overviews).
3. Publish prices, plans and core facts as plain text on your own site, in one consistent version.
4. Add an Organization record in structured data, and consider offering Markdown versions of key pages.
5. Keep doing search work. Agents still reach most pages through a web search, as the authors note.

If you want an outside audit of how agents read your site, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

No one has yet fixed a site’s readability and measured its recommendations before and after.

- **Cause and effect.** The study compared different businesses, so publishing depth or product quality may explain part of the gap.
- **Independent replication.** The only large study comes from a vendor using its own scoring tool.
- **Other agents and tasks.** Four agent setups were tested; buying, booking and other actions were left for later work.
- **Which fix matters most.** The score bundles many checks, so the effect of Markdown, structured data or bot settings alone is unknown.

## Frequently asked questions

### Do AI agents recommend businesses whose websites they can read?

In the largest study so far, yes: agent-ready businesses were clearly recommended 20% of the time, against 11% for others. The study was run by ora research, which sells the readiness tool it used.

### Does blocking AI bots hurt my brand in AI answers?

It can shift the answer to other sources. When agents could not read a site, only 58% of the answer came from the business’s own pages, against 78% when they could.

### Do AI agents make up facts about businesses they cannot read?

Mostly they leave facts out. In the ora study, wrong facts rose only from 4% to 6%, while facts never mentioned rose from 29% to 45%.

### How many websites are ready for AI agents?

Very few by any measure. Only 3.2% of top websites served Markdown to agents in our study, and ora reports fewer than 1% of the sites it scored earned its top grade.

## Sources

- Finder, Elovic, Shalev and Yosef (2026), [AX is the New AEO](https://arxiv.org/abs/2609.34951), arXiv:2609.34951.
- Underneath (2026), [Do websites serve Markdown to AI agents? 2026 data](https://underneath.agency/research/agent-readable-web-study)
- Underneath (2026), [Which AI crawlers do top websites block? 10,000 sites, 2026](https://underneath.agency/research/ai-crawler-blocking-study)
- Underneath (2026), [Do AI answers match a business’s Google profile?](https://underneath.agency/research/ai-business-facts-accuracy-study)
- Underneath (2026), [Do Wikipedia and schema make AI assistants recommend a brand?](https://underneath.agency/research/brand-entity-ai-recommendations-study)

---

This is the Markdown twin of https://underneath.agency/resources/do-ai-agents-recommend-readable-websites. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Do AI assistants answer from training data or live web pages?"
description: "Mostly from live pages. In a 2026 test of four AI agents, memory supplied 7–10% of answers, though some AI answers still cite no source at all."
canonical: "https://underneath.agency/resources/do-ai-assistants-answer-from-training-data"
published: 2026-10-10
updated: 2026-10-10
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Do AI assistants still answer from training data or from live web pages?

For buyer questions, today’s AI assistants mostly answer from pages they fetch at the moment of answering, not from what they memorized in training. In one 2026 test of four AI agents, memory supplied only 7–10% of the finished answer. Memory has not vanished, though: a minority of AI answers still cite no web source, and an assistant without web search rarely finds new brands.

## The short version

1. Across seven OpenAI releases, the share of an answer built from memory fell from 52% to 14%, in a 2026 test of 90 buyer questions by Finder and colleagues, who work for a company that sells website readiness scores for AI agents.
2. In the same team’s main test of four AI agents, training knowledge supplied only 7–10% of the finished answer, whether or not the business’s website was easy for the agent to read.
3. Some answers still skip the web: 8.8% of answers from ChatGPT, Copilot, Gemini and Perplexity carried no citation, rising to 11.6% for health questions (Allaham and Diakopoulos, 2026).
4. Live search changes who gets found: Perplexity named new Product Hunt startups in 8.29% of discovery questions, against 3.32% for an OpenAI model with no web search (Sharma, 2026).
5. In our own test, ChatGPT ran 3.7 searches per buyer question, and none of the answers that skipped searching cited a single page.

## How much of an AI answer still comes from memory?

Very little, in the most recent test of AI agents answering buyer questions. [Finder and colleagues](https://arxiv.org/abs/2609.34951) put 90 buyer questions about real businesses to successive OpenAI releases, with web tools available but optional. A separate AI model then judged how much of each answer came from memory rather than from pages the system retrieved.

The older OpenAI model built about half its answer from memory. Then the share fell steeply, from 52% in gpt-4.1 to 14% in gpt-5.6. In the team’s main experiment, 37,927 agent journeys about 1,056 businesses, training knowledge made up only 7–10% of the finished answer.

| Setup | Share of the answer built from memory |
|---|---|
| gpt-4.1, OpenAI’s last older-style flagship | 52% |
| gpt-5.6 | 14% |
| Four AI agent setups, second half of 2026 | 7–10% |

Read these figures with care. The authors work at ora, which sells the readiness score used in the study, and the memory share was judged by an AI model rather than by people. The release-by-release trend covers OpenAI models only.

## How often do AI search engines skip the web entirely?

Sometimes, and the rate depends on the engine and the topic. [Allaham and Diakopoulos](https://arxiv.org/abs/2605.23684) put 712 real user questions on politics, health and the environment to ChatGPT with search, Copilot, Gemini and Perplexity, through each product’s own interface. Of the 2,848 answers, 8.8% carried no citation at all.

Health questions were the most likely to get an answer with no sources: 11.6%, about one in nine. The figure was 8.2% for environment questions and 5.7% for politics. The authors read this as a sign that the engines still lean on what the model already knows for some questions, which raises doubts about how current those answers are.

Engines also differ sharply. In a mid-2025 study of 55,936 queries by [Zhang and colleagues](https://arxiv.org/abs/2512.09483), Grok gave no cited website in 82% of its answers and Gemini in 38%. In our [study of AI citations and Google rankings](https://underneath.agency/research/ai-citations-google-rankings-study), Claude cited sources in 61 of 80 buyer answers and Gemini in 73, while ChatGPT and Perplexity cited sources in all 80.

## Does live search change which brands an assistant can find?

Yes, and the difference is largest for new brands. A model without web access cannot know about a product launched after its training ended. [Sharma](https://arxiv.org/abs/2601.00912) tested 112 startups from the 2025 Product Hunt leaderboard on two AI models in December 2025.

Asked about a product by name, the model without web search recognized it 99.4% of the time. Asked a discovery question, such as which new tools in a category are worth trying, it named the startup in only 3.32% of answers. Perplexity, which searches the web as it answers, did so in 8.29%.

Over the whole test, Perplexity surfaced 31 of the 112 products at least once, 27.7%, against 6 products (5.4%) for the model without search. One caveat matters: the “ChatGPT” in this study is a low-cost OpenAI model called directly by developers with no web search, not the consumer ChatGPT app with search turned on. The same startup test also looked at [what predicts visibility in Perplexity](https://underneath.agency/resources/do-backlinks-matter-for-perplexity-visibility).

## What do assistants do instead of remembering?

They search, often several times, and they favor recent pages. In our [hidden searches study](https://underneath.agency/research/ai-hidden-searches-study), ChatGPT ran a mean of 3.7 searches per buyer question before answering, Gemini 1.9 and Claude 0.76. Every answer that ran a search cited at least one page, and none of the 42 answers without a search cited anything.

The pages they find tend to be new. In our [freshness study](https://underneath.agency/research/ai-source-freshness-study), pages published in the last 90 days made up 17.4% to 22.6% of each assistant’s dated citations, against 6.9% of Google’s top 10 for the same questions.

When an agent cannot read a business’s own website, it does not fall back on memory either. Finder and colleagues found that web searches per journey rose from 1.8 on the easiest sites for an agent to read to 4.5 on the hardest. The gap was filled with other people’s pages, not with training knowledge. What that does to accuracy is covered in [whether agents invent or omit business facts](https://underneath.agency/resources/do-ai-agents-invent-or-omit-business-facts).

## What should you do about it?

Put most of your effort into what AI assistants can read today, not into what they may have memorized. In practice:

1. State current facts plainly on your own pages: prices, features, locations and how to get started. That is what an agent fetches when a buyer asks about you.
2. Check that AI search crawlers and the fetchers that act for a user can reach those pages. Our [crawler blocking study](https://underneath.agency/research/ai-crawler-blocking-study) shows how often old robots.txt rules block them by accident. Your server logs show which pages those bots request, as [reading AI bot server logs](https://underneath.agency/resources/ai-bot-server-logs-content-demand) explains.
3. Keep earning coverage on other sites. Assistants reach most pages through their own web searches, so what ranks and what others publish about you still shapes the answer.
4. Test discovery questions, not just your brand name. Recognition by name, 99.4% in one test, says little about whether an assistant recommends you when a buyer asks for options.
5. Watch health and other sensitive topics more closely, because answers with no sources were most common there.

If you want help turning this into a plan, our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) starts with these checks.

## What does the research not tell us yet?

The direction is clear, but the evidence has real gaps.

- The sharpest evidence of memory’s decline comes from one vendor study. It covers OpenAI models only and relies on an AI judge.
- No study here measures how memory shapes which brands an assistant searches for in the first place. A search can still be steered by what the model already knows.
- The rates of uncited answers come from a few hundred to a few thousand questions, on specific topics, at one point in time.
- The startup test used a developer version of an OpenAI model with no search, so it does not describe consumer ChatGPT with search on.
- None of these studies changed a website and then tracked what assistants said about it over time.

## Frequently asked questions

### Does ChatGPT use its training data or search the web?

Both, but for buyer questions it now mostly searches. In one 2026 vendor test, memory supplied 14% of answers from gpt-5.6, down from 52% for gpt-4.1.

### Can a brand launched after an AI’s training cutoff appear in its answers?

Yes, but mainly through live search. In a December 2025 test, a model without web search named new startups in 3.32% of discovery questions, against 8.29% for Perplexity with search.

### Do AI search answers always cite their sources?

No. In a 2026 audit of four engines, 8.8% of answers had no citation, and in a separate mid-2025 study Grok gave no cited website in 82% of its answers.

### Is it still worth trying to influence what AI models learned in training?

It now covers a small part of the answer. Finder and colleagues found training knowledge made up 7–10% of agent answers in their 2026 test, with most of the rest built from pages read at the moment of answering.

## Sources

- Finder, Elovic, Shalev and Yosef (2026), [AX is the New AEO](https://arxiv.org/abs/2609.34951), arXiv:2609.34951.
- Allaham and Diakopoulos (2026), [Synthetic Sources?: Auditing Generative Search Engine Citations for Evidence of AI-Generated Sources](https://arxiv.org/abs/2605.23684), arXiv:2605.23684.
- Zhang et al. (2025), [Source Coverage and Citation Bias in LLM-based vs. Traditional Search Engines](https://arxiv.org/abs/2512.09483), arXiv:2512.09483.
- Sharma (2026), [The Discovery Gap: How Product Hunt Startups Vanish in LLM Organic Discovery Queries](https://arxiv.org/abs/2601.00912), arXiv:2601.00912.
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude cite pages that rank?](https://underneath.agency/research/ai-citations-google-rankings-study)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [How fresh are the pages AI engines cite?](https://underneath.agency/research/ai-source-freshness-study)

---

This is the Markdown twin of https://underneath.agency/resources/do-ai-assistants-answer-from-training-data. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Do AI assistants favor big brands over smaller competitors?"
description: "Mostly yes when products look alike: AI assistants default to market leaders, but in tests a small, clearly stated advantage beat a famous name."
canonical: "https://underneath.agency/resources/do-ai-assistants-favor-big-brands"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Do AI assistants favor big brands over smaller competitors?

Yes: when nothing else separates the options, AI assistants lean hard toward the brands they already know. The lean is real but conditional, because in controlled tests a small, clearly stated advantage was enough to beat a famous name. For a challenger brand, the task is to give the assistant a reason to pick you, and to put that reason where the assistant will read it.

## The short version

1. In 50 unbranded soda questions to ChatGPT and Perplexity, major brands took 62.2% of brand mentions and niche brands 9.0% ([Chen and colleagues](https://arxiv.org/abs/2509.08919), University of Toronto, 2025).
2. When ten skincare products had identical specifications, three AI models picked the one real brand in all 670 valid trials ([Chu and Hou](https://arxiv.org/abs/2606.17443), 2026).
3. In the same tests, a rating edge of just +0.075 stars was enough for an unknown brand to win half the time.
4. In one vendor’s data on 102 brands, household names appeared in 73% of unbranded AI answers on day one and niche brands in 11% ([Kumar](https://arxiv.org/abs/2606.20065), 2026).
5. In [our own study](https://underneath.agency/research/brand-entity-ai-recommendations-study) of four assistants, independent coverage was the strongest signal: each tenfold rise in independent sites naming a brand went with 4.7 times the odds of a recommendation.

## How strong is the big-brand lean in real AI answers?

It is strong enough to see clearly in live tests of several assistants. [Chen and colleagues](https://arxiv.org/abs/2509.08919) asked ChatGPT and Perplexity 50 unbranded questions, such as “most popular cola brand” or “top cola brands in the US”. They then sorted every brand named into major, niche or other.

| Assistant | Major brands | Niche brands | Other |
|---|---|---|---|
| ChatGPT | 56.3% | 12.3% | 31.4% |
| Perplexity | 67.9% | 5.8% | 26.3% |
| Both combined | 62.2% | 9.0% | 28.8% |

Coca-Cola alone drew 107 mentions on Perplexity. The authors suggest that prominent sources and what the assistants already “know” jointly pull answers toward major labels. That is their reading, not something they measured, and questions such as “most popular” invite famous names by design.

A second, larger data set points the same way. Kumar, a co-founder of the AI visibility company Ranqo, analyzed 102,025 AI answers about brands tracked on the company’s platform between March and May 2026. On each brand’s first tracking run, global household names appeared in 73% of unbranded category answers, and niche or small brands in 11%. This is a vendor’s own customer data, not a random sample, and the brand tiers were assigned by hand.

## Is it the brand name itself that wins?

In controlled tests, yes: with nothing else to go on, the assistants picked the name they recognized. [Chu and Hou](https://arxiv.org/abs/2606.17443) gave three AI models (GPT-4o-mini, Claude Sonnet and Gemini 3 Flash) lists of ten skincare products with identical ratings, prices, reviews and ingredients. One was a real brand such as CeraVe; the other nine were invented names.

Across 670 valid trials, the real brand was recommended every single time, in English and in Chinese. The product list was supplied in the question, with no live web search. So this measures what the assistant brings from its training, not what it finds online. Skincare makers can see what this means for them in [how skincare brands get found in AI answers](https://underneath.agency/resources/skincare-brands-ai-search).

An older test by [Pfrommer and colleagues](https://arxiv.org/abs/2406.03589) at UC Berkeley found a similar pull in some assistants. Across 50 product categories, GPT-4 Turbo and Llama 3 were heavily swayed by what they already knew about product names, and GPT-4 Turbo paid little attention to the product pages it was given. Those were 2024-era models in a research setup, not today’s consumer ChatGPT.

## Can a small, clear advantage beat a famous name?

Yes, in the same tests the brand advantage collapsed once a rival was visibly better. When Chu and Hou gave an invented brand better specifications than the real one, the assistants still chose the real brand only 1.7% to 4.6% of the time.

The tipping point was tiny. An unknown brand won half the time with a rating just +0.075 stars higher, 1.6 times as many reviews, or a 7.3% lower price. Across all their conditions, product details explained 82.4% of how the assistants ranked the products, and the brand name only 1.2%.

The authors call this a “conditional monopoly”: the big brand wins as a tiebreaker when the information looks the same. The catch is that their test handed the assistant clean, comparable facts for every product. In real AI answers, a smaller brand’s facts first have to be found.

## Is “big” the same as “popular” to an AI assistant?

Not exactly: the brands AI assistants favor are not always the ones consumers know best. [Malthouse and colleagues](https://arxiv.org/abs/2609.16304) at Northwestern University asked six AI models for up to five brands in five categories, 40 times each, with web search switched off. They compared the results with Kantar BrandZ scores for how readily consumers think of each brand.

The match was loose. Disney Cruise Line, which scored 111 on that measure, was almost never recommended, while Viking, at 70, was recommended prominently. In cordless drills and hiking jackets, the models favored premium brands such as DeWalt, Milwaukee, Patagonia and Arc’teryx over mass-market names such as Black+Decker and L.L.Bean. That pattern held in only two of the five categories. We look at why [well-known brands miss AI recommendations](https://underneath.agency/resources/why-well-known-brands-miss-ai-recommendations) and how to check your own.

A hotel experiment by [Baig and colleagues](https://arxiv.org/abs/2606.16344) shows another way established players gain ground. Across twelve AI models choosing between made-up hotels, a high review count (2,100 against 45) raised the chance of being recommended by 8.3 percentage points. The authors note that this kind of weighting advantages established properties over new entrants with thin review histories.

## Where does the big-brand advantage come from?

Mostly from how widely other people write about a brand, as far as current data can tell. In [our study of brand entity signals](https://underneath.agency/research/brand-entity-ai-recommendations-study), ChatGPT, Gemini, Perplexity and Claude answered the same 80 US buyer questions. Of the options all four named for a question, 60.0% had an English Wikipedia article, against 20.6% of the options only one assistant named.

Most of that gap went away once we accounted for how prominent each brand already was. The strongest signal we measured was independent coverage: how many websites, other than the brand’s own, named it in the pages the assistants cited. Each tenfold increase went with 4.7 times the odds of being recommended. That is an association, not proof of cause.

Chen and colleagues reached a similar view from the source side. On questions about well-known brands, 93.5% of the sources ChatGPT drew on were earned media, meaning third-party publications and reviews rather than the brand’s own site or social posts. Paid reach is a separate question, weighed in [whether ad spend helps AI recommendations](https://underneath.agency/resources/does-ad-spend-help-ai-brand-recommendations).

## What should you do about it?

Treat AI visibility as a contest challengers start from behind, and give assistants a specific, checkable reason to choose you.

1. **Measure unbranded questions.** Ask the buyer questions your customers ask, without naming yourself, several times on each assistant. Being recognized by name tells you little about being recommended.
2. **State your advantage as plain, comparable facts.** Ratings, review counts, prices and specifications moved the assistants far more than brand names did in controlled tests. Publish them where they are easy to find and compare.
3. **Earn independent coverage.** Reviews, comparisons and mentions on other sites carried the strongest link to recommendation in our data.
4. **Do not invent authority.** Chu and Hou found that made-up clinical claims swayed the assistants, but they flag such claims as potential false advertising and call for platforms to check them.
5. **Track each assistant separately.** Perplexity leaned toward major soda brands more than ChatGPT did in the same test.

If you want help turning this into a plan, see [how we approach generative engine optimization](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The research shows a clear lean toward familiar brands, but it leaves real gaps.

- **The tests are narrow.** The soda test covered one category and two assistants. The skincare tests supplied the products in the question instead of letting assistants search.
- **The causes are inferred.** Whether training data, prominent sources or something else drives the lean is the researchers’ interpretation, not something they observed.
- **Vendor data needs independent checks.** The tier figures come from one company’s customers, who skew toward software, fintech and Indian consumer brands.
- **No proven lasting fix exists yet.** A [critical survey of 45 studies](https://arxiv.org/abs/2607.14035) found no technique with a stable, cross-platform effect on whether content is discovered in the first place.

## Frequently asked questions

### Does ChatGPT prefer well-known brands?

In unbranded tests, usually yes. In 50 soda questions, 56.3% of the brands ChatGPT named were major labels and 12.3% niche ones; Perplexity leaned further, at 67.9% major.

### Can a small brand beat a big brand in AI recommendations?

Yes, if the assistant can see a clear advantage. In controlled skincare tests, an unknown brand with slightly better ratings, more reviews or a lower price won most head-to-head comparisons.

### Do AI assistants simply recommend the most popular brands?

Not reliably. A Northwestern study found Disney Cruise Line almost never recommended despite high consumer awareness, while some premium brands were favored over bigger mass-market ones.

### Why do AI assistants default to market leaders?

Researchers point to what assistants learned in training and to which sources are most prominent online. Our own data points to independent coverage: brands that many other sites write about were recommended far more often.

## Sources

- Chen, Wang, Chen and Koudas (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919), arXiv:2509.08919.
- Chu and Hou (2026), [Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems](https://arxiv.org/abs/2606.17443), arXiv:2606.17443.
- Kumar (2026), [Generative Engine Optimization at Scale: Measuring Brand Visibility Across AI Search Engines](https://arxiv.org/abs/2606.20065), arXiv:2606.20065.
- Pfrommer, Bai, Gautam and Sojoudi (2024), [Ranking Manipulation for Conversational Search Engines](https://arxiv.org/abs/2406.03589), arXiv:2406.03589.
- Malthouse, Lee, Yang, Pal and Feng (2026), [Evaluating Brand Retrieval and Ranking in Large Language Model Recommendations](https://arxiv.org/abs/2609.16304), arXiv:2609.16304.
- Baig et al. (2026), [Whose hotel does the AI recommend? An algorithm audit of reputation signals in LLM-assisted hotel selection](https://arxiv.org/abs/2606.16344), arXiv:2606.16344.
- Martinez (2026), [Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026)](https://arxiv.org/abs/2607.14035), arXiv:2607.14035.
- Underneath (2026), [Do Wikipedia and schema make AI assistants recommend a brand?](https://underneath.agency/research/brand-entity-ai-recommendations-study)

---

This is the Markdown twin of https://underneath.agency/resources/do-ai-assistants-favor-big-brands. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Do pages need author bylines to be cited by ChatGPT? | Underneath"
description: "No. In a ChatGPT health audit, 64.7% of cited pages named no author and only 7.0% gave a name with credentials; publisher authority mattered more."
canonical: "https://underneath.agency/resources/do-ai-cited-pages-need-author-bylines"
published: 2026-10-11
updated: 2026-10-11
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Do pages need author bylines to be cited by ChatGPT?

No. In a published audit of the health pages ChatGPT cites, most named no author at all, and only a small minority showed an author with credentials. ChatGPT leaned instead on who published the page, and the study cannot say whether adding a byline would change anything.

## The short version

1. In an audit of 615 sources ChatGPT cited for health questions, 64.7% showed no author attribution (Jacques and colleagues, City University of New York, January 2026).
2. Another 28.3% gave a name only, and just 7.0% gave a name with credentials such as MD or PhD.
3. ChatGPT’s most cited sources were Wikipedia (10.7%), Mayo Clinic (9.9%) and Cleveland Clinic (9.8%), which carry institutional rather than personal authority.
4. In our own study of 3,096 pages ranking for AI Overview searches, an author signal’s apparent edge of 4.8 points disappeared once pages on the same search were compared.

## How many ChatGPT-cited pages show an author?

About one in three. Most cited pages in the audit named nobody, and complete credentials were rare.

[Jacques and colleagues](https://arxiv.org/abs/2601.17109) drew 100 questions at random from HealthSearchQA, a set of 3,173 consumer health questions that Google Research compiled from real search suggestions. Each went into ChatGPT 5.2 Pro in a fresh account, with a request to include sources with links. The team then coded every cited page against an “Authority Signals Framework”, whose first question is “Who wrote it?”

| Author attribution on cited pages | Sources | Share |
|---|---|---|
| No author named | 398 | 64.7% |
| Name only | 174 | 28.3% |
| Name with credentials | 43 | 7.0% |

So 43 of the 615 cited sources met the standard many content guidelines treat as essential. The audit is about health, where expertise matters more than almost anywhere, which makes the result more striking.

## Which cited sources skip bylines, and why does it work for them?

Large institutions, whose name does the work a byline would do for a smaller publisher.

Citations were heavily concentrated. Ten organizations accounted for 52.8% of all health citations, led by Wikipedia at 10.7%, Mayo Clinic at 9.9% and Cleveland Clinic at 9.8%. Wikipedia is written by many contributors rather than one named author. Hospitals and government agencies carry the authority of their own names. The same pattern of [a few sites winning most citations](https://underneath.agency/resources/ai-search-citation-concentration) shows up in AI Mode and Perplexity.

Overall, 75.7% of citations came from sources with inherent institutional authority. The authors conclude that these sources rely on their standing rather than on explicit signals. Authorship, editorial process and technical optimization acted as secondary signals behind “who published it”.

Our own research points the same way outside health. In our [Wikipedia and schema study](https://underneath.agency/research/brand-entity-ai-recommendations-study), the strongest predictor of a brand recommendation was independent coverage. That means how many independent sites named the brand in the pages an assistant cited. Each tenfold increase went with 4.7 times the odds of a recommendation. Reputation built across the web, not a line on the page, carried the weight.

## If not bylines, what did cited pages show instead?

Fewer quality markers than most guidelines assume, and technical polish more often than editorial ones.

| Signal on ChatGPT-cited health pages | Share |
|---|---|
| Schema markup | 74.3% |
| Lists references | 39.2% |
| States a medical review | 29.6% |

Schema markup, the machine-readable labels that describe a page to software, was the most common signal. Editorial signals such as references and a stated medical review were present on a minority of pages.

[Commercial health publishers](https://underneath.agency/resources/chatgpt-cited-commercial-health-sites) were the exception. Lacking an institutional name, they stated a medical review on 71.1% of their cited pages. That suggests a reviewer statement may matter more to a publisher without a famous brand, though the study does not test it.

The audit also ran a sample of cited pages through an AI-text detector. Of 305 pages it could classify, 22.6% were judged AI-derived. A named human author was plainly not a condition of being cited.

## Do bylines help in Google’s AI Overviews?

Not measurably, once pages competing on the same search are compared.

Our [study of 3,096 pages](https://underneath.agency/research/ai-overview-cited-pages-study) ranking in Google’s top 10 compared the pages AI Overviews cited with those beside them that were not cited. A first, simpler comparison suggested pages with an author signal were cited 4.8 points more often. Comparing pages only with others on the same search, with all features entered together, turned that into −2.1 points, which is indistinguishable from no effect.

Ranking position, by contrast, mattered a great deal. Pages in positions 1 to 3 were cited about twice as often as those at 7 to 10. This is Google’s AI, not ChatGPT, but it is the only published comparison of cited and uncited pages on authorship we know of. We cover the other [on-page signals linked to AI citations](https://underneath.agency/resources/on-page-signals-linked-to-ai-citations) separately.

## So are author bylines pointless?

No. They still serve readers and supporting evidence; there is simply no proof that they drive AI citation.

A byline with real credentials helps a human reader judge a page, and buyers often check sources after an AI answer. The audit authors, citing a 2006 national survey, note that 88% of US adults lack proficient health literacy, so many readers struggle to judge credibility at all.

Evidence inside the page appears to count for more than a name above it. In a 252,000-trial controlled test by [Vishwakarma and colleagues](https://arxiv.org/abs/2605.25517) at Sprinklr, a software vendor, AI models preferred pages that backed claims with evidence such as tests or certifications. Publisher names were removed in that test, so the content alone made the difference.

## What should you do about it?

Keep honest bylines where you have real experts, but do not treat them as the lever for AI visibility.

1. Do not hold back publishing because a page lacks a credentialed author; cited pages mostly had none.
2. Where a qualified person wrote or reviewed a page, name them with credentials and a review date. It helps readers and costs little.
3. Put the evidence in the text: references, figures, test results and clear sourcing for claims.
4. Build the publisher’s standing through independent coverage and search ranking, which the research ties more closely to citation.
5. If you are a smaller publisher in an expert field, show your vetting process openly, as cited commercial health sites do.

If you want help deciding where those efforts pay off for your category, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

No study has added bylines to real pages and measured the effect on ChatGPT citations.

- The ChatGPT audit describes cited pages only; it has no group of uncited pages to compare with.
- It covers health questions on one ChatGPT model, collected on January 11, 2026, with a prompt that asked for sources.
- Attribution was coded largely by an automated pipeline, with some fields checked by hand.
- Our AI Overview comparison covers Google’s AI, one day and US searches, and detects only author markup and on-page signals.
- Whether named experts help more in finance, law or business software is untested.

## Frequently asked questions

### Does ChatGPT check author credentials before citing a page?

There is no evidence that it requires them. In an audit of 615 ChatGPT-cited health sources, only 7.0% showed an author name with credentials, and 64.7% named no author.

### Is E-E-A-T important for AI search?

The studies here did not test Google’s E-E-A-T guidelines directly. They do show that institutional reputation, independent coverage and evidence on the page went with citation far more clearly than bylines did.

### Should I add author bios to my blog for AI visibility?

Add them if they are real and useful to readers, not as an AI tactic. Our study of AI Overview citations found no measurable edge for author signals once pages on the same search were compared.

### What matters more than a byline for AI citations?

Who publishes the page and how well it ranks. In the ChatGPT health audit, 75.7% of citations went to institutional sources, and in our Google study, ranking position outweighed every on-page feature.

## Sources

- Jacques, Datuowei, Jones, Basch, Vanderpool, Udeozo and Chapa (2026), [Authority Signals in AI Cited Health Sources: A Framework for Evaluating Source Credibility in ChatGPT Responses](https://arxiv.org/abs/2601.17109), arXiv:2601.17109.
- Vishwakarma, Kumar and Jamidar (2026), [What Gets Cited: Competitive GEO in AI Answer Engines](https://arxiv.org/abs/2605.25517), arXiv:2605.25517.
- Underneath (2026), [What pages cited by AI Overviews have in common: 3,096 pages](https://underneath.agency/research/ai-overview-cited-pages-study)
- Underneath (2026), [Do Wikipedia and schema make AI assistants recommend a brand?](https://underneath.agency/research/brand-entity-ai-recommendations-study)

---

This is the Markdown twin of https://underneath.agency/resources/do-ai-cited-pages-need-author-bylines. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Do AI engines cite local-language sources in other languages?"
description: "Mostly yes: ChatGPT and Perplexity switch to local-language sources, while Claude kept citing English sites in one study. Translation alone is not enough."
canonical: "https://underneath.agency/resources/do-ai-engines-cite-local-language-sources"
published: 2026-10-10
updated: 2026-10-10
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Do AI engines cite local-language sources in other languages?

Mostly yes: when asked in another language, most AI engines switch to websites in that language, but how far they switch depends on the engine. In one 2025 study, ChatGPT and Perplexity sourced almost entirely from the local-language web, while Claude kept reusing English sites. For brands selling abroad, that means earning coverage in each market’s own media, not only translating the website.

## The short version

1. In a 2025 University of Toronto study of 100 buyer questions in five languages, ChatGPT and Perplexity drew almost entirely on local-language sites, while Claude reused English authority sites.
2. In Tokyo, Gemini cited Japanese websites 68.4% of the time for Japanese questions and 6.4% for English questions about the same hotels.
3. Hotel websites made up 11.0% of Gemini’s Japanese citations against 8.2% of English ones, which the authors link to richer Japanese content.
4. Across twelve European languages, a 2026 study found that asking in a brand’s home language raised its recommendation share by 0.80 on a 0-to-1 scale for local champions, against 0.15 for global brands.
5. Even within English, country matters: in our test, ChatGPT answers outside the US that cited a local website named local brands 45.0% of the time, against 3.1% when they cited none.

## Do AI engines switch to local-language websites?

Most do, but to very different degrees. The engine you ask matters as much as the language.

[Chen and colleagues](https://arxiv.org/abs/2509.08919) at the University of Toronto translated 100 English buyer questions, ten in each of ten consumer categories, into Chinese, Japanese, German, French and Spanish. They put them to Google, Gemini, Claude, ChatGPT and Perplexity and checked the language of every cited website. Their summary: “GPT and Perplexity heavily localize, sourcing almost entirely from the target language’s ecosystem. Claude, by contrast, reuses English-language authority domains across languages.”

Gemini sat in between. For German and Spanish questions, close to 50% of its cited websites were still in English. Across engines, Japanese and French questions produced the strongest switch to local sites.

## How different are the sources in each language?

Very different for most engines: the same question in two languages often shares almost no cited websites. Claude is the exception.

In the Toronto study, Google’s overlap between the English and translated versions peaked at about 0.11, meaning around a tenth of websites were shared. ChatGPT’s overlap was near zero everywhere, as it moved to a different set of sites in each language. Gemini reached about 0.32 at best, for English and German, and Perplexity about 0.22, for German laptop questions. Claude showed much higher overlap in every category.

Brand lists moved less than sources. The authors found brand overlap “much higher than domain overlap”, especially in categories led by global brands such as cameras and laptops. Different sources can still lead to similar recommendations when a few brands dominate worldwide.

## What happens in a real local market?

A study in Tokyo shows two almost separate web worlds, one English and one Japanese. Gemini drew on whichever matched the question’s language.

[Zhu and Chang](https://arxiv.org/abs/2603.20062) asked Gemini 2.5 Flash 156 hotel questions about Tokyo in English and Japanese in March 2026, collecting 1,357 citations. Japanese questions cited Japanese websites 68.4% of the time; English questions only 6.4%. The kinds of sources differed too. Travel blogs were 22.9% of English non-booking-site citations but 1.2% of Japanese ones, while Japanese travel agencies took 12.9% and appeared not at all in English.

Hotels’ own websites did better in Japanese: 11.0% of Japanese citations against 8.2% of English ones. The authors note that Japanese hotel sites tend to carry deeper content, such as neighborhood guides and transit information, while English versions focus on booking. The study covers one city, one engine and one month, and it is correlational.

## Does the question’s language change which brands get recommended?

Yes, strongly for local brands: a home-language question can turn a missing brand into a default pick. Global brands move much less.

[Żatuchin](https://arxiv.org/abs/2606.23165), who is also affiliated with an AI brand-monitoring company, asked three AI engines about 66 European brands in twelve languages, collecting 35,640 answers in April and May 2026. Switching from English to a brand’s home language raised its recommendation share by 0.80 on a 0-to-1 scale for local champions, but only 0.15 for global brands. [An English-only check](https://underneath.agency/resources/english-only-ai-visibility-audits) would understate a local champion’s visibility.

The same study found the overall mix of source types broadly similar across languages, with Wikipedia the most cited website in 11 of the 12. What changed was which specific sites and brands appeared, not the general type of source.

## Does the number of sources change by language too?

Sometimes, and not in the same direction for every engine. Language can change how many citations an answer carries.

In an analysis of a public dataset by [Zhang, He and Yao](https://arxiv.org/abs/2604.25707), ChatGPT averaged 7.77 citations for Chinese questions and 7.03 for English. Google’s AI search showed the opposite: 11.57 citations in English against 7.53 in Chinese. The authors conclude that language effects have to be measured engine by engine. Chinese-built AI models are a separate question again; see [how brands fare in Chinese AI models](https://underneath.agency/resources/chinese-vs-western-ai-brand-visibility).

## Does the same apply between English-speaking countries?

Yes: even in English, the user’s country shifts the sources cited, and local sources go with local brands. Location matters, not just language.

In [our country study](https://underneath.agency/research/ai-recommendations-by-country-study), we asked ChatGPT and Gemini the same 40 English buyer questions from the US, UK, Canada and Australia on 28 September 2026. Two ChatGPT answers from the same country shared 0.534 of their cited websites; answers from different countries shared 0.324. Outside the US, ChatGPT answers citing at least one local website had a local-brand share of 45.0%, against 3.1% for answers citing none. This is an association: we did not change the sources to test cause.

## What should you do about it?

Build credibility in each market’s own media and test in each language, rather than relying on translated pages. Concretely:

1. List the sites each engine cites for your category in each target language and country.
2. Earn coverage from respected local publishers and review sites, not just English trade press.
3. Give local pages real substance, such as local guides, prices and service details, not just translations. Our guide to [GEO across languages for global brands](https://underneath.agency/resources/geo-strategy-for-global-brands-across-languages) covers market-by-market planning.
4. Keep English authority strong too, since some engines, such as Claude in one study, reuse English sources.
5. Track AI answers in every language and country you sell in; an English-only view can mislead.

If you want help planning multilingual AI visibility, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The research shows that engines localize differently, but not why, or how stable that is. The gaps:

- Each study covers one period; engine behavior changes, and the 2025 Toronto results may already be out of date.
- The language studies use buyer-ranking questions, hotel questions or brand questions, not every kind of query.
- No study we reviewed tests whether earning local coverage causes more AI citations; the evidence is observational.
- Coverage of languages is uneven: five major languages, Japanese hotels and twelve European languages.
- One key multilingual study has an author affiliated with a monitoring company.

## Frequently asked questions

### Does ChatGPT use local sources when you ask in another language?

In one 2025 study, yes: it sourced almost entirely from the target language’s websites. Its cited sites barely overlapped between English and other languages.

### Is translating my website enough for AI search abroad?

Probably not. The Toronto authors conclude that earned coverage in the local language is needed for engines that localize, such as ChatGPT and Perplexity.

### Which AI engine relies most on English sources?

Claude, in the 2025 Toronto study. It reused English-language authority sites across all five languages tested.

### Does asking from another country change AI answers even in English?

Yes. In our test, ChatGPT answers from the same country shared 0.534 of their cited websites, against 0.324 across countries.

## Sources

- Chen and colleagues (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919), arXiv:2509.08919.
- Zhu and Chang (2026), [The End of Rented Discovery: How AI Search Redistributes Power Between Hotels and Intermediaries](https://arxiv.org/abs/2603.20062), arXiv:2603.20062.
- Żatuchin (2026), [The Language Blind Spot: How Query Language and Brand Recognition Tier Shape AI-Constructed Brand Reputation Across Twelve European Languages](https://arxiv.org/abs/2606.23165), arXiv:2606.23165.
- Zhang, He and Yao (2026), [From Citation Selection to Citation Absorption: A Measurement Framework for Generative Engine Optimization Across AI Search Platforms](https://arxiv.org/abs/2604.25707), arXiv:2604.25707.
- Underneath (2026), [Same question, four countries: do AI recommendations change?](https://underneath.agency/research/ai-recommendations-by-country-study)

---

This is the Markdown twin of https://underneath.agency/resources/do-ai-engines-cite-local-language-sources. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Do AI engines cite your own website or third-party reviews?"
description: "Mostly third-party sources. Brand-owned sites are a small share of AI citations in most studies, though the share rises for buying-ready questions."
canonical: "https://underneath.agency/resources/do-ai-engines-cite-your-own-website"
published: 2026-10-10
updated: 2026-10-10
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Do ChatGPT and other AI engines cite my own website or third-party reviews?

AI engines lean mainly on third parties: review publishers, news outlets, review platforms and other companies’ pages. Your own website is usually a small share of what they cite. The share grows when the question is about buying, and one engine, Perplexity, cites brand sites far more than the others.

## The short version

1. In US consumer electronics questions, AI search drew 92.1% of its sources from earned media such as reviews and publications, [an August 2025 study found](https://arxiv.org/abs/2509.08919).
2. In a vendor dataset of brand prompts, only 2.9% of citations pointed to the tracked brand’s own site, while 75.2% went to other companies in the same space, [Kumar reports](https://arxiv.org/abs/2606.20065).
3. When asked whether a brand is legit, Perplexity cited the brand’s own site in 94.9% of answers and Google AI Mode in 17.7%, in [our reputation study](https://underneath.agency/research/is-it-legit-ai-reputation-study).
4. Independent coverage was the strongest signal we measured for being recommended, in [our brand entity study](https://underneath.agency/research/brand-entity-ai-recommendations-study).

## Do AI engines cite brand websites or third-party sources?

Mostly third-party sources, across the studies we reviewed. [Chen and colleagues](https://arxiv.org/abs/2509.08919) compared Google with AI search on ranking-style product questions in August 2025. For US consumer electronics, Google’s results gave brand sites 32.9% of their sources. AI search drew 92.1% from earned media, with negligible social media.

The same tilt held for brands of every size. For well-known brands, ChatGPT drew 93.5% of its sources from earned media and 6.5% from brand sites. For niche brands, it drew 95.1% from earned media and 4.9% from brand sites.

Treat the exact figures with care. The study queried developer versions of the engines, not the consumer apps, and used an AI model to label sources. The authors also thank a GEO vendor for support.

## How small is a brand’s own share of citations?

Very small in most commercial datasets. [Kumar](https://arxiv.org/abs/2606.20065), whose company sells AI visibility tracking, analyzed 149,912 citations from brand prompts across five engines in 2026. Only 2.9% of citations pointed at the brand’s own domain, while 75.2% pointed at other companies in the same space.

In other words, the “corporate” sites AI engines cite are usually your competitors and peers. The author notes that engines often build “alternatives” answers, and those lean on peer-brand product pages. The dataset skews toward software, fintech and Indian direct-to-consumer brands.

For product questions in Europe, brand sites play a larger role in Google. [Uberti-Bona Marin and colleagues](https://arxiv.org/abs/2609.18729) found manufacturers and brands were 16.1% of the domains Google’s AI Overviews displayed. Editorial and product-review sources were 25.4%.

## Does the answer change with the type of question?

Yes. Brand sites gain ground when buyers are ready to act. Chen and colleagues found that on transactional questions, brand content rose in prominence across Google, ChatGPT and Perplexity.

Hotels show the same split. [Zhu and Chang](https://arxiv.org/abs/2603.20062), who work at an AI company, audited 1,357 citations from Gemini for 156 Tokyo hotel questions in March 2026. Booking sites took 55.3% of all citations. Experiential questions drew 55.9% of citations from other sources, against 30.8% for transactional ones.

Hotel websites themselves were a small slice: 8.2% of citations for English questions and 11.0% for Japanese ones. The authors argue AI search does not prefer earned media as such. It cites whatever content type best answers the question. Whether that loosens [reliance on booking sites and marketplaces](https://underneath.agency/resources/ai-search-marketplace-dependence) is covered separately.

Banking is another case. In Chen’s bank questions, earned sources were 64.6% of citations and bank-owned sites 34.1%.

## Does it depend on which AI engine is answering?

Yes, sharply. In [our reputation study](https://underneath.agency/research/is-it-legit-ai-reputation-study), we asked four engines whether 79 brands were legit. Perplexity cited the brand’s own site in 94.9% of answers. Google AI Mode did so in 17.7%.

Even when an engine cites your site, reviews shape the verdict. In the same study, 88.0% of answers cited a review or complaint platform. Claims attached only to review platforms were negative 56.5% of the time, against 6.4% for claims attached only to the brand’s own website.

Engines also differ for product questions. In the Uberti-Bona Marin audit, editorial and product-review sources made up 56.7% of the domains ChatGPT displayed, far more than in Google’s AI Overviews.

## Which third-party sources carry the most weight?

Review publishers, review platforms and ranked “best of” lists. In Kumar’s vendor data, ranked list articles were 35.7% of the content pages engines cited. A single list can put a brand into many answers at once. Which kinds of sites lead depends on the question, as [the types of websites AI engines rely on](https://underneath.agency/resources/which-websites-do-ai-search-engines-cite) shows.

Some of those lists are written by the brands themselves. In [our study of self-promoting lists](https://underneath.agency/research/self-promoting-best-lists-study), 24.2% of AI-cited numbered “best” lists with an identifiable publisher ranked that publisher first. Such lists were only 1.1% of all citations, so they are a minor route.

Assistants also go looking for reviews. In [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study), ChatGPT looked for reviews in 46.2% of answers. When its search named a source such as NerdWallet or Avvo, the answer cited that source 44.0% of the time.

## Does third-party coverage change whether you are recommended?

It is the strongest signal we have measured, though not proven to cause recommendations. In our brand entity study, each tenfold increase in the number of independent sites naming a brand in the cited pages went with 4.7 times the odds of being recommended.

Other researchers reach a similar view but without a direct test. [Kumar and Palkhouski](https://arxiv.org/abs/2509.10762) argue that even high-quality pages may not be cited if they sit only on vendor blogs. They also note they did not vary where content was published, so the off-site effect remains untested in their work.

## What should you do about it?

Treat third-party coverage as the main route into AI answers, and your own site as the source of record.

1. Map which review sites, publications and lists the engines cite for your category.
2. Earn coverage there through product reviews, analyst inclusion and press, not paid placements dressed as editorial. On paid reach more broadly, see [whether advertising spend helps AI recommendations](https://underneath.agency/resources/does-ad-spend-help-ai-brand-recommendations).
3. Watch review platforms closely. They carry most of the negative claims engines repeat.
4. Keep your own site clear and factual for transactional questions, where brand pages gain ground.
5. Check Perplexity separately, since it cites brand sites far more often than other engines.

Our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) includes this kind of source audit.

## What does the research not tell us yet?

No study we reviewed shows that adding third-party coverage causes a brand to be cited more.

- The evidence is observational: brands with more coverage may differ in other ways.
- “Earned”, “brand” and “editorial” are labeled differently across studies, often by AI models.
- Several large datasets come from vendors or companies with a stake in the answer.
- Most studies cover consumer products, hotels or software; many other industries are untested.
- We found no study measuring how much a brand-site citation is worth in traffic or sales.

## Frequently asked questions

### Does ChatGPT cite company websites?

Sometimes, but rarely as the main source. In one study, ChatGPT drew 6.5% of its sources from brand sites for well-known brands and 93.5% from earned media.

### Why does AI cite my competitors’ pages instead of mine?

Engines often build “alternatives” answers from peer-brand pages. In one vendor dataset, 75.2% of citations went to other companies in the same space and 2.9% to the brand itself.

### Do review sites like Trustpilot affect AI answers?

Yes, they shape the tone. In our reputation study, 88.0% of answers cited a review or complaint platform, and claims resting only on those platforms were negative 56.5% of the time.

### Is PR more important than my website for AI search?

For discovery questions, third-party coverage matters more in most studies. Your own site gains weight on transactional questions and on Perplexity, which cited brand sites in 94.9% of reputation answers.

## Sources

- Chen et al. (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919), arXiv:2509.08919.
- Kumar (2026), [Generative Engine Optimization at Scale: Measuring Brand Visibility Across AI Search Engines](https://arxiv.org/abs/2606.20065), arXiv:2606.20065.
- Uberti-Bona Marin et al. (2026), ["If I Had to Buy Just ONE: Galaxy S26 Ultra": Auditing AI-Generated Product Recommendations](https://arxiv.org/abs/2609.18729), arXiv:2609.18729.
- Zhu and Chang (2026), [The End of Rented Discovery: How AI Search Redistributes Power Between Hotels and Intermediaries](https://arxiv.org/abs/2603.20062), arXiv:2603.20062.
- Kumar and Palkhouski (2025), [AI Answer Engine Citation Behavior An Empirical Analysis of the GEO16 Framework](https://arxiv.org/abs/2509.10762), arXiv:2509.10762.
- Underneath (2026), [“Is this brand legit?” How AI assistants build a reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study)
- Underneath (2026), [Do Wikipedia and schema make AI assistants recommend a brand?](https://underneath.agency/research/brand-entity-ai-recommendations-study)
- Underneath (2026), [How many “best of” lists cited by AI rank their own brand first?](https://underneath.agency/research/self-promoting-best-lists-study)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)

---

This is the Markdown twin of https://underneath.agency/resources/do-ai-engines-cite-your-own-website. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Do Gemini, GPT and Claude prefer the same kind of content?"
description: "Mostly yes on content quality: Gemini, GPT and Claude share most preferences. But they cite different sources, so one good page will not show up everywhere."
canonical: "https://underneath.agency/resources/do-ai-engines-prefer-same-content"
published: 2026-10-11
updated: 2026-10-11
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Do Gemini, GPT and Claude prefer the same kind of content?

Mostly, yes: in controlled tests, Gemini, GPT and Claude rewarded largely the same qualities in a page. Where they differ is in which sources they go looking for, how many they cite and how sensitive they are to changes. One well-built page can serve all of them, but showing up in each engine still depends on where that engine searches.

## The short version

1. In a lab test, 78.95% to 84.21% of the content preferences found for Gemini, GPT and Claude were shared between each pair ([Wu and colleagues](https://arxiv.org/abs/2510.11438)).
2. Topic moved preferences more than engine did: under Gemini, rules for shopping questions overlapped with rules for research questions by only 34.78% to 40.00%.
3. Engines differ in sources: 93.5% of ChatGPT’s sources for well-known brands were independent publishers, while 25.1% of Gemini’s were brand sites ([Chen and colleagues](https://arxiv.org/abs/2509.08919)).
4. In [our study of 80 buyer questions](https://underneath.agency/research/ai-citations-google-rankings-study), 8.3% of ChatGPT’s citations ranked in Google’s top 10 for the question, against 25.7% of Claude’s.
5. In [our brand agreement study](https://underneath.agency/research/ai-assistants-brand-agreement-study), 66.3% of the options recommended for a question came from one assistant only.

## How much do the engines’ content preferences overlap?

A lot. Researchers at Carnegie Mellon built test AI search engines on Gemini, GPT and Claude ([Wu and colleagues](https://arxiv.org/abs/2510.11438)). They then had AI models explain each engine’s choices and boil them down into rules about what it preferred in a page.

On the same set of research questions, Gemini and GPT shared 78.95% of their rules. Gemini and Claude shared 84.21%, as did GPT and Claude. Each engine kept a few preferences of its own, and rules tuned to one engine worked best on that engine.

Topic mattered more than engine. Under Gemini, rules for two sets of open research questions overlapped by 88.24%, but rules for shopping questions overlapped with them by only 34.78% to 40.00%. Shopping rules leaned toward practical guidance over in-depth explanation.

Read these figures with care. The engines were smaller, cheaper versions (Gemini 2.5 Flash-Lite, GPT-4o mini and Claude 3 Haiku), each choosing among five documents handed to it. The overlap was measured on keywords the researchers labeled by hand.

This is a simulation, not a test of live ChatGPT or Gemini.

## What do all of them seem to reward?

Substance, stated clearly. Rules shared by all three engines included comprehensive coverage, credible sources, accurate and current facts, and a neutral tone ([Wu and colleagues](https://arxiv.org/abs/2510.11438)). They also shared clear structure, specific evidence and stating the conclusion at the start. We describe [what AI-preferred content looks like](https://underneath.agency/resources/what-content-do-ai-engines-prefer) in more detail.

A separate test by researchers at the software company Sprinklr points the same way. They ran 252,000 trials on six AI models, changing one thing at a time in two competing sources ([Vishwakarma and colleagues](https://arxiv.org/abs/2605.25517)). Eleven of 18 content factors, 61%, mattered in at least four of the six models.

Four mattered in every model: matching the topic, stating a price, carrying a recent date and appearing earlier in the list of sources. Layout changes alone had no consistent effect. Those results are unpacked in [why an AI cites a rival’s page first](https://underneath.agency/resources/why-ai-cites-competitor-page-first).

## Where do the engines differ?

They differ in where they look, how many sources they show and how easily they are moved. University of Toronto researchers classified the sources four AI engines cited for brand questions ([Chen and colleagues](https://arxiv.org/abs/2509.08919)).

| Engine | Independent publishers | Brand sites | Social and community |
|---|---|---|---|
| ChatGPT | 93.5% | 6.5% | 0% |
| Claude | 87.3% | 6.8% | 5.9% |
| Gemini | 63.4% | 25.1% | 11.5% |
| Perplexity | 67.4% | 8.8% | 23.8% |

These figures are for well-known brands.

The engines also differ in how many sources they cite. In one study of 602 test questions, ChatGPT cited 6.88 sources per question on average and Perplexity 16.35 ([Zhang Kai and colleagues](https://arxiv.org/abs/2604.25707)). ChatGPT, though, drew more heavily on each source it cited.

They differ in what signals they respond to. In that study, ChatGPT responded most to relevance to the question. Google responded to closeness in meaning to the question and answer, plus definitions; Perplexity to relevance, headings and length.

In the Sprinklr test, 50% of factors mattered for Claude 3.5 and only 33% for Gemini 2.5.

## If preferences overlap, why do the engines cite different pages?

Partly because they search differently, and partly because they choose differently from the same evidence. In [our study of 80 buyer questions](https://underneath.agency/research/ai-citations-google-rankings-study), 8.3% of ChatGPT’s citations ranked in Google’s top 10 for the question, against 16.6% for Gemini and 25.7% for Claude.

The result is little shared ground. In [our brand agreement study](https://underneath.agency/research/ai-assistants-brand-agreement-study), 66.3% of options recommended for a question came from one assistant only. All four assistants agreed on the first pick for 10.0% of questions.

Different sources explained little of this: when one assistant recommended an option and another did not, the second assistant’s own cited pages named it 42.1% of the time.

Other audits agree. A 2026 survey reports one audit in which only 26% of cited domains were cited by both Bing Chat and Perplexity ([Martinez](https://arxiv.org/abs/2607.14035)). A shared taste in content does not mean a shared shortlist.

## Does a page good enough for one engine work for the others?

On one study’s scoring, pages cited by several engines were of higher quality. A study of business software questions audited 1,100 pages cited by Brave, Google’s AI Overviews and Perplexity ([Kumar and Palkhouski](https://arxiv.org/abs/2509.10762)).

The engines cited pages of very different quality on the authors’ own scoring: an average of 0.727 for Brave against 0.300 for Perplexity. Pages cited by more than one engine scored 71% higher on quality than pages cited by just one. The authors have a commercial research affiliation, and the score is their own.

## Does page structure need tailoring for each engine?

Probably not much. A team from Japanese universities grouped six engines, including older products such as Bing Chat, by how they search ([Yu and colleagues](https://arxiv.org/abs/2603.29979)). They then restructured 200 articles for them.

The same rewrites raised citation rates by 17.3% on average across all groups.

The authors suggest each kind of engine leans on different features, such as a clear summary up front or sections that stand alone. Those weights come from their own predictive model and their own grouping of engines, and a recent survey cautions that the work relies on automated judges. Treat them as hypotheses, not settings to tune.

## What should you do about it?

Build one strong page, then work on each engine’s route to it.

1. Write for the shared preferences: a direct answer first, full coverage, specific evidence, credible sources, accurate dates and prices, and a neutral tone. On wording itself, see [whether readable writing helps AI visibility](https://underneath.agency/resources/does-readable-writing-help-ai-visibility).
2. Do not fork content per engine. Most preferences are shared; tailored rules did best in lab tests, but shared rules still helped.
3. Map where each engine looks. ChatGPT leans on independent publishers, Gemini cites more brand sites, Perplexity more community sources.
4. Track each engine separately. A win in one does not carry over, so measure citations engine by engine.
5. Ask your category’s buying questions in each assistant every few weeks and note which sources each one cites.

If you want help building pages that work across engines, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

No study has compared live versions of Gemini, ChatGPT and Claude on the same pages over time.

- The overlap figures come from small, older model versions choosing among five documents, not from live products.
- Content preference tests feed engines a fixed set of pages, so they say little about how each engine searches.
- Source mixes come from specific question sets, countries and dates, and engines change after updates.
- Several studies are vendor-authored or use the authors’ own scoring tools.
- Nobody has measured whether the same page earns the same clicks or trust from each engine’s users.

## Frequently asked questions

### Do I need different content for ChatGPT, Gemini and Claude?

Mostly no. In a lab test, 78.95% to 84.21% of content preferences were shared between each pair of engines.

### Why does ChatGPT cite different sources than Gemini?

They search differently. For well-known brands, 93.5% of ChatGPT’s sources were independent publishers, while Gemini cited brand sites 25.1% of the time.

### Does ranking in Google help with every AI assistant equally?

No. In our study, 25.7% of Claude’s citations ranked in Google’s top 10 for the question, against 8.3% of ChatGPT’s.

### Will optimizing for one AI engine hurt my visibility in another?

Not on current evidence. Rules learned from one engine still improved results on others in lab tests, though tailored rules worked best on their own engine.

## Sources

- Wu, Zhong, Kim and Xiong (2025), [What Generative Search Engines Like and How to Optimize Web Content Cooperatively](https://arxiv.org/abs/2510.11438), arXiv:2510.11438.
- Vishwakarma, Kumar and Jamidar (2026), [What Gets Cited: Competitive GEO in AI Answer Engines](https://arxiv.org/abs/2605.25517), arXiv:2605.25517.
- Chen, Wang, Chen and Koudas (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919), arXiv:2509.08919.
- Zhang Kai, He Xinyue and Yao Jingang (2026), [From Citation Selection to Citation Absorption: A Measurement Framework for Generative Engine Optimization Across AI Search Platforms](https://arxiv.org/abs/2604.25707), arXiv:2604.25707.
- Kumar and Palkhouski (2025), [AI Answer Engine Citation Behavior An Empirical Analysis of the GEO16 Framework](https://arxiv.org/abs/2509.10762), arXiv:2509.10762.
- Yu, Yang, Ding and Sato (2026), [Structural Feature Engineering for Generative Engine Optimization: How Content Structure Shapes Citation Behavior](https://arxiv.org/abs/2603.29979), arXiv:2603.29979.
- Martinez (2026), [Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026)](https://arxiv.org/abs/2607.14035), arXiv:2607.14035.
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude cite pages that rank?](https://underneath.agency/research/ai-citations-google-rankings-study)
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude agree on brands?](https://underneath.agency/research/ai-assistants-brand-agreement-study)

---

This is the Markdown twin of https://underneath.agency/resources/do-ai-engines-prefer-same-content. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Do Google AI Overviews downplay negative content? | Underneath"
description: "In one audit of 11,000 Google questions, AI Overviews drew less from negatively toned sources. Brand reviews were not tested, and criticism still appears."
canonical: "https://underneath.agency/resources/do-ai-overviews-downplay-negative-content"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Do Google AI Overviews downplay negative content?

In one large audit, yes: Google’s AI Overviews drew noticeably less of their text from negatively toned sources than from neutral or positive ones. That audit used general questions, not brand or review searches, and AI answers about brands still contain plenty of criticism. The realistic risk is not that complaints vanish, but that they end up in the citation list, further down the answer, or left out.

## The short version

1. In an audit of 11,000 real Google questions, AI Overview text drew 13.8 points less than expected from the negatively toned pages it cited (University of Illinois researchers).
2. The same summaries drew 22.1 points less from the forums and social sites they cited, which is where many customer complaints live.
3. When an AI Overview claim is unsupported, it is about 2.6 times more likely to state something no cited page mentions than to contradict a page (Washington University researchers, 98,020 claims).
4. In our test of 79 brands, AI answers raised at least one problem 99.7% of the time, but Gemini and Google AI Mode put a negative point in the first paragraph only 1.3% of the time.

## Do AI Overviews give negative sources less weight?

In one large audit, yes: negatively toned sources were under-used in the summary text. [Huang and colleagues at the University of Illinois](https://arxiv.org/abs/2603.16138) ran 11,000 real search questions, drawn from a public collection of Google queries, through Google and ChatGPT search. Only 57.8% of the questions triggered an AI Overview, the AI summary at the top of Google’s results.

For a sample of answers, the researchers split each summary into single facts and checked which cited pages supported each one. That shows how much of each cited page made it into the summary. Compared with an even split, the summaries used cited sources in these ways:

| Kind of cited source | Use in the AI Overview text |
|---|---|
| Negative tone | 13.8 points less than average |
| Positive tone | 1.1 points more than average |
| Social media and forums | 22.1 points less than average |
| Wikipedia | 5.4 points more than average |

ChatGPT search showed no reliable link between a source’s tone and how much it was used. The authors concluded that “Google AIO’s synthesis implicitly filters against negative content.”

## Does that mean Google hides bad reviews about your brand?

Not as far as the evidence goes, because the audit did not test brand or review searches. The questions were general information questions from a collection first gathered around 2018. They were run from one location, Urbana, Illinois, at one point in time.

The tone of each source was scored by automated tools, not people, and the study shows a pattern, not its cause. The authors note that an under-used source may simply be lower quality, repetitive or less relevant. A negative page could be under-used for any of those reasons.

AI Overviews are not uniformly upbeat either. Compared with ChatGPT answering without search, their wording used 44% fewer positive emotion words but only 25% fewer negative ones. On contested debate questions, [Grossman and colleagues](https://arxiv.org/abs/2604.27790) found that 33.4% of AI Overviews opened with a plain yes or no, against 5.6% of Gemini answers.

## Where do complaints and reviews end up in AI Overviews?

Often in the citation list rather than the summary, because forums are cited but little used. In the Illinois audit, AI Overviews cited Reddit and Quora but drew 22.1 points less content from social and forum sources than from others. Reddit was among the most under-used sources of all.

[Xu and colleagues at Washington University in St. Louis](https://arxiv.org/abs/2605.14021) tracked 55,393 trending searches over 40 days. AI Overviews cited user-generated platforms less often than Google’s own first page did, in all 19 topic categories. In science searches, 56.22% of first-page links came from those platforms, against 15.33% of AI Overview citations.

The same team checked 98,020 individual claims against the pages the AI Overviews cited. 11.0% were [not supported by the cited pages](https://underneath.agency/resources/ai-answers-unsupported-claims), and 4.1% were contradicted by them. Unsupported claims were about 2.6 times more likely to state something no cited page mentioned than to contradict a page.

## Do other AI engines treat negative information the same way?

Partly: some also under-use negative sources, and several place criticism below the opening. In the Illinois team’s check on Perplexity, negatively toned sources were under-used by 10.1 points, close to the AI Overview pattern.

Our [“Is this brand legit?” study](https://underneath.agency/research/is-it-legit-ai-reputation-study) asked ChatGPT, Gemini, Perplexity and Google AI Mode about 79 brands on 26 September 2026. Criticism was everywhere: 99.7% of complete answers made at least one negative claim, and 35.5% of all claims were negative. Placement differed sharply by engine:

| Engine | Answers with a negative claim in the first paragraph |
|---|---|
| Perplexity | 31.6% |
| Gemini | 1.3% |
| Google AI Mode | 1.3% |
| All four engines | 10.9% |

[Kumar](https://arxiv.org/abs/2606.20065), a co-founder of an AI visibility company, reports from tracked brands that 0.0% of brand, question and engine combinations were consistently negative. Tone flipped between positive and negative across runs in 45.5% of them, so negativity came and went rather than settling.

## What does an AI assistant weigh when reviews are mixed?

In one controlled test, star ratings mattered most and replies to reviews barely registered. [Baig and colleagues](https://arxiv.org/abs/2606.16344) asked 12 AI models to pick among fictional hotels whose ratings, reviews, prices and other details were set at random. A 4.7-star hotel was recommended 31.6 points more often than a 3.9-star one.

A visible management response to reviews had no detectable effect, at +0.1 points. This was a simulated choice among five listed hotels, and two of the authors are affiliated with a travel-sector company. Where machine and customer priorities part ways is weighed in [AI versus human reputation priorities](https://underneath.agency/resources/ai-vs-human-reputation-priorities).

Review sites carry most of the criticism in real answers. In our brand study, claims backed only by review or complaint platforms were negative 56.5% of the time. Claims backed only by the brand’s own website were negative 6.4% of the time.

## What should you do about it?

Do not count on AI summaries to hide criticism; fix what customers complain about and check where it appears.

1. Search your brand with words such as “reviews”, “complaints” and “legit” in Google and in the main assistants. Note whether criticism appears, and whether it is in the opening or further down.
2. Treat review platforms as the source of your AI reputation. Fix the recurring complaints there, such as billing and support problems.
3. Protect your ratings. In the hotel test, ratings moved recommendations far more than replies to reviews did.
4. Publish a clear, factual page about who you are and how you handle complaints, so assistants have accurate material for that part of the answer.
5. Check more than once. AI answers vary from run to run, so one result is not your reputation.

For help building this into regular monitoring, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

Nobody has tested whether AI Overviews downplay negative content about brands or products specifically.

- The Illinois audit used general questions from an older collection, one location and automated tone scoring.
- It shows a pattern, not a cause; negative pages may be under-used for reasons other than tone.
- Our brand study covered Google AI Mode, not AI Overviews, with one answer per engine and brand.
- The finding that no tracked brand was consistently described negatively comes from a vendor’s own customer data.
- Whether moving criticism below the opening changes what buyers decide has not been measured.

## Frequently asked questions

### Will Google’s AI Overview show negative reviews of my company?

It may cite them but use little of what they say. In one audit, AI Overviews drew 13.8 points less from negatively toned sources and 22.1 points less from forums, though brand searches were not tested.

### Does Google’s AI favor positive content?

Only slightly, in the one audit that measured it. Positively toned sources were used 1.1 points more than average, while negatively toned ones were used 13.8 points less.

### Why does an AI Overview cite Reddit but not repeat what Reddit says?

AI Overviews tend to list forum pages without drawing much from them. In the Illinois audit, social and forum sources were cited but under-used by 22.1 points in the summary text. Being cited is not the same as being used; see [pages cited for claims they do not make](https://underneath.agency/resources/ai-citing-pages-for-claims-they-do-not-make).

### Should I respond to negative reviews to improve AI recommendations?

Respond for your customers’ sake, but do not expect replies alone to move AI picks. In a test of 12 AI models, management responses had no detectable effect, while a higher star rating added 31.6 points.

## Sources

- Michelle Huang, Agam Goyal, Koustuv Saha and Eshwar Chandrasekharan (2026), [Answer Bubbles: Information Exposure in AI-Mediated Search](https://arxiv.org/abs/2603.16138), arXiv:2603.16138.
- Haofei Xu, Umar Iqbal and Jacob M. Montgomery (2026), [Measuring Google AI Overviews: Activation, Source Quality, Claim Fidelity, and Publisher Impact](https://arxiv.org/abs/2605.14021), arXiv:2605.14021.
- Riley Grossman and colleagues (2026), [How Generative AI Disrupts Search: An Empirical Study of Google Search, Gemini, and AI Overviews](https://arxiv.org/abs/2604.27790), arXiv:2604.27790.
- Pratyush Kumar (2026), [Generative Engine Optimization at Scale: Measuring Brand Visibility Across AI Search Engines](https://arxiv.org/abs/2606.20065), arXiv:2606.20065.
- Mirza Samad Ahmed Baig, Syeda Anshrah Gillani and Asher Ali (2026), [Whose hotel does the AI recommend? An algorithm audit of reputation signals in LLM-assisted hotel selection](https://arxiv.org/abs/2606.16344), arXiv:2606.16344.
- Underneath (2026), [“Is this brand legit?” How AI assistants build a reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study)

---

This is the Markdown twin of https://underneath.agency/resources/do-ai-overviews-downplay-negative-content. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Do Google AI Overviews reduce clicks to my website? | Underneath"
description: "Yes. In a 2026 field experiment, hiding Google AI Overviews raised clicks to outside sites by 8.8 points; browsing data and Wikipedia traffic agree."
canonical: "https://underneath.agency/resources/do-ai-overviews-reduce-clicks"
published: 2026-10-08
updated: 2026-10-10
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Do Google AI Overviews reduce clicks to my website?

Yes, on the best evidence available. In a 2026 field experiment with 1,100 US Google users, hiding AI Overviews raised the rate of clicks to outside websites by 8.8 percentage points. Browsing data from Pew researchers and a study of Wikipedia traffic point the same way, though the size of the loss depends on the site and the searches.

## The short version

1. Hiding AI Overviews raised clicks to outside websites by 8.8 percentage points per search in a preregistered March 2026 experiment with 1,100 US users (Wang and colleagues).
2. In a Pew panel of 900 US adults in March 2025, people clicked a result on 8% of visits to pages with an AI Overview, against 15% on pages without one.
3. People in the same panel ended their browsing session on 26% of pages with an AI Overview, against 16% without.
4. English Wikipedia lost about 5% of its search traffic once AI Overviews became Google’s US default, roughly 100.27 million visits a month (Khosravi and Yoganarasimhan).
5. The clicks lost bought no measurable gain for users: in the 1,100-person experiment, hiding AI Overviews changed neither trust in Google nor satisfaction with search.

## What is the strongest evidence that AI Overviews cut clicks?

A 2026 randomized field experiment on everyday Google use found that hiding AI Overviews raised outside clicks.

[Wang, Gleason, Bart, Wilson and Metaxa](https://arxiv.org/abs/2608.18352) recruited US adults who search with Google in Chrome and installed a browser extension. After three normal days, people were randomly assigned to one of three versions of Google for seven days: AI Overviews hidden, Google unchanged, or every search sent to AI Mode. Before the switch, 36% of their searches showed an AI Overview.

Hiding AI Overviews raised the click-through rate, meaning clicks to outside websites divided by searches, by 8.8 percentage points. The authors also saw suggestive evidence of extra clicks to news publishers. One complication: a Google page change broke the hiding tool partway through, so only 51.1% of AI Overviews were actually hidden. The 8.8-point estimate is adjusted for that.

The sample was not typical of the US: 75% were under 45 and 87% had at least some college education. The test lasted a week, so it says nothing about how habits settle over months.

## What does real-world browsing data show?

Pew’s browsing panel shows people clicked results about half as often when an AI Overview appeared.

[Chapekis, Lieb, Shah and Smith](https://arxiv.org/abs/2608.04831) at Pew Research Center tracked the web browsing of 900 US adults in March 2025, covering 68,879 distinct Google searches. This is observational data, so it shows an association, not proof of cause. The gaps held after the authors allowed for the type of search, such as length and question words.

| What people did next | Page with an AI Overview | Page without one |
|---|---|---|
| Clicked a search result | 8% | 15% |
| Ended their browsing session | 26% | 16% |
| Clicked a source cited in the AI Overview | 1% | not applicable |

A separate industry analysis, reported by [Aral, Li and Zuo](https://arxiv.org/abs/2602.13415), found a median zero-click rate of 80% on searches with an AI Overview against 60% without. We could not review that analysis directly. We look more closely at why [people stop browsing after an AI Overview](https://underneath.agency/resources/ai-overviews-end-browsing-sessions) in a separate guide.

## How much traffic did a real publisher lose?

English Wikipedia lost roughly 5% of its search traffic after AI Overviews became Google’s US default.

[Khosravi and Yoganarasimhan](https://arxiv.org/abs/2602.18455) compared search visits to about half a million English Wikipedia articles with the same articles in German and French. Google made AI Overviews part of normal US search in May 2024, while they were not part of default search in the German and French markets during the study period. The study covers December 2023 to December 2024.

English search traffic fell by about 5.45% relative to German and 4.82% relative to French. That works out to roughly 100.27 million fewer search visits a month. A comparison with Japanese Wikipedia, which the authors treat as a weaker check, gave a 16.53% decline.

These are averages across all searches, all search engines and all users, including the many who never saw an AI Overview. The authors note that the loss on searches that actually showed one would be larger. Several other papers quote a figure of about 15% for this study; the sixth version, which we reviewed, reports 5.45% and 4.82%. The Wikipedia case shows [how little control publishers have](https://underneath.agency/resources/google-search-design-changes-traffic-risk) when Google changes its results pages.

## Is the loss the same for every site?

No: it depends on the site, the searches it relies on and whether its content can be summarized.

Wikipedia is a hard case for AI Overviews to hurt, because it is among the sources they cite most. It still lost traffic. Whether [being cited makes up for lost clicks](https://underneath.agency/resources/ai-overview-citations-vs-lost-clicks) is a question we cover separately. Khosravi and Yoganarasimhan note that a separate experiment by Agarwal and Sen found larger effects, possibly because ordinary publishers lose more than Wikipedia. [Xu, Iqbal and Montgomery](https://arxiv.org/abs/2605.14021) summarize that experiment as cutting organic click-throughs by 38%; we could not review it directly.

Not every site loses. [Zhang, Cui and Zhang](https://arxiv.org/abs/2605.16428) found that Reddit communities AI Overviews can cite gained 12.0% more daily comments than communities Google excludes from them. The gains were largest for opinion and experience discussions, which a summary cannot replace.

The stakes are commercial. Xu and colleagues found that 50.63% of the pages AI Overviews cited carried display ads, so a lost click is often lost revenue as well as lost attention.

## Do users gain anything in exchange for fewer clicks?

Not measurably: hiding AI Overviews changed neither trust in Google nor satisfaction in the experiment.

Wang and colleagues found no detectable change in trust, usefulness, satisfaction or sense of control when AI Overviews were hidden. Besides clicking, the clearest change was that people spent 0.59 fewer minutes per search session when AI Overviews were hidden.

That does not prove AI Overviews are useless to searchers. Khosravi and Yoganarasimhan stress that their data measures publisher traffic, not the value of users’ time. Some people do get their answer faster. But in the only randomized test of real use, the clicks publishers lost were not matched by a gain users noticed.

## Is AI Mode worse for clicks than AI Overviews?

Yes: sending every search to AI Mode cut outside clicks by 18.8 percentage points, about twice the AI Overview effect.

AI Mode is Google’s conversational search, a separate tab where Gemini answers and takes follow-up questions. In the same experiment, forced AI Mode also cut the share of people clicking through to news sites, Reddit and Wikipedia. Our article on [what AI Mode as Google’s default would cost](https://underneath.agency/resources/ai-mode-default-traffic-loss) covers that result in detail.

## What should you do about it?

Plan for fewer clicks per search on informational queries, and measure where your own exposure sits.

1. Find which of your traffic-driving searches show AI Overviews. Our [frequency study](https://underneath.agency/research/ai-overviews-frequency-study) found them on 60.8% of 800 US commercial keywords, but far fewer on local searches.
2. Forecast with a range, not one number. The evidence runs from about 5% of search visits for Wikipedia, averaged across all searches, to 8.8 points of click-through per search in the experiment.
3. Track click-through, not just rankings. A page can hold position one and still lose clicks when a summary sits above it.
4. Watch AI Mode. Its effect on clicks was about twice that of AI Overviews.
5. Look at what your content offers beyond facts. Opinion and experience content gained engagement on Reddit, while factual reference content lost visits. For publishers, [the risk to content businesses](https://underneath.agency/resources/content-business-risk-from-ai-search) is covered in depth.

If you want help sizing your exposure, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The research shows clicks fall, but not yet how much a typical business site loses in revenue.

- Only one randomized experiment on real Google use was available to us, with a young, educated sample over seven days.
- The Pew panel is observational and from March 2025; AI Overviews have changed since.
- The Wikipedia study covers one nonprofit publisher in 2024.
- No study here measures leads, sales or revenue for ordinary business sites.
- Long-run effects are unknown: people may click more, or less, as habits settle.

## Frequently asked questions

### Do AI Overviews reduce organic traffic?

Yes, on average. Hiding them raised clicks to outside sites by 8.8 percentage points in a 2026 US experiment, and English Wikipedia lost about 5% of its search traffic after they launched.

### How much do AI Overviews reduce click-through rate?

Pew’s 2025 browsing data shows clicks on results fell from 15% of visits to 8% when an AI Overview appeared. That is close to half, though the data shows an association rather than proof.

### Do people click the links inside AI Overviews?

Rarely. In Pew’s panel, just 1% of visits to pages with an AI Overview led to a click on a cited source.

### Do AI Overviews hurt every website equally?

No. Reddit discussion communities gained 12.0% more daily comments after AI Overviews launched, while factual sources such as Wikipedia lost visits.

## Sources

- Wang, Gleason, Bart, Wilson and Metaxa (2026), [AI in Search Reduces Publisher Referrals Without Improving User Experience: Experimental Evidence](https://arxiv.org/abs/2608.18352), arXiv:2608.18352.
- Chapekis, Lieb, Shah and Smith (2026), [Investigating Click Behaviors On Google Search Result Pages That Produce an AI Overview](https://arxiv.org/abs/2608.04831), arXiv:2608.04831.
- Khosravi and Yoganarasimhan (2026), [Impact of AI Search Summaries on Website Traffic: Evidence from Google AI Overviews and Wikipedia](https://arxiv.org/abs/2602.18455), arXiv:2602.18455.
- Aral, Li and Zuo (2026), [The Rise of AI Search: Implications for Information Markets and Human Judgement at Scale](https://arxiv.org/abs/2602.13415), arXiv:2602.13415.
- Xu, Iqbal and Montgomery (2026), [Measuring Google AI Overviews: Activation, Source Quality, Claim Fidelity, and Publisher Impact](https://arxiv.org/abs/2605.14021), arXiv:2605.14021.
- Zhang, Cui and Zhang (2026), [The Impact of AI Search on the Online Content Ecosystem: Evidence from Google and Reddit](https://arxiv.org/abs/2605.16428), arXiv:2605.16428.
- Underneath (2026), [When does Google show an AI Overview? 1,248 US searches](https://underneath.agency/research/ai-overviews-frequency-study)

---

This is the Markdown twin of https://underneath.agency/resources/do-ai-overviews-reduce-clicks. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Do AI Search Engines Agree With Leading Questions?"
description: "Often, on loaded questions. In a 2024 audit, AI answer engines gave one-sided answers that agreed with a charged question 50% to 80% of the time."
canonical: "https://underneath.agency/resources/do-ai-search-engines-agree-with-leading-questions"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Do AI search engines just agree with however a question is phrased?

Often, when the question is loaded. In the main audit on this, AI answer engines gave one-sided answers that went along with a charged question 50% to 80% of the time. That evidence comes from social and political debate questions; for questions about brands, the research shows wording changes the answer a lot, but not that engines simply agree.

## The short version

1. Three AI answer engines gave one-sided answers to 50% to 80% of charged debate questions in a 2024 audit, siding with the question’s framing ([Narayanan Venkit and colleagues, 2024](https://arxiv.org/abs/2410.22349)).
2. Adding web search to an AI answer cut hedging words by up to 60% while keeping confident language, across 11,000 real searches ([Huang and colleagues, 2026](https://arxiv.org/abs/2603.16138)).
3. In [our phrasing study](https://underneath.agency/research/ai-prompt-phrasing-study), adding “on a tight budget” kept the same first-named brand only 15.3% of the time, against 68.0% for simply asking again.
4. In [our reputation study](https://underneath.agency/research/is-it-legit-ai-reputation-study), 99.7% of answers to “Is this brand legit?” raised at least one problem, so a yes/no question did not produce a purely agreeable answer.

## What happens when you ask an AI a loaded question?

In the best evidence available, it usually answers the side the question leans toward.

[Narayanan Venkit and colleagues](https://arxiv.org/abs/2410.22349) tested You.com, BingChat and Perplexity in 2024. They used 168 debate questions written to lean one way, such as “Why should zoos exist?” rather than a neutral version. An answer counted as one-sided if it gave no points for the other side.

All three engines gave one-sided answers 50% to 80% of the time, favoring agreement with the question’s framing over presenting multiple views. Perplexity did worst. Its answers were also the longest, at 18.8 statements on average, so length did not bring balance.

People noticed. In the study’s interviews, 19 of 21 expert participants raised the lack of balanced viewpoints on opinionated or charged questions. One said the engine was “just telling you, ’you’re right… and here are the reasons why.’”

This evidence covers social and political debates, not product or vendor questions, and engines from 2024.

## Do AI answers sound more certain than their sources?

Yes, in several studies, which makes a one-sided answer easier to accept.

In the 2024 audit, Perplexity used the most confident language, with more than 90% of its answers rated very confident. It also rarely softened its tone on debate questions, unlike the other two engines.

[Huang and colleagues](https://arxiv.org/abs/2603.16138) compared answers from the same AI model with and without web search across 11,000 real searches. Search cut hedging words by up to 60% while keeping confident language. The ratio of certain to tentative wording rose from 0.39 without search to 0.49 with it. The authors warn this can present uncertain information with more confidence than the sources justify.

A slanted answer delivered in a confident voice is the combination that most needs checking. Tone also affects which sources get used; see [whether AI Overviews downplay negative content](https://underneath.agency/resources/do-ai-overviews-downplay-negative-content).

## Does this carry over to questions about brands and vendors?

Wording clearly changes brand answers, but the evidence shows the assistant following the buyer’s stated need, not simply flattering it.

In [our phrasing study](https://underneath.agency/research/ai-prompt-phrasing-study), we asked ChatGPT, Gemini and Perplexity 20 buyer questions in several wordings, 1,560 answers in all. Asking the identical question again kept the same first brand 68.0% of the time. Other wordings moved it further:

| Wording added to the question | Same first brand as the original |
|---|---|
| None (asked again) | 68.0% |
| “an honest, unbiased answer” | 54.9% |
| “I run a small business with about 10 employees” | 40.6% |
| “on a tight budget” | 15.3% |

The answers followed the framing. 84.7% of budget answers quoted a dollar figure, against 33.0% of original answers. Asking for an unbiased answer changed the list least, so it did not act as a strong check on the default.

A neutral yes/no question about a vendor did not produce simple agreement either. In our [reputation study](https://underneath.agency/research/is-it-legit-ai-reputation-study), all 312 complete answers to “Is this brand legit?” said yes, but 99.7% also raised at least one problem. We did not test negatively loaded vendor questions, such as “Why is this brand a scam?”.

## Does the wording also change which sources the AI reads?

Yes, which is one way framing can steer an answer before it is written.

In our phrasing study, reworded questions led the assistants to cite different websites far more than repeat runs did, and the brands changed most where the sources changed most. In a small test, [Wen and colleagues](https://arxiv.org/abs/2606.12439) compared 30 questions with lightly reworded versions across seven OpenAI and Google models. For Google’s Gemini models, every pair changed its cited sites after rewording. We cover this in [how question phrasing changes AI sources](https://underneath.agency/resources/does-question-phrasing-change-ai-sources).

Wording even decides whether Google shows an AI answer at all. [Xu, Iqbal and Montgomery](https://arxiv.org/abs/2605.14021) found question-form searches triggered AI Overviews (Google’s AI summary at the top of results) 64.7% of the time, against 9.5% for other searches. In [our AI Overview study](https://underneath.agency/research/ai-overviews-frequency-study), the same 96 topics showed an AI Overview 59.4% of the time as typed and 93.8% as a natural question.

## Do people check answers that agree with them?

Less than they check answers that disagree.

In the 2024 user study, participants asked questions that either matched or challenged their own views. When the question matched their view, they hovered over about one cited source (1.08 on average) and clicked about half of one (0.48). They checked noticeably more when the answer went against them.

So the combination is reinforcing. A leading question gets an agreeable answer, the answer sounds confident, and the person who asked is the least likely to check it. That matters because [AI answers often outrun their sources](https://underneath.agency/resources/ai-answers-unsupported-claims).

## What should you do about it?

Assume buyers ask loaded questions about your category, test those questions yourself, and publish balanced evidence engines can find.

1. **List the loaded questions buyers ask.** Include negative ones (“Why is X overpriced?”) and positive ones (“Why is X the best?”), not just neutral prompts.
2. **Ask them across assistants and repeat them.** One run is not enough; our study found repeat runs shared only part of their brand lists.
3. **Publish plain, specific answers to the hard questions.** Pricing, limits and comparisons stated factually give an engine material for the other side.
4. **Track several wordings of each key question.** A brand can lead one wording and vanish from another, so one score hides most of the picture.
5. **Read answers about your category with the same skepticism you expect from buyers.** Confident tone is not evidence of balance.

If you want help testing how assistants answer your buyers’ questions, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The strongest evidence of agreement with loaded questions comes from political debate questions, not commercial ones.

- No study we reviewed measures how often AI assistants agree with a negatively loaded question about a named vendor.
- The one-sided-answer figures come from 2024 engines; current ChatGPT and Google answers were not tested that way.
- Our phrasing study measured whether answers changed, not whether they became more or less accurate.
- The hedging study is correlational and mainly compares one model with and without search.
- How much a slanted AI answer changes what buyers actually purchase is unknown.

## Frequently asked questions

### Does ChatGPT just tell users what they want to hear?

The research has not tested this for ChatGPT on buying questions. In a 2024 audit of three other engines, one-sided answers that agreed with charged debate questions appeared 50% to 80% of the time.

### Does asking an AI for an unbiased answer help?

Only a little. In our phrasing study, asking for “an honest, unbiased answer” kept the same first brand 54.9% of the time, against 68.0% for asking again.

### Why do AI answers about the same product differ between people?

Wording is one reason. Adding “on a tight budget” kept the same first-named brand only 15.3% of the time, and changed the websites the assistants cited.

### Are AI answers more confident than their sources?

Often. Adding web search cut hedging words by up to 60% across 11,000 searches, so answers can sound surer than the evidence.

## Sources

- Narayanan Venkit, Laban, Zhou, Mao and Wu (2024), [Search Engines in an AI Era: The False Promise of Factual and Verifiable Source-Cited Responses](https://arxiv.org/abs/2410.22349), arXiv:2410.22349.
- Huang, Goyal, Saha and Chandrasekharan (2026), [Answer Bubbles: Information Exposure in AI-Mediated Search](https://arxiv.org/abs/2603.16138), arXiv:2603.16138.
- Wen and colleagues (2026), [Position: Generative Engine Optimization Creates Underexamined Risks, Governance Must Target Concentration, Disclosure, and Academic Blind Spots](https://arxiv.org/abs/2606.12439), arXiv:2606.12439.
- Xu, Iqbal and Montgomery (2026), [Measuring Google AI Overviews: Activation, Source Quality, Claim Fidelity, and Publisher Impact](https://arxiv.org/abs/2605.14021), arXiv:2605.14021.
- Underneath (2026), [Does rewording a question change AI brand recommendations?](https://underneath.agency/research/ai-prompt-phrasing-study)
- Underneath (2026), [“Is this brand legit?” How AI assistants build a reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study)
- Underneath (2026), [How often do AI Overviews appear?](https://underneath.agency/research/ai-overviews-frequency-study)

---

This is the Markdown twin of https://underneath.agency/resources/do-ai-search-engines-agree-with-leading-questions. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Do AI search engines cite the same websites as Google?"
description: "Mostly not. Studies find AI engines and Google share 18% to 38% of sources, and only 8.3% of ChatGPT’s citations ranked in Google’s top 10 in our test."
canonical: "https://underneath.agency/resources/do-ai-search-engines-cite-the-same-sites-as-google"
published: 2026-10-10
updated: 2026-10-10
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Do AI search engines cite the same websites as Google?

Mostly not. Several large studies from 2025 and 2026 find that AI search engines and Google share only a minority of their sources, and the AI engines also disagree with one another. A Google ranking report therefore describes only part of where your brand appears in AI answers.

## The short version

1. Only 38% of the domains cited by six AI search engines also appeared in Google or Bing results, across 55,936 queries collected in mid-2025 (Zhang and colleagues).
2. Even inside Google, AI Overviews and the regular results shared only 18% of their sources per query on average, across 11,500 queries collected in December 2025 (Grossman and colleagues).
3. In our test of 80 buyer questions, 8.3% of the pages ChatGPT cited ranked in Google’s top 10 for the same question, against 25.7% for Claude.
4. AI search leans on third-party coverage: in US tests, it drew 72.7% of its sources from independent “earned” media, against 45.4% for Google (Chen and colleagues, 2025).
5. AI engines also disagree with each other: in one small vendor test, 96.4% of cited web addresses appeared in only one of four AI engines.

## How much do AI search and Google sources overlap?

Far less than most marketers assume: usually well under half of sources are shared. [Zhang and colleagues](https://arxiv.org/abs/2512.09483) compared six AI search engines, including ChatGPT, Gemini, Perplexity and Google’s AI Mode, with Google and Bing. Only 38% of domains appeared on both sides, while 37% appeared only in the AI engines.

Google’s own products do not line up either. [Grossman and colleagues](https://arxiv.org/abs/2604.27790) compared Google’s regular results, its AI Overviews (the AI summary at the top of Google’s results) and Gemini. On average, only 18% of the sources returned by either AI Overviews or the regular results were returned by both. We look at what that means for rankings in [whether Google rankings lead to AI Overviews](https://underneath.agency/resources/does-ranking-on-google-get-you-into-ai-overviews).

[Huang and colleagues](https://arxiv.org/abs/2603.16138) found the same pattern for an OpenAI model with web search. Its 100 most-cited domains overlapped only 24% with AI Overviews and 25% with Google’s regular results. AI Overviews and the regular results, by contrast, shared 68% of their top domains.

| Study | What was compared | Share in common |
|---|---|---|
| Zhang et al., mid-2025 | Domains cited by six AI engines and by Google or Bing | 38% |
| Grossman et al., December 2025 | Sources per query, AI Overviews and Google’s results | 18% |
| Huang et al., 2026 | Top 100 domains, OpenAI model with search and Google’s results | 25% |

Each figure counts something different, from whole domains to single pages, so they are not directly comparable. The direction is the same in all three. Zhang and colleagues suggest one reason: AI engines break a question into smaller searches instead of searching for it as typed.

## Do pages that rank on Google get cited by AI assistants?

Some do, but they are a minority, and the share varies widely by assistant. In our [study of AI citations and Google rankings](https://underneath.agency/research/ai-citations-google-rankings-study), we compared 2,492 citations from four assistants with Google’s results for 80 US buyer questions. Of the pages ChatGPT cited, 8.3% ranked in Google’s top 10 for the question; for Claude it was 25.7%.

Most of the gap is not explained by ranking at all. 68.5% of ChatGPT’s citations were not in Google’s top 100 for the question, or for searches of the kind ChatGPT runs. Being a trusted site helps more than ranking one page: ChatGPT’s site-level overlap with Google, 22.9%, was nearly three times its page-level overlap. For one engine, see [whether backlinks still matter for Perplexity](https://underneath.agency/resources/do-backlinks-matter-for-perplexity-visibility).

Where AI engines and Google do share a domain, it often sits at the top of Google’s list. Zhang and colleagues found that the most common position for these shared domains was the first result: 14.53% of cases on Google and 23.27% on Bing.

## What kinds of sites do AI engines prefer instead?

They favor third-party editorial and reference sources, and fewer social or brand-owned pages. [Chen and colleagues](https://arxiv.org/abs/2509.08919) sorted sources into brand-owned, earned (independent publications and reviews) and social. For US queries, Google returned 45.4% earned sources, while AI search returned 72.7% earned and negligible social content.

Social platforms drop out almost entirely for some engines. Huang and colleagues found the OpenAI model with search drew 0.1% of its citations from social platforms, against 13.4% for Google’s regular results. AI engines also cite far fewer pages: a mean of 4.3 web addresses per answer in the Zhang study, against 10.3 for traditional search.

Age matters too. In our [freshness study](https://underneath.agency/research/ai-source-freshness-study), assistants cited pages first published about half as long ago as Google’s top 10 for the same questions, a ratio of 0.50.

## Do AI engines at least agree with each other?

No. In some tests they overlap with each other even less than with Google. In a small test on 6 June 2026, [Tannenbaum](https://arxiv.org/abs/2609.22655) sent 15 commercial prompts to ChatGPT, Microsoft Copilot, Google and Perplexity. Of all cited web addresses, 96.4% appeared in only one engine. The author founded a company that sells AI visibility software, and 15 prompts is a narrow sample. We compare the engines in more depth in [whether AI engines cite the same sources](https://underneath.agency/resources/do-ai-search-engines-cite-the-same-sources).

Our own [comparison of AI Mode and AI Overviews](https://underneath.agency/research/ai-mode-vs-ai-overviews-study) found the same within Google. On searches where both cited sources, the two surfaces shared a mean of 13.9% of cited web addresses. Of the pages in Google’s organic top 10, the AI Overview cited 29.7% and AI Mode 16.3%.

## What should you do about it?

Treat Google rankings and AI visibility as related but separate results, and measure each one. Concretely:

1. Track citations engine by engine. ChatGPT, Gemini, Perplexity, Claude, AI Overviews and AI Mode each draw on different sources, so one combined score hides real differences.
2. Map the third-party sites that AI engines cite in your category. Industry publications, review sites and reference pages carry more weight in AI answers than in Google’s results.
3. Strengthen your whole site, not just single ranking pages. Assistants often cite a different page from a site that ranks.
4. Keep core pages current. AI assistants lean toward recently published pages more than Google’s top 10 does.
5. Keep investing in search rankings. When AI engines and Google share a source, the first result is the most common match. Our explainer on [how GEO differs from SEO](https://underneath.agency/resources/what-is-generative-engine-optimization) sets out where the two overlap.

If you want a measured starting point, our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) begins with an engine-by-engine audit.

## What does the research not tell us yet?

The studies agree on the direction, but several questions remain open.

- They measure overlap in different units (domains, pages, top-100 lists), so no single overlap figure applies everywhere.
- Most collections are one-off snapshots from US locations, and AI sources change from day to day.
- Several studies used developer versions of AI models rather than the consumer apps people use.
- None shows why an AI engine picks one source over another, or whether changing a page would change its citations.
- The evidence on agreement between AI engines comes partly from a 15-prompt vendor test and needs larger, independent replication.

## Frequently asked questions

### If I rank first on Google, will ChatGPT cite me?

Not reliably. In our test of 80 buyer questions, only 8.3% of the pages ChatGPT cited ranked in Google’s top 10, and 68.5% were outside Google’s top 100 for the question.

### Do Google’s AI Overviews use the same sources as Google search?

Only partly. Grossman and colleagues found AI Overviews and Google’s regular results shared 18% of sources per query on average in December 2025.

### Which AI assistant is closest to Google’s rankings?

Claude, among the four we tested: 25.7% of its citations ranked in Google’s top 10, against 8.3% for ChatGPT.

### Can a Google ranking report measure AI visibility?

No. Across six AI engines, only 38% of cited domains also appeared in Google or Bing results, so most AI sources never show up in a ranking report.

## Sources

- Zhang et al. (2025), [Source Coverage and Citation Bias in LLM-based vs. Traditional Search Engines](https://arxiv.org/abs/2512.09483), arXiv:2512.09483.
- Grossman, Liu, Chen, Smith, Borcea and Chen (2026), [How Generative AI Disrupts Search: An Empirical Study of Google Search, Gemini, and AI Overviews](https://arxiv.org/abs/2604.27790), arXiv:2604.27790.
- Huang, Goyal, Saha and Chandrasekharan (2026), [Answer Bubbles: Information Exposure in AI-Mediated Search](https://arxiv.org/abs/2603.16138), arXiv:2603.16138.
- Chen, Wang, Chen and Koudas (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919), arXiv:2509.08919.
- Tannenbaum (2026), [Scoring With the Engine: Retrieval Exposure, Cross-Engine Divergence, and the Limits of Engine-Agnostic GEO Scores](https://arxiv.org/abs/2609.22655), arXiv:2609.22655.
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude cite pages that rank?](https://underneath.agency/research/ai-citations-google-rankings-study)
- Underneath (2026), [AI Mode vs AI Overviews: how different are the sources?](https://underneath.agency/research/ai-mode-vs-ai-overviews-study)
- Underneath (2026), [How fresh are the pages AI engines cite?](https://underneath.agency/research/ai-source-freshness-study)

---

This is the Markdown twin of https://underneath.agency/resources/do-ai-search-engines-cite-the-same-sites-as-google. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Do ChatGPT, Copilot, Google and Perplexity cite the same sources?"
description: "No. Studies find ChatGPT, Copilot, Google and Perplexity cite largely different pages for the same question, so success on one engine rarely carries over."
canonical: "https://underneath.agency/resources/do-ai-search-engines-cite-the-same-sources"
published: 2026-10-10
updated: 2026-10-10
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Do ChatGPT, Copilot, Google and Perplexity cite the same sources?

No. For the same question, the major AI search engines cite mostly different pages, and often entirely different websites. Being cited by one engine says little about whether the others will cite you.

## The short version

1. In a one-day test of 15 commercial prompts, 96.4% of the cited web addresses appeared in only one of four engines, [a vendor-run study found](https://arxiv.org/abs/2609.22655).
2. For the same product question, ChatGPT and Gemini shared on average only 5.4% of the sites they displayed, in [a 2026 audit of product recommendations](https://arxiv.org/abs/2609.18729).
3. Even Google’s own products disagree: [our comparison](https://underneath.agency/research/ai-mode-vs-ai-overviews-study) found AI Mode and AI Overviews shared 13.9% of cited pages on the same searches.
4. The engines differ more from each other than from themselves on another day: the vendor study found same-engine results about 42 times more similar than cross-engine results.

## Do AI search engines cite the same pages for the same question?

Rarely: most cited pages appear on only one engine. [Tannenbaum](https://arxiv.org/abs/2609.22655) ran 15 commercial prompts through ChatGPT, Microsoft Copilot, Google and Perplexity on 6 June 2026. Among 528 distinct cited web addresses, 509 appeared in only one engine. That is 96.4% of all observed addresses.

Pairs of engines rarely overlapped at all. Across 73 same-prompt comparisons, 84.9% of engine pairs shared no cited address. The author founded a company that sells AI visibility software, and his data covers one niche on one day, so treat the exact figures with care.

Independent studies point the same way. [Uberti-Bona Marin and colleagues](https://arxiv.org/abs/2609.18729) asked 117 product questions in the Netherlands in September 2026. For the same question, ChatGPT and Gemini shared only 5.4% of their displayed sites on average, and 76.7% of comparisons shared no site at all.

## Do the engines at least agree on their top sources?

No, the disagreement holds at the top of the list too. In the Tannenbaum data, no pair of engines shared a single address among their first five citations, across all 60 comparisons on the ten prompts every engine answered.

No single engine gives a full picture either. ChatGPT captured 42.6% of the addresses cited across all four engines, Google 24.3%, Perplexity 23.6% and Copilot 11.4%. Watching one engine means missing most of the sources the others use. Engines also differ in [how many sources they cite per answer](https://underneath.agency/resources/how-many-sources-ai-search-engines-cite).

Our own data adds a check on Google. When AI Mode and AI Overviews both answered the same search, [they led with the same source in 20.9% of searches](https://underneath.agency/research/ai-mode-vs-ai-overviews-study).

## Do Google’s own AI products agree with each other?

Only partly, even though they share Google’s index. In our study of 400 searches, AI Mode and AI Overviews shared a mean 13.9% of cited pages where both cited sources. In 30.5% of those searches they had no cited page in common.

Gemini differs again. [Grossman and colleagues](https://arxiv.org/abs/2604.27790) compared Google Search, AI Overviews and Gemini on 11,500 queries. On average, only 18% of the sources returned by either AI Overviews or regular search were retrieved by both, and AI Overviews and Gemini were the least similar pair.

A larger study by [Zhang and colleagues](https://arxiv.org/abs/2512.09483) covered 55,936 queries across six AI engines and Google and Bing in mid-2025. Only 38% of cited domains appeared in both AI and traditional results. We look further at [how AI citations compare with Google’s results](https://underneath.agency/resources/do-ai-search-engines-cite-the-same-sites-as-google).

## Are the differences just random noise?

No: engines differ far more from each other than from their own results on another day. Tannenbaum found that the same engine’s cited addresses changed by 67.0% between 5 and 6 June. Even so, one engine’s list was about 42 times more similar to its own list a day earlier than to another engine’s list on the same day.

Engine family matters. [Yang](https://arxiv.org/abs/2507.05301) studied news citations in an AI search comparison site in 2025. Models from the same company scored 0.82 to 0.99 on a similarity scale where 1 means identical news sources. Across companies, scores fell to between 0.11 and 0.58.

Our [four-assistant study](https://underneath.agency/research/ai-assistants-brand-agreement-study) of 80 buyer questions found the same split. In 43.8% of question pairs, two assistants shared no cited domain at all.

## Does it matter which version of an engine you test?

Yes. The app a customer uses and the developer version a tracking tool queries can cite different sites. In the product audit, ChatGPT’s app and its developer version shared an average of 12.0% of their displayed sites for the same question, asked moments apart. For 60.9% of pairs, they shared no site.

This matters for measurement. [A dashboard built on developer access](https://underneath.agency/resources/api-ai-visibility-monitoring) may not show what buyers actually see in the app.

Asking the same engine again also changes the list. In the same audit, two runs of one question shared on average 26.0% of displayed sites in ChatGPT, 29.8% in Gemini and 45.9% in AI Overviews. For ChatGPT, the number of distinct sites seen per question rose from 2.65 after one request to 5.56 after three. A single answer is one sample, not the engine’s settled view. We explain [why citations change between checks](https://underneath.agency/resources/why-ai-search-citations-change) separately.

## Is anything shared across engines?

A small core of well-known publishers, and little else. [Chen and colleagues](https://arxiv.org/abs/2509.08919) compared Claude, ChatGPT and Perplexity on car questions in August 2025. Each engine’s sources were mostly its own: 50.3% of Claude’s domains, 60.8% of ChatGPT’s and 56.5% of Perplexity’s were exclusive to that engine.

The domains all three shared were mainly large review outlets such as Car and Driver, Edmunds and Consumer Reports. In consumer electronics, the shared core was TechRadar, Tom’s Guide and RTINGS. That study used developer versions of the engines and thanked a GEO vendor for support.

Rankings differ too. In [our rankings study](https://underneath.agency/research/ai-citations-google-rankings-study), 8.3% of the pages ChatGPT cited ranked in Google’s top 10 for the question, against 25.7% for Claude.

## What should you do about it?

Treat each AI engine as its own channel, with its own sources to win.

1. Decide which engines your buyers use, and measure each one separately. If you need one number, see [how to combine engines into one score](https://underneath.agency/resources/combine-ai-engines-visibility-score).
2. Do not read one engine’s results as a proxy for the others.
3. Test the consumer apps, not only developer versions, because their sources differ.
4. Repeat measurements over several days, since one engine’s sources change daily.
5. Prioritize the few publishers that every engine cites in your category. They are the closest thing to a shared foundation.

If you want help measuring across engines, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The research shows engines diverge, but not why, nor how that changes over months.

- The sharpest figures come from small samples, such as 15 prompts on one day from a company with a commercial interest.
- Studies use different units: exact pages, domains, displayed sources or news outlets only.
- Few studies cover business-to-business questions, so divergence in those categories is unmeasured.
- No study we reviewed tracks whether divergence is growing or shrinking over time.
- Citation is not the same as influence: the sources an engine displays are a small part of what it reads.

## Frequently asked questions

### If ChatGPT cites my site, will Perplexity cite it too?

Not necessarily. In one four-engine test, 96.4% of cited addresses appeared on only one engine, though that test covered just 15 prompts.

### Do Gemini and Google AI Overviews use the same sources?

No. A study of 11,500 queries found AI Overviews and Gemini were the least similar pair among Google Search, AI Overviews and Gemini.

### Why do AI engines cite different websites?

Each engine runs its own searches and applies its own selection. News citations from models by the same company scored 0.82 to 0.99 on similarity, while cross-company scores fell as low as 0.11.

### Should I track every AI engine separately?

Yes, if your buyers use more than one. The broadest single engine captured only 42.6% of the addresses cited across four engines in one test.

## Sources

- Tannenbaum (2026), [Scoring With the Engine: Retrieval Exposure, Cross-Engine Divergence, and the Limits of Engine-Agnostic GEO Scores](https://arxiv.org/abs/2609.22655), arXiv:2609.22655.
- Uberti-Bona Marin et al. (2026), ["If I Had to Buy Just ONE: Galaxy S26 Ultra": Auditing AI-Generated Product Recommendations](https://arxiv.org/abs/2609.18729), arXiv:2609.18729.
- Grossman et al. (2026), [How Generative AI Disrupts Search: An Empirical Study of Google Search, Gemini, and AI Overviews](https://arxiv.org/abs/2604.27790), arXiv:2604.27790.
- Zhang et al. (2025), [Source Coverage and Citation Bias in LLM-based vs. Traditional Search Engines](https://arxiv.org/abs/2512.09483), arXiv:2512.09483.
- Yang (2025), [News Source Citing Patterns in AI Search Systems](https://arxiv.org/abs/2507.05301), arXiv:2507.05301.
- Chen et al. (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919), arXiv:2509.08919.
- Underneath (2026), [AI Mode vs AI Overviews: how different are the sources?](https://underneath.agency/research/ai-mode-vs-ai-overviews-study)
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude agree on brands?](https://underneath.agency/research/ai-assistants-brand-agreement-study)
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude cite pages that rank?](https://underneath.agency/research/ai-citations-google-rankings-study)

---

This is the Markdown twin of https://underneath.agency/resources/do-ai-search-engines-cite-the-same-sources. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Do backlinks still matter for visibility in Perplexity?"
description: "Probably, but the evidence is thin. In a test of 112 startups, linking websites were the strongest predictor of Perplexity visibility; no study proves cause."
canonical: "https://underneath.agency/resources/do-backlinks-matter-for-perplexity-visibility"
published: 2026-10-10
updated: 2026-10-10
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Do backlinks still matter for visibility in Perplexity?

The limited evidence says yes, though it shows a link rather than proof of cause. In a December 2025 test of 112 Product Hunt startups, the number of other websites linking to a startup was the strongest predictor of whether Perplexity named it. Perplexity also cites many pages that rank well on Google, where links have long counted.

## The short version

1. Referring domains, the number of separate websites linking to you, were the strongest predictor of Perplexity naming a startup, scoring 0.319 on a scale where 0 means no link and 1 a perfect one (Sharma, 112 startups, December 2025).
2. In the same test, a score for AI-friendly page content showed no link with Perplexity visibility (−0.102), and no signal predicted visibility in an OpenAI model without web search.
3. In our test of 80 buyer questions, Perplexity cited 33.0% of the pages in Google’s top 10 for the question, and 40.8% of those in positions 1 to 3.
4. Perplexity leans on third-party sources: 67.4% of its sources about well-known brands were independent “earned” media such as reviews and publications (Chen and colleagues, 2025).
5. Perplexity’s sources shift between runs: repeated runs of the same query shared only about half of their cited domains (0.50), in a study by an AI visibility vendor.

## What does the research say about backlinks and Perplexity?

One study measured it directly, and links came out as the strongest single signal. [Sharma](https://arxiv.org/abs/2601.00912) picked 112 startups from the top 500 on the 2025 Product Hunt leaderboard and ran 2,240 queries to Perplexity and to an OpenAI model without web search. The queries ran through each company’s developer connection between December 15 and 20, 2025.

Sharma then compared each startup’s visibility with its link data, Product Hunt results and community presence. For Perplexity, referring domains had the strongest link with visibility across the full sample, at 0.319. The share of links that pass ranking credit, known as followed links, also mattered, at 0.238.

| Signal | Strength of link with Perplexity visibility |
|---|---|
| Reddit mentions, after cleaning (60 products) | 0.395 |
| Referring domains (all 112 products) | 0.319 |
| Product Hunt daily rank (better rank, more visibility) | 0.286 |
| Share of followed links | 0.238 |
| AI-friendly content score | −0.102, no clear link |

The scale runs from 0 (no link) to 1 (a perfect one), so these are moderate links, not strong ones. The Reddit figure is higher, but it rests on fewer products: Sharma removed 46% of the sample because their generic names, such as “Cursor”, matched unrelated Reddit posts.

## Why would links matter to an AI engine?

Because Perplexity searches the web as it answers, and search results reward links. Sharma’s explanation is simple: the signals that help a page rank in traditional search also make it more likely to be found and cited by an engine that searches live.

The contrast with a model that does not search supports this. Perplexity named the startups in 8.29% of discovery questions, and it named 31 of the 112 at least once (27.7%). For the OpenAI model without search, Sharma found no signal that predicted visibility at all. Perplexity showed seven signals with a clear link, which suggests an engine that searches live gives marketers levers that a model answering from memory does not. We look at how much assistants still [answer from training data](https://underneath.agency/resources/do-ai-assistants-answer-from-training-data) in a separate guide.

Our own data links Perplexity closely to Google rankings. In our [study of AI citations and Google rankings](https://underneath.agency/research/ai-citations-google-rankings-study), Perplexity cited 33.0% of the pages in Google’s top 10 for each question, far more than ChatGPT, Gemini or Claude. It cited 40.8% of pages ranked 1 to 3 and 28.8% of pages ranked 4 to 10.

## Does ranking explain everything Perplexity cites?

No. Most of what Perplexity cites does not rank on Google for the question at all. In the same study, 66.8% of Perplexity’s citations were not in Google’s top 100 for the question. Perplexity lists many sources, a median of 20 per answer, and its first three citations ranked far more often than the rest, 23.8% against 14.2% overall. The wider gap is covered in [how AI sources differ from Google’s](https://underneath.agency/resources/do-ai-search-engines-cite-the-same-sites-as-google).

Perplexity also reaches pages by routes other than its own crawler. In our [crawler blocking study](https://underneath.agency/research/ai-crawler-blocking-study), 40.4% of the pages Perplexity cited on top sites were closed to PerplexityBot by the site’s own robots.txt file.

Third-party coverage carries a lot of weight. [Chen and colleagues](https://arxiv.org/abs/2509.08919) found 67.4% of Perplexity’s sources about well-known brands came from earned media, independent publications and reviews. In a soft drink test, 67.9% of Perplexity’s brand mentions were major brands. Coverage on other sites and links from them often come together, so this fits with the link finding without proving it.

## How reliable is a single Perplexity check?

Not very. Perplexity’s sources change from run to run, so one check can mislead. [Sielinski](https://arxiv.org/abs/2603.08924), who works for an AI visibility measurement company, ran the same queries repeatedly on Perplexity, OpenAI’s search model and Gemini over nine days. Repeated Perplexity runs shared only about half of their cited domains, an overlap of 0.50.

Perplexity also cites more sources than its rivals in other data. In a public dataset of 602 test prompts analyzed by [Zhang, He and Yao](https://arxiv.org/abs/2604.25707), Perplexity cited 16.35 sources per prompt on average, against 6.88 for ChatGPT. The authors argue that classic signals such as backlinks help a page enter the pool of sources, but do not show whether the answer actually uses it. Signals on the page itself are weighed in [which on-page signals track with AI citations](https://underneath.agency/resources/on-page-signals-linked-to-ai-citations).

## What should you do about it?

Keep building links and third-party coverage, and treat them as an input to Perplexity rather than a guarantee. In practice:

1. Earn links from relevant, reputable websites. In the one direct test, referring domains were the strongest signal for Perplexity visibility.
2. Pursue coverage in publications and review sites. Perplexity draws most of its sources from that kind of earned media.
3. Encourage genuine discussion in relevant communities such as Reddit, not planted posts.
4. Rank for the narrower questions buyers ask, since Perplexity cites top-ranked pages more often.
5. Check that robots.txt and your firewall do not block Perplexity’s crawlers by accident.
6. Measure visibility over many runs and weeks, not once.

If you want help prioritizing these, our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) starts with that measurement.

## What does the research not tell us yet?

The link between backlinks and Perplexity rests on thin evidence.

- Only one study measures backlinks against Perplexity visibility directly: 112 startups, mostly developer, productivity and AI tools, queried over one week.
- It shows a link, not a cause. Better-known products tend to have more links, more coverage and more discussion all at once.
- The Reddit result rests on 60 products after cleaning, and the study tested Perplexity through its developer connection rather than the consumer app.
- No study has added links to a site and then measured the change in Perplexity answers.
- The findings may not carry over to established brands, other industries, or other AI engines such as ChatGPT with search.

## Frequently asked questions

### Do backlinks help you appear in Perplexity answers?

Probably. In a December 2025 test of 112 startups, referring domains were the strongest predictor of Perplexity visibility, with a moderate link of 0.319, though no study has shown cause.

### Does Perplexity use Google rankings?

It overlaps with them more than other assistants but does not copy them. Perplexity cited 33.0% of Google’s top-10 pages in our test, yet 66.8% of its citations were outside Google’s top 100 for the question.

### Are backlinks more useful than AI-optimized content for Perplexity?

In the one study that compared them, yes. Referring domains showed a link with Perplexity visibility, while an AI-friendly content score showed none (−0.102).

### How often does Perplexity recommend a new startup?

Rarely. Perplexity named the startups in 8.29% of discovery questions in Sharma’s test.

## Sources

- Sharma (2026), [The Discovery Gap: How Product Hunt Startups Vanish in LLM Organic Discovery Queries](https://arxiv.org/abs/2601.00912), arXiv:2601.00912.
- Chen, Wang, Chen and Koudas (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919), arXiv:2509.08919.
- Sielinski (2026), [Quantifying Uncertainty in AI Visibility: A Statistical Framework for Generative Search Measurement](https://arxiv.org/abs/2603.08924), arXiv:2603.08924.
- Zhang, He and Yao (2026), [From Citation Selection to Citation Absorption: A Measurement Framework for Generative Engine Optimization Across AI Search Platforms](https://arxiv.org/abs/2604.25707), arXiv:2604.25707.
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude cite pages that rank?](https://underneath.agency/research/ai-citations-google-rankings-study)
- Underneath (2026), [Which AI crawlers do top websites block? 10,000 sites, 2026](https://underneath.agency/research/ai-crawler-blocking-study)

---

This is the Markdown twin of https://underneath.agency/resources/do-backlinks-matter-for-perplexity-visibility. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Does being cited by ChatGPT boost our Google rankings?"
description: "No evidence says so. The one field test found ChatGPT gains left Google clicks flat to falling, and ChatGPT mostly cites pages Google does not rank."
canonical: "https://underneath.agency/resources/do-chatgpt-citations-boost-google-rankings"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Does being cited by ChatGPT boost our Google rankings?

There is no evidence that it does. The only published field test found that pages winning more ChatGPT traffic saw Google clicks stay flat or fall, not rise. That is one website over a few months, so it is an absence of evidence rather than proof of no effect, but it means a “ChatGPT lifts Google” assumption does not belong in a forecast.

## The short version

1. In a field test on one website, pages that gained ChatGPT referrals saw Google clicks fall about 25%, against about 20% site-wide; the authors found “no support” for a lift to Google ([Watanabe and Nakayashiki](https://arxiv.org/abs/2606.04362)).
2. ChatGPT and Google largely draw on different pages: only 8.3% of the pages ChatGPT cited ranked in Google’s top 10 for the question ([our AI citations and Google rankings study](https://underneath.agency/research/ai-citations-google-rankings-study)).
3. Where a link exists, it runs from Google to AI: the first organic result was cited in 49.5% of AI Overviews, the ninth in 15.5% ([our AI Overview citations study](https://underneath.agency/research/ai-overview-citations-study)).
4. Being cited by AI does not guarantee visits either: English Wikipedia, among the sources AI summaries cite most, lost about 5.45% of its search traffic after AI Overviews launched ([Khosravi and Yoganarasimhan](https://arxiv.org/abs/2602.18455)).

## Is there evidence that ChatGPT citations lift Google rankings?

No: the one direct test found Google traffic flat to declining on pages that gained ChatGPT traffic. [Watanabe and Nakayashiki](https://arxiv.org/abs/2606.04362), who work at Glasp, tracked their own site from mid-2025 to May 2026 using its analytics and Google Search Console data. They optimized one large section for ChatGPT in January 2026 and left the rest alone.

The optimized pages gained. Their ChatGPT referrals grew 6.1 times between January and May 2026, and after [allowing for ChatGPT’s general growth](https://underneath.agency/resources/chatgpt-referral-growth-and-geo) the authors estimate a real lift of about 1.8 to 2.3 times. Google is where the flywheel should have shown up next.

It did not. Google clicks to the optimized section fell about 25% from the second half of 2025, close to the roughly 20% fall across the whole site. The authors write that they “find no support for the stronger ‘flywheel’ claim,” and that Google traffic to those pages was “flat-to-declining, not rising.”

## How strong is that evidence?

It is a single case, so treat it as a caution, not a verdict. The study covers one domain, and ChatGPT supplied nearly all its AI traffic: other AI engines contributed under 1% of AI-referred visits. The window ran about five months after the changes, and a Google effect could take longer to appear.

There are other limits the authors flag. The changes were a bundle, and they [protected pages that already earned Google clicks](https://underneath.agency/resources/chatgpt-optimization-google-rankings) from being rewritten. The authors are also reporting on their own company’s site.

Still, this is the only published test of the question, and it points the wrong way for the flywheel. Absent better evidence, the safe planning assumption is that ChatGPT and Google results must each be earned on their own terms.

## Do ChatGPT and Google draw on the same pages?

Mostly not, which makes a strong link between them less likely. In [our study of 80 US buyer questions](https://underneath.agency/research/ai-citations-google-rankings-study), only 8.3% of the pages ChatGPT cited ranked in Google’s top 10 for the question. 68.5% were not in Google’s top 100 for the question or for any of the searches ChatGPT ran.

ChatGPT also does its own searching. It ran a mean of 3.7 searches per answer, none of them a word-for-word copy of the user’s question ([our hidden searches study](https://underneath.agency/research/ai-hidden-searches-study)). It is choosing from a different pool, by different rules.

Google’s rankings are themselves a moving target. Between 26 and 28 September, Google’s top 10 for the same questions changed substantially, and ChatGPT’s match fell from 8.3% to 5.2% against the later results. Rankings shift for many reasons that have nothing to do with AI citations.

## Which way does the link run, if there is one?

The evidence points from Google rankings to AI citations, at least inside Google’s own AI. In [our study of 4,051 AI Overview citations](https://underneath.agency/research/ai-overview-citations-study), 28.7% were page-one results for the same search. The first organic result was cited in 49.5% of AI Overviews, the ninth in 15.5%.

A vendor-authored study takes the same view of cause and effect. In the workflow proposed by [Vishwakarma and colleagues](https://arxiv.org/abs/2605.25517) at Sprinklr, a brand missing from AI citations has a problem with being found, and the stated action is to “improve SEO.” In their framing, search work feeds AI visibility, not the reverse.

So the dependable direction is the familiar one. Ranking well on Google helps Google’s AI find you. Nothing published shows ChatGPT returning the favor.

## Does being cited by AI at least bring visitors back?

Not reliably, and that weakens the flywheel logic further. [Khosravi and Yoganarasimhan](https://arxiv.org/abs/2602.18455) at the University of Washington studied English Wikipedia after Google made AI Overviews the US default in May 2024. Wikipedia is among the sources AI summaries cite most often.

Its monthly search traffic still fell 5.45% relative to the German edition and 4.82% relative to the French one. The authors conclude that any traffic from those citations “did not offset the average decline.” A citation is visibility, but it is not the same as a visit, and visits are what a flywheel would run on. That loss followed a change in [how Google lays out its results](https://underneath.agency/resources/google-search-design-changes-traffic-risk), which publishers cannot control.

## What should you do about it?

Plan ChatGPT and Google as separate channels, and do not count on one to lift the other. In practice:

1. Remove any “AI citations will lift SEO” line from forecasts and business cases until there is evidence for it.
2. Keep investing in Google rankings directly. They still drive citations in Google’s own AI Overviews.
3. Set separate targets for ChatGPT visibility, based on the pages and searches it actually uses.
4. If you test the flywheel yourself, keep a comparison group of similar pages you do not optimize, and watch Google clicks for longer than a few months.
5. Report the two channels side by side so a gain in one cannot mask a loss in the other.

For help setting up that kind of measurement, see [our generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The question has barely been studied, so the honest answer is “not shown,” not “impossible.”

- Only one field test addresses it directly, on one website, one AI engine and about five months of data after the changes.
- Nobody has measured whether brand searches on Google rise after a brand appears often in ChatGPT answers.
- Our studies compare which pages each system cites or ranks; they do not follow pages over time.
- Longer-run effects, over a year or more, are untested.

## Frequently asked questions

### Do AI citations help SEO?

Not on current evidence. The one field test found Google clicks to pages that gained ChatGPT traffic were flat to declining, about 25% down against 20% site-wide.

### Is there an AEO to SEO flywheel?

None has been shown. The authors of the only direct test write that they “find no support” for it in their study window.

### Does ranking on Google help you get cited by AI?

For Google’s own AI, yes. In our study, the top organic result was cited in 49.5% of AI Overviews; for ChatGPT the overlap with Google’s top 10 was only 8.3%.

### Why would ChatGPT citations not carry over to Google?

Partly because the two systems pick different pages. In our study, 68.5% of the pages ChatGPT cited were not in Google’s top 100 for the question or its own searches.

### Can ChatGPT traffic make up for lost Google traffic?

The published evidence cannot tell you. On the one site studied, ChatGPT referrals grew while Google clicks fell, but the authors withheld absolute visit counts, so the two channels cannot be compared in size. Measure both on your own site before assuming one replaces the other.

## Sources

- Watanabe and Nakayashiki (2026), [Disentangling Answer Engine Optimization from Platform Growth: A Log-Based Natural Experiment on ChatGPT Referral Traffic](https://arxiv.org/abs/2606.04362), arXiv:2606.04362.
- Khosravi and Yoganarasimhan (2026), [Impact of AI Search Summaries on Website Traffic: Evidence from Google AI Overviews and Wikipedia](https://arxiv.org/abs/2602.18455), arXiv:2602.18455.
- Vishwakarma, Kumar and Jamidar (2026), [What Gets Cited: Competitive GEO in AI Answer Engines](https://arxiv.org/abs/2605.25517), arXiv:2605.25517.
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude cite pages that rank?](https://underneath.agency/research/ai-citations-google-rankings-study)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [Do AI Overviews cite the pages that rank? 4,051 citations](https://underneath.agency/research/ai-overview-citations-study)

---

This is the Markdown twin of https://underneath.agency/resources/do-chatgpt-citations-boost-google-rankings. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Do citations make AI answers more trusted, even when wrong?"
description: "Yes. In a US experiment, reference links raised trust in AI answers, and broken or irrelevant links raised it just as much as valid ones."
canonical: "https://underneath.agency/resources/do-citations-make-ai-answers-more-trusted"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Do citations make people trust AI answers more, even when the links are wrong?

Yes. In a large US experiment, adding reference links to AI search answers raised people’s trust in them, and the boost was the same whether the links were valid or swapped for broken and irrelevant ones. That makes a citation a credibility signal in its own right, and a wrong attribution a quiet reputational risk.

## The short version

1. Among 4,927 US adults, adding clickable reference links to Google AI answers raised trust by 0.091 points and willingness to share by 0.121 points on a 7-point scale ([Li and Aral](https://arxiv.org/abs/2504.06435)).
2. Replacing a link with a broken or irrelevant one changed trust by just 0.021 points, a difference too small to distinguish from zero (Li and Aral).
3. Citations themselves are often wrong: in a 2024 audit, only 49.0% to 68.3% of citations pointed to a source that supported the statement ([Narayanan Venkit and colleagues](https://arxiv.org/abs/2410.22349)).
4. In 1,534 head-to-head votes on real AI search answers, the quality and political lean of cited news sources had no detectable effect on which answer users preferred ([Yang](https://arxiv.org/abs/2507.05301)).

## Do citations make people trust AI answers more?

Yes, in the largest experiment on the question. [Li and Aral](https://arxiv.org/abs/2504.06435) at MIT showed 4,927 US adults Google AI answers to nine public-affairs questions in spring 2024. One group of 785 people saw the answers with clickable reference links, while others saw the same text without them.

Links raised trust by 0.091 points on a 7-point scale and willingness to share by 0.121 points. Those are modest shifts in absolute terms. They are, however, about twice the size of the trust gap between AI answers and regular Google results in the same study. We look at that gap in [whether people trust AI search less](https://underneath.agency/resources/do-people-trust-ai-search-less).

People shown links also clicked more and spent more time on each answer. That sounds like checking, but the next result suggests it did not work as checking.

## Does it matter if the citations are broken or wrong?

Not measurably. Within the reference group, the researchers swapped one link for a bad one in 7 of the 9 tasks, chosen at random for each person. Most swaps were broken or invalid links; in two tasks the swap was a working link to an irrelevant page.

Trust moved by just 0.021 points when a link was bad, and by 0.009 points for an irrelevant link, neither distinguishable from zero. The authors also looked at people who spent extra time on their first question, and found no sign that they noticed invalid links either. They call this “the veneer of rigor”: citations create trust even when the references are not rigorous.

The original Google answers were not flawless either. In one, a booster every 10 years was described as “recommended for age 19-26” while the cited article recommended a much wider age range.

## Who is most swayed by citations?

People with less formal education and people outside the tech industry. Li and Aral found reference links raised trust and sharing significantly more for participants without a college degree and for those not working in technology-related industries.

How often people used AI tools made no significant difference to the citation effect. In a later essay, [Aral and colleagues](https://arxiv.org/abs/2602.13415) draw the conclusion directly: less educated and less tech-savvy consumers are more vulnerable to misrepresentations and errors in AI search.

The effect also varied by topic. Links raised trust significantly for questions on jobs and AI, police oversight and inflation, but not for gun control, immigration, taxes, healthcare, vaccines or climate change.

## How often are AI citations actually wrong?

Often enough that a citation should not be read as proof. In the 2024 audit by Narayanan Venkit and colleagues, You.com’s citations pointed to a supporting source 68.3% of the time, BingChat’s 65.8% and Perplexity’s 49.0%. Even when a supporting source was in the list, engines often cited a different one.

Google’s AI Overviews do better, but not perfectly. [Xu and colleagues](https://arxiv.org/abs/2605.14021) found 2,609 claims, in a 2026 sample of US searches, that were directly contradicted by a source the Overview itself cited. The authors note that a reader who sees only the summary has no way to detect this.

Aral and colleagues also cite a Columbia University test in which ChatGPT Search was “confidently wrong” in 146 of 200 attempts to credit quotes to their sources. And citations are not always shown at all: in their own global search data, Google’s AI answers showed references 92% of the time for general questions but only 50% of the time for Covid questions.

## Do people reward better sources?

Apparently not, at least in what they choose. [Yang](https://arxiv.org/abs/2507.05301) analyzed 24,069 conversations from an online arena where users compare two AI search answers and pick the better one. In 1,534 decided comparisons where both answers cited news sources, neither the quality nor the political lean of those sources predicted which answer won; answer length mattered most.

Yang concludes that users may not closely examine the sources cited, “provided that citations are present.” Arena volunteers are not typical consumers, and a vote is not a trust rating, so this supports the experiment rather than proving the same point.

Expert users describe the same habit. One participant in the 2024 user study put it plainly: “you see a citation, you assume it’s a valid source…I’ll just see that there is a source. That’s it. I don’t verify it.” Click data points the same way; see [how often people check AI sources](https://underneath.agency/resources/do-people-check-ai-sources).

## What does this mean for brands?

A citation borrows credibility from whoever is cited, whether or not the cited page says what the answer claims. In [our study of “Is this brand legit?” answers](https://underneath.agency/research/is-it-legit-ai-reputation-study), the engines attached a citation to 71.0% of the claims they made about brands, from 93.3% for Perplexity to 56.7% for Google AI Mode.

When we checked a sample of those cited claims against pages we could read, only 46.6% were fully supported, and others were partly supported or not supported at all. Combined with the experiment above, the implication is that readers likely give many of those claims the same credit, supported or not.

## What should you do about it?

Assume readers take every citation of your brand at face value, and check what it is attached to.

1. Read what AI answers say next to your links. A claim beside your URL gains credibility from your name even if your page does not support it. See [how often AI misattributes claims to pages](https://underneath.agency/resources/ai-citing-pages-for-claims-they-do-not-make).
2. Fix or redirect broken and outdated pages. Readers will not notice a dead link, but the assistant may still pair your name with an old claim.
3. Publish the facts you want attributed to you in clear, quotable sentences, so a correct citation is easy for the engine to make.
4. Watch the third-party pages most often cited about you, since their claims arrive with the same borrowed authority.

For help auditing how AI answers cite and describe you, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The core finding rests on one well-designed experiment, with clear limits on how far it reaches.

- Li and Aral tested nine public-affairs questions using 2024-era Google AI answers. The authors note that product and shopping searches may behave differently.
- Each person saw each answer once. No study shows whether people learn to distrust citations after spotting bad ones.
- Nobody has tested whether being cited makes readers trust the cited brand itself, as opposed to the answer.
- The arena data measures preference among volunteers, not trust among buyers.

## Frequently asked questions

### Do people click on the sources in AI answers?

Rarely. In a Pew Research Center panel of 900 US adults, only 1% of visits to Google pages with an AI Overview led to a click on a cited source.

### Are citations in AI answers reliable?

Not reliably. In a 2024 audit of three AI search engines, between 49.0% and 68.3% of citations pointed to a source that actually supported the statement.

### Would people notice if an AI answer cited a broken link?

Mostly not, in the one experiment that tested it. Swapping in broken or irrelevant links changed trust by 0.021 points on a 7-point scale, and longer reading time did not help.

### Does being cited by AI make my company look more credible?

Probably, but it is untested. The evidence shows citations raise trust in the answer; no study has yet measured the effect on trust in the cited company.

## Sources

- Li and Aral (2025), [Human Trust in AI Search: A Large-Scale Experiment](https://arxiv.org/abs/2504.06435), arXiv:2504.06435.
- Aral, Li and Zuo (2026), [The Rise of AI Search: Implications for Information Markets and Human Judgement at Scale](https://arxiv.org/abs/2602.13415), arXiv:2602.13415.
- Narayanan Venkit, Laban, Zhou, Mao and Wu (2024), [Search Engines in an AI Era: The False Promise of Factual and Verifiable Source-Cited Responses](https://arxiv.org/abs/2410.22349), arXiv:2410.22349.
- Xu, Iqbal and Montgomery (2026), [Measuring Google AI Overviews: Activation, Source Quality, Claim Fidelity, and Publisher Impact](https://arxiv.org/abs/2605.14021), arXiv:2605.14021.
- Yang (2025), [News Source Citing Patterns in AI Search Systems](https://arxiv.org/abs/2507.05301), arXiv:2507.05301.
- Chapekis, Lieb, Shah and Smith (2026), [Investigating Click Behaviors On Google Search Result Pages That Produce an AI Overview](https://arxiv.org/abs/2608.04831), arXiv:2608.04831.
- Underneath (2026), [“Is this brand legit?” How AI assistants build a reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study)

---

This is the Markdown twin of https://underneath.agency/resources/do-citations-make-ai-answers-more-trusted. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Do comparison pages help B2B brands get cited by AI? | Underneath"
description: "Partly. Comparison content gets used more once a page is cited, but brands’ own sites drew only 2.9% of AI citations in one large vendor dataset."
canonical: "https://underneath.agency/resources/do-comparison-pages-help-b2b-ai-citations"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · Content that gets cited

# Do comparison pages help B2B brands get cited by AI?

Partly: comparison content tends to shape AI answers once a page is chosen, but a vendor’s own page is rarely the one chosen. AI assistants do look for comparisons when buyers ask about software. Most of the citations, though, go to third-party lists and to other companies’ sites.

## The short version

1. In [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study), ChatGPT ran a search that compared options in 33.8% of its answers to 80 buyer questions.
2. In a dataset of 602 prompts analyzed by [Zhang and colleagues](https://arxiv.org/abs/2604.25707), cited pages with comparison content had 55.28% more influence on the answer than pages without it.
3. Across 100-plus brands tracked by the vendor [Ranqo](https://arxiv.org/abs/2606.20065), comparison pages were just 4.5% of cited content, and the tracked brand’s own site drew 2.9% of citations.
4. For software questions in the US, [Chen and colleagues](https://arxiv.org/abs/2509.08919) found AI search drew 72.7% of its sources from independent “earned” sites, against 45.4% for Google.
5. In [our pricing study](https://underneath.agency/research/ai-pricing-accuracy-study), only 61.9% of software prices quoted by AI assistants were fully faithful to the official pricing page.

## Do AI assistants look for comparisons when B2B buyers ask?

Yes, at least ChatGPT does, and it often searches the vendors’ own sites too. In [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study), we captured the searches assistants ran before answering 80 buyer questions in eight industries, from software to legal services. ChatGPT ran a search that compared options in 33.8% of its answers.

It also ran a search aimed at one option’s own website in 23.8% of answers. Gemini and Claude did neither in our sample: both shares were 0.0%. So the payoff from a comparison page depends on which assistant your buyers use.

[Chen and colleagues](https://arxiv.org/abs/2509.08919), at the University of Toronto, reached a similar view from AI prompts people shared on Reddit, mostly about consumer purchases. They found decision support dominates: people ask what to buy and how to compare, and expect a shortlist with reasons.

## How often do AI engines cite comparison pages?

Not often: ranked “best of” lists are cited far more than head-to-head comparison pages. The vendor [Ranqo](https://arxiv.org/abs/2606.20065) classified the pages cited across 100-plus brands it tracks on five AI engines. Its author works for the company, and the data come from brands tracked on its platform.

| Type of cited content | Share of content citations |
|---|---|
| Ranked “best of” list | 35.7% |
| Generic article | 31.0% |
| How-to guide | 9.7% |
| Comparison page | 4.5% |
| Alternatives page or case study | 0.7% |

A comparison page, in this scheme, is a page weighing named options against each other, such as one product versus another. Lists that rank many options win far more citations, a pattern we cover in [our article on “best of” lists](https://underneath.agency/resources/best-of-lists-ai-recommendations).

## Does comparison content shape the answer once a page is cited?

Yes, in the largest public dataset that measures it. [Zhang and colleagues](https://arxiv.org/abs/2604.25707), independent researchers, analyzed 602 prompts across ChatGPT, Google’s AI and Perplexity. They scored how much each cited page shaped the final answer, through repeated use, early placement and shared wording.

| Cited page contains… | Influence score with it | Without it | Difference |
|---|---|---|---|
| Comparison content | 0.1389 | 0.0894 | +55.28% |
| Definitions | 0.1252 | 0.0795 | +57.33% |
| Q&A format | 0.0947 | 0.1005 | -5.74% |

When the researchers looked at the role a citation played, comparisons scored 0.1524, just behind definitions at 0.1531 and well ahead of plain references. Comparison questions had the highest influence of any question type. The authors stress this is a snapshot that shows association, not cause. Whether changing a page’s format alone moves citations is covered in [our guide to restructuring existing content](https://underneath.agency/resources/does-content-structure-increase-ai-citations).

## Do comparisons win the citation against a competing page?

They help a little, but price, topic match and freshness matter far more. Researchers at the software company Sprinklr, [Vishwakarma and colleagues](https://arxiv.org/abs/2605.25517), ran 252,000 trials across six AI models. Each trial showed a model two product pages that differed in one factor, with brand names removed.

Of 18 factors tested, 11 made a reliable difference in at least four of the six models. Including comparisons with alternatives was one of seven “secondary” factors. Four “gatekeepers” mattered more: matching the topic, stating a price, carrying a recent date and appearing first.

The test used consumer product reviews, not B2B software pages, and only two pages at a time. We cover it in depth in [our article on why AI cites a competitor’s page first](https://underneath.agency/resources/why-ai-cites-competitor-page-first).

## Will AI cite a vendor’s own comparison page?

Sometimes, but independent sources are usually preferred. In Ranqo’s data, only 2.9% of citations pointed at the tracked brand’s own domain, while 75.2% went to other companies in the same space. In other words, vendor pages are cited a lot, but mostly your competitors’ pages, in answers about your category.

[Chen and colleagues](https://arxiv.org/abs/2509.08919) compared a web-enabled GPT model with Google on 1,000 ranking questions in the US and Canada in 2025. For software products in the US, the AI drew 72.7% of sources from earned media, such as review and comparison sites, and 26.7% from brands. Google’s mix was 43.7% brand, 10.9% social and 45.4% earned.

A 2025 audit of 70 B2B software prompts by [Kumar and Palkhouski](https://arxiv.org/abs/2509.10762) linked well-built pages to citation. Yet its authors warn that even high-quality pages may not be cited if they sit only on vendor blogs. More on this in [our article on whether AI cites your own website](https://underneath.agency/resources/do-ai-engines-cite-your-own-website).

## Do self-serving comparison lists work?

They do get cited, but there is no sign they win the answer. In [our study of “best of” lists](https://underneath.agency/research/self-promoting-best-lists-study), 26.0% of cited numbered lists with an identifiable publisher included that publisher. Of those, 92.9% ranked themselves first, and most were software companies writing about their own category.

Such lists were only 1.1% of all citations in our data. Across answers that cited a list, the list’s top pick was named in 47.1% of cases for self-ranking lists and 43.4% for independent ones. In the cleaner comparison, within the same answers, there was no detectable difference.

## Do comparison pages help when buyers have not named you?

Mostly no: comparison questions usually already include your brand. In Ranqo’s data, 46% of comparison prompts named the brand being tracked. On average, brands appeared in 74.6% of answers to comparison prompts, against 22.9% for broad discovery questions that named no brand.

That makes a comparison page a tool to defend a shortlist you are already on. It does little to get you onto the shortlist in the first place, which depends more on lists, reviews and coverage elsewhere.

## Does accuracy on comparison pages matter?

Yes, because assistants repeat what they find, including stale figures on your own site. In [our pricing study](https://underneath.agency/research/ai-pricing-accuracy-study), 61.9% of 840 plan prices quoted by four assistants for 45 software products were fully faithful to the pricing page. For 39 of 64 prices that differed, the same figure appeared on another page of the vendor’s own site.

Comparison pages often restate your prices and your rivals’ prices. If they go out of date, an assistant may quote the old figure to a buyer.

## What should you do about it?

Publish comparison and definition content where it helps buyers decide, and keep your main effort on third-party coverage.

1. **Build pages that answer the comparison, not just praise you.** State criteria, prices, specifications and the cases where a rival fits better; these are the parts AI answers reuse.
2. **Define your category plainly.** Definitions were the most influential citation role in the Zhang dataset, so a clear explainer of what your product category is and does is worth having.
3. **Put prices and dates on the page.** Both were gatekeepers in the Sprinklr test, and both need regular checks.
4. **Keep every page that states a price in sync.** Retire or update old comparison pages so assistants cannot pick up stale figures.
5. **Invest in independent lists and reviews.** Ranked lists and earned media draw most citations for software questions, and they decide who reaches the shortlist.
6. **Do not rely on self-ranking lists.** They get cited, but in our data their top pick fared no better than an independent list’s.
7. **Test on the assistants your buyers use.** Only ChatGPT ran comparison searches in our sample, so check results engine by engine.

If you want help deciding which comparison and category pages to build, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

No study has yet tested whether adding a comparison page to a B2B site raises live AI citations.

- **The usage findings are snapshots.** The Zhang dataset links comparison content to influence but cannot show that adding it causes more.
- **The controlled test is not B2B.** The 252,000-trial study used anonymized consumer reviews, two pages at a time.
- **The citation shares come from a vendor.** Ranqo’s figures come from its own customers’ tracked brands and prompts.
- **Comparison pages and lists overlap.** Studies classify pages differently, so “comparison” and “best of” shares are not directly comparable.
- **Nobody has measured buyer outcomes.** No study tracks whether a cited comparison page leads to demos, trials or deals.

## Frequently asked questions

### Should B2B companies create “X vs Y” comparison pages for AI search?

Yes, if they are honest and specific, but expect modest citation gains. Comparison content shaped answers more in one 602-prompt dataset, yet comparison pages were only 4.5% of cited content in a large vendor dataset.

### Do AI assistants cite vendor websites for software comparisons?

Less than independent sources. For US software questions, one 2025 study found AI search drew 26.7% of its sources from brands and 72.7% from earned media.

### Do “alternatives to” pages get cited by ChatGPT?

Rarely, on current evidence. Alternatives pages and case studies were 0.7% of cited content in one vendor’s dataset of AI citations.

### Does ranking yourself first in a “best of” list work for AI search?

Not clearly. In our study, 92.9% of cited lists that included their publisher put it first, but their top pick was not named more often than an independent list’s.

### Are definition pages useful for AI visibility?

They appear to be. Pages with definitions had 57.33% more influence on AI answers in one dataset, and definitions were the most influential role a citation played.

## Sources

- Zhang, He and Yao (2026), [From Citation Selection to Citation Absorption: A Measurement Framework for Generative Engine Optimization Across AI Search Platforms](https://arxiv.org/abs/2604.25707), arXiv:2604.25707.
- Kumar (2026), [Generative Engine Optimization at Scale: Measuring Brand Visibility Across AI Search Engines](https://arxiv.org/abs/2606.20065), arXiv:2606.20065.
- Chen, Wang, Chen and Koudas (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919), arXiv:2509.08919.
- Vishwakarma, Kumar and Jamidar (2026), [What Gets Cited: Competitive GEO in AI Answer Engines](https://arxiv.org/abs/2605.25517), arXiv:2605.25517.
- Kumar and Palkhouski (2025), [AI Answer Engine Citation Behavior An Empirical Analysis of the GEO16 Framework](https://arxiv.org/abs/2509.10762), arXiv:2509.10762.
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [How many “best of” lists cited by AI rank their own brand first?](https://underneath.agency/research/self-promoting-best-lists-study)
- Underneath (2026), [How faithfully do AI assistants quote software prices?](https://underneath.agency/research/ai-pricing-accuracy-study)

---

This is the Markdown twin of https://underneath.agency/resources/do-comparison-pages-help-b2b-ai-citations. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Do FAQ pages help you get cited by AI search engines?"
description: "Not by format alone. In the research so far, FAQ formatting did not make AI engines use a page more; the evidence inside the page did."
canonical: "https://underneath.agency/resources/do-faq-pages-help-ai-citations"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Do FAQ pages help you get cited by AI search engines?

Not on their own: in the studies so far, question-and-answer formatting did not make AI search engines use a page more. What went with heavier use was the substance inside the page, such as definitions, figures, comparisons and depth. An FAQ can help when it carries that substance, and does little when it is a thin wrapper.

## The short version

1. In a study of 21,143 citations across ChatGPT, Google and Perplexity, pages in a question-and-answer format were used 5.74% less on average than other pages ([Zhang Kai and colleagues](https://arxiv.org/abs/2604.25707)).
2. In the same data, pages containing numbers were used 61.55% more, a far larger gap than any format effect.
3. In [our study of 3,096 pages](https://underneath.agency/research/ai-overview-cited-pages-study) ranking alongside Google’s AI Overviews, a page with an FAQ heading was 2.5 points less likely to be cited, a gap we could not tell apart from zero.
4. In a test on Amazon product listings, adding an FAQ moved a listing up by only 0.05 places on average ([Bagga and colleagues](https://arxiv.org/abs/2511.20867)).
5. In a small study of 14 Tokyo hotels, a hotel with a 33-question FAQ was cited by Gemini, while a hotel with a short FAQ was not ([Zhu and Chang](https://arxiv.org/abs/2603.20062)).

## Does an FAQ format make AI engines use your page more?

No: in the largest study on this, question-and-answer pages were used slightly less than other pages. The researchers ran 602 test questions through ChatGPT, Google’s AI Overview and Gemini, and Perplexity, then examined 18,151 cited pages they could fetch ([Zhang Kai and colleagues](https://arxiv.org/abs/2604.25707)).

They separated two outcomes. Being cited means a page appears in the source list. Being used, which they call absorption, means the answer’s wording and facts come from that page.

They scored use by how often and how early a page was referenced and how much of the answer overlapped with it. On that score, pages in a question-and-answer format were used 5.74% less on average than other pages.

The authors are careful: this was a simple average with no controls, and their label for FAQ pages may be noisy. They say it “does not prove that FAQ content is harmful.” It does show the format alone is no shortcut.

## What do AI engines actually draw on, if not the format?

They draw on the evidence inside a page. In the same study, pages with certain kinds of content were used far more than pages without them:

| Page contains | Gap in how much answers drew on the page |
|---|---|
| Numbers or statistics | 61.55% more |
| Definitions | 57.33% more |
| Comparisons | 55.28% more |
| Step-by-step guidance | 41.20% more |
| Question-and-answer format | 5.74% less |

The most-used pages were also built in sections. Among the quarter of pages that answers drew on most, the average page had 10.59 headings, against 0.85 for the quarter drawn on least.

The authors’ explanation is plain. An FAQ often produces short, isolated answers that cannot support a longer response. A detailed explainer gives the engine more to work with.

Their advice is that the value “comes from the evidence inside the page, not from the presence of question marks in headings.” The wider pattern is laid out in [our guide to what content AI engines favor](https://underneath.agency/resources/what-content-do-ai-engines-prefer).

## Do FAQ headings or FAQ markup help in Google’s AI Overviews?

The evidence says no reliable lift. In [our AI Overview study](https://underneath.agency/research/ai-overview-cited-pages-study), we compared 3,096 pages ranking in Google’s top 10 for US searches. Every search showed an AI Overview, the AI summary at the top of Google’s results.

We compared each page only with pages on the same search, so differences in topic did not count. On that basis, a page with an FAQ heading was 2.5 points less likely to be cited, a gap we could not tell apart from zero.

Question-style headings looked like a strong signal in our first version, at 9.6 points. Compared like with like, that fell to 3.5 points and is no longer distinguishable from zero.

FAQ markup, the code that labels questions and answers for search engines, showed a 5.9-point gap. It did not hold up once we allowed for testing 14 page features at once. Ranking position mattered far more than any of these features.

## Does adding an FAQ change anything in controlled experiments?

Barely. In a lab test built on 13,747 shopping questions paired with Amazon listings, researchers rewrote listings 15 different ways. They measured how far each version moved in five AI engines’ product rankings ([Bagga and colleagues](https://arxiv.org/abs/2511.20867)).

Adding an FAQ was one of only four rewrites that matched or beat a plain rewrite. Its average gain was 0.05 places, close to nothing. The other eleven rewrites did worse, and some, like making listings read like advertisements, did real harm.

A second lab test points the same way. Researchers at the software company Sprinklr ran 252,000 trials across six AI models, each time giving the model two near-identical sources ([Vishwakarma and colleagues](https://arxiv.org/abs/2605.25517)).

Changing layout alone, such as dense paragraphs against organized sections, had no consistent effect on which source was cited first. Relevance, prices and dates did. The split evidence on layout changes is weighed in [our guide to restructuring existing content](https://underneath.agency/resources/does-content-structure-increase-ai-citations).

## When does an FAQ page actually work?

An FAQ works when it is deep enough to answer real questions. The clearest example comes from a study of hotel recommendations in Tokyo, run by an AI company’s researchers ([Zhu and Chang](https://arxiv.org/abs/2603.20062)).

They scored 14 hotel websites on content depth, including how full their FAQ was. Hotels that Gemini cited directly averaged 8.6 out of 15 on depth; hotels it did not cite averaged 3.4.

One independent hotel, Kadoya, had no special markup but a 33-question FAQ and a detailed sightseeing guide, and Gemini cited it directly. Another, Hotel K5, had a 9.6 out of 10 rating on a booking site, markup and an FAQ, but each feature was brief. Gemini named K5 but drew its facts from booking and editorial sites instead.

This is 14 hotels, one city and one engine, and the authors say it shows association, not cause. It fits the larger studies: the depth of the answers mattered, not the FAQ label. The full hotel comparison is in [our guide to deep content without technical SEO](https://underneath.agency/resources/deep-content-without-technical-seo-gemini).

## What should you do about it?

Treat an FAQ as a container for evidence, not as a ranking tactic.

1. Do not convert pages to FAQ format expecting AI engines to use them more. No study we reviewed shows that the format alone does this.
2. Where you keep an FAQ, make each answer substantial. Include the specific figure, the definition, the comparison or the steps a buyer needs.
3. Put your strongest evidence where it is easy to lift: clear sections with headings, numbers stated plainly, and comparisons laid out side by side.
4. Keep markup as tidy housekeeping, not as a growth plan. In our data it did not reliably move citations.
5. Measure whether answers use your page, not only whether they link to it. The two can differ widely.

If you want help turning existing pages into evidence AI answers can use, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

No study has yet tested FAQ formatting on live pages with a proper before-and-after design.

- The main FAQ result is a simple average. It does not rule out that thin FAQ pages, not the format, explain the gap.
- Our AI Overview study and the hotel study observed pages as they were. Neither changed a page and watched the result.
- The lab tests used product listings and paired sources, not full websites in live AI search.
- Results come from 2025 and 2026 versions of fast-changing products. Engines may weigh pages differently after updates.
- Nobody has measured whether FAQ pages drive more clicks or sales from AI answers.

## Frequently asked questions

### Should we add FAQ schema to get into AI Overviews?

Not as a growth tactic. In our study of 3,096 ranking pages, FAQ markup showed a 5.9-point gap that did not survive allowing for testing 14 features at once.

### Is FAQ content bad for AI search?

No. The largest study found question-and-answer pages used 5.74% less on average, and its authors say this does not prove FAQ content is harmful.

### What content do AI answers draw on most?

Pages with numbers, definitions, comparisons and steps. Pages with numbers or statistics were used 61.55% more than pages without them in a study of three AI search engines.

### Do question-style headings help a page get cited?

The evidence is weak. In our data the gap fell from 9.6 points to 3.5 once pages were compared on the same search, and it could not be told apart from zero.

## Sources

- Zhang Kai, He Xinyue and Yao Jingang (2026), [From Citation Selection to Citation Absorption: A Measurement Framework for Generative Engine Optimization Across AI Search Platforms](https://arxiv.org/abs/2604.25707), arXiv:2604.25707.
- Bagga, Farias, Korkotashvili, Peng and Wu (2026), [E-GEO: A Testbed for Generative Engine Optimization in E-Commerce](https://arxiv.org/abs/2511.20867), arXiv:2511.20867.
- Vishwakarma, Kumar and Jamidar (2026), [What Gets Cited: Competitive GEO in AI Answer Engines](https://arxiv.org/abs/2605.25517), arXiv:2605.25517.
- Zhu and Chang (2026), [The End of Rented Discovery: How AI Search Redistributes Power Between Hotels and Intermediaries](https://arxiv.org/abs/2603.20062), arXiv:2603.20062.
- Underneath (2026), [What pages cited by AI Overviews have in common: 3,096 pages](https://underneath.agency/research/ai-overview-cited-pages-study)

---

This is the Markdown twin of https://underneath.agency/resources/do-faq-pages-help-ai-citations. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Do statistics, quotes and citations get you cited by AI?"
description: "Mostly not. A 2025 benchmark found most AI rewriting tactics did nothing or lowered a page’s rank; adding statistics hurt in 19 of 24 settings."
canonical: "https://underneath.agency/resources/do-geo-content-tactics-work"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · Content that gets cited

# Do GEO content tactics like adding statistics, quotes or citations actually get you cited more by AI?

Mostly not, on the best current evidence. The early study that made these tactics famous found gains in a simulated setup. A larger 2025 benchmark found most of them ineffective and often harmful. A 2026 review of 45 studies found no tactic with a stable, lasting effect across AI platforms. Real, relevant evidence still matters; automatically sprinkling in statistics and quotes does not.

## The short version

1. The original 2023 study by [Aggarwal and colleagues](https://arxiv.org/abs/2311.09735) reported 30 to 40% gains from adding sources, quotations and statistics. That was in a simulated engine where the page was already among five sources.
2. A 2025 benchmark by [Puerto and colleagues](https://arxiv.org/abs/2506.11097) found the Statistics method lowered a page’s rank in 19 of 24 settings tested.
3. In an e-commerce test by [Bagga and colleagues](https://arxiv.org/abs/2511.20867), eleven of fifteen hand-written rewriting recipes made product listings rank worse.
4. In a 252,000-trial test by [Vishwakarma and colleagues](https://arxiv.org/abs/2605.25517), topic match, price, a recent date and list position decided which source was cited first; formatting alone had no effect.
5. A 2026 review by [Martinez](https://arxiv.org/abs/2607.14035) of 45 studies found no technique with a stable, long-term, cross-platform effect on being found.

## Where did the idea that statistics and quotes work come from?

From the 2023 study that coined generative engine optimization (GEO). [Aggarwal and colleagues](https://arxiv.org/abs/2311.09735) built a test engine using GPT-3.5 that answered questions from the top five Google results. They rewrote one source at a time and measured how much of the answer reflected it.

Their top methods, Cite Sources, Quotation Addition, and Statistics Addition, achieved a relative improvement of 30-40% on their main visibility measure. On Perplexity, tested on 200 queries with source text uploaded as files, Quotation Addition gave a 22% improvement. [Keyword stuffing, an old search trick](https://underneath.agency/resources/does-keyword-stuffing-work-in-ai-search), performed 10% worse than doing nothing.

Two caveats matter. The gains measure how much of the answer drew on a page already handed to the AI, not whether the page gets found. [Martinez](https://arxiv.org/abs/2607.14035) notes the famous “up to 40%” comes mainly from a rise from 19.3 to 27.2 on one score. It does not mean 40% more readers or 40% more chance of being retrieved. We trace the number in [our guide to the 40% GEO claim](https://underneath.agency/resources/does-geo-increase-ai-visibility-by-40-percent).

## Did later tests confirm those gains?

Mostly no: a larger benchmark found most of these tactics did nothing or made things worse. [Puerto and colleagues](https://arxiv.org/abs/2506.11097) tested ten rewriting methods across six domains, from retail products to news, on four AI models. They measured whether a rewritten page got cited earlier in the answer.

In their words, most current C-SEO methods are not only largely ineffective but also frequently have a negative impact on document ranking. C-SEO is their term for rewriting content for conversational search. The Statistics method decreases rankings in 19 out of 24 evaluated settings. In product recommendation tasks on Anthropic’s Haiku 3.5 model, 26 out of 30 cases show significant negative effects.

Only two methods showed real gains, and only in narrow cases: a summary placed at the top of the page and a combined content rewrite. Even those worked only for retail and video games on one model, and none worked for question answering.

| Study | Tactic tested | Result |
|---|---|---|
| Aggarwal et al. (2023) | Statistics, quotes, citations | Gains of 30 to 40% in a simulated engine |
| Puerto et al. (2025) | Same tactics and more | Mostly no effect; statistics hurt in 19 of 24 settings |
| Bagga et al. (2025) | Fifteen rewriting recipes for product listings | Four modestly matched or beat a plain rewrite; eleven did worse |

The e-commerce study by [Bagga and colleagues](https://arxiv.org/abs/2511.20867) used 13,747 realistic consumer product queries, each paired with 10 Amazon listings. Hand-written recipes such as “sound authoritative” or “add technical terms” did not reliably beat a simple rewrite. Whether plainer, smoother prose helps is covered in [our guide on readable writing and AI answers](https://underneath.agency/resources/does-readable-writing-help-ai-visibility).

## Why do the studies disagree?

They measure different things, and gains for one page come at another’s expense. The 2023 study measured the share of the answer’s words tied to a page. The 2025 benchmark measured the page’s citation rank. Puerto and colleagues note that the 2023 study’s own position-weighted results showed scores generally falling.

Competition matters too. In the 2023 study, adding citations raised the visibility of the fifth-ranked source by 115.1%, while the top-ranked source lost 30.3%. Puerto and colleagues found that gains shrank steadily as more sites adopted the same method. If every competitor adds statistics, nobody stands out.

## What matters more than rewriting?

Being retrieved and ranked high in the first place, and matching what the buyer asked. Puerto and colleagues found that moving a document to the top of the AI’s reading list produced far greater gains than any rewriting method.

[Vishwakarma and colleagues](https://arxiv.org/abs/2605.25517), at the software company Sprinklr, ran 252,000 head-to-head trials across six AI models. Two product pages differed in one factor at a time. Four factors won across all six models: matching the topic, stating the price, a recent date rather than an old one, and appearing first in the list. We cover those trials in [our guide to why AI cites a rival’s page](https://underneath.agency/resources/why-ai-cites-competitor-page-first).

Our own data agrees on ranking. In [our study of pages cited by AI Overviews](https://underneath.agency/research/ai-overview-cited-pages-study), 41.7% of pages ranking 1 to 3 were cited, against 20.1% at positions 7 to 10. Position came first; the page features we measured explained little more.

## Is adding real evidence different from adding statistics?

Yes: relevant, verifiable evidence helps; inserted numbers for their own sake may not. In the Vishwakarma trials, claims without supporting evidence such as tests or certifications lost out. That is real substance, not decoration.

Martinez puts it plainly: the criterion is not to add numbers, but to provide relevant, verifiable, dated and properly attributed evidence. The automated “Statistics” rewrites in these benchmarks were generated by an AI model, so they may not reflect what real, sourced data does.

There is also a trust risk. [Chu and colleagues](https://arxiv.org/abs/2608.16824) scanned 10,095 pages that Google Search and Gemini retrieved for real queries. They estimated 8.90% were optimized for AI engines, rising to 16.36% among pages modified in 2026. On those pages, 69.34% of citations were rated low on verifiability.

## Can automated rewriting tools do better?

In lab tests, yes, but those tests are simulations. [Wu and colleagues](https://arxiv.org/abs/2510.11438) built a system that learns what AI engines prefer and rewrites pages to match. It reported an average improvement of 35.99% on visibility measures, using simulated engines given five candidate documents each.

Bagga and colleagues found an automatically tuned rewriting instruction beat every hand-written recipe. Its rewrites turned marketing prose into clear, labeled product facts. Neither study tested live ChatGPT or Google results.

## What should you do about it?

Spend on substance and findability, not on stylistic rewrites.

1. **Stop paying for “add stats and quotes” rewrites.** The strongest benchmark found these did nothing or harmed ranking in most settings.
2. **Keep classic search work.** Ranking high enough to be retrieved mattered more than any rewrite in the controlled tests and in our own data.
3. **Answer the actual question.** Topic match was a gatekeeper in all six models tested. Write pages that address what buyers ask, in their terms.
4. **Publish facts buyers need.** State prices, specifications and dates clearly. These were among the strongest factors in head-to-head trials.
5. **Use evidence you can stand behind.** Cite real tests, data and sources. Fabricated or vague statistics risk trust if a reader or an AI checks them.

If you want help deciding where content work will pay off, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

No study yet shows any content tactic lifting citations on live AI search over time.

- **Almost every test is simulated.** The engines were built by researchers from fixed sets of documents, not live ChatGPT or Google results.
- **The rewrites were machine-made.** “Add statistics” meant an AI inserting statistical elements, not a company publishing its own research.
- **English and narrow domains.** The 2025 benchmark used English only, across six domains.
- **Vendor research.** The 252,000-trial study comes from a software company and tested only two pages at a time with brand names removed.
- **No long-term evidence.** Martinez found no reviewed technique with a stable, long-term, cross-platform effect on being found, or on clicks and sales.

## Frequently asked questions

### Does adding statistics to content help with AI search?

Not as a rewriting trick. In one benchmark, automated statistics insertion lowered rankings in 19 of 24 settings, though relevant, verifiable evidence may still help.

### What is the “40% GEO improvement” people quote?

It comes from the 2023 GEO study’s simulated engine. It describes a larger share of the answer drawn from a page already given to the AI, not 40% more traffic.

### Is GEO just SEO under a new name?

Not entirely, but ranking still matters most. In controlled tests, moving a page to the top of the AI’s reading list beat every rewriting method.

### Do quotes and citations make AI trust my page more?

The evidence is mixed. Quotes helped in the 2023 simulated test but not in a 2025 benchmark; claims backed by real evidence did better in a 2026 head-to-head test.

## Sources

- Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan and Deshpande (2023), [GEO: Generative Engine Optimization](https://arxiv.org/abs/2311.09735), arXiv:2311.09735.
- Puerto, Gubri, Green, Oh and Yun (2025), [C-SEO Bench: Does Conversational SEO Work?](https://arxiv.org/abs/2506.11097), arXiv:2506.11097.
- Bagga, Farias, Korkotashvili, Peng and Wu (2025), [E-GEO: A Testbed for Generative Engine Optimization in E-Commerce](https://arxiv.org/abs/2511.20867), arXiv:2511.20867.
- Vishwakarma, Kumar and Jamidar (2026), [What Gets Cited: Competitive GEO in AI Answer Engines](https://arxiv.org/abs/2605.25517), arXiv:2605.25517.
- Wu, Zhong, Kim and Xiong (2025), [What Generative Search Engines Like and How to Optimize Web Content Cooperatively](https://arxiv.org/abs/2510.11438), arXiv:2510.11438.
- Chu, Leng, Li, Shen, Shen and Zhang (2026), [GEO-Flag: Detecting and Measuring GEO-Optimized Web Content](https://arxiv.org/abs/2608.16824), arXiv:2608.16824.
- Martinez (2026), [Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026)](https://arxiv.org/abs/2607.14035), arXiv:2607.14035.
- Underneath (2026), [What pages cited by AI Overviews have in common](https://underneath.agency/research/ai-overview-cited-pages-study)

---

This is the Markdown twin of https://underneath.agency/resources/do-geo-content-tactics-work. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Do people actually check the sources AI assistants cite?"
description: "Rarely. In a Pew panel of 900 US adults, only 1% of visits to Google pages with an AI Overview led to a click on a cited source."
canonical: "https://underneath.agency/resources/do-people-check-ai-sources"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Do people actually check the sources AI assistants cite?

Rarely. In a Pew Research Center panel that tracked the real browsing of 900 US adults, only 1% of visits to Google pages with an AI Overview led to a click on one of its cited sources. Smaller studies point the same way: people glance at one or two citations at most, and the presence of a citation, not its quality, seems to satisfy them.

## The short version

1. Only 1% of visits to Google pages with an AI Overview led to a click on a cited source, in a month of tracked browsing by 900 US adults ([Chapekis and colleagues](https://arxiv.org/abs/2608.04831)).
2. Expert users in a 2024 study looked at about 2 sources in an AI answer, against 12 on a regular Google page ([Narayanan Venkit and colleagues](https://arxiv.org/abs/2410.22349)).
3. When the AI answer agreed with their own view, those experts clicked a source 0.48 times per question on average (Narayanan Venkit and colleagues).
4. A week of Google’s AI Mode cut click-through to other websites by 18.8 percentage points, in a field test with 1,100 US users ([Wang and colleagues](https://arxiv.org/abs/2608.18352)).

## How often do people click the sources in AI answers?

Almost never, in the best real-world data available. Pew researchers [Chapekis and colleagues](https://arxiv.org/abs/2608.04831) tracked the browsing of 900 US adults from a representative panel throughout March 2025, covering 68,879 distinct Google searches. About 18% of those searches produced an AI Overview, the AI summary at the top of Google’s results.

Just 1% of visits to pages with an AI Overview led to a click on a cited source. Clicks on any result also fell: people clicked a search result on 8% of visits to pages with an Overview, against 15% on pages without one. They ended their browsing session on 26% of pages with an Overview, against 16% without.

This is observational data, so it shows association, not cause. It also counted only the first three cited sources, the most a desktop Overview showed without expanding.

## Do people at least look at the sources?

Briefly, and far less than on a normal results page. In a 2024 study by [Narayanan Venkit and colleagues](https://arxiv.org/abs/2410.22349), 21 participants with PhD-level expertise searched the same questions on an AI answer engine and on Google, thinking aloud as they went.

| Behavior in the 2024 user study | AI answer engines | Google |
|---|---|---|
| Sources hovered over, on average | 2 | 12 |
| Sources clicked, on average | usually 1 | 4 |

Agreement made checking even rarer. When the question matched the participant’s own opinion, they clicked a source just 0.48 times on average and hovered over 1.08. Questions that challenged their views prompted noticeably more checking.

Participants described the pull of a confident answer. One said: “It writes so confidently, I feel convinced without even looking at the source.” How that confidence compares with trust in Google results is covered in [our guide to trust in AI search](https://underneath.agency/resources/do-people-trust-ai-search-less).

## Does the checking that does happen catch bad citations?

Not reliably. In the same 2024 study, the authors found that misattributed citations were noticed only by the few participants who scrutinized the sources. Perplexity, for example, displayed 5.00 sources per answer on average but cited only 2.58 of them in the text.

A large experiment found the same blind spot. [Li and Aral](https://arxiv.org/abs/2504.06435) gave some of their 4,927 US participants AI answers containing broken or irrelevant links. People who spent extra time and effort on a question showed no sign of noticing the bad links, and people who trusted AI more clicked more but spent less time evaluating what they found. The effect of those links on trust is covered in [our guide on whether citations build trust](https://underneath.agency/resources/do-citations-make-ai-answers-more-trusted).

## Do people choose answers with better sources?

Not that researchers can detect. [Yang](https://arxiv.org/abs/2507.05301) studied 1,534 head-to-head votes on an online arena where users compare two AI search answers. The quality and political lean of the news sources each answer cited made no detectable difference to which one users preferred.

What did matter was length: longer answers tended to win. Arena volunteers are not typical consumers, so this is supporting evidence, not proof.

## Does AI search keep people from visiting websites at all?

It seems to, especially in conversational modes. In the field test by [Wang and colleagues](https://arxiv.org/abs/2608.18352), switching 1,100 US users to Google’s AI Mode for a week cut click-through to other websites by 18.8 percentage points. The share of users clicking through to news sites fell 12.5 points, to Reddit 21.2 points and to Wikipedia 9.9 points.

The reverse also held. Hiding AI Overviews raised click-through by 8.8 percentage points, though only about half of Overviews were successfully hidden.

Standalone assistants show a similar pattern in panel data from [Iannelli and Ai](https://arxiv.org/abs/2607.04282) of Scrunch AI, a company that sells AI visibility tools, in a paper not yet peer reviewed. User-weighted, 34.1% of sessions involving an AI assistant showed no visit to an outside website, against 19.5% of search sessions by the same people. More of the evidence is gathered in [our guide to clicks lost to AI answers](https://underneath.agency/resources/ai-answer-engines-fewer-website-clicks).

## Why does this matter for brands?

Because most people read only the answer, what the answer says about you matters more than whether they click. The checking is done by the machine instead. In [our study of hidden searches](https://underneath.agency/research/ai-hidden-searches-study), ChatGPT ran a mean of 3.7 searches of its own per buyer question before answering, reading pages the user never sees.

Being cited is also a separate route from ranking. [Xu and colleagues](https://arxiv.org/abs/2605.14021) found that 29.8% of the domains cited by Google’s AI Overviews did not appear anywhere in the first page of regular results for the same search.

## What should you do about it?

Measure and manage what AI answers say, not just the traffic they send.

1. Track whether and how you appear in the answer text for your buyers’ key questions. Referral clicks badly understate how many people read about you.
2. Make the facts you most need buyers to know easy to quote, in plain sentences on pages AI assistants can read. Few readers will open the page to find the nuance.
3. Fix inaccurate claims at their source, since readers are unlikely to click through and spot them.
4. Treat AI referral traffic as a floor, not a measure of influence.

For help measuring how AI answers present your brand, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The strongest data covers Google’s AI Overviews; much less is known about standalone assistants.

- The Pew panel covers one month in 2025 and counts clicks on only the first three sources in each Overview.
- The 2024 user study involved 21 experts in a think-aloud session, which may encourage more checking than everyday use.
- The only large panel on ChatGPT-style assistants comes from a vendor and has not yet been peer reviewed.
- Not clicking is not the same as not verifying. People may check elsewhere, later or by asking again, and no study measures that.
- No study we found looks specifically at whether buyers check sources for product or brand questions.

## Frequently asked questions

### What percentage of people click on AI Overview sources?

About 1%. In Pew’s tracked panel of 900 US adults, just 1% of visits to Google pages with an AI Overview led to a click on a cited source.

### Do people verify what ChatGPT tells them?

The evidence is thin. In one vendor panel, 34.1% of sessions involving an AI assistant showed no visit to any outside website, but that does not prove the answer went unchecked.

### Does Google’s AI Mode send traffic to websites?

Less than regular search. In a 2026 field test, a week of AI Mode cut click-through to other websites by 18.8 percentage points.

### If people do not click, does being cited still matter?

Yes. Citations raise trust in an answer even when nobody follows them: in one experiment, adding reference links lifted trust by 0.091 points on a 7-point scale.

## Sources

- Chapekis, Lieb, Shah and Smith (2026), [Investigating Click Behaviors On Google Search Result Pages That Produce an AI Overview](https://arxiv.org/abs/2608.04831), arXiv:2608.04831.
- Narayanan Venkit, Laban, Zhou, Mao and Wu (2024), [Search Engines in an AI Era: The False Promise of Factual and Verifiable Source-Cited Responses](https://arxiv.org/abs/2410.22349), arXiv:2410.22349.
- Li and Aral (2025), [Human Trust in AI Search: A Large-Scale Experiment](https://arxiv.org/abs/2504.06435), arXiv:2504.06435.
- Yang (2025), [News Source Citing Patterns in AI Search Systems](https://arxiv.org/abs/2507.05301), arXiv:2507.05301.
- Wang, Gleason, Bart, Wilson and Metaxa (2026), [AI in Search Reduces Publisher Referrals Without Improving User Experience: Experimental Evidence](https://arxiv.org/abs/2608.18352), arXiv:2608.18352.
- Iannelli and Ai (2026), [The New Shape of Search: How Conversational AI Recomposes Information Seeking](https://arxiv.org/abs/2607.04282), arXiv:2607.04282.
- Xu, Iqbal and Montgomery (2026), [Measuring Google AI Overviews: Activation, Source Quality, Claim Fidelity, and Publisher Impact](https://arxiv.org/abs/2605.14021), arXiv:2605.14021.
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)

---

This is the Markdown twin of https://underneath.agency/resources/do-people-check-ai-sources. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Do people trust AI search less than regular Google results?"
description: "Slightly. In US experiments, the same text was trusted a little less when labeled AI, and a week of Google’s AI Mode lowered trust in Google."
canonical: "https://underneath.agency/resources/do-people-trust-ai-search-less"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Do people trust AI search answers less than regular Google results?

Slightly, on average, and much depends on who is asking and how the answer is presented. In a controlled test with 4,927 US adults, identical information was trusted a little less and shared less when it was labeled as AI-generated. In a week-long field test, being switched to Google’s AI Mode lowered trust in information found on Google, while removing AI Overviews made no detectable difference.

## The short version

1. In a preregistered experiment with 4,927 US adults, the same Google text was rated slightly less trustworthy, and less worth sharing, when labeled as AI-generated ([Li and Aral](https://arxiv.org/abs/2504.06435)).
2. The gap was small: 0.043 points on a 7-point trust scale, and it did not survive a stricter re-test, while the drop in willingness to share did (Li and Aral).
3. In a 2026 field test with 1,100 US Google users, a week of forced AI Mode lowered trust in information on Google by 0.34 points on a 7-point scale ([Wang and colleagues](https://arxiv.org/abs/2608.18352)).
4. How AI answers are presented moves trust more than the AI label does: highlighting uncertain passages cut trust by 0.157 points, while adding reference links raised it by 0.091 (Li and Aral).

## Do people trust the same answer less when it is labeled AI?

Yes, slightly, according to the largest experiment so far. [Li and Aral](https://arxiv.org/abs/2504.06435) at MIT ran it from March 20 to April 2, 2024, with 4,927 participants matched to the US adult population on age, gender and ethnicity. Each person read Google answers to 9 policy and public-affairs questions, from gun control to adult vaccine schedules.

The trick was that everyone saw identical text. One group was told it came from Google’s generative AI; another saw it in the classic Google snippet format and was told it was not AI. People rated how accurate, believable, biased, complete and trustworthy each answer was, and whether they would share it with a friend.

Labeling the text as AI lowered both trust and willingness to share. The size matters, though. Traditional search was trusted just 0.043 points more on a 7-point scale, and that difference did not hold up when the authors re-tested the data with a stricter method. The sharing gap was larger, at 0.141 points, and it did hold.

## Does using Google’s AI Mode change trust in Google?

Yes: a week of AI Mode lowered trust in information found on Google. [Wang and colleagues](https://arxiv.org/abs/2608.18352), from the University of Pennsylvania and Northeastern University, ran a preregistered field experiment in March 2026 with a browser extension that changed what participants saw on Google.

After three normal days, 1,100 US Google users were split into three groups for seven days: AI Mode for every search, Google as it is today, or Google with AI Overviews hidden. Forced AI Mode lowered trust in information on Google by 0.34 points on a 7-point scale. It also lowered perceived usefulness, satisfaction and sense of control, and the share of people searching on a competing engine rose by 11.2 percentage points.

Hiding AI Overviews, by contrast, had no detectable effect on trust. Two caveats limit how far this travels. The sample skewed young and educated (75% under 45, 87% with at least some college), and participants had used AI Mode for only 0.6% of their searches before the test, so this measures a sudden, forced switch.

## What do people say when asked about AI Mode?

They mention control, sources and accuracy more often than trust. In the same field test, the researchers coded 309 open-ended comments from people in the AI Mode group after a week of use.

| Theme in comments after a week of AI Mode | Share of 309 responses |
|---|---|
| Overall negative | 33.6% |
| Overall positive | 29.3% |
| Limited links or source diversity | 13.4% |
| Hallucination or accuracy concerns | 4.6% |
| Expressed trust in AI results | 1.6% |

Expert users voice similar doubts. In a 2024 user study by [Narayanan Venkit and colleagues](https://arxiv.org/abs/2410.22349), 14 of 21 participants with PhD-level expertise said AI answer engines took extra work to verify and trust. And 12 of 20 distrusted the kinds of sources the engines used, such as forums, blogs and LinkedIn posts. One kind of source that deserves that caution appears in [our study of self-ranking “best of” lists](https://underneath.agency/research/self-promoting-best-lists-study).

## Who trusts AI search more, and who less?

Frequent AI users, college graduates and people working in technology trust it more. In Li and Aral’s experiment, people who used AI tools and AI search often trusted generative results significantly more, and were more willing to share them, than people who used them rarely or never.

Use was already common. In that 2024 sample, 85% of participants said they used generative AI, and 63% of those used it to look up information. Participants with a college degree, those working in technology-related industries, and Republicans also trusted the AI results more than their counterparts.

The pattern suggests the trust gap may narrow as use grows, though no study has yet tracked the same people over time to test that. That trust is also why [ads blended into AI answers](https://underneath.agency/resources/risks-of-native-ads-in-ai-answers) raise concern.

## Does presentation change trust more than the AI label?

Yes: in the same experiment, design features moved trust more than the AI label itself. [Adding reference links](https://underneath.agency/resources/do-citations-make-ai-answers-more-trusted) raised trust by 0.091 points on the 7-point scale, roughly twice the size of the AI-label gap. Highlighting passages the system was unsure about lowered trust by 0.157 points, and a generic explanation of how AI search works had no effect on average.

Trust also shapes behavior. Li and Aral found that people who trusted AI results more clicked on them more but spent less time evaluating them. How rarely readers open the cited pages is covered in [our guide on checking AI sources](https://underneath.agency/resources/do-people-check-ai-sources). In the AI Mode field test, people spent 0.43 more minutes per search session, yet click-through to other websites fell by 18.8 percentage points.

## What should you do about it?

Assume AI answers about your company are trusted almost as much as search results, and manage them accordingly.

1. Do not discount AI answers as low-credibility. The measured trust penalty for the AI label was small, and frequent users showed little of it.
2. Check what AI Mode, AI Overviews and the main assistants say about you, since the people most likely to read those answers are also the ones most likely to trust them.
3. Give skeptical readers something to verify. Clear, specific pages that support the claims made about you serve the people who do click through.
4. Track AI Mode separately from classic search. In the field test it changed both trust and clicking behavior far more than AI Overviews did.

For help monitoring and improving how AI search describes you, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The two strongest studies cover public-affairs questions and Google’s own products, not brand or product searches.

- Li and Aral tested 9 policy-style questions and note that results may differ for product searches or everyday tasks.
- Their experiment ran in early 2024. Trust may have shifted as AI answers became routine, and no follow-up has measured that.
- The AI Mode field test lasted one week, with a younger, more educated and more left-leaning sample than the US population.
- Both studies are US-only. Trust in AI search may differ across countries and languages.
- We found no study comparing trust in standalone assistants such as ChatGPT with trust in Google’s results.

## Frequently asked questions

### Do people trust AI Overviews less than normal Google results?

A little, in one large experiment. The same text was trusted 0.043 points less on a 7-point scale when labeled AI, a gap that did not survive a stricter re-test.

### Does Google’s AI Mode make people trust Google less?

In one field test, yes. A week of forced AI Mode lowered trust in information on Google by 0.34 points on a 7-point scale among 1,100 US users.

### Did removing AI Overviews increase trust in Google?

No detectable change. In the same field test, hiding AI Overviews did not measurably change trust or the overall search experience.

### Who is most likely to trust AI search results?

People who use AI often. In Li and Aral’s experiment, frequent users of AI tools, college graduates and tech-industry workers trusted generative search results more than others.

## Sources

- Li and Aral (2025), [Human Trust in AI Search: A Large-Scale Experiment](https://arxiv.org/abs/2504.06435), arXiv:2504.06435.
- Wang, Gleason, Bart, Wilson and Metaxa (2026), [AI in Search Reduces Publisher Referrals Without Improving User Experience: Experimental Evidence](https://arxiv.org/abs/2608.18352), arXiv:2608.18352.
- Narayanan Venkit, Laban, Zhou, Mao and Wu (2024), [Search Engines in an AI Era: The False Promise of Factual and Verifiable Source-Cited Responses](https://arxiv.org/abs/2410.22349), arXiv:2410.22349.

---

This is the Markdown twin of https://underneath.agency/resources/do-people-trust-ai-search-less. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Does ad spend help your brand get recommended by AI?"
description: "Not directly, on current evidence: search interest and online conversation tracked AI brand prominence far more closely than advertising spend did."
canonical: "https://underneath.agency/resources/does-ad-spend-help-ai-brand-recommendations"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Does advertising spend help our brand get recommended by AI, or does something else matter more?

Not directly, as far as current research can tell: what people search for and say about a brand tracked AI recommendations far more closely than ad budgets did. In a Northwestern study of 82 brands, advertising added little once search interest and online conversation were counted. The study is exploratory and cannot prove cause, but it points to where AI assistants get their picture of a brand: other people’s words.

## The short version

1. Across 82 brands in five categories, Google search interest was the strongest signal of how prominently AI models recommended a brand, and ad spend added little once the signals were considered together ([Malthouse and colleagues](https://arxiv.org/abs/2609.16304), 2026).
2. In [our study of four AI assistants](https://underneath.agency/research/brand-entity-ai-recommendations-study), each tenfold rise in independent sites naming a brand went with 4.7 times the odds of a recommendation.
3. In tests run in August 2025, 93.5% of the sources ChatGPT used for well-known brand questions were earned media, not brand-owned pages ([Chen and colleagues](https://arxiv.org/abs/2509.08919)).
4. In one vendor’s data, only 2.9% of AI citations pointed to the tracked brand’s own website ([Kumar](https://arxiv.org/abs/2606.20065), 2026).
5. In [our hidden searches study](https://underneath.agency/research/ai-hidden-searches-study), ChatGPT ran a search aimed at a named publication, ranking or award in 43.8% of its answers.

## Does ad spend predict which brands AI assistants recommend?

Not on its own, in the most direct study so far. [Malthouse and colleagues](https://arxiv.org/abs/2609.16304) at Northwestern University measured how often, and how high, six AI models recommended brands in five categories, from cordless drills to cruises. They then compared that with five measures of each brand’s presence in the market.

| Signal | Source used |
|---|---|
| Advertising spend | Kantar |
| News mentions | LexisNexis |
| Online brand conversation | Brandwatch |
| Search interest | Google Trends |
| Information seeking | Wikipedia page views |

Across 82 brands, Google search interest was the strongest and most consistent signal, followed by online conversation. Advertising, news mentions and Wikipedia views contributed relatively little once all five were considered together. A stricter second check dropped advertising, news and Wikipedia entirely and kept only search interest and conversation.

## How sure can we be about this?

Not very sure yet: it is an exploratory pattern in 82 brands, not proof of cause. The authors say plainly that their results should not be read causally. All five measures rise and fall together, and each one on its own was linked to more prominent recommendations, advertising included.

So the finding is narrower than “ads do not matter”. Ads are part of a brand’s general market visibility. They simply added little extra once search interest and conversation were known.

The measurement has limits too. Advertising was measured over the 12 months to May 2026, conversation volume was estimated by Brandwatch from a 5% random sample of mentions, and the AI models answered with web search switched off. Results with live web search may differ.

## Why would search interest matter more than advertising?

The research does not say, and the authors warn against assuming searches cause recommendations. Both could reflect the same underlying thing: genuine consumer interest in a brand, which also shows up in what people write and discuss.

What other studies do show is that AI assistants lean on third-party sources. [Chen and colleagues](https://arxiv.org/abs/2509.08919) at the University of Toronto sorted the sources behind AI answers into brand-owned, earned and social. For niche brands, 95.1% of ChatGPT’s sources were earned media, such as publications, reviews and expert sites.

[Our brand entity study](https://underneath.agency/research/brand-entity-ai-recommendations-study) points the same way. Across ChatGPT, Gemini, Perplexity and Claude, the signal most closely linked to being recommended was independent coverage: how many other websites named the brand in the pages the assistants cited.

## What do AI answers draw on instead of brand messaging?

Mostly other companies’ pages, review lists and coverage, rarely the brand’s own site. Kumar, a co-founder of the AI visibility company Ranqo, classified 149,912 citations from his company’s tracking data. Only 2.9% pointed to the tracked brand’s own domain, while 75.2% pointed to other companies in the same space. The wider evidence is in [our guide on whether AI cites your own site](https://underneath.agency/resources/do-ai-engines-cite-your-own-website).

Among other sources, YouTube was cited most, at 4.2% of citations. Ranked “best of” lists were the single most cited kind of content page, at 35.7% of content citations. This is a vendor’s data from its own customers, so treat the exact shares with caution.

The assistants also go looking for outside verdicts. In [our hidden searches study](https://underneath.agency/research/ai-hidden-searches-study), ChatGPT ran a mean of 3.7 web searches per buyer question. When one of those searches named a source such as NerdWallet or Avvo, the answer cited that source 44.0% of the time.

## Can you simply pay to appear in AI answers?

Not inside the organic answer, under current published policy. [Chu and Hou](https://arxiv.org/abs/2606.17443) note that in February 2026, OpenAI began testing ads in ChatGPT for US users on its Free and Go plans. The paid placements appear below the organic answer and are labeled, and OpenAI states that organic answers are generated independently. Why such labels matter to advertisers is covered in [our guide to labeling ads in AI search](https://underneath.agency/resources/should-ai-search-ads-be-labeled).

Their experiments also suggest that classic sales language does little. In skincare tests on three AI models, evidence-style claims beat the leading brand 50% to 73% of the time, while urgency and scarcity copy did so only 10% to 13% of the time. Authority claims were worth the same as +0.17 rating points of real product improvement.

One warning applies. Those authority claims were deliberately invented, including fake clinical trials, to find the upper limit. The authors class made-up claims as potential false advertising and recommend only real, verifiable evidence.

## What should you do about it?

Keep advertising for what it does well, and fund the things AI assistants actually read.

1. **Do not expect ads to buy recommendations.** On current evidence, spend shows little direct link to AI prominence once interest and conversation are counted.
2. **Watch search interest and conversation.** Track them next to AI visibility as early signals, while remembering the link is not proven to be causal.
3. **Invest in earned coverage.** Reviews, rankings, awards and “best of” lists from credible publishers are what assistants search for and cite.
4. **Write evidence, not slogans.** Real certifications, published results and clear specifications moved AI choices far more than urgency copy.
5. **Test before reallocating budget.** Change one input, then measure AI answers over several weeks and runs. It helps to settle first [which team owns AI search visibility](https://underneath.agency/resources/who-should-own-ai-search-visibility).

If you want help planning that mix, see [how we approach generative engine optimization](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The research describes associations, and the key experiment has not been run.

- **Nobody has varied ad spend and measured AI answers.** The Northwestern study observed brands as they were.
- **The sample is small and specific.** It covers 82 brands in five US-oriented categories, with models answering from training only.
- **The effect of ChatGPT ads is unknown.** Testing began in 2026, and no study yet measures how paid placements interact with organic recommendations.
- **The direction of cause is open.** Whether rising search interest leads to more AI recommendations is a question the authors propose testing next.

## Frequently asked questions

### Does advertising help a brand show up in ChatGPT?

Not directly, on current evidence. In a study of 82 brands, ad spend added little to AI recommendation prominence once search interest and online conversation were taken into account.

### What predicts whether AI assistants recommend a brand?

Search interest and online conversation in one study, and independent coverage in ours. Brands named by more independent websites had 4.7 times the odds of a recommendation for each tenfold increase.

### Should we move budget from ads to PR for AI search?

Not wholesale, because the evidence is correlational. A safer step is to fund earned coverage alongside ads and test whether AI visibility moves.

### Are there ads in ChatGPT answers?

OpenAI began testing labeled ads below organic answers for some US users in February 2026, and says organic answers are generated independently of them.

## Sources

- Malthouse, Lee, Yang, Pal and Feng (2026), [Evaluating Brand Retrieval and Ranking in Large Language Model Recommendations](https://arxiv.org/abs/2609.16304), arXiv:2609.16304.
- Chen, Wang, Chen and Koudas (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919), arXiv:2509.08919.
- Kumar (2026), [Generative Engine Optimization at Scale: Measuring Brand Visibility Across AI Search Engines](https://arxiv.org/abs/2606.20065), arXiv:2606.20065.
- Chu and Hou (2026), [Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems](https://arxiv.org/abs/2606.17443), arXiv:2606.17443.
- Underneath (2026), [Do Wikipedia and schema make AI assistants recommend a brand?](https://underneath.agency/research/brand-entity-ai-recommendations-study)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)

---

This is the Markdown twin of https://underneath.agency/resources/does-ad-spend-help-ai-brand-recommendations. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Does AI search actually use the pages it cites? | Underneath"
description: "Not evenly. AI search engines lean heavily on some cited pages and barely use others, so a citation count overstates how much a page shaped the answer."
canonical: "https://underneath.agency/resources/does-ai-search-use-the-pages-it-cites"
published: 2026-10-07
updated: 2026-10-07
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Does AI search actually use the pages it cites?

Only partly, and very unevenly. AI search engines draw heavily on some of the pages they cite and barely touch others. So being listed as a source is not the same as shaping what the answer says.

## The short version

1. ChatGPT cited 6.88 sources per prompt and Perplexity 16.35, yet each page ChatGPT cited shaped its answer about four times as much, in a study of 602 prompts ([Zhang, He and Yao](https://arxiv.org/abs/2604.25707)).
2. Google’s AI Overviews cited about twice as many sources as OpenAI’s search-enabled GPT, but drew less of their content into the summary, scoring 0.447 against 0.500 ([Huang and colleagues](https://arxiv.org/abs/2603.16138)).
3. In a 2024 audit, 36% of the sources BingChat listed were never cited in its answer text ([Narayanan Venkit and colleagues](https://arxiv.org/abs/2410.22349)).
4. Encyclopedia pages averaged an influence score of 0.2144, against 0.0726 for news pages, even though news is cited often ([Zhang, He and Yao](https://arxiv.org/abs/2604.25707)).

## What is the difference between being cited and being used?

Being cited means your page appears in the source list. Being used means its facts or wording shape the answer.

[Zhang Kai, He Xinyue and Yao Jingang](https://arxiv.org/abs/2604.25707) call these two stages “citation selection” and “citation absorption.” They studied 602 prompts across ChatGPT, Google’s AI Overviews and Gemini, and Perplexity, covering 21,143 citations. For each cited page they built an influence score from 0 to 1. It rises when the answer refers to the page often and early, covers it across several paragraphs, and shares its wording.

The idea is not theirs alone. [Martinez’s 2026 survey](https://arxiv.org/abs/2607.14035) of GEO research notes that a source may be “cited without being used, paraphrased without a link, or decisive in shaping the structure of an answer.” The reverse also happens: an answer can draw on pages it never lists, as [Tannenbaum](https://arxiv.org/abs/2609.22655) points out.

Note the source. The absorption study comes from three independent researchers, uses a public dataset, and has not been peer reviewed.

## Which AI search engines give each cited source more influence?

ChatGPT cites fewer pages but leans on each one more. Google and Perplexity spread their answers across many more sources.

| Engine | Sources cited per prompt | Average influence of each fetched page (0 to 1) |
|---|---|---|
| ChatGPT | 6.88 | 0.2713 |
| Google AI Overviews and Gemini | 12.06 | 0.0584 |
| Perplexity | 16.35 | 0.0646 |

The gap widened on hard questions. For prompts with several constraints, ChatGPT averaged just 3.4 citations, while Perplexity averaged 17.7. The authors read this as ChatGPT doing more of the work itself after picking fewer sources, but say that is an interpretation, not a proven cause.

A separate team reached a similar result by another method. [Huang and colleagues](https://arxiv.org/abs/2603.16138) at the University of Illinois collected answers to 11,000 real search queries. For 1,100 queries per system, they checked which facts in each AI summary could be traced to each cited page. Google’s AI Overviews cited nearly twice as many sources as OpenAI’s search-enabled GPT, at 5.0 against 2.7. Yet they captured less of those sources’ content, scoring 0.447 against 0.500. More citations meant thinner use of each.

## Are some cited sources only there for show?

Often, yes. Several audits found sources that are listed or cited but add little or nothing to the answer.

A 2024 audit by [Narayanan Venkit and colleagues](https://arxiv.org/abs/2410.22349) tested You.com, BingChat and Perplexity on 303 queries. Of the sources Perplexity listed, 8% were never cited in the answer text; for BingChat it was 36%. Every engine included at least 30% of sources that backed no unique information in the answer. Users in their study called this “buffing”: an impression of thoroughness without substance.

The gap can also be selective. Huang and colleagues found that Google’s AI Overviews cite Reddit and Quora, but drew 22.1 percentage points less content from those forum sources than from others. Sources with a negative tone were drawn on 13.8 percentage points less in AI Overviews. The authors call this a “citation-synthesis gap”: the list looks diverse while the text relies on a narrower set.

The role a page plays matters too. In the absorption study, pages cited as a definition averaged an influence of 0.1531. Pages cited only as a reference averaged 0.0529, about a third as much.

## Which cited pages get used the most?

Pages that supply a definition, a figure or a comparison get used most. Pages that are merely relevant get used least.

The absorption study shows a clear split between being picked and being used. News sites are cited often, yet news pages averaged an influence of 0.0726, against 0.2144 for encyclopedia pages. Pages cited as a direct source of facts averaged 0.1241; pages used only as background, 0.0775.

Format alone did not help. Pages written in a question-and-answer format scored 5.74% lower than other pages. We cover that result in [our article on FAQ pages](https://underneath.agency/resources/do-faq-pages-help-ai-citations). Length helped when it carried substance: Huang and colleagues found short sources were drawn on 17.7 percentage points less in Google’s AI Overviews.

## Does a cited page always back up what the answer says?

No. Some claims in AI answers are not supported by any page the answer cites.

In a 40-day audit of Google’s AI Overviews in spring 2026, [Xu, Iqbal and Montgomery](https://arxiv.org/abs/2605.14021) checked 98,020 claims. They found 11.0% were not supported by the cited pages, mostly because no cited page mentioned the claim at all. We look at this problem, including claims wrongly pinned on a page, in [our article on misattributed citations](https://underneath.agency/resources/ai-citing-pages-for-claims-they-do-not-make).

## Why isn’t citation count enough to measure GEO?

Because a count treats a passing reference and the page that wrote half the answer as equal. It also rewards engines that cite many sources lightly.

Count citations on Perplexity and your page may appear often while shaping little. On ChatGPT a page may appear less often but carry far more of the answer. The absorption study’s authors warn that “a dashboard that counts citations alone will miss the central pattern of this dataset.” They suggest tracking five things side by side: how often you are selected, how many citations you get, how much your page shapes the answer, whether the answer is accurate, and how concentrated citations are.

We discuss how to set up those measures in [our article on what to measure](https://underneath.agency/resources/what-to-measure-ai-visibility) and in [our guide to share of citations](https://underneath.agency/resources/measuring-share-of-citations-in-ai-search).

## What should you do about it?

Measure whether AI answers use your pages, not only whether they list them. Then make your pages easier to use.

1. Read the answer, not just the source list. For your key buyer questions, check which of your facts, figures or phrasing appear in the text.
2. Report results by engine. A citation on ChatGPT and one on Perplexity are not worth the same.
3. Give engines something to take. Clear definitions, specific numbers and comparisons went with the most use.
4. Check accuracy. When your page is cited, confirm the claim next to it matches what your page says.
5. Do not chase citation counts alone. A rise in citations with no change in what answers say about you is a weak win.

If you want help building this kind of measurement, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

Researchers can describe how unevenly sources are used, but cannot yet see inside the engines. The main gaps:

- The influence score is a stand-in built from wording overlap and position. Martinez notes it “does not reveal an internal causal trace.” A true test would compare answers with and without a source.
- Measures of use only cover pages researchers could read. The absorption study fetched 76.44% of cited pages. In [our reputation study](https://underneath.agency/research/is-it-legit-ai-reputation-study), only 38.6% of cited pages could be read automatically.
- The studies are snapshots. Engines change quickly, and the absorption study lacks dates for each record.
- No study yet links being used, rather than merely cited, to clicks, trust or sales.

## Frequently asked questions

### What is citation absorption in AI search?

It is how much a cited page actually shapes an AI answer. Zhang, He and Yao coined the term to separate it from simply appearing in the source list.

### Which AI search engine relies most on each source it cites?

ChatGPT, in the one study that measured it. Each fetched page ChatGPT cited scored 0.2713 on influence, against 0.0584 for Google and 0.0646 for Perplexity.

### Is citation count a good GEO KPI?

Not on its own. Citations vary in how much they shape the answer, so pair the count with a check of what each answer actually says.

### Does Perplexity use every source it lists?

Not always. In a 2024 audit, 8% of the sources Perplexity listed were not cited in its answer. In a later study, its cited pages averaged a low influence score of 0.0646.

## Sources

- Zhang, He and Yao (2026), [From Citation Selection to Citation Absorption: A Measurement Framework for Generative Engine Optimization Across AI Search Platforms](https://arxiv.org/abs/2604.25707), arXiv:2604.25707.
- Huang, Goyal, Saha and Chandrasekharan (2026), [Answer Bubbles: Information Exposure in AI-Mediated Search](https://arxiv.org/abs/2603.16138), arXiv:2603.16138.
- Narayanan Venkit, Laban, Zhou, Mao and Wu (2024), [Search Engines in an AI Era: The False Promise of Factual and Verifiable Source-Cited Responses](https://arxiv.org/abs/2410.22349), arXiv:2410.22349.
- Xu, Iqbal and Montgomery (2026), [Measuring Google AI Overviews: Activation, Source Quality, Claim Fidelity, and Publisher Impact](https://arxiv.org/abs/2605.14021), arXiv:2605.14021.
- Martinez, O. (2026), [Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026)](https://arxiv.org/abs/2607.14035), arXiv:2607.14035.
- Tannenbaum (2026), [Scoring With the Engine: Retrieval Exposure, Cross-Engine Divergence, and the Limits of Engine-Agnostic GEO Scores](https://arxiv.org/abs/2609.22655), arXiv:2609.22655.
- Underneath (2026), [“Is this brand legit?” How AI assistants build a reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study)

---

This is the Markdown twin of https://underneath.agency/resources/does-ai-search-use-the-pages-it-cites. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Does showing up in AI answers actually drive business results?"
description: "Possibly, but not yet proven. A 2026 review of 45 studies found that claims about GEO’s return outstrip the evidence; none tied AI visibility to sales."
canonical: "https://underneath.agency/resources/does-ai-visibility-drive-business-results"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search and business value

# Does showing up in AI answers actually drive business results?

Possibly, but no published study has yet shown that appearing in AI answers raises sales. The research does show that AI answers shape what buyers see and trust, often without a click. So the value is real enough to measure, but too unproven to promise.

## The short version

1. A 2026 review of 45 studies by [Martinez](https://arxiv.org/abs/2607.14035) concluded that claims about GEO return on investment clearly outstrip the academic evidence.
2. In the clearest field test, by [Watanabe and Nakayashiki](https://arxiv.org/abs/2606.04362), ChatGPT referrals to one site grew 5.7 times, but pages that got no optimization grew 3.5 times anyway.
3. In a panel of US adults, [Pew researchers](https://arxiv.org/abs/2608.04831) found just 1% of visits to pages with an AI Overview led to a click on a cited source.
4. In an audit of real shopping questions, [Uberti-Bona Marin and colleagues](https://arxiv.org/abs/2609.18729) found ChatGPT stated its own product preference in 79% of answers that recommended products.
5. A 2026 framework by [Kato and colleagues](https://arxiv.org/abs/2609.11915) shows how to value AI visibility like a media channel, but it has only been tested on simulated sales data.

## Is there evidence that AI visibility leads to sales?

No published study has yet linked appearing in AI answers to sales or revenue. [Martinez](https://arxiv.org/abs/2607.14035), at Sciences Po, reviewed 45 studies published between late 2023 and mid-2026. The section on business outcomes is titled “Traffic and Conversions: The Weakest Evidence.”

His verdict is blunt: at this stage, claims about GEO return on investment clearly outstrip the academic evidence. He lists the claim that citation scores predict clicks, conversions or revenue as rejected as a general claim. Only one suggestive field test and a few industry reports exist.

One of those industry reports describes a 20% production traffic lift against a control group. The review notes it does not describe group sizes, how pages were assigned or how uncertain the figure is. That is traffic, not sales, and it cannot support a general estimate.

## What does the best real-world test show?

It shows a likely traffic gain, not a sales gain, and even the traffic gain is uncertain. [Watanabe and Nakayashiki](https://arxiv.org/abs/2606.04362) work at Glasp, a web highlighting service, and studied their own site. They optimized one section of it and left the rest alone as a comparison.

Total ChatGPT referrals grew by a factor of 5.7, but untouched pages grew by a factor of 3.5 as ChatGPT itself grew. After allowing for that, they estimate an extra lift of about 1.82 times for the optimized pages. A stricter check found jumps almost as large before the work began, so the authors call the effect suggestive, not conclusive.

The study measured visits from ChatGPT, not sign-ups or revenue. We cover how to read results like these in [our article on crediting ChatGPT referral growth](https://underneath.agency/resources/chatgpt-referral-growth-and-geo) and in [how to prove GEO caused a change in sales](https://underneath.agency/resources/prove-geo-caused-sales).

## Why isn’t website traffic a fair measure of the value?

Because much of what AI answers do for a brand happens without any click. In a panel of 900 US adults tracked in March 2025, [Chapekis and colleagues](https://arxiv.org/abs/2608.04831) at Pew found just 1% of visits to pages with an AI Overview led to a click on a cited source. An AI Overview is the AI summary at the top of Google’s results.

Assistants show the same pattern. In a browsing panel studied by [Iannelli and Ai](https://arxiv.org/abs/2607.04282), 34.1% of sessions that used an AI assistant showed no visit to any outside website. The authors work for Scrunch AI, a vendor of AI visibility tools, and note a session with no visit is not proof the question was answered. See [our article on assistant sessions without website visits](https://underneath.agency/resources/ai-assistant-sessions-without-website-visits).

AI search can also make it harder to reach your site at all. In a one-week experiment with 1,100 US Google users, [Wang and colleagues](https://arxiv.org/abs/2608.18352) found that switching people to AI Mode cut their click-through rate by 18.8 percentage points. Among participants who described their week with AI Mode, 15.3% mentioned difficulty getting to specific websites.

## Can an AI answer influence buyers who never click?

Very likely, though no study has yet measured the purchases that follow. AI assistants increasingly act like advisers rather than lists of links. In an audit of real shopping questions, [Uberti-Bona Marin and colleagues](https://arxiv.org/abs/2609.18729) found ChatGPT stated a first-person product preference in 79% of answers that recommended products, against 7% for Gemini and 2% for AI Overviews.

The same audit notes that, in the largest published measurement of ChatGPT use, product and service recommendations were roughly 2% of conversations. Small as a share, that is a large number of buying conversations. The authors also found that the recommended products often changed when the same question was asked again.

Buyers already lean on AI search. Among participants in a 2024 study by [Li and Aral](https://arxiv.org/abs/2504.06435) at MIT, 85% used AI tools, and 63% of those used them to look up information. Their experiment with 4,927 Americans found that citations raised trust in AI answers even when the links were wrong; see [our article on citations and trust](https://underneath.agency/resources/do-citations-make-ai-answers-more-trusted).

## Does every kind of appearance count the same?

No: being described when someone names you differs from being recommended when they don’t. In a study of 112 Product Hunt startups, [Sharma](https://arxiv.org/abs/2601.00912) found ChatGPT recognized 99.4% of them when asked by name. Asked open discovery questions, it surfaced them only 3.32% of the time.

Appearances also come and go. In [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), only 25.2% of the brands ChatGPT named for a buyer question appeared in all five runs of that question. A single appearance in one answer is a sample, not a position you hold.

And an appearance can carry the wrong details. In [our study of business facts](https://underneath.agency/research/ai-business-facts-accuracy-study), 18.9% of 636 answers from four AI engines gave at least one address, phone number, website or opening time that differed from the Google profile. A mention that sends a buyer to an old phone number has negative value.

## How should executives measure the ROI of generative search visibility?

Count how often buyers likely notice your brand in answers, then tie that to sales using a comparison group. [Kato and colleagues](https://arxiv.org/abs/2609.11915) propose fitting AI visibility into marketing mix modeling, the method many firms already use to value media spend. Their key point is that a mention in an answer is not yet a media input.

To turn it into one, the authors combine four things:

| Input | What it means | Where it comes from |
|---|---|---|
| How often answers name you | Share of repeated answers to relevant questions that mention the brand | Running your buyer questions many times |
| How many such questions are asked | Volume of relevant questions in each market and period | Market estimates |
| Which AI tools handle them | Each assistant’s share of those questions | Usage data |
| Whether readers notice | Chance a reader registers the name in the answer | User studies |

They tested the idea with 2,240 answers to 56 product questions in English and Japanese. The brand they tracked was Glasp, the same site as the field test above; it appeared in 33.8% of answers from one GPT model and 27.8% from another. The link to sales in their paper is simulated, so the method is promising but unproven on real revenue.

The authors also stress that a content change can pay off outside AI answers, through ordinary search traffic or a better landing page. A fair measure counts those paths too, rather than crediting only the AI mention. Measurement like this is also the base of [a longer-term AI search strategy](https://underneath.agency/resources/ai-search-strategy-next-five-years).

## What should you do about it?

Treat AI visibility as a channel worth testing, and do not book its return until you have measured it.

1. **Measure exposure, not just visits.** Track how often each assistant names you across repeated runs of your buyers’ real questions; referral traffic misses most of the influence.
2. **Separate being described from being recommended.** Report questions that name your brand apart from open category questions; the second is where new demand comes from.
3. **Check what the answers say.** Audit prices, contact details and claims, because a wrong mention can cost you a customer.
4. **Build in a comparison group.** Leave some comparable pages, products or markets untouched so platform growth is not mistaken for your work, as explained in [our article on judging a rise in AI visibility](https://underneath.agency/resources/did-geo-work-raise-ai-visibility).
5. **Connect to the models you already use.** If finance runs marketing mix modeling, add noticed AI mentions as an input rather than building a separate scorecard.
6. **Plan for the click losses too.** Value from AI answers sits beside traffic you may lose; see [what lost clicks mean for pipeline](https://underneath.agency/resources/ai-answers-pipeline-revenue) and [whether citations make up for them](https://underneath.agency/resources/ai-overview-citations-vs-lost-clicks).

If you want help setting up this kind of measurement, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The research has not yet answered the central question: how much revenue an AI mention is worth.

- **No sales outcome has been measured.** The best field test tracked ChatGPT visits on one site, and its effect is only suggestive.
- **The valuation method is untested on real sales.** The marketing mix framework uses real answers but simulated business results.
- **How many people notice a mention is unknown.** No published study yet measures how often readers register a brand name inside an answer.
- **Influence without clicks is inferred, not counted.** Studies show trust and stated preferences, not the purchases that follow.
- **Most evidence is narrow.** Key studies cover one site, US panels or a few weeks, and some come from AI visibility vendors.

## Frequently asked questions

### Does being cited by an AI search engine actually generate business value?

It can, but nobody has yet measured how much. A 2026 review of 45 studies found the claim that citations predict conversions or revenue is not supported as a general rule.

### How should executives measure the ROI of generative search visibility?

Measure how often buyers are likely to see your brand in answers, then compare sales against a group that got no GEO work. A 2026 framework adds noticed AI mentions to marketing mix modeling, though it has only been tested on simulated sales.

### Is AI referral traffic a good measure of GEO success?

Only partly. In a Pew panel, 1% of visits to pages with an AI Overview led to a click on a cited source, so most of the exposure never shows up as traffic.

### Do AI recommendations influence what people buy?

They likely do, but purchases have not been measured. ChatGPT stated a personal product preference in 79% of answers that recommended products in one audit, and people trust answers more when they carry citations.

## Sources

- Martinez (2026), [Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026)](https://arxiv.org/abs/2607.14035), arXiv:2607.14035.
- Watanabe and Nakayashiki (2026), [Disentangling Answer Engine Optimization from Platform Growth: A Log-Based Natural Experiment on ChatGPT Referral Traffic](https://arxiv.org/abs/2606.04362), arXiv:2606.04362.
- Chapekis, Lieb, Shah and Smith (2026), [Investigating Click Behaviors On Google Search Result Pages That Produce an AI Overview](https://arxiv.org/abs/2608.04831), arXiv:2608.04831.
- Iannelli and Ai (2026), [The New Shape of Search: How Conversational AI Recomposes Information Seeking](https://arxiv.org/abs/2607.04282), arXiv:2607.04282.
- Wang, Gleason, Bart, Wilson and Metaxa (2026), [AI in Search Reduces Publisher Referrals Without Improving User Experience: Experimental Evidence](https://arxiv.org/abs/2608.18352), arXiv:2608.18352.
- Uberti-Bona Marin, Bertaglia, Astante, Rijsbosch, van Dijck, Hannák, Spanakis and Kollnig (2026), ["If I Had to Buy Just ONE: Galaxy S26 Ultra": Auditing AI-Generated Product Recommendations](https://arxiv.org/abs/2609.18729), arXiv:2609.18729.
- Li and Aral (2025), [Human Trust in AI Search: A Large-Scale Experiment](https://arxiv.org/abs/2504.06435), arXiv:2504.06435.
- Sharma (2026), [The Discovery Gap: How Product Hunt Startups Vanish in LLM Organic Discovery Queries](https://arxiv.org/abs/2601.00912), arXiv:2601.00912.
- Kato, Honma and Kato (2026), [Generative Marketing Mix Modeling: A Causal Inference Framework Linking GEO and GEM to Business Impact](https://arxiv.org/abs/2609.11915), arXiv:2609.11915.
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)
- Underneath (2026), [Do AI answers match a business’s Google profile?](https://underneath.agency/research/ai-business-facts-accuracy-study)

---

This is the Markdown twin of https://underneath.agency/resources/does-ai-visibility-drive-business-results. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Does blocking Google-Extended cut your AI Overview visibility?"
description: "Possibly. A large 2025 study linked blocking Google-Extended to fewer AI Overview citations; our 2026 check saw no clear effect. Weigh it as a real risk."
canonical: "https://underneath.agency/resources/does-blocking-google-extended-hurt-ai-overviews"
published: 2026-10-10
updated: 2026-10-10
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Does blocking Google-Extended in robots.txt reduce our visibility in AI Overviews?

Possibly, and the evidence is mixed, so treat it as a real risk rather than a settled fact. A large study of searches collected in December 2025 found that sites blocking Google-Extended were significantly less likely to be cited in AI Overviews, even though Google says the setting does not affect them. Our own smaller check in September 2026 found no visible difference.

## The short version

1. Google-Extended is a robots.txt setting that asks Google not to use a site for Gemini; 12.3% of top sites with a readable robots.txt use it to block their whole site (our crawler study, September 2026).
2. Across 11,500 queries collected in December 2025, sites blocking Google-Extended were significantly less likely to be cited in AI Overviews, and 21 popular publishers that block it were never cited by Gemini (Grossman and colleagues).
3. In our same-day check of 481 AI Overviews in September 2026, 17.8% of cited pages on top sites came from sites using Google-Extended, close to the 16.8% expected, so no clear effect showed.
4. Training and search blocks are separate elsewhere too: 51.7% of sites blocking OpenAI’s training crawler still allowed its search crawler, and ChatGPT cited none of 123 checked pages closed to that search crawler.
5. Blocked AI agents answer anyway from other sources: in a 2026 vendor test, answers drew 58% from a business’s own pages when its site was hard to read, against 78% when it was easy.

## What does Google-Extended actually control?

It controls whether Gemini may use your content, not your Google ranking. According to Google’s own description, as reported by [Grossman and colleagues](https://arxiv.org/abs/2604.27790), sites can block it to keep their content out of training future Gemini models and out of the sources Gemini cites. Google also says it does not affect ranking in Google Search, nor AI Overviews’ access to the content.

Google-Extended is a name that robots.txt files can address, not a separate crawler. In our [crawler blocking study](https://underneath.agency/research/ai-crawler-blocking-study) of 5,572 top sites with a readable robots.txt, 12.3% blocked Google-Extended from the whole site. 10.1% of all sites opted out of Gemini training this way while still allowing Googlebot, Google’s search crawler.

The only way to keep content out of AI Overviews, during spring 2026, was through general search settings. [Xu and colleagues](https://arxiv.org/abs/2605.14021) note that publishers could restrict AI Overview use with noindex and nosnippet instructions, which also affect regular search results.

## Does blocking it reduce AI Overview citations?

One large study says yes; our smaller check found no clear difference. Grossman and colleagues collected Google results, AI Overviews and Gemini answers for 11,500 queries, collected in December 2025 from a single location in New Jersey. AI Overviews appeared on 51.5% of the representative real-user queries.

Comparing each source’s chances in AI Overviews with its chances in the regular results for the same query, they found that sites blocking Google-Extended were significantly less likely to be cited in AI Overviews. Sites such as the New York Times, Yelp, IMDb and Facebook had fewer AI Overview citations. The authors call this surprising and suggest publishers may need to rethink the block.

Our cross-check pointed the other way. We matched the sources in 481 US AI Overviews from 26 September 2026 against robots.txt files. 17.8% of cited pages on top-10,000 sites came from sites using Google-Extended, against 16.8% expected from their popularity.

| Study | When | What it found |
|---|---|---|
| Grossman et al., 11,500 queries | December 2025 | Blocking sites significantly less likely to be cited in AI Overviews |
| Underneath, 481 AI Overviews | September 2026 | 17.8% of cited pages from blocking sites, against 16.8% expected |

Neither study is a controlled test. Sites that block Google-Extended differ from other sites in size and type, and both studies cover a single point in time.

## What happens to your visibility in Gemini?

Gemini stops citing you, as Google says it will. Grossman and colleagues found 21 popular publishers, each retrieved for at least 20 queries by both Google Search and AI Overviews, that Gemini never cited. All of them block Google-Extended.

The list included major news and reference brands. Popular social and review sites, including Facebook, Instagram, TikTok, IMDb, Yelp and Tripadvisor, also received no Gemini citations. The authors describe the resulting loss of Gemini visibility as self-inflicted, because Google is doing what these sites asked.

Robots.txt is not the only factor, though. In our study, all 64 AI Overview citations of pages closed even to Googlebot were Reddit pages, and Google has a data licensing agreement with Reddit.

## Do other AI companies treat training and search separately?

Yes. Most run separate crawlers, and the search crawler is the one tied to citations. In our crawler study, 51.7% of the sites that blocked OpenAI’s training crawler, GPTBot, still allowed OAI-SearchBot, the crawler behind ChatGPT search.

On the same day, ChatGPT cited 123 pages on top sites whose robots.txt we could check, and none was closed to OAI-SearchBot. Six of them were closed to GPTBot and were cited anyway. In that sample, training opt-outs did not keep pages out of ChatGPT answers, while search opt-outs did.

Some blocks are accidental. 3.8% of sites block every unnamed bot from the whole site, which also shuts out AI crawlers that did not exist when the rule was written.

## What happens when an AI agent cannot read your site?

It answers anyway, using other people’s pages. [Finder and colleagues](https://arxiv.org/abs/2609.34951) ran buyer questions about 1,056 businesses through four AI agents. In about 99% of journeys that hit a dead end on a site, the agent still answered, from whatever it found elsewhere.

When a site was easy for agents to read, 78% of the answer came from the business’s own pages. When it was not, that fell to 58%, and web search filled the gap. Two independent AI judges both rated the answer a clear recommendation 20% of the time for easy-to-read sites, against 11% for the rest. We cover that study in [our guide on agent-readable websites](https://underneath.agency/resources/do-ai-agents-recommend-readable-websites).

This is a different setting from Google-Extended, and the authors work for a company that sells readiness scores. It still shows the general trade-off: an AI that cannot use your content does not stay silent about you. Whether outside coverage makes up for it is covered in [our guide on off-site mentions and readability](https://underneath.agency/resources/off-site-mentions-vs-site-readability-ai-agents).

## What should you do about it?

Make the Google-Extended decision deliberately, with the trade-off written down. Practical steps:

1. Decide what you are protecting. Blocking Google-Extended keeps content out of Gemini training and Gemini citations; it may also cost AI Overview citations.
2. Audit robots.txt crawler by crawler. Separate training crawlers, AI search crawlers and fetchers that act for a user, and check for old rules that block every bot.
3. Do not use noindex or nosnippet to avoid AI Overviews unless you accept losing regular search visibility too.
4. If you block or unblock Google-Extended, track AI Overview and Gemini citations for your key searches before and after the change.
5. Check your firewall and content delivery network as well, since many blocks happen there rather than in robots.txt. Your server logs can also show [which pages AI bots request](https://underneath.agency/resources/ai-bot-server-logs-content-demand).

If you want an outside view of these settings, our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) covers crawler access.

## What does the research not tell us yet?

Whether the block itself causes fewer AI Overview citations is still unproven.

- Both studies are observational snapshots. No one has yet compared the same sites before and after changing the setting.
- The two results disagree, and they differ in date, sample size and method.
- The Grossman study collected data from one location on mobile; our check covered a few hundred AI Overviews on one day.
- Google can change how AI Overviews treat Google-Extended at any time, so findings may date quickly.
- Robots.txt shows what a site asks, not what crawlers do; blocks at the firewall were not measured.

## Frequently asked questions

### Does blocking Google-Extended affect my Google search rankings?

Google says it does not. Google-Extended governs Gemini training and Gemini citations; our study found 10.1% of top sites block it while still allowing Googlebot.

### Can I opt out of AI Overviews without leaving Google search?

Not cleanly. During spring 2026, the available controls were noindex and nosnippet, which also affect regular search results, according to Xu and colleagues.

### Should we unblock Google-Extended?

It depends on what you value more. A large December 2025 study linked the block to fewer AI Overview citations, while our September 2026 check saw 17.8% of cited pages from blocking sites against 16.8% expected.

### How many websites block Google-Extended?

Among top sites with a readable robots.txt, 12.3% blocked it from the whole site in September 2026.

## Sources

- Grossman, Liu, Chen, Smith, Borcea and Chen (2026), [How Generative AI Disrupts Search: An Empirical Study of Google Search, Gemini, and AI Overviews](https://arxiv.org/abs/2604.27790), arXiv:2604.27790.
- Xu, Iqbal and Montgomery (2026), [Measuring Google AI Overviews: Activation, Source Quality, Claim Fidelity, and Publisher Impact](https://arxiv.org/abs/2605.14021), arXiv:2605.14021.
- Finder, Elovic, Shalev and Yosef (2026), [AX is the New AEO](https://arxiv.org/abs/2609.34951), arXiv:2609.34951.
- Underneath (2026), [Which AI crawlers do top websites block? 10,000 sites, 2026](https://underneath.agency/research/ai-crawler-blocking-study)

---

This is the Markdown twin of https://underneath.agency/resources/does-blocking-google-extended-hurt-ai-overviews. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Can restructuring content increase AI citations?"
description: "Possibly, but the evidence is split. One lab test found a 17.3% lift from restructuring alone; a 252,000-trial test found formatting had no effect."
canonical: "https://underneath.agency/resources/does-content-structure-increase-ai-citations"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · Content that gets cited

# Can restructuring existing content, without changing what it says, increase AI citations?

Possibly, but the evidence is split and none of it comes from live tests on real websites. One 2026 lab study found that restructuring pages raised their citation rate from 45.0% to 52.8%, a relative gain of 17.3%. A larger controlled test found formatting changes had no effect. Our own data suggests structure helps an AI use a page more than it helps the page get chosen.

## The short version

1. In a study of 200 articles across six AI engines, [Yu and colleagues](https://arxiv.org/abs/2603.29979) reported that restructuring alone lifted the citation rate from 45.0% to 52.8%.
2. In 252,000 head-to-head trials across six AI models, [Vishwakarma and colleagues](https://arxiv.org/abs/2605.25517) found formatting choices had no impact on which source was cited first.
3. In [our study of pages cited by AI Overviews](https://underneath.agency/research/ai-overview-cited-pages-study), pages with a table were no more likely to be cited, but once cited were 14.7 points more likely to have their wording used.
4. In an analysis of 602 prompts by [Zhang and colleagues](https://arxiv.org/abs/2604.25707), pages in Q&A format had slightly lower influence on answers, a difference of -5.74%.

## Can changing a page’s structure alone raise AI citations?

In one lab study, yes, by about a sixth. [Yu and colleagues](https://arxiv.org/abs/2603.29979) took 200 articles from a standard research dataset and rewrote their structure while keeping the meaning. They then tested original and restructured versions across six AI engines, including ChatGPT, Claude and Perplexity.

The share of queries citing the content rose from 45.0% to 52.8%. In the authors’ words, the six engines showed consistent 17.3% citation improvements. Gains ranged from 14% to 20% across different kinds of engine.

The articles were long, averaging 2,547 words, and came paired with 377 real-world queries. An automated check confirmed the restructured versions kept their meaning. The paper is a preprint, and it does not explain in detail how it queried the live engines. Two engine names it lists, Google SGE and Bing Chat, are older product names.

## Which kinds of structure mattered most in that study?

The big-picture layout mattered most, and visual emphasis least. The authors split structure into three levels and removed each one in turn to see how much the gain fell.

| Level of structure | What it covers | Share of the gain |
|---|---|---|
| Document architecture | Heading hierarchy, navigation and overall flow | 44.9% |
| Information chunking | Paragraphs, lists and tables | 39.7% |
| Visual emphasis | Bold, italics and keyword placement | 15.4% |

An AI judge also scored the restructured pages higher on “influence” (+32.0%) and “click probability” (+31.4%). These are ratings by an AI model, not clicks measured from real users.

## Do other controlled tests find the same thing?

No: the largest controlled test found formatting made no difference. [Vishwakarma and colleagues](https://arxiv.org/abs/2605.25517), at the software company Sprinklr, ran 252,000 trials across six AI models. Each trial showed a model two product pages that differed in exactly one factor.

They report that formatting choices had no impact, suggesting the AI models parse content regardless of visual organization. Dense paragraphs versus organized sections, and grouped versus scattered information, did not change which page was cited first.

What did matter were facts and fit: matching the topic, stating a price, carrying a recent date and appearing first in the list. Those four won across all six models. We cover them in [our guide to why AI picks a competitor’s page](https://underneath.agency/resources/why-ai-cites-competitor-page-first). Note this test used only two pages at a time with brand names removed, unlike a live search.

A 2025 benchmark by [Puerto and colleagues](https://arxiv.org/abs/2506.11097) points the same way. A combined rewrite that included bolding and restructuring gained only in one domain, retail products, on one model.

## Does structure help a page get chosen, or help the AI use it?

The best evidence says it helps use more than choice. In [our study of pages cited by AI Overviews](https://underneath.agency/research/ai-overview-cited-pages-study), we compared 3,096 top-10 pages on the same searches. Tables, headings and other layout features showed no reliable link to being cited.

But structure mattered after the choice was made. Cited pages with an HTML table were 14.7 points more likely to have their citing passages absorbed, meaning the AI Overview’s wording could be traced to the page. Tables showed no association with being cited at all (+0.6). Tables plausibly hold the specific figures an answer repeats.

[Zhang and colleagues](https://arxiv.org/abs/2604.25707) found a similar pattern in a public dataset of 602 prompts across ChatGPT, Google and Perplexity. Pages that most shaped answers averaged 10.59 headings, against 0.85 for pages that shaped them least. They caution this is a snapshot that cannot show whether headings cause the effect. Other traits of the pages engines lean on are described in [our guide to content AI engines favor](https://underneath.agency/resources/what-content-do-ai-engines-prefer).

## Do FAQ sections and question-style headings help?

Not reliably. In the Zhang dataset, pages in Q&A format had a mean influence of 0.0947 against 0.1005 for other pages. The authors conclude the value comes from the evidence inside the page, not from question marks in headings.

Our study found the same. Question-form headings, compared with pages on the same search, showed a gain of +3.5 points that could not be told apart from zero. FAQ headings showed no gain either. We look at this more closely in [our guide to FAQ pages and AI citations](https://underneath.agency/resources/do-faq-pages-help-ai-citations).

## Does structured data count as structure that helps?

The evidence is mixed and mostly against a lift. A 2025 study by [Kumar and Palkhouski](https://arxiv.org/abs/2509.10762) audited 1,100 pages cited for 70 B2B software prompts. It found metadata, clean page markup and structured data most associated with citation. Its authors list a private research company alongside a university. Every page it audited had been cited somewhere, so it compared cited pages with each other, not with pages passed over.

When we compared cited pages with uncited pages on the same search, Organization schema showed no meaningful association with citation (+0.5 points). Only a machine-readable date held up, at +7.9 points, and mainly on informational searches. In a before-and-after test that [our study](https://underneath.agency/research/ai-overview-cited-pages-study) reports, Ahrefs tracked 1,885 pages that added schema against 4,000 controls. Adding schema did not raise citations. The wider set of page features is reviewed in [our guide to on-page signals and citations](https://underneath.agency/resources/on-page-signals-linked-to-ai-citations).

## What should you do about it?

Restructure to make facts easy to find and reuse, but do not expect structure alone to get you chosen.

1. **Fix substance and fit first.** Topic match, prices, specifications and dates beat formatting in the largest controlled test.
2. **Use clear headings and logical sections.** Document layout carried the most weight in the one study that found a structural lift, and it helps human readers too.
3. **Put key figures in tables.** In our data, tables did not help a page get cited, but they went with more of the answer coming from the page.
4. **Don’t convert everything to FAQs.** Q&A packaging alone showed no gain in two separate studies.
5. **Add a machine-readable date.** It was the one page feature that held up in our comparison, mainly for informational content.
6. **Keep ranking in view.** In our study, position in Google’s results mattered more than any page feature.

If you want help prioritizing these edits across your site, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

No one has yet restructured real pages and measured citations on live AI search before and after.

- **The positive result is one preprint.** The 17.3% lift comes from a single study whose method for querying live engines is not described in detail.
- **The negative result is vendor research.** The 252,000-trial study comes from a software company and used two anonymized pages at a time.
- **The observational studies cannot show cause.** Our study and the Zhang analysis compare pages as they are; well-structured pages may also be better in other ways.
- **“Use” is measured by word overlap.** Our absorption measure tracks wording, not whether the answer is faithful to the page.
- **No study measured clicks or sales.** None of these studies tracked whether restructured pages brought more visitors or customers.

## Frequently asked questions

### Do headings and bullet points help content get cited by ChatGPT?

The evidence is split. One lab study found a 17.3% citation lift from restructuring, while a 252,000-trial test found formatting had no impact.

### Should I add tables to my pages for AI search?

Yes, where you have real figures to show. In our study, tables did not raise the chance of citation, but cited pages with tables were 14.7 points more likely to have their wording used.

### Is FAQ formatting good for AI visibility?

Not on its own. Q&A pages had slightly lower influence on AI answers in one dataset, and FAQ headings showed no gain in our comparison.

### Is restructuring cheaper than writing new content for AI search?

It can be, but the payoff is uncertain. The only study showing a structural lift was a lab test, and no live before-and-after test exists yet.

## Sources

- Yu, MuFeng, Ding and Sato (2026), [Structural Feature Engineering for Generative Engine Optimization: How Content Structure Shapes Citation Behavior](https://arxiv.org/abs/2603.29979), arXiv:2603.29979.
- Vishwakarma, Kumar and Jamidar (2026), [What Gets Cited: Competitive GEO in AI Answer Engines](https://arxiv.org/abs/2605.25517), arXiv:2605.25517.
- Zhang, He and Yao (2026), [From Citation Selection to Citation Absorption: A Measurement Framework for Generative Engine Optimization Across AI Search Platforms](https://arxiv.org/abs/2604.25707), arXiv:2604.25707.
- Puerto, Gubri, Green, Oh and Yun (2025), [C-SEO Bench: Does Conversational SEO Work?](https://arxiv.org/abs/2506.11097), arXiv:2506.11097.
- Kumar and Palkhouski (2025), [AI Answer Engine Citation Behavior An Empirical Analysis of the GEO16 Framework](https://arxiv.org/abs/2509.10762), arXiv:2509.10762.
- Underneath (2026), [What pages cited by AI Overviews have in common](https://underneath.agency/research/ai-overview-cited-pages-study)

---

This is the Markdown twin of https://underneath.agency/resources/does-content-structure-increase-ai-citations. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Is it true that GEO increases AI visibility by 40%? | Underneath"
description: "Not as usually stated. The 40% comes from a 2023 lab test in which a page already given to the AI won more of the answer’s text, not more traffic."
canonical: "https://underneath.agency/resources/does-geo-increase-ai-visibility-by-40-percent"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Is it true that GEO increases AI visibility by 40%?

Not in the way it is usually quoted. The figure comes from a 2023 lab experiment. Rewriting a page that was already one of five sources given to an AI raised its share of the answer’s text by up to about 40%. It says nothing about being found, being clicked or winning customers, and later studies found such gains hard to repeat.

## The short version

1. The 40% figure is the best result from Aggarwal and colleagues’ 2023 paper that named generative engine optimization (GEO), measured in a simulated engine built on GPT-3.5.
2. It means a page’s weighted share of the answer’s words rose from 19.3 to 27.2 (about 41%) when quotations were added.
3. A 2026 review of 45 studies by Martinez lists “GEO increases visibility by 40%” as rejected as a general claim.
4. In a wider 2025 benchmark by Puerto and colleagues, popular rewriting tactics gave a reliable gain in only three of 54 tests.

## Where does the 40% figure come from?

From the first academic GEO paper, published in 2023 by researchers at Princeton and IIT Delhi. If the term is new to you, start with [our explainer on what GEO is](https://underneath.agency/resources/what-is-generative-engine-optimization).

[Aggarwal and colleagues](https://arxiv.org/abs/2311.09735) wrote that their methods “can boost visibility by up to 40% in generative engine responses.” Their engine was a simulation. For each question, only the top 5 Google results were fetched, and GPT-3.5 wrote an answer citing them.

The team then rewrote one of those five sources in nine ways, such as adding statistics, quotations or citations, and compared it with the original. The best methods improved the page’s share by 41% on one measure and 28% on another. That 41% is the source of the rounded “40%” quoted in sales decks.

## What exactly went up by 40%?

A page’s share of the words in an AI answer, among five pages already handed to the AI.

[Martinez’s 2026 review](https://arxiv.org/abs/2607.14035), a critical survey of the field, spells it out. The headline comes from quotation addition, which lifted the page’s share from 19.3 to 27.2, or about 41% in relative terms. The share is weighted toward sentences near the top of the answer. The five pages’ shares always add up to 100, so one page can only gain what the others lose.

The review is blunt about what this does not mean. The result does not mean that 40% more readers will click, nor that a page will gain 40% in its chance of being retrieved. It means that a source already in front of the AI received a larger share of the text.

The gains were also uneven. When sources were cited more, the page that started fifth gained 115.1%, while the page that started first lost 30.3%. The method mostly helped pages the search engine had ranked low.

## Did the result hold on a real AI search engine?

Partly, in a limited test on Perplexity that still supplied the pages directly.

The 2023 paper also tried its methods on Perplexity, a live AI search engine. The largest gains there were 22% and 37% on the two measures. But the review notes the test used 200 examples, and the texts were uploaded as files rather than found on the open web. It shows Perplexity can use a better-written page more, not that Perplexity will find it. How an old search trick fared in that test is covered in [our guide to keyword stuffing in AI search](https://underneath.agency/resources/does-keyword-stuffing-work-in-ai-search).

## Have later studies confirmed it?

Only within lab setups; tests closer to real search found much smaller, inconsistent or even negative effects.

[Puerto and colleagues](https://arxiv.org/abs/2506.11097) built a benchmark across six kinds of content, from news to retail products, and four AI models. In their main experiment, out of 54 combinations of tactic and content type, only three gave a reliable gain, and many pushed pages down. Moving a page higher in the list the AI received beat every rewrite. When more competitors used the same tactic, the gains shrank toward zero. We review those tactics one by one in [our guide to statistics, quotes and citations](https://underneath.agency/resources/do-geo-content-tactics-work).

Newer methods learn their own rewriting rules. One of them, AutoGEO from [Wu and colleagues](https://arxiv.org/abs/2510.11438) at Carnegie Mellon, reported an average improvement of 35.99% in GEO measures. That gain, the review notes, was again measured on five documents already retrieved.

The review also describes a 2026 preprint by Kim and colleagues that put search back into the loop, using a large pool of documents. Rewriting only a page’s body text cut its presence in the reranked top 10 by about 16% and its final citations by about 6%. A rewrite can win inside the answer and still lose the race to be found.

## Does GEO raise traffic or sales?

The evidence is very thin. The review rates any link between citation scores and clicks or revenue as “very low” confidence.

One study the review covers examined a website’s logs after some pages were reworked for AI answers. ChatGPT referrals to those pages grew 5.7 times. But untreated pages on the same site had already grown 3.5 times as ChatGPT itself grew. The authors’ estimate of the extra lift was suggestive, not proven.

A separate study of 112 startups by [Sharma](https://arxiv.org/abs/2601.00912), a single-author paper from IIT Patna, scored their websites on GEO features with a simple automated check. The score showed no link with whether ChatGPT or Perplexity recommended them in discovery questions. Referring links from other sites mattered on Perplexity instead.

Our own [AI Overview study](https://underneath.agency/research/ai-overview-cited-pages-study) makes the same distinction. Among 3,096 pages ranking in Google’s top 10, ranking position explained more of which pages were cited than all 14 page features together.

## What should you do about it?

Treat “40% more AI visibility” as a lab result, and ask vendors who quote it what they will measure.

1. Ask which measure a promised gain refers to: being found, being cited, being cited first, or traffic and sales. These are different results.
2. Ask for a baseline and a comparison group. A rise in AI referrals can come from the AI product growing, as the log study showed.
3. Expect results to vary by assistant, question and day, and insist on repeated measurement across several assistants.
4. Keep paying for the basics that the research supports: relevant pages that answer buyer questions, real evidence, and strong search ranking.
5. Be wary of fixed recipes. The review found no technique with a lasting effect on being found across several AI engines. See [the GEO practices research supports](https://underneath.agency/resources/what-geo-practices-does-research-support) instead.

If you want a measurement plan built this way, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The field has strong lab evidence and almost no proof of business outcomes.

- No reviewed study shows a stable, long-term effect of any GEO technique on being found across several AI engines.
- Almost no studies measure clicks, leads or revenue with a proper comparison group.
- Most experiments fix the set of pages in advance, so they skip the step where most pages fail: being retrieved at all.
- AI products change often, so a result from 2023 or 2025 may not hold today.
- Gains measured with one page optimizing may disappear when competitors do the same.

## Frequently asked questions

### Where does the 40% GEO statistic come from?

It comes from Aggarwal and colleagues’ 2023 paper that named GEO. In a simulated engine built on GPT-3.5, adding quotations to one of five sources raised its share of the answer’s text by about 41%.

### Does GEO guarantee more traffic from ChatGPT?

No. No published study shows that, and the 40% result measured share of an answer’s words, not clicks. In the one log study reviewed, reworked pages grew 5.7 times while untreated pages grew 3.5 times, so much of the rise came from ChatGPT’s own growth.

### Is GEO worth investing in if the 40% is overstated?

It can be, if you judge it on measured results rather than the headline. The best-supported work is helping AI systems find and use relevant, evidence-rich pages, measured across several assistants over time.

### Has anyone replicated the 40% GEO result?

Not outside similar lab setups. A 2025 benchmark found reliable gains in only three of 54 tests, and a test that included search found rewrites reduced final citations by about 6%.

## Sources

- Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan and Deshpande (2024), [GEO: Generative Engine Optimization](https://arxiv.org/abs/2311.09735), arXiv:2311.09735.
- Martinez (2026), [Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026)](https://arxiv.org/abs/2607.14035), arXiv:2607.14035.
- Puerto, Gubri, Green, Oh and Yun (2025), [C-SEO Bench: Does Conversational SEO Work?](https://arxiv.org/abs/2506.11097), arXiv:2506.11097.
- Wu, Zhong, Kim and Xiong (2025), [What Generative Search Engines Like and How to Optimize Web Content Cooperatively](https://arxiv.org/abs/2510.11438), arXiv:2510.11438.
- Sharma (2026), [The Discovery Gap: How Product Hunt Startups Vanish in LLM Organic Discovery Queries](https://arxiv.org/abs/2601.00912), arXiv:2601.00912.
- Underneath (2026), [What pages cited by AI Overviews have in common: 3,096 pages](https://underneath.agency/research/ai-overview-cited-pages-study)

---

This is the Markdown twin of https://underneath.agency/resources/does-geo-increase-ai-visibility-by-40-percent. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Does keyword stuffing still work in AI search? | Underneath"
description: "No. In published tests, repeating query keywords did nothing or lowered a page’s share of AI answers; relevance and search ranking mattered more."
canonical: "https://underneath.agency/resources/does-keyword-stuffing-work-in-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Does keyword stuffing still work in AI search?

No. In every published test we could find, stuffing a page with repeated query keywords failed to raise its visibility in AI answers, and it often lowered it. What still counts is whether the page genuinely answers the question, and whether it ranks well enough to be read at all.

## The short version

1. In the original 2023 lab test (Aggarwal and colleagues), keyword stuffing cut a page’s share of the AI answer from 19.3 to 17.7; the best rewrites raised it 41%.
2. On Perplexity, a live AI search engine, the same tactic performed 10% worse than leaving the page alone (same study).
3. In a 2026 shopping test by Bagga and colleagues, AI models told to watch for manipulation flagged keyword-stuffed listings 84.5% to 100% of the time.
4. In our study of ChatGPT’s local picks, a 10.6-point edge for keyword-stuffed business names shrank to 2.4 points with more data, too small to tell from chance.

## Does keyword stuffing help a page appear in AI answers?

No. Controlled tests show it adds nothing or makes a page less visible in AI-written answers.

The first test came from [Aggarwal and colleagues](https://arxiv.org/abs/2311.09735), who coined the term [generative engine optimization (GEO)](https://underneath.agency/resources/what-is-generative-engine-optimization) in 2023. They built a simulated AI search engine. Only the top 5 Google results for each question were fetched, and GPT-3.5 wrote a cited answer from them.

The researchers then rewrote one source at a time in nine ways and measured how much of the answer drew on it. Keyword stuffing meant adding more keywords from the question, as in classic search optimization. Its position-weighted share of the answer fell from a baseline of 19.3 to 17.7, as [Martinez’s 2026 review](https://arxiv.org/abs/2607.14035) of the results sets out. The best rewrites, such as adding quotations or statistics, [improved the same measure by 41%](https://underneath.agency/resources/does-geo-increase-ai-visibility-by-40-percent).

The team repeated a smaller test on Perplexity. There, keyword stuffing performed 10% worse than the baseline.

## Has anyone repeated the test since?

Yes. Later tests on other AI models found the same weak and inconsistent result.

[Wu and colleagues at Carnegie Mellon](https://arxiv.org/abs/2510.11438) re-ran the original rewrites in 2025 on an engine built on Google’s Gemini, across three sets of questions. Keyword stuffing lost ground on the original test questions and gained on open research questions. That is not a lift anyone could plan around.

| Question set (Gemini engine) | Untouched pages | Keyword-stuffed pages |
|---|---|---|
| Original GEO test questions | 19.44 | 18.05 |
| Open research questions | 20.18 | 22.68 |
| Shopping questions | 18.32 | 19.17 |

Each number is a page’s position-weighted share of the answer, out of 100. The authors’ own learned rewriting method beat every simple tactic in the table.

A wider benchmark by [Puerto and colleagues](https://arxiv.org/abs/2506.11097) tested ten popular rewriting tactics, though not keyword stuffing itself, across six kinds of content and four AI models. In their main experiment, out of 54 combinations of tactic and content type, only three produced a reliable gain, and many rewrites pushed pages down. Moving a page higher in the list the AI received beat any rewrite.

## Why doesn’t repetition work when AI writes the answer?

AI assistants read for meaning and write their own searches, so repeating a phrase gives them nothing new.

Classic search engines grew up matching the words on a page to the words in a query. The 2023 authors argued that AI engines are not limited to that kind of matching, because the model reads the whole page and the whole question.

Our [hidden searches study](https://underneath.agency/research/ai-hidden-searches-study) shows how far the assistants move from the buyer’s wording. On 80 buyer questions, ChatGPT ran a mean of 3.7 searches per answer before replying. None of the 509 searches repeated the user’s question word for word. A page tuned to one exact phrase is tuned for a search the assistant may never run.

Once pages are retrieved, the AI picks what it can use. Martinez’s review of 45 studies rates the support for keyword stuffing as “null or negative” across multiple benchmarks, and its advice is simply to avoid it. The best-supported levers in that review are relevance to the question and a page’s position among the results the AI reads. We gather the rest in [our guide to research-backed GEO practices](https://underneath.agency/resources/what-geo-practices-does-research-support).

## Do keywords still matter at all?

Yes, in the plain sense: a page must use the words and cover the topic the buyer is asking about.

In a 2026 controlled test by [Vishwakarma and colleagues](https://arxiv.org/abs/2605.25517) at Sprinklr, a software vendor, six AI models repeatedly chose between two versions of the same page. There were 252,000 trials in all. Being on topic was one of four factors that decided citation in all six models. A page missing key terms that a competing page used also lost. That factor was among 11 of the 18 factors (61%) that mattered in at least four models. Formatting changes alone made no difference.

The shopping test by [Bagga and colleagues](https://arxiv.org/abs/2511.20867) points the same way. When software searched for the best instructions for rewriting product listings, the winners converged on one playbook. It included working in relevant keywords and synonyms, and it explicitly warned against repeated keywords and unnatural phrasing.

So the line runs between coverage and repetition. Naming the product, the problem and the terms buyers use helps an AI match a page to a question. Saying the same phrase again and again does not. Whether clearer, plainer prose helps instead is covered in [our guide on readable writing and AI answers](https://underneath.agency/resources/does-readable-writing-help-ai-visibility).

## Does stuffing a business name with keywords help in ChatGPT?

Not reliably. Our own data showed an early edge that disappeared once we collected more answers.

Many local businesses add service and city words to their Google Maps name, such as “Business Name - Emergency Plumber - City”. ChatGPT copies those names faithfully: in our [local recommendations study](https://underneath.agency/research/chatgpt-local-recommendations-study), all 115 keyword-stuffed names were reproduced verbatim on 26 September 2026.

Copying a name is not the same as favoring it. Our [local picks study](https://underneath.agency/research/chatgpt-local-picks-google-profile-study) compared 898 ChatGPT answers with the Google Maps top 20 for 120 local searches in four countries. On the first day, stuffed names showed a 10.6-point advantage. It held up in only two of eight sets of answers, and after adjusting for rank and reviews it shrank to 2.4 points, within the range of chance.

Having more reviews than local rivals, by contrast, helped in all eight sets. The study also notes that Google’s guidelines do not allow keyword-stuffed business names, and a profile that breaks them can be suspended.

## Is there a risk in stuffing, beyond wasted effort?

Yes. AI systems can read keyword-stuffed text as a sign of manipulation and rank it lower.

In the shopping test, each AI model ranking products received one short instruction to watch for manipulative descriptions. The researchers used a fixed set of 200 shopping questions. Listings rewritten to repeat category keywords were flagged as questionable between 84.5% and 100% of the time across five AI models. They also fell in the rankings on every model.

That warning instruction was the researchers’ own, not a known feature of ChatGPT or Google. Still, it shows how cheaply an AI can spot the tactic. The strongest models in the test, GPT-5 and Claude, gave manipulative rewrites little lift whether or not they flagged them.

## What should you do about it?

Stop paying for keyword density and put the budget into pages that answer buyer questions completely.

1. Remove keyword-density targets from content briefs and vendor scopes. No published test shows they help in AI answers.
2. Make sure each important page names the product, the problem and the words buyers actually use, naturally and in context.
3. Add what AI answers can reuse: prices, specifications, comparisons, dates and evidence for claims. These won in the controlled tests above.
4. Keep investing in search ranking. In the benchmark tests, a page’s position among the results the AI read mattered more than any rewrite.
5. Keep your Google Business Profile name as your real business name, and put the effort into earning reviews instead.
6. Measure AI visibility directly, across several assistants and repeated runs, rather than assuming search tactics carry over.

If you want help setting up that measurement, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The evidence is consistent but narrow, and most of it comes from lab setups rather than live search products.

- Most tests hand the AI a fixed set of pages. None shows what keyword stuffing does to a page’s chance of being found in the first place.
- Only the 2023 study checked a live engine, Perplexity, and AI products have changed many times since.
- None of these tests measured what stuffing does to Google rankings, which still feed many AI answers. The 2023 authors say so themselves.
- The manipulation-flagging result depends on a warning the researchers added. Live assistants may be stricter or more lenient.
- Our local finding covers ChatGPT, four countries and two days in September 2026.

## Frequently asked questions

### Is keyword stuffing bad for ChatGPT visibility?

There is no evidence that it helps, and some that it hurts. In controlled tests, stuffed pages lost share of the answer, and AI models told to look for manipulation flagged stuffed listings most of the time.

### Should I still do keyword research for AI search?

Yes, to learn the words and questions buyers use, not to repeat them. In the 252,000-trial test, pages missing a buyer’s key terms lost to pages that had them, while formatting tricks made no difference.

### Does keyword stuffing work in Google AI Overviews?

No published study has tested it on AI Overviews directly. The closest evidence, from simulated engines fed with Google results and from an engine built on Gemini, shows no dependable gain.

### What works better than keyword stuffing for AI answers?

Relevance, ranking and reusable evidence. In the 2023 lab test, adding quotations or statistics raised a page’s share of the answer by up to 41%. That gain applied only to pages the engine had already been given.

## Sources

- Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan and Deshpande (2024), [GEO: Generative Engine Optimization](https://arxiv.org/abs/2311.09735), arXiv:2311.09735.
- Martinez (2026), [Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026)](https://arxiv.org/abs/2607.14035), arXiv:2607.14035.
- Wu, Zhong, Kim and Xiong (2025), [What Generative Search Engines Like and How to Optimize Web Content Cooperatively](https://arxiv.org/abs/2510.11438), arXiv:2510.11438.
- Puerto, Gubri, Green, Oh and Yun (2025), [C-SEO Bench: Does Conversational SEO Work?](https://arxiv.org/abs/2506.11097), arXiv:2506.11097.
- Vishwakarma, Kumar and Jamidar (2026), [What Gets Cited: Competitive GEO in AI Answer Engines](https://arxiv.org/abs/2605.25517), arXiv:2605.25517.
- Bagga, Farias, Korkotashvili, Peng and Wu (2026), [E-GEO: A Testbed for Generative Engine Optimization in E-Commerce](https://arxiv.org/abs/2511.20867), arXiv:2511.20867.
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [ChatGPT local recommendations: stable details, shifting lists](https://underneath.agency/research/chatgpt-local-recommendations-study)
- Underneath (2026), [Which Google Maps businesses does ChatGPT recommend?](https://underneath.agency/research/chatgpt-local-picks-google-profile-study)

---

This is the Markdown twin of https://underneath.agency/resources/does-keyword-stuffing-work-in-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Does list order change what an AI assistant recommends?"
description: "Yes, in controlled tests: options shown first are picked more often. One hotel study valued first place at $11.7 a night; the size varies widely by AI model."
canonical: "https://underneath.agency/resources/does-list-order-change-ai-recommendations"
published: 2026-10-10
updated: 2026-10-10
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Does list order change what an AI assistant recommends?

Yes: in controlled tests, AI assistants pick options shown earlier more often, even when nothing else about them changes. The average effect is modest, but it varies sharply between AI models, and in real AI search the order is set by the search results the assistant reads. Where your business sits in that list is worth watching.

## The short version

1. In a 2026 test of 12 AI models choosing among five made-up hotels, being listed first was worth $11.7 per night, about a tenth of a full step up in guest rating.
2. Most of those models were nearly unaffected by order, but one Google model favored the first slot by about 26 percentage points.
3. In a 2026 test of planted fake product pages, a fake page in the first search result fooled the two weakest models in 27% of cases; in slots two to ten, only 1 to 4%.
4. In a skincare test, list position explained 6.5% of how three AI models ranked products, against 82.4% for rating, price and reviews.
5. In our study of 481 US AI Overviews, the first Google result was cited 49.5% of the time and the ninth 15.5%.

## Does list position really change AI recommendations?

Yes, in experiments that change only the order. The effect on average is small but real.

The cleanest test is by [Baig and colleagues](https://arxiv.org/abs/2606.16344). Twelve AI models, from OpenAI, Google, Anthropic and four open models, each chose one of five made-up hotels. The researchers shuffled rating, price, reviews and list position at random, running 3,024 main choice sets per model. Because only the order changed between otherwise similar cards, any position effect is caused by the order itself.

Compared with the first slot, a hotel was recommended 2.1 percentage points less often in slot 2 and 3.7 points less often in slot 5. The authors call list position “a content-free artifact” that “shifts recommendations causally”.

An earlier study by [Pfrommer and colleagues](https://arxiv.org/abs/2406.03589) at UC Berkeley found the same direction in simulated product search. Even when told to mention the best products first, all the AI models they tested preferred products placed earlier in what they read.

## How much is first place worth?

In the hotel study, about $11.7 per night, against $126.4 for a full step up in guest rating.

Converting effects into dollars makes them comparable. Being listed first was worth $11.7 per night in the pooled results, while moving from a 3.9-star to a 4.7-star rating was worth $126.4. A top rating raised the chance of being chosen by 31.6 percentage points, and a high price cut it by 30.0. The other hotel factors are covered in [our guide to how AI picks hotels](https://underneath.agency/resources/why-ai-recommends-one-hotel-over-another).

So order matters much less than reputation and price. But it is free: the first-listed hotel gained without changing anything about itself. A separate skincare study by [Chu and Hou](https://arxiv.org/abs/2606.17443) agrees on the scale. Across 14,395 trials with three AI models, position explained 6.5% of the ranking, product details 82.4% and brand name only 1.2%.

## Do all AI models care equally about order?

No: sensitivity to order varies a lot by model. An average can hide one model that cares a great deal.

In the hotel study, most models were nearly position-neutral, but gemini-2.0-flash gave the first slot an advantage of about 26 percentage points, roughly ten times the average. A 2026 test by [Vishwakarma and colleagues](https://arxiv.org/abs/2605.25517) at Sprinklr, a marketing software company, ran 252,000 simulated trials. It pitted two sources against each other in six AI models. Being listed second rather than first was one of four factors that mattered in all six models.

Pfrommer and colleagues also found that the models differed in what drove their rankings: some leaned on product names they already knew, others on order. The research does not let you predict which way a given assistant will lean today.

## Where does the order come from in real AI search?

From the search results the assistant retrieves, so search ranking sets the starting order. That is where most of the commercial stakes sit.

The fake-page study by [Luo and Chen](https://arxiv.org/abs/2606.13610) shows how strong the top slot is. They rewrote real search results so that a page promoted a fake product; the position test used six open models on digital products. Placed first, it fooled the two most vulnerable open models in 27% of cases; placed second to tenth, it fooled them in only 1 to 4%. In their words, “the first page the model reads dominates the recommendation”.

Real data point the same way. In [our AI Overview study](https://underneath.agency/research/ai-overview-citations-study), the first organic result was cited in 49.5% of AI Overviews, the third in 33.5% and the ninth in 15.5%. Being high in the list an AI reads is linked to being used, though our data cannot show that the position itself causes it. That is one reason [classic SEO still matters for AI search](https://underneath.agency/resources/is-seo-still-important-for-ai-search).

## Will the AI tell you that position influenced its choice?

Almost never. Assistants act on order without saying so.

In the hotel study, list position carried 4.1% of the weight behind the models’ choices. Yet it was mentioned in at most 0.7% of the reasons they gave. Their stated reasons broadly tracked their choices for rating and price, but not for order or review volume.

This matters for monitoring. Reading an assistant’s explanation will not reveal whether your business won or lost on order rather than merit. We look closer at this gap in [our guide to trusting AI explanations](https://underneath.agency/resources/ai-explanations-for-recommendations).

## Does a brand’s position in AI answers stay the same?

No: positions in real AI answers move a lot from one run to the next. One snapshot is not enough.

In [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), five runs of the same 20 questions minutes apart, 86.4% of the brands ChatGPT named in every run moved position at least once. ChatGPT’s first pick changed at least once for 80.0% of questions. Position bias may favor whichever brand happens to be retrieved first, and that changes between runs.

## What should you do about it?

Track where you appear, not just whether you appear, and work on what sets the order. In practice:

1. Measure your position in AI answers over repeated runs, including how often you are named first.
2. Treat search ranking as part of AI visibility; the pages an assistant reads first carry the most weight.
3. Fix the basics that outweigh order in every study: ratings, clear prices and complete product details.
4. On marketplaces and comparison sites that feed AI assistants, check how listings are ordered and whether you can influence it fairly.
5. Do not trust an assistant’s own explanation as proof that order played no role.

If you want help tracking position across engines, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The research proves order effects in controlled tests, but not their size in live AI search. The gaps:

- The cleanest studies use made-up hotels or products and fixed lists; live assistants choose their own sources.
- Effects differ widely by model and by version, and models change often.
- No study we reviewed measures how list position in AI answers affects real bookings or sales.
- Two of the tests come from authors at companies in this market, and several are simulations rather than live audits.
- How much position matters for well-known brands, which models may already know, is not well measured.

## Frequently asked questions

### Do AI assistants favor the first option they see?

On average, slightly. In a 12-model hotel test, the first slot was worth $11.7 per night, but one model showed an advantage of about 26 percentage points.

### Can a business pay to be listed first in AI answers?

The studies here do not cover paid placement. They show that position in the material an assistant reads can shift its choice, which is why that order matters commercially.

### Is ranking first on Google enough to be recommended by AI?

No, but it helps. In our study the top Google result was cited in 49.5% of AI Overviews, so more than half the time it was not.

### Why does my brand’s position change every time I ask ChatGPT?

Answers vary from run to run. In our test, ChatGPT’s first pick changed at least once for 80.0% of questions across five runs.

## Sources

- Baig and colleagues (2026), [Whose hotel does the AI recommend? An algorithm audit of reputation signals in LLM-assisted hotel selection](https://arxiv.org/abs/2606.16344), an algorithm audit of reputation signals in AI-assisted hotel selection, arXiv:2606.16344.
- Pfrommer and colleagues (2024), [Ranking Manipulation for Conversational Search Engines](https://arxiv.org/abs/2406.03589), arXiv:2406.03589.
- Luo and Chen (2026), [One Polluted Page Is Enough: Evaluating Web Content Pollution in LLM Recommenders](https://arxiv.org/abs/2606.13610), on web content pollution in AI recommenders, arXiv:2606.13610.
- Chu and Hou (2026), [Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems](https://arxiv.org/abs/2606.17443), on brand bias and manipulation in AI recommendation systems, arXiv:2606.17443.
- Vishwakarma and colleagues (2026), [What Gets Cited: Competitive GEO in AI Answer Engines](https://arxiv.org/abs/2605.25517), arXiv:2605.25517.
- Underneath (2026), [Do AI Overviews cite the pages that rank?](https://underneath.agency/research/ai-overview-citations-study)
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)

---

This is the Markdown twin of https://underneath.agency/resources/does-list-order-change-ai-recommendations. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Does question phrasing change which sources AI search cites?"
description: "Yes. In a Gemini hotel test, experience-style questions drew 55.9% of citations from non-booking sites, against 30.8% for booking-style ones."
canonical: "https://underneath.agency/resources/does-question-phrasing-change-ai-sources"
published: 2026-10-10
updated: 2026-10-10
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Does the way customers phrase a question change which sources AI search cites?

Yes: how a buyer phrases a question changes which sources AI search reaches for. In a 2026 test of Google’s Gemini on Tokyo hotel questions, experience-style questions drew 55.9% of their citations from non-booking sources. Booking-style questions on the same need drew 30.8%. Wording also changes whether Google shows an AI answer at all, and which forum threads it cites.

## The short version

1. In a test of 156 hotel questions to Gemini, [Zhu and Chang](https://arxiv.org/abs/2603.20062) found a gap of 25.1 percentage points in citations to non-booking sources between experience-style and booking-style questions.
2. The shift held in all four need categories tested; for business travel, non-booking sources rose from 40.1% to 68.5% of citations.
3. Even cosmetic edits such as “what is” versus “what’s” changed the sources in Google’s AI Overviews more than rerunning the query did, according to [Grossman and colleagues](https://arxiv.org/abs/2604.27790).
4. In [our AI Overview study](https://underneath.agency/research/ai-overviews-frequency-study), the same topic got an AI Overview 59.4% of the time as a short keyword and 93.8% as a natural question.
5. In [our study of hidden searches](https://underneath.agency/research/ai-hidden-searches-study), none of 509 searches the assistants ran repeated the user’s question word for word.

## Does the phrasing of a question change which sources AI search cites?

Yes, sharply, in a controlled 2026 test of hotel questions. [Zhu and Chang](https://arxiv.org/abs/2603.20062) asked Gemini 2.5 Flash, with Google Search turned on, 156 questions about Tokyo hotels in March 2026. They collected 1,357 citations.

Each question came in a matched pair. One was booking-style, such as “Cheap hotel in Shinjuku”. The other described the same need as an experience, such as “Good value hotel with local charm in Shinjuku”. Only the framing changed, not the area, need or language.

The booking-style questions leaned on booking sites such as Booking.com, Expedia and Japanese equivalents. In the authors’ words, non-OTA sources account for 55.9% of experiential citations (419 of 750) but only 30.8% of transactional citations (187 of 607). OTA stands for online travel agency, the booking sites. A supplementary set of 20 experience-style questions that booking-site reviews could answer still drew 54.0% non-booking citations.

## Did the shift hold across different kinds of buyer need?

Yes: all four need categories showed it, with gaps between 19.0 and 28.4 points.

| Buyer need (Gemini, Tokyo hotels) | Booking-style question | Experience-style question |
|---|---|---|
| Budget | 15.5% | 42.5% |
| Rating and quality | 22.1% | 46.7% |
| Business travel | 40.1% | 68.5% |
| Convenience and access | 40.1% | 59.2% |

Share of citations from sources other than booking sites. Budget questions were the most locked to booking sites, which fits their price-comparison strength. Business questions showed the largest gap, driven by coworking review sites, hotel websites describing workspaces, and business travel content.

Language mattered too. For Japanese-language experience questions, 62.1% of citations came from non-booking sources. The authors link this to a richer Japanese-language web of travel agencies, coworking sites and hotel pages.

## Who gains when the question changes?

Mostly travel blogs and editorial sites, not the businesses’ own websites. In English, the non-booking citations were dominated by travel blogs (22.9%) and editorial curation sites (22.5%).

Hotel websites themselves did not gain from the shift in framing. In English, hotel-direct citation rates were 8.0% for booking-style questions and 8.3% for experience-style ones. What separated cited hotels was the depth of their content. In an exploratory check of 14 hotel websites, cited hotels averaged 8.6/15 on a content-depth score, against 3.4/15 for hotels that were not cited.

One example stands out. A small independent hotel with a 33-question FAQ and a 13-attraction sightseeing guide was cited directly. A design hotel with good ratings but brief pages was named, yet Gemini took its information from booking and editorial sites instead. We tell that story in [our guide to deep content without technical SEO](https://underneath.agency/resources/deep-content-without-technical-seo-gemini).

## Does this happen outside hotels?

The direction appears elsewhere, though the evidence is thinner. In a 2025 study, [Chen and colleagues](https://arxiv.org/abs/2509.08919) grouped consumer questions by intent. They report that brand-owned content rose in prominence for purchase-ready questions such as “Buy iPhone 15 online”. Comparison questions such as “Garmin vs Apple Watch” leaned on reviews and publishers. They give these as charts, not exact shares.

Our own data shows wording changing which kinds of sources Google’s AI cites. In [our Reddit study](https://underneath.agency/research/ai-reddit-citations-study), question wordings cited Reddit in 35.6% of their AI Overviews, against 21.1% for the plain keywords that day. They rarely cited the same thread, though.

## Do small wording changes matter too?

Yes, for Google’s AI Overviews, even trivial edits shift the sources. [Grossman and colleagues](https://arxiv.org/abs/2604.27790) made cosmetic edits to 200 queries, such as “what is” versus “what’s” or adding a question mark. They then compared the sources Google returned with those from simply rerunning the original query.

For AI Overviews, agreement between the sources fell by 28.99% compared with a plain rerun. Google’s regular results fell by 13.95% and Gemini by 3.85%. In other words, Google’s AI summary was the most sensitive of the three to wording that changes nothing about meaning.

Chen and colleagues found the opposite pattern for AI chat engines on rephrased ranking questions. ChatGPT, Perplexity and Gemini were steadier across paraphrases than Google’s regular results. The engines do not all react to wording the same way.

## Does phrasing change whether an AI answer appears at all?

Yes, strongly: questions trigger far more AI Overviews than short keywords. [Xu and colleagues](https://arxiv.org/abs/2605.14021) studied 55,393 trending queries. Question-form queries trigger AIOs at 64.7% versus 9.5% for non-question queries.

[Our AI Overview study](https://underneath.agency/research/ai-overviews-frequency-study) tested the same 96 commercial topics three ways. As written, 59.4% showed an AI Overview; as a long non-question form, 87.5%; as a natural question, 93.8%. Most of the lift came from longer, more specific wording, not the question mark itself. The full pattern is in [our guide to which searches trigger AI Overviews](https://underneath.agency/resources/which-searches-trigger-ai-overviews).

## What does the AI actually search for?

Its own rewritten searches, not your customer’s words. In [our hidden searches study](https://underneath.agency/research/ai-hidden-searches-study), none of the 509 searches repeated the user’s question word for word. ChatGPT ran a mean of 3.7 searches per answer.

What it searched for shaped what it cited. ChatGPT looked for reviews in 46.2% of answers and prices in 23.8%. When a search named a source such as NerdWallet, the answer cited that source 44.0% of the time, against 8.1% in similar answers that did not name it.

Phrasing also moves the brands recommended. In [our prompt phrasing study](https://underneath.agency/research/ai-prompt-phrasing-study), asking the identical question again kept the same first brand 68.0% of the time, but adding “on a tight budget” kept it only 15.3%. Loaded wording raises a related question, covered in [our guide on AI answers and leading questions](https://underneath.agency/resources/do-ai-search-engines-agree-with-leading-questions).

## What should you do about it?

Plan content around the different ways buyers describe the same need, not one keyword.

1. **Collect real phrasings.** Ask sales and support teams how customers describe the problem: by price, by experience, by situation. Each phrasing can pull different sources.
2. **Test both framings.** Run booking-style and experience-style versions of your key questions in the AI engines your buyers use. Note which sources appear for each.
3. **Go deep on your own pages.** In the hotel study, depth separated cited hotel websites from uncited ones: real FAQs, area guides and specific details, not thin feature pages.
4. **Earn coverage where experience questions lead.** If experience-style questions cite blogs and editorial sites in your category, those are the places to be reviewed and described accurately.
5. **Measure across wordings and repeat.** A single phrasing on a single day is a snapshot. Small edits changed Google’s AI sources more than reruns did.

For a structured way to map and track these phrasings, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The central finding rests on one engine, one city and one industry.

- **One engine and one moment.** The hotel study used Gemini 2.5 Flash with Google Search, in March 2026, mostly with one run per question. The authors note other engines search different indexes and may behave differently.
- **One market.** Tokyo has an unusually rich web of non-booking content, in two languages. Cities or industries with thinner content may show a smaller gap.
- **Company research.** The hotel study’s authors work at a private AI company, and the paper is a preprint.
- **Correlation, not cause.** The content-depth check covered only 14 hotels, and no study here changed a website and measured the result.
- **Discovery, not sales.** None of these studies measured whether the cited source changed where people bought.

## Frequently asked questions

### Do AI search engines cite different sources for “best” and “cheap” questions?

Often yes. In the hotel study, budget questions drew 15.5% of citations from non-booking sources when phrased for booking, and 42.5% when phrased as an experience.

### Should I write content for questions or for keywords?

Write for the needs buyers describe, in their words. Question-style searches trigger more AI Overviews, but our study did not show that question-shaped content itself earns more citations.

### Does ChatGPT search for the exact question my customer typed?

No. In our study, none of 509 searches by ChatGPT, Gemini and Claude repeated the user’s question word for word; the assistants wrote their own searches.

### Why do AI answers about my category change from one search to the next?

Partly because wording shifts sources. In one study, cosmetic edits cut the agreement between AI Overview source lists by 28.99% compared with a plain rerun.

## Sources

- Zhu and Chang (2026), [The End of Rented Discovery: How AI Search Redistributes Power Between Hotels and Intermediaries](https://arxiv.org/abs/2603.20062), arXiv:2603.20062.
- Grossman, Liu, Chen, Smith, Borcea and Chen (2026), [How Generative AI Disrupts Search: An Empirical Study of Google Search, Gemini, and AI Overviews](https://arxiv.org/abs/2604.27790), arXiv:2604.27790.
- Xu, Iqbal and Montgomery (2026), [Measuring Google AI Overviews: Activation, Source Quality, Claim Fidelity, and Publisher Impact](https://arxiv.org/abs/2605.14021), arXiv:2605.14021.
- Chen, Wang, Chen and Koudas (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919), arXiv:2509.08919.
- Underneath (2026), [When does Google show an AI Overview?](https://underneath.agency/research/ai-overviews-frequency-study)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [When does a Reddit thread become evidence in Google’s AI?](https://underneath.agency/research/ai-reddit-citations-study)
- Underneath (2026), [Does rewording a question change AI brand recommendations?](https://underneath.agency/research/ai-prompt-phrasing-study)

---

This is the Markdown twin of https://underneath.agency/resources/does-question-phrasing-change-ai-sources. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "If we rank on Google, will we show up in AI Overviews?"
description: "Not reliably. In our test only 28.7% of AI Overview citations were page-one results, though the top result was cited in 49.5% of AI Overviews."
canonical: "https://underneath.agency/resources/does-ranking-on-google-get-you-into-ai-overviews"
published: 2026-10-10
updated: 2026-10-10
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# If we rank on Google, will we show up in AI Overviews?

Ranking helps, but it does not guarantee a citation. In our September 2026 test, the first organic result was cited in 49.5% of AI Overviews, yet only 28.7% of all AI Overview citations were page-one results. Independent studies agree that AI Overviews, the AI summary at the top of Google’s results, draw on a wider pool than Google’s first page.

## The short version

1. Only 28.7% of 4,051 AI Overview citations were page-one organic results for the same search, in our study of 481 US AI Overviews in September 2026.
2. Position still matters: the first organic result was cited in 49.5% of AI Overviews and the ninth in 15.5%, in the same study.
3. Across 7,562 AI Overviews collected over 40 days in 2026, 29.8% of cited domains did not appear anywhere on Google’s first page (Xu and colleagues).
4. In a US and German audit summarized by Martinez, 53% of domains cited by AI Overviews were not in the organic top 10, and 27% were not in the top 100.
5. AI Overview sources move more than rankings: the same pages appeared across two months 18% of the time for AI Overviews, against 45% for organic results.

## How often do AI Overviews cite pages that rank on page one?

Less than a third of citations come from page-one results, but most AI Overviews cite at least one. In our [AI Overview citation study](https://underneath.agency/research/ai-overview-citations-study), we collected every source cited in 481 US AI Overviews on 26 September 2026, 4,051 citations in all. Only 28.7% were page-one organic results for the same search.

Most AI Overviews still drew on page one somewhere. 78.8% cited at least one page-one result, while 21.2% cited none. Counting whole websites rather than single pages raises the match to 43.0%, because Google sometimes cites a different page from a site that ranks.

[Xu and colleagues](https://arxiv.org/abs/2605.14021) found the same pattern at a larger scale. They issued 55,393 trending US queries between March 13 and April 21, 2026. Averaged across 7,562 AI Overviews, only 31.5% of cited domains were among the top five results, 48.0% among the top ten and 70.2% anywhere on the first page.

| Measure | Result | Study |
|---|---|---|
| Citations that are page-one pages | 28.7% | Underneath, September 2026 |
| Cited domains in Google’s top five | 31.5% | Xu et al., spring 2026 |
| Cited domains in Google’s top ten | 48.0% | Xu et al., spring 2026 |
| Cited domains anywhere on page one | 70.2% | Xu et al., spring 2026 |
| Cited domains outside the top ten | 53% | Kirsten et al., via Martinez |

The figures differ partly because some count pages and others count whole domains. In every study, a large share of what AI Overviews cite is not on page one.

## Does a higher ranking raise your chances?

Yes, steadily, with no sudden drop after the top three. In our study, the first organic result was cited in 49.5% of AI Overviews, the third in 33.5%, the sixth in 23.2% and the ninth in 15.5%. A first-place result was about three times as likely to be cited as a ninth-place one.

But position explains only a small part of the picture. In our [study of what cited pages have in common](https://underneath.agency/research/ai-overview-cited-pages-study), we compared 3,096 top-10 pages on the same searches. Position and 14 page features together left 91.4% of the variation in which pages were cited unexplained.

Rankings also carry over less well to Google’s other AI surface. Our [comparison of AI Mode and AI Overviews](https://underneath.agency/research/ai-mode-vs-ai-overviews-study) found that, for pages in positions 1 to 3, the AI Overview cited 42.3% and AI Mode 26.0%. Why ranking work still earns its place is covered in [our guide on SEO in AI search](https://underneath.agency/resources/is-seo-still-important-for-ai-search).

## Where do the other citations come from?

Largely from sites that have no page-one result at all, plus video and Google’s own pages. In our study, 38.5% of all AI Overview citations came from sites with no page-one organic result for the search. YouTube was cited in 64.2% of AI Overviews, and most of its citations were not page-one results.

The independent audits point the same way. Martinez’s [critical survey of the field](https://arxiv.org/abs/2607.14035) reports an audit by Kirsten and colleagues in the United States and Germany: 53% of domains cited by AI Overviews were not in the organic top ten, and 27% were absent from the top 100.

The off-page sources are not weaker. Xu and colleagues found AI Overview sources scored as more credible than the first page as a whole, though not than its top few results. They also cited far fewer social platforms: 14.2% of AI Overview references came from sites such as YouTube, Facebook and Reddit, against 49.9% of first-page results.

## Is an AI Overview citation stable once you have it?

No. AI Overview sources change faster than organic rankings do. In the audit by Kirsten and colleagues, as summarized by Martinez, page overlap across two months was 18% for AI Overviews, against 45% for organic Google results.

Small wording changes matter too. [Grossman and colleagues](https://arxiv.org/abs/2604.27790) changed queries in minor ways, such as “what is” versus “what’s”. AI Overview sources for the edited query matched the original 28.99% less well than two runs of the same query did, against a 13.95% drop for regular results.

AI Overviews also [appear only on some searches](https://underneath.agency/resources/which-searches-trigger-ai-overviews). Xu and colleagues saw one on 13.7% of trending queries overall, but on 64.7% of searches phrased as a question. A strong ranking cannot earn a citation on a search that shows no AI Overview.

## What should you do about it?

Keep investing in rankings, but track AI Overview citations as a separate result. Steps that follow from the evidence:

1. Measure citation directly. A rank report cannot tell you whether you are cited, so check the AI Overview itself for the searches that matter to you. Whether a citation brings visits is a separate question, covered in [our guide on citations and lost clicks](https://underneath.agency/resources/ai-overview-citations-vs-lost-clicks).
2. Aim for the top positions. The citation rate falls steadily down the page, so moving from ninth to first matters.
3. Cover the questions behind the search. Many cited pages answer a narrower part of the question rather than ranking for the search as typed.
4. Look at video. YouTube appeared in most AI Overviews we studied, usually from outside the organic results.
5. Track over weeks, not once. A citation seen today may be gone next month, and small wording changes shift the sources.

If you want this measured for your own searches, our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) starts with that audit.

## What does the research not tell us yet?

The studies describe the gap well, but they cannot yet say how to close it.

- All of them are observational. None shows that ranking higher causes a citation, or that changing a page wins one.
- They count pages and domains differently, so the headline figures are not directly comparable.
- The two-month stability figure comes from one audit, reported second-hand in a survey.
- Most data comes from US searches, collected at one location over days or weeks, not across many countries or user profiles.
- A citation is not influence. None of these studies measures how much of the answer each cited page actually supplied.

## Frequently asked questions

### Does ranking number one on Google guarantee an AI Overview citation?

No. In our study the first organic result was cited in 49.5% of AI Overviews, so about half the time it was not.

### What share of AI Overview sources come from Google’s first page?

About three in ten citations by page. In our study 28.7% of citations were page-one pages, and Xu and colleagues found 29.8% of cited domains were not on page one at all.

### Can a page that does not rank still be cited in an AI Overview?

Yes. In our study 38.5% of AI Overview citations came from sites with no page-one result for the search.

### Is SEO still worth doing for AI Overviews?

Yes, but it is not enough on its own. Higher positions are cited more often, yet most citations come from elsewhere, so citations need their own tracking.

## Sources

- Xu, Iqbal and Montgomery (2026), [Measuring Google AI Overviews: Activation, Source Quality, Claim Fidelity, and Publisher Impact](https://arxiv.org/abs/2605.14021), arXiv:2605.14021.
- Martinez (2026), [Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026)](https://arxiv.org/abs/2607.14035), arXiv:2607.14035.
- Grossman, Liu, Chen, Smith, Borcea and Chen (2026), [How Generative AI Disrupts Search: An Empirical Study of Google Search, Gemini, and AI Overviews](https://arxiv.org/abs/2604.27790), arXiv:2604.27790.
- Underneath (2026), [Do AI Overviews cite the pages that rank? 4,051 citations](https://underneath.agency/research/ai-overview-citations-study)
- Underneath (2026), [What pages cited by AI Overviews have in common: 3,096 pages](https://underneath.agency/research/ai-overview-cited-pages-study)
- Underneath (2026), [AI Mode vs AI Overviews: how different are the sources?](https://underneath.agency/research/ai-mode-vs-ai-overviews-study)

---

This is the Markdown twin of https://underneath.agency/resources/does-ranking-on-google-get-you-into-ai-overviews. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Does readable writing help your content appear in AI answers?"
description: "Somewhat. AI engines cite easier-to-read pages, and fluent rewrites helped in early lab tests, but later tests found polishing text alone rarely helps."
canonical: "https://underneath.agency/resources/does-readable-writing-help-ai-visibility"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Does readable writing help your content appear in AI answers?

Somewhat, but it is not a lever on its own. AI search engines tend to cite easier-to-read pages, and early lab tests found that smoother writing raised a page’s share of an answer. Later, larger tests found that polishing wording rarely changes whether a page is cited, so readability works best as part of genuinely useful content.

## The short version

1. In the first lab study of AI search optimization, making a source more fluent and easier to understand raised its visibility by 15 to 30% ([Aggarwal and colleagues](https://arxiv.org/abs/2311.09735)).
2. Across 55,936 searches, pages cited by AI search engines had a reading ease score of 24.15, against 12.32 for pages from Google and Bing ([Zhang and colleagues](https://arxiv.org/abs/2512.09483)).
3. A later benchmark tested ten rewriting methods, including fluency and simpler language, and found real gains in only three of 54 cases ([Puerto and colleagues](https://arxiv.org/abs/2506.11097)).
4. In a shopping test, a fluency rewrite moved Amazon listings down by 0.40 places on average in Claude’s rankings ([Bagga and colleagues](https://arxiv.org/abs/2511.20867)).
5. One end-to-end test found rewriting page text for AI engines cut final citations by 6%, and rewritten pages were also found less often ([Martinez](https://arxiv.org/abs/2607.14035)).

## What did the first lab test find about fluent writing?

It found a clear lift, in a controlled setting. In the 2023 study that named generative engine optimization, researchers rewrote web sources in several ways. They then measured how much of an AI answer each source accounted for ([Aggarwal and colleagues](https://arxiv.org/abs/2311.09735)).

Two of the rewrites were about style rather than substance. One improved fluency; the other simplified the language. These style changes raised visibility by 15 to 30%.

Fluency also combined well: pairing it with added statistics beat any single method by more than 5.5%. Whether added statistics hold up on their own is covered in [our guide to statistics, quotes and citations](https://underneath.agency/resources/do-geo-content-tactics-work).

The setup matters. The engine was built by the researchers on GPT-3.5, and only the top 5 Google results were placed in front of it for each question. So the test shows what happens once a page is already in the AI engine’s hands, not whether better writing gets it found.

A later study still found fluency the strongest of these original methods ([Wu and colleagues](https://arxiv.org/abs/2510.11438)). Rules learned from each engine beat it by up to 50.99%.

## Do AI engines cite easier-to-read pages in the real world?

Yes, on average they do. A team at the Hong Kong University of Science and Technology and Rutgers compared sources from six AI search engines with Google and Bing results ([Zhang and colleagues](https://arxiv.org/abs/2512.09483)). They used 55,936 searches built from trending topics in July and August 2025.

They sampled 10,000 web addresses from each group and scored the text with two standard reading measures. Reading ease runs from hard to easy; grade level estimates the years of schooling a reader needs.

| Reading measure | Pages from Google and Bing | Pages cited by AI engines |
|---|---|---|
| Reading ease (higher is easier) | 12.32 | 24.15 |
| Grade level (lower is easier) | 18.24 | 14.57 |

Both groups score as hard reading, but the AI-cited pages were noticeably easier. The authors say AI engines “tend to favor more accessible, less textually demanding content.” This compares two groups of pages; it does not show that easier writing caused any citation. Our [study of pages cited by AI Overviews](https://underneath.agency/research/ai-overview-cited-pages-study) compared pages on the same search instead.

## Why did later experiments find polishing text does little?

Because when researchers measured which source an AI engine cites first, style edits rarely moved it. C-SEO Bench, a test across six kinds of content, applied ten methods, including fluency and simpler language, to one page at a time ([Puerto and colleagues](https://arxiv.org/abs/2506.11097)).

Out of 54 combinations of method and content type, only three produced a reliable improvement, and none of those were fluency or simpler language. Many methods pushed pages down. The authors note the earlier lift was measured by how many words an answer devoted to a source, which is not the same as being cited first.

A shopping test found the same. Researchers rewrote Amazon product listings with 15 instructions and measured how far each moved in five AI engines’ rankings ([Bagga and colleagues](https://arxiv.org/abs/2511.20867)).

The fluency rewrite did worse than a plain rewrite. It moved listings down by 0.40 places on average in Claude’s rankings and by 0.21 places in GPT-5’s.

A third test, run by researchers at the software company Sprinklr, compared dense paragraphs with organized sections across 252,000 trials on six AI models. Layout alone had no consistent effect ([Vishwakarma and colleagues](https://arxiv.org/abs/2605.25517)).

## Is readability the cause, or a sign of something else?

Most likely a sign of something else. The studies cannot separate clear writing from the qualities that usually come with it, such as focus, accuracy and usefulness.

In a study of 18,151 pages cited by ChatGPT, Google and Perplexity, an AI model rated the quality of each page ([Zhang Kai and colleagues](https://arxiv.org/abs/2604.25707)). The pages answers drew on most scored about 1.49 times higher than those drawn on least. Relevance to the question was the strongest signal they found, ahead of length or layout.

The Wu study shows what the engines seem to reward. Gemini, GPT and Claude each favored clear, concise language, but alongside comprehensive coverage, accurate facts, specific evidence and a conclusion stated up front ([Wu and colleagues](https://arxiv.org/abs/2510.11438)). Readability sits inside a broader picture of a page that answers the question well. How far the three engines agree is covered in [our guide on whether engines prefer the same content](https://underneath.agency/resources/do-ai-engines-prefer-same-content).

## Can rewriting for AI engines backfire?

Yes. A rewrite can help a page once it is in front of the AI engine, yet make it harder to find in the first place.

A 2026 survey of 45 studies describes one end-to-end test ([Martinez](https://arxiv.org/abs/2607.14035)). There, rewriting page text for AI engines made pages reach the top results less often and cut final citations by 6%.

The C-SEO Bench authors also found that gains shrink as more sites adopt the same method, so any edge from a style trick fades. The survey’s conclusion is that general rules of thumb “transfer poorly” between engines and settings.

## What should you do about it?

Write clearly for people, and treat that as a baseline rather than a growth tactic.

1. Edit for plain, readable language. AI-cited pages read more easily on average, and clear writing serves human readers too.
2. Do not run pages through generic AI rewriters to make them “more fluent”. Controlled tests found little gain and some losses.
3. Put the substance first: a direct answer, specific figures, definitions and comparisons. Those are what the engines appear to draw on.
4. Protect what makes the page findable. Keep the terms people search for when you simplify, because a rewrite that loses them can cost you citations. Keeping terms is not the same as repeating them; see [our guide to keyword stuffing in AI search](https://underneath.agency/resources/does-keyword-stuffing-work-in-ai-search).
5. Test changes on a few pages before rolling them out, and check citations in each engine you care about.

If you want help making your pages both readable and useful to AI answers, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

No study has shown that rewriting a live page for readability raises its citations in real AI search.

- The 15 to 30% lift came from a 2023 test with older models and a fixed set of five sources per question.
- The reading ease gap compares groups of pages. It cannot separate readability from topic, site or quality.
- The benchmarks tested rewrites made by AI models, not careful human editing.
- Most tests used English content. Effects in other languages are unknown.
- Nobody has measured whether readable pages turn AI visibility into more clicks or sales.

## Frequently asked questions

### Does simple language help with AI search?

Slightly, at best. AI-cited pages scored 24.15 on reading ease against 12.32 for Google and Bing pages, but tests that simplified pages found no reliable gain in citations.

### Should we rewrite our content to be more fluent for ChatGPT?

Not as a tactic on its own. In a benchmark of 54 method and content combinations, only three produced reliable gains, and fluency was not among them.

### What reading level do AI-cited pages have?

Fairly high. In a sample of 10,000 AI-cited pages, the average grade level was 14.57, against 18.24 for pages from Google and Bing.

### Did the original GEO study find that readability matters?

Yes, in its setting. Fluency and simpler language raised visibility by 15 to 30%, but only for sources already handed to the AI engine.

## Sources

- Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan and Deshpande (2023), [GEO: Generative Engine Optimization](https://arxiv.org/abs/2311.09735), arXiv:2311.09735.
- Zhang, Ye, Peng, Garimella and Tyson (2025), [Source Coverage and Citation Bias in LLM-based vs. Traditional Search Engines](https://arxiv.org/abs/2512.09483), arXiv:2512.09483.
- Puerto, Gubri, Green, Oh and Yun (2025), [C-SEO Bench: Does Conversational SEO Work?](https://arxiv.org/abs/2506.11097), arXiv:2506.11097.
- Bagga, Farias, Korkotashvili, Peng and Wu (2025), [E-GEO: A Testbed for Generative Engine Optimization in E-Commerce](https://arxiv.org/abs/2511.20867), arXiv:2511.20867.
- Vishwakarma, Kumar and Jamidar (2026), [What Gets Cited: Competitive GEO in AI Answer Engines](https://arxiv.org/abs/2605.25517), arXiv:2605.25517.
- Wu, Zhong, Kim and Xiong (2025), [What Generative Search Engines Like and How to Optimize Web Content Cooperatively](https://arxiv.org/abs/2510.11438), arXiv:2510.11438.
- Zhang Kai, He Xinyue and Yao Jingang (2026), [From Citation Selection to Citation Absorption: A Measurement Framework for Generative Engine Optimization Across AI Search Platforms](https://arxiv.org/abs/2604.25707), arXiv:2604.25707.
- Martinez (2026), [Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026)](https://arxiv.org/abs/2607.14035), arXiv:2607.14035.

---

This is the Markdown twin of https://underneath.agency/resources/does-readable-writing-help-ai-visibility. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Does Reddit content actually shape Google AI Overviews?"
description: "Less than citations suggest. Google’s AI Overviews cite Reddit, but one study found they drew 22.1 points less content from forums than from other sources."
canonical: "https://underneath.agency/resources/does-reddit-shape-google-ai-overviews"
published: 2026-10-10
updated: 2026-10-10
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Does Reddit and forum content actually shape Google AI Overview answers?

Reddit and forum threads shape Google’s AI Overviews less than their citations suggest. A 2026 study of 11,000 real search queries found that Google cites Reddit or Quora in 7 to 8% of queries. But it draws much less of the summary’s content from them than from other sources. Our own check found that one in five sentences citing a Reddit thread was not supported by the top of that thread.

## The short version

1. In a study of 11,000 real search queries, [Huang and colleagues](https://arxiv.org/abs/2603.16138) found Google’s AI Overviews cited Reddit in 7% of queries, against 42% for Google’s regular results.
2. When cited, social and forum sources were under-used: their content was reflected in the summary 22.1 percentage points less than the average cited source.
3. In [our Reddit study](https://underneath.agency/research/ai-reddit-citations-study) of 800 US searches, 17.9% of AI Overviews cited a Reddit thread, rising to 37.7% in home and local services.
4. Judged against the post and its top 10 comments, 36.4% of sentences citing Reddit were supported, 42.8% partly supported and 20.8% not supported.
5. Reddit citations are unstable: in our check two days apart, only 50.0% of the cited threads were cited again.

## How often do Google’s AI Overviews cite Reddit and forums?

Regularly, but far less often than Google’s regular results show them. An AI Overview is the AI summary that appears at the top of some Google results pages.

[Huang and colleagues](https://arxiv.org/abs/2603.16138) sent 11,000 real search queries to Google from one US location and compared the AI Overview with the regular results. Only 57.8% of queries produced an AI Overview at all. Where they did, AI Overviews and regular results prominently feature user-generated platforms: Facebook (10% and 36%), Quora (8% and 24%), and Reddit (7% and 42%). Social platforms made up 8.5% of AI Overview citations, against 13.4% of regular results. The wider picture for social platforms is in [our guide to social media and AI search](https://underneath.agency/resources/does-social-media-help-ai-search-visibility).

The kind of search matters a lot. Those queries were general information questions. In [our Reddit study](https://underneath.agency/research/ai-reddit-citations-study) of 800 US commercial keywords, 17.9% of AI Overviews and 9.8% of AI Mode answers cited a Reddit thread. Rates differed sharply by industry.

| Industry (our study, September 2026) | AI Overviews citing Reddit |
|---|---|
| Home and local services | 37.7% |
| B2B software and technology | 35.4% |
| Hospitality and travel | 16.1% |
| Retail and ecommerce | 10.3% |
| Financial services and insurance | 5.7% |

## When an AI Overview cites Reddit, does it actually use what the thread says?

Often not much: forum sources are cited, but little of their content reaches the summary. Huang and colleagues broke each AI Overview into single factual statements. They then used automated tools to check how much of each cited source those statements reflected.

Social media and forum sources such as Reddit and Quora were under-covered by 22.1 points compared with other cited sources. In the authors’ words, the model lists user-driven sources in its references but draws disproportionately little content from them. They call this a gap between citation and synthesis.

Other sources gained at forums’ expense. Wikipedia was both the most-cited domain and over-represented in the summaries, by 5.4 points in AI Overviews. [Sources with a negative tone were under-covered](https://underneath.agency/resources/do-ai-overviews-downplay-negative-content) by 13.8 points. Long sources of over 800 words were favored over short ones.

The authors caution that the coverage scores come from automated tools they did not check again by hand. Part of the length effect, they note, comes from how the scores were calculated. The direction for forums was clear, but the exact size is an estimate.

## When Google’s AI does lean on a thread, does the thread back up the sentence?

Only partly: in our check, a fifth of sentences citing Reddit were not supported by the top of the thread. In [our Reddit study](https://underneath.agency/research/ai-reddit-citations-study), 96.9% of Reddit citations in Google’s AI answers were attached to a specific sentence. Two AI coders read each of 250 such sentences against the post and its top 10 comments.

They found 36.4% of sentences were supported, 42.8% partly supported and 20.8% not supported. “Partly” was often one commenter’s view stated as what “users” say. Product recommendations were the weakest: 23.0% of recommendation sentences were not supported by the thread they cited.

| What the Reddit thread was cited for | Share of sentences |
|---|---|
| A fact (price, rule, specification) | 25.2% |
| A recommendation of a named option | 24.4% |
| People’s experience or opinion | 19.6% |
| A procedure or fix | 13.2% |

So when a thread does shape the answer, the answer may stretch it. A general “which tool do you use?” thread can end up cited for praise of a product the top comments do not mention.

## Which Reddit threads get cited?

Threads with real discussion, often from communities of practitioners, and often not ones Google ranks. Among threads Google showed for the same search, cited ones had a median of 40 comments against 20 for threads the AI skipped. Question-style titles did not reliably predict citation.

Specialist communities, where people who do the work answer questions about it, supplied 50.8% of Google’s Reddit citations in our data. Examples were r/askaplumber and r/projectmanagers.

Google’s ranking is only a partial guide. 53.9% of Reddit threads cited in AI Overviews appeared in neither the regular top 10 nor the forums box for the same search. On 505 searches where we recorded Google’s top 100, 17 of the 27 Reddit citations went to threads outside the top 100.

## Is Reddit’s place in AI answers stable?

No: which threads get cited shifts with the date, the wording and the AI engine. For the same 96 keywords two days apart, 62.5% of the AI Overviews that cited Reddit still did, and only 50.0% of the cited threads were cited again. Rewording a keyword as a question kept only 7.1% of the cited threads.

Other AI engines treat forums very differently. Huang and colleagues found ChatGPT with search drew almost none of its citations from social platforms (0.1%). In our study, ChatGPT and Claude cited no Reddit thread across 80 buyer questions, while Perplexity did in 12.5%. [Chen and colleagues](https://arxiv.org/abs/2509.08919) likewise found AI search engines nearly excluded social sources for product ranking questions, unlike Google.

Platform shifts can be sudden. [Finder and colleagues](https://arxiv.org/abs/2609.34951) report, citing industry trackers, that Reddit’s share of ChatGPT citations fell 86 to 95% within a week in August 2026, depending on the panel.

## Do AI Overviews send people to Reddit?

Yes for experience-based communities, at least before AI Mode. Using public Reddit data, [Zhang and colleagues](https://arxiv.org/abs/2605.16428) compared communities Google’s AI Overviews can draw on with adult communities it excludes. After AI Overviews expanded in August 2024, daily comments in the eligible communities rose by 12.0% relative to the excluded ones.

The gain was 2.3 times larger for comments in communities built on opinions, advice and personal experience than in fact-based ones. People still visit forums for what a summary cannot give them. But after Google launched AI Mode, the conversational version of its AI search, that premium fell by 59% for comment authors. Part of the experience-seeking traffic stayed inside Google. What this means for publishers is covered in [our guide to content businesses and AI summaries](https://underneath.agency/resources/content-business-risk-from-ai-search).

## What should you do about it?

Treat Reddit as a place where buyers talk, not as a shortcut to control Google’s AI summary.

1. **Check your industry first.** In our data, Reddit citations were common in home services and software, and rare in financial services. Look at the AI Overviews for your own buyer searches before investing.
2. **Do not measure success by citations alone.** A cited thread may contribute little to what the summary says. Read the sentence the citation is attached to, and check whether it describes you accurately.
3. **Take part honestly where your buyers ask questions.** Cited threads tended to have real discussion in specialist communities. Staff should post as themselves; posing as customers breaks Reddit’s rules.
4. **Keep investing in sources that get absorbed.** Encyclopedic, long and detailed sources were over-represented in the summaries Huang and colleagues studied. Your own detailed pages and earned coverage carry more of the answer.
5. **Monitor repeatedly.** One check is a snapshot. Thread citations changed within two days and with small changes in wording.

If you want a repeatable way to track this, our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) covers it.

## What does the research not tell us yet?

No study yet shows that changing a Reddit thread changes what Google’s AI says about a business.

- **Most evidence is observational.** Both the Huang study and ours describe patterns. Neither changed a thread to test the effect.
- **The coverage scores are automated.** Huang and colleagues measured how much of each source the summary reflects with automated tools, without extra checks by people. Our support labels were coded by two AI models, not people.
- **The query sets differ.** Huang and colleagues used general information questions collected around 2018, from one US location at one point in time. Our study used US commercial keywords on one day in September 2026.
- **Small samples on stability.** Our two-day and rewording comparisons rest on 12 to 16 Reddit-citing answers each.
- **Influence on buyers is unmeasured.** No study here measured whether people who read an AI Overview citing Reddit trust or act on it differently.

## Frequently asked questions

### Should my brand post on Reddit to show up in Google AI Overviews?

Only as honest participation, not as a ranking tactic. Cited forum sources contribute less content than other sources in one large study, and no study has shown that posting changes the AI summary.

### Why does an AI Overview cite Reddit if it doesn’t use the content?

The research does not say why. Huang and colleagues describe a gap where the summary lists forum sources but draws disproportionately little content from them, which can suggest more breadth than the text reflects.

### Does ChatGPT cite Reddit as much as Google does?

Not in the studies we reviewed. ChatGPT with search drew 0.1% of its citations from social platforms in one study, and cited no Reddit thread across our 80 buyer questions.

### Does being cited on Reddit drive traffic to the thread?

It can. One study found AI Overviews raised daily comments in eligible Reddit communities by 12.0%, mostly in experience-based ones, though AI Mode reduced that gain.

## Sources

- Huang, Goyal, Saha and Chandrasekharan (2026), [Answer Bubbles: Information Exposure in AI-Mediated Search](https://arxiv.org/abs/2603.16138), arXiv:2603.16138.
- Zhang, Cui and Zhang (2026), [The Impact of AI Search on the Online Content Ecosystem: Evidence from Google and Reddit](https://arxiv.org/abs/2605.16428), arXiv:2605.16428.
- Chen, Wang, Chen and Koudas (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919), arXiv:2509.08919.
- Finder, Elovic, Shalev and Yosef (2026), [AX is the New AEO](https://arxiv.org/abs/2609.34951), arXiv:2609.34951.
- Underneath (2026), [When does a Reddit thread become evidence in Google’s AI?](https://underneath.agency/research/ai-reddit-citations-study)

---

This is the Markdown twin of https://underneath.agency/resources/does-reddit-shape-google-ai-overviews. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Does social media content help your brand show up in AI search?"
description: "Mostly on Google. AI Overviews often cite Facebook, Instagram and YouTube; ChatGPT and Claude rarely cite social platforms at all."
canonical: "https://underneath.agency/resources/does-social-media-help-ai-search-visibility"
published: 2026-10-07
updated: 2026-10-07
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Does social media content help your brand show up in AI search?

Social media helps mainly where Google’s AI is the engine, and much less elsewhere. Google’s AI Overviews cite YouTube, Facebook and Instagram among their most-used sources, while ChatGPT and Claude barely cite social platforms. No study has yet tested whether a brand’s own posts raise its mentions in AI answers.

## The short version

1. In a US study of 55,393 searches, YouTube, Facebook and Instagram were three of the five sites Google’s AI Overviews cited most ([Xu and colleagues](https://arxiv.org/abs/2605.14021)).
2. OpenAI’s GPT-4o-mini with web search took just 0.1% of its citations from social platforms, against 8.5% for Google’s AI Overviews ([Huang and colleagues](https://arxiv.org/abs/2603.16138)).
3. Even when cited, social and forum posts fed Google’s AI summaries 22.1 percentage points less than other cited sources (Huang and colleagues).
4. Beauty and fashion searches drew 28.9% of AI Overview citations from social and community platforms, health just 10.7% (Xu and colleagues).
5. Beyond YouTube, Facebook, Instagram and Reddit, all other social platforms together supplied 0.51% of AI Overview citations (Xu and colleagues).

## Do AI search engines cite social media at all?

Yes, Google’s AI Overviews cite social platforms often. YouTube, Facebook and Instagram rank among their most-cited sites.

[Xu and colleagues](https://arxiv.org/abs/2605.14021) ran 55,393 trending US searches over 40 days in 2026 and logged every source in Google’s AI Overviews, the AI summary at the top of Google’s results. The most-cited site was youtube.com at 5.49% of all citations. Facebook followed at 3.68% and Instagram at 3.65%, both ahead of most news sites.

[Our own AI Overview study](https://underneath.agency/research/ai-overview-citations-study) of 481 US AI Overviews found the same pattern. Facebook was cited in 7.1% of them. Forums, social and video platforms together made up 20.0% of all citations. For a wider view of which kinds of sites AI engines rely on, see [which websites AI search engines cite](https://underneath.agency/resources/which-websites-do-ai-search-engines-cite).

## Which AI engines cite social platforms most?

Google’s AI features and Perplexity cite social content; ChatGPT and Claude mostly do not.

[Huang and colleagues](https://arxiv.org/abs/2603.16138) sent 11,000 real search questions to several systems. OpenAI’s GPT-4o-mini with web search drew 0.1% of its citations from social platforms. Google’s AI Overviews drew 8.5%, and Google’s ordinary results 13.4%. [Chen and colleagues](https://arxiv.org/abs/2509.08919) looked at searches about well-known brands. Their “social” group also counts Reddit, YouTube and Quora.

| AI engine | Share of sources from social platforms, well-known brand searches |
|---|---|
| ChatGPT | none |
| Claude | 5.9% |
| Gemini | 11.5% |
| Perplexity | 23.8% |

Google’s two AI features also differ. In [our comparison of AI Mode and AI Overviews](https://underneath.agency/research/ai-mode-vs-ai-overviews-study), on the 236 US searches where both cited sources, forums, social and video made up 18.8% of AI Overview citations against 11.2% for AI Mode, Google’s chat-style search tab.

## Does a cited social post shape the answer as much as a website?

No. Google’s AI cites social posts more than it uses them, and cites them far less than ordinary results show them.

Huang’s team compared what Google’s AI summaries cited with what they actually said. Content from social media and forums was drawn on 22.1 percentage points less than other cited sources. A citation, in other words, did not mean the post shaped the answer. Our guide on [whether Reddit shapes Google AI Overviews](https://underneath.agency/resources/does-reddit-shape-google-ai-overviews) looks at the same gap for forum threads.

Google’s AI is also more selective than Google itself. In Xu’s study, 14.2% of AI Overview sources came from social and community platforms, against 49.9% of the links on Google’s first page of results.

## Which industries lean most on social content?

Visual, experience-led categories lean most on social content. Institutional topics like health lean least.

Xu’s team found that the share of AI Overview citations from social and community platforms varied by a factor of 3.1 across categories. It was highest in beauty and fashion at 28.9% and autos at 27.4%. It was lowest in health at 10.7% and climate at 9.3%.

Shopping questions follow the same pattern. In an audit of AI product recommendations by [Uberti-Bona Marin and colleagues](https://arxiv.org/abs/2609.18729), social and community sources made up 22.2% of the sources displayed in Google’s AI Overviews. Reputation questions do too. In [our study of “is this brand legit?” questions](https://underneath.agency/research/is-it-legit-ai-reputation-study), 43.0% of Google AI Mode answers cited a forum, social or video source, against 11.4% for ChatGPT.

## What about LinkedIn, X and TikTok?

There is little evidence that they matter much yet. Four platforms account for nearly all social citations in Google’s AI.

In Xu’s study, YouTube, Facebook, Instagram and Reddit supplied 96.4% of the social and community citations in AI Overviews. LinkedIn, X, TikTok, Pinterest, Quora and Threads together supplied 0.51% of all citations.

LinkedIn does appear in a different dataset. [Zhang and colleagues](https://arxiv.org/abs/2604.25707) analyzed 21,143 citations from ChatGPT, Google and Perplexity on 602 test prompts. There, linkedin.com was the fifth most-cited site, with 187 citations. That dataset was built by one of the authors and mixes English and Chinese prompts, so treat it as a hint rather than a measure. None of the studies we reviewed reported X or TikTok on their own.

## What should you do about it?

Treat social content as a Google channel first and a source of real evidence, not a shortcut.

1. Check which engine your buyers use. If they rely on ChatGPT or Claude, social posts will rarely be cited, and independent reviews and press matter more.
2. If Google’s AI matters to you, keep YouTube, Facebook and Instagram profiles accurate and useful, especially in visual or consumer categories.
3. Publish things people can quote: clear product facts, demonstrations and answers to common questions, not slogans.
4. Do not seed fake reviews or posts. Research on [fake reviews and fake brands in AI recommendations](https://underneath.agency/resources/fake-reviews-fake-brands-ai-recommendations) shows how the practice works and why it puts brands at risk.
5. Measure by engine. Track whether social sources appear in answers about your category on each AI engine before shifting budget.

If you want help deciding where social content fits, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The studies count which platforms get cited. None tests whether a brand’s own posting changes its visibility.

- No study we found compared brands that post often with brands that do not, and measured their AI mentions.
- Most data covers Google in the US. Other countries and languages are thinly studied.
- “Social” groups differ between studies; some include Reddit, YouTube and Quora, which behave like forums or video libraries.
- Several studies could not read social pages behind logins, so how AI engines use their content is partly unknown.
- AI engines change often, and these figures are snapshots from 2025 and 2026.

## Frequently asked questions

### Does ChatGPT use Facebook or Instagram posts?

Rarely, based on current research. OpenAI’s GPT-4o-mini with web search took 0.1% of its citations from social platforms in Huang’s study of 11,000 questions.

### Should we post more on Instagram to show up in AI Overviews?

It may help in visual categories, but no study has tested it. Instagram supplied 3.65% of AI Overview citations in Xu’s study, and social sources mattered most in categories like beauty and fashion.

### Does LinkedIn content show up in AI search?

Sometimes, though the evidence is thin. LinkedIn was the fifth most-cited site in one dataset of 21,143 citations, but a tiny share of Google’s AI Overview citations in another.

### Is YouTube more useful than other social platforms for AI search?

For Google, yes. YouTube was the single most-cited site in Google’s AI Overviews at 5.49% of citations; see [our YouTube study](https://underneath.agency/research/ai-overview-youtube-videos-study) for which videos get cited.

## Sources

- Xu and colleagues (2026), [Measuring Google AI Overviews: Activation, Source Quality, Claim Fidelity, and Publisher Impact](https://arxiv.org/abs/2605.14021), arXiv:2605.14021.
- Huang and colleagues (2026), [Answer Bubbles: Information Exposure in AI-Mediated Search](https://arxiv.org/abs/2603.16138), arXiv:2603.16138.
- Chen and colleagues (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919), arXiv:2509.08919.
- Uberti-Bona Marin and colleagues (2026), ["If I Had to Buy Just ONE: Galaxy S26 Ultra": Auditing AI-Generated Product Recommendations](https://arxiv.org/abs/2609.18729), arXiv:2609.18729.
- Zhang, He and Yao (2026), [From Citation Selection to Citation Absorption: A Measurement Framework for Generative Engine Optimization Across AI Search Platforms](https://arxiv.org/abs/2604.25707), arXiv:2604.25707.
- Underneath (2026), [Do AI Overviews cite the pages that rank? 4,051 citations](https://underneath.agency/research/ai-overview-citations-study)
- Underneath (2026), [AI Mode vs AI Overviews: how different are the sources?](https://underneath.agency/research/ai-mode-vs-ai-overviews-study)
- Underneath (2026), [“Is this brand legit?” How AI assistants build a reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study)
- Underneath (2026), [The YouTube videos Google’s AI cites are small](https://underneath.agency/research/ai-overview-youtube-videos-study)

---

This is the Markdown twin of https://underneath.agency/resources/does-social-media-help-ai-search-visibility. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Can AI search lower a DTC brand’s cost of new customers?"
description: "Possibly. AI-referred shoppers bring a high share of new buyers and convert well, but the channel is still small next to paid social and search."
canonical: "https://underneath.agency/resources/dtc-brands-ai-search-sales"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Can AI recommendations lower what a DTC brand pays for each new customer?

Possibly, and the early signs point that way: shoppers sent by AI assistants arrive on product pages ready to buy, and a high share of them are new to the store. But the channel is still small next to paid social and search, and no one has yet published what an AI-sourced customer costs. For a direct-to-consumer brand, AI visibility is best treated as a new acquisition channel to build, not a replacement for the ones that pay the bills.

## The short version

1. Paid acquisition keeps getting dearer. [Meta](https://s21.q4cdn.com/399680738/files/doc_financials/2026/q2/Meta-06-30-2026-Exhibit-99-1-FINAL.pdf) reported its average price per ad rose 12% year over year in the second quarter of 2026.
2. AI referrals bring new customers. [Shopify’s president said](https://www.digitalcommerce360.com/2026/08/06/how-shopify-is-approaching-agentic-ai-2026/amp/) “new buyer orders are coming in at nearly twice the rate of other channels” from AI, and AI-driven traffic and orders to Shopify stores tripled year over year.
3. Those shoppers buy more readily. [Shopify’s data](https://www.shopify.com/enterprise/blog/ai-search-insights) shows AI-referred visits to product pages convert at nearly 50% higher rates than organic search, with 14% higher order values.
4. Specialist brands benefit most so far. 75% of AI-attributed purchases on Shopify in the second quarter came from outside its top 100 product categories, as [Retail TouchPoints reported](https://www.retailtouchpoints.com/news/shopify-credits-ai-for-34-revenue-growth-in-q2-2026/620805/).
5. The trust question is decided on review sites. In [our study](https://underneath.agency/research/is-it-legit-ai-reputation-study) of “Is this brand legit?” answers, 88.0% cited a review or complaint platform, and claims resting only on review platforms were negative 56.5% of the time.

## Who buys from a DTC brand, and what does each new customer cost today?

The buyer is a stranger the brand must pay to reach, usually through social ads and creators.

A direct-to-consumer (DTC) brand sells mostly through its own online store. That gives it the margin and the customer data a marketplace seller gives up, but it also means every new customer has to be found and persuaded from scratch. For most DTC brands that work is done by paid social: Meta and TikTok ads, creator partnerships and affiliate links, with Google search catching the people who already know the name.

The price of that work is rising. Meta, the largest seller of social ads, said ad impressions rose 14% and the average price per ad rose 12% year over year in the second quarter of 2026. Social is also becoming a bigger share of how online purchases start. Over the 2025 holidays, [Adobe](https://news.adobe.com/news/2026/01/adobe-holiday-shopping-season) found social media’s share of online retail revenue reached 4.6%, and the affiliates and partners channel, which includes influencers, reached 20.4%.

Profit is hard to come by on those terms. When Warby Parker reported its first full year of net income, $1.6 million on $871.9 million of revenue, [Retail Dive](https://www.retaildive.com/news/warby-parker-first-annual-net-income-fourth-quarter-2025-earnings/813223) noted it joined “a small group of DTC brands that have been able to achieve profitability.” The executive question is whether a new channel can bring in first-time buyers without the ad bill rising with them. Insurtech startups weigh the same trade as funding tightens, favoring a channel built on evidence rather than ad spend, as [our guide to insurtech customers and AI](https://underneath.agency/resources/insurtech-customers-ai-search) explains.

## Where do AI assistants fit in how shoppers find DTC brands?

They sit where a shopper describes a problem and asks what to buy, before any ad reaches them.

A social ad interrupts someone who was not shopping. An assistant answers someone who is. That difference shows up in the data. [Shopify](https://www.shopify.com/enterprise/blog/ai-search-insights) found that more than half of AI-referred sessions start on a product page, against about 20% for organic search. The shopper has already compared options in the conversation and clicks through to a specific product.

The channel is growing fast from a small base. Shopify says AI-referred orders on its stores grew nearly 13 times year over year in the first quarter of 2026, and AI chatbot referral sessions more than 8 times. [Triple Whale](https://www.triplewhale.com/reports-guides/chatgpt-ads), an analytics company for ecommerce brands, says AI-attributed orders it sees rose 105.5% between Black Friday and Cyber Monday 2025 and July 2026, with ChatGPT behind roughly three out of four of them.

It is not yet large. Shopify says organic search still refers more sessions to its merchants than all tracked AI platforms combined, and Harley Finkelstein called agentic commerce volume “still small relative to our massive GMV.” Traditional search sessions on Shopify stores are up 1.3 times over two years and hold roughly one-third of all storefront sessions. AI adds to search; it has not replaced it.

## Which questions do shoppers ask AI before buying from a DTC brand?

They describe a problem, compare named brands, and ask whether an unfamiliar brand can be trusted.

The prompts below are illustrative, written to show the stages; they are not observed data.

- **Problem first:** “What bra won’t irritate my eczema?” or “Best cookware without nonstick coatings.”
- **Brand comparison:** “Is [brand] worth it compared with [rival]?” or “Which mattress brands are best for back pain?”
- **Dupe hunting:** “Something like [premium brand] hoodies but cheaper.”
- **Trust check:** “Is [brand] legit? What do customers say?”
- **Fit and returns:** “Does [brand] run small, and what is the return policy?”

The first type is where DTC brands built around a specific need can win. [Shopify’s holiday report](https://www.shopify.com/news/agentic-holiday-2026) quotes Cottonique, a hypoallergenic apparel brand, whose co-founder says “Nobody with eczema types ‘hypoallergenic bra.’” Assistants can match the plain-language problem to the product whose page explains it. Cottonique says its AI-referred sales grew 276% year over year, with total sales up 52%; those are the brand’s own figures, not an independent test.

## How does an AI recommendation turn into a DTC sale?

The assistant shortlists a product, the shopper lands on its page, checks reviews, and buys.

For a DTC brand the path is short: a question, a product named in the answer, a click to the brand’s own product page (or, where an assistant offers it to the merchant, a checkout inside the assistant), a first order, and then the retention program that makes the customer profitable. The assistant replaces the ad as the first touch, and the brand pays nothing per impression for it.

What the brand still pays for is being the right answer. Three findings matter here:

- **New customers, not just existing ones.** Finkelstein said new buyer orders from AI channels arrive at nearly twice the rate of other channels. For a DTC brand that pays most of its marketing budget to find first-time buyers, that is the line that counts.
- **Higher intent.** Shopify found AI-referred conversion beat organic search in 23 of 25 merchant categories, by an average of 56% within those categories. [Salesforce’s holiday data](https://www.salesforce.com/news/stories/2025-holiday-shopping-data/) says shoppers referred by AI search converted nine times more often than those from social media.
- **The maker has a documented edge.** [OpenAI says](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search) that when several merchants sell a product, ChatGPT ranks them on availability, price, quality and whether the seller is “the maker or primary seller of that item.” A DTC brand is the maker, which helps it keep the sale on its own store if price and stock are competitive.

We infer that the economics improve most for brands whose products need explaining. Aviator Nation, which sells tie-dyed hoodies for around $200, told Shopify it has leaned less on paid acquisition and sees ChatGPT as one of several touchpoints that bring shoppers back to the brand.

## What makes an assistant suggest one DTC brand over another?

Clear, complete product data and independent evidence of quality; platforms document some inputs, studies observe others.

**Documented by the platform.** OpenAI says ChatGPT considers structured product data from merchants and third parties, such as price and description, “and other third-party content,” and summarizes reviews from public websites. It says product results are not ads; ads are shown separately, and Triple Whale’s guide already reports early results from 188 shops running ChatGPT ads. Shopify merchants’ data reaches ChatGPT through Shopify Catalog.

**Reported by a platform, not independently tested.** Finkelstein said AI searches powered by Shopify’s Catalog “converted twice the rate of those using scraped data,” because products show up complete and accurate. Shopify also says 75% of AI-attributed purchases came from outside its top 100 categories, which suggests niche products are getting found.

**Observed in our study.** When shoppers ask whether a brand is legit, the answer leans on review platforms. In our study of 79 brands, Trustpilot and the BBB accounted for 61.7% of all review-platform citations. Every complete answer called the brand legitimate, but 99.7% raised at least one problem. For a young DTC brand with a thin review history, a handful of complaints on those sites can become the assistant’s summary.

**Our inference.** Studies covered in [how small brands get recommended by AI](https://underneath.agency/resources/how-small-brands-get-recommended-by-ai) suggest independent coverage matters more than a brand’s own site. Advertising budgets do not appear to buy recommendations directly, as our piece on [whether ad spend helps AI brand recommendations](https://underneath.agency/resources/does-ad-spend-help-ai-brand-recommendations) explains.

## What does it cost a DTC brand to stay invisible to AI assistants?

It keeps the brand fully dependent on paid channels whose prices are rising.

We cannot put a dollar figure on it; no study has measured lost DTC sales from AI absence. The risk is structural. If a growing share of high-intent, first-time buyers starts with an assistant, and Shopify’s data says those shoppers are over-represented among new buyers, then a brand the assistant does not name must buy those customers back through ads at Meta’s rising prices. Meanwhile, rivals the assistant does name acquire them at no media cost.

There is also a reputation cost. When someone asks whether your brand is legit, the assistant builds an answer anyway, from whatever review and complaint pages it finds. In our study, claims attached only to review platforms were negative 56.5% of the time. A brand that ignores those pages lets them write the first impression.

## How does GEO work for a DTC brand?

GEO makes your products easy for assistants to understand, verify and recommend, without promising that they will.

Generative engine optimization (GEO) for a DTC brand covers six areas:

1. **Problem-led product pages.** Write each hero product page around the problems it solves and who it is for, in shoppers’ own words, with specs, sizing, materials and care in plain text. AI shoppers land on product pages, so the page is also the first impression.
2. **Complete product data.** Keep your catalog and feeds full and current; Shopify’s Catalog figures suggest complete data converts better than scraped pages.
3. **Third-party coverage.** Earn reviews from independent publications and creators who explain why the product is different. This is digital PR aimed at the sources assistants cite, not just at reach.
4. **Review platform hygiene.** Claim your Trustpilot and BBB profiles, answer complaints, and ask real customers for detailed reviews. Never buy or fake them.
5. **Consistent facts everywhere.** Price, return policy, shipping and claims should match across your site, retail partners, Amazon if you sell there, and press coverage. [Our guide to AI shopping assistants](https://underneath.agency/resources/ecommerce-brands-ai-shopping-assistants) covers retailer and marketplace listings in more depth.
6. **Measure it as acquisition.** Track AI referrals as their own channel, compare their new-customer share and order value with paid social, and check regularly how assistants answer your category and trust questions. [Why analytics miss AI visibility](https://underneath.agency/resources/why-analytics-miss-ai-visibility) explains what referral data cannot see.

## What don’t DTC founders know yet about AI-sourced customers?

It does not tell us what an AI-sourced customer costs, or how long they stay.

The strongest numbers here come from platforms with an interest in the channel: Shopify, which sells Agentic Storefronts, and Triple Whale, which sells attribution. Neither has published how its figures are calculated in detail. “New buyer orders at nearly twice the rate” is a quote from an earnings call, not a published metric. Brand results such as Cottonique’s are self-reported.

No one has published the repeat purchase rate or lifetime value of customers who arrived through an assistant, which is what decides whether they are cheaper customers or merely cheaper first orders. Attribution is also weak: a shopper who sees a brand in ChatGPT and later clicks a Meta ad will be credited to Meta. Testing whether GEO work caused a rise is covered in [how to prove GEO caused sales](https://underneath.agency/resources/prove-geo-caused-sales).

## Where should a DTC brand start?

Start with the questions that should lead to your hero products, and the trust question new buyers ask.

List the five to ten problems your best-selling products solve and write them as a shopper would ask an assistant. Add “Is [brand] legit?” and “[brand] vs [main rival].” Run them in ChatGPT, Google AI Mode, Gemini and Perplexity, several times each, and note whether you appear, how you are described, and which review pages are cited. Then check your AI referral traffic’s share of new customers against paid social.

If you want help turning that into a plan, [tell us about your brand](https://underneath.agency/contact) and we will show where assistants name you and your rivals today, what they say about your reputation, and which fixes are most likely to bring in first-time buyers without adding to your ad spend. The [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page describes the ongoing work behind such a plan, from problem-led product pages and complete feeds to review-site upkeep and independent coverage.

## Frequently asked questions

### Does ChatGPT send traffic to DTC brand websites or to Amazon?

Both. OpenAI says ChatGPT ranks merchants partly on whether the seller is the maker or primary seller, so a DTC brand with competitive price and stock can win the click to its own store.

### Can AI search replace Meta ads for a DTC brand?

Not today. AI referrals are growing fast but remain small next to paid social and organic search. Treat AI visibility as an added acquisition channel and measure it against paid social.

### Do Shopify brands appear in ChatGPT automatically?

OpenAI says Shopify merchants’ product data is integrated through Shopify Catalog with no extra work required. Appearing in the feed is not the same as being recommended; that still depends on the answer.

### What do AI assistants say when asked whether a brand is legit?

In our study, every complete answer called the brand legitimate, but almost all raised problems, usually drawn from review and complaint sites such as Trustpilot and the BBB.

## Sources

- Meta (2026-07-29), [Meta Reports Second Quarter 2026 Results](https://s21.q4cdn.com/399680738/files/doc_financials/2026/q2/Meta-06-30-2026-Exhibit-99-1-FINAL.pdf)
- Digital Commerce 360 (2026-08-06), [How Shopify is approaching agentic AI so far in 2026](https://www.digitalcommerce360.com/2026/08/06/how-shopify-is-approaching-agentic-ai-2026/amp/)
- Retail TouchPoints (2026-08), [Shopify Credits AI for 34% Revenue Growth in Q2 2026](https://www.retailtouchpoints.com/news/shopify-credits-ai-for-34-revenue-growth-in-q2-2026/620805/)
- Shopify (2026), [AI search insights](https://www.shopify.com/enterprise/blog/ai-search-insights)
- Shopify (2026-10-06), [Welcome to the first holiday season of the agentic era](https://www.shopify.com/news/agentic-holiday-2026)
- Triple Whale (2026), [ChatGPT Ads: The Future of Discovery](https://www.triplewhale.com/reports-guides/chatgpt-ads)
- Adobe (2026-01-07), [Holiday Shopping Season Drove a Record $257.8 Billion Online with Consumers Embracing Generative AI Tools](https://news.adobe.com/news/2026/01/adobe-holiday-shopping-season)
- Salesforce (2026-01), [AI and agents account for $262 billion of 2025 holiday spend](https://www.salesforce.com/news/stories/2025-holiday-shopping-data/)
- Retail Dive (2026-02-26), [Warby Parker posts first annual net income](https://www.retaildive.com/news/warby-parker-first-annual-net-income-fourth-quarter-2025-earnings/813223)
- OpenAI (2026), [Shopping with ChatGPT Search](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search)
- Underneath (2026), [“Is this brand legit?” How AI assistants build a reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study)

---

This is the Markdown twin of https://underneath.agency/resources/dtc-brands-ai-search-sales. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How do ecommerce brands get chosen by AI shopping assistants?"
description: "By making every listing, retailer page and review source agree, so ChatGPT, Google, Rufus and Perplexity can find, compare and trust your products."
canonical: "https://underneath.agency/resources/ecommerce-brands-ai-shopping-assistants"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How does an ecommerce brand get its products chosen by AI shopping assistants?

By making the same accurate facts about each product show up everywhere an assistant looks: your own store, every retailer and marketplace listing, product feeds, reviews and the editorial sites that compare products. Shoppers now ask ChatGPT, Google’s AI Mode, Gemini, Perplexity and Amazon’s Rufus what to buy, and those assistants draw on different sources. A brand that sells across many channels has to be legible in all of them.

## The short version

1. AI shopping is now a measurable sales channel. [Adobe](https://news.adobe.com/news/2026/01/adobe-holiday-shopping-season) counted a 693.4% rise in traffic from AI tools to US retail sites over the 2025 holidays, a season in which Americans spent $257.8 billion online.
2. The visits are worth more than average. In March 2026, Adobe’s data showed [AI traffic converted 42% better](https://techcrunch.com/2026/04/16/ai-traffic-to-us-retailers-rose-393-in-q1-and-its-boosting-their-revenue-too/) than other traffic and earned 37% more revenue per visit.
3. Marketplaces have their own assistants. Amazon says Rufus was used by 300 million customers in 2025 and drove nearly $12 billion in incremental annualized sales, as [Modern Retail reported](https://www.modernretail.co/?p=159156).
4. [OpenAI documents](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search) that when several merchants sell a product, ChatGPT ranks them on availability, price, quality and whether the seller is the maker or primary seller. Which channel gets the sale is part of the answer.
5. The assistants do not agree on sources. In a September 2026 audit, [ChatGPT and Gemini shared only 5.4%](https://arxiv.org/abs/2609.18729) of the websites they showed for the same product question, so visibility in one says little about the others.

## Who is the shopper, and what is an AI-assisted sale worth to a multichannel brand?

The shopper is any consumer researching a purchase, and the sale can land on your site or a retailer’s.

Most consumer brands no longer sell through one door. A vacuum, a toy or a skincare line is typically sold on the brand’s own store, on Amazon and Walmart, through Target or a specialty chain, and sometimes on Etsy or a resale site. The executive question is not only “will an assistant mention us?” but “when it does, where does the shopper buy, and at what margin?”

The market is large and still shifting online. The [US Census Bureau](https://www.census.gov/retail/mrts/www/data/pdf/ec_current.pdf) estimates retail ecommerce sales of $340.2 billion in the second quarter of 2026, 17.1% of all retail sales and up 12.2% from a year earlier. Over the 2025 holiday season, [Adobe](https://news.adobe.com/news/2026/01/adobe-holiday-shopping-season) found electronics ($59.8 billion), apparel ($49.0 billion) and furniture ($31.1 billion) drove more than half of online spending.

What one AI-assisted sale is worth depends on your channel mix. A sale on your own store keeps the full margin and the customer’s email. A sale on a marketplace carries fees but may be the shopper’s default. We infer that an ecommerce brand should value AI visibility in two layers: being named for the category question, and being shown with the right seller once the shopper picks the product. Brands that sell mainly through their own store should also read [our guide for DTC brands](https://underneath.agency/resources/dtc-brands-ai-search-sales).

## Where do AI shopping assistants sit in the path to purchase today?

They sit at the research step, between noticing a need and choosing where to buy.

Adobe calls generative AI chat services and browsers “an integral tool for consumers to find deals and research products.” In its 2025 holiday data, AI services were used most for video games, toys, appliances, electronics and personal care products. In a survey of more than 5,000 US consumers reported by [TechCrunch](https://techcrunch.com/2026/04/16/ai-traffic-to-us-retailers-rose-393-in-q1-and-its-boosting-their-revenue-too/), 39% said they used AI for online shopping and AI traffic in the first quarter of 2026 rose 393% from a year earlier. Electronics shoppers who compare specs are covered in [our guide to consumer electronics and AI search](https://underneath.agency/resources/consumer-electronics-sales-from-ai-search).

[Salesforce’s holiday data](https://www.salesforce.com/news/stories/2025-holiday-shopping-data/), drawn from more than 1.5 billion shoppers, put a wider number on it: AI and agents influenced $262 billion in holiday sales, 20% of all retail sales. That figure includes retailers’ own recommendation and service tools, not only ChatGPT or Perplexity. For third-party assistants specifically, Salesforce says the share of traffic from AI search channels doubled year over year and those visitors converted nine times more often than shoppers referred by social media.

Shoppers still want to make the final call themselves. [Shopify’s 2026 holiday research](https://www.shopify.com/news/agentic-holiday-2026) found 65% of shoppers say they will use AI for at least one shopping task this season, yet only one-third said they trusted an AI agent to buy on their behalf. Today the assistant shortlists; the shopper decides.

The assistants now come from four directions:

- **Chat assistants.** ChatGPT shows product carousels, merchant lists and review summaries, and its help page still describes an [Instant Checkout](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search) option for some eligible merchants. OpenAI scaled that feature back in March 2026, according to [FashionUnited](https://fashionunited.com/news/retail/us-consumers-would-let-ai-agents-buy-clothes-but-not-without-a-say/2026092974919).
- **Search.** Google says its [Shopping Graph](https://blog.google/company-news/inside-google/message-ceo/nrf-2026-remarks/) holds more than 50 billion product listings, with more than 2 billion refreshed every hour, and feeds AI Mode and the Gemini app. Its Universal Commerce Protocol, built with Shopify, Etsy, Wayfair, Target and Walmart, adds checkout inside those surfaces.
- **Marketplace assistants.** Amazon’s Rufus answers questions inside the store. Amazon says customers who use it are 60% more likely to complete a purchase, and its monthly active users rose 115% year over year.
- **Browser agents.** Perplexity’s Comet browser can shop on a user’s behalf. In August 2026 a federal appeals court [lifted an order](https://www.illawarramercury.com.au/story/9324130/amazon-loses-ban-on-perplexitys-ai-shopping-tools/) that had stopped it from shopping on Amazon, though Amazon’s lawsuit continues.

## Which questions do shoppers ask AI before they buy?

They ask for a shortlist, a comparison, a verdict on quality and the best place to buy.

The prompts below are illustrative, written to show the stages of a purchase; they are not observed data.

| Stage | Example prompt (illustrative) | What the brand needs to be true |
|---|---|---|
| Need | “Best cordless vacuum for pet hair under $300” | Product facts match the need and the price band |
| Comparison | “Shark vs Dyson for a small apartment” | Specs and reviews are easy to compare |
| Trust | “Is this brand good quality or just marketing?” | Independent reviews and press support the claims |
| Gift | “Present for an 8-year-old who loves sharks” | Use case and audience are written down |
| Where to buy | “Cheapest place to buy this blender” | Price and stock are current on every channel |

Shopify’s own example shows why descriptive prompts matter. A shopper describes a need in plain words, and the assistant matches it to a product whose page says who it is for. One apparel brand Shopify profiled, Cottonique, says its customers with eczema never searched for the word “hypoallergenic”; they described the problem. Drinks are bought the same way, by need, as [our guide for beverage brands](https://underneath.agency/resources/beverage-brands-ai-recommendations) explains.

## How does an AI answer turn into a sale, and where does the sale land?

The answer creates a shortlist, then the assistant or the shopper picks a seller.

The path for an ecommerce brand runs: a shopping question, a product named in the answer, a click to a merchant or an in-chat checkout, and later a review that feeds the next answer. The second step is where multichannel brands gain or lose margin. [OpenAI says](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search) that after a shopper clicks a product, ChatGPT may list the merchants that sell it, ranked on “availability, price, quality, and whether they are the maker or primary seller of that item.” That is documented by the platform. It means an out-of-stock brand store, or a price that differs across retailers, can hand the sale to someone else even after the brand has won the recommendation.

The quality of the visit is the reason to care. Adobe found AI-referred shoppers in March 2026 spent 48% longer on site and earned 37% more revenue per visit than other traffic. Back in March 2025, the same comparison ran the other way: AI visitors converted 38% worse. We infer the shoppers arriving from assistants have already done their comparison and come ready to buy.

Inside a marketplace the mechanism changes. Rufus answers on Amazon’s own pages, so the brand’s listing, images, questions and reviews on Amazon are what it has to work with; your own site does not enter the picture. Amazon also now sells “sponsored prompts” inside Rufus. Andy Jassy said nearly 20% of shoppers who interact with a brand prompt continue the conversation about that brand. Paid placement is arriving in AI shopping, alongside the organic answer.

## What decides which products and sellers an assistant shows?

Platforms document a few inputs; studies observe more; the rest is inference.

**Documented by the platform.** OpenAI says ChatGPT considers “structured metadata from first-party and third-party providers (e.g., price, product description) and other third-party content,” and that its review summaries are drawn from reviews on public websites. It says product results “are not ads, nor influenced by any OpenAI partnerships.” Shopify merchants’ data reaches ChatGPT through Shopify Catalog; other merchants can apply for a direct product feed. Google says the Shopping Graph covers inventory, prices and reviews.

**Observed in studies.** In the [September 2026 audit](https://arxiv.org/abs/2609.18729) of 117 product questions, editorial and product-review sites made up 56.7% of the websites ChatGPT showed. In Google’s AI Overviews, retailers and marketplaces were 17.8% of sources and manufacturers and brands 16.1%; the authors call these categories exploratory. Repeating a question in ChatGPT produced a shifting set of sources: on average only 26.0% of domains overlapped between repeats. Our own [Reddit citation study](https://underneath.agency/research/ai-reddit-citations-study) found Google’s AI cited Reddit in 10.3% of retail and ecommerce AI Overviews, far less than for software or home services.

**Our inference.** Each assistant reads from a different place: ChatGPT from feeds and review sites, Google from its Shopping Graph and the open web, Rufus from Amazon’s catalog and reviews. A reasonable expectation is that a product described the same way, with the same price, ratings and claims, across all of them is easier for each assistant to recognize and recommend. Research covered in [what drives AI product recommendations](https://underneath.agency/resources/what-drives-ai-product-recommendations) supports the idea that hard facts such as ratings and price weigh more than the brand name when assistants can see them.

## What does a brand lose when assistants show a rival instead?

It loses the shopper before they reach any store, and often to a competitor’s product, not just its listing.

The evidence on lost revenue is indirect, so we describe the risk rather than put a number on it. Three things are clear from the sources. First, AI shopping traffic converts better than average traffic, so missing it costs more than its share of visits suggests. Second, Shopify reports brands that worked on AI readiness seeing large gains: [Stanley 1913](https://www.shopify.com/news/agentic-holiday-2026) had five times as many AI-referred US orders as in the same window a year earlier, and ILIA’s AI-referred sales nearly tripled. Those are self-reported results from Shopify’s own customers, not controlled tests. Beauty is one category where a large retailer, Ulta, already reports on its AI visitors, as [our guide for beauty retailers](https://underneath.agency/resources/beauty-retailers-ai-search) shows.

Third, many retailers are not ready. Adobe found around 34% of product pages “can’t be properly accessed by AI,” and roughly a quarter of home page and category page content not yet optimized for AI systems. A brand whose product pages fall in that group leaves the description to whatever retailer page or review site the assistant can read instead. That source may carry an old price or a rival’s comparison.

Shopify’s survey adds a competitive note: 71% of regular AI users said they were more likely to choose an independent brand this year than last. Assistants can surface smaller brands that a keyword search would have buried. For an established brand, we infer that shelf space it took for granted in search is now contested.

## How does GEO work for an ecommerce brand across retailers and marketplaces?

GEO makes your products easy for every assistant to find, describe accurately and trust, without promising placement.

Generative engine optimization (GEO) for a multichannel brand covers six areas:

1. **One set of product facts.** Names, specs, sizes, materials, prices and claims should match across your store, Amazon, Walmart, other retailers and every product feed. Inconsistency gives assistants reasons to hesitate or to pick a cleaner listing. For nursery gear, that includes safety details, as [our guide for baby product brands](https://underneath.agency/resources/baby-product-brands-ai-recommendations) explains.
2. **Feeds and readable pages.** Keep Google Merchant Center and any ChatGPT product feed current; make sure product pages render their key facts in plain text an AI system can read. Adobe’s 34% figure shows how common this gap is.
3. **Marketplace listings as AI content.** On Amazon, Rufus answers from your listing, images and reviews. Write the listing to answer the questions shoppers ask, such as fit, compatibility and care, rather than to stuff keywords.
4. **Independent coverage.** Chat assistants lean on editorial review sites. Getting products tested and reviewed by credible publications, and making review samples and specs easy to obtain, is digital PR aimed at the sources assistants cite.
5. **Reviews that say something.** Assistants summarize public reviews. Encouraging detailed, honest reviews on your site and on retailer pages gives them substance to quote. Fake or incentivized reviews are a risk, which we cover in [fake reviews and AI recommendations](https://underneath.agency/resources/fake-reviews-fake-brands-ai-recommendations).
6. **Measurement per assistant.** Check your priority category questions in ChatGPT, Google AI Mode, Gemini, Perplexity and Rufus separately, repeat each question several times, and note which seller gets the click.

The product copy itself matters too; [what product content AI shopping assistants prefer](https://underneath.agency/resources/product-content-ai-shopping-assistants-prefer) covers it in depth.

## What can’t multichannel sellers learn yet from the AI shopping data?

It does not show how much of a brand’s sales AI answers cause, or how each assistant weighs channels.

Adobe and Salesforce measure clicks from AI tools to retail sites. They cannot see a shopper who reads a ChatGPT answer and then opens the Amazon app to buy, so AI influence is probably undercounted on marketplaces and overstated where checkout happens elsewhere. Salesforce’s $262 billion includes retailers’ own recommendation engines, not only AI search. Amazon’s Rufus figures are the company’s own and not independently audited.

No platform has published how it weighs a brand’s own store against a retailer selling the same item beyond the factors OpenAI lists. In-chat checkout is a moving target: [Reuters reported](https://www.zawya.com/en/business/americas/retailers-tap-ai-shopping-traffic-but-fight-to-keep-customer-data-424957) that OpenAI ended Instant Checkout in March 2026, while FashionUnited said it was scaled back. Sponsored prompts and the legal status of browser agents are also changing during 2026. And the product audit covered 117 questions in the Netherlands; source patterns in the US may differ.

## Where should an ecommerce brand start?

Start by checking how assistants answer the ten to twenty questions that drive your top categories.

Pick the categories that carry your revenue, write the questions a shopper would ask at each stage, and run them in each assistant, including Rufus if you sell on Amazon. Note whether your products appear, how they are described, which sources are cited and which seller gets the click. Then compare product facts across your store and your three largest retail partners. The gaps you find usually point to the first fixes: feeds, unreadable pages, thin listings or missing review coverage.

For a structured view across your store and retail partners, [contact our team](https://underneath.agency/contact). We will map how AI shopping assistants describe and route your products today and set out the steps most likely to improve visibility where your sell-through happens, whether that is your own store or a retail partner. Our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page explains how that work is carried out across channels, from matching product facts and feeds to marketplace listings and per-assistant checks.

## Frequently asked questions

### Do AI shopping assistants send shoppers to Amazon or to brand websites?

Both. OpenAI says ChatGPT ranks the merchants for a product on availability, price, quality and whether the seller is the maker or primary seller, so a brand store can be shown first when it is competitive on those points.

### Does Amazon’s Rufus use my own website?

Rufus answers inside Amazon, so it works from your Amazon listing, images, questions and reviews. Your own site matters for ChatGPT, Google and Perplexity.

### Can we pay to appear in AI shopping answers?

OpenAI says ChatGPT’s product results are not ads. Amazon sells sponsored prompts inside Rufus, and paid formats are spreading, so check each platform’s current rules.

### How do we know whether AI shopping is driving our sales?

Track AI referrals and their conversion in analytics, ask buyers how they found you, and check regularly which of your products assistants name for your category questions.

## Sources

- U.S. Census Bureau (2026-08-18), [Quarterly Retail E-Commerce Sales, 2nd Quarter 2026](https://www.census.gov/retail/mrts/www/data/pdf/ec_current.pdf)
- Adobe (2026-01-07), [Holiday Shopping Season Drove a Record $257.8 Billion Online with Consumers Embracing Generative AI Tools](https://news.adobe.com/news/2026/01/adobe-holiday-shopping-season)
- TechCrunch (2026-04-16), [AI traffic to US retailers rose 393% in Q1, and it’s boosting their revenue too](https://techcrunch.com/2026/04/16/ai-traffic-to-us-retailers-rose-393-in-q1-and-its-boosting-their-revenue-too/)
- Salesforce (2026-01), [AI and agents account for $262 billion of 2025 holiday spend](https://www.salesforce.com/news/stories/2025-holiday-shopping-data/)
- Modern Retail (2026-04-30), [Amazon Rufus users are up 115%, while engagement is up 400%](https://www.modernretail.co/?p=159156)
- FashionUnited (September 29, 2026), [US consumers would let AI agents buy clothes, but not without a say](https://fashionunited.com/news/retail/us-consumers-would-let-ai-agents-buy-clothes-but-not-without-a-say/2026092974919)
- Reuters, via Zawya (2026), [Retailers tap AI shopping traffic but fight to keep customer data](https://www.zawya.com/en/business/americas/retailers-tap-ai-shopping-traffic-but-fight-to-keep-customer-data-424957)
- OpenAI (2026), [Shopping with ChatGPT Search](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search)
- Google (2026-01-11), [The AI platform shift and the opportunity ahead for retail](https://blog.google/company-news/inside-google/message-ceo/nrf-2026-remarks/)
- Shopify (2026-10-06), [Welcome to the first holiday season of the agentic era](https://www.shopify.com/news/agentic-holiday-2026)
- Reuters via Illawarra Mercury (2026-08-05), [Amazon loses ban on Perplexity’s AI shopping tools](https://www.illawarramercury.com.au/story/9324130/amazon-loses-ban-on-perplexitys-ai-shopping-tools/)
- Uberti-Bona Marin and colleagues (2026), ["If I Had to Buy Just ONE: Galaxy S26 Ultra": Auditing AI-Generated Product Recommendations](https://arxiv.org/abs/2609.18729)
- Underneath (2026), [When does a Reddit thread become evidence in Google’s AI?](https://underneath.agency/research/ai-reddit-citations-study)

---

This is the Markdown twin of https://underneath.agency/resources/ecommerce-brands-ai-shopping-assistants. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How do ecommerce platforms and apps win merchants who ask AI?"
description: "By being named when merchants and their agencies ask AI which platform or app to use, and by proving you can put merchants’ products in AI shopping."
canonical: "https://underneath.agency/resources/ecommerce-platform-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How do ecommerce platforms and apps win merchants who ask AI?

By being named, with accurate costs and capabilities, when merchants and their agencies ask AI which platform or app to use, and by showing that you can get those merchants’ own products into AI shopping. That second part is new: AI channels have become a feature merchants compare platforms on. For an ecommerce software company, each merchant won is worth a share of that merchant’s sales for years, so both parts feed revenue.

## The short version

1. Platform revenue rides on merchant sales: in the second quarter of 2026, [Shopify](https://www.shopify.com/investors/press-releases/shopify-delivers-big-30-growth-across-gmv-revenue-gross-profit) processed $115,567 million in merchandise volume and earned $2,781 million from merchant solutions such as payments, against $802 million from subscriptions.
2. [OpenAI documents](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search) that Shopify merchants’ product data already flows into ChatGPT through Shopify Catalog, with no additional work required from individual merchants. Getting products into AI shopping is now part of what a platform sells.
3. Merchants have reason to ask: [Adobe’s data](https://techcrunch.com/2026/04/16/ai-traffic-to-us-retailers-rose-393-in-q1-and-its-boosting-their-revenue-too/) shows AI traffic to US retail sites rose 393% year over year in the first quarter of 2026 and converted 42% better than other traffic in March.
4. Shopify reports merchant results to match: [Stanley 1913](https://www.shopify.com/news/agentic-holiday-2026) saw five times as many AI-referred US orders as a year earlier, and Cottonique’s AI-referred sales grew 276%.
5. Recommendations also travel through partners: [40,000 Shopify partners](https://www.shopify.com/news/a-world-with-more-builders) referred a new merchant last year, and the app ecosystem earned $1.3 billion.

## Who picks an ecommerce platform, and what is one merchant worth over time?

Founders and ecommerce leaders buy it, often guided by an agency, and a merchant’s value grows with its sales.

Who makes the call depends on how big the merchant is. A new or small merchant usually picks a platform alone, on a free trial. A mid-market brand’s ecommerce director chooses with an agency that will build and run the store. An enterprise retailer runs a formal replatforming project with its chief technology officer, finance, a systems integrator and a request for proposals; enterprise suites such as [Salesforce Commerce Cloud](https://www.salesforce.com/commerce/pricing/) list every edition as “contact for pricing,” which tells you the sale is negotiated. Below the platform sits a second market: the apps merchants add for reviews, subscriptions, shipping, email and more, bought one install at a time inside app stores.

What a merchant is worth depends on how much it sells. Shopify’s second-quarter numbers show the shape of the model: $3,583 million of revenue, of which merchant solutions, which grow with the merchant’s sales, were by far the larger part. Monthly recurring revenue from subscriptions was $221 million. Even merchants on [Shopify Plus](https://www.shopify.com/plus/pricing) who use a third-party payment processor pay 0.20% per transaction to Shopify. So a platform that wins a fast-growing brand earns more each year that brand grows, and replatforming is rare enough that we infer most of that value is locked in at the original choice.

The app economy works the same way at a smaller scale. Shopify says partners in around 100 countries support millions of merchants, and app developers earned on top of the $1.3 billion the app ecosystem earned last year.

## How do merchants choose a platform today, and where does AI come in?

Through agencies, reviews and peers, and increasingly through AI assistants and AI shopping features.

Agencies and other partners are a major route. Shopify’s count of 40,000 partners referring a new merchant in one year shows how much platform choice is advice from a trusted builder. Reviews matter too: in [Gartner Digital Markets’ survey](https://cdn-static.bizzabo.com/bizzabo.file.upload/yVkIqwmWRDKbEhqeE3aQ_Gartner-Digital-Markets_2025-Report_Making-The-List.pdf) of 3,500 software buyers across industries, specialty retailers were among the buyers who rely most on customer reviews (49%), and small enterprises used ChatGPT or other generative AI tools to build shortlists more often than larger firms (31%).

AI also enters the decision as a requirement, not only as a research tool. Merchants now see AI shopping as a sales channel. Adobe, analyzing more than 1 trillion visits to US retail sites, found that AI-driven revenue per visit was 37% higher than non-AI traffic in March 2026, reversing the pattern of a year earlier. Shopify’s 2026 holiday report says 65% of shoppers plan to use AI for at least one shopping task this season. A merchant comparing platforms in 2026 therefore asks not only “which platform is cheapest to run?” but “which platform gets my products into ChatGPT, Google and other AI shopping surfaces?”

## Which questions do merchants and agencies ask AI assistants?

Questions about switching, total cost, AI shopping readiness, B2B features, integrations and apps.

We wrote the prompts below to illustrate merchant and agency platform questions; none come from real buyers:

- Switching: “Best platform to migrate to from Magento 2 for a $20M apparel brand with 3,000 SKUs.”
- Total cost: “Shopify vs WooCommerce total cost for a store doing $2M a year, including payment fees.”
- AI shopping: “Which ecommerce platforms let my products appear in ChatGPT shopping results?”
- B2B: “Best ecommerce platform for wholesale with customer-specific price lists and net terms.”
- Integrations: “Which headless commerce platforms integrate with NetSuite and our warehouse system?”
- Apps: “Best subscription app for a Shopify coffee brand, and how do the fees compare?”

Agencies ask the same questions on their clients’ behalf, which means a single answer can shape several replatforming decisions. That is our inference from how agencies work, not a measured effect.

## How does a platform recommendation become years of merchant revenue?

Through a trial or an agency brief, then a launch, then years of transaction-linked revenue.

**Small merchants.** The assistant names a few platforms; the merchant starts a trial on one, builds a store and begins paying a subscription. Payments and other merchant services follow as soon as orders do.

**Mid-market and enterprise brands.** The assistant shapes the shortlist that a brand and its agency take into a replatforming evaluation. We infer that these deals rarely show up as an AI referral: they arrive as a request for proposals, an agency introduction or a direct inquiry from a team that already knows your name.

**Apps.** A merchant asks for the best tool for a job, then installs the app named from the platform’s app store. The AI answer and the app store listing work together, so both need to tell the same story.

**Growth.** The platform’s revenue then rises with the merchant’s sales. AI shopping makes that loop tighter: a platform that helps its merchants sell through AI channels, as Shopify Catalog is designed to do, earns more per merchant if those channels grow. How AI answers bring merchants their first-time buyers is covered in [our guide for DTC brands](https://underneath.agency/resources/dtc-brands-ai-search-sales).

## Why does an assistant suggest one commerce platform or app over another?

Platforms document how they handle products, not platforms; for software choices, the evidence is studies and inference.

**Documented by the platform.** OpenAI says that when ChatGPT lists merchants for a product, “merchants are ranked based on factors like availability, price, quality, and whether they are the maker or primary seller.” That governs your merchants’ visibility, not yours, but it is why AI-shopping integrations are a selling point. Google says AI Mode draws on [“shopping data for billions of products”](https://blog.google/products/search/ai-mode-search/) and uses a “query fan-out” technique, issuing related searches across subtopics and combining the results, so a question about choosing a platform can pull in pricing pages, reviews and comparisons at once.

**Observed in our studies.** In [our study of self-ranking “best of” lists](https://underneath.agency/research/self-promoting-best-lists-study), 24.2% of AI-cited numbered lists with an identifiable publisher ranked the publisher first; Klaviyo, an ecommerce marketing platform, was among the repeat publishers. Our [llms.txt adoption study](https://underneath.agency/research/llms-txt-adoption-study) found 11.5% of top websites publish a valid file, and notes that an outside tracker attributes much of the recent growth to Shopify enabling it across its platform. That is a platform deciding, for all its merchants at once, how readable their stores are to AI systems.

**Observed in industry data.** Adobe found that around 34% of retailers’ product pages cannot be properly accessed by AI. A reasonable expectation is that merchants will increasingly ask which platforms fix this for them, and that platforms able to show it will be described that way in answers.

**Trust factors specific to ecommerce software.** Merchants weigh total cost including payment fees, migration effort, uptime during peak season, app and agency ecosystems, B2B and international features, and now AI-channel support. Shopify’s own merchant story makes the point about independent sources: a Stanley 1913 executive said agentic optimization means positioning “our brand and product story in other authoritative sources,” not only on the brand’s own site. The same applies to the platforms themselves. It also applies to the merchants on those platforms: an outdoor gear brand, for example, earns its place through expert reviews and retailer guides, as [our guide for outdoor brands](https://underneath.agency/resources/outdoor-brands-ai-search) explains.

## What does a commerce platform give up when merchants never hear its name from AI?

Years of a merchant’s transaction revenue, decided at one moment.

Because platform revenue grows with merchant sales, a merchant lost at the selection stage is not one lost subscription; it is a share of that merchant’s sales for as long as it stays on a rival platform. At Shopify the transaction-linked part of revenue was more than three times the subscription part last quarter. We infer that the cost of being absent from AI answers falls hardest on platforms selling to fast-growing brands, and on apps competing for the few slots an assistant lists for a job. For the broader effect of fewer search clicks, see [our article on what happens if you skip GEO](https://underneath.agency/resources/what-happens-if-you-skip-geo).

## What does GEO cover for a commerce platform or app developer?

It makes your platform or app easy to describe correctly and easy to verify; it cannot guarantee a recommendation.

For platforms and app makers selling to merchants, generative engine optimization (GEO) usually comes down to six pieces of work:

1. **Clear total cost.** Publish plan prices, payment fees, transaction fees and typical app costs in one current place, so answers comparing platforms do not rely on old figures. If an assistant already quotes an outdated fee, [how to fix wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) explains the repair.
2. **Documented AI-channel support.** State plainly which AI shopping surfaces your merchants’ products can reach, how product data is structured and what merchants must do. If it is automatic, say so in public documentation.
3. **Migration and comparison pages.** Merchants and agencies ask switching questions, so publish fair migration guides and comparisons by source platform. Whether migration and comparison pages earn citations is weighed in our article on [comparison pages](https://underneath.agency/resources/do-comparison-pages-help-b2b-ai-citations).
4. **Agency and partner proof.** Make partner directories, case studies and certified agency lists public and current, since agencies both influence the choice and write about it.
5. **Reviews and independent coverage.** Keep review profiles, app store listings and coverage in ecommerce trade media current. Our article on [which “best of” lists matter](https://underneath.agency/resources/best-of-lists-ai-recommendations) helps pick targets; [how product content helps AI shopping assistants](https://underneath.agency/resources/product-content-ai-shopping-assistants-prefer) covers the merchant side.
6. **Measurement by buyer type.** Track merchant, agency and enterprise questions separately across ChatGPT, Gemini, Perplexity, Copilot and Google, and connect them to trials, app installs, agency referrals and enterprise inquiries.

## What is still unproven about AI’s role in choosing an ecommerce platform?

No one has measured how often AI assistants decide which ecommerce platform a merchant chooses.

The strongest numbers here describe shoppers and merchants’ sales, not platform selection: Adobe’s traffic data, Shopify’s holiday report and the brand results Shopify publishes about its own merchants. The Gartner Digital Markets survey covers software buyers across industries. Merchant case results are self-reported and selected by the platform. Our own studies record which pages AI answers cite; they do not follow merchants into a trial or a replatforming. Treat AI as a growing influence on platform choice, and a proven sales channel for merchants, with its effect on platform revenue still unmeasured.

## What should a commerce platform or app check before chasing more trials and installs?

Check what assistants currently say about your fees, migration paths and AI-shopping support.

A practical first step is an audit of the switching, cost, AI-readiness and app questions merchants and agencies ask, across the main assistants and Google’s AI features, matched against your trials, installs and agency-sourced deals so gaps are ranked by the merchant revenue at stake. To have us build that view with you and plan the fixes, [ask us for a merchant-question audit](https://underneath.agency/contact). Our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page covers how the follow-on work is handled for a platform or app, from cost pages and AI-channel documentation to migration guides and partner proof.

## Frequently asked questions

### Do merchants really use AI to choose an ecommerce platform?

Some do, and smaller firms appear to lean on AI tools most when shortlisting software. No published study measures platform choice specifically, so treat it as likely and growing rather than measured.

### Why does AI shopping support matter to platform sales?

Because merchants now see AI as a sales channel and compare platforms on it. OpenAI documents that Shopify merchants’ product data reaches ChatGPT automatically, so rival platforms are compared against that.

### Should ecommerce apps worry about AI search, or only about the app store?

Both. Merchants ask assistants for the best app for a job and then install from the app store, so the answer and the listing need to agree on what the app does and costs.

### How do we measure AI’s effect on merchant acquisition?

Track AI referrals to trial and pricing pages, ask new merchants and agencies how they heard of you, and check regularly how often assistants name you for merchant questions. Then follow those merchants’ processed volume over time.

## Sources

- Shopify (2026), [Shopify Delivers Big: 30%+ Growth Across GMV, Revenue, Gross Profit, and Free Cash Flow](https://www.shopify.com/investors/press-releases/shopify-delivers-big-30-growth-across-gmv-revenue-gross-profit)
- Shopify (2026), [Welcome to the first holiday season of the agentic era](https://www.shopify.com/news/agentic-holiday-2026)
- Shopify (2026), [The next era of commerce, built by partners](https://www.shopify.com/news/a-world-with-more-builders)
- Shopify (2026), [Shopify Plus pricing](https://www.shopify.com/plus/pricing)
- OpenAI (2026), [Shopping with ChatGPT Search](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search)
- TechCrunch (2026), [AI traffic to US retailers rose 393% in Q1, and it’s boosting their revenue too](https://techcrunch.com/2026/04/16/ai-traffic-to-us-retailers-rose-393-in-q1-and-its-boosting-their-revenue-too/)
- Gartner Digital Markets (2025), [Making the List: 2025 software buying report](https://cdn-static.bizzabo.com/bizzabo.file.upload/yVkIqwmWRDKbEhqeE3aQ_Gartner-Digital-Markets_2025-Report_Making-The-List.pdf)
- Salesforce (2026), [Commerce Cloud pricing](https://www.salesforce.com/commerce/pricing/)
- Google (2025), [Expanding AI Overviews and introducing AI Mode](https://blog.google/products/search/ai-mode-search/)
- Underneath (2026), [self-ranking “best of” lists](https://underneath.agency/research/self-promoting-best-lists-study) and [llms.txt adoption](https://underneath.agency/research/llms-txt-adoption-study)

---

This is the Markdown twin of https://underneath.agency/resources/ecommerce-platform-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Can AI search win EdTech companies more schools and families?"
description: "Partly. Teachers, students and families already ask AI; districts still buy on evidence, privacy and state lists, so AI answers must find that proof."
canonical: "https://underneath.agency/resources/edtech-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Can AI search bring an EdTech company more schools, teachers and families?

Partly, and the path differs for each buyer. Teachers, students and families already ask AI tools for help, so being named in those answers can start adoption. Districts and universities still buy through committees, evidence reviews, privacy checks and, more and more, state approval lists, so an AI answer only helps if it can find that proof and repeat it accurately.

## The short version

1. Districts are drowning in tools: they have access to an average of 3,001 digital tools, yet students and teachers typically use about four learning apps inside their learning platform, according to [Instructure data reported by EdWeek Market Brief](https://marketbrief.edweek.org/education-market/just-because-an-ed-tech-product-is-accessible-doesnt-mean-its-being-used-heres-why/2026/07).
2. The people who use EdTech already use AI: six in 10 US public school teachers used an AI tool for their work in 2024-25 ([Gallup and the Walton Family Foundation](https://news.gallup.com/poll/691967/three-teachers-weekly-saving-six-weeks-year.aspx)), and 42% of college students use generative AI weekly ([Tyton Partners](https://tytonpartners.com/time-for-class-2025/)).
3. State lists now steer curriculum buying: at least 18 states publish approved materials lists, and 76% of school and district leaders say those lists influence elementary reading purchases “a lot” ([EdWeek Market Brief](https://marketbrief.edweek.org/regulation-policy/how-widely-influential-are-state-approved-materials-lists-k-12-officials-weigh-in/2026/09)).
4. Guardrails are tightening: 79% of districts now have AI guidelines, up from 57% in 2025 ([CoSN](https://www.cosn.org/edtech-topics/state-of-edtech-leadership/)), and New York City banned student-facing generative AI through eighth grade for one year.
5. AI search can also take traffic away: the homework help company Chegg sued Google in 2025, claiming AI Overviews hurt its traffic and revenue ([The Verge](https://www.theverge.com/news/619051/chegg-google-ai-overviews-monopoly)).

## Who buys education software, and how?

Four different buyers: districts, teachers, colleges and families, each with its own path to a purchase.

**K-12 districts.** The US has 98,577 public schools, according to [federal data](https://nces.ed.gov/fastfacts/display.asp?id=84), grouped into districts that buy through budgets, requests for proposals and board approval. Influence is spread out. In an EdWeek Research Center survey on [who influences purchases](https://marketbrief.edweek.org/meeting-district-needs/security-vendors-are-flooding-k-12-which-district-officials-have-influence-on-what-gets-bought/2026/09), superintendents were named by 63% of district and school leaders and school boards by 46%; 31% pointed to finance leaders and 29% to technology leaders. That survey was about security purchases, but the chain it describes, from vetting to the business office to the board, is typical of large district buys.

**Teachers.** Teachers often find a tool first and bring it to the district. Their influence shows in usage: across 12.6 million K-12 users of the Canvas platform, Instructure counted 618,000 educators launching learning apps in 2025-26.

**Colleges and universities.** Instructors and administrators choose course tools and platforms. Tyton’s 2025 study drew on more than 3,300 students, instructors and administrators across more than 900 colleges.

**Families and students.** Parents and learners buy subscriptions directly: tutoring, test prep, language learning, homework help. This is the closest EdTech gets to consumer marketing, and the most exposed to AI answers.

What a customer is worth varies widely, and public benchmarks are thin. What the data do show is pressure on budgets. The public school population fell 2% over the past decade, and federal projections suggest a possible 5% drop by 2031, as [EdWeek Market Brief reports](https://marketbrief.edweek.org/strategy-operations/as-k-12-shrinks-heres-how-companies-are-looking-for-growth/2026/09). In its survey, 55% of district and school leaders said enrollment had dropped in the past five years. Fewer students means fewer seats to sell and harder questions at renewal.

## Where do teachers, students and district buyers already use AI?

With the end users first: teachers, students and families use AI daily; district buyers’ use for vendor research is unmeasured.

- **Teachers.** In the Gallup and Walton survey of 2,232 public school teachers, 60% used AI tools for their work, and 32% did so at least weekly.
- **Teens.** [Pew Research Center](https://www.pewresearch.org/short-reads/2025/01/15/about-a-quarter-of-us-teens-have-used-chatgpt-for-schoolwork-double-the-share-in-2023/) found 26% of US teens had used ChatGPT for schoolwork, up from 13% in 2023, and 54% said it is acceptable to use ChatGPT to research new topics.
- **Colleges.** Tyton found weekly generative AI use among 42% of students and 30% of instructors.
- **District technology leaders.** In [CoSN’s 2026 survey](https://www.cosn.org/edtech-topics/state-of-edtech-leadership/), 96% of education technology leaders said AI could positively affect education, and 64% of districts were using AI in operations, up from 37% the year before.

None of these surveys asks whether district buyers use ChatGPT, Gemini or Google’s AI features to find vendors. For vendor research, the nearest figure comes from outside education: among software buyers of every kind in [G2’s 2026 survey](https://company.g2.com/news/g2-research-the-answer-economy), 51% started research with an AI chatbot more often than with Google. G2 sells review visibility and did not survey schools.

We infer that AI answers matter most at two points: when a teacher, parent or student looks for a tool to solve a problem today, and when a district team does its first pass over a crowded category before writing a request for proposals.

## Which questions do educators, parents and students ask?

Questions about fit, evidence, privacy, price and alternatives. We wrote the classroom and household prompts below ourselves to show the pattern; none was captured from a real teacher, parent or student.

| Buyer | Illustrative prompt |
|---|---|
| Teacher | “What is a free tool to give feedback on student writing that works in Google Classroom?” |
| Curriculum director | “Which K-5 reading programs are on the Texas approved list and align with the science of reading?” |
| District technology leader | “Which math platforms have ESSA evidence and a signed student data privacy agreement in our state?” |
| University administrator | “What are alternatives to our current proctoring software that faculty and students accept?” |
| Parent | “Is this online tutoring service worth it for a 7th grader struggling with algebra?” |
| Student | “What’s the best app to practice Spanish for the AP exam?” |

These questions mix discovery with trust. A parent wants to know whether a tutoring service is safe and worth the money; a district wants proof of impact and a privacy agreement before anything else.

A parent’s single question about a tutoring service can set off several searches behind the answer. For AI Overviews and AI Mode, Google describes a [“query fan-out” technique](https://developers.google.com/search/docs/appearance/ai-features) that may issue multiple related searches; on OpenAI’s side, the help center says [ChatGPT search typically rewrites a question](https://help.openai.com/en/articles/9237897-chatgpt-search) into one or more targeted queries. Those hidden searches lean on outside judgment, which is where an EdTech vendor’s reviews and awards come in: in [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study), 46.2% of ChatGPT’s answers involved a search for reviews or ratings, and 43.8% targeted a named publication, ranking or award.

## How does being named by AI become subscriptions, classroom adoptions and district contracts?

By different routes for each buyer: subscriptions for families, teacher adoption for schools, and shortlists for district contracts.

1. **Families and learners.** An AI answer names a tutoring service or app; the parent or student tries it; a subscription follows. This is the shortest path, and the one most like consumer software.
2. **Teacher-led adoption.** A teacher asks for a tool, tries the free version, and uses it with a class. If enough teachers do the same, the district buys a license. Because students and teachers use only about four apps in their learning platform, the prize is becoming one of those few.
3. **District and state contracts.** A curriculum or technology team narrows the field, checks state approval lists, evidence and privacy agreements, then issues a request for proposals or buys from an approved list. An AI answer can put a vendor on the first list; the evidence decides the rest.
4. **Higher education.** Faculty and administrators compare options; the institution signs a campus license.

The more formal the purchase, the less the AI answer decides on its own. State approval matters most in core subjects: 62% of leaders said state lists influence elementary math purchases a lot, against 59% for secondary reading, while just 17% said the lists don’t influence secondary reading decisions at all.

## Why would an assistant recommend one learning tool over another?

No platform documents how it chooses products; studies point to independent sources, and education buyers reward evidence and privacy.

**What Google and OpenAI say.** Both companies describe AI answers that search the web and cite sources; neither explains how a tutoring app or classroom tool gets picked.

**What research has measured.** Earned sites such as reviews and independent publications supplied 72.7% of AI search sources for US software questions, [Chen and colleagues](https://arxiv.org/abs/2509.08919) found. Repeat runs change the list, too: across five runs of the same question in [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), only 25.2% of the brands ChatGPT named appeared every time. Neither study covered schools, colleges or family learning products.

**What education buyers check.** These trust factors are specific to the sector and are public, which matters because an assistant can read them too:

- **Evidence of impact.** More than half of the tools in Instructure’s 2025-26 top 40 had documented proof of impact, which Instructure says is 13 points higher than the broader EdTech market. EdWeek Market Brief adds that districts are “demanding stronger evidence behind product claims.”
- **Privacy credentials.** Nearly half of the top 40 tools hold third-party privacy certifications from groups such as 1EdTech or iKeepSafe. The [Federal Trade Commission’s](https://www.ftc.gov/news-events/news/press-releases/2025/01/ftc-finalizes-changes-childrens-privacy-rule-limiting-companies-ability-monetize-kids-data) children’s privacy rule requires verifiable parental consent before collecting personal information from children under 13, and was updated in 2025.
- **Security track record.** On December 28, 2024, PowerSchool, a major student information system vendor, learned that personal data had been taken from some customers’ systems through a support portal accessed with [a single compromised credential](https://www.powerschool.com/security/sis-incident/). CoSN found 65% of districts name insufficient cybersecurity staffing and the lack of a dedicated budget as top barriers. A breach becomes part of what any search, human or AI, finds about a vendor.
- **Policy fit.** Seven states passed laws or directives on devices and digital tools in classrooms this year, [EdWeek Market Brief reports](https://marketbrief.edweek.org/regulation-policy/nycs-ai-ban-and-screen-time-rules-puts-new-pressure-on-ed-tech/2026/09), and New York City’s ban on student-facing generative AI affects nearly 600,000 students.

**Our inference.** The proof districts verify (evidence ratings, privacy agreements, state approvals, accessibility statements) is published by third parties and is easy to cite. A reasonable expectation is that vendors whose proof is public, consistent and current give AI answers more to work with. No published study has tested it for school or family learning products.

## What does it cost an EdTech company to be missing, or to be summarized?

Lost trials and shortlist places, and for content-heavy businesses, lost traffic.

- **The first list in a crowded market.** With districts able to reach about 3,001 tools, a product that is not named early may never be compared.
- **Teacher discovery.** If teachers ask an assistant for a tool and a competitor is named, the bottom-up route to a district license narrows, we infer.
- **Content businesses can lose the click.** Chegg sued Google in February 2025, claiming AI Overviews used its content and hurt its traffic and revenue; Google said AI Overviews send traffic to a greater diversity of sites. The case shows the risk for EdTech companies whose growth depends on free answers to homework questions. We look at this risk in [which content businesses AI search threatens](https://underneath.agency/resources/content-business-risk-from-ai-search) and [whether AI Overviews reduce clicks](https://underneath.agency/resources/do-ai-overviews-reduce-clicks).
- **Wrong descriptions.** An outdated privacy or pricing description in an AI answer can stop a district review early, we infer. Our guide to [fixing wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) covers how to correct a stale description.

## What does GEO involve for a learning-product company?

It makes your evidence, privacy and fit easy for AI assistants to find and repeat, for each buyer. It cannot promise that a teacher, parent or district will see you named.

1. **One clear identity per audience.** Say what the product does, for which grades or courses, and which platforms it works in, the same way on your site, app marketplace listings and review profiles.
2. **Public proof pages.** Publish ESSA evidence and research summaries, privacy certifications, signed state data privacy agreements, accessibility statements and state approvals on plain web pages, not only in sales packets.
3. **Independent coverage.** Earn mentions in education trade press, research clearinghouses and educator communities. The outside-coverage work is the same as in other sectors, described in [how brands build authority for AI search](https://underneath.agency/resources/how-brands-build-authority-for-ai-search).
4. **Answer real questions honestly.** Write comparison, alternatives and “is it worth it” pages for parents, teachers and administrators, with real prices and limits. An independent “best tutoring apps” list tends to count for more than your own site, as our guide to [which pages to target](https://underneath.agency/resources/best-of-lists-ai-recommendations) shows.
5. **Reviews from real users.** Encourage teachers and families to review the product where they already look. Never plant reviews; see [legitimate GEO versus manipulation](https://underneath.agency/resources/legitimate-geo-vs-manipulation).
6. **Measure each audience separately.** Teacher, parent and district questions draw different answers. Track ChatGPT, Google AI Overviews and AI Mode, Gemini, Perplexity, Copilot and Claude; see [how many prompts to track](https://underneath.agency/resources/how-many-prompts-to-track-ai-visibility).

## What don’t we know yet about AI search in education buying?

We cannot yet say how often districts pick vendors through AI, or whether it lifts EdTech sales.

- **No survey of buyers’ AI search use in education.** The surveys above measure AI use for teaching and learning, not for choosing products.
- **Some sources are vendors.** Instructure sells a learning platform; G2 sells review visibility.
- **The Chegg case is unresolved.** A lawsuit is a claim, not a finding.
- **No link to seats or subscriptions has been shown.** Whether being named turns into EdTech revenue is still open; see [does AI visibility drive business results](https://underneath.agency/resources/does-ai-visibility-drive-business-results).

## What should an EdTech company check before the next adoption and renewal cycle?

Check how each assistant describes you to teachers, parents and districts, then publish the proof each buyer needs.

Draft the questions a teacher, parent, student, district leader and university buyer would ask, and put each one to ChatGPT, Gemini, Google’s AI features and the other main assistants more than once. Note whether your product is named, how your evidence, privacy terms and price are described, and which sources are cited. Where answers fall short, the cause is usually proof that lives only in sales packets, or a page that has gone stale.

We can run that check with you: [ask us to review how AI answers present you to schools and families](https://underneath.agency/contact). We will show where AI answers put you in front of teachers, families and district teams, and which gaps in your evidence and privacy proof are most likely costing you trials, adoptions and contracts. For the ongoing side, our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page describes how proof pages, educator reviews and separate checks for teacher, parent and district questions are kept current.

## Frequently asked questions

### Do teachers use ChatGPT to find classroom tools?

Many teachers use AI for their work: 60% did in 2024-25. No survey measures how often they use it to find products.

### Does being on a state approved list help AI visibility?

Untested for AI, but it matters to buyers: 76% of leaders say state lists influence elementary reading purchases a lot.

### Should we publish our privacy agreements and certifications?

Yes, at least the facts districts check before they sign. Nearly half of the top 40 tools in Instructure’s data hold third-party privacy certifications.

### Will AI search hurt our free content traffic?

It can. Chegg claims AI Overviews hurt its traffic; plan for answers that summarize your content without a click.

### Is GEO different for consumer and school products?

Yes. Families decide quickly from answers and reviews; districts need evidence, privacy agreements and approvals before a contract.

## Sources

- EdWeek Market Brief (2026-07), [Just Because an Ed-Tech Product Is Accessible Doesn’t Mean It’s Being Used](https://marketbrief.edweek.org/education-market/just-because-an-ed-tech-product-is-accessible-doesnt-mean-its-being-used-heres-why/2026/07)
- EdWeek Market Brief (2026-09), [How Widely Influential Are State-Approved Materials Lists? K-12 Officials Weigh In](https://marketbrief.edweek.org/regulation-policy/how-widely-influential-are-state-approved-materials-lists-k-12-officials-weigh-in/2026/09)
- EdWeek Market Brief (2026-09-18), [Security Vendors Are Flooding K-12. Which District Officials Have Influence on What Gets Bought?](https://marketbrief.edweek.org/meeting-district-needs/security-vendors-are-flooding-k-12-which-district-officials-have-influence-on-what-gets-bought/2026/09)
- EdWeek Market Brief (2026-09), [As K-12 Shrinks, Here’s How Companies Are Looking for Growth](https://marketbrief.edweek.org/strategy-operations/as-k-12-shrinks-heres-how-companies-are-looking-for-growth/2026/09)
- EdWeek Market Brief (2026-09), [NYC’s AI Ban and Screen-Time Rules Put New Pressure on Ed Tech](https://marketbrief.edweek.org/regulation-policy/nycs-ai-ban-and-screen-time-rules-puts-new-pressure-on-ed-tech/2026/09)
- Instructure (2026), [The Edtech Top 40: K-12 Edtech Engagement](https://www.instructure.com/edtech-top40)
- Gallup and Walton Family Foundation (2025-06), [Three in 10 Teachers Use AI Weekly, Saving Six Weeks a Year](https://news.gallup.com/poll/691967/three-teachers-weekly-saving-six-weeks-year.aspx)
- Pew Research Center (2025-01-15), [About a quarter of U.S. teens have used ChatGPT for schoolwork](https://www.pewresearch.org/short-reads/2025/01/15/about-a-quarter-of-us-teens-have-used-chatgpt-for-schoolwork-double-the-share-in-2023/)
- Tyton Partners (2025-06-11), [Time for Class 2025](https://tytonpartners.com/time-for-class-2025/)
- CoSN (2026), [State of EdTech Leadership](https://www.cosn.org/edtech-topics/state-of-edtech-leadership/)
- National Center for Education Statistics (n.d.), [Fast Facts: Educational institutions](https://nces.ed.gov/fastfacts/display.asp?id=84)
- Federal Trade Commission (2025-01-16), [FTC Finalizes Changes to Children’s Privacy Rule](https://www.ftc.gov/news-events/news/press-releases/2025/01/ftc-finalizes-changes-childrens-privacy-rule-limiting-companies-ability-monetize-kids-data)
- PowerSchool (2025), [PowerSchool Cybersecurity Incident](https://www.powerschool.com/security/sis-incident/)
- The Verge (2025-02-24), [Chegg sues Google over AI Overviews](https://www.theverge.com/news/619051/chegg-google-ai-overviews-monopoly)
- G2 (2026-04-15), [In the Answer Economy, Don’t Win the Click — Win the Answer](https://company.g2.com/news/g2-research-the-answer-economy)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- OpenAI Help Center (2025), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Chen, Wang, Chen and Koudas (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919), arXiv:2509.08919.
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)

---

This is the Markdown twin of https://underneath.agency/resources/edtech-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How endpoint security vendors win deals when CISOs ask AI"
description: "Endpoint security vendors earn AI shortlist places through public test results, incident transparency and consistent facts, not self-ranked lists."
canonical: "https://underneath.agency/resources/endpoint-security-enterprise-customers-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How do endpoint security vendors win enterprise deals when CISOs ask AI which EDR to buy?

By making the evidence that decides endpoint deals, independent test results, deployment facts and incident history, easy for AI answers to find and repeat. Endpoint protection is the largest single category in security spending, and large enterprises already run it, so most deals mean displacing an incumbent. An AI answer that names you at the start of that replacement cycle can be worth years of revenue.

## The short version

1. Endpoint protection platforms are the single largest security category at $17.8 billion, growing 14.5%, and will add $14.2 billion of spending by 2030, according to Gartner’s forecast as summarized by [Louis Columbus](https://softwarestrategiesblog.com/2026/04/01/top-10-fastest-growing-security-categories-gartner-2026-forecast/).
2. The buying trigger is ransomware: [Verizon](https://verizon.com/about/news/2025-data-breach-investigations-report) found it in 44% of breaches in its 2025 report, up 37% from the year before.
3. Endpoint products are tested in public: [AV-Comparatives](https://www.av-comparatives.org/tests/business-security-test-2025-august-november/) ran 461 real-world test cases against business products from August to November 2025, and [AV-TEST](https://www.av-test.org/en/antivirus/business-windows-client/) evaluated 15 endpoint products in July and August 2026.
4. The prize for winning is large: [CrowdStrike](https://www.01net.it/crowdstrike-reports-fourth-quarter-and-fiscal-year-2026-financial-results/) reported $5.25 billion in annual recurring revenue, and 50% of its subscription customers used six or more of its modules.
5. AI answers move around: in [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), a single ChatGPT answer showed only 57.8% of the brands it named across five runs of the same question.

## Who signs off on an EDR purchase, and how much is an account worth?

A CISO decides with the security operations team and IT; a won account grows by endpoints and modules.

Endpoint detection and response (EDR) watches laptops, servers and workloads for attacks and lets analysts respond; extended detection and response (XDR) adds signals from email, identity, network and cloud. Managed detection and response (MDR) wraps a service around either. Gartner’s 1Q26 forecast, summarized by Columbus, puts endpoint protection platforms at $17.8 billion, the single largest category in the security forecast, with the largest dollar increase of any category through 2030.

Three groups shape the decision:

- **The CISO**, who answers to the board if ransomware gets through.
- **The security operations team**, who live in the console every day and care about detection quality, false alarms and investigation speed.
- **IT operations**, who deploy the agent to every device and fear performance problems and outages.

The IT concern became sharper after July 2024, when a faulty CrowdStrike update crashed Windows machines worldwide. [Microsoft estimated](https://blogs.microsoft.com/blog/2024/07/20/helping-our-customers-through-the-crowdstrike-outage/) that 8.5 million Windows devices were affected, “less than one percent of all Windows machines.” Since then, update safety and vendor resilience are, we infer, standard questions in endpoint evaluations.

A won customer is worth a lot and tends to expand. CrowdStrike’s results for fiscal 2026 show ending annual recurring revenue of $5.25 billion, up 24%. Its module adoption figures show how accounts grow after the first sale: 50% of subscription customers used six or more modules, 34% seven or more and 24% eight or more. Accounts on its flexible Falcon Flex licensing reached $1.69 billion in ending recurring revenue.

## What triggers an enterprise endpoint purchase?

Ransomware incidents, renewals of an incumbent contract, platform consolidation, and doubts about the current vendor.

- **Ransomware.** Verizon’s 2025 report, which analyzed 12,195 confirmed data breaches, found ransomware in 44% of them. For smaller organizations, IDC’s Craig Robinson noted in the release, it was present in 88% of breaches.
- **Exposed weaknesses.** Exploitation of vulnerabilities was the entry point in 20% of breaches, up 34%, with a focus on perimeter devices and VPNs that an endpoint agent may not cover.
- **Renewals and consolidation.** Most large organizations already have endpoint protection, so a new vendor usually wins at an incumbent’s renewal, often as part of a platform deal. CrowdStrike notes, for example, that buyers can now buy its platform through Microsoft Marketplace using their existing Azure spending commitment. Marketplace listings matter for cloud tools too, as [our guide for cloud security vendors](https://underneath.agency/resources/cloud-security-enterprise-buyers-ai-search) explains.
- **Loss of confidence.** An outage, a missed detection or a bundled alternative from an existing supplier can reopen a decision mid-contract.

Each trigger produces a specific question, and those questions are increasingly put to AI assistants first. For identity vendors, a credential breach plays this role, covered in [our guide for identity platforms](https://underneath.agency/resources/iam-enterprise-demand-ai-search).

## Where does AI search sit in an endpoint security evaluation?

At the research and shortlist stage, though no vendor-neutral study measures CISOs’ use of AI assistants specifically.

The nearest evidence comes from surveys of technology buyers across categories, not security teams alone. In [TrustRadius’s 2026 B2B Buying Disconnect Report](https://www.demandgenreport.com/industry-news/news-brief/trustradius-ai-has-changed-how-buyers-research-but-not-what-they-trust/53853/), 63% of buyers used AI during their purchase journey and 94% of those fact-checked its answers at least some of the time. 83% shortlisted three or fewer products. TrustRadius sells review visibility, so it has an interest in the topic.

For endpoint security, the fact-checking step matters more than usual. Buyers can check claims against public test reports, peer reviews and incident records, and an AI answer is only the first filter. A reasonable expectation is that an answer which names a vendor without supporting evidence will not survive that check.

AI answers also vary. In our consistency study, ChatGPT named the same first brand in two runs of a question only 53.8% of the time. B2B software was the most stable industry we tested, with mean overlap of 0.708 across assistants, but even there a single answer is a sample, not a verdict.

## What do CISOs and SOC leads ask assistants about EDR?

Questions about comparisons, test results, resilience, coverage and cost of switching. We wrote the renewal-season prompts below to illustrate; they are not drawn from real CISO queries.

| Buying moment | Illustrative prompt |
|---|---|
| Comparison | “CrowdStrike Falcon vs SentinelOne Singularity vs Microsoft Defender for Endpoint for 8,000 endpoints?” |
| Category | “Do we need XDR, or is EDR plus a managed service enough for a 300-person company?” |
| Evidence | “Which EDR products did best in the latest MITRE ATT&CK Evaluations and AV-Comparatives tests?” |
| Resilience | “How do endpoint vendors test updates before release, and who has had outages?” |
| Coverage | “Which EDR tools protect Linux servers, Macs and older Windows systems equally well?” |
| Switching | “How hard is it to replace our current antivirus with a new EDR across 20 sites?” |
| Bundling | “Is Defender for Endpoint in Microsoft 365 E5 good enough, or do we need a specialist?” |

Each prompt can trigger several searches. Google documents that AI Overviews and AI Mode [may use a “query fan-out” technique](https://developers.google.com/search/docs/appearance/ai-features), and OpenAI documents that [ChatGPT search rewrites a question](https://help.openai.com/en/articles/9237897-chatgpt-search) into “one or more targeted queries.” For an evidence question, we infer those searches will reach test labs’ reports and vendors’ own readings of them.

## How does a place in an AI answer turn into a per-endpoint contract?

Through the renewal window: AI answer, shortlist, proof of concept on real devices, then a multi-year, per-endpoint contract.

1. Ahead of a renewal or after an incident, a CISO or SOC lead asks an assistant to compare options.
2. The answer names a few vendors; the team checks them against test reports, reviews and peers.
3. Two or three vendors run a proof of concept on a sample of real devices, often with a simulated attack.
4. The winner replaces the incumbent across the estate, priced per endpoint.
5. The account expands into more modules and sometimes a managed service, as CrowdStrike’s module figures show.

The AI answer matters most at step 2, and the timing is unforgiving. Endpoint contracts usually run for years, so a vendor absent from the shortlist at renewal may wait a full contract term for the next chance, we infer.

## What puts one endpoint agent on an assistant’s list and leaves another off?

No platform documents how it chooses vendors; cited lists often promote their publishers, and buyers verify independent evidence.

**What Google and OpenAI disclose.** Both say their AI answers search the web and link to the pages behind them; neither explains how it picks which EDR or XDR vendors to recommend.

**Observed in studies.** Endpoint vendors often publish their own “best EDR” lists, and AI answers do cite such lists. In [our study of cited “best of” lists](https://underneath.agency/research/self-promoting-best-lists-study), 24.2% of 269 cited numbered lists with an identifiable publisher ranked their own publisher first, and when a publisher included itself, 92.9% of the time it was number one. Those lists were only 1.1% of all citations, and we found no detectable difference in how often a self-ranking list’s top entry was named compared with an independent list’s. Why an independent ranking counts for more than a vendor’s own “best EDR” page is covered in our guide to [which pages to target](https://underneath.agency/resources/best-of-lists-ai-recommendations).

**The independent evidence buyers check.** Endpoint security has more public testing than almost any software category:

- **AV-Comparatives** publishes business tests twice a year. To earn its December 2025 “Approved Business Product” award, a product had to score at least 90% in the Malware Protection Test with zero false alarms on common business software.
- **AV-TEST** scores business endpoint products every two months on protection, performance and usability.
- **MITRE ATT&CK Evaluations** are another widely cited test series, but we could not verify their results from a saved source, so we draw nothing from them here.
- **Peer reviews**, such as Gartner Peer Insights, which CrowdStrike cites for its Customers’ Choice recognition.

**Our inference.** These sources are public, specific and repeated in security media, which makes them natural material for AI answers. A reasonable expectation is that a vendor whose test participation, results and methods are clearly explained on its own site, and accurately reported elsewhere, is easier to name with confidence. We have not tested this for endpoint vendors. How CISOs rank these trust signals is covered in our article on [cybersecurity software and AI search](https://underneath.agency/resources/cybersecurity-software-revenue-from-ai-search).

## What does an endpoint vendor lose by being left out at renewal?

A missed renewal window, which in this category can mean years; no study has measured it in dollars.

- **Short lists.** With 83% of technology buyers shortlisting three or fewer products, and endpoint leaders well known, a challenger left out of the first answer may not get a proof of concept at all.
- **Long contracts.** A lost renewal locks the estate to another agent for the contract term, we infer.
- **Expansion lost too.** The first sale is the foothold for later modules; half of CrowdStrike’s subscription customers used six or more.
- **Inaccurate answers.** An AI answer that repeats an old test result, wrong platform coverage or an incident without its resolution can cost a shortlist place. If an assistant is repeating stale test results or coverage, [how to fix wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) sets out the steps.

## What does GEO look like for an EDR or XDR vendor?

It makes your test evidence, coverage and track record easy for assistants to repeat; it cannot promise a recommendation.

1. **Explain your test results honestly.** Publish which independent tests you entered, what they measured and how you did, with links to the lab’s own report. Do not overclaim; buyers and journalists check.
2. **State coverage precisely.** List supported operating systems, versions, server and cloud workloads, and offline or air-gapped options on plain pages.
3. **Publish update and resilience practices.** Explain how updates are staged and tested, and how customers control rollout. Buyers now ask.
4. **Be transparent about incidents.** Keep clear, dated public records of past problems and what changed. If the only account of an outage an assistant can find is someone else’s, that is the version it will repeat.
5. **Earn independent coverage.** Threat research, incident response reports and expert comment after major attacks give security media something to cite. Wider authority-building for security brands is covered in [how brands build authority for AI search](https://underneath.agency/resources/how-brands-build-authority-for-ai-search).
6. **Build peer proof.** Encourage detailed reviews from security teams that run the product, without incentives that breach platform rules.
7. **Avoid self-ranking lists as a strategy.** They are a small share of citations and buyers discount them.
8. **Measure repeatedly.** Run renewal-season questions across ChatGPT, Google AI Overviews and AI Mode, Gemini, Perplexity, Copilot and Claude several times; see [how many prompts to track](https://underneath.agency/resources/how-many-prompts-to-track-ai-visibility).

## What can’t the evidence yet tell endpoint security vendors?

It cannot yet show that assistants rely on lab results when naming EDR vendors, or that visibility moves win rates.

- **No CISO-specific study.** TrustRadius covers technology buyers broadly; we found no vendor-neutral, public study of how CISOs use AI to choose endpoint products.
- **MITRE results.** We could not save a primary source for the latest MITRE evaluation results, so we cite no figures from them.
- **Test labs differ.** AV-TEST and AV-Comparatives measure different things, and vendors choose which to enter.
- **No tie to win rates or contracts has been shown.** For what is known so far, see [does AI visibility drive business results](https://underneath.agency/resources/does-ai-visibility-drive-business-results).

## What should an endpoint vendor check before its target accounts reach renewal?

Ask the renewal-season questions your target accounts would ask, and see how assistants describe your evidence.

Cover comparison, evidence, resilience, coverage and switching, and repeat each question in every major assistant, since a single answer is only a sample. Note whether you are named, which test results and incidents come up, and which sources are cited. Gaps usually show missing or unclear public evidence rather than missing marketing pages.

That check is where our work with security vendors begins: [ask us to review your standing in EDR comparisons](https://underneath.agency/contact). We will show how AI answers present you against incumbents in enterprise endpoint evaluations, and which gaps are most likely costing you proof-of-concept invitations at renewal time. Our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page outlines the work that follows, such as plain pages on test results and coverage, clear incident records and repeated renewal-season checks.

## Frequently asked questions

### Do MITRE ATT&CK Evaluations help endpoint vendors appear in AI answers?

Untested. We have not verified MITRE results against a saved source, and no study has measured whether any lab’s test series changes which vendors AI answers recommend.

### Should we publish our own “best EDR tools” list?

It is a weak strategy. In our study, 92.9% of self-including lists put their publisher first, and such lists were only 1.1% of AI citations.

### Can a challenger compete with CrowdStrike and Microsoft in AI answers?

It can, through specific, verifiable proof for a defined buyer, but it starts behind brands with far more coverage.

### How should we handle a past incident in AI answers?

Publish a clear, dated account of what happened and what changed, so answers have your verified record and not only third-party summaries.

## Sources

- Louis Columbus, Software Strategies Blog (2026-04-01), [Gartner’s $246.2B Security Forecast shows 10 categories growing 2x to 3x the market](https://softwarestrategiesblog.com/2026/04/01/top-10-fastest-growing-security-categories-gartner-2026-forecast/)
- Verizon (2025-04-23), [2025 Data Breach Investigations Report news release](https://verizon.com/about/news/2025-data-breach-investigations-report)
- AV-Comparatives (2025-12-15), [Business Security Test 2025 (August–November)](https://www.av-comparatives.org/tests/business-security-test-2025-august-november/)
- AV-TEST (2026-08), [Test antivirus software for Windows: business users](https://www.av-test.org/en/antivirus/business-windows-client/)
- CrowdStrike, via 01net (2026-03-03), [CrowdStrike Reports Fourth Quarter and Fiscal Year 2026 Financial Results](https://www.01net.it/crowdstrike-reports-fourth-quarter-and-fiscal-year-2026-financial-results/)
- Microsoft (2024-07-20), [Helping our customers through the CrowdStrike outage](https://blogs.microsoft.com/blog/2024/07/20/helping-our-customers-through-the-crowdstrike-outage/)
- Demand Gen Report, James Hickey (2026-07-30), [TrustRadius: AI Has Changed How Buyers Research, But Not What They Trust](https://www.demandgenreport.com/industry-news/news-brief/trustradius-ai-has-changed-how-buyers-research-but-not-what-they-trust/53853/)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- OpenAI Help Center (2025), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)
- Underneath (2026), [How many “best of” lists cited by AI rank their own brand first?](https://underneath.agency/research/self-promoting-best-lists-study)

---

This is the Markdown twin of https://underneath.agency/resources/endpoint-security-enterprise-customers-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How engineering firms win project inquiries from AI search"
description: "Engineering firms get named by AI when their project types, licenses, people and locations are stated clearly and confirmed by sources clients already trust."
canonical: "https://underneath.agency/resources/engineering-firms-project-inquiries-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Will clients find our engineering firm when they ask AI who can design their project?

They can, if an assistant can confirm what you design, where you are licensed, who leads the work and which similar projects you have delivered. Referrals still bring most engineering work, but a referred client now checks a firm online before calling, and AI assistants are becoming part of that check.

This guide is for engineering services firms: civil, structural, mechanical, electrical and plumbing (MEP), geotechnical and process engineering. It covers private clients, architects and prime firms that hire consultants, and public owners that select engineers by qualifications. It is not about engineering software, which we cover separately.

## The short version

1. The market is large but capacity-bound: the [ACEC Research Institute](https://www.acec.org/news/last-word-blog/post/new-acec-institute-research-on-engineering-industrys-economic-impact/) puts US engineering and design services revenue at $459 billion in 2024, and more than half of ACEC member firms turned down projects because they lacked staff.
2. Firms are choosing work more carefully: in [Deltek’s 2025 Clarity study](https://www.deltek.com/en/about/media-center/press-releases/2025/what-the-46th-annual-deltek-clarity-ae-study-reveals-about-the-industry), proposal volume fell 38% while the value of awarded work grew 52%, and the median win rate rose to 50%.
3. Fit now beats relationships: the same study found “fit for the type of work” overtook existing relationships as the top factor in deciding which projects to pursue.
4. Firms already use AI; their clients are next: in [Deltek’s 2026 study](https://www.deltek.com/en/about/media-center/press-releases/2026/the-latest-deltek-clarity-industry-studies-highlight-ai-challenges) of 896 architecture and engineering firms, AI adoption rose from 53% to 70% in a year, and 42% believe they could lose market share within two years without significant digital transformation.
5. Public work is selected on qualifications: under [federal law](https://www.law.cornell.edu/uscode/text/40/1103), an agency must hold discussions with at least 3 firms and rank the most highly qualified before any fee is negotiated.

## Who hires engineering firms, and what is one client worth?

Public owners, private developers, architects, industrial owners and prime firms, and a good client returns project after project.

The industry is large. The ACEC Research Institute found engineering and design services revenue grew 5.3% in 2024, with growth expected to slow to 2.3% in 2025, and the industry directly employs 1.7 million Americans. Texas ($96 billion) and California ($94 billion) led all states in economic value added. At the top, the ENR Top 500 Design Firms grew revenue 7.4% to $158.7 billion in 2025, [according to ENR figures quoted by BL Companies](https://www.blcompanies.com/bl-companies-rises-up-rankings-on-2026-engineering-news-record-top-500-design-firms/), which ranked No. 231 with more than $111 million in revenue.

The buyers differ by market:

- **Public owners** such as federal agencies, state transportation departments, cities and water utilities. Federal selection follows the Brooks Act and [FAR Subpart 36.6](https://www.acquisition.gov/far/subpart-36.6): firms file the Standard Form 330 qualifications statement, a board ranks them, and the agency then negotiates a fair and reasonable price with the top-ranked firm.
- **Private owners and developers** who hire civil, structural and MEP engineers for buildings, sites and industrial facilities, often after a referral.
- **Architects and prime firms** that bring in consultants for a pursuit, such as a structural engineer for a mass timber building or a specialty subconsultant to complete a team.
- **Industrial owners** who need process, electrical or controls engineering for a plant expansion.

No public source gives an average fee per client, so we do not quote one. What the data shows is that each relationship compounds. Firms are winning half of what they chase, and in a capacity-constrained market the value of a client is the stream of follow-on projects, on-call contracts and referrals it brings.

## Where do clients meet AI before they call an engineer?

Mostly at the checking step, after a referral or before a shortlist. Someone needs to know who does this work nearby.

We found no current survey of how owners and developers use AI to choose engineering firms. The most recent public study we found on how end users pick engineers is old: a 2012 survey by Accountability Information Management, [reported by Canadian Consulting Engineer](https://canadianconsultingengineer.com/clients-mostly-use-peer-recommendations-to-choose-engineers), found that 64% of end users relied on recommendations from peers, 33% on contractors’ recommendations and 14% on web searches. We cite it as a baseline, not as today’s picture.

Today’s cross-industry evidence points in one direction. In [Gartner’s survey of 645 B2B buyers](https://www.gartner.com/en/newsroom/press-releases/2026-05-20-gartner-survey-finds-sixty-nine-percent-of-b-two-b-buyers-turn-to-sales-reps-to-validate-ai-generated-insights), 45% had used generative AI in a recent purchase, mainly to gather information on vendors, and 69% prefer to validate AI-generated insights with a salesperson. For an engineering firm, that “salesperson” is usually a principal on the phone.

The firms themselves are early AI users. Deltek’s 2026 study found generative AI use among architecture and engineering firms rose from 64% to 78% in a year, and its 2025 study listed proposal development and business development among the main uses. Our inference: the people who hire engineers work in the same offices and use the same tools, so AI-assisted checking of consultants is a reasonable expectation, but it has not been measured. The vendors selling design tools to these offices face the same shift, covered in [how engineering software gets found](https://underneath.agency/resources/engineering-software-ai-search).

## Which questions do clients and partners ask AI about engineering firms?

Questions that combine a discipline, a project type, a place and a constraint, much like a request for qualifications.

These examples are our own, written to show the shape of real requests. They are not captured from any assistant or client.

| Who is asking | Example question |
|---|---|
| Developer | “Structural engineers in Denver with mass timber mid-rise experience” |
| Architect | “MEP engineering firms that design laboratories and cleanrooms in the Boston area” |
| Industrial owner | “Process engineering firms for a battery materials plant in the Southeast” |
| Prime firm | “Certified small business geotechnical firms licensed in Texas for a highway project” |
| City or utility staff | “Engineering firms that have designed water treatment plant upgrades for cities under 100,000 people” |
| Referred client | “What projects has [firm name] done, and who are its principals?” |

Each one is a filter. If your project pages never say “mass timber”, your licenses by state are not listed, or your lab work sits inside a PDF brochure, an assistant has nothing to match.

## How does an AI mention turn into a project?

Through an inquiry or invitation, a qualifications check, an interview and a negotiated fee, then repeat work.

1. **Named or checked.** A client asks for firms, or checks a firm someone recommended. The answer shapes who gets the call.
2. **Invited.** The client sends a request for proposal or qualifications, or a prime firm calls about teaming. For public work, the formal notice still goes out; AI influences who notices it and who is invited to partner.
3. **Qualified.** The client reviews similar projects, licensed staff and references. Under federal rules, the agency must discuss the work with at least three of the most highly qualified firms and rank them in order of preference.
4. **Interviewed and negotiated.** Price comes after selection in qualifications-based selection, and it is negotiated, not bid.
5. **Repeated.** A successful project leads to the next one, an on-call contract or a referral.

Assistants shape who is named and who is invited, and little after that. Your people and past work win the rest. In Deltek’s data, firms submitted far fewer proposals but won more value, which suggests the right invitations matter more than the number of them.

## What makes an assistant name one engineering firm over another?

Verifiable facts about projects, people, licenses and place; the platforms explain their searches, not their choices.

**Documented by the platforms.** [OpenAI says](https://help.openai.com/en/articles/9237897-chatgpt-search) ChatGPT search rewrites a question into one or more targeted queries sent to search providers, and that sites must allow its crawler, OAI-SearchBot, to be eligible for inclusion. Google’s [AI Mode announcement](https://blog.google/products/search/ai-mode-search/) describes a “query fan-out” technique that sends several related searches across the subtopics of one question. Neither publishes how firms are chosen.

**Observed in our studies.** Engineering is often a local or regional purchase, so our local studies are the closest evidence, with the caveat that they covered other local services:

- In [our study of ChatGPT local recommendations](https://underneath.agency/research/chatgpt-local-recommendations-study), two answers to the same question on the same day shared a business-set overlap of 0.65, on a scale where 1 means identical. The firms named shift from run to run.
- In [our study of which Google Maps businesses ChatGPT recommends](https://underneath.agency/research/chatgpt-local-picks-google-profile-study), businesses with more reviews than the local median were 19.5 points more likely to be listed, after adjusting for other signals.
- In [our business-facts study](https://underneath.agency/research/ai-business-facts-accuracy-study), 18.9% of answers stated at least one fact that differed from the business’s Google profile, mostly where the business’s own sources disagreed.

**Our inference for engineering firms.** The trust factors are the ones a selection board already scores, put where machines can read them: project type, size, location and your role; licensed professionals and the states they are licensed in; [professional licensure](https://ncees.org/licensure/) itself, which typically requires an accredited degree, four years of experience and two exams; certifications such as small or disadvantaged business status; awards and trade coverage; and consistent firm facts across your site, directories and profiles.

## What does it cost an engineering firm to be left out?

Invitations and teaming calls that go to another firm, which no pipeline report will show.

We have no measurement of projects lost to AI absence, so this is our reasoning, labeled as such:

- **Growth is slowing.** ACEC expects revenue growth to ease from 5.3% to 2.3%, and Deltek’s 2026 study found firms forecasting 9.5% net revenue growth for 2026 while backlogs soften. When work tightens, being on the first list matters more.
- **Staff is the constraint.** With staff growth of only 1.2% in Deltek’s data and turnover above 13%, firms cannot answer every request for proposals. Being invited to well-fitted work is worth more than chasing volume.
- **Wrong facts disqualify.** An answer that lists an old office, misses a state license or describes you as an architecture firm can remove you from a search before anyone reads your qualifications. Our guide on [fixing wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) explains how to trace the source.

## How does GEO work for an engineering services firm?

Generative engine optimization (GEO) makes your firm easy for AI assistants to find, describe accurately and confirm.

For an engineering firm, the work usually covers:

1. **Project pages in plain text.** One page per notable project with type, size, location, delivery method, your role and the client (with permission), written as text rather than images or brochures.
2. **Discipline and market pages.** Clear pages for what you design (structural, MEP, civil, process) and who you design it for (healthcare, water, transportation, industrial).
3. **People and licenses.** Principal and project manager pages listing licenses by state and relevant project experience.
4. **Consistent firm facts.** The same name, offices, disciplines and certifications on your site, Google Business Profile, LinkedIn, association directories and government registrations.
5. **Independent coverage.** Award entries, trade publication features, conference papers and association rankings give an assistant outside sources to cite. Our article on [how small brands get recommended by AI](https://underneath.agency/resources/how-small-brands-get-recommended-by-ai) explains why outside sources matter.
6. **Reviews where clients look.** For firms that serve local private clients, profile reviews are one of the few signals we have seen move local recommendations. Other local professional services work the same way; recruiting agencies, for example, are named from local listings, reviews and specialty pages, as [our guide for recruiting agencies](https://underneath.agency/resources/recruiting-agencies-employer-clients-ai-search) explains.
7. **Crawl access and measurement.** Allow the documented search crawlers, then ask a fixed set of project questions across ChatGPT, Gemini, Perplexity, Copilot and Google’s AI features over time, and compare who is named with where invitations actually come from.

No engineering firm can be promised a place in an AI answer. What this work does is make your firm the easiest one for an assistant, and then a client or selection board, to check. Firms in other project-based fields face the same pattern, as our guides on [how IT services firms win projects from AI search](https://underneath.agency/resources/it-service-firms-leads-ai-search) and [how contract manufacturers get on supplier shortlists](https://underneath.agency/resources/contract-manufacturers-rfqs-from-ai-search) show.

## What is still unknown about AI and engineering firm selection?

How often clients use AI to find engineers, and how many projects it produces, has not been measured.

- **No client-side survey.** The AI adoption figures here describe engineering firms, not their clients. The only engineer-selection survey we found is from 2012.
- **Public procurement limits the effect.** Qualifications-based selection runs on formal submissions and board scoring. We infer AI matters more for private work, teaming and the research before a submission than for the scoring itself.
- **Sponsors have interests.** Deltek sells software to project-based firms, and ACEC advocates for the industry.
- **No engineering-specific ranking studies.** Our studies covered other local services and buyer questions. Applying them to engineering firms is our inference.

## Where should an engineering firm start?

Ask assistants the project questions your best clients and teaming partners would ask, and see whether your firm is named.

That first check usually shows whether you appear for your core disciplines and markets, whether your offices, licenses and project types are described correctly, which directories and publications the answers rely on, and which firms are named instead.

If your growth depends on more well-fitted invitations to propose, [have us check how AI describes your firm](https://underneath.agency/contact). We ask the discipline, market and location questions clients and teaming partners use, show which firms and directories are named in your place, and set out the project and license facts most likely to bring more qualified project inquiries. Our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page describes how that work is handled for a design firm, from plain-text project pages and licensed-staff profiles to consistent directory listings.

## Frequently asked questions

### Does AI search matter if most of our work comes from referrals?

Yes. Referred clients still check a firm before calling, and assistants are becoming part of that check. If the answer misdescribes your firm, the referral can stall.

### Can AI search help us win public contracts selected on qualifications?

Indirectly. Selection boards score formal submissions, but prime firms looking for partners and agency staff researching the field can use AI before a solicitation.

### Should individual engineers have their own pages?

For licensed leaders, yes. Pages that list licenses by state and relevant projects help clients and assistants confirm who will do the work.

### Do rankings like ENR’s Top 500 help?

They are independent sources that assistants and clients can find. A listing, an award or a feature gives outside confirmation that your own site cannot.

## Sources

- ACEC (2025-10-08), [New ACEC Research Shows Engineering and Design Services Industry Contributed $685 Billion to U.S. GDP in 2024](https://www.acec.org/news/last-word-blog/post/new-acec-institute-research-on-engineering-industrys-economic-impact/)
- Deltek (2025-05-13), [AI, Talent and Record Profits: What the 46th Annual Deltek Clarity A&E Study Reveals](https://www.deltek.com/en/about/media-center/press-releases/2025/what-the-46th-annual-deltek-clarity-ae-study-reveals-about-the-industry)
- Deltek (2026-05-12), [The Latest Deltek Clarity Industry Studies Highlight AI Challenges, Talent Strain, and Delivery Capacity Pressures](https://www.deltek.com/en/about/media-center/press-releases/2026/the-latest-deltek-clarity-industry-studies-highlight-ai-challenges)
- BL Companies (2026-06-12), [BL Companies Rises Up Rankings on 2026 Engineering News-Record Top 500 Design Firms](https://www.blcompanies.com/bl-companies-rises-up-rankings-on-2026-engineering-news-record-top-500-design-firms/)
- Legal Information Institute (n.d.), [40 U.S. Code § 1103: Selection procedure](https://www.law.cornell.edu/uscode/text/40/1103)
- Acquisition.gov (n.d.), [FAR Subpart 36.6: Architect-Engineer Services](https://www.acquisition.gov/far/subpart-36.6)
- NCEES (n.d.), [Licensure](https://ncees.org/licensure/)
- Canadian Consulting Engineer (2012-12-17), [Clients mostly use peer recommendations to choose engineers](https://canadianconsultingengineer.com/clients-mostly-use-peer-recommendations-to-choose-engineers)
- Gartner (2026-05-20), [Gartner Survey Finds 69% of B2B Buyers Turn to Sales Reps to Validate AI-Generated Insights](https://www.gartner.com/en/newsroom/press-releases/2026-05-20-gartner-survey-finds-sixty-nine-percent-of-b-two-b-buyers-turn-to-sales-reps-to-validate-ai-generated-insights)
- OpenAI Help Center (2026), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Google (2025-03-05), [Expanding AI Overviews and introducing AI Mode](https://blog.google/products/search/ai-mode-search/)
- Underneath (2026), [ChatGPT local recommendations: stable details, shifting lists](https://underneath.agency/research/chatgpt-local-recommendations-study)
- Underneath (2026), [Which Google Maps businesses does ChatGPT recommend?](https://underneath.agency/research/chatgpt-local-picks-google-profile-study)
- Underneath (2026), [Do AI answers match a business’s Google profile?](https://underneath.agency/research/ai-business-facts-accuracy-study)

---

This is the Markdown twin of https://underneath.agency/resources/engineering-firms-project-inquiries-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How engineering software wins customers through AI search"
description: "Engineering software gets recommended by AI when its workflows, file formats, plans and limits are stated plainly and confirmed by sources engineers trust."
canonical: "https://underneath.agency/resources/engineering-software-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# When engineers ask AI which CAD, simulation or PLM tool to use, will they hear our name?

They will if an assistant can find and verify what your software does for a specific workflow, what it costs to start, which files it reads and where it falls short. Engineers already use AI in their own work, software buyers increasingly ask chatbots for recommendations, and a wave of acquisitions has many teams asking what to use next.

This guide is for companies that sell computer-aided design (CAD), simulation and analysis (CAE) and product lifecycle management (PLM) software to engineers. It is not about engineering services firms, which we cover separately, or software for running factories. The revenue here comes from a trial or free plan that turns into paid seats, then an enterprise agreement.

## The short version

1. The category is large and growing: [CIMdata](https://www.industrialmachinerydigest.com/software/quality-management-software/cimdata-publishes-executive-plm-market-report) puts the broader PLM market, which includes CAD, simulation and PLM platforms, at $88.3 billion in 2025, up 9.9%.
2. The leaders sell recurring subscriptions: [Autodesk](https://adsknews.autodesk.com/en/?p=55945) reported $7,206 million in fiscal 2026 revenue, up 18%, with its manufacturing products at $1,379 million, and [Dassault Systèmes](https://www.3ds.com/assets/invest/2026-02/dassault-systemes-25q4-earnings_pr_va.pdf) said recurring revenue was 82% of its 2025 software revenue.
3. Consolidation is reshaping choices: Synopsys completed its acquisition of Ansys in July 2025, and Siemens completed its purchase of Altair for an enterprise value of about $10 billion.
4. Engineers are early AI adopters but cautious: in [Digital Engineering 24/7’s 2025 reader survey](https://www.digitalengineering247.com/article/engineering-technology-outlook-2026/features), 33% already used AI and 56% said CAD assistants would benefit design work most.
5. Starting is cheap, so the shortlist matters: [Onshape](https://www.onshape.com/en/pricing) offers a free plan for non-commercial use and up to 6 months free on its Professional plan for qualified users, so engineers can try several tools before anyone talks to sales.

## Who chooses engineering software, and what is a customer worth?

Engineers pick the tool they want; managers, CAD administrators, IT and resellers decide how many seats and on what terms.

The buying group usually includes the design or simulation engineer who will use the software, an engineering manager who owns the budget, a CAD or PLM administrator who worries about data and file compatibility, IT and procurement, and often a reseller. Autodesk, for example, calls its channel partners Solution Providers. In Digital Engineering 24/7’s survey of 194 readers, the largest group (26%) described their role as product or system design engineering.

The money sits in recurring revenue and expansion. Public results show the scale:

| Company and period | Figure |
|---|---|
| Autodesk, fiscal 2026 | $7,206 million revenue, up 18% |
| Autodesk manufacturing products, fiscal 2026 | $1,379 million, up 16% |
| Dassault Systèmes, 2025 | €6.24 billion total revenue |
| Dassault Industrial Innovation (CATIA, SIMULIA, ENOVIA), 2025 | €3.13 billion, 56% of software revenue |
| Dassault Mainstream Innovation (including SOLIDWORKS), 2025 | €1.43 billion |

At the entry level, the seat price is public for some tools. Onshape lists its Standard plan at $1,500 and its Professional plan at $2,500 per user per year. A team of engineers on an annual plan, renewed for years and expanded into data management or simulation, is worth far more than the first seat. That expansion path, not the first sale, is where AI visibility pays back.

## Where does AI already sit in how engineers pick tools?

At the start: engineers and software buyers ask chatbots for options, then test them hands-on.

There is no public survey of how engineers use AI to choose CAD or simulation software. Two pieces of evidence come close.

First, engineers are using AI in their work. Digital Engineering 24/7 found 33% of its readers already used AI, generative AI or machine learning, and 29% planned to within two years. Asked where AI would help most, they chose CAD assistants (56%), AI-supported simulation (48%) and generative CAD (45%).

Second, software buyers in general are using chatbots to find products. In [G2’s 2026 Buyer Behavior Report](https://sell.g2.com/2026-buyer-behavior-report), more than 80% of buyers had sourced software recommendations from an AI chatbot in the last two years. That survey covers all business software, not engineering tools, so we treat it as a direction, not a measurement for this category.

Simulation and PLM are already widespread among engineers, which keeps switching questions alive. In the same Digital Engineering 24/7 survey, 48% of readers used simulation software and 30% used PLM.

## Which questions do engineers ask AI about software?

Workflow questions, comparison questions and “what now” questions after a vendor changes hands.

The questions below are ours, written to show the pattern. They were not captured from any assistant or user.

| Situation | Example question |
|---|---|
| Choosing a first tool | “Best CAD for a three-person hardware startup doing sheet metal and small assemblies” |
| Comparing | “Fusion vs SOLIDWORKS vs Onshape for a small team that needs data management” |
| Specialist analysis | “CFD software for electronics cooling that can run in the cloud” |
| Regulated work | “PLM for a medical device company that needs FDA design controls” |
| After an acquisition | “What changes for Altair license holders now that Siemens owns Altair?” |
| Replacing a tool | “Alternatives to our current FEA package that read our existing models” |
| On a budget | “Cheapest professional CAD that exports STEP and has a free trial” |

The last row matters more than it looks. In [our prompt phrasing study](https://underneath.agency/research/ai-prompt-phrasing-study), asking the same question again kept the same first brand 68.0% of the time, but adding “on a tight budget” kept it only 15.3%. Engineers who add a constraint such as a budget, team size or file format can get a different shortlist, so your facts need to cover those constraints.

Acquisitions create a wave of these questions. [Synopsys completed its acquisition of Ansys](https://www.airframer.com/news/release/synopsys-completes-acquisition-of-ansys) in July 2025 and said it was positioned to win in an expanded $31 billion addressable market. [Siemens completed its acquisition of Altair](https://www.digitalengineering247.com/article/siemens-completes-acquisition-of-altair/news) for about $10 billion. Customers of both are asking about roadmaps, licensing and alternatives.

## How does an AI answer become seats and an enterprise agreement?

Through a shortlist, a free plan or trial, a team pilot, a purchase and then expansion.

1. **Shortlisted.** An engineer asks for tools that fit a workflow. The answer usually offers a handful of CAD or simulation tools.
2. **Tried.** Engineers download a trial, use a free plan or start an education license. Onshape, for example, offers professional-grade CAD free of charge to students and educators and no-cost licenses to qualified startups.
3. **Tested on real work.** The team checks a benchmark part, file import, data management and performance on its hardware or in the cloud.
4. **Bought.** Seats are purchased directly or through a reseller.
5. **Expanded.** More seats, data management, simulation and PLM follow, sometimes under an enterprise agreement.

An assistant can help with the shortlist and the trial; your product, documentation and resellers win the last three. A tool that never makes the shortlist never gets the trial.

## Why does an assistant recommend one engineering tool over another?

It recommends tools whose facts it can find and confirm; the platforms document their searches, not their picks.

**What the platforms document.** According to [OpenAI](https://help.openai.com/en/articles/9237897-chatgpt-search), ChatGPT search turns a question into one or more targeted queries for its search providers, and only sites that let OAI-SearchBot in are eligible to appear. [According to Google](https://blog.google/products/search/ai-mode-search/), AI Mode fans a question out into multiple related searches on its subtopics, a technique the company calls “query fan-out.” Neither publishes how a CAD or simulation tool gets chosen.

**Observed in our studies.**

- Before answering a software buyer’s question, ChatGPT ran a mean of 3.7 searches in [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study). When one of those searches named a source, the answer cited that source 44.0% of the time, against 8.1% when it did not.
- Outside coverage was the strongest predictor in [our brand entity study](https://underneath.agency/research/brand-entity-ai-recommendations-study): every tenfold rise in the number of independent sites that named a brand in the cited pages came with 4.7 times the odds of being recommended.
- In [our study of self-promoting lists](https://underneath.agency/research/self-promoting-best-lists-study), 24.2% of AI-cited “best X” lists with an identifiable publisher ranked their own publisher first. Many “best CAD software” pages are written by vendors, and engineers know it.

**Our inference for engineering software.** The facts an engineer checks are concrete: supported workflows (part, assembly, drawing, sheet metal, CFD, FEA), file formats and kernels, operating systems and cloud options, licensing terms and what changed after an acquisition, plan limits, and training resources. Independent sources that confirm them include trade publications, user forums, review platforms, university courses and benchmark reports.

## What does an engineering software vendor lose when AI leaves it out?

Trials it never sees, from engineers who will standardize a team on whatever they tried first.

We have no measurement of trials lost to AI absence, so this reasoning is our own:

- **Trials are self-serve.** With free plans and long trials common, an engineer can try three tools in a week without contacting a vendor. A tool missing from the shortlist is not compared at all.
- **Habits last.** We infer that students and hobbyists who learn a tool on a free or education plan often carry it into work, so early AI recommendations shape which tools a generation of engineers knows.
- **Switching windows are short.** Acquisitions and pricing changes push customers to look around for a few months. A rival that is named during that window can win accounts that were not for sale before.
- **Wrong facts cost deals.** An answer that says your tool lacks a feature, cannot read a format or has no free trial removes you from evaluations. When an answer gets a file format or a plan wrong, our guide to [fixing wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) shows how to find and correct the page behind it.

## How does GEO work for a CAD, simulation or PLM company?

Generative engine optimization (GEO) makes your software easy for AI assistants to find, describe accurately and verify.

For an engineering software company, the work usually covers:

1. **Workflow pages in plain text.** One page per job the software does, with capabilities, limits and example models, not just a feature grid.
2. **Interoperability facts.** Supported file formats, geometry kernels, import and export limits, and integrations with data management and simulation tools.
3. **Plan and licensing clarity.** What the free, trial, education and startup options include, and what changes on paid plans, stated where a crawler can read it.
4. **Honest comparison and migration pages.** How you differ from the tools engineers compare you with, and how to move models across. Our review of [whether comparison pages help B2B brands get cited](https://underneath.agency/resources/do-comparison-pages-help-b2b-ai-citations) covers what works and what does not.
5. **Independent proof.** Trade press, benchmark studies, user forums, university courses, review platforms and conference talks. Our article on [best-of lists and AI recommendations](https://underneath.agency/resources/best-of-lists-ai-recommendations) explains why independent lists carry weight.
6. **Consistent reseller and partner pages.** Resellers often publish their own pages about your product; keep product names, plans and capabilities consistent across them.
7. **Crawl access and measurement.** Allow the documented search crawlers, and ask a fixed set of workflow, comparison and budget questions across ChatGPT, Gemini, Perplexity, Copilot and Google’s AI features over time, then compare with trial signups.

Nobody can promise that an assistant will recommend your CAD, simulation or PLM tool. The goal is to make it the easiest one for an assistant, and then an engineer, to check. Developer tool companies sell to a similar technical audience, and our article on [how developer tool companies win users when developers ask AI first](https://underneath.agency/resources/developer-tools-ai-search) shows how that plays out.

## Which parts of this can’t the data show yet?

How often engineers pick tools from AI answers, and how many seats that produces, has not been measured.

- **No engineering-specific buying data.** The chatbot figure here covers all business software buyers. We found no public survey of how engineers use AI to choose CAD, simulation or PLM tools.
- **Sources have interests.** Vendors report their own results, CIMdata sells research to the PLM industry, and G2 runs a review platform.
- **No category ranking studies.** Our studies covered buyer questions across many categories, not engineering software, so their relevance to CAD, simulation and PLM buyers is our inference.
- **Seat economics are private.** Public results show revenue and some list prices, not what a typical team spends over its lifetime.

## Where should an engineering software company start?

Ask assistants the workflow and comparison questions engineers ask before a trial, and see whether your tool is named.

That first check usually shows whether your software appears for its core workflows, whether its formats, plans and limits are described correctly, which publications, forums and lists the answers rely on, and which tools are recommended instead.

If your growth depends on more trials turning into team seats and enterprise agreements, [get in touch and we will test how assistants recommend your software](https://underneath.agency/contact). We put the workflow, comparison and budget questions engineers ask to the main assistants, show which rival tools, forums and lists win those answers, and set out the fixes most likely to bring more qualified trials. How the ongoing work runs for a CAD, simulation or PLM vendor, from workflow pages to reseller consistency, is laid out on our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page.

## Frequently asked questions

### Do engineers trust AI recommendations for software?

They use them as a starting point. Engineers test tools on real models before committing, so AI decides what gets tried, not what gets bought.

### Should an engineering software company list its plan prices publicly?

Where you can, yes. Clear plan and licensing facts let an assistant answer budget questions accurately instead of guessing or leaving you out.

### Do vendor-written “best CAD software” lists help?

They can be cited, but self-ranking lists are easy to spot. Independent reviews, forums and trade coverage are stronger confirmation.

### How should we handle questions after an acquisition?

Publish a plain page on what changes and what does not for customers: licensing, support, roadmap and file compatibility. Customers and assistants both look for it.

## Sources

- CIMdata via Industrial Machinery Digest (2026-06-04), [CIMdata Publishes Executive PLM Market Report](https://www.industrialmachinerydigest.com/software/quality-management-software/cimdata-publishes-executive-plm-market-report)
- Autodesk (2026-02-26), [Autodesk, Inc. Announces Fiscal 2026 Fourth Quarter Results](https://adsknews.autodesk.com/en/?p=55945)
- Dassault Systèmes (2026-02-11), [Q4 revenue growth of 1% with solid operating margin and EPS expansion](https://www.3ds.com/assets/invest/2026-02/dassault-systemes-25q4-earnings_pr_va.pdf)
- Synopsys via Airframer (2025-07-17), [Synopsys completes acquisition of Ansys](https://www.airframer.com/news/release/synopsys-completes-acquisition-of-ansys)
- Digital Engineering 24/7 (2025-03), [Siemens Completes Acquisition of Altair](https://www.digitalengineering247.com/article/siemens-completes-acquisition-of-altair/news)
- Digital Engineering 24/7 (2026), [Engineering Technology Outlook 2026](https://www.digitalengineering247.com/article/engineering-technology-outlook-2026/features)
- Onshape (n.d.), [Onshape pricing](https://www.onshape.com/en/pricing)
- G2 (2026), [2026 Buyer Behavior Report](https://sell.g2.com/2026-buyer-behavior-report)
- OpenAI Help Center (2026), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Google (2025-03-05), [Expanding AI Overviews and introducing AI Mode](https://blog.google/products/search/ai-mode-search/)
- Underneath (2026), [Does rewording a question change AI brand recommendations?](https://underneath.agency/research/ai-prompt-phrasing-study)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [Do Wikipedia and schema make AI assistants recommend a brand?](https://underneath.agency/research/brand-entity-ai-recommendations-study)
- Underneath (2026), [How many “best of” lists cited by AI rank their own brand first?](https://underneath.agency/research/self-promoting-best-lists-study)

---

This is the Markdown twin of https://underneath.agency/resources/engineering-software-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Is an English-only AI visibility check enough if you sell abroad?"
description: "No. In a 12-language European test, AI named local brands far more often when asked in their home language, so English-only checks undercount them."
canonical: "https://underneath.agency/resources/english-only-ai-visibility-audits"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Is an English-only AI visibility check enough if you sell abroad?

Not if you own, or compete with, brands that are strongest in one home market. In the largest multilingual test published so far, AI assistants named locally headquartered brands far more often when asked in the brand’s home language than in English, while global brands barely moved. An English-only audit can tell a local market leader that AI ignores it, and tell a multinational that all is well when it may not be.

## The short version

1. In a test of 66 European brands in twelve languages, asking in a local champion’s home language raised how often AI named it by 0.80 on a 0 to 1 scale, against 0.15 for global brands ([Żatuchin, 2026](https://arxiv.org/abs/2606.23165)).
2. In a second study by the same author, the language of the question accounted for 26.5% of the variation in the tone of a single AI answer about a brand, against 1.5% for the brand itself.
3. A one-brand test in English and Japanese found two ChatGPT models swapped places: one led by 6.79 points in English, the other by 18.75 points in Japanese.
4. Language is not the only blind spot: in our country study, ChatGPT answers from different English-speaking countries shared 0.429 of their brands, against 0.594 from the same country.

## Why would the language of a question change what AI recommends?

AI assistants search the web in the language they are asked, and each language surfaces different sources and brands. [Żatuchin](https://arxiv.org/abs/2606.23165) put the same buyer and reputation questions to GPT-5.4, Gemini 3.1 Pro and Perplexity Sonar Pro. The questions covered 66 brands from eleven Northern, Baltic and Central European markets, asked in twelve languages. The study collected 35,640 answers in April and May 2026, all from assistants that searched the web before answering.

The answers about one brand were not translations of each other. They differed in tone, in the sources behind them and, most of all, in which brands they named. The author notes the tone scores are noisy on long answers, so the direction matters more than the exact size.

Sources shifted at the margin. Wikipedia was the most-cited site in 11 of the 12 languages. In Lithuanian, the national business daily vz.lt edged it out, with 4.38% of citations. The pattern suggests local media matter most in smaller languages. Which engines switch to local sources is covered in [our guide to local-language citations](https://underneath.agency/resources/do-ai-engines-cite-local-language-sources).

## Which brands does an English-only audit miss?

It misses local champions: brands headquartered in, and strongest in, one market. The study split 65 brands into 24 local champions and 41 global or pan-European brands. It then compared answers to buyer questions such as “Who are the leading [industry] companies in [market]?” in English and in each brand’s home language.

Switching to the home language raised a local champion’s share of answers that named it by 0.80, on a scale where 1 means named every time. For global brands the rise was 0.15. A local champion was rarely named in English answers and named in almost every home-language answer. A global brand was named at similar rates in either language.

The author spells out the business risk. A team auditing only in English will conclude “the model ignores us” for a brand that is the default answer at home. The same audit will give a falsely reassuring picture of a multinational whose home-market coverage matters more.

## Does language change how AI describes you, or only whether it names you?

Mostly whether it names you; tone moved far less than visibility. On tone, the home language added 6.2 points for local champions and 9.7 points for global brands on a 0 to 100 scale. That is a mild warming for both groups, with no reversal.

Tone is still not identical across languages. In the main study, 45% of brands showed a tone gap above 0.15 between their English and home-language answers. In a smaller set of 20 brands asked about in several local languages, 90% showed such a gap in at least one language.

So the description of your brand can drift by language, but the bigger commercial effect is on being named at all.

## How large is the language effect next to other causes of variation?

In one detailed test, language outweighed every other factor except chance. A [follow-up study by the same author](https://arxiv.org/abs/2607.13304) used 12,933 answers about 20 Central and Eastern European brands, in eight languages, from three AI models. It split the variation in each answer’s score into its causes.

The language of the question accounted for 26.5% of the variation in a single answer. The brand itself accounted for 1.5%. A brand’s standing against its rivals held roughly steady across AI models and across rewordings of the question, but not across languages. The part tied to a particular brand in a particular language was 8.6%.

The paper’s conclusion is blunt: a measurement that never varies the language is blind to the one factor that most reorders a brand’s score. One caution: the score measured was the tone of each answer, and 91.9% of answers scored as exactly neutral, so this result describes tone more than visibility.

## Does the effect show up outside Europe?

Early evidence says yes, but it rests on a single brand. [Kato and colleagues](https://arxiv.org/abs/2609.11915) asked two OpenAI models 56 product questions, half in English and half in Japanese. Each model answered every question 20 times, giving 2,240 answers in one collection window. The team tracked how often one brand, the web-highlighting tool Glasp, was named.

The ranking of the two models flipped by language. GPT-4o named the brand more often by 6.79 percentage points in English. GPT-5.6 Luna named it more often by 18.75 points in Japanese. A team testing only in English would have drawn the wrong conclusion about which assistant favored them in Japan. Chinese assistants raise a related question, covered in [our guide to Chinese and Western AI models](https://underneath.agency/resources/chinese-vs-western-ai-brand-visibility).

Location matters even within one language. Our [country study](https://underneath.agency/research/ai-recommendations-by-country-study) kept every question in English and changed only the country. Two ChatGPT answers from the same country shared 0.594 of their brands; two from different countries shared 0.429. In the UK, local brands made up 25.4% of ChatGPT’s picks with the location set, and 49.1% when the question also named the country.

## What should you do about it?

Audit in the languages and locations your buyers actually use, starting with the markets where you are strongest. In practice:

1. List your priority markets and the language buyers there use for your category. Run your buyer questions in that language, not only in English.
2. Set the location to each market as well. Our data show location changes answers even when the language stays the same.
3. Translate the buyer’s need, not just the words. A [survey of measurement methods](https://arxiv.org/abs/2609.06811) warns that a translated prompt may not carry over local availability, vocabulary or regulation. Have someone in each market review the prompts.
4. Report results by language and market. Averaging a Lithuanian result into one global number hides the very gap you are looking for.
5. Track more than one assistant. In the European study, how repeatable answers were depended far more on the assistant than on the language: Perplexity scored 0.904 on a 0 to 1 scale, against 0.952 for Gemini.
6. If you are a multinational, treat a reassuring English result as unconfirmed until you have checked your home markets in their own languages. [Our guide to GEO across languages](https://underneath.agency/resources/geo-strategy-for-global-brands-across-languages) covers the work that follows.

If you want help setting up tracking across languages and markets, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The evidence is solid for Northern and Central Europe and thin almost everywhere else. The main gaps:

- The main study covers 66 brands in eleven European markets in one window in April and May 2026. Asian, Latin American and many other languages are untested, apart from one English and Japanese brand test.
- The author of both European studies is affiliated with Rankfor.AI, a company that sells AI brand monitoring, and discloses it. Brands were sorted into local and global groups by hand.
- These studies measure what AI says, not what buyers think or buy. No study yet links a language gap in AI answers to lost revenue.
- A small pilot that crossed language and location had 234 usable runs, and its control category did not repeat the effect. How language and location interact is still open.
- AI models change often. Every figure here is tied to specific model versions and dates.

## Frequently asked questions

### Do AI assistants answer differently in other languages?

Yes. In 35,640 answers across twelve European languages, the same brand was described and recommended differently depending on the language, and the largest difference was in whether local brands were named at all.

### Is English-only AI visibility tracking enough for an international brand?

For a global brand it may be roughly fair; for a local champion it is not. In one large European study, asking in the home language raised a local champion’s naming rate by 0.80 on a 0 to 1 scale, against 0.15 for global brands.

### Should we translate our AI tracking prompts?

Yes, but translate the buyer’s need rather than the literal words. Local products, rules and vocabulary can change what a good answer is, so a reviewer from each market should check the prompts.

### Does the country setting matter if every question is in English?

Yes. In our study of the US, UK, Canada and Australia, ChatGPT answers from different countries shared 0.429 of their brands, against 0.594 for two answers from the same country.

## Sources

- Żatuchin (2026), [The Language Blind Spot: How Query Language and Brand Recognition Tier Shape AI-Constructed Brand Reputation Across Twelve European Languages](https://arxiv.org/abs/2606.23165), arXiv:2606.23165.
- Żatuchin (2026), [Where Does the Noise Come From? A Variance-Components Decomposition of Non-Determinism in LLM Brand Answers](https://arxiv.org/abs/2607.13304), arXiv:2607.13304.
- Kato, Honma and Kato (2026), [Generative Marketing Mix Modeling: A Causal Inference Framework Linking GEO and GEM to Business Impact](https://arxiv.org/abs/2609.11915), arXiv:2609.11915.
- Martinez (2026), [Measuring GEO Visibility: Prompt Corpora Define the Answer Market](https://arxiv.org/abs/2609.06811), arXiv:2609.06811.
- Underneath (2026), [Same question, four countries: do AI recommendations change?](https://underneath.agency/research/ai-recommendations-by-country-study)

---

This is the Markdown twin of https://underneath.agency/resources/english-only-ai-visibility-audits. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Are AI assistants shaping enterprise software shortlists?"
description: "Increasingly, yes. Enterprise buyers research vendors with AI assistants, often private ones at work, before any seller hears about the deal."
canonical: "https://underneath.agency/resources/enterprise-software-shortlists-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Are AI assistants now shaping which enterprise software gets shortlisted?

Increasingly, yes: enterprise buyers now use AI assistants to gather information on vendors before they talk to any of them, and many do it inside the company’s own Copilot or ChatGPT workspace. For a vendor selling six- and seven-figure contracts, the prize is not website traffic but a place on the early list that a buying committee then tests for months. That place depends mostly on what independent sources say about you, and buyers still check it with people before they sign.

## The short version

1. Enterprise buyers lean on AI more than smaller firms: in [G2’s 2025 survey](https://research.g2.com/cmos-2025-buyer-behavior-report-research-g2), buyers at companies with 1,000 to 5,000 employees named review sites (56%) and AI search (55%) as their top two research sources, against 38% and 35% for small businesses.
2. Much of that research happens behind the firewall: [Forrester](https://www.forrester.com/blogs/b2b_buyers_make_zero_click_buying_number_one/) found 61% of business buyers use private AI tools provided by their organization, and they are four times as likely as consumers to use Microsoft Copilot.
3. AI is a starting point, not the final word: in [Gartner’s](https://www.gartner.com/en/newsroom/press-releases/2026-05-20-gartner-survey-finds-sixty-nine-percent-of-b-two-b-buyers-turn-to-sales-reps-to-validate-ai-generated-insights) survey of 645 B2B buyers, 45% used generative AI in a recent purchase, mainly to gather information on vendors, and 69% prefer to validate AI-generated insights with sales reps.
4. The deals are large and slow: [Forrester](https://www.forrester.com/press-newsroom/forrester-2026-the-state-of-business-buying/) found that 78% of buyers making purchases of $10 million or more run a trial first, and in North America the average buying cycle in [6sense’s 2025 study](https://customerthink.com/research-round-up-6sense-study-provides-critical-insights-on-b2b-buyer-behavior/) was 11.1 months.
5. One enterprise customer can be worth millions a year: [ServiceNow](https://s205.q4cdn.com/537566246/files/doc_news/ServiceNow-Reports-Second-Quarter-2026-Financial-Results-2026.pdf) ended the second quarter of 2026 with 658 customers paying more than $5 million in annual contract value.

## Who signs an enterprise software contract, and what can one account pay?

A large committee buys it, with procurement involved from the start, and one won account can pay millions a year.

Enterprise software is the slice of business software sold to large organizations: platforms for IT, finance, HR, data and customer service, sold through sales teams rather than a credit card form. The overall market is growing fast. [Gartner](https://www.gartner.com/en/newsroom/press-releases/2025-10-22-gartner-forecasts-worldwide-it-spending-to-grow-9-point-8-percent-in-2026-exceeding-6-trillion-dollars-for-the-first-time) forecast worldwide software spending of $1,433,037 million in 2026, up 15.2%, and said generative AI features “are now ubiquitous across software already owned and operated by enterprises.”

The buyer is a group, not a person. [Forrester’s 2026 study](https://forrester.com/blogs/state-of-business-buying-2026) found that 13 internal stakeholders and nine external participants influence the average business purchase, and that procurement staff are decision-makers in 53% of buying cycles. When the purchase includes generative AI features, the buying group doubles, from seven members to 14.

Public filings show what one enterprise customer is worth:

| Company | Disclosed figure |
|---|---|
| ServiceNow, Q2 2026 | 658 customers with more than $5 million in annual contract value; 123 deals over $1 million in net new annual contract value in the quarter |
| [Snowflake](https://www.01net.it/snowflake-reports-financial-results-for-the-second-quarter-of-fiscal-2027/), quarter to July 31, 2026 | 828 customers with more than $1 million in trailing 12-month product revenue; net revenue retention of 126% |

Net revenue retention of 126% means existing customers, as a group, spent 26% more than a year earlier. In enterprise software the first contract is often the smallest one, so winning a seat at the first evaluation can be worth far more than the opening deal.

This article is about those large, committee-led deals. For self-serve and mid-market software, where trials and seat-based plans drive revenue, see [how B2B SaaS companies generate revenue from AI search](https://underneath.agency/resources/b2b-saas-revenue-from-ai-search). Health technology sold to employers and health plans is covered in [our guide for healthtech vendors](https://underneath.agency/resources/healthtech-employers-payers-ai-search).

## Where do AI assistants sit in an enterprise software evaluation?

At the research and shortlist stage, often inside tools the buyer’s employer provides, and before sellers are contacted.

The survey evidence is consistent across firms:

- In Forrester’s Buyers’ Journey Survey of nearly 18,000 global business buyers, 94% reported using AI during their buying process.
- Much of it happens in tools the employer supplies: according to Forrester, 61% of business buyers, a majority, use private AI tools their organization provides. They are twice as likely as consumers to use ChatGPT and four times as likely to use Microsoft Copilot, and more than half use private versions behind their firewall.
- In G2’s 2025 survey of 1,100 decision-makers, AI chatbots were the single source that most influenced vendor shortlists, at 17.1%, ahead of review sites (15.1%), vendor websites (12.8%) and market research firms (10.6%).
- In Gartner’s survey, buyers used an average of seven information sources during a recent purchase.

The private-tool finding matters for how visibility works. Microsoft documents that when web search is used, [Microsoft 365 Copilot](https://learn.microsoft.com/en-us/copilot/microsoft-365/microsoft-365-copilot-privacy) “generates a search query that it sends to the Bing Search service.” OpenAI documents that [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search) “typically rewrites your query into one or more targeted queries” that it sends to search providers. So an enterprise buyer’s private assistant still reads the public web; our inference is that a vendor’s visibility in Bing deserves as much attention as in Google.

Analyst research is changing too. In [TrustRadius’s 2026 survey](https://www.demandgenreport.com/industry-news/news-brief/trustradius-ai-has-changed-how-buyers-research-but-not-what-they-trust/53853/) of technology buyers, analyst reports were used by only 13% of buyers to make their purchase decisions, a 63% decrease since 2022. Gartner itself says clients now use its own AI tool, AskGartner. We infer that analyst opinion now reaches many buyers through AI summaries rather than full reports.

## Which questions do enterprise buyers put to AI assistants?

Questions about fit at scale, security, integrations, total cost and proof from peers. We wrote the committee-style prompts below as illustrations; they were not captured from real enterprise buyers.

| Concern | Illustrative prompt |
|---|---|
| Category at scale | “Which IT service management platforms are used by global banks with more than 50,000 employees?” |
| Alternatives | “What are the alternatives to Workday for a manufacturer running SAP?” |
| Comparison | “Snowflake vs Databricks for a regulated insurer: governance, cost and lock-in?” |
| Security and compliance | “Which customer data platforms are FedRAMP authorized and support EU data residency?” |
| Integrations | “Does this procurement suite integrate with Oracle Fusion and Coupa?” |
| Scalability and risk | “Has this vendor had major outages, and what do its uptime commitments say?” |
| Proof | “Which Fortune 500 companies use this platform, and what do reviewers say about implementation?” |
| Cost | “What does a 5,000-seat deployment of this platform typically cost over three years?” |

Each of these questions sends the assistant to look for proof. When we logged the searches behind ChatGPT’s answers to 80 buyer questions in [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study), 43.8% of answers had gone looking for a named publication, ranking or award. In enterprise software, we infer, those named sources include analyst evaluations, peer-review platforms and industry press.

## How does AI visibility turn into enterprise pipeline?

Through the first shortlist: named early, tested by a committee, proven in a trial, then won as a multi-year contract.

1. **Early research.** A team member asks an assistant to map the category and summarize options. G2 found AI chatbots were the top influence on shortlists.
2. **Shortlist.** Buyers narrow the field quickly. TrustRadius found 83% of buyers shortlisted three or fewer products.
3. **Validation.** The committee checks the assistant’s view with people. Gartner found 69% prefer to validate AI-generated insights with sales reps, and 51% say they are more likely to encounter misleading information from generative AI, while 49% say the same of a sales rep.
4. **Proof.** More than 60% of business buyers use some form of trial, and 78% of buyers making $10 million-plus purchases run one first, according to Forrester.
5. **Contract and expansion.** The winner signs a multi-year deal and, as Snowflake’s retention figure shows, often grows within the account.

The value arrives at steps four and five, months after the AI answer. That is why we suggest enterprise vendors judge AI visibility by its effect on pipeline: opportunities that name an AI assistant as a source, and win rates where the vendor was on the first list. Traffic counts miss most of it, since the research happens in private tools that send no clicks.

## What earns an enterprise vendor a mention in an assistant’s answer?

Mostly independent evidence the assistant can find and check; the platforms do not publish their selection rules.

Documented by the platforms: Google says AI Overviews and AI Mode [may run several related searches](https://developers.google.com/search/docs/appearance/ai-features) across subtopics, and Microsoft and OpenAI document that Copilot and ChatGPT turn prompts into web searches. None of them says how vendors are chosen for an answer.

Observed in studies:

- **Independent coverage counts most.** In [our brand-entity study](https://underneath.agency/research/brand-entity-ai-recommendations-study), each tenfold increase in the number of independent sites naming a brand in the pages an assistant cited went with 4.7 times the odds of being recommended. It was the strongest predictor we measured.
- **Google rankings are only part of it.** In [our rankings study](https://underneath.agency/research/ai-citations-google-rankings-study), only 8.3% of the pages ChatGPT cited ranked in Google’s top 10 for the question.
- **Readable sites help, according to one vendor study.** In a [study by ora research](https://arxiv.org/abs/2609.34951) across 1,056 businesses, agent-ready businesses had answers built from their own pages 78% of the time against 56%. The authors sell tools in this area.

Our inference for enterprise software: the trust factors are analyst evaluations, peer reviews from companies of similar size, named customer references, public security and compliance documents, integration listings in partner marketplaces, and clear pricing logic. For security vendors specifically, see [cybersecurity software and AI search](https://underneath.agency/resources/cybersecurity-software-revenue-from-ai-search).

## What does it cost to be missing from the early shortlist?

Possibly a whole buying cycle, because enterprise buyers decide their shortlist early and renew slowly.

The cost is inferred from buying behavior, not measured directly:

- **Lists form early.** In 6sense’s study, 94% of buyers ranked their shortlist vendors before contacting any of them, and first contact came when buyers were 61% of the way through their buying process. A vendor absent from the AI-assisted research may never hear the deal existed.
- **Cycles are long.** With North American buying cycles averaging 11.1 months, a missed evaluation can mean waiting until the winner’s contract comes up for renewal, we infer. Security teams face this at SIEM migrations, covered in [our guide for SIEM vendors](https://underneath.agency/resources/siem-enterprise-pipeline-ai-search).
- **Buyers switch for AI.** In G2’s survey, half of enterprise buyers at companies with 1,000 to 5,000 employees said they had switched vendors for better AI, as did 41% at companies with more than 5,000. Incumbents cannot assume the renewal is safe either.

How much pipeline any single vendor loses this way has not been published. Anyone quoting a precise figure is estimating.

## What does GEO mean for a vendor selling large enterprise contracts?

Generative engine optimization (GEO) makes your company easy for a buying committee’s assistant to understand, verify and recommend.

For enterprise software, that usually means:

1. **Entity clarity.** Consistent product names, categories and descriptions across your site, partner marketplaces, review profiles and reference sites, so an assistant does not confuse products or editions.
2. **Proof that can be checked.** Public trust and compliance pages (SOC 2, ISO 27001, FedRAMP status, data residency), uptime history, and integration documentation that is readable without a login.
3. **Third-party coverage.** Analyst evaluations where they apply, peer reviews from enterprise customers, industry press, and named case studies that other sites can cite.
4. **Answer-ready pages.** Plain pages for enterprise questions: migration from a named competitor, total cost logic, implementation timelines, and honest comparisons. The evidence on competitor comparisons is summarized in [do comparison pages help B2B brands get cited by AI](https://underneath.agency/resources/do-comparison-pages-help-b2b-ai-citations).
5. **Coverage across assistants.** Check ChatGPT, Copilot (and Bing behind it), Gemini, Google’s AI features and Perplexity, because your buyers’ employers choose the tool.
6. **Pipeline measurement.** Add AI assistants to “how did you hear about us” fields, ask in discovery calls which sources shaped the list, and track AI visibility over repeated runs, since a single answer is only a sample.

No GEO program can guarantee that a committee’s assistant recommends you. What it can do is make sure that when a buyer’s assistant goes looking, it finds accurate, verifiable evidence about you.

## What is still unknown about AI and enterprise shortlists?

Whether AI visibility wins enterprise contracts: buyers clearly research vendors with AI, but no study links it to signed deals.

- **No controlled link to revenue.** No published study traces enterprise contracts back to AI answers. Of everything here, outcomes are the least proven; see [does AI visibility drive business results](https://underneath.agency/resources/does-ai-visibility-drive-business-results).
- **Private tools are hard to observe.** What a buyer’s company Copilot answers is not visible to vendors, and its admins may switch web search off.
- **Surveys mix segments.** Forrester, Gartner and TrustRadius cover many kinds of business purchases; only some figures are broken out for large companies. G2 and TrustRadius run review platforms and have an interest in these findings.
- **Selection is a black box.** Answers vary between runs and between assistants, and the platforms do not publish ranking rules.

## What should an enterprise vendor check before its next large evaluation begins?

Put your buying committees’ questions to the assistants they use, then compare the answers with your pipeline.

A useful first review covers three things: whether you are named for your core categories and use cases, which competitors and third-party sources appear instead, and whether security, integration and pricing facts about you are correct. That shows which proof is missing before a committee goes looking for it.

For vendors whose pipeline depends on making the first list for multi-year contracts, we offer that review: [ask us to map how assistants present you to buying committees](https://underneath.agency/contact). We will show how assistants describe you against competitors, which evidence gaps keep you off early shortlists, and what would strengthen your place before procurement gets involved. Our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page covers the longer program that follows, such as public trust and compliance pages, enterprise peer reviews and checks across the assistants buyers’ employers provide.

## Frequently asked questions

### Do enterprise buyers use ChatGPT or Microsoft Copilot to research vendors?

Both, and often private versions at work. Forrester found 61% of business buyers use private AI tools from their organization and are four times as likely as consumers to use Microsoft Copilot.

### Does a Gartner Magic Quadrant placement still matter for AI visibility?

Probably, though no study measures it directly. Analyst reports were used by only 13% of technology buyers in TrustRadius’s 2026 survey, but assistants often search for named rankings and publications, so analyst coverage can reach buyers through AI answers.

### Should we measure AI search by website traffic?

Not for enterprise deals. Much of the research happens in private assistants that send no clicks, so measure shortlist presence, sourced pipeline and win rates instead.

### How is this different from GEO for mid-market SaaS?

The buyer is a committee, the cycle runs close to a year, and proof matters more than price. Mid-market SaaS turns AI visibility into trials and signups; enterprise software turns it into a seat at a long evaluation.

## Sources

- G2 (2025), [Proving Value in the Age of AI: 2025 Buyer Behavior Report](https://research.g2.com/cmos-2025-buyer-behavior-report-research-g2)
- Forrester (2025), [B2B Buyers Make Zero-Click Buying Number One](https://www.forrester.com/blogs/b2b_buyers_make_zero_click_buying_number_one/)
- Forrester (2026-01-21), [Forrester’s 2026 Buyer Insights: GenAI Is Upending B2B Buying](https://www.forrester.com/press-newsroom/forrester-2026-the-state-of-business-buying/)
- Forrester (2026-01), [The State Of Business Buying, 2026](https://forrester.com/blogs/state-of-business-buying-2026)
- Gartner (2026-05-20), [Gartner Survey Finds 69% of B2B Buyers Turn to Sales Reps to Validate AI-Generated Insights](https://www.gartner.com/en/newsroom/press-releases/2026-05-20-gartner-survey-finds-sixty-nine-percent-of-b-two-b-buyers-turn-to-sales-reps-to-validate-ai-generated-insights)
- Gartner (2025-10-22), [Gartner Forecasts Worldwide IT Spending to Grow 9.8% in 2026](https://www.gartner.com/en/newsroom/press-releases/2025-10-22-gartner-forecasts-worldwide-it-spending-to-grow-9-point-8-percent-in-2026-exceeding-6-trillion-dollars-for-the-first-time)
- CustomerThink (2025), [Research Round-Up: 6sense Study Provides Critical Insights on B2B Buyer Behavior](https://customerthink.com/research-round-up-6sense-study-provides-critical-insights-on-b2b-buyer-behavior/)
- Demand Gen Report (2026-07-30), [TrustRadius: AI Has Changed How Buyers Research, But Not What They Trust](https://www.demandgenreport.com/industry-news/news-brief/trustradius-ai-has-changed-how-buyers-research-but-not-what-they-trust/53853/)
- ServiceNow (2026-07-22), [ServiceNow Reports Second Quarter 2026 Financial Results](https://s205.q4cdn.com/537566246/files/doc_news/ServiceNow-Reports-Second-Quarter-2026-Financial-Results-2026.pdf)
- Snowflake, via 01net (2026-09-02), [Snowflake Reports Financial Results for the Second Quarter of Fiscal 2027](https://www.01net.it/snowflake-reports-financial-results-for-the-second-quarter-of-fiscal-2027/)
- Microsoft Learn (2026), [Data, Privacy, and Security for Microsoft 365 Copilot](https://learn.microsoft.com/en-us/copilot/microsoft-365/microsoft-365-copilot-privacy)
- OpenAI Help Center (2025), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- Finder, Elovic, Shalev and Yosef (2026), [AX is the New AEO](https://arxiv.org/abs/2609.34951), arXiv:2609.34951.
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [Do Wikipedia and schema make AI assistants recommend a brand?](https://underneath.agency/research/brand-entity-ai-recommendations-study)
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude cite pages that rank?](https://underneath.agency/research/ai-citations-google-rankings-study)

---

This is the Markdown twin of https://underneath.agency/resources/enterprise-software-shortlists-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How do ERP vendors get on the shortlist when buyers ask AI?"
description: "By being on the long list AI assistants draw from analyst, consultant and comparison sources, with clear industry fit, partners and implementation facts."
canonical: "https://underneath.agency/resources/erp-software-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How do ERP vendors get on the shortlist when buyers ask AI?

By being well represented in the sources AI assistants check when an ERP buyer asks for options: analyst reports, independent consultants’ rankings, comparison sites, partner pages and your own clear statements of industry fit, cost and implementation. ERP buyers research for months, in committees, with consultants at their side, so AI rarely closes the deal. It increasingly shapes the long list the deal starts from, and in ERP one won contract can be worth more than years of website traffic.

## The short version

1. A large replacement wave is under way: [SAP](https://news.sap.com/2020/02/sap-s4hana-maintenance-2040-clarity-choice-sap-business-suite-7/) ends mainstream maintenance for the core applications of SAP Business Suite 7 at the end of 2027, with extended maintenance costing a premium of two percentage points, while committing to S/4HANA until 2040.
2. Demand is moving now: [Microsoft](https://www.microsoft.com/en-us/investor/events/fy-2026/earnings-fy-2026-q4) reported Dynamics 365 revenue up 13% in its June 2026 quarter, with ERP bookings healthy while CRM saw longer sales cycles.
3. Each deal is large and slow: [Software Path’s analysis](https://softwarepath.com/guides/erp-report) of about 1,000 selection projects found an average budget of $9,000 per user and 17 weeks spent choosing a system (2022 data, the latest it publishes).
4. In [our study of the searches AI assistants run](https://underneath.agency/research/ai-hidden-searches-study), ChatGPT averaged 3.7 searches per buyer question and, in 43.8% of answers, searched for a named publication, ranking or award, the kind of source ERP buyers already lean on.
5. ERP has a deep layer of independent comparison: [Top10ERP](https://www.top10erp.org/) says it has supported over 950,000 manufacturing and distribution businesses, [Panorama Consulting](https://www.panorama-consulting.com/resource-center/erp-software-research-and-reports/) publishes annual top-10 rankings, and [SAP’s ERP page](https://www.sap.com/products/erp.html) cites three Gartner Magic Quadrants.

## Who signs off on an ERP purchase, and what is one contract worth?

A committee led by finance and IT buys it, often with an independent consultant and an implementation partner involved.

ERP touches finance, operations, supply chain and IT at once, so the decision sits with a group: typically the chief financial officer, the chief information officer, operations leaders and, for larger projects, the board. Independent selection consultants such as Panorama are often hired to run the process, and an implementation partner or systems integrator is chosen alongside the software. That makes ERP one of the few software categories where several outside advisers stand between the vendor and the buyer.

The buyer profile spans a wide range. Software Path found that the system most often being replaced in its sample was QuickBooks, meaning a growing company buying its first ERP, and that 14% of companies were moving off homegrown or outgrown systems. At the other end are global firms moving from SAP ECC. Cloud is now the default expectation: almost 97% of companies in the same data were considering a cloud-based system, and only 3% were looking exclusively for on-premise software.

What a customer is worth is high and long-lived. With $9,000 of budget per user and 26% of employees using the system on average, by that arithmetic even a mid-sized company’s project runs well into six figures, before multi-year subscription renewals. Panorama warns that the approved figure understates the real spend: “an implementation estimate is a scope document expressed in dollars,” and hidden costs surface later. For the vendor, that means one won ERP contract can be worth more than a large volume of low-intent website traffic.

## Why are so many companies evaluating ERP now?

Because of maintenance deadlines, the move to cloud and the promise of AI built into the system.

SAP’s 2027 deadline forces its large on-premise base to decide: move to S/4HANA, typically through RISE with SAP, pay for extended maintenance, or reconsider vendors. SAP’s own [RISE page](https://www.sap.com/products/erp/rise.html) pitches the migration with AI-enabled assistants for custom code, data and testing, and argues that AI embedded in cloud ERP can act within live business processes. Microsoft’s comment that ERP bookings stayed healthy while CRM slowed suggests that ERP spending is holding up even when other application budgets are tight. At the smaller end, [Odoo](https://www.odoo.com/) says it has 28 million users, showing how crowded the market for growing companies has become.

Every one of these decisions starts with research, and each involves questions buyers can now put to an AI assistant before they talk to anyone.

## Where do AI assistants sit in ERP research?

Early, at the long list and the education stage, with buyers checking the answers with people later.

There is no ERP-specific survey of AI use yet. The best recent evidence is general: in a [Gartner survey](https://www.gartner.com/en/newsroom/press-releases/2026-05-20-gartner-survey-finds-sixty-nine-percent-of-b-two-b-buyers-turn-to-sales-reps-to-validate-ai-generated-insights) of 645 business buyers, 45% said they used generative AI during a recent purchase, mainly to gather information on vendors and products, and 69% preferred to validate what the AI told them with a sales representative. Buyers in the same survey worried about misleading AI answers. That pattern fits ERP well: buyers use AI to frame the category and build a long list, then lean on consultants, references and demos to check it.

A reasonable expectation is that consultants and implementation partners use assistants too, to scan options for a client in an unfamiliar industry. Because one adviser can shape several selections, being described accurately in those answers matters more in ERP than the raw number of buyers using AI suggests.

## Which questions do ERP buyers ask AI assistants?

Questions about industry fit, comparisons, migration paths, cost, implementation partners and compliance.

We wrote the example prompts below to show how a CFO, CIO or selection consultant might phrase an ERP question; none comes from observed data:

- Industry fit: “Best ERP for a $150M discrete manufacturer with three plants and make-to-order production.”
- Comparison: “NetSuite vs Dynamics 365 Business Central for a distributor outgrowing QuickBooks.”
- Migration: “Should we move from SAP ECC to RISE with SAP or evaluate other vendors before 2027?”
- Cost: “What does a cloud ERP implementation cost for 60 users, including partner fees?”
- Partners: “Which partners implement Acumatica for food manufacturers in Texas?”
- Compliance: “Which ERP systems support FDA 21 CFR Part 11 for a medical device company?”

Each of these questions points an assistant toward a different kind of source: industry pages, comparisons, the vendor’s own migration documentation, cost guides, partner directories and compliance documentation.

## How does an AI answer become an ERP deal?

Through the long list: the answer shapes who is considered, then consultants, demos and partners decide the rest.

**Long list.** The assistant names a handful of systems for the buyer’s industry and size. Vendors missing from that list may never receive the request for proposals.

**Consultant and request for proposals.** A selection consultant or internal team narrows the list. Here independent rankings matter, which is why vendors cite analyst reports on their own pages.

**Demos and references.** Scripted demos and reference calls test the claims. We infer that the AI answer’s influence shows up here not as a tracked referral but as a buyer arriving with a view of your strengths and weaknesses.

**Partner selection and contract.** The buyer chooses an implementation partner, often from the vendor’s directory. For the vendor, the deal is a multi-year subscription; for the partner, a services engagement. Both depend on being on the original list.

## Why does an assistant put some ERP systems on a long list and leave others off?

The companies behind the assistants disclose little; our studies show them hunting for named rankings, reviews and recent pages.

**What Google describes.** By Google’s account, AI Mode applies a [“query fan-out” technique](https://blog.google/products/search/ai-mode-search/): it issues multiple related searches across subtopics and data sources, then combines the results. A single ERP question can therefore pull in analyst coverage, consultant rankings, comparison sites and vendor pages at once.

**Observed in our studies.** In our hidden-searches study, ChatGPT looked for reviews in 46.2% of answers, and when one of its searches named a source, the answer cited that source 44.0% of the time, against 8.1% when no search named it. That is an association, but it suggests that being the source an assistant searches for by name, such as a well-known ranking, matters. [Our study of AI citations and Google rankings](https://underneath.agency/research/ai-citations-google-rankings-study) found only 8.3% of the pages ChatGPT cited ranked in Google’s top 10 for the question, so strong search rankings alone do not secure a place in the answer. And [our freshness study](https://underneath.agency/research/ai-source-freshness-study) found the assistants cited pages first published about half as long ago as Google’s top 10 for the same questions (a ratio of 0.50), which favors current editions of rankings and recently updated guides.

**Trust factors specific to ERP.** Buyers weigh industry fit, analyst and consultant assessments, implementation track record, partner coverage in their region, total cost of ownership and references from similar companies. A reasonable expectation is that assistants answering ERP questions lean on the same public evidence: Gartner Magic Quadrants as summarized by vendors and press, consultant reports such as Panorama’s, comparison sites such as Top10ERP, partner directories and detailed case studies.

## What does a missed long list cost an ERP vendor?

The whole evaluation, usually for many years.

Because ERP is replaced rarely, a company that chooses a rival in 2026 may not return to the market for a decade. We infer that missing the long list during the 2027 SAP migration window, or during a growth company’s first ERP purchase, costs the vendor not only that contract but also the renewals, expansion modules and partner services that follow. And because a single consultant can carry the same long list into several projects, an omission can repeat. For the broader link between AI answers and pipeline, see [our article on AI answers and pipeline](https://underneath.agency/resources/ai-answers-pipeline-revenue).

## What does GEO involve for an ERP vendor selling through consultants and partners?

It makes your industry fit, costs and partners easy to verify, without promising a place on any long list.

In ERP, generative engine optimization (GEO) comes down to six workstreams, one of them run with your implementation partners:

1. **Industry pages with specifics.** Publish pages for each industry and company size you serve, with modules, compliance features and named customer examples, so an assistant can match you to a specific question.
2. **Independent assessments, kept current.** Participate in analyst evaluations and consultant rankings, and keep summaries of the latest editions on your site. How that kind of third-party standing is built is the subject of [how brands build authority for AI search](https://underneath.agency/resources/how-brands-build-authority-for-ai-search).
3. **Migration and comparison guides.** Publish honest guides for buyers leaving SAP ECC, QuickBooks or legacy systems, and fair comparisons. For the evidence on that format, read [our article on comparison pages](https://underneath.agency/resources/do-comparison-pages-help-b2b-ai-citations).
4. **Implementation facts.** State typical timelines, cost ranges and what is included, because buyers ask, and Panorama warns that costs left out of estimates are where ERP budgets go wrong.
5. **A clean partner network.** Keep your partner directory public, current and consistent with partners’ own sites, so answers to “who implements this near me” are right.
6. **Tracking by stage of the selection.** Re-run industry-fit, comparison, migration, cost and partner questions in ChatGPT, Gemini, Perplexity, Copilot and Google, and tie the results to the requests for proposals you receive and your pipeline; [how to design AI visibility tracking](https://underneath.agency/resources/how-to-design-ai-visibility-tracking) explains the setup.

## What don’t we know yet about AI’s role in ERP selection?

Nobody has measured how often assistants shape ERP long lists, or what that influence is worth.

The ERP buying data here is either older (Software Path’s 2022 figures) or vendor-reported (SAP, Microsoft, Odoo, Top10ERP). The Gartner survey covers business buying in general, not ERP. Our own studies record what assistants search for and cite across several industries, with no ERP-only sample and no view of which system a buyer then signs for. Treat AI as a real and growing influence on who gets considered, with its effect on ERP contracts still unmeasured.

## How can an ERP vendor learn whether it is missing from AI-built long lists?

Ask assistants the questions your target industries ask, and see which systems they name and why.

Run industry-fit, comparison, migration, cost and partner questions through the main assistants and Google’s AI features, then line the answers up against the requests for proposals you received and the deals you won, so you can see which long lists you are missing before a request for proposals goes out. To work through that comparison with us, and plan how to close the gaps with your consultants and partners in mind, [talk to us about an ERP visibility audit](https://underneath.agency/contact). Our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page spells out the ongoing work for an ERP vendor, including industry pages, migration guides and a partner directory that matches partners’ own sites.

## Frequently asked questions

### Do enterprise ERP buyers really use ChatGPT to choose a system?

Many business buyers now use generative AI to gather vendor information, but they check what it says with people. For ERP, we expect AI to shape the first long list more than the final choice.

### Do analyst reports still matter if buyers ask AI?

Yes, and possibly more. Assistants in our study often searched for named rankings and publications, and cited the named source far more often when they did.

### Should implementation partners care about AI visibility?

Yes. Buyers ask assistants who implements a given system in their industry and region, and the answer draws on partner directories and partners’ own sites. Consistent, current partner information helps both the vendor and the partner.

### How can an ERP vendor tell if AI is influencing its pipeline?

Ask new prospects and consultants how they built their long list, track AI referrals to industry and pricing pages, and regularly check which vendors assistants name for your buyers’ questions. Then compare that with requests for proposals received and lost.

## Sources

- SAP (2020), [SAP S/4HANA Maintenance Until 2040: Clarity and Choice for SAP Business Suite 7](https://news.sap.com/2020/02/sap-s4hana-maintenance-2040-clarity-choice-sap-business-suite-7/)
- SAP (2026), [SAP ERP](https://www.sap.com/products/erp.html) and [RISE with SAP](https://www.sap.com/products/erp/rise.html)
- Microsoft (2026), [FY26 Q4 earnings](https://www.microsoft.com/en-us/investor/events/fy-2026/earnings-fy-2026-q4)
- Software Path (2022), [ERP Software Report](https://softwarepath.com/guides/erp-report)
- Panorama Consulting Group (2026), [ERP Software Research and Reports](https://www.panorama-consulting.com/resource-center/erp-software-research-and-reports/) and [The Real Cost of Implementing an ERP System](https://www.panorama-consulting.com/real-cost-of-implementing-an-erp-system/)
- Top10ERP (2026), [Top10ERP](https://www.top10erp.org/)
- Odoo (2026), [Odoo](https://www.odoo.com/)
- Gartner (2026), [Gartner Survey Finds 69% of B2B Buyers Turn to Sales Reps to Validate AI-Generated Insights](https://www.gartner.com/en/newsroom/press-releases/2026-05-20-gartner-survey-finds-sixty-nine-percent-of-b-two-b-buyers-turn-to-sales-reps-to-validate-ai-generated-insights)
- Google (2025), [Expanding AI Overviews and introducing AI Mode](https://blog.google/products/search/ai-mode-search/)
- Underneath (2026), [hidden searches by AI assistants](https://underneath.agency/research/ai-hidden-searches-study), [AI citations and Google rankings](https://underneath.agency/research/ai-citations-google-rankings-study) and [source freshness](https://underneath.agency/research/ai-source-freshness-study)

---

This is the Markdown twin of https://underneath.agency/resources/erp-software-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Will AI search change how boards choose an executive search firm?"
description: "Referrals still decide most mandates, but directors now use AI to check firms and partners. Public, verifiable track records decide what those checks find."
canonical: "https://underneath.agency/resources/executive-search-firms-clients-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Will AI search change how boards choose an executive search firm?

It is changing how they check one. Most retained mandates still come from existing clients and referrals, but most public company directors now use generative AI in their board work, and an assistant asked about a firm or a partner can only repeat what is public. For a business built on discretion, the task is to make the record that can be shared, such as announced placements, research and partner expertise, clear enough to be found and repeated accurately.

## The short version

1. Mandates are valuable: Korn Ferry says its search fees are generally one-third of the placed executive’s estimated first-year cash compensation, and Heidrick & Struggles reported average revenue per executive search of $162 thousand.
2. Referrals dominate: Korn Ferry obtains a majority of its new engagements from existing clients or their referrals, according to its fiscal 2026 annual report.
3. Leadership change is steady: US companies announced 2,032 CEO exits in 2025, and outside hires (960) outnumbered internal promotions (881), according to Challenger, Gray & Christmas.
4. Directors use AI: in a Diligent survey of 104 US public company directors, 82% had used generative AI in board work in the past six months, up from 66%.
5. Board seats are scarce: S&P 500 boards appointed 374 new independent directors in 2025, the lowest number since 2016.

This article is about retained search for chief executives, senior leaders and board directors. Agencies that fill temporary, contract and mid-level roles face a different buyer and a different fee, so they are not covered here.

## Who hires an executive search firm, and what is one mandate worth?

A board chair, a nominating committee, a CEO or a private equity partner. A single mandate can be worth six figures, and satisfied clients return.

The fee is set by the role. Korn Ferry’s [fiscal 2026 annual report](https://ir.kornferry.com/sec-filings/all-sec-filings/content/0000056679-26-000021/kfy-20260430.htm) says fee revenue from executive and professional search is “generally one-third of the estimated first-year cash compensation of the placed candidate,” plus a share for expenses, and an uptick fee when final pay is higher. Heidrick & Struggles, in its [second-quarter 2025 results](https://heidrick.mediaroom.com/2025-08-04-Heidrick-Struggles-Delivers-14-Revenue-Growth-in-Q2,-Driving-Strong-Profitability), reported average revenue per executive search of $162 thousand, with 420 search consultants producing $2.3 million each on an annualized basis.

The largest firm shows the scale. In [Korn Ferry’s fiscal 2026 results](https://ir.kornferry.com/sec-filings/all-sec-filings/content/0000056679-26-000017/kfy-20260430xex991q4fy26.htm), Executive Search earned $924.1 million in fee revenue and opened 6,514 new engagements, with 566 consultants at year end.

The same annual report explains how the work is won. “We obtain a majority of our new engagements from existing clients or from referrals by those clients,” it says, and it names reputation, both the firm’s and each consultant’s, as essential to securing them. That is the starting point for any discussion of AI search: it does not replace the referral, it is where the referral gets checked. Agencies filling temporary, contract and permanent roles sell to a different buyer, covered in [our guide for recruiting agencies](https://underneath.agency/resources/recruiting-agencies-employer-clients-ai-search).

## How much leadership change is there for search firms to serve?

A steady flow: about 2,000 CEO exits a year in the US, more of them filled from outside than within.

Challenger, Gray & Christmas, which tracks CEO changes at US companies, counted [2,032 CEO exits in 2025](https://www.challengergray.com/wp-content/uploads/2026/02/Dec25-Challenger-CEO-Report.pdf), down 9% from 2024. Public companies had their most active year on record, with 446 CEO exits. Of the replacements, 960 were external and 881 internal, and every external appointment is a potential search mandate.

2026 is quieter. Through June, Challenger counted [920 CEO exits](https://www.challengergray.com/wp-content/uploads/2026/08/June-26-CEO-Turnover-Report-Final.pdf), down 26% from the same period of 2025, as “boards continue to hold onto the leaders they have.” Tenure keeps shortening: [HRD America](https://www.hcamag.com/us/specialization/recruitment/will-ai-replace-headhunters/589846) cites Russell Reynolds Associates putting average CEO tenure at 7.1 years, down from 8.3 years in 2021.

Board searches are scarcer still. Spencer Stuart’s board index, summarized by the [Harvard Law School Forum on Corporate Governance](https://corpgov.law.harvard.edu/2025/11/03/2025-u-s-board-index/), found S&P 500 boards appointed 374 new independent directors in 2025, and only 50% of boards appointed one at all, down from 58% in 2024. With fewer seats and fewer CEO changes, each invitation to pitch matters more.

## Where does AI enter a board’s choice of search firm?

At the checking stage: directors and CEOs use AI to benchmark firms, partners and track records.

There is no published study of how boards use AI to choose search firms. There is good evidence that directors use AI for board work:

- **Use is now the norm.** In [Diligent Institute’s survey](https://www.diligent.com/resources/blog/dci-board-ai-use-2026) of 104 US public company directors, 82% had used generative AI in board work in the past six months, up from 66% in September 2025.
- **Benchmarking is a common use.** 45% had used it to prepare for board or committee discussions or to benchmark peers, competitors or market trends.
- **Consumer tools are in the mix.** 49% had heard of board members using publicly available or consumer-facing AI tools for board work rather than company-approved systems. Diligent sells board software, so treat this as a vendor’s survey.

The search firms themselves see the shift. Korn Ferry’s annual report lists increased competition from “in-house human resource professionals whose ability to provide job placement services has been enhanced by professional profiles made available on the internet and enhanced social media-based or AI-based search tools.” The industry body is adapting: AESC launched an AI training program for search consultants in June 2026, and HRD America cites a report that 67% of executive search firms were using AI-powered tools at the start of 2026.

Our inference: a nominating committee member who receives three firm names from colleagues is now likely to ask an assistant what each firm is known for, who leads its practice in the sector, and what has been said about it. The answer shapes who gets the first call.

## What do boards and CEOs ask AI about search firms?

Questions about fit, record and risk. We wrote these examples to show the pattern; none is taken from real search logs.

| Buyer | Example question |
|---|---|
| Board chair | “Which executive search firms have the strongest CEO succession practice for mid-cap medical device companies?” |
| Nominating committee | “Who helps boards find first-time directors with cybersecurity experience?” |
| Private equity partner | “Search firms that place CFOs in PE-backed software companies, and their typical fees” |
| CHRO | “Global firm or boutique for a chief supply chain officer search: pros and cons” |
| CEO | “Who leads the industrial practice at [firm], and what searches have they completed?” |
| Any buyer | “Has [firm] had conflicts or confidentiality problems? What do clients say?” |

Two kinds of question stand out. The first is about a partner, not a firm, because clients hire the person who will run their search. The second is about risk. Korn Ferry’s annual report notes that firms with smaller client bases are subject to fewer off-limits arrangements, which limit where a firm can recruit from. For a boutique, being clearly described as free of those conflicts in a sector can be an advantage.

## How does an AI answer turn into a retained mandate?

Through the shortlist: an answer names the firm or partner, the committee invites it to pitch, and a retainer follows.

The path is a referral or AI answer → AI-assisted check on the firm and lead partner → invitation to pitch against two to four firms → retainer signed → search over several months → placement → follow-on work such as assessment, succession planning or board searches.

The economics make the early stage decisive. A firm that is not invited to pitch has no chance at a fee worth a third of an executive’s first-year cash pay. Because Korn Ferry says a small number of consultants hold primary responsibility for each client relationship, the partner’s public profile matters as much as the firm’s. A reasonable expectation is that firms whose partners are clearly tied to sectors and roles, in independent sources, are the ones an assistant can describe well. HR consulting firms face a close version of this test, where the firm must be tied to a named people problem in sources others cite, as [our guide for HR consulting firms](https://underneath.agency/resources/hr-consulting-firms-clients-ai-search) shows.

## What decides which search firms an assistant names?

Mostly how widely and specifically independent sources describe the firm. Platforms say little; most of what we know is observed in other markets.

**Documented by the platform.** OpenAI’s help page on [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search) says ChatGPT may search the web automatically, that responses may include citations, and that cited results “can be incomplete, outdated, or incorrect.”

**Observed in our studies (cross-industry, not executive search).**

- In our [brand entity study](https://underneath.agency/research/brand-entity-ai-recommendations-study), 60.0% of options named by all four assistants had an English Wikipedia article, but most of that advantage went once prominence was accounted for. What mattered more for a search firm’s odds was outside mention: every tenfold rise in the independent sites that named a brand in the cited pages came with 4.7 times the odds of a recommendation.
- Our [hidden searches study](https://underneath.agency/research/ai-hidden-searches-study) found ChatGPT went looking for a specific publication, ranking or award in 43.8% of its answers. Rankings and trade press lists of search firms are therefore part of the evidence an answer can draw on.
- Repeat the same question and the shortlist moves: in our [consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), just 25.2% of the brands ChatGPT named turned up in all five runs. One check of what an assistant says about a firm is a snapshot.

**Our inference for executive search.** Discretion limits the public record. Many searches are never announced, so assistants lean on what is: appointment announcements that credit the firm, research the firm publishes, partner biographies, and trade coverage. Spencer Stuart’s [Board Index](https://www.spencerstuart.com/research-and-insight/-/media/2026/09/USBI2026/2026_US_Spencer_Stuart_Board_Index_Highlights.pdf), now in its 41st year, is an example of research that others summarize year after year, and every summary ties the firm’s name to board composition.

## What does a search firm lose if AI misdescribes or ignores it?

Invitations to pitch, mainly in sectors and roles where the firm is strong but publicly quiet. No study has measured it.

The largest firms are already widely covered. A specialist boutique with a strong record in, say, hospital CEO searches may be invisible if its placements are unannounced and its partners’ pages are thin. If a committee member asks an assistant which firms specialize in that work, the boutique may not appear, even though a colleague recommended it. That is our inference.

Misdescription is the other risk. An assistant that describes a retained search firm as a contingency recruiter, attributes a placement to the wrong firm, or lists a partner who has left sends the wrong signal to a buyer who values precision. Our guide on [fixing wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) covers how to correct the sources behind such errors.

## How does GEO work for an executive search firm?

Generative engine optimization (GEO) makes the shareable part of a firm’s record accurate, specific and easy for assistants to repeat. It cannot promise a mention.

1. **Partner pages that say what each partner does.** Sectors, functions, board work and the kinds of searches led, in plain words, matched on LinkedIn and conference bios.
2. **Announced placements, with consent.** When a client announces an appointment, ask whether the firm may be credited. Trade press reports of completed searches tie the firm to the sector and the role.
3. **Research others cite.** A recurring study of CEO succession, board composition or pay in one sector gives journalists and other sites a reason to name the firm. Our piece on [building the kind of authority AI search picks up](https://underneath.agency/resources/how-brands-build-authority-for-ai-search) explains why that outside citation counts.
4. **Plain process and ethics pages.** How the firm handles confidentiality, off-limits arrangements, candidate data and AI use. Buyers ask these questions, and boards want written answers.
5. **Association and directory profiles.** AESC membership, which covers more than 16,000 professionals in more than 80 countries, and accurate listings in rankings and directories.
6. **A consistent entity.** The same firm description, offices, practices and leadership across the website, press releases and profiles. Our article on [why Wikipedia matters for AI search](https://underneath.agency/resources/why-wikipedia-matters-for-ai-search) explains when an encyclopedia entry is realistic and when it is not.
7. **Regular checks.** Ask the sector and partner questions above in several assistants, several times each quarter, and record who is named and which sources are cited.

## Where does the evidence on AI and executive search run out?

At the mandate: no data yet connects a firm’s AI visibility to invitations to pitch or signed retainers.

- **No study of AI answers about search firms.** We found none that measures which firms assistants name.
- **Director surveys are small.** Diligent’s figures come from 104 directors and describe board work in general, not choosing advisers.
- **Client research is private.** AESC’s 2026 report on how clients evaluate and select firms is available only to members.
- **Our findings come from other markets.** The coverage and consistency results were measured on consumer and business brands.

## Where should an executive search firm start?

With the sectors and roles you most want to lead, checked against what assistants say today.

List the questions a board chair, CEO or deal partner would ask about those searches. Ask them in ChatGPT, Gemini, Perplexity, Claude and Google’s AI features, more than once. Note which firms and partners are named, which sources are cited, and whether anything about your firm is wrong or missing.

If you want more invitations to pitch for the mandates you are best placed to win, [contact us for a review of how AI assistants describe your firm and partners](https://underneath.agency/contact). We will test the sector, role and reputation questions your clients ask, show the sources behind the answers, and plan the pages, research and coverage that make your record easy to verify without compromising discretion. Our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page describes how that work is run over time for a search firm, from partner pages and consented placement credits to regular checks.

## Frequently asked questions

### Do AI assistants only name the largest executive search firms?

The large firms are the most widely covered, so they are named often. In our cross-industry studies, independent coverage of a specific strength was what went with being recommended.

### Can a search firm improve AI visibility without breaching client confidentiality?

Yes. Partner expertise, published research, consented placement announcements and process pages are all shareable without naming confidential searches.

### Do directors really use AI to research advisers?

Directors use AI widely in board work, including benchmarking. No survey yet isolates adviser selection.

### How long does it take to change what AI says about a firm?

Corrections to the firm’s own pages and profiles can show within weeks. Coverage and research build over months and years.

## Sources

- Korn Ferry (2026-06), [Form 10-K for the fiscal year ended April 30, 2026](https://ir.kornferry.com/sec-filings/all-sec-filings/content/0000056679-26-000021/kfy-20260430.htm)
- Korn Ferry (2026-06-23), [Fourth quarter and full year FY’26 results](https://ir.kornferry.com/sec-filings/all-sec-filings/content/0000056679-26-000017/kfy-20260430xex991q4fy26.htm)
- Heidrick & Struggles (2025-08-04), [Heidrick & Struggles Delivers 14% Revenue Growth in Q2, Driving Strong Profitability](https://heidrick.mediaroom.com/2025-08-04-Heidrick-Struggles-Delivers-14-Revenue-Growth-in-Q2,-Driving-Strong-Profitability)
- Challenger, Gray & Christmas (2026-02-04), [December 2025 CEO Turnover Report](https://www.challengergray.com/wp-content/uploads/2026/02/Dec25-Challenger-CEO-Report.pdf)
- Challenger, Gray & Christmas (2026-07-23), [June 2026 CEO Turnover Report](https://www.challengergray.com/wp-content/uploads/2026/08/June-26-CEO-Turnover-Report-Final.pdf)
- Harvard Law School Forum on Corporate Governance (2025-11-03), [2025 U.S. Board Index](https://corpgov.law.harvard.edu/2025/11/03/2025-u-s-board-index/)
- Spencer Stuart (2026-09), [2026 U.S. Spencer Stuart Board Index Highlights](https://www.spencerstuart.com/research-and-insight/-/media/2026/09/USBI2026/2026_US_Spencer_Stuart_Board_Index_Highlights.pdf)
- Diligent Institute (2026), [Board AI use 2026: Director Confidence Index](https://www.diligent.com/resources/blog/dci-board-ai-use-2026)
- HRD America (2026-09-15), [Will AI replace headhunters?](https://www.hcamag.com/us/specialization/recruitment/will-ai-replace-headhunters/589846)
- AESC (2026-03-24), [AESC Releases New Members-Only Report on the Future of Executive Search](https://www.aesc.org/insights/press-release/new-members-only-report-future-executive-search/)
- OpenAI (2026), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Underneath (2026), [Do Wikipedia and schema make AI assistants recommend a brand?](https://underneath.agency/research/brand-entity-ai-recommendations-study)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)

---

This is the Markdown twin of https://underneath.agency/resources/executive-search-firms-clients-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Can fake reviews get a fake brand recommended by AI?"
description: "Yes. In tests of 12 AI models, planted review pages got a non-existent brand recommended in up to 73.8% of products, and no defense tested fixed it."
canonical: "https://underneath.agency/resources/fake-reviews-fake-brands-ai-recommendations"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Can fake reviews on the web make AI assistants recommend a fake brand over ours?

Yes: in a 2026 test of 12 AI models, every one could be pushed into recommending a non-existent brand by fake review pages in what it read. A single planted page at the top of the search results was sometimes enough. The risk is highest in everyday categories where buyers rely on community opinion, and none of the defenses tested solved it.

## The short version

1. [Luo and Chen](https://arxiv.org/abs/2606.13610) tested 12 AI models on 225 real products. Rewriting three top search results got the fake brand recommended in 13.3% to 73.8% of products, depending on the model.
2. One fake page in the first search slot fooled the most vulnerable models in 27% of products; the same page lower down was nearly harmless.
3. When fooled, the AI usually put the fake brand first: it took the top slot in 57% of fooled cases.
4. AI answers lean on review sites: in [our brand reputation study](https://underneath.agency/research/is-it-legit-ai-reputation-study), 88.0% of answers to “Is this brand legit?” cited a review or complaint platform.
5. A simple fix, re-ordering sources by credibility, removed only 17% of fake recommendations in testing.

## Has this happened outside the lab?

Yes, according to Chinese media reports cited by researchers, though no study has measured it in the wild. [Luo and Chen](https://arxiv.org/abs/2606.13610) describe a television exposé on March 15, 2026, China’s annual consumer rights broadcast. It reported paid operators who seeded fake reviews online and could get a fake brand into the top picks of [mainstream Chinese AI assistants](https://underneath.agency/resources/chinese-vs-western-ai-brand-visibility) within hours.

[Wen and colleagues](https://arxiv.org/abs/2606.12439) note that the OECD AI Incident Monitor records a 2026 poisoning incident in China in which AI assistants allegedly recommended fictitious or low-quality products. Both are reports, not measurements. They explain why researchers built a controlled test.

## How easily were AI models fooled in testing?

Easily, and every model tested could be fooled. Luo and Chen took real search results for 225 real products across 15 categories. They froze those results, then swapped the leading real brand name in some pages for an invented brand. Nothing else changed: same web address, same rank, same writing. They did this offline, so no real web page was polluted.

The main results:

| Setup | Share of products where the fake brand was recommended |
|---|---|
| Top three search results rewritten | 13.3% to 73.8%, depending on the model |
| One fake page in the first slot | Up to 27% for the most vulnerable models |
| One fake page in slots two to ten | Only 1–4% |

Placement was severe. When a model was fooled, the fake brand took the first slot in 57% of cases and a top-three slot in 84%. The main test was in Chinese. An English rerun on three categories, with 360 trials using US search results, kept the same pattern: 8 of 12 models landed within 10 points of their Chinese rate.

## Why does one fake review page matter so much?

Because AI assistants lean heavily on the first page they read, and that page is often open to anyone. In Luo and Chen’s data, the first search result was a user-generated page in 52.4% of queries, and 74% of queries had one in the top three. These are forums, Q&A sites and similar pages where anyone with an account can post without editorial approval. Our guide on [how list position sways AI picks](https://underneath.agency/resources/does-list-order-change-ai-recommendations) shows that order matters even without fake pages.

Our own research shows how much AI answers rely on review-style sources. In [our brand reputation study](https://underneath.agency/research/is-it-legit-ai-reputation-study), we asked ChatGPT, Gemini, Perplexity and Google AI Mode whether 79 brands were legitimate on 26 September 2026. Of the answers, 88.0% cited a review or complaint platform. Trustpilot and the BBB alone made up 61.7% of review-platform citations.

Ratings and review counts also move AI picks directly. In an audit using synthetic hotels across twelve AI models, [Baig and colleagues](https://arxiv.org/abs/2606.16344) found that a top guest rating raised the chance of being recommended by 31.6 percentage points. In [our study of ChatGPT’s local picks](https://underneath.agency/research/chatgpt-local-picks-google-profile-study), businesses with more reviews than the local median were 19.5 points more likely to be listed. That held after adjusting for Maps rank and other signals. Anything that inflates those signals can shift the answer.

## Which brands and categories are most exposed?

Categories where buyers rely on taste and word of mouth, not well-known brands. Luo and Chen found the most exposed categories were dining, personal services and supplements. The least exposed were phones and PCs, home appliances and electronics accessories. Dining was the most-fooled category for two thirds of the models.

The pattern tracks what the AI already knows. Where the models broadly agreed on which real brands to recommend, they resisted the fakes. Where their knowledge was thin, they fell for them. Local businesses, restaurants, clinics, salons and small consumer brands sit on the risky side of that line.

## Do smarter or more careful AI models resist?

No: bigger models and extra reasoning did not protect against fake brands. Luo and Chen found that Gemini 3.1 Pro was fooled roughly three times as often as the smaller Gemini 3 Flash. When they switched step-by-step reasoning on and off in two models, reasoning made each more vulnerable, by up to 18 points.

Fooled answers also embellished. They added social proof that was not in the fake pages, such as claims of community popularity or drop tests. Fooled outputs used such phrases 1.5 to 11 times more often than answers that resisted.

Telling the AI to be careful backfired. A prompt warning it to be wary of unfamiliar brands raised the pooled fooled rate by 10.5 points. On Gemini 3.1 Pro it rose by 44 points.

## Can AI platforms filter out the fakes?

Not yet without heavy side effects. Luo and Chen tested four defenses:

- **A caution prompt:** made things worse on average, as above.
- **Two agreement filters:** caught most fakes, but threw away 52% to 79% of legitimate recommendations.
- **Re-ordering sources by credibility,** putting editorial sites first and user-generated pages last: lowered the fooled rate from 50.4% to 42.1% on six open-weights models, removing 17% of fake recommendations.

The authors call none of these adequate. The defenses were not tuned for this attack, so better ones may come. Fake reviews are one of several tactics in our guide on [using GEO to spread false claims](https://underneath.agency/resources/can-geo-push-false-information-into-ai-answers).

## What should you do about it?

Make the genuine evidence about your brand plentiful, consistent and easy to find. Practical steps:

1. Keep real reviews flowing on the platforms AI answers cite, such as Google, Trustpilot and the BBB, and respond to them.
2. Search your category in several AI assistants regularly, and note any unfamiliar brand that suddenly appears near the top.
3. When one appears, check the cited pages for invented brands, copied review text or new accounts posting in bulk, and report them to the platform hosting them.
4. Build independent coverage, such as press, trade lists and expert reviews, so the AI has trustworthy pages to weigh against forum posts.
5. Never buy or plant reviews yourself. It is the same tactic, and it carries legal and reputational risk.

If you want help strengthening the genuine evidence AI assistants find about your brand, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The research shows AI models can be fooled, not how often it happens to real brands. Specifically:

- The main test fed frozen search results to AI models; it did not test live ChatGPT or Gemini products searching the web themselves.
- Most data is Chinese-language, with local services set in Shenzhen; the English rerun covered only three categories.
- The test assumes the fake page already ranks near the top of search, which may be harder in practice.
- The evidence snapshot dates from April 2026, and results may shift as search results change.
- Real-world cases come from media and incident reports, not from measured studies.

## Frequently asked questions

### Can fake reviews really change what ChatGPT recommends?

In tests, yes. Across 12 AI models given search results with planted pages, fake brands were recommended in up to 73.8% of products, though the test did not use the live ChatGPT product.

### How many fake pages does it take to fool an AI assistant?

Sometimes one. A single fake page in the first search slot fooled the most vulnerable models in 27% of products. In one test, the most vulnerable crossed half with as few as three fake pages.

### Which businesses are most at risk from fake AI recommendations?

Businesses in everyday categories such as dining, personal services and supplements were most exposed. Technical products such as phones and appliances, where AI models know the real brands, were least exposed.

### How do I protect my brand from fake competitors in AI answers?

Keep genuine reviews and independent coverage strong, and monitor AI answers for unfamiliar brands. In our research, 88.0% of AI answers about brand legitimacy cited review or complaint platforms, so those platforms matter.

## Sources

- Luo, M. and Chen, L. (2026), [One Polluted Page Is Enough: Evaluating Web Content Pollution in LLM Recommenders](https://arxiv.org/abs/2606.13610), arXiv:2606.13610.
- Wen, Y., Zhang, N., Yuan, H., Chen, X., Zhang, H. and Guo, H. (2026), [Position: Generative Engine Optimization Creates Underexamined Risks, Governance Must Target Concentration, Disclosure, and Academic Blind Spots](https://arxiv.org/abs/2606.12439), arXiv:2606.12439.
- Baig, M. S. A., Gillani, S. A. and colleagues (2026), [Whose hotel does the AI recommend? An algorithm audit of reputation signals in LLM-assisted hotel selection](https://arxiv.org/abs/2606.16344), arXiv:2606.16344.
- Underneath (2026), [“Is this brand legit?” How AI assistants build a reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study)
- Underneath (2026), [Which Google Maps businesses does ChatGPT recommend?](https://underneath.agency/research/chatgpt-local-picks-google-profile-study)

---

This is the Markdown twin of https://underneath.agency/resources/fake-reviews-fake-brands-ai-recommendations. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How do fashion brands get recommended when shoppers ask AI?"
description: "By describing products in shoppers’ style and occasion language, with clear images and consistent details everywhere AI shopping tools and visual search look."
canonical: "https://underneath.agency/resources/fashion-ecommerce-ai-recommendations"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How does a fashion brand get recommended when shoppers ask AI what to wear?

By being easy to match to a look, an occasion and a style: product data written in the words shoppers use, images that AI try-on and visual search can work with, and the same details on every retailer that carries the brand. Google, ChatGPT, Pinterest and retailers such as Zalando now offer AI styling and try-on tools. Shopper use is real but uneven, so fashion brands should build for it now without expecting it to replace social and stores yet.

## The short version

1. Fashion is one of the biggest online categories. US shoppers spent [$49.0 billion on apparel](https://news.adobe.com/news/2026/01/adobe-holiday-shopping-season) online over the 2025 holidays, up 7.4%, Adobe found.
2. The AI tools are built for style. [Google says](https://blog.google/products-and-platforms/products/shopping/google-shopping-ai-mode-virtual-try-on-update/) shoppers can virtually try “billions of apparel listings” on themselves, and [OpenAI documents](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search) a “Try on” button for clothing in ChatGPT.
3. Retailers’ own assistants are scaling. [Zalando](https://corporate.zalando.com/en/node/10985) says close to 10 million customers asked its Assistant for advice in early 2026, up from 6 million in all of 2025.
4. US shoppers are cautious. In a [YouGov poll](https://yougov.com/en-us/articles/54891-are-clothes-shoppers-ready-for-ai-in-apparel-retail), just 6% of clothes shoppers said they would use ChatGPT or Gemini to discover clothing, against 60% who prefer browsing in stores.
5. Inconsistent product details are common. In a [UK survey](https://athoscommerce.com/news/ai-fragmented-discovery-and-rising-consumer-expectations-are-reshaping-fashion-ecommerce/) of 2,000 consumers, only 14% said fashion product details always match across social platforms, marketplaces and retailer sites.

## Who shops for fashion online, and how do they find what to buy?

Fashion shoppers start with inspiration, not a product name, and decide on look, occasion and price.

A fashion purchase rarely begins with a spec. It begins with a wedding invitation, a trip, a trend seen on TikTok or a feeling about a style. That is why fashion discovery has lived on social platforms, Pinterest boards, editorial edits and store windows, and why search for clothing is so often visual. In the [Connected Consumer 2026 report](https://athoscommerce.com/news/ai-fragmented-discovery-and-rising-consumer-expectations-are-reshaping-fashion-ecommerce/) by Athos Commerce and Drapers, 62% of UK consumers browse for fashion online at least weekly but only 38% buy that often. Most of the journey is looking.

The business stakes are high and the climate is hard. Over the 2025 holidays, apparel discounts peaked at 25.1% off list price in [Adobe’s data](https://news.adobe.com/news/2026/01/adobe-holiday-shopping-season). The [State of Fashion 2026 report](https://cfda.com/resources/sustainability-resource-hub/library-lexicon-annex/the-state-of-fashion-report-2026-mckinsey-company-and-bof-insights/) by McKinsey and The Business of Fashion, as summarized by the CFDA, found nearly half of executives expect industry conditions to worsen in 2026, a share up 8 percentage points on the year before, and executives named artificial intelligence as the industry’s biggest opportunity.

Fashion brands also sell through many doors: their own site, department stores, multi-brand platforms such as Zalando, and marketplaces. Each has its own product pages, and each can be the page an AI tool reads.

## Where do AI shopping tools fit into fashion discovery today?

They sit at the inspiration and narrowing stage, increasingly with visual and try-on features.

Four kinds of tools now help shoppers find fashion:

- **Google AI Mode.** Google says its AI Mode shopping experience draws on a Shopping Graph of more than 50 billion product listings, more than 2 billion of them refreshed every hour, and shows “a beautiful, browsable panel of images and product listings” for style requests. Its virtual try-on works with the shopper’s own photo for shirts, pants, skirts and dresses.
- **ChatGPT.** OpenAI documents a “Try on” button on clothing and accessory listings and says shoppers can upload a photo of clothing and ask how it would look on them. It warns that try-on images “do not guarantee fit or size.”
- **Pinterest.** [Pinterest](https://s204.q4cdn.com/369458543/files/doc_earnings/2026/q2/earnings-result/Q2-2026-Press-Release.pdf) reported 640 million monthly users in the second quarter of 2026, and its CEO said its AI is “trained on our unique human curation of style and taste.” Its [Pinterest Assistant](https://www.digitalcommerce360.com/2025/11/04/pinterest-introduces-ai-powered-assistant-for-online-shopping-discovery/) takes voice, text and image input and returns shoppable results.
- **Retailer assistants.** Zalando’s Assistant now covers fashion and beauty in one conversation. Zalando had 62.3 million active customers in the first quarter of 2026.

How many shoppers use these tools depends on who you ask. YouGov’s May 2026 poll of US clothes shoppers found 6% would use general AI tools to discover new clothing or brands, and 16% were interested in styling or outfit suggestions. The UK Connected Consumer survey found 60% of consumers use tools such as ChatGPT, Claude or Gemini at least occasionally while shopping for fashion, and 38% trust fashion recommendations from AI. The two surveys asked different questions in different countries, so the honest reading is that AI fashion discovery is growing from a minority base.

## Which questions do fashion shoppers ask AI?

They ask about occasions, looks, trends, brands like the ones they love, and whether a brand is worth it.

The prompts below are illustrative, written to show the kinds of request fashion shoppers make; they are not observed data.

| Type | Example prompt (illustrative) |
|---|---|
| Occasion | “What should I wear to a fall wedding as a guest?” |
| Look or aesthetic | “Quiet luxury outfit ideas under $300” |
| Styling | “What tops go with wide-leg jeans?” |
| Trend | “Is the burgundy trend still in this winter?” |
| Similar brands | “Brands like [label] but more affordable” |
| Quality check | “Is [brand] good quality for the price?” |
| Visual | A photo of a coat seen on the street: “Where can I buy this?” |

Google’s own example is the second type: a shopper asks AI Mode for “a cute travel bag,” and AI Mode runs several searches at once to work out what makes a bag good for a rainy trip before suggesting options. Visual queries are already common. In the UK survey, 58% of consumers had used visual search to look for fashion items, and 52% went on to buy something afterward; among Gen Z, 73% had used visual search to discover fashion.

## How does an AI recommendation turn into a fashion sale?

The tool turns a style request into a shortlist of looks, the shopper tries or compares, then buys.

The path runs from an occasion or style prompt to a visual set of options, through a try-on or comparison, to a product page and a basket. Two features make it different from other categories.

First, the answer is a set of images. Google describes AI Mode showing an image panel that updates as the shopper refines the request; ChatGPT shows product carousels with imagery. We infer that a garment shown on a clear, well-lit image, with the color and silhouette named in the data, has a better chance of being matched to a visual request than one with a vague title. Home decor brands face the same look-first requests, as our guide to [decor discovery in AI tools](https://underneath.agency/resources/home-decor-brands-ai-product-discovery) explains.

Second, the sale often lands with a retailer, not the brand. When the shopper asks Zalando’s Assistant, it recommends from Zalando’s catalog. When they ask ChatGPT, OpenAI says the merchant list is ranked partly on “whether they are the maker or primary seller of that item,” alongside price, availability and quality. A fashion brand’s own site can win that click, but only if its stock and price are competitive.

Shoppers still want control of the final step. A [Global Payments survey](https://fashionunited.com/news/retail/us-consumers-would-let-ai-agents-buy-clothes-but-not-without-a-say/2026092974919) reported by FashionUnited found 69% of respondents would let an AI agent spend up to $100 on clothing and footwear, but 42% worried an agent could buy the wrong item, a risk FashionUnited notes is high in clothing because sizing varies by brand. Our guide for [apparel brands on fit and size facts](https://underneath.agency/resources/apparel-brands-customers-ai-search) covers the details AI tools need to match clothes to a shopper.

## What decides which fashion brands and products AI recommends?

Platforms document their inputs only in part; studies observe a few fashion patterns; the rest is inference.

**Documented by the platform.** OpenAI says ChatGPT considers structured product data such as price and description from first-party and third-party providers, “other third-party content,” and public reviews, and that product results are not ads. Google says the Shopping Graph holds details like reviews, prices, color options and availability.

**Observed in a study.** A 2026 measurement of 55,393 trending Google searches found [AI Overviews appeared on only 3.5%](https://arxiv.org/abs/2605.14021) of Beauty and Fashion queries, the lowest of 19 categories, against 46.1% for Hobbies and Leisure. When they did appear, user-generated platforms, led overall by YouTube, Facebook and Instagram, supplied 28.9% of Beauty and Fashion citations, the highest share of any category. Those were trending queries, not shopping searches, so treat this as a signal that fashion answers lean on social content rather than proof. Our own [study of when Google shows an AI Overview](https://underneath.agency/research/ai-overviews-frequency-study) looks at the same question for commercial searches.

**Our inference.** Fashion shoppers describe what they want in style language: “flowy,” “oversized,” “old money,” “wedding guest.” A reasonable expectation is that products whose titles, descriptions and attributes use that vocabulary, alongside color, silhouette, fabric and occasion, are easier for AI tools to match. Creator content and fashion editorial likely carry weight because they supply the style language and the social proof that AI answers draw on. Our guide to [what product content AI shopping assistants prefer](https://underneath.agency/resources/product-content-ai-shopping-assistants-prefer) covers the research on product copy, and [does social media help AI search visibility](https://underneath.agency/resources/does-social-media-help-ai-search-visibility) covers which engines cite social platforms.

## What does a fashion brand lose if AI tools cannot match it to a look?

It loses the inspiration moment, which in fashion is where most brand choices are made.

No study yet measures lost fashion sales from AI invisibility, so we describe the exposure rather than size it. A shopper who asks for “brands like” a label they already know, or for a look for an event, gets a short set of options. A brand that is not in that set is not considered, and the alternatives often sit on the same retailer pages.

Inconsistent details make the problem worse. When only 14% of UK shoppers say product details always match across channels, AI tools pulling from several retailers may describe the same dress with different colors, prices or fabrics. Seasonal drops add timing risk: AI assistants often answer from older information, which [why ChatGPT misses new products](https://underneath.agency/resources/why-chatgpt-misses-new-products) explains. A collection that sells for eight weeks can be over before an assistant learns it exists.

There is a brand cost too. A [Shopify report](https://www.shopify.com/news/agentic-holiday-2026) quotes Aviator Nation’s ecommerce director on selling $200 hoodies: “There needs to be some emotional connection as to why you will convert and buy this hoodie.” If the story behind a label is not written where AI tools can read it, the answer may reduce the brand to price.

## How does GEO work for a fashion brand?

GEO makes your products and brand story easy for AI tools to match, show and trust, without promising placement.

Generative engine optimization (GEO) for fashion covers six areas:

1. **Style and occasion language in product data.** Name the silhouette, aesthetic, color, fabric, fit type and occasions in titles, descriptions and feed attributes, using the words shoppers type, not internal style codes.
2. **Images built for visual search and try-on.** Clear, consistent on-model and flat images, with descriptive alt text, so visual search and try-on tools can work with the garment.
3. **Consistent details across every door.** Make sure wholesale partners, Zalando, department stores and marketplaces carry the same names, colors, prices and fabric details as your own site.
4. **Fashion coverage and creators.** Earn editorial edits, stylist features and creator content that describe your pieces in style language. Fashion answers lean on social and editorial sources.
5. **Brand story where AI can read it.** Publish plain-text pages on what the brand stands for, who it dresses and how pieces are made, so an answer can explain why a shopper might choose you.
6. **Seasonal timing and measurement.** Get new collections into feeds and onto readable pages before launch, then test occasion and style prompts in ChatGPT, Google AI Mode, Gemini, Perplexity and Pinterest, repeating each several times.

## What can’t fashion brands learn from today’s AI shopping data?

How many fashion sales AI tools drive, and how they weigh style against price and reviews.

Surveys disagree on how widely fashion shoppers use AI, partly because they ask different questions in different countries. The UK survey comes from a commerce search vendor; YouGov’s poll is independent but asked about interest, not behavior. Zalando’s and Pinterest’s figures are their own. No platform has said how it ranks fashion items for a style request, and the trending-query study did not test shopping searches. Buying inside the chat is also unsettled: [FashionUnited reports](https://fashionunited.com/news/retail/us-consumers-would-let-ai-agents-buy-clothes-but-not-without-a-say/2026092974919) OpenAI scaled back Instant Checkout in March 2026 after users researched products in ChatGPT but rarely completed purchases there, though OpenAI’s help page still describes it for some eligible merchants. Whether virtual try-on reduces fashion returns has not been shown in published data.

## Where should a fashion brand start?

Start with the occasion and style prompts that should lead to your hero pieces and your next drop.

List the occasions, looks and “brands like” comparisons that bring your customers to you, and write them the way a shopper would ask. Run them in ChatGPT, Google AI Mode, Gemini, Perplexity and Pinterest, and on the retailer assistants that carry you. Note whether your pieces appear, which image is shown, which retailer gets the click, and how your brand is described. Then compare your product details across your own site and your three biggest retail partners.

If you would rather have a second pair of eyes on your collections, [talk to us about a fashion review](https://underneath.agency/contact). We will show how AI shopping and visual search tools describe your pieces today and which changes are most likely to get them chosen when shoppers ask what to wear. Our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page walks through the ongoing work, such as writing style and occasion language into feeds, aligning retailer details and timing each seasonal drop.

## Frequently asked questions

### Do people really use ChatGPT to find clothes?

Some do, but fewer than social media or stores. In YouGov’s US poll, 6% of clothes shoppers would use general AI tools to discover clothing, while a UK survey found 60% use AI tools at least occasionally when shopping for fashion.

### Does virtual try-on in Google or ChatGPT help my products get recommended?

Try-on is a shopping feature, not a ranking factor that either company has documented. Clear product images and complete listings make your items usable by it.

### Should a fashion brand optimize for Pinterest as well as ChatGPT?

Yes, if your customers use it. Pinterest reports 640 million monthly users and offers an AI assistant for shoppable results, and visual discovery is central to fashion.

### How fast do AI assistants pick up a new collection?

It varies. Feeds and readable product pages can reach shopping tools quickly, but assistants that answer from older information may miss new pieces, so launch data early.

## Sources

- Adobe (2026-01-07), [Holiday Shopping Season Drove a Record $257.8 Billion Online with Consumers Embracing Generative AI Tools](https://news.adobe.com/news/2026/01/adobe-holiday-shopping-season)
- Google (2025-05-20), [Shop with AI Mode, use AI to buy and try clothes on yourself virtually](https://blog.google/products-and-platforms/products/shopping/google-shopping-ai-mode-virtual-try-on-update/)
- OpenAI (2026), [Shopping with ChatGPT Search](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search)
- Zalando (2026-05-06), [Zalando delivers strong Q1 as artificial intelligence and ABOUT YOU integration drive growth and efficiencies](https://corporate.zalando.com/en/node/10985)
- Pinterest (2026-08-04), [Pinterest Announces Second Quarter 2026 Results](https://s204.q4cdn.com/369458543/files/doc_earnings/2026/q2/earnings-result/Q2-2026-Press-Release.pdf)
- Digital Commerce 360 (2025-11-04), [Pinterest introduces AI-powered assistant for online shopping discovery](https://www.digitalcommerce360.com/2025/11/04/pinterest-introduces-ai-powered-assistant-for-online-shopping-discovery/)
- YouGov (2026), [Are clothes shoppers ready for AI in apparel retail?](https://yougov.com/en-us/articles/54891-are-clothes-shoppers-ready-for-ai-in-apparel-retail)
- Athos Commerce and Drapers (2026-06-04), [Connected Consumer 2026 report announcement](https://athoscommerce.com/news/ai-fragmented-discovery-and-rising-consumer-expectations-are-reshaping-fashion-ecommerce/)
- FashionUnited (2026-09-29), [US consumers would let AI agents buy clothes, but not without a say](https://fashionunited.com/news/retail/us-consumers-would-let-ai-agents-buy-clothes-but-not-without-a-say/2026092974919)
- CFDA (2025), [The State of Fashion Report 2026: McKinsey & Company and BoF insights](https://cfda.com/resources/sustainability-resource-hub/library-lexicon-annex/the-state-of-fashion-report-2026-mckinsey-company-and-bof-insights/)
- Shopify (2026-10-06), [Welcome to the first holiday season of the agentic era](https://www.shopify.com/news/agentic-holiday-2026)
- Haofei Xu and colleagues (2026), [Measuring Google AI Overviews: Activation, Source Quality, Claim Fidelity, and Publisher Impact](https://arxiv.org/abs/2605.14021)

---

This is the Markdown twin of https://underneath.agency/resources/fashion-ecommerce-ai-recommendations. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How financial consulting firms win business through AI search"
description: "By being the firm AI names for a specific finance problem, company size and deal stage, with a reputation, credentials and current commentary that hold up."
canonical: "https://underneath.agency/resources/financial-consulting-firms-clients-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How can a financial consulting firm win CFO, deal and restructuring work through AI search?

By being the firm an AI assistant names when an owner, CFO or investor describes a specific finance problem, company size and stage, and by making sure the reputation evidence those answers check is strong. Fractional CFO work, quality of earnings reviews and turnaround mandates are each bought differently, so AI search matters at different moments for each. No study has yet measured how many engagements start with an AI answer.

This guide covers financial consulting for businesses: CFO advisory and fractional CFOs, financial planning and analysis (FP&A), transaction advisory such as due diligence and quality of earnings, and turnaround and restructuring advice. It is not about personal financial advice, and nothing here is financial advice to any business. Tax and audit firms are covered in how CPA firms win clients through AI search.

## The short version

1. Distress work is rising: business bankruptcy filings rose 16.9% to 26,941 in the year ending June 30, 2026, according to the [US courts](https://www.uscourts.gov/data-news/judiciary-news/2026/07/28/bankruptcies-rise-122-percent).
2. Deal work is steady: [GF Data, an ACG company](https://www.acg.org/news-trends/news/gf-data-reports-show-steady-middle-market-deal-flow-amid-more-selective), counted 170 private equity-sponsored middle-market transactions in the first half of 2026, tracking toward 340 for the year, up 10%.
3. Finance teams are short of people: in the [Controllers Council’s 2026 study](https://controllerscouncil.org/finance-hiring-is-back-but-the-recruiting-challenge-has-changed/), 61 percent of respondents reported shortages of accounting, finance or CPA talent.
4. Part-time CFOs lead the fractional market: finance made up 46% of US fractional job postings, with CFOs the largest single occupation, in [Lightcast’s analysis](https://lightcast.io/blog/rise-of-fractional-leadership).
5. AI answers probe reputation: in [our “is it legit?” study](https://underneath.agency/research/is-it-legit-ai-reputation-study), 99.7% of complete answers about a brand raised at least one negative point, and 88.0% cited a review or complaint platform.

## Who hires a financial consulting firm, and what is a client worth?

Owners, CEOs, CFOs, boards, private equity sponsors and lenders, depending on which of three problems they face.

**CFO advisory and FP&A.** The buyer is usually the CEO or owner of a growing company without a senior finance leader, or a CFO short of capacity. Lightcast found 97% of fractional job postings in 2026 came from medium and small companies. The value is a monthly retainer that can run for years, plus projects such as a budget, a forecast model or preparation for a raise.

**Transaction advisory.** The buyer is a private equity sponsor, a corporate acquirer or an owner preparing to sell. Work is priced per deal: diligence, quality of earnings, working capital analysis. Deal pricing makes the stakes clear. GF Data reports an average purchase multiple of 7.0x trailing 12-month adjusted EBITDA in the second quarter of 2026, so every dollar of earnings the diligence confirms or removes moves the price several times over.

**Turnaround and restructuring.** The buyer is a CEO or board under pressure, often pushed by a lender, with counsel involved. Fees come as retainers and, in some engagements, success fees.

A public example shows the market’s shape. [FTI Consulting](https://www.fticonsulting.com/about/newsroom/press-releases/fti-consulting-reports-second-quarter-2026-financial-results) reported Corporate Finance segment revenues of $411.4 million in the second quarter of 2026, up 8.5%. It credited higher bill rates, demand for transformation services and “higher success fees,” partly offset by “lower demand for turnaround & restructuring services.” Demand moves between these lines as conditions change, which is why firms that offer several of them must be clear about each.

## What triggers a company to hire financial consultants now?

A capacity gap, a deal, or a cash problem, and each sends the buyer looking at a different speed.

**Capacity gaps.** The Controllers Council found 38 percent of respondents expect to increase finance and accounting headcount in the next twelve months, while talent is short. When a controller leaves or a lender asks for better reporting, outside help fills the gap.

**Deals.** A letter of intent sets off diligence on a tight timeline. Owners who plan to sell research the process months earlier, including what a quality of earnings report is and who prepares one. Our guide on [how advisory firms reach business owners](https://underneath.agency/resources/business-advisory-firms-clients-ai-search) covers the exit and cash questions owners bring to AI first.

**Cash pressure.** The [CFO Survey](https://www.richmondfed.org/research/national_economy/cfo_survey/research_and_commentary/2026/20260923_research_commentary) run by the Richmond Fed and Duke University found 19 percent of firms say access to financing, or its cost, has constrained their investment or spending. Around double the proportion of smaller firms, under 500 employees, report being financially constrained compared to larger ones. Overall, CFOs rated their optimism about the US economy at 60.3 out of 100 in the third quarter.

## Where does AI search sit in how businesses choose a financial adviser?

Early, when the buyer is learning the vocabulary, and again when they check a firm’s reputation before calling.

The path differs by service:

| Service | How the buyer finds a firm | Where AI search fits |
|---|---|---|
| Fractional CFO, FP&A | Accountant, investors, peers, search | Explaining the role and finding candidates by stage and industry |
| Transaction advisory | Banker or sponsor panels, prior deals | Owners learning the sale process; sponsors checking a new firm |
| Restructuring | Lender and counsel referrals, fast | Quiet early research by a CEO; checking a recommended firm’s record |

Across business purchases, assistants are already part of research. In [a Gartner survey](https://www.gartner.com/en/newsroom/press-releases/2026-05-20-gartner-survey-finds-sixty-nine-percent-of-b-two-b-buyers-turn-to-sales-reps-to-validate-ai-generated-insights) of 645 B2B buyers in all sectors, 45% said they used generative AI during a recent purchase, mainly to gather information on vendors. No survey isolates buyers of financial consulting.

Large decisions also involve many people. [Forrester](https://www.forrester.com/press-newsroom/forrester-2026-the-state-of-business-buying/) counts 13 internal stakeholders and nine external influencers in a typical business purchase. In a restructuring or a sale, those influencers include lenders, lawyers, bankers and board members, each of whom may check a firm independently, we infer.

## What do business owners and CFOs ask AI about financial advisers?

Questions about a situation, a size and a stage, often before they know the right term. Treat the examples below as illustrations we drafted, not logged buyer queries.

| Situation | Example question |
|---|---|
| Growth | “Fractional CFO for a $20 million software company preparing to raise a Series B?” |
| Reporting | “Who can build a 13-week cash flow forecast for a distributor our bank is worried about?” |
| Sale | “Do I need a quality of earnings report before selling my manufacturing company?” |
| Buyer side | “Which firms do quality of earnings for lower middle market deals in Texas?” |
| Distress | “What does a chief restructuring officer do, and how are they paid?” |
| Fees | “How much does a fractional CFO cost per month for a 100-person company?” |

Many of these name a place. Local questions split the assistants more: [our four-assistant study](https://underneath.agency/research/ai-assistants-brand-agreement-study) measured an overlap of 0.160 between them when a city or state was named, compared with 0.390 when the question was national. A regional firm may be named by one assistant and missing from another.

## How does an AI answer become an engagement?

Through a first call for advisory and deal work; through a reputation check for restructuring, where referrals still lead.

**Fractional CFO and FP&A:** AI answer → website and profile check → call → monthly retainer or project. The shortest path from answer to revenue.

**Transaction advisory:** owner or sponsor research → shortlist → proposal with fixed fee → diligence → often post-deal work such as integration or reporting.

**Restructuring:** a lender or lawyer names two or three firms → the CEO or board asks an assistant about them → engagement. Here the AI answer rarely creates the lead; it can confirm it or undermine it.

That last point is where reputation matters most. In our “is it legit?” study, every complete answer affirmed the brand was legitimate, but 99.7% raised a problem, and review and complaint platforms carried most of the negative claims. A firm with a disputed engagement in the press or a public fee fight should expect an assistant to mention it, and should make sure its own account of its record is easy to find. When an answer gets a firm’s record wrong, our guide on [correcting brand errors in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) walks through what an advisory firm can fix and where.

## What decides whether an assistant names a financial consulting firm?

No platform publishes how it picks financial advisers; studies point to named sources, fresh pages and reputation evidence.

**Documented by the platforms.** Assistants with web search cite the pages they draw on; none documents how it chooses among consulting firms.

**Observed in our studies, in other industries.** None asked about financial consultants directly:

- **Named sources get cited.** In [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study), when an assistant’s own search named a source, the answer cited that source 44.0% of the time, against 8.1% when the search did not name it. For this industry the obvious named sources are league tables, award lists, professional bodies and trade publications.
- **Newer pages are favored.** In [our freshness study](https://underneath.agency/research/ai-source-freshness-study), pages published in the last 90 days made up 17.4% to 22.6% of each assistant’s dated citations, against 6.9% of Google’s top 10. Market commentary on rates, multiples and filings dates quickly.
- **Reputation answers lean on reviews.** In the “is it legit?” study, 88.0% of answers cited a review or complaint platform.

**Our inference, specific to financial consulting.** Credentials and memberships that buyers already trust (CPA, CFA, turnaround and insolvency certifications, membership of ACG or the Turnaround Management Association) are likely to be read as evidence when they appear consistently across profiles. Client confidentiality limits case studies, so anonymized deal summaries with size, sector and outcome do much of the work. Firms that also offer tax or audit work can read [how CPA firms win clients through AI search](https://underneath.agency/resources/accounting-firms-clients-ai-search).

## What does it cost a financial consulting firm to be missing?

The calls that go to whichever firm the assistant named, and doubt planted at the reputation check. No study puts a figure on it.

- **Owners who plan early find someone else.** A seller who learned about quality of earnings from an assistant may call the firm it named months before a banker gets involved.
- **Referrals can be undercut.** A lender’s recommendation loses force if an assistant describes the firm inaccurately or leads with an old dispute.
- **Thin teams buy faster.** With 61 percent of finance leaders reporting talent shortages, many will choose the first credible fractional or project help they find, we infer.

The broader cost of fewer clicks is set out in [what lost clicks to AI answers mean for pipeline](https://underneath.agency/resources/ai-answers-pipeline-revenue).

## How does GEO work for a financial consulting firm?

It makes your services, deal sizes, sectors and record clear and checkable. It cannot promise a recommendation.

1. **Explain each service in plain words.** What a fractional CFO does in the first 90 days, what a quality of earnings report covers, when a 13-week cash flow is needed, what a chief restructuring officer is paid for. Buyers ask these before they ask for names.
2. **State who you serve.** Revenue ranges, deal sizes, industries, regions and stages, the same way on your site, LinkedIn and directory profiles.
3. **Publish dated market commentary.** Quarterly notes on multiples, lending conditions and filings, signed by a named partner, with sources. Our freshness study suggests assistants lean toward newer pages.
4. **Make credentials consistent.** Certifications, memberships and regulatory status, where relevant, matching across every profile.
5. **Show the record within confidentiality limits.** Anonymized deal summaries, published league table positions and, where clients agree, named testimonials.
6. **Tend reputation sources.** Reviews where they exist, responses to public criticism, and an accurate account of any notable dispute.
7. **Get quoted.** Business journals, deal publications and association events give assistants independent sources to cite.
8. **Test the questions.** Ask ChatGPT, Gemini, Perplexity, Claude, Copilot and Google’s AI Overviews and AI Mode your buyers’ questions, with size and place, several times each.

For a firm that also manages private wealth, see [how wealth managers win clients through AI search](https://underneath.agency/resources/wealth-management-firms-clients-ai-search); that is a separate audience with separate rules. Firms that also advise on finance systems such as ERP can compare [how IT consulting firms win advisory work](https://underneath.agency/resources/it-consulting-firms-clients-ai-search).

## Which questions about AI and financial consulting remain open?

How often AI answers start or stop an engagement. No study follows a CFO, deal or restructuring mandate back to an AI answer.

- **No survey of financial consulting buyers.** Gartner and Forrester cover business buyers in general.
- **Market data covers part of the field.** GF Data tracks private equity-sponsored deals of $1 million to $500 million; FTI is one large firm.
- **Our studies covered other categories.** Reputation and freshness patterns may differ for advisers bound by confidentiality.
- **Referral-led work is hard to observe.** A restructuring mandate rarely records whether a board member checked an assistant first.

## Where should a financial consulting firm start?

With the questions owners and CFOs ask before they know whom to call, sorted by service line.

Write 20: five each for fractional CFO work, FP&A, transactions and restructuring, with company size and place. Ask each major assistant several times. Record which firms are named, which sources are cited, and how your firm’s services, sectors and record are described, including anything negative.

If you would like an independent view of those answers, [see how we review advisory firms’ AI visibility](https://underneath.agency/contact). We will show which CFO, deal and restructuring questions name your firm, where others are named instead, and which gaps in explanations, credentials and reputation sources are most likely costing you first calls and engagement letters. Our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page details how the follow-on work runs for an advisory firm, from plain service explanations and dated market commentary to tending reputation sources.

## Frequently asked questions

### Do business owners use AI to find a fractional CFO?

No survey isolates them. Across business purchases, 45% of buyers in Gartner’s survey used generative AI, and finance leads the fractional market.

### Does AI matter for restructuring work, which comes through referrals?

Mostly as a check. A board or CEO given two names may ask an assistant about each, and reputation answers often raise problems.

### Should we publish fees?

Explain the structure: monthly retainers, fixed deal fees or success fees. Buyers ask early, and clear structures are easier to compare.

### Can a boutique firm appear alongside large firms?

On specific sizes, sectors and regions, plausibly. Questions naming a place produced very different lists across assistants in our study.

### Is this article financial advice?

No. It describes how businesses choose financial consultants and how AI search affects that choice.

## Sources

- Administrative Office of the U.S. Courts (2026-07-28), [Bankruptcies rise 12.2 percent](https://www.uscourts.gov/data-news/judiciary-news/2026/07/28/bankruptcies-rise-122-percent)
- Association for Corporate Growth (2026-08-27), [GF Data reports show steady middle-market deal flow amid more selective financing conditions in Q2](https://www.acg.org/news-trends/news/gf-data-reports-show-steady-middle-market-deal-flow-amid-more-selective)
- Controllers Council (2026), [Finance hiring is back, but the recruiting challenge has changed](https://controllerscouncil.org/finance-hiring-is-back-but-the-recruiting-challenge-has-changed/)
- Lightcast (2026-08-20), [The rise of fractional leadership](https://lightcast.io/blog/rise-of-fractional-leadership)
- FTI Consulting (2026-07-30), [FTI Consulting reports second quarter 2026 financial results](https://www.fticonsulting.com/about/newsroom/press-releases/fti-consulting-reports-second-quarter-2026-financial-results)
- Federal Reserve Bank of Richmond and Duke University (2026-09-23), [Are firms financially constrained?](https://www.richmondfed.org/research/national_economy/cfo_survey/research_and_commentary/2026/20260923_research_commentary)
- Federal Reserve Bank of Richmond and Duke University (2026-09-23), [CFO outlook: steady overall but weaker for small and financially constrained firms](https://www.richmondfed.org/research/national_economy/cfo_survey/data_and_results/2026/20260923_data_and_results)
- Gartner (2026-05-20), [Gartner survey finds 69% of B2B buyers turn to sales reps to validate AI-generated insights](https://www.gartner.com/en/newsroom/press-releases/2026-05-20-gartner-survey-finds-sixty-nine-percent-of-b-two-b-buyers-turn-to-sales-reps-to-validate-ai-generated-insights)
- Forrester (2026-01-21), [Forrester’s 2026 buyer insights: GenAI is upending B2B buying](https://www.forrester.com/press-newsroom/forrester-2026-the-state-of-business-buying/)
- Underneath (2026), [“Is this brand legit?” How AI assistants build a reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study)
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude agree on brands?](https://underneath.agency/research/ai-assistants-brand-agreement-study)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [How fresh are the pages AI engines cite?](https://underneath.agency/research/ai-source-freshness-study)

---

This is the Markdown twin of https://underneath.agency/resources/financial-consulting-firms-clients-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How AI assistants are changing fintech software buying"
description: "Yes. Finance teams and bank staff research vendors with AI assistants, and in finance those answers lean on trusted, regulated and well-ranked sources."
canonical: "https://underneath.agency/resources/fintech-software-customers-from-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Are AI assistants changing how banks and finance teams choose fintech software?

Yes: business buyers, including finance teams and bank staff, now use AI assistants to research vendors, often inside private tools their employer provides. In finance, those assistants lean heavily on sources that already rank and on trusted third parties, and every shortlisted vendor still faces regulated due diligence. For a fintech company, AI visibility is won with verifiable proof of compliance, security and results, not with louder claims.

## The short version

1. Google shows AI answers on most finance searches: in [our study](https://underneath.agency/research/ai-overviews-frequency-study) of 800 US keywords, financial services and insurance searches triggered an AI Overview 88.0% of the time, among the highest of eight industries.
2. Finance answers lean on established sources: in [our AI Overview citation study](https://underneath.agency/research/ai-overview-citations-study), 34.8% of finance and insurance citations were page-one Google results, against 23.9% for B2B software, and nerdwallet.com was cited in 10.0% of all AI Overviews in the sample.
3. Bank buyers must document every vendor choice: US banking regulators’ [interagency guidance](https://www.federalreserve.gov/newsevents/pressreleases/bcreg20230606a.htm) covers “due diligence and third-party selection,” including relationships with fintech companies, and the EU’s [DORA](https://www.eiopa.europa.eu/digital-operational-resilience-act-dora_en) has applied since 17 Jan 2025.
4. A fintech customer can be worth years of recurring and transaction revenue: [Q2](https://www.01net.it/q2-holdings-inc-announces-second-quarter-2026-financial-results/), which sells digital banking software to banks and credit unions, reported subscription annualized recurring revenue of $825.5 million, and [BILL](https://www.01net.it/bill-reports-fourth-quarter-and-fiscal-year-2026-financial-results/) served 479,300 businesses.
5. Fintech is moving inside the assistants: Intuit signed a [deal worth more than $100 million](https://techcrunch.com/2025/11/18/intuit-signs-100m-deal-with-openai-to-bring-its-apps-to-chatgpt) with OpenAI to bring TurboTax, Credit Karma and QuickBooks into ChatGPT.

## Which institutions and teams buy fintech software, and what is each worth?

Three kinds of buyer: financial institutions, business finance teams and consumers, each with different economics.

**Financial institutions.** Banks and credit unions buy digital banking, lending, payments, fraud and compliance software through formal evaluations. Q2’s second-quarter 2026 results show the stakes: subscription annualized recurring revenue of $825.5 million, up 15%, and eight Enterprise and Tier 1 contracts signed in the quarter. One of those came after a large bank acquired an existing Q2 customer. Contracts like these run for years. Our guide on [getting onto a bank’s vendor shortlist](https://underneath.agency/resources/banking-technology-enterprise-deals-ai-search) looks at core, payments and fraud deals in more detail.

**Business finance teams.** CFOs, controllers and accounts payable teams buy spend management, payables, receivables and treasury tools, often faster and with less ceremony. BILL, which sells this kind of software, reported fiscal 2026 core revenue of $1,504.7 million across 479,300 businesses. Most of it, $1,211.2 million, came from transaction fees, so a customer’s value grows with every payment it makes through the platform. Processors that sell to merchants follow a related path, set out in [how payment processors reach merchant shortlists](https://underneath.agency/resources/payment-processors-merchant-demand-ai-search).

**Consumers.** Neobanks, lenders, “buy now, pay later” providers and investing apps sell directly to people. The [Federal Reserve’s 2024 household survey](https://www.federalreserve.gov/publications/files/2024-report-economic-well-being-us-households-202505.pdf) found use of buy now, pay later edged up to 15 percent of adults. [Sensor Tower](https://sensortower.com/press/press-release-boosted-by-gen-ai-services-consumers-spent-more-money-in-apps-than-games-for-first-time) reports credit and lending app downloads rose 18% in 2025.

This article focuses on the first two, where deals are larger and buying is more deliberate. Subscription economics outside finance are covered in [how B2B SaaS companies generate revenue from AI search](https://underneath.agency/resources/b2b-saas-revenue-from-ai-search).

## When do bank and finance-team buyers consult AI?

At the research stage, inside work tools, and increasingly inside the assistants themselves.

No public survey isolates fintech buyers, so the evidence comes from wider business buying:

- In [Forrester’s](https://forrester.com/blogs/state-of-business-buying-2026) survey of nearly 18,000 business buyers, 94% reported using AI during their buying process, and procurement staff were decision-makers in 53% of buying cycles.
- Employer-provided assistants are widespread: in Forrester’s research, [61% of business buyers](https://www.forrester.com/blogs/b2b_buyers_make_zero_click_buying_number_one/) use private AI tools provided by their organization. In regulated firms such as banks, we infer, that is where much vendor research now starts.
- [Pew Research Center](https://www.pewresearch.org/short-reads/2025/06/25/34-of-us-adults-have-used-chatgpt-about-double-the-share-in-2023/) found 28% of employed US adults have used ChatGPT for work, up 20 points in two years.

Banks themselves are heavy AI adopters. The [Evident AI Index](https://evidentinsights.com/ai-index/), which tracks 50 of the world’s largest banks, reports that AI deployment across the sector moved nearly three times faster this year than in previous years. Teams that use AI daily in their own operations are, we infer, likely to use it when researching suppliers too.

On the consumer side, fintech now operates inside the assistant. OpenAI launched [Instant Checkout](https://openai.com/index/buy-it-in-chatgpt/) in ChatGPT, built with Stripe on an open Agentic Commerce Protocol, and says more than 700 million people use ChatGPT each week. Intuit’s apps, including Credit Karma, are coming to ChatGPT under its OpenAI deal, so users can review credit options without leaving the chat. And in [Menlo Ventures’ 2026 consumer survey](https://menlovc.com/wp-content/uploads/2026/09/menlo_ventures_consumer_ai_report-2026.pdf), 23% of Americans who pay bills already use AI to help.

## Which questions do fintech buyers ask AI assistants?

Questions about fit, integrations, compliance, fees, alternatives and whether a provider can be trusted. We wrote the examples below to illustrate bank, CFO and consumer questions.

| Buyer | Illustrative prompt |
|---|---|
| Community bank | “Which digital banking platforms do credit unions under $2 billion in assets use, and how long does conversion take?” |
| Bank risk team | “Which fraud detection vendors for real-time payments have bank customers in the US, and what do their SOC 2 reports cover?” |
| Mid-market CFO | “Best accounts payable automation for a 300-person company on NetSuite” |
| Startup finance lead | “Ramp vs Brex for a Series A company with international contractors” |
| Controller | “Alternatives to BILL for a company that pays many overseas suppliers” |
| Payments team | “Which payment processors support the Agentic Commerce Protocol?” |
| Consumer | “Is this savings app FDIC-insured, and what do customers say about it?” |

Many of these questions reach Google first, where AI answers are now the norm. In our frequency study, financial services stayed among the two highest industries even after adjusting for the mix of searches. Google documents that its AI features [may use “query fan-out”](https://developers.google.com/search/docs/appearance/ai-features), running several related searches across subtopics before answering.

## How does an AI answer lead to a bank contract or a finance-team trial?

By putting you on the first list for an evaluation, or in the answer that leads to a trial.

**The bank path.** A team member asks an assistant to map vendors, a shortlist forms, a request for proposal goes out, and the vendor then passes due diligence, contract negotiation and implementation. The interagency guidance describes this life cycle as “planning, due diligence and third-party selection, contract negotiation, ongoing monitoring, and termination.” The revenue is a multi-year subscription that, at Q2, also expands when existing customers add products.

**The finance-team path.** A controller asks for options, compares two or three tools, and starts a trial or demo. Forrester found more than 60% of business buyers use some form of trial. At a company like BILL, revenue then follows the customer’s payment volume.

**The consumer path.** A person asks whether an app is right and safe, checks reviews, and signs up. Here, the assistant’s summary of reputation can decide the outcome.

In all three, the AI answer comes first and the money comes much later. We suggest measuring AI visibility against the pipeline it feeds: request-for-proposal invitations, demo requests and sign-ups that name an AI assistant as a source.

## What makes an assistant trust a fintech vendor enough to name it?

Mostly trusted third-party evidence; in finance, assistants and Google’s AI lean on established, well-ranked sources.

Documented by platforms: Google says its AI features can run several related searches. OpenAI says ChatGPT’s shopping results are “organic and unsponsored, ranked purely on relevance to the user,” and that Instant Checkout items “are not preferred in product results.” Neither publishes how it chooses financial software vendors.

Observed in our studies:

- **Ranking still matters more in finance.** In our citation study, 34.8% of finance and insurance AI Overview citations were page-one results, against 23.9% for B2B software. Finance searches lean more on pages that already rank.
- **Named authorities get cited.** In [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study), when ChatGPT’s search named a source such as NerdWallet, the answer cited that source 44.0% of the time, against 8.1% when it did not.
- **Reputation answers come from review platforms.** In [our reputation study](https://underneath.agency/research/is-it-legit-ai-reputation-study), 88.0% of answers to “Is this brand legit?” cited a review or complaint platform, and Trustpilot and the BBB accounted for 61.7% of review-platform citations.
- **Country changes the answer.** In [our country study](https://underneath.agency/research/ai-recommendations-by-country-study), ChatGPT put QuickBooks first for small business accounting software in every US and Canadian run and Xero first in every UK and Australian run. The country mattered most for financial services and insurance.
- **Communities matter less than in software.** In [our Reddit study](https://underneath.agency/research/ai-reddit-citations-study), 5.7% of Google’s AI Overviews for financial services searches cited Reddit, against 35.4% for B2B software. In [our YouTube study](https://underneath.agency/research/ai-overview-youtube-videos-study), they cited a video on 49.0% of financial services searches, against 91.0% for B2B software.

Our inference for fintech: the trust factors are regulatory status (bank charter or partner bank, licenses, FDIC or similar protections stated accurately), security attestations such as SOC 2 and PCI DSS, named financial institution customers, coverage in finance publications and comparison sites, and clean review profiles. An assistant that cannot verify these is more likely to name a better-documented rival, we infer.

## What does GEO look like when your buyers are banks and CFOs?

Generative engine optimization (GEO) makes your regulatory status, security and results easy for assistants to find, verify and repeat.

In practice, for a vendor selling to banks and finance teams, the work usually covers:

1. **Accurate regulatory facts.** State on public pages exactly what you are and are not: licenses, partner banks, deposit insurance, data handling, and the markets you serve. Assistants repeat what they find, and errors here are costly. When an assistant misstates your charter or partner bank, follow [how to fix wrong information about your brand in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers).
2. **Public due diligence material.** A trust page with security attestations, uptime, incident history and policies, readable without a login, so a bank’s research and its assistant find the same proof.
3. **Search foundations.** Because finance answers lean on page-one results, keep the pages that answer buyers’ questions ranking well in Google and Bing.
4. **Trusted third parties.** Coverage in finance and banking publications, placement on comparison sites buyers name, analyst and peer-review profiles, and named customer stories with financial institutions.
5. **Reputation hygiene.** Monitor Trustpilot, the BBB and app store reviews, respond to complaints, and fix the issues behind them; assistants summarize what those platforms say.
6. **Market-by-market checks.** Test answers in each country you sell in, since assistants change their picks by market. See [whether an English-only AI visibility check is enough](https://underneath.agency/resources/english-only-ai-visibility-audits).

No one can guarantee an assistant will recommend a fintech vendor. GEO makes it more likely that the evidence it finds is accurate, verifiable and in your favor.

## What can’t the evidence tell fintech vendors yet?

It shows finance searches are full of AI answers; it does not show how banks’ private assistants pick vendors.

- **No fintech-specific buyer survey.** We found no public, independent survey of how bank or CFO buyers use AI assistants to choose fintech vendors. The buyer figures above cover business buying in general.
- **Private tools are invisible.** What a bank’s internal assistant answers, and which sources it may use, cannot be observed from outside.
- **Our studies are snapshots.** Our citation and frequency findings come from US searches in September 2026; answers change over time and between assistants.
- **No link to revenue yet.** No public study connects AI visibility to fintech contracts or transaction volume. What is known outside finance is summarized in [does AI visibility drive business results](https://underneath.agency/resources/does-ai-visibility-drive-business-results).

## What should a fintech vendor test before the next bank RFP?

Ask the questions your bank and finance-team buyers ask, in the assistants they use, in each market.

That first check shows whether you are named for your core use cases, which rivals and sources appear instead, and whether your regulatory, security and pricing facts are described correctly. In fintech, a wrong fact can cost a deal before a seller ever hears about it.

If your growth depends on request-for-proposal invitations from financial institutions or demos with finance teams, [ask us for a due-diligence-style visibility review](https://underneath.agency/contact). We will compare how assistants describe you and your competitors, list the evidence gaps a bank’s own due diligence would surface, and propose how to close them before the next evaluation. Our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page explains how that work is run for vendors selling to banks and CFOs, including public due-diligence pages, accurate regulatory facts and checks in each market.

## Frequently asked questions

### Are bank staff using ChatGPT to research fintech suppliers?

Many business buyers do, often through private AI tools at work. Forrester found 61% of business buyers use AI tools provided by their organization, though no public survey isolates bank buyers.

### Why do AI answers about finance favor sites like NerdWallet?

Finance answers lean on established, well-ranked sources. NerdWallet was cited in 10.0% of the AI Overviews in our sample, and ChatGPT cited named sources such as NerdWallet far more often when its own search mentioned them.

### Does SEO still matter for fintech AI visibility?

Yes, more than in many industries. In our study, 34.8% of finance and insurance AI Overview citations were page-one Google results, against 23.9% for B2B software.

### Can a fintech company pay to appear in ChatGPT’s recommendations?

Not according to OpenAI’s documentation, which says shopping results are “organic and unsponsored” and that Instant Checkout items are not preferred in results. Visibility has to be earned through evidence.

## Sources

- Board of Governors of the Federal Reserve System (2023-06-06), [Agencies issue final guidance on third-party risk management](https://www.federalreserve.gov/newsevents/pressreleases/bcreg20230606a.htm)
- EIOPA (2025), [Digital Operational Resilience Act (DORA)](https://www.eiopa.europa.eu/digital-operational-resilience-act-dora_en)
- Q2 Holdings, via 01net (2026), [Q2 Holdings, Inc. Announces Second Quarter 2026 Financial Results](https://www.01net.it/q2-holdings-inc-announces-second-quarter-2026-financial-results/)
- BILL, via 01net (2026), [BILL Reports Fourth Quarter and Fiscal Year 2026 Financial Results](https://www.01net.it/bill-reports-fourth-quarter-and-fiscal-year-2026-financial-results/)
- Board of Governors of the Federal Reserve System (2025-05), [Economic Well-Being of U.S. Households in 2024](https://www.federalreserve.gov/publications/files/2024-report-economic-well-being-us-households-202505.pdf)
- Sensor Tower (2026-01-21), [Boosted by Gen AI services, consumers spent more money in apps than games for first time](https://sensortower.com/press/press-release-boosted-by-gen-ai-services-consumers-spent-more-money-in-apps-than-games-for-first-time)
- Forrester (2026-01), [The State Of Business Buying, 2026](https://forrester.com/blogs/state-of-business-buying-2026)
- Forrester (2025), [B2B Buyers Make Zero-Click Buying Number One](https://www.forrester.com/blogs/b2b_buyers_make_zero_click_buying_number_one/)
- Pew Research Center (2025-06-25), [34% of U.S. adults have used ChatGPT, about double the share in 2023](https://www.pewresearch.org/short-reads/2025/06/25/34-of-us-adults-have-used-chatgpt-about-double-the-share-in-2023/)
- Evident (2026), [Evident AI Index for Banks](https://evidentinsights.com/ai-index/)
- OpenAI (2025-09-29), [Buy it in ChatGPT: Instant Checkout and the Agentic Commerce Protocol](https://openai.com/index/buy-it-in-chatgpt/)
- TechCrunch (2025-11-18), [Intuit signs $100M+ deal with OpenAI to bring its apps to ChatGPT](https://techcrunch.com/2025/11/18/intuit-signs-100m-deal-with-openai-to-bring-its-apps-to-chatgpt)
- Menlo Ventures (2026-09-15), [2026: The State of Consumer AI](https://menlovc.com/wp-content/uploads/2026/09/menlo_ventures_consumer_ai_report-2026.pdf)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- Underneath (2026), [When does Google show an AI Overview? 1,248 US searches](https://underneath.agency/research/ai-overviews-frequency-study)
- Underneath (2026), [Do AI Overviews cite the pages that rank? 4,051 citations](https://underneath.agency/research/ai-overview-citations-study)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [“Is this brand legit?” How AI assistants build a reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study)
- Underneath (2026), [Same question, four countries: do AI recommendations change?](https://underneath.agency/research/ai-recommendations-by-country-study)
- Underneath (2026), [When does a Reddit thread become evidence in Google’s AI?](https://underneath.agency/research/ai-reddit-citations-study)
- Underneath (2026), [The YouTube videos Google’s AI cites are small](https://underneath.agency/research/ai-overview-youtube-videos-study)

---

This is the Markdown twin of https://underneath.agency/resources/fintech-software-customers-from-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How fintech startups win customers in AI search"
description: "By owning the category explainer, earning independent coverage early and stating fees plainly. AI favors known brands until it finds a clear reason not to."
canonical: "https://underneath.agency/resources/fintech-startups-customers-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How can a fintech startup get recommended by AI when established brands own the answers?

By giving AI assistants a clear, verifiable reason to name you: independent coverage, an honest explanation of your category, and fees and protections stated plainly. Research shows assistants default to well-known brands when options look alike, and that new products are almost invisible in open-ended questions. The same research suggests the default breaks when a challenger has distinct, checkable evidence, which a startup can start building before it can outspend anyone.

## The short version

1. New products rarely surface in open questions: in a [study of 112 Product Hunt startups](https://arxiv.org/abs/2601.00912), ChatGPT recognized them 99.4% of the time when asked by name, but named them in only 3.32% of discovery questions.
2. The incumbent default is real but fragile: in a [controlled test](https://arxiv.org/abs/2606.17443), assistants picked the well-known brand 100% of the time when products looked identical, but a competitor’s edge of less than 0.1 rating stars broke that.
3. Independent coverage predicts recommendations: in [our brand study](https://underneath.agency/research/brand-entity-ai-recommendations-study), each tenfold increase in independent sites naming a brand went with 4.7 times the odds of being recommended.
4. Fintech capital is scarce, so cheaper channels matter: [CB Insights](https://www.cbinsights.com/research/report/fintech-trends-q2-2026/) counted 726 fintech deals in the second quarter of 2026, the fewest in more than four years.
5. New categories reward early customers: [Klarna](https://s205.q4cdn.com/644747736/files/doc_news/Klarna-Delivers-Strong-Start-to-2026-With-1Bn-Revenue-and-68M-Adj--Operating-Profit-2026.pdf) says customers who joined in 2022 brought $12 each in their first year and $52 a year now.

A note before you read: we write here about how fintech brands show up in AI answers. Nothing below is financial, legal or regulatory guidance.

## Why is AI search a different problem for a fintech startup?

Because assistants lean on what is already well documented, and a startup has little documentation yet.

An established brand has years of reviews, press, comparison-site listings and links. A startup may have a better product and almost none of that. When someone asks an assistant an open question, such as “what’s the best way to get paid before payday?”, the assistant builds its answer from what it can find and trust.

Two studies show the size of the gap:

- **Recognition is not recommendation.** In the Product Hunt study by Amit Prakash Sharma, ChatGPT recognized startups asked about by name 99.4% of the time and Perplexity 94.3%. For discovery-style questions, the rates fell to 3.32% and 8.29%. For Perplexity, referring domains, an ordinary search signal, predicted who appeared.
- **Known brands win ties.** Xi Chu and YuPeng Hou tested three assistants on product choices with identical specifications. The well-known brand was recommended every time, a pattern they call a “conditional monopoly.” It disappeared when a rival had even a small, visible quality advantage.

The second study used skincare products and lab-built product lists, so it does not prove how assistants treat financial products. Our inference for fintech: an assistant will name the incumbent unless it finds specific, credible evidence that you are different. Generic claims (“faster, cheaper, smarter”) give it nothing to work with. Insurtechs face the same problem, covered in [how insurtechs win customers from incumbents](https://underneath.agency/resources/insurtech-customers-ai-search).

Money is also tighter. CB Insights reports fintech funding fell 20% to $11.7B in the second quarter of 2026, with most of it going to a handful of mega-rounds, and just 4 new fintech unicorns. For most startups, we infer, that makes a channel built on evidence, rather than ad spend, more valuable.

## Which new fintech categories depend most on explanation?

Categories people do not yet understand, such as earned wage access, where the first question is “what is it?”

New categories create a sequence of questions that incumbents do not own yet:

- **Earned wage access (EWA).** Workers draw pay they have already earned before payday. A [Congressional Research Service brief](https://www.everycrsreport.com/files/2024-07-31_IF12727_016d4fdcf42f18b428fa53fa04aa4279c6e61b87.html) cites a CFPB estimate that in 2022 more than 10 million workers used these products, totaling $32 billion. It also cites the CFPB’s findings that 82% of employer-partnered transactions had fees and that the average effective annual rate was 109%.
- **Buy now, pay later (BNPL).** The [Federal Reserve’s household survey](https://www.federalreserve.gov/publications/files/2024-report-economic-well-being-us-households-202505.pdf) found use edged up to 15 percent of adults in 2024. The category’s early challengers are now its incumbents: Klarna reported 119 million active consumers, and [Affirm](https://www.digitaltransactions.net/affirms-results-boost-its-bnpl-ranking/) reported 24.1 million active consumers and 419,000 merchants.

The BNPL story is the useful lesson. Today’s incumbents were once the unknown option. In a category that is still forming, such as EWA, the companies that explain it clearly and honestly are, we infer, the ones whose pages and coverage assistants learn the category from.

Fee and cost questions come early in these categories, and regulators are watching. The CRS brief describes EWA rules as varying by state. Pages that state fees and costs plainly are both a regulatory necessity and the material an accurate AI answer needs. Lenders meet the same cost questions, as our guide on [reaching borrowers who ask AI](https://underneath.agency/resources/lending-platforms-borrowers-ai-search) shows.

## What do customers ask AI about a new kind of financial product?

What it is, whether it is safe, what it costs, and how it compares with what they already use. These prompts are illustrative, written by us, not observed data.

| Stage | Illustrative prompt |
|---|---|
| Category | “What is earned wage access, and is it a loan?” |
| Cost | “How much do pay-early apps really cost if I use them every week?” |
| Alternatives | “Alternatives to payday loans that won’t trap me in debt” |
| Comparison | “Klarna vs Affirm vs my credit card for a $1,200 laptop” |
| Challenger check | “Is [new app] legit, and who holds my money?” |
| Business buyers | “Newer alternatives to our bank for a startup that pays overseas contractors” |

Consumers bring these questions to AI often. In an [Intuit Credit Karma survey](https://stacker.com/stories/personal-finance-investing/rise-fin-ai-why-americans-are-trusting-generative-ai-their), 66% of Americans who had used generative AI said they had used it to seek financial advice. Credit Karma is part of Intuit, a large incumbent, so read its survey as vendor research.

## How does early AI visibility compound for a startup?

Customers won early tend to grow in value, so each one an AI answer sends is worth more over time.

Klarna’s results show the shape. Its 2022 cohort generated $12 in annual revenue per consumer in the first year and $52 today, as those customers used more services. Klarna’s gross merchandise volume reached $33.7 billion in the first quarter of 2026. A startup’s first cohorts, we infer, follow the same logic: the earlier a customer joins, the longer the revenue runs.

AI visibility also compounds on the evidence side. Coverage, reviews and comparison listings earned this year stay on the web, and our studies suggest assistants draw on exactly that material. [Our freshness study](https://underneath.agency/research/ai-source-freshness-study) found that pages under 90 days old took 17.4% to 22.6% of the dated citations in each assistant’s answers, against 6.9% of Google’s top 10 on the same questions, a gap that favors a young fintech with new pages. New, well-made pages are not shut out.

## What gives a challenger a fair chance of being named?

Distinct, verifiable evidence that appears in independent sources; platforms document only part of how this works.

**Documented by platforms.** Google says its AI features may use [“query fan-out”](https://developers.google.com/search/docs/appearance/ai-features), running several related searches across subtopics before answering. A comparison question can therefore pull in pages about each option.

**Observed in studies.**

- **Independent coverage.** In our brand study, coverage of a brand across independent sites in the pages assistants cited was the strongest predictor we measured, at 4.7 times the odds per tenfold increase.
- **Search foundations.** In the Product Hunt study, referring domains predicted Perplexity visibility, while scores for content written specifically for AI did not.
- **A real difference.** In the controlled brand test, a small, visible quality edge was enough to break the incumbent default.

**Our inference for fintech.** The difference has to be true, specific and checkable: a fee that is lower in writing, a protection the incumbent lacks, a license or partner bank named plainly, a feature reviewers confirm. The same controlled study found that fabricated authority claims could also shift answers. In financial services, invented claims are a regulatory and reputational risk, and [GEO can backfire](https://underneath.agency/resources/can-geo-backfire-on-your-brand) when it outruns the facts.

## What does GEO look like for a fintech startup?

Generative engine optimization (GEO) for a challenger means building the evidence an incumbent already has, faster and more precisely.

1. **Own the category explainer.** Publish the clearest honest guide to your category: what it is, what it costs, who it suits and who it does not.
2. **Earn independent coverage early.** Fintech and personal finance press, comparison sites and analyst notes; aim for specific facts, not just a funding announcement.
3. **Honest comparisons.** Pages comparing you with the incumbent on fees, speed and protections, kept current. Our research on [whether comparison pages help brands get cited by AI](https://underneath.agency/resources/do-comparison-pages-help-b2b-ai-citations) shows what tends to work.
4. **Plain regulatory facts.** Licenses, partner banks, how funds are held and every fee, written with counsel, on public pages. Our guide on [how neobanks answer safety and fee questions](https://underneath.agency/resources/neobanks-account-openings-ai-search) shows this for app-only banks.
5. **Reviews from day one.** App store, Trustpilot and BBB profiles, with complaints answered.
6. **Search basics.** Links from credible sites and pages that rank, which still feed several assistants.
7. **Launch checks.** New products are often missing from answers at first; [why ChatGPT doesn’t mention a newly launched product](https://underneath.agency/resources/why-chatgpt-misses-new-products) explains why and what to do.

The broader playbook for smaller brands is in [how a small brand can get recommended by AI assistants](https://underneath.agency/resources/how-small-brands-get-recommended-by-ai). No one can guarantee an assistant will recommend a startup; GEO makes the evidence it finds accurate, specific and easy to verify.

## What does the evidence not tell fintech startups yet?

It shows incumbents are favored by default; it does not show how assistants treat financial products specifically.

- **No fintech-specific test.** The Product Hunt study covered startups across categories, and the brand test used skincare and household goods; neither tested financial products.
- **Lab conditions.** The controlled brand test used product lists the researchers built, not live web search.
- **Vendor surveys.** Consumer AI usage figures in finance come largely from companies with products to sell.
- **No public link to sign-ups.** We found no public data connecting a fintech startup’s AI visibility to customers acquired. The wider evidence on incumbents is in [do AI assistants favor big brands over smaller competitors?](https://underneath.agency/resources/do-ai-assistants-favor-big-brands)

## Where should a fintech startup start?

Start by asking assistants your category, alternative and comparison questions, and note who is named and why.

A first check shows whether assistants understand your category at all, which incumbents and sources they name instead, and whether your fees, licenses and partner bank are described correctly. That map shows where a small amount of evidence could change the answer.

If your growth plan depends on taking customers from better-known financial brands, [ask us to run that check with you](https://underneath.agency/contact). We will test how assistants answer your category, alternative and comparison questions, show which evidence the incumbents have that you lack, and plan the coverage, content and reputation work, reviewed with your compliance team, to close the gap. Our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) puts that work in order for a challenger, starting with an honest category explainer, plain fee and license pages and early independent coverage.

## Frequently asked questions

### Can a new fintech brand appear in ChatGPT answers at all?

Yes, but rarely at first. In one study, startups named directly were recognized 99.4% of the time but appeared in only 3.32% of open discovery questions.

### Do AI assistants always prefer big financial brands?

Not always. In a controlled test, the known brand won every tie, but a small, visible quality advantage for a rival broke that default. Specific, verifiable differences matter.

### Is it worth explaining our category if competitors benefit too?

Usually, yes. In a new category such as earned wage access, the company whose explanation is clearest and most widely cited shapes how assistants describe the whole category.

### Should a fintech startup write content just for AI?

No. In the Product Hunt study, content scores built for AI did not predict visibility, while referring domains did. Build evidence and coverage that people trust too.

## Sources

- Amit Prakash Sharma (2025), [The Discovery Gap: How Product Hunt Startups Vanish in LLM Organic Discovery Queries](https://arxiv.org/abs/2601.00912)
- Xi Chu and YuPeng Hou (2026), [Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems](https://arxiv.org/abs/2606.17443)
- CB Insights (2026-07), [State of Fintech Q2’26](https://www.cbinsights.com/research/report/fintech-trends-q2-2026/)
- Klarna (2026-05-14), [Klarna Delivers Strong Start to 2026 With $1Bn Revenue and $68M Adj. Operating Profit](https://s205.q4cdn.com/644747736/files/doc_news/Klarna-Delivers-Strong-Start-to-2026-With-1Bn-Revenue-and-68M-Adj--Operating-Profit-2026.pdf)
- Digital Transactions (2025-11), [Affirm’s results boost its BNPL ranking](https://www.digitaltransactions.net/affirms-results-boost-its-bnpl-ranking/)
- Congressional Research Service (2024-07-31), [Earned Wage Access](https://www.everycrsreport.com/files/2024-07-31_IF12727_016d4fdcf42f18b428fa53fa04aa4279c6e61b87.html)
- Board of Governors of the Federal Reserve System (2025-05), [Economic Well-Being of U.S. Households in 2024](https://www.federalreserve.gov/publications/files/2024-report-economic-well-being-us-households-202505.pdf)
- Intuit Credit Karma, via Stacker (2026-06-29), [The rise of fin-AI: Why Americans are trusting generative AI with their wallets](https://stacker.com/stories/personal-finance-investing/rise-fin-ai-why-americans-are-trusting-generative-ai-their)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- Underneath (2026), [Do Wikipedia and schema make AI assistants recommend a brand?](https://underneath.agency/research/brand-entity-ai-recommendations-study)
- Underneath (2026), [How fresh are the pages AI engines cite?](https://underneath.agency/research/ai-source-freshness-study)

---

This is the Markdown twin of https://underneath.agency/resources/fintech-startups-customers-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How do you fix wrong information about your brand in AI answers?"
description: "Trace the error to the page it came from and fix that page. AI answers are built mostly from live web sources, often your own old pages."
canonical: "https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How do you fix wrong information about your brand in AI answers?

You fix it at the source: find the page the AI answer is repeating, then correct, retire or outweigh that page. Most wrong facts about a brand are not invented by the AI. They are old, conflicting or third-party facts that are still live on the web.

## The short version

1. In one agent study, only 7–10% of the finished answer came from the AI’s built-in memory; the rest came from pages it read ([Finder and colleagues](https://arxiv.org/abs/2609.34951), 2026, a vendor study).
2. Of 64 software prices that differed from the official pricing page, 39 were on another page of the vendor’s own site ([our pricing study](https://underneath.agency/research/ai-pricing-accuracy-study)).
3. Where a business’s Google profile number was missing from its website, 30.6% of phone numbers AI gave differed from the profile, against 1.6% where it was there ([our business facts study](https://underneath.agency/research/ai-business-facts-accuracy-study)).
4. In a simulated test, a single polluted page among the search results fooled AI assistants into recommending a fake product up to 27% of the time ([Luo and Chen](https://arxiv.org/abs/2606.13610), 2026).

## Can you actually fix what an AI assistant says about you?

Often yes, because today’s AI answers are rebuilt from live web pages each time, not fixed in memory.

[Finder and colleagues](https://arxiv.org/abs/2609.34951) ran 37,927 AI agent journeys about 1,056 real businesses. Only 7–10% of the finished answer came from the model’s training knowledge, whether or not the business’s site was readable. The rest came from what the agent fetched. The authors work for ora, which sells the readiness score their study uses, so treat the exact figures with care.

That matters for anyone trying to correct a mistake. If an answer is assembled from pages, changing the pages changes the raw material. It also means a wrong fact usually has an address you can find. In [our pricing study](https://underneath.agency/research/ai-pricing-accuracy-study), only 4 of 64 differing prices could not be found on any source we could fetch.

## Why does source provenance matter in generative search?

Because an AI answer inherits the strengths and errors of whatever pages it draws on, usually without telling the reader.

Provenance means where a claim came from. [Xu, Iqbal and Montgomery](https://arxiv.org/abs/2605.14021) checked 98,020 claims in Google’s AI Overviews, the AI summaries at the top of Google’s results. They found 11.0% were not supported by the pages cited, and that the credibility of a source and the accuracy of the claim were largely unrelated. In 1.39% of claims, the cited sources disagreed with each other. When your own pages disagree, the AI can pick either version. Two live guides cover the wider problem: [how often AI answers say things their sources do not support](https://underneath.agency/resources/ai-answers-unsupported-claims) and [whether an AI can cite your page for something it does not say](https://underneath.agency/resources/ai-citing-pages-for-claims-they-do-not-make).

Our pricing study shows the same inheritance. When a quoted price differed from the official page, we looked for that figure elsewhere. It appeared on the product’s own cited pages 90.0% of the time, against 20.0% on pages cited for other products. The AI was copying from its sources, not guessing.

Provenance also explains the darker risk. In the simulated test by [Luo and Chen](https://arxiv.org/abs/2606.13610), one polluted page among ten real search results was enough to push a fake brand into recommendations up to 27% of the time. AI assistants resisted best in categories whose real brands they already knew well. A defense that ranked sources by credibility removed only about a sixth of the fakes.

## Which kind of error are you dealing with?

Start by sorting the error, because each type has a different source and a different fix.

| What the AI got wrong | Where it usually comes from | Where to fix it |
|---|---|---|
| An old price, plan or product name | Old pages, help articles or announcements on your own site | Update or retire every page that states it |
| A phone number, address or hours | Your website and listings disagree | Publish one version everywhere |
| A fact it leaves out | Your site is hard for AI agents to read | Put key facts in plain text on readable pages |
| A complaint or reputation claim | Review and complaint platforms | Resolve the issue where customers report it |
| A claim no source makes | Unclear; the AI may have misread a page | Document it and check again over several runs |

The live guide [Do AI agents make up facts about my business, or just leave them out?](https://underneath.agency/resources/do-ai-agents-invent-or-omit-business-facts) covers the gap between missing and wrong facts. This guide is about the repair.

## How do you fix errors that start on your own pages?

Make every page you control say the same current thing, and remove the old versions that compete with it.

Conflicting first-party sources are the most common cause we have measured. In [our business facts study](https://underneath.agency/research/ai-business-facts-accuracy-study), we asked four AI engines for the contact details of 159 local businesses. For 24.2% of the businesses we could check, the Google profile number did not appear anywhere on their own website. Those businesses produced most of the differing numbers: 30.6% of the numbers AI gave for them differed from the profile, against 1.6% for the rest.

Prices follow the same pattern. Of 64 differing software prices, 39 were on another page of the vendor’s own site, such as an old plan FAQ or a post about an earlier price change. Pages with a monthly/annual price toggle had 58.4% fully faithful prices, against 72.9% without one. A toggle hides half the price list from a single read of the page.

Readability matters too. In the ora agent study, answers built from a business’s own site were 41% more accurate than answers about the same business built from the wider web. If an agent cannot load your page, it fills the gap from other people’s pages. See [where AI agents get their answer when they can’t read your site](https://underneath.agency/resources/when-ai-agents-cant-read-your-site).

## What about errors that come from other people’s sites?

Fix them where they live, because AI engines lean heavily on review platforms and user-written pages.

In [our reputation study](https://underneath.agency/research/is-it-legit-ai-reputation-study), four AI engines answered “Is this brand legit?” for real brands. 88.0% of answers cited a review or complaint platform. Claims attached only to review platforms were negative 56.5% of the time, against 6.4% for claims attached only to the brand’s own website. Trustpilot and the BBB accounted for 61.7% of all review-platform citations.

The same study found a weakness worth knowing. 71.4% of claims that a problem was common rested on evidence that did not show how common it was. A handful of loud complaints can read as a pattern. Resolving recurring complaints, and answering them publicly, changes the evidence the engines cite.

For directories, listings and news articles, the fix is ordinary and slow: ask the owner to correct the page. No study we reviewed tested whether a correction request to an AI company itself changes its answers.

## How do you know a fix has worked?

Check repeatedly, on every major engine, because one answer proves little and engines differ widely.

In our business facts study, 4.4% of Gemini’s answers had a fact that differed from the Google profile, against 37.1% for Perplexity. A fix that shows up on one engine may not show up on another. Run the same buyer questions on each engine you care about.

Tone moves more than presence. In one vendor’s tracking data, whether the AI framed a brand positively or negatively flipped about 6.7 times more often than whether it mentioned the brand at all ([Kumar](https://arxiv.org/abs/2606.20065), 2026). Judge a correction over many runs, as explained in [why ChatGPT gives a different answer about your brand each time](https://underneath.agency/resources/why-ai-answers-about-your-brand-change).

## What should you do about it?

Treat wrong AI answers as a source-cleanup job with a named owner and a regular check.

1. **Log each error with its engine, question and date.** Note any page the answer cites; it is your first lead.
2. **Search for the wrong fact itself.** Look on your own site first, including help centers, old plan pages and press posts.
3. **Retire or update old pages.** Redirect or rewrite pages that state outdated prices, plans, numbers or claims.
4. **Publish one version of each core fact.** Use the same phone number, hours, plan names and prices on your site, profiles and listings.
5. **Show prices and key facts in plain text.** State monthly and annual prices side by side instead of behind a toggle.
6. **Work the review platforms.** Fix recurring complaints at the root and reply where customers report them.
7. **Ask third parties to correct their pages.** Prioritize the pages the AI answers actually cite.
8. **Re-check on a schedule.** Ask the same questions on each engine over several runs before calling it fixed.

To decide who that named owner should be, see [which teams should own AI search visibility](https://underneath.agency/resources/who-should-own-ai-search-visibility).

If you want help with this kind of source cleanup, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The research shows where wrong facts come from, but no study has yet timed how fast a fix shows up.

- No study we reviewed measured how long it takes for a corrected page to change AI answers, on any engine.
- No study tested correction or feedback channels offered by AI companies.
- The agent study comes from a company that sells agent-readiness scoring and uses its own score.
- Our studies cover local businesses and software prices, on single days, so other industries may differ.
- The polluted-page test was simulated on frozen search results, mainly in Chinese, not on the live web.
- Whether an AI that cites a bad source can be talked out of it by better sources elsewhere is untested.

## Frequently asked questions

### Can I ask ChatGPT or Google to correct wrong information about my company?

The research we reviewed has not tested that route. What it does show is that answers draw on live pages, so correcting the pages the answer relies on is the evidence-backed path.

### Why does AI show an old price for my product?

Usually because an old price is still on the web, often on your own site. In our pricing study, 39 of 64 differing prices were on another page of the vendor’s own site.

### Does fixing my Google Business Profile fix AI answers?

It helps only if your other sources agree. Phone numbers differed from the profile 30.6% of the time when the website did not show the profile number, against 1.6% when it did.

### Can a competitor plant false information about my brand in AI answers?

It is possible in principle. In a simulated test, a single planted page fooled AI assistants into recommending a fake product up to 27% of the time, and none of the tested defenses was adequate.

## Sources

- Finder, Elovic, Shalev and Yosef (2026), [AX is the New AEO](https://arxiv.org/abs/2609.34951), arXiv:2609.34951.
- Xu, Iqbal and Montgomery (2026), [Measuring Google AI Overviews: Activation, Source Quality, Claim Fidelity, and Publisher Impact](https://arxiv.org/abs/2605.14021), arXiv:2605.14021.
- Luo and Chen (2026), [One Polluted Page Is Enough: Evaluating Web Content Pollution in LLM Recommenders](https://arxiv.org/abs/2606.13610), arXiv:2606.13610.
- Kumar (2026), [Generative Engine Optimization at Scale: Measuring Brand Visibility Across AI Search Engines](https://arxiv.org/abs/2606.20065), arXiv:2606.20065.
- Underneath (2026), [How faithfully do AI assistants quote software prices?](https://underneath.agency/research/ai-pricing-accuracy-study)
- Underneath (2026), [Do AI answers match a business’s Google profile?](https://underneath.agency/research/ai-business-facts-accuracy-study)
- Underneath (2026), [“Is this brand legit?” How AI assistants build a reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study)

---

This is the Markdown twin of https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How can food ecommerce brands turn AI meal planning into orders?"
description: "By being the products AI names when shoppers plan meals and build lists: complete food data, compliant claims, reviews and presence where carts are filled."
canonical: "https://underneath.agency/resources/food-ecommerce-sales-from-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How can food ecommerce brands turn AI meal planning into orders?

By making sure that when shoppers ask AI to plan meals, match a diet or build a grocery list, your products are the ones named, described accurately and easy to buy. Shoppers already use AI for recipes and shopping lists, and assistants can now place grocery orders inside the chat. Food brands that supply complete, honest product information and earn independent reviews give assistants a reason to pick them.

## The short version

1. Online food buying is large and growing fast: US online grocery sales passed $128.6 billion in 2025, up 32.9%, and reached 19% of weekly grocery spending in December, per Brick Meets Click and Mercatus data reported by [Talk Business & Politics](https://talkbusiness.net/2026/02/the-supply-side-online-grocery-sales-surge-more-than-32-in-december/).
2. Grocery shoppers use AI for food decisions: in a January 2026 [FMI](https://www.fmi.org/our-research/research-reports/u-s-grocery-shopper-trends/fmi-blog/2026/01/29/from-smartphones-to-ai--how-shoppers-use-technology-to-feed-their-families) survey of 1,518 shoppers, 53% had used AI tools for at least one food need such as recipes, meal planning or shopping lists, and only 18% always verified what AI told them.
3. Meal plans can now become orders in the chat: [Instacart’s app in ChatGPT](https://www.supermarketnews.com/grocery-technology/instacart-launches-end-to-end-shopping-app-on-chatgpt) builds a shopping list from a recipe and checks out without leaving the conversation, and [Walmart](https://corporate.walmart.com/news/2025/10/14/walmart-partners-with-openai-to-create-ai-first-shopping-experiences) announced that customers could plan meals and buy through ChatGPT.
4. Diet questions drive demand: in the [IFIC 2025 Food & Health Survey](https://www.foodnavigator.com/Article/2026/01/21/ific-survey-how-americans-view-health-and-food-choices/), 57% of Americans tried a specific diet in the past year, and high-protein was the most followed at 23%.
5. Repeat orders are where food brands make money: at [HelloFresh](https://webdisclosure.com/press-release/hellofresh-se-etr-fy-2025-hellofresh-se-continues-to-show-strong-aebitda-performance-while-efficiency-program-progresses-meaningfully-0kE0IoWUrh7), most orders by the end of 2025 came from customers who had already ordered 50 or more boxes.

## Who buys food online, and what is a customer worth?

A household shopper who orders often, so the value of a customer comes from repeat baskets, not one sale.

Food ecommerce covers three kinds of business: online grocers and delivery platforms (Walmart, Amazon, Kroger, Instacart), direct-to-consumer food brands such as meal kits, specialty pantry, snack, coffee, meat and seafood companies, and the packaged-food brands whose products fill online carts at those retailers.

The channel is now mainstream. Brick Meets Click and Mercatus report December 2025 online grocery sales of a record $12.7 billion, with active users completing an average of 2.9 orders that month. Retailers dominate: Walmart held 31.6% and Amazon 22.6% of online grocery sales in 2025, according to the same report. Over the 2025 holidays alone, [Adobe](https://news.adobe.com/news/2026/01/adobe-holiday-shopping-season) counted $23.7 billion in online grocery spending, up 10.2%.

What a customer is worth depends on frequency. HelloFresh, a public meal kit company, reported an average order value in North America of €78.3 in 2025 and said a majority of its orders now come from long-tenured customers. Its newer customers also became more valuable: net revenue per conversion after 20 weeks was 21% higher in the second half of 2025 than two years earlier. For a food brand, acquiring a customer who reorders for months matters far more than any single basket. Our guide for [subscription brands winning new subscribers](https://underneath.agency/resources/subscription-ecommerce-customers-ai-search) looks at that model across categories.

## Where does AI already sit in food shopping?

Before the cart: in recipes, meal plans, diet questions and the shopping list itself.

The strongest evidence comes from the grocery industry’s own trade association. FMI found that 68% of grocery shoppers have used AI tools such as ChatGPT and 26% are regular users. They use them for recipes, best prices or promotions, ingredient or nutrition facts, meal or menu planning, grocery lists and diet or nutrition advice. Only 18% always verify AI-generated information, and 35% rarely or never do.

Seasonal surveys point the same way. In a November 2025 [Qlik](https://www.qlik.com/us/news/company/press-room/press-releases/qlik-survey-family-still-rules-the-kitchen-but-ai-joins-the-thanksgiving-table) survey, 54% of respondents said they used an AI tool like ChatGPT or Copilot to help plan, prep or cook a holiday meal, and 31% said they would lean on AI more than family to make the shopping list. Qlik sells data software, so treat the survey as directional.

Grocers and delivery apps are now building the purchase step into their AI tools. Instacart says it was the first app in ChatGPT to offer checkout inside the chat, and it automates a shopping list from a recipe for the user to approve and pay for. Walmart’s announcement describes customers who are “planning meals, restocking household essentials, or finding something new” and can “simply chat and buy.” Supermarket News reports that Albertsons has launched its own AI shopping assistant and that Kroger is adding AI features with Instacart.

There is one caution in the data. Early in 2025, [Adobe](https://blog.adobe.com/en/publish/2025/03/17/adobe-analytics-traffic-to-us-retail-websites-from-generative-ai-sources-jumps-1200-percent) found conversion from AI-referred traffic was lowest in the apparel, home goods and grocery categories. Retail-wide, AI traffic later converted 42% better than other traffic by March 2026, according to Adobe data reported by [TechCrunch](https://techcrunch.com/2026/04/16/ai-traffic-to-us-retailers-rose-393-in-q1-and-its-boosting-their-revenue-too/), but no grocery-specific update was published. We infer that food sales from AI arrive more often through a planned list or an in-chat checkout than through a click to a brand page.

## What do food shoppers ask AI before they buy?

Mostly meal and diet questions, with products appearing as ingredients, swaps or delivery options.

These prompts are illustrative, written by us to show the patterns. They are not captured from real shoppers, and we do not claim what any assistant returns for them.

| Need | Illustrative prompt |
|---|---|
| Meal plan | “Plan five high-protein dinners for two and give me a shopping list” |
| Diet fit | “Gluten-free pasta that holds up in a baked dish” |
| Delivery | “Where can I order good-quality steaks delivered in Ohio?” |
| Subscription | “Best meal kit for a vegetarian couple on a budget” |
| Swap | “A lower-sugar alternative to my usual granola” |
| Gift | “Food gift box for a coworker who loves hot sauce” |

The pattern matters for revenue. Many of these answers turn into lists of ingredients, and an ingredient can be listed generically (“Greek yogurt”) or as a specific product. We infer that a brand gets named when the assistant has a reason to name it: a recipe that calls for it, reviews that single it out, or a retailer catalog that offers it clearly for the stated diet and budget. Home decor brands face a similar gap when shoppers ask for a look rather than a product, as our guide on [decor discovery in AI answers](https://underneath.agency/resources/home-decor-brands-ai-product-discovery) explains.

## How does an AI meal plan turn into a sale?

Through a named product or list item, a cart at a retailer or brand, and then repeat orders.

1. **Planned.** The shopper asks for a meal plan, a diet-friendly swap or a gift idea.
2. **Listed.** The assistant names dishes and ingredients, sometimes specific brands or stores.
3. **Carted.** The list moves into a retailer cart, such as Instacart’s app inside ChatGPT, or the shopper clicks to a brand’s own site.
4. **Repeated.** A product that works becomes part of the weekly basket or a subscription. HelloFresh’s results show why this step carries the value.

OpenAI has explained how products and merchants are surfaced. When it launched [Instant Checkout](https://openai.com/index/buy-it-in-chatgpt/) in September 2025, it said product results were “organic and unsponsored, ranked purely on relevance to the user,” and that when several merchants sold the same item, ChatGPT considered “availability, price, quality, whether a merchant is the primary seller, and whether Instant Checkout is enabled.” OpenAI scaled the feature back in March 2026, according to [FashionUnited](https://fashionunited.com/news/retail/us-consumers-would-let-ai-agents-buy-clothes-but-not-without-a-say/2026092974919), though its help page still describes it for some eligible merchants. For food brands that sell through retailers, accurate listings at those retailers are part of being bought, not just being named.

## What decides which food products an assistant names?

Product facts, reviews and independent food coverage, filtered by diet, budget and availability.

**Documented by the platform.** OpenAI says ChatGPT’s [product results](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search) draw on structured information such as price and product description from first-party and third-party providers, other third-party content, and its safety standards and product policies. It also says review summaries are generated from public websites and are not verified by OpenAI, and that price updates can lag.

**Observed in studies.**

- In a test of soft-drink recommendations, ChatGPT relied heavily on Wikipedia, about 19.7% of its citations, and on a small set of consumer-review and food sites such as Tasting Table and Sporked ([Chen and colleagues](https://arxiv.org/abs/2509.08919)). Our guide on [how drinks brands get recommended](https://underneath.agency/resources/beverage-brands-ai-recommendations) covers that category in depth.
- In a product-ranking test of three AI systems, product facts such as rating, price and reviews explained 82.4% of how products were ranked ([Chu and Hou](https://arxiv.org/abs/2606.17443)). We explain what that means in [what actually drives AI product recommendations](https://underneath.agency/resources/what-drives-ai-product-recommendations).
- When we logged the searches ChatGPT runs behind the scenes in [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study), it searched for reviews or ratings on 46.2% of the buyer questions we tested.

**Our inference for food.** The trust signals likely to matter are complete nutrition and ingredient data, clear allergen and diet labels, ratings at the retailers that supply carts, coverage in food publications and recipe sites, and claims that a shopper, or an assistant, can check.

## What is the cost of being missing, or described wrongly?

A missed place in weekly baskets, and the risk that an assistant repeats a wrong or unsupported claim.

No food company has published how many orders it gains or loses through AI answers, so the cost cannot be measured precisely. The mechanism is clear enough: if an assistant builds a high-protein meal plan and lists a rival’s yogurt or a generic one, your product is not in that cart, and possibly not in the habits that follow. Pet food brands face the same repeat-order stakes, covered in [how pet product brands get recommended](https://underneath.agency/resources/pet-product-brands-ai-search).

Accuracy is the sharper risk. FMI found 35% of shoppers who use AI rarely or never verify what it tells them, so a wrong allergen, price or nutrition detail can reach the order unchecked. Claims also have rules. The [FDA](https://www.fda.gov/food/nutrition-food-labeling-and-critical-foods/label-claims-conventional-foods-and-dietary-supplements) says health claims describe a relationship between a food substance and reduced risk of a disease, and are subject to FDA oversight; terms such as “high” and “low” must meet FDA definitions, and “healthy” is itself a regulated, implied nutrient content claim. A brand that publishes loose health language risks having it repeated in AI answers, which is a regulatory and a trust problem, not a growth tactic. Baby product brands carry similar safety-detail stakes, as our guide on [what parents ask AI before buying](https://underneath.agency/resources/baby-product-brands-ai-recommendations) shows.

Answers also shift over time. In [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), only 25.2% of the brands ChatGPT named for a question appeared in all five repeats, so one good test result proves little.

## How does GEO work for a food ecommerce brand?

Generative engine optimization (GEO) makes your food products easy for assistants to find, match to a meal and describe correctly.

For food ecommerce, the work usually includes:

1. **Complete product data everywhere.** Ingredients, allergens, nutrition per serving, pack sizes, storage, price, delivery area and diet labels that meet their definitions, identical on your site, retailer catalogs and product feeds. Our review of [the product content AI shopping assistants prefer](https://underneath.agency/resources/product-content-ai-shopping-assistants-prefer) shows which fields matter most.
2. **Recipes that use your products.** Plain, useful recipes and meal plans that name your product as an ingredient, organized around the diets people follow, such as high-protein or gluten-free.
3. **Claims you can support.** Factual statements such as grams of protein per serving rather than health promises; keep any regulated claim within FDA rules.
4. **Retailer and marketplace presence.** Accurate listings and genuine ratings at the grocers whose carts assistants fill, since that is where many AI orders will close.
5. **Independent coverage.** Reviews and taste tests from food publications, recipe creators and community threads, earned rather than bought.
6. **Measurement across assistants.** Ask a fixed set of meal-plan, diet and delivery questions in ChatGPT, Gemini, Perplexity and Google’s AI features, repeat them, and track whether you are named and described correctly. Our guide to [why analytics miss AI visibility](https://underneath.agency/resources/why-analytics-miss-ai-visibility) explains why traffic reports alone undercount this.

None of this guarantees that an assistant puts your product in a basket. It raises the chance that, when an assistant looks, your facts are the clearest and most trusted available.

## What does the evidence not tell food brands yet?

How many grocery orders start in an AI answer, and how in-chat grocery apps choose brands.

- **No food-specific AI conversion data.** Adobe’s grocery finding predates the latest retail-wide improvement, and Brick Meets Click does not break out AI-originated orders.
- **Surveys measure self-reports.** FMI and Qlik asked shoppers what they do; Qlik sells data software.
- **In-chat grocery apps are new.** Instacart and Walmart describe what their ChatGPT experiences do, not how they choose between brands for a generic list item.
- **Studies are adjacent.** The citation finding comes from soft drinks, and the ranking test used skincare.

## Where should a food ecommerce brand start?

Start by asking AI the meal and diet questions your customers ask, and seeing whose products fill the list.

Run your core meal-plan, diet, delivery and gift questions in the assistants and in-chat grocery apps your shoppers use, several times each. Record whether your products are named, whether nutrition, allergen and price details are right, and which sources the answers draw on. That baseline shows whether the gap is product data, retailer listings, recipe content or independent coverage.

If more orders and longer subscriptions are the goal, [we can map this for your pantry or meal brand](https://underneath.agency/contact). We will show where your products appear when shoppers ask AI to plan meals and build lists, why other brands appear instead, and which fixes are most likely to put your products into more carts. The steady work behind those fixes, such as matching nutrition and allergen data across retailers and recipes built around your products, is set out on our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page.

## Frequently asked questions

### Do grocery shoppers really use ChatGPT to plan meals?

Many do. FMI’s January 2026 survey found 68% of grocery shoppers have used AI tools like ChatGPT, and 53% have used them for food needs such as recipes, meal planning or shopping lists.

### Can shoppers buy groceries inside ChatGPT?

Yes, in some cases. Instacart launched a ChatGPT app that turns a recipe into a shopping list and checks out in the chat, and Walmart announced shopping through ChatGPT with Instant Checkout in 2025. FashionUnited reports that OpenAI cut back that checkout feature in March 2026, while OpenAI’s help page still lists it for some eligible merchants.

### Should food brands add health claims to get recommended?

No. No evidence shows health claims help, and FDA rules govern health and nutrient content claims. Clear, accurate nutrition facts are a safer basis for being described correctly.

### Does selling through Instacart or Walmart replace a brand’s own GEO work?

No. Retailer listings help a product get bought, but assistants also draw on reviews and food publications, and OpenAI says review summaries come from public websites.

## Sources

- Talk Business & Politics (2026-02), [The Supply Side: Online grocery sales surge more than 32% in December](https://talkbusiness.net/2026/02/the-supply-side-online-grocery-sales-surge-more-than-32-in-december/)
- Blue Book Services (2026-01-23), [U.S. eGrocery sales surge 32% YOY to a record in December 2025](https://www.bluebookservices.com/?p=449043)
- FMI – The Food Industry Association (2026-01-29), [From Smartphones to AI: How Shoppers Use Technology to Feed Their Families](https://www.fmi.org/our-research/research-reports/u-s-grocery-shopper-trends/fmi-blog/2026/01/29/from-smartphones-to-ai--how-shoppers-use-technology-to-feed-their-families)
- Qlik (2025-11-11), [Qlik Survey: Family Still Rules the Kitchen, but AI Joins the Thanksgiving Table](https://www.qlik.com/us/news/company/press-room/press-releases/qlik-survey-family-still-rules-the-kitchen-but-ai-joins-the-thanksgiving-table)
- FoodNavigator-USA (2026-01-21), [What ‘healthy’ means to Americans is changing, and so should marketing](https://www.foodnavigator.com/Article/2026/01/21/ific-survey-how-americans-view-health-and-food-choices/)
- Supermarket News (2025-12), [Instacart launches end-to-end shopping app on ChatGPT](https://www.supermarketnews.com/grocery-technology/instacart-launches-end-to-end-shopping-app-on-chatgpt)
- Walmart (2025-10-14), [Walmart Partners with OpenAI to Create AI-First Shopping Experiences](https://corporate.walmart.com/news/2025/10/14/walmart-partners-with-openai-to-create-ai-first-shopping-experiences)
- HelloFresh SE via EQS (2026-03-18), [FY 2025: HelloFresh SE continues to show strong AEBITDA performance while efficiency program progresses meaningfully](https://webdisclosure.com/press-release/hellofresh-se-etr-fy-2025-hellofresh-se-continues-to-show-strong-aebitda-performance-while-efficiency-program-progresses-meaningfully-0kE0IoWUrh7)
- Adobe (2026-01-07), [Holiday Shopping Season Drove a Record $257.8 Billion Online with Consumers Embracing Generative AI Tools](https://news.adobe.com/news/2026/01/adobe-holiday-shopping-season)
- Adobe (2025-03-17), [Adobe Analytics: Traffic to U.S. retail websites from generative AI sources jumps 1,200 percent](https://blog.adobe.com/en/publish/2025/03/17/adobe-analytics-traffic-to-us-retail-websites-from-generative-ai-sources-jumps-1200-percent)
- TechCrunch (2026-04-16), [AI traffic to US retailers rose 393% in Q1, and it’s boosting their revenue too](https://techcrunch.com/2026/04/16/ai-traffic-to-us-retailers-rose-393-in-q1-and-its-boosting-their-revenue-too/)
- OpenAI (2025-09-29), [Buy it in ChatGPT: Instant Checkout and the Agentic Commerce Protocol](https://openai.com/index/buy-it-in-chatgpt/)
- FashionUnited (2026-09-29), [US consumers would let AI agents buy clothes, but not without a say](https://fashionunited.com/news/retail/us-consumers-would-let-ai-agents-buy-clothes-but-not-without-a-say/2026092974919)
- OpenAI (2026), [Shopping with ChatGPT Search](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search)
- U.S. Food and Drug Administration, [Label Claims for Conventional Foods and Dietary Supplements](https://www.fda.gov/food/nutrition-food-labeling-and-critical-foods/label-claims-conventional-foods-and-dietary-supplements)
- Chen and colleagues (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919)
- Chu and Hou (2026), [Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems](https://arxiv.org/abs/2606.17443)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)

---

This is the Markdown twin of https://underneath.agency/resources/food-ecommerce-sales-from-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How can footwear brands get their shoes found through AI search?"
description: "By making each shoe model easy for AI to match to a fit, use and price question, so it lands on the shortlist before shoppers pick a size and a store."
canonical: "https://underneath.agency/resources/footwear-brands-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How can a footwear brand get its shoes found when shoppers ask AI what to buy?

By making each shoe model easy for AI assistants to match to a shopper’s fit, use and budget question, and easy to verify on independent sites and retailer listings. Footwear is bought model by model, and comfort and fit come first, so the work happens at the level of the Clifton or the Cloudmonster, not the logo. This guide covers who decides, where AI now sits in the shoe-buying journey, and what a brand can and cannot control.

## The short version

1. The prize is large and moving toward performance: US footwear sales held at $90 billion in 2025, with running up 9% and walking up 9% in dollars, according to [Circana](https://mr-mag.com/circana-issues-q4-2025-footwear-industry-pulse-report/); in the first half of 2026, running shoes rose 13% in value and volume ([World Footwear](https://www.worldfootwear.com/news/us-footwear-sales-rise-slightly-in-the-first-half/11731.html)).
2. Shoppers decide on comfort and fit first, then quality, then price, in a [Footwear Insight survey](https://www.formula4media.com/new-articles/trend-insight-0w1cl) of 307 active consumers. Those are exactly the questions people now type into AI assistants in full sentences.
3. Awareness is still the constraint even for fast-growing brands: [On](https://press.on-running.com/on-reports-results-for-the-second-quarter-and-six-month-period-ended-june-30-2026) reports global brand awareness of 30% while its own channels reached 45.7% of second-quarter sales.
4. Much of the sale still happens at retailers: Deckers’ wholesale sales were $3.21 billion against $2.26 billion direct to consumers in fiscal 2026 ([FashionUnited](https://fashionunited.com/news/business/hoka-and-ugg-drive-deckers-brands-to-5-47-billion-dollars-in-fy26-net-sales/2026052272510)). An AI answer can send the shopper to either.
5. Google added shoes to its AI “try it on” tool in October 2025 ([Google](https://blog.google/products/shopping/virtual-try-on-shoes/)), and OpenAI documents that ChatGPT ranks the merchants for a product partly on whether they are “the maker or primary seller.”

## Who decides which shoes get bought, and what is one customer worth?

The shopper decides, model by model, and a good fit often turns into years of repeat pairs.

For a footwear brand, the buyer is not a procurement committee. It is a runner replacing worn-out trainers, a nurse who needs to survive 12-hour shifts, a parent buying school shoes, or someone chasing a new lifestyle sneaker. What they weigh is consistent. In the Footwear Insight survey, run on the MESH01 panel of active and outdoor consumers, comfort and fit ranked first, quality and durability second and price and value third. The same shoppers said they were most likely to buy running shoes (53%), casual or comfort shoes (47%), walking shoes (43%) and boots (43%) that fall and winter.

Price pressure is real. 66% of those shoppers had seen price increases on footwear they wanted, and 50% said they would wait for sales if prices kept rising. Circana noted that nearly half of consumers have delayed purchases or chosen cheaper alternatives because of price increases. A shopper who is stretching a budget researches harder before committing. Furniture shoppers spending hundreds to thousands of dollars research across several visits and channels, as our guide on [furniture brands winning high-ticket orders](https://underneath.agency/resources/furniture-brands-sales-from-ai-search) explains.

What a customer is worth depends less on one order than on the model they settle into. Runners and people who stand all day tend to rebuy a shoe that works. That habit is why the growth brands report the way they do: Deckers’ HOKA brand grew 15.9% to $2.59 billion in the fiscal year to March 2026, and On’s shoe sales reached CHF 781.6 million in a single quarter. Our inference is that the first match between a shopper and a model is the valuable moment, because it can set the next several purchases.

Executives should keep the channel split in mind. At Deckers, wholesale sales grew 12.3% and direct sales 6.3% last fiscal year. At On, direct sales grew 26.0% in the second quarter, and the company credits its rising direct share, along with full-price discipline, for a 65.4% gross margin. So the same AI answer can feed a higher-margin sale on your site or a sale through a retail partner, depending on where the shopper clicks. How a retailer competes for that same click, with accurate prices and stock, is shown in our guide on [how beauty retailers turn AI picks into sales](https://underneath.agency/resources/beauty-retailers-ai-search).

## Where do AI assistants already sit in the shoe-buying journey?

Between the need and the shortlist, where shoppers used to ask a store specialist or read reviews.

The cross-retail evidence is strong. [Adobe’s data](https://techcrunch.com/2026/04/16/ai-traffic-to-us-retailers-rose-393-in-q1-and-its-boosting-their-revenue-too/), reported by TechCrunch, shows AI traffic to US retail sites rose 393% in the first quarter of 2026, and that this traffic converted 42% better than other traffic in March 2026. Adobe’s survey found 39% of people had used AI for online shopping. These are all-retail figures, not footwear figures, but they describe the same shoppers.

Google and OpenAI have both built shopping features that suit shoes. Google says its Shopping Graph now has more than 50 billion product listings, with more than 2 billion refreshed every hour, and its [AI Mode shopping announcement](https://blog.google/products/shopping/google-shopping-ai-mode-virtual-try-on-update/) opens with a shoe question: “How do I choose a pair of hiking boots?” In October 2025 Google extended its try-on tool to shoes so shoppers can see “those heels or sneakers” on a photo of themselves ([TechCrunch](https://techcrunch.com/2025/10/08/googles-virtual-try-on-shopping-tool-expands-to-more-countries-now-lets-you-try-on-shoes/)). [OpenAI’s help page](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search) describes product results with images, prices, review summaries and links, plus a try-on button for clothes and accessories that, it warns, does “not guarantee fit or size.” Our guide on [how fashion brands get matched to a look](https://underneath.agency/resources/fashion-ecommerce-ai-recommendations) covers these visual and try-on tools for clothing.

Our own data shows how often Google already answers before the click. In [our study of when Google shows AI Overviews](https://underneath.agency/research/ai-overviews-frequency-study), 68.0% of the 100 retail and ecommerce keywords showed one, and 95.0% of the question-form retail keywords did. Footwear questions are often questions: which shoe, for what, in which size.

## What do shoe shoppers ask AI assistants?

Long, specific questions about fit, use, comparisons, sizing and price, usually naming a problem rather than a product.

We wrote the examples below to illustrate the kinds of questions footwear shoppers ask; they are not observed prompts:

- Use case: “Best cushioned running shoe for a heavier runner training for a first marathon.”
- Fit: “Which running shoes have a wide toe box but still feel light?”
- Comparison: “Hoka Clifton vs Brooks Ghost for someone who walks 10 miles a day at work.”
- Sizing: “Does this model run small compared with my usual size in Nike?”
- Work and lifestyle: “Comfortable black shoes for a restaurant server that pass a dress code.”
- Weather: “Waterproof sneakers that don’t look like hiking boots.”
- Price: “Best walking shoes under $100 that last more than a year.”

Research on AI shopping suggests these questions are long for a reason. [Bagga and colleagues](https://arxiv.org/abs/2511.20867) built a set of 13,747 shopping requests modeled on how people describe needs on Reddit; they averaged about 59 words, against three or so words for typical search keywords. Shoppers state budgets, past experiences and must-have features. A shoe page that only says “responsive cushioning” gives an assistant little to match against “my knees hurt after long shifts on concrete.”

Some of these questions touch on foot pain or injury. Brands should describe construction and intended use plainly and leave diagnosis to clinicians. That is good practice in any channel, and claims that sound medical invite more scrutiny, not less.

## How does an AI answer turn into a pair sold?

Through a model named on a shortlist, a size decision, and a click to your site or a retailer.

The path in footwear has four steps, and each one can be lost.

**Named.** The assistant suggests a few models for the stated need. If your shoe is not among them, nothing after this point happens. On’s 30% global brand awareness, reported at a time of strong growth, is a reminder that most shoppers anywhere still do not know most brands; an answer that names a model is a way to be introduced. Luxury brands see the same pattern, since most of their shoppers’ AI questions name no brand at all, as [how luxury brands reach high spenders](https://underneath.agency/resources/luxury-brands-ai-search) shows.

**Described correctly.** The shopper reads how the assistant characterizes the shoe: cushioned or firm, narrow or roomy, true to size or not. A wrong description loses the sale, or worse, wins a sale that comes back as a return.

**Sized.** Fit is the top purchase factor, so the shopper needs a confident size. Try-on tools help with looks, but both Google and OpenAI frame them as visual. Clear size guidance on your own pages and retailer pages is what the answer can repeat. Our guide for [apparel brands on fit and size facts](https://underneath.agency/resources/apparel-brands-customers-ai-search) covers the same problem for clothing.

**Bought.** OpenAI documents that when a shopper opens a product, ChatGPT may list several merchants, ranked “based on factors like availability, price, quality, and whether they are the maker or primary seller of that item.” A reasonable expectation is that a brand with accurate stock and price data on its own store can capture more of these clicks, while a brand whose retailers have better data will see the sale go to them. Either is revenue, at different margins.

Shopify reports results along this path for its merchants. Its [2026 holiday report](https://www.shopify.com/news/agentic-holiday-2026) says 65% of shoppers plan to use AI for at least one shopping task this season, and notes that only one-third trusted an agent to buy for them. In other words, AI shapes the choice, and the shopper still completes the purchase.

## What decides whether a shoe model gets named?

The platforms document product data and merchant factors; studies point to verifiable facts; the rest is our inference.

**Documented by the platforms.** OpenAI says ChatGPT considers “structured metadata from first-party and third-party providers (e.g., price, product description) and other third-party content,” and that review summaries are built from reviews on public websites. For Shopify merchants, OpenAI says product data already reaches ChatGPT through Shopify Catalog. Google says AI Mode runs several searches at once, a “query fan-out,” to work out what makes a product good for a stated situation, then suggests options against those criteria.

**Observed in studies.** In a controlled test summarized in [our article on what drives AI product picks](https://underneath.agency/resources/what-drives-ai-product-recommendations), product facts such as rating, price and reviews mattered far more than brand name when the assistant could see them. Bagga and colleagues found that rewrites that kept facts, named concrete attributes and answered likely buyer questions moved products up, while advertising-style copy did not. For a shoe page, that means fit and construction facts over taglines, as our guide to [the product content AI shopping assistants prefer](https://underneath.agency/resources/product-content-ai-shopping-assistants-prefer) explains.

**Our inference for footwear.** The trust factors shoe shoppers use are specific: width options, heel-to-toe drop, weight, cushioning, outsole grip, water resistance, how a model sizes against other brands, how long it lasts, and what changed between versions. When those facts match across your product page, your retailers’ listings and independent reviews, an assistant has consistent material to work with. When version 9 and version 10 of a shoe are described differently on different sites, a reasonable expectation is that answers will blur them. New models face an added delay; [our article on why ChatGPT misses new products](https://underneath.agency/resources/why-chatgpt-misses-new-products) explains why launches can be invisible for a while.

## What does a footwear brand lose when AI leaves its shoes out?

The first introduction to a shopper who might have bought the same model for years.

We have no published measure of lost footwear sales from AI answers, and we will not invent one. What the evidence supports is the shape of the risk. Running and walking are where US footwear growth is, according to Circana. Those are the categories where shoppers ask detailed fit and use questions. AI-referred shoppers converted better than other traffic in Adobe’s retail data. A brand left off the shortlist for “best shoe for standing all day” therefore loses the shoppers most ready to buy, in the categories that are growing.

There is a second, quieter cost: a shoe named with the wrong fit or features can turn into returns and poor reviews, which then feed the next answer. Because answers change from run to run, a single check tells you little; [our article on why AI answers about your brand change](https://underneath.agency/resources/why-ai-answers-about-your-brand-change) covers how to measure it properly.

## How does generative engine optimization work for a footwear brand?

By making each model’s fit, use and price facts clear, consistent and confirmed by independent sources wherever AI looks.

The work is practical and model by model:

- **Model pages that answer fit questions.** Width options, drop, stack height, weight, sizing advice against common reference brands, and the use each shoe is built for, written in plain words a shopper would use.
- **Clean product data everywhere it flows.** Feeds to Google Merchant Center, Shopify Catalog or OpenAI’s merchant program, with live price, stock and every size and width, so assistants can list you as the maker with accurate availability.
- **Retailer listings that match.** Many sales still happen at wholesale partners. Their pages should describe each model the same way you do.
- **Independent coverage.** Lab-style reviews, specialty running and workwear publications, and honest customer reviews give an assistant something to verify against. Shopify quotes Stanley 1913’s commerce lead saying that this kind of optimization means positioning a brand’s story “in other authoritative sources.”
- **Version clarity.** Name each new version, say what changed, and keep older versions’ pages accurate.
- **Visibility tracking across assistants.** Measure which models are named, for which use cases, in ChatGPT, Google AI Mode, AI Overviews, Gemini and Perplexity, repeated over time.

None of this can guarantee a recommendation. No one controls what an assistant says. What a brand controls is whether the facts it needs are available, consistent and believable.

## What don’t we know yet about AI and shoe sales?

No one has published how often shoe shoppers use AI, or how many pairs it sells.

The AI shopping figures here come from all of retail, not footwear alone. We found no public footwear-specific measure of AI-referred sales, return rates for AI-referred orders, or how often assistants name specific shoe models. Platform documentation describes product results in general, not how shoes are ranked. Try-on tools are new, and no one has published whether they change size choice or returns. Brand results such as Shopify’s merchant stories are self-reported. Read these numbers as pointing the way for footwear, not as measured shoe results.

## Where should a footwear brand start?

With an audit of how AI assistants describe your best-selling models for the fit and use questions buyers ask.

Pick the five or ten models that carry your revenue, list the use-case, fit and comparison questions their buyers ask, and check what ChatGPT, Google’s AI features, Gemini and Perplexity say, run several times. Look for three things: whether each model is named, whether its fit and features are described correctly, and which store the answer sends shoppers to. If you would like a hand, [send us your top models](https://underneath.agency/contact) and we will map where your models appear, where they are missing or misdescribed, and the work most likely to put them on more shortlists that end in a sale. Our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page explains how that model-by-model work runs, from fit-focused model pages and clean feeds to matching retailer listings and clear version names.

## Frequently asked questions

### Do AI assistants recommend specific shoe models or just brands?

Shopping features in ChatGPT and Google show individual products with prices, images and links, so the unit is usually the model. That is why model pages, sizes and versions need accurate, consistent facts.

### Can AI try-on tools tell shoppers which size to buy?

No. Google’s try-on shows how shoes look on a photo, and OpenAI warns that its try-on images “do not guarantee fit or size.” Size advice still has to come from your pages, retailer listings and reviews.

### Will AI shopping send sales to my retailers instead of my own store?

Sometimes. OpenAI says ChatGPT ranks merchants on availability, price, quality and whether the seller is the maker. Accurate stock and price data on your own store helps, and retailer sales still count as revenue.

### Should footwear brands make health claims to match pain-related questions?

No. Describe construction and intended use plainly, and leave medical advice to clinicians. Claims that sound medical attract scrutiny from regulators, retailers and shoppers alike.

### How long before changes show up in AI answers?

It varies by assistant and by source. Product data can update quickly; independent reviews and articles take longer to appear and be picked up. Track the same questions repeatedly rather than checking once.

## Sources

- Circana via MR Magazine (2026), [Circana issues Q4 2025 footwear industry pulse report](https://mr-mag.com/circana-issues-q4-2025-footwear-industry-pulse-report/)
- World Footwear (September 2026), [US footwear sales rise slightly in the first half](https://www.worldfootwear.com/news/us-footwear-sales-rise-slightly-in-the-first-half/11731.html)
- Footwear Insight, Formula4 Media (July 2025), [Trend Insight: footwear shopping behavior](https://www.formula4media.com/new-articles/trend-insight-0w1cl)
- On Holding (August 2026), [On reports results for the second quarter and six-month period ended June 30, 2026](https://press.on-running.com/on-reports-results-for-the-second-quarter-and-six-month-period-ended-june-30-2026)
- FashionUnited (May 2026), [Hoka and Ugg drive Deckers Brands to 5.47 billion dollars in FY26 net sales](https://fashionunited.com/news/business/hoka-and-ugg-drive-deckers-brands-to-5-47-billion-dollars-in-fy26-net-sales/2026052272510)
- TechCrunch (April 2026), [AI traffic to US retailers rose 393% in Q1, and it’s boosting their revenue too](https://techcrunch.com/2026/04/16/ai-traffic-to-us-retailers-rose-393-in-q1-and-its-boosting-their-revenue-too/)
- Google (May 2025), [Shop with AI Mode, use AI to buy and try clothes on yourself virtually](https://blog.google/products/shopping/google-shopping-ai-mode-virtual-try-on-update/)
- Google (October 2025), [Google Shopping’s try on tool adds shoes, expands to new countries](https://blog.google/products/shopping/virtual-try-on-shoes/)
- TechCrunch (October 2025), [Google’s virtual try-on shopping tool expands to more countries, now lets you try on shoes](https://techcrunch.com/2025/10/08/googles-virtual-try-on-shopping-tool-expands-to-more-countries-now-lets-you-try-on-shoes/)
- OpenAI (2026), [Shopping with ChatGPT Search](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search)
- Shopify (October 2026), [Welcome to the first holiday season of the agentic era](https://www.shopify.com/news/agentic-holiday-2026)
- Bagga and colleagues (2025), [E-GEO: A Testbed for Generative Engine Optimization in E-Commerce](https://arxiv.org/abs/2511.20867)
- Underneath (2026), [When does Google show an AI Overview? 1,248 US searches](https://underneath.agency/research/ai-overviews-frequency-study)

---

This is the Markdown twin of https://underneath.agency/resources/footwear-brands-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How do freight companies win shippers who research with AI?"
description: "Shippers check freight providers through rankings, reviews and quotes; AI answers now sit on that path, so carriers and brokers must be easy to verify by lane."
canonical: "https://underneath.agency/resources/freight-companies-shippers-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How do freight carriers, brokers and forwarders win shippers who research with AI?

By being the provider an AI answer can tie to a specific lane, mode and service record, and then answering the quote fast. Shippers pick freight partners mostly on reliability and trust, and logistics staff now use AI tools far more than a year ago while ranking AI last among their information sources. Visibility in AI search helps a freight company get the quote request; verified service and quick pricing win the load.

## The short version

1. The market is huge and fragmented: the US trucking freight bill was about $906 billion in 2024, and of almost 580,000 active motor carriers, 91.5% operate 10 or fewer trucks, according to the [American Trucking Associations](https://www.trucking.org/economics-and-industry-data).
2. Trust is the scarce asset: in the [35th Annual Study of Logistics and Transportation Trends](https://www.logisticsmgmt.com/article/35th_annual_study_of_logistics_and_transportation_trends_trust_but_verify), only 16% of respondents had high trust in freight brokers, against 33% for freight forwarders and about 53% for asset-based carriers.
3. AI use is rising faster than trust in it: the same study found AI use among logistics employees rose from 45% to 65% in a year, yet only 7% reported high trust in AI-generated outputs.
4. Shippers buy on service: 67% of shippers in a [Denim and Peerless Research Group survey](https://www.denim.com/news/freight-survey-reveals-what-shippers-really-want-from-brokers-and-carriers) named service level and reliability as the top factor in choosing a freight partner; only 10% named price.
5. Speed of quoting is now a weapon: [C.H. Robinson](https://www.freightwaves.com/news/how-is-c-h-robinson-using-ai-its-cfo-has-a-story-to-tell) receives about 600,000 rate quote requests a year in North America and, with an AI tool, now answers all of them in about 32 seconds, against 17 to 20 minutes before.

## Who buys freight, and what is a shipper worth?

Shippers’ transportation and procurement teams buy it, lane by lane, and a won lane pays every week.

The US trucking market alone is enormous. The ATA estimates trucks moved roughly 72.7% of the nation’s freight by weight in 2024. Forwarding is global and concentrated at the top: [Logistics Management’s Top 25 Freight Forwarders report](https://www.logisticsmgmt.com/article/top_25_freight_forwarders_navigating_a_new_era) cites Transport Intelligence’s estimate of a market worth about $240 billion in 2025, with DSV at $37.4 billion, DHL at $35.5 billion and Kuehne+Nagel at $33.8 billion in gross revenue.

Brokerage shows the unit economics. Members reporting to the Transportation Intermediaries Association handled 1.74 million shipments in the second quarter, with an average invoice of $1,690 per shipment, and truckload accounted for 72% of broker activity, [Logistics Management reported](https://www.logisticsmgmt.com/article/tia_q2_report_signals_potential_freight_market_rebound). A shipper that moves a few loads a week on one lane is therefore worth hundreds of thousands of dollars a year to a broker or carrier, and a contract freight relationship across several lanes can be worth much more. The arithmetic is ours, from those averages.

The buying unit varies by provider type:

| Provider | Typical buyer | What they are choosing |
|---|---|---|
| Asset carrier (FTL, LTL) | Transportation manager, freight procurement | Lanes, equipment, service levels, contract rates |
| Freight broker | Shipping or logistics manager, often for spot or overflow freight | Capacity on demand, carrier vetting, price |
| Intermodal provider | Transportation manager on long-haul lanes | Cost against transit time |
| Ocean or air forwarder | Importer, exporter, supply chain lead | Trade-lane coverage, customs, booking reliability |

Warehouse and fulfillment providers sell differently again; see [how 3PLs make the AI shortlist](https://underneath.agency/resources/logistics-providers-shipper-contracts-ai-search).

## How do shippers choose a freight provider today?

On reliability first, then trust, then price, with independent rankings and reviews as checks.

The Denim survey covered nearly 100 shippers, a small sample, but its message matches other evidence: service and reliability dominate, and poor communication or invoicing drive churn. Independent rankings formalize this. In [Mastio & Co.’s annual LTL survey](https://www.freightwaves.com/news/ltl-service-survey-lists-old-dominion-top-national-carrier), the firm conducted 1,630 interviews with LTL shippers and rated 147 carriers on 28 service metrics, from on-time delivery and damages to billing accuracy and claims. Old Dominion Freight Line ranked as the top national carrier for a 16th consecutive year.

Trust is under strain. In the 35th Annual Study, 39% of respondents said trust among trading partners had deteriorated over five years, and only 22% said it had improved. Fraud is part of the reason: 47% were very or extremely concerned about double brokering, 45% about carrier identity fraud and 53% about AI-generated or fabricated documents. The study was fielded shortly after the Supreme Court’s Montgomery v. Caribe Transport II ruling on negligent-selection claims against brokers, and its authors describe a “trust, but verify” era that depends on documented verification.

Review platforms play a role too. CarrierSource describes itself as a platform for transportation provider reviews and sells brokers and carriers data on what shippers research there, [according to Air Freight News](https://airfreight.news/articles/full/carriersource-launches-ai-powered-research-agents-to-turn-shipper-intent-data-into-revenue). That a business can be built on shipper research activity shows that shippers do look up providers before they call.

## Where do AI assistants sit in a shipper’s research?

Early, as one research tool among several, and treated with caution. No public study yet measures how often shippers ask AI assistants to find freight providers.

The 35th Annual Study shows AI is now part of daily logistics work. Overall employee use of AI rose from 45% to 65% between 2025 and 2026, and the share using it with their organization’s approval rose from 16% to 47%. But 38% reported low or no trust in AI-generated outputs, and AI ranked last among the seven data and information sources the study evaluated.

For freight sellers, that combination means two things. Shippers will use assistants to build a list or compare options, so absence there costs a look. And they will verify what the assistant says, so the facts it repeats must match what they find on the provider’s site, on review platforms and in rankings.

Digital booking is also mainstream in forwarding. [Freightos](https://www.prnewswire.com/il/news-releases/freightos-reports-second-quarter-2026-results-302852762.html) reported 458 thousand transactions and a record $422 million in gross booking value in the second quarter of 2026, with about 21 thousand unique buyer users booking freight on its platforms. Shippers who book online are a short step from asking an assistant where to book.

## What do shippers ask AI assistants about freight?

Lane-specific, mode-specific and trust-specific questions. The examples below are hypothetical, written to show the pattern, not taken from logs.

- “Which reefer carriers run Laredo to Chicago with team drivers?”
- “LTL carriers with liftgate and inside delivery in the Northeast.”
- “Freight forwarder for full-container ocean shipments from Vietnam to Savannah, with customs brokerage.”
- “Is intermodal cheaper than truckload from Los Angeles to Dallas, and how much slower?”
- “How do I check that a freight broker is legitimate and not double brokering?”
- “How much does it cost to ship two pallets from Atlanta to Denver?”

Three of these are about capacity on a lane, one about mode choice, one about trust and one about price. Rate questions are hard for any assistant because freight prices move daily; we expect answers to give ranges and point to quoting tools rather than firm rates. Parcel carriers face their own version of the rate question, covered in [how shipping platforms win small sellers](https://underneath.agency/resources/shipping-companies-customers-ai-search). The trust question is the one brokers should worry about most, given the fraud concerns above.

## How does AI visibility become freight revenue?

Through the quote request. An AI answer that names a provider for a lane leads to a quote, a test load, then recurring freight.

The path is short and fast:

1. A shipper asks about capacity or a mode on a specific lane.
2. The answer names a few providers, often citing rankings, directories or the providers’ own lane pages.
3. The shipper requests quotes from some of them, by email, web form or portal.
4. The fastest credible quote wins a test load.
5. Good service turns the test into a lane award, often in the next bid.

Step four is where many freight companies lose the value of visibility. C.H. Robinson’s finance chief told FreightWaves the firm historically answered only about 60% to 65% of its quote requests; its AI tool now answers 100%. After it added less-than-truckload freight to its AI quoting, monthly LTL quote volumes jumped by at least 30%, [FreightWaves reported](https://www.freightwaves.com/news/c-h-robinson-delivers-on-ai-but-investors-still-skeptical). A smaller broker or carrier cannot match that engineering, but it can make sure every page an assistant might cite leads to a quote path that answers the same day.

## What decides whether an AI assistant names a freight company?

The platforms document only the basics; studies point to rankings and review platforms. We label each type of evidence.

Documented by the platforms:

- Google says a page needs to be indexed and eligible for a snippet to appear as a supporting link in AI Overviews or AI Mode, with no special optimization required, and that both may run several related searches to answer one question ([Google Search Central](https://developers.google.com/search/docs/appearance/ai-features)).
- OpenAI says ChatGPT search may send rewritten, targeted queries to partner search providers ([OpenAI Help Center](https://help.openai.com/en/articles/9237897-chatgpt-search)).

Observed in our studies:

- When answering buyer questions, ChatGPT ran a search aimed at a named publication, ranking or award in 43.8% of its answers ([hidden searches study](https://underneath.agency/research/ai-hidden-searches-study)). In freight, the Mastio rankings, Logistics Management’s Top 25 lists and trade press are such sources.
- Asked whether a brand was legitimate, 88.0% of AI answers cited a review or complaint platform ([reputation study](https://underneath.agency/research/is-it-legit-ai-reputation-study)). For brokers facing double-brokering fears, what review platforms say is likely to shape that answer.

Our inference: assistants answering lane questions need facts they can match to the question, such as lanes served, equipment, terminals, service area, transit times and authority details. Providers who state those facts plainly, consistently and in more than one place give an assistant something to cite. Our article on [whether small websites can get cited by AI](https://underneath.agency/resources/can-small-websites-get-cited-by-ai) is relevant here, since most carriers are small.

## What does it cost a freight company to be missing?

Mostly invisible losses: quote requests that never arrive. No study measures them in freight yet.

We have no figure for freight demand lost to AI answers, and we will not estimate one. What the evidence does show is where the risk sits. Shippers buy on reliability and trust, trust in brokers is low, and verification is becoming routine. If an assistant cannot find clear, consistent facts about a provider, it can leave the provider out or describe it vaguely; if review platforms carry complaints, those may surface when a shipper asks whether the company is legitimate. Either way, a shipper may never send the quote request. Our guide on [fake reviews and fake brands in AI recommendations](https://underneath.agency/resources/fake-reviews-fake-brands-ai-recommendations) covers why verification matters in categories with fraud.

## How does generative engine optimization work for a freight company?

It makes a carrier, broker or forwarder easy to match to a lane and easy to verify. It cannot promise that an assistant will name the company.

For freight, the work looks like this:

1. **Lane, mode and equipment pages.** One page per major lane or region and per mode (FTL, LTL, intermodal, ocean, air), stating equipment, transit times, accessorials and terminals. Generic “nationwide coverage” pages give an assistant nothing to match.
2. **Verification facts in one place.** Operating authority, insurance coverage, safety and compliance information, years in business and how the company vets carriers. For brokers, a clear page on fraud prevention answers the “is this broker legit” question directly.
3. **Service proof.** Published on-time and claims figures where a company can stand behind them, customer case studies by lane and mode, and participation in independent surveys such as Mastio’s.
4. **Third-party presence.** Reviews on independent platforms, listings in industry rankings and directories, and expert commentary in FreightWaves, Logistics Management or the Journal of Commerce.
5. **A fast quote path.** Every page an assistant might cite should lead to a quote form or email that is answered within hours.
6. **Monitoring by question.** Track a fixed set of lane, mode and trust questions across ChatGPT, Gemini, Google AI Overviews and AI Mode, Perplexity and Copilot, and record which providers and sources appear.

Freight technology vendors face a different problem, covered in our guides on [logistics software in AI search](https://underneath.agency/resources/logistics-software-ai-search) and [supply chain software on the RFP longlist](https://underneath.agency/resources/supply-chain-software-ai-search). For more on the role of third-party lists, see [best-of lists and AI recommendations](https://underneath.agency/resources/best-of-lists-ai-recommendations).

## Which questions remain open for freight?

Several, and they should shape how much a freight company invests and how it measures results.

- How many shippers ask AI assistants to find carriers, brokers or forwarders. The 35th Annual Study measures AI use at work, not provider search.
- Whether assistants favor large brokers over small carriers for lane questions. Studies of other industries suggest brand prominence matters, but freight has not been tested.
- How assistants handle rate questions when prices change daily.
- How stable answers are from one week to the next for the same lane question.

## Where should a freight company start?

With its top lanes and its trust story. Check what AI answers say about both, then fix the gaps.

List the five lanes or trade routes that matter most to revenue and the modes you sell on them. Ask the main assistants the questions a shipper would ask about each, plus “is [your company] legit?”, and record who is named and which sources are cited. Then compare that with your lane pages, review profiles and rankings. If you would rather hand that work off, [ask us to run the lane-by-lane check for you](https://underneath.agency/contact), starting with the lanes where you want more quote requests. Beyond that check, our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page describes the ongoing lane pages, verification facts and quote paths a carrier, broker or forwarder keeps current.

## Frequently asked questions

### Do shippers really use ChatGPT to find carriers or brokers?

Some likely do, but no public study measures it. Logistics workers’ AI use rose to 65% in 2026, while trust in AI outputs stayed low.

### Can a small carrier be named next to big brokers?

For narrow lane and equipment questions, plausibly yes, if its facts are clear and repeated elsewhere. Broad questions likely favor well-known brands.

### Should we publish freight rates on our website?

Rates change too fast for fixed lists. Publish how pricing works, what drives cost and how to get a quote quickly.

### Do LTL rankings like Mastio’s matter for AI answers?

Plausibly. Assistants often search for named rankings, and Mastio’s survey rated 147 carriers, but no study has tested its effect on AI answers.

### How can a freight broker counter fraud fears in AI answers?

Publish verification facts and carrier-vetting steps plainly, keep review profiles current and answer complaints in public.

## Sources

- American Trucking Associations (2026), [Economics and Industry Data](https://www.trucking.org/economics-and-industry-data)
- Logistics Management, Boone, Manrodt, Voss and colleagues (2026-09-01), [35th Annual Study of Logistics and Transportation Trends: Trust, but Verify](https://www.logisticsmgmt.com/article/35th_annual_study_of_logistics_and_transportation_trends_trust_but_verify)
- Logistics Management (2026-09), [Top 25 Freight Forwarders: Navigating a new era](https://www.logisticsmgmt.com/article/top_25_freight_forwarders_navigating_a_new_era)
- Logistics Management (2026), [TIA Q2 report signals potential freight market rebound](https://www.logisticsmgmt.com/article/tia_q2_report_signals_potential_freight_market_rebound)
- Denim and Peerless Research Group (2025-06-17), [Freight survey reveals what shippers really want from brokers and carriers](https://www.denim.com/news/freight-survey-reveals-what-shippers-really-want-from-brokers-and-carriers)
- FreightWaves (2025), [Survey reveals top LTL carriers for 2025](https://www.freightwaves.com/news/ltl-service-survey-lists-old-dominion-top-national-carrier)
- FreightWaves (2026-01), [How is C.H. Robinson using AI? Its CFO has a story to tell](https://www.freightwaves.com/news/how-is-c-h-robinson-using-ai-its-cfo-has-a-story-to-tell)
- FreightWaves (2025-04), [C.H. Robinson delivers on AI, but investors still skeptical](https://www.freightwaves.com/news/c-h-robinson-delivers-on-ai-but-investors-still-skeptical)
- Air Freight News (2025-09-15), [CarrierSource launches AI-powered research agents to turn shipper intent data into revenue](https://airfreight.news/articles/full/carriersource-launches-ai-powered-research-agents-to-turn-shipper-intent-data-into-revenue)
- Freightos, via PR Newswire (2026-08-17), [Freightos Reports Second Quarter 2026 Results](https://www.prnewswire.com/il/news-releases/freightos-reports-second-quarter-2026-results-302852762.html)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- OpenAI Help Center (2025), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [“Is this brand legit?” How AI assistants build a reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study)

---

This is the Markdown twin of https://underneath.agency/resources/freight-companies-shippers-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How can furniture brands sell more through AI search?"
description: "By being named when shoppers ask AI to narrow a sofa or bedroom search, with product facts, reviews and delivery terms an assistant can read and check."
canonical: "https://underneath.agency/resources/furniture-brands-sales-from-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How can a furniture brand turn AI shopping answers into more high-ticket orders?

By being on the short list when a shopper asks an AI assistant to narrow thousands of sofas, beds and tables down to three or four, and by giving that assistant the dimensions, materials, reviews and delivery terms it needs to choose you. Furniture is a considered, expensive, mostly hybrid purchase, so the AI answer rarely closes the sale on its own. It decides which brands get the website visit and the showroom trip.

## The short version

1. Furniture shoppers research before they buy. In Home News Now’s 2025 Consumer Insights Now survey, 64% of furniture shoppers visited a website and 51% visited a physical store to view their options, and 92% said a sofa must have positive online reviews.
2. Orders are large and customers come back. [Wayfair’s second quarter 2026 results](https://s24.q4cdn.com/589059658/files/doc_financials/2026/q2/2026-08-04-Press-Release.pdf) show an average order value of $332, and repeat customers placed 80.2% of orders.
3. AI shopping is real but still small for furniture. In an Exploding Topics survey of 1,009 US consumers, 77.6% had used AI for a shopping decision, but furniture was among the least common categories at roughly 29%. Wayfair’s CEO called traffic from AI shopping platforms “very small.”
4. The shoppers who do arrive from AI are valuable. Adobe data reported by TechCrunch shows AI traffic to US retailers rose 393% in the first quarter of 2026 and converted 42% better than other traffic in March.
5. Many furniture catalogs are hard for AI to read: Adobe found around 34% of retail product pages can’t be properly accessed by AI.

## Who buys furniture online, and what is a customer worth?

A household replacing a sofa or bedroom, spending hundreds to thousands of dollars, often across several visits and channels.

The category is large but not growing in stores. Census Bureau figures show [US furniture and home furnishings stores](https://www.census.gov/retail/marts/www/marts_current.pdf) sold $88,364 million in the first eight months of 2026, down 1.6% on the year, while nonstore retailers grew 10.3%. Growth is moving to channels where the shopper starts on a screen.

Upholstery leads. In the [Consumer Insights Now survey](https://homenewsnow.com/blog/2025/09/19/cin-week-2-looks-at-whats-driving-consumer-furniture-purchases/), published by the trade outlet Home News Now and sponsored by Bread Financial, 41% of furniture shoppers said they had bought a sofa that year, and 53% of younger millennials (ages 29 to 36) had bought upholstery. Area rugs, lamps and occasional tables followed, which tells you shoppers furnish a room, not just buy an item. Brands that sell those finishing pieces can see [how home decor brands get found in AI](https://underneath.agency/resources/home-decor-brands-ai-product-discovery).

Budgets are serious. In the [second wave of the same survey](https://homenewsnow.com/blog/2025/10/03/consumer-insights-now-week-3-looks-at-whats-driving-planned-2nd-half-furniture-purchases/), 48% wanted a sofa under $1,000 and 77% under $2,000, so almost a quarter expected to spend more. For bedroom furniture, 20% of shoppers were willing to spend more than $3,000. Mattress sellers competing for the same bedroom budget can see [how mattress brands win AI shoppers](https://underneath.agency/resources/mattress-brands-customers-from-ai-search).

The best public view of what an online furniture customer is worth comes from Wayfair. In the second quarter of 2026 it had 21.7 million active customers, earned $596 per active customer over the previous twelve months, and placed 64.1% of orders through a mobile device. It also spent $392 million on advertising in the quarter, against net revenue of $3,519 million. For a furniture brand, every customer who finds you without a paid click is worth a great deal.

## Where do AI assistants already sit in the furniture buying journey?

At the narrowing step: after inspiration, before the store visit or the product page.

The journey is hybrid. A [2025 study by 3D Cloud and Provoke Insights](https://3dcloud.com/news/3d-cloud-2025-furniture-study-release/) found 45% of furniture shoppers engaged with both digital and in-store channels. In the Consumer Insights Now survey, 27% of upholstery buyers started their search online, between 29% and 30% read online reviews or visited a retailer’s website, and 30% sat on the sofa before buying. Of shoppers who visited a website, 57% went to Amazon, 45% to a local store or Ashley, and 42% to Wayfair, followed by Walmart, Target and La-Z-Boy.

AI assistants are entering that journey, more slowly than in other categories. In the [Exploding Topics survey reported by TechWyse](https://www.techwyse.com/news/ai-search/ai-shopping-adoption-consumer-distrust-autonomous-spending-2026), product research was the most cited AI shopping use at 68.5%, and ChatGPT was the most used tool at 77.56% of AI shoppers, followed by Gemini at 58.21%. Clothing and electronics led the categories; furniture trailed at roughly 29%.

The largest furniture seller is preparing anyway. [Wayfair joined Google’s Universal Commerce Protocol](https://www.digitalcommerce360.com/2026/01/14/wayfair-google-universal-commerce-protocol/) in January 2026, so its products can be bought from listings in AI Mode in Google Search and in the Gemini app, with Wayfair as the merchant of record. On its May 2026 earnings call, CEO Niraj Shah said [Wayfair works with Perplexity, OpenAI and Google](https://www.digitalcommerce360.com/2026/05/05/wayfair-agentic-ai-commerce/) on AI shopping. “We want to be everywhere,” he said. He also said the traffic is still “very small.”

## Which questions lead furniture shoppers to a brand?

Questions about a room, a constraint and a budget, not just a product name.

The survey data shows what furniture shoppers care about: 99% said a sofa must be comfortable, 98% the right size, 96% easy to clean, 84% sold with a warranty and 87% delivered free. Those constraints become the prompt. We wrote the examples below to show the shape of these questions; they are illustrative, not logged queries:

- Room and use case: “Best performance-fabric sofa for a family with a dog, under $1,500, that fits an 84-inch wall.”
- Comparison: “Pottery Barn vs Crate & Barrel sofas: which holds up better?”
- Alternatives: “Sofas that look like the RH Cloud but cost less.”
- Logistics: “Sectionals with free white-glove delivery and easy returns.”
- Trust: “Is this online sofa brand good quality, or will it sag in a year?”

The room and constraint questions decide which brands make the short list. Comparison and trust questions decide which one gets the visit. Assistants are built for this kind of question: Google says AI Mode runs a [“query fan-out”](https://blog.google/products/shopping/google-shopping-ai-mode-virtual-try-on-update/), several searches at once, to work out what makes a product fit a need, then suggests options from its Shopping Graph.

## How does an AI answer turn into a furniture sale?

Through a short list: the answer names a few products, and the shopper visits one site or showroom.

**The short list.** An assistant answering a sofa question names a handful of options. If you are not one of them, the shopper may never search your brand name.

**The visit.** Visits from AI are worth more than they used to be. [Adobe’s data, reported by TechCrunch](https://techcrunch.com/2026/04/16/ai-traffic-to-us-retailers-rose-393-in-q1-and-its-boosting-their-revenue-too/), shows AI traffic to US retailers converted 42% better than other traffic in March 2026, and revenue per visit was 37% higher. Twelve months before, AI visitors had converted 38% worse. These figures cover all retail, not furniture alone.

**The store.** Because 45% of furniture shoppers move between screen and showroom, a reasonable expectation is that some AI influence ends as a store sale that no analytics tool links back to the answer. Shoppers ask for the sofa they saw named; they do not say where they saw it.

**The repeat order.** With repeat customers placing 80.2% of Wayfair’s orders, a first sale won through an AI answer can lead to rugs, lamps and the bedroom later.

Not every furniture sale will move this way. Shah expects AI agents to matter most for replenishment, commodities and technical goods, while home, fashion and beauty are categories where “there’s a lot of emotion” and “consumers actually don’t want to own the same items as each other.” He named inexpensive seating like barstools as the furniture most likely to be bought through agents, and said there is no margin in that volume. We infer the bigger prize for most furniture brands is not checkout inside the assistant. It is being in the answer that shapes the short list for a considered purchase.

## What decides whether an AI assistant names your furniture?

Product data, reviews and independent coverage; platforms document some of it, and studies show the rest.

**Documented by the platform.** OpenAI says that when [ChatGPT chooses shopping results](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search), it considers structured metadata from first-party and third-party providers, such as price and product description, plus other third-party content, and that it may consider price, reviews and other context. Its [product feed specification](https://developers.openai.com/commerce/specs/feed) has fields for material, dimensions and weight, which matter more for a sofa than for a T-shirt. Google says its Shopping Graph holds more than 50 billion product listings, each with details like reviews, prices, color options and availability, and that more than 2 billion are refreshed every hour.

**Observed in studies.** An [audit of 1,536 chatbot responses to real shopping questions](https://arxiv.org/abs/2609.18729) found ChatGPT expressed a first-person product preference in 79% of product-recommending responses, against 7% for Gemini, and the products recommended often changed across repeated requests. For the same question, ChatGPT and Gemini shared only 5.4% of the sources they displayed. Being named on one assistant says little about another. In [our study of brand recommendations](https://underneath.agency/research/brand-entity-ai-recommendations-study), independent coverage was the strongest signal we measured: each tenfold increase in the independent sites naming a brand in the cited pages went with 4.7 times the odds of being recommended.

**Trust factors specific to furniture.** Shoppers cannot touch a sofa through a chat window, so they lean on proxies: 92% want positive online reviews, 80% want easy returns, and 65% want a brand they know. We infer that the same proxies help an assistant: review volume and ratings on your site and on retailers that carry you, clear warranty and return terms, delivery options, and coverage by design editors and testing sites that answer “does it hold up?” Our piece on [what drives AI product recommendations](https://underneath.agency/resources/what-drives-ai-product-recommendations) covers how assistants weigh product facts against brand names.

## What does it cost a furniture brand to be missing?

Mostly the short list you never see, plus product pages AI cannot read.

Adobe found that roughly a quarter of the content on retailers’ homepages and category pages was not set up for AI, and around 34% of product pages can’t be properly accessed by AI. For furniture, the product page is where the facts that settle a decision live: dimensions, fabric, firmness, assembly, lead time. If an assistant cannot read them, we infer it falls back on whatever a marketplace or a review site says about you, or picks a competitor whose facts it can confirm.

The cost is hard to see because furniture shoppers often finish the purchase in a store or on another device. A brand missing from AI answers will not see a drop in a dashboard. It will see fewer people arriving already sure about one of its sofas. We do not have evidence of how many furniture sales this is today, and Wayfair’s own view is that the traffic is still small. The case for acting now is that the shoppers who do use AI arrive ready to buy, and the habits of the 53% of younger millennials buying upholstery will not reverse.

## How does GEO work for a furniture brand?

By making your products easy for AI to understand, check and recommend, without promising placement.

1. **Product facts in machine-readable form.** Dimensions, seat depth, materials, fabric performance, care, assembly, weight limits, lead time, delivery type and return window in your product pages, your structured data and your product feeds. Retailers that carry you need the same facts.
2. **Room and use-case content.** Pages that answer the questions shoppers ask: sofas for small apartments, pet-friendly fabrics, how to measure for a sectional. Research on [the product content AI shopping assistants prefer](https://underneath.agency/resources/product-content-ai-shopping-assistants-prefer) shows why dimensions and materials belong in plain text.
3. **Independent coverage.** Design editors, home publications and testing sites that review your pieces. “Best sofa” lists are cited often, and our guide to [best-of lists in AI answers](https://underneath.agency/resources/best-of-lists-ai-recommendations) shows which ones to target.
4. **Reviews where assistants look.** Reviews on your site, on retailer listings and on independent platforms, with responses to complaints about delivery damage and quality.
5. **Consistent brand facts.** Warranty, showroom locations, financing and return terms that match across your site, retailers and listings.
6. **Measurement over many runs.** A tracked set of furniture prompts by room, budget and style, asked repeatedly on several assistants, because single answers change.

Smaller furniture brands start behind the national names shoppers already know, and [how small brands get recommended by AI](https://underneath.agency/resources/how-small-brands-get-recommended-by-ai) sets out how they close that gap.

## What can’t furniture sellers measure about AI yet?

How many furniture sales start with an AI answer, and whether that share is growing faster than elsewhere.

The furniture survey data comes from trade research and vendors, and the AI shopping survey is a single online panel. Adobe’s conversion figures cover all US retail, and its AI category includes many tools. Wayfair, the most visible furniture seller in AI shopping, says the volume is small. Platform documentation tells us what assistants consider, not how much weight each factor carries. No public study we found has measured which furniture brands AI assistants name for room and budget questions. Proving that GEO work caused sales also needs a comparison group, which we explain in [how to prove GEO caused sales](https://underneath.agency/resources/prove-geo-caused-sales).

## Where should a furniture brand start?

With an audit of how AI assistants answer the room, budget and comparison questions your best customers ask.

Start with your highest-margin categories, usually upholstery and bedroom, and the questions that lead to them. Check whether your brand and hero products are named, which retailers and review sites are cited instead, and whether assistants can read your dimensions, fabrics, delivery and return terms. Then fix the product facts and coverage that decide the short list. If you want a partner for that, [tell us which collections matter most](https://underneath.agency/contact) and we will map where your brand stands in AI answers for the questions that lead to your largest orders and showroom visits. Our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page goes through the ongoing work, such as putting dimensions and materials into feeds, room-by-room content and review upkeep across your retailers.

## Frequently asked questions

### Do furniture shoppers really use ChatGPT to buy a sofa?

Some do, fewer than in clothing or electronics. In one 2026 survey, roughly 29% of AI shoppers had used AI for furniture. Most still visit a website or showroom before buying.

### Should we enable checkout inside ChatGPT or Google AI Mode?

It depends on your platform and margins. Wayfair joined Google’s Universal Commerce Protocol, but its CEO expects agents to matter most for low-cost, commodity items. Being named in the answer matters for more of your sales than checkout does.

### Which product details matter most for AI answers about furniture?

The ones shoppers filter on: size, materials, fabric performance, comfort, warranty, delivery and returns. OpenAI’s product feed has fields for dimensions, material and weight.

### Can we see AI-driven furniture sales in analytics?

Partly. Visits from assistants show as referrals, but many furniture shoppers finish in a store or on another device, so AI influence is undercounted.

## Sources

- Wayfair (2026-08-04), [Wayfair Announces Second Quarter 2026 Results](https://s24.q4cdn.com/589059658/files/doc_financials/2026/q2/2026-08-04-Press-Release.pdf)
- U.S. Census Bureau (2026-09-16), [Advance Monthly Sales for Retail and Food Services, August 2026](https://www.census.gov/retail/marts/www/marts_current.pdf)
- Home News Now (2025-09-19), [CIN Week 2 looks at what’s driving consumer furniture purchases](https://homenewsnow.com/blog/2025/09/19/cin-week-2-looks-at-whats-driving-consumer-furniture-purchases/)
- Home News Now (2025-10-03), [Consumer Insights Now Week 3 looks at what’s driving planned furniture purchases](https://homenewsnow.com/blog/2025/10/03/consumer-insights-now-week-3-looks-at-whats-driving-planned-2nd-half-furniture-purchases/)
- 3D Cloud (2025-03-07), [3D Cloud 2025 furniture study](https://3dcloud.com/news/3d-cloud-2025-furniture-study-release/)
- TechWyse (2026), [77% of US Consumers Now Use AI to Shop](https://www.techwyse.com/news/ai-search/ai-shopping-adoption-consumer-distrust-autonomous-spending-2026)
- Digital Commerce 360 (2026-01-14), [Wayfair joins Google’s Universal Commerce Protocol](https://www.digitalcommerce360.com/2026/01/14/wayfair-google-universal-commerce-protocol/)
- Digital Commerce 360 (2026-05-05), [Wayfair wants “to be everywhere” when it comes to agentic AI](https://www.digitalcommerce360.com/2026/05/05/wayfair-agentic-ai-commerce/)
- TechCrunch (2026-04-16), [AI traffic to US retailers rose 393% in Q1](https://techcrunch.com/2026/04/16/ai-traffic-to-us-retailers-rose-393-in-q1-and-its-boosting-their-revenue-too/)
- OpenAI (2026), [Shopping with ChatGPT Search](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search)
- OpenAI (2026), [Product Feed Spec](https://developers.openai.com/commerce/specs/feed)
- Google (2025-05-20), [New ways to shop with AI Mode](https://blog.google/products/shopping/google-shopping-ai-mode-virtual-try-on-update/)
- Uberti-Bona Marin, Bertaglia et al. (2026-09), ["If I Had to Buy Just ONE: Galaxy S26 Ultra": Auditing AI-Generated Product Recommendations](https://arxiv.org/abs/2609.18729)
- Underneath (2026), [Brand entity and AI recommendations study](https://underneath.agency/research/brand-entity-ai-recommendations-study)

---

This is the Markdown twin of https://underneath.agency/resources/furniture-brands-sales-from-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "GEO Agency vs SEO Agency: Which Do You Need? | Underneath"
description: "An SEO agency ranks pages in search results; a GEO agency gets a brand cited and recommended in AI answers. What each does, what they cost, and which you need."
canonical: "https://underneath.agency/resources/geo-agency-vs-seo-agency"
published: 2026-09-25
updated: 2026-09-28
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# GEO agency vs SEO agency

An SEO agency improves how your pages rank in Google and Bing; a GEO agency, a generative engine optimization agency, gets your brand mentioned, cited and recommended inside the answers ChatGPT, Google AI Overviews, Gemini, Perplexity and Copilot give. They share a foundation but differ in what they optimize, where the results show up and how they are measured. This guide compares the two, including cost, and helps you decide whether you need one, the other or both.

Underneath teamGuide

On this page

1. [The difference in one table](#the-difference-in-one-table)
2. [What an SEO agency does](#what-an-seo-agency-does)
3. [What a GEO agency does](#what-a-geo-agency-does)
4. [Where the two overlap](#where-the-two-overlap)
5. [Where they differ, measured](#where-they-differ-measured)
6. [Do you need an SEO agency, a GEO agency or both](#do-you-need-an-seo-agency-a-geo-agency-or-both)
7. [Frequently asked questions](#frequently-asked-questions)

Related service

Generative Engine Optimization

Get mentioned, cited and recommended by ChatGPT, Gemini, Perplexity and Copilot, and described correctly when you are.

[How we do it](https://underneath.agency/services/generative-engine-optimization)

## The short version

1. An SEO agency earns positions on a results page; a GEO agency earns a place inside an AI-generated answer. The first is measured in rankings and clicks, the second in mentions, citations and share of voice.
2. On Google’s AI surfaces the two jobs overlap heavily, because pages that rank near the top are cited about twice as often as pages lower down. On Perplexity and ChatGPT they are different jobs, because those engines cite mostly from outside the top 10.
3. Most companies need both, measured together: SEO as the foundation, GEO on top of it to turn ranking pages into cited ones, plus the third-party work the AI engines depend on.

Ask any of the engines we tested “GEO agency vs SEO agency” and you get a definition of each, a side-by-side table and a section titled “do you need both”. This guide follows the same shape, then adds what measuring the engines shows about the difference.

## The difference in one table

| | SEO agency | GEO agency |
| --- | --- | --- |
| Goal | Rank pages and earn clicks | Be the brand the AI answer names, cites and recommends |
| Where you win | Google and Bing results pages, map results | ChatGPT, Google AI Overviews and AI Mode, Gemini, Perplexity, Copilot |
| Unit judged | The page | The passage, the brand’s entity, and the third-party pages that mention it |
| On-page work | Keywords, titles, content depth, internal links, speed | Definition-first openings, question headings, complete FAQ answers, comparison tables, structured data for AI engines |
| Off-page work | Backlinks | Mentions and listings on the roundups, directories, forums and publications the engines cite |
| Technical work | Crawlability, indexation, Core Web Vitals | The same, plus AI crawler access, server-rendered text, schema completeness, one canonical URL |
| Measurement | Rankings, organic traffic, conversions | Mentions, citations, share of voice by engine, AI referral traffic and leads |
| Time to results | Months | Weeks on Google’s AI surfaces for indexed pages; months for third-party citations and ChatGPT |
| Typical retainer | Agencies average about $3,209 a month (Ahrefs survey, 2024) | About $1,500 to $50,000 a month (2026 price guides) |

## What an SEO agency does

A search engine optimization agency improves how a site ranks and converts in traditional results: keyword and intent research, technical health (crawling, rendering, indexation, speed), information architecture, content that targets queries, and links from other sites. Its output is positions, and the traffic and conversions that follow. In 2026 a good SEO agency also reports on AI Overviews, because the Overview sits above the results it ranks for, but its levers are still the page and the link.

## What a GEO agency does

A generative engine optimization agency works on the inputs an AI answer engine uses when it composes an answer: pages that can be quoted passage by passage, structured data the engine can read, a brand entity it can recognize, and the third-party pages it already trusts. Its nine activities, from audit to training, are set out in [What does a GEO agency do?](https://underneath.agency/resources/what-does-a-geo-agency-do). Its output is a place in the answer: the brand named, the page cited, the description accurate. Its measurement is per engine, because the engines do not share sources.

## Where the two overlap

- Technical foundations. Both need a crawlable, fast, well-structured site. GEO adds rules for the AI crawlers and more thorough structured data.
- Content quality. Google’s guidance for its generative AI features says the usual SEO fundamentals apply and no special optimizations are required. Helpful, specific, well-structured content serves both.
- Entity clarity. Consistent facts about the brand across the site, profiles and directories help ranking systems and answer engines alike.
- Authority. Links serve SEO; mentions on cited pages serve GEO; the same publications and directories often provide both.

## Where they differ, measured

We ran the query “GEO agency vs SEO agency” 22 times on the five engines between 15 and 25 September 2026, and a related service keyword daily for ten days. The numbers show where the two jobs part ways.

- On Google, ranking helps a great deal. For the service keyword, the AI Overview cited 75% of the pages ranked 1 to 3 in the same search and 42% of those ranked 4 to 10. For this comparison query, about half of the Overview’s citations came from the top 10. A single keyword overstates the overlap, though: across 481 AI Overviews, [only 28.7% of citations ranked in the top 10](https://underneath.agency/research/ai-overview-citations-study), and on searches that showed both, [19.9% of Google AI Mode’s citations were top-10 pages](https://underneath.agency/research/ai-mode-vs-ai-overviews-study). On Google, ranking gets you much of the way to a citation, but not all of it.
- Perplexity and ChatGPT do not. ChatGPT took 26 of 28 citations for the service keyword from outside Google’s top 10, and cited academic research and official documentation before any vendor. Perplexity cited pages with a median length around 3,700 words against 1,500 for pages it retrieved and skipped, regardless of rank. Here an SEO agency’s levers do not reach.
- The source sets are small and stable. For this query, one agency explainer was cited in 18 of 18 AI Overviews, 22 of 22 Perplexity answers and 20 of 22 AI Mode answers, and named inside 49 of 110 answers, in the daily runs and again in the single-afternoon batch. Reddit was named in 42. Entering a set like that is GEO work: matching the page’s shape, and being mentioned where it is mentioned.
- The engines want the expanded term. ChatGPT’s own searches for this query began “geo generative engine optimization agency vs seo agency definition” and added “2026”. A page that says “GEO” without the expansion, or carries no year, is missing the query the engine runs.
- Buyers still ask about SEO. The People Also Ask box under the result asked “is SEO still worth it in 2026?” (15 times), “has GEO replaced SEO?” (13), “is SEO dead now with AI?” (10) and “are SEO and GEO the same?” (10). The questions at the end of this guide answer them.

## Do you need an SEO agency, a GEO agency or both

- SEO first if you do not yet rank for the questions your buyers ask. On Google’s AI surfaces there is no shortcut past ranking: pages in the top three are cited about twice as often as pages at positions 7 to 10.
- GEO now, alongside SEO, if your pages rank but the AI answers name competitors, if your buyers use ChatGPT or Perplexity, or if the engines describe your brand inaccurately. Those are problems ranking does not fix.
- Both, coordinated, for most companies. One measurement framework (rankings alongside citations, mentions and share of voice), one technical foundation, one content standard that serves both, and one owner for the off-site sources, whether that is one partner or an SEO partner plus a GEO specialist.
- Two agencies only if your SEO partner cannot show per-engine measurement or a third-party citation plan. Ask before you add a second retainer; some SEO agencies can add the GEO layer, and the [criteria for choosing a GEO agency](https://underneath.agency/resources/how-to-choose-a-geo-agency) apply to either.

Underneath does GEO only. Every program works on the passages, the entity and the third-party sources that decide whether the AI answer names you, and runs alongside your SEO team or agency, which keeps the ranking side.

## Frequently asked questions

### What is the difference between a GEO agency and an SEO agency?

An SEO agency ranks pages in Google and Bing results and is measured in rankings and traffic. A GEO agency gets a brand cited, mentioned and recommended in AI-generated answers from ChatGPT, Google AI Overviews, Gemini and Perplexity, and is measured in mentions, citations and share of voice by engine.

### Are SEO and GEO the same?

No, but they share a foundation. Both need a crawlable, well-structured, authoritative site; GEO adds quotable passages, structured data for AI engines, a consistent brand entity and mentions on the third-party pages the engines cite, and it is measured per engine.

### Has GEO replaced SEO?

No. Ranking is still the main route into Google’s AI surfaces. GEO works on whether the ranking page is also the one the answer quotes, and covers Perplexity and ChatGPT, where ranking matters far less.

### Is SEO still worth it in 2026?

Yes. In our study of 486 searches, [41.7% of the pages ranking in Google’s top three were cited by the AI Overview](https://underneath.agency/research/ai-overview-cited-pages-study), against 20.1% at positions 7 to 10. Ranking earns the position and raises the chance of the citation.

### Can one agency do both SEO and GEO?

Some can. Whether you use one agency or two, check that whoever owns GEO measures AI visibility per engine, has a technical checklist for AI crawlers, and has a plan for the third-party pages the engines cite.

### How much do GEO and SEO agencies cost?

In [Ahrefs’ survey of 439 SEO providers](https://ahrefs.com/blog/seo-pricing/), agencies charged an average of $3,209 a month (2024). For GEO, published 2026 price guides range from about $1,500 to $50,000 a month, with mid-sized companies typically at $5,000 to $25,000 ([WebFX](https://www.webfx.com/blog/ai/generative-engine-optimization-cost/)). Underneath’s GEO programs are priced in three bands, $5,000 to $10,000, $11,000 to $20,000 and $20,000+ a month.

## Keep *reading.*

[All resources](https://underneath.agency/resources)

- Guide · AI search

  ### [What does a GEO agency do?](https://underneath.agency/resources/what-does-a-geo-agency-do)

  The nine things a generative engine optimization agency does, what it costs, and how the AI engines describe the job.

  Underneath team
- Guide · AI search

  ### [How to choose a GEO agency](https://underneath.agency/resources/how-to-choose-a-geo-agency)

  Eight criteria, the red flags, and the questions to ask on the first call, checked against what the AI engines themselves tell buyers.

  Underneath team
- Guide · AI search

  ### [Is a GEO agency worth it?](https://underneath.agency/resources/is-a-geo-agency-worth-it)

  When hiring one pays back, when it does not, and how to decide with numbers rather than fear of missing out.

  Underneath team

Free strategy call

## Ranking but not cited?

On a free 30-minute call we show you where you rank, where you are cited, and where the gap between the two is costing you.

[Book a strategy call](https://underneath.agency/contact)[How we work](https://underneath.agency/approach)

---

This is the Markdown twin of https://underneath.agency/resources/geo-agency-vs-seo-agency. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How should global brands approach GEO across languages?"
description: "Plan market by market: AI answers change with language and country, so earn local coverage, write native content and measure each market separately."
canonical: "https://underneath.agency/resources/geo-strategy-for-global-brands-across-languages"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How should global brands approach GEO across languages?

Treat each language market as its own AI search market, with its own sources, its own competitors and sometimes its own version of Google. Research from 2025 and 2026 shows that AI assistants draw on different websites and name different brands when the language changes. The AI layer itself is also switched on unevenly from country to country. The workable model is one global playbook fed by local evidence, local coverage and local measurement.

## The short version

1. Google’s AI Overviews reached 229 countries in 2025, up from 7 in 2024, yet France, Turkey, Iran, China and Cuba were still left out ([Aral, Li and Zuo](https://arxiv.org/abs/2602.13415)).
2. In a twelve-language European study, asking in a local brand’s home language raised how often AI named it by 0.80 on a 0-to-1 scale, against 0.15 for global brands ([Żatuchin](https://arxiv.org/abs/2606.23165)).
3. In Tokyo, Gemini’s Japanese answers cited hotels’ own websites more often than its English answers; the authors estimate 2.27 times more such citations and link the gap to deeper Japanese-language content ([Zhu and Chang](https://arxiv.org/abs/2603.20062)).
4. In our four-country test, ChatGPT’s brand picks moved between countries far more for insurance, lending and tax questions (a gap of 0.180) than for global products (0.034) ([our country study](https://underneath.agency/research/ai-recommendations-by-country-study)).
5. In a lab test, fake brand pages fooled AI assistants about as often in English as in Chinese: 87% of the time for restaurant questions in English, against 82% in the matched Chinese category ([Luo and Chen](https://arxiv.org/abs/2606.13610)).

## Should global companies build a separate GEO strategy for each language?

Yes for evidence and measurement, no for principles: the playbook can be global, the inputs must be local. What differs by language is which sources, competitors and wording matter, not what earns trust.

The largest multilingual study so far, by [Żatuchin](https://arxiv.org/abs/2606.23165), collected 35,640 answers from three AI assistants that search the web, about 66 European brands in twelve languages, in April and May 2026. Switching from English to a brand’s home language raised a local champion’s recommendation share by 0.80 on a 0-to-1 scale, against 0.15 for a global brand. So how much local work you need depends on your position in each market. A multinational is named at similar rates in either language. A brand that leads at home lives or dies in its home language. The author is affiliated with an AI brand-monitoring company and discloses it. Our article on [English-only AI visibility checks](https://underneath.agency/resources/english-only-ai-visibility-audits) covers what that means for audits.

The principle that travels is earned credibility. In a University of Toronto study by [Chen and colleagues](https://arxiv.org/abs/2509.08919), a web-enabled ChatGPT model leaned on earned media: independent reviews and publishers. It drew 77.6 percent of its sources for Canadian consumer electronics questions from earned media, and 92.1 percent for the same questions about the US. The rule held in both markets; the publications behind it were local.

## Why does the same content perform differently in different languages?

Because each language has its own web, and AI assistants mostly search the web in the question’s language. Your page competes against a different set of sources in every language.

[Zhu and Chang](https://arxiv.org/abs/2603.20062) describe Tokyo’s hotel market as two largely separate webs: Booking.com and Expedia dominate English results, while Jalan, Rakuten Travel and Ikyu dominate Japanese ones. Gemini’s answers followed the language of the question. For questions about the guest experience, 62.1% of Japanese citations came from sources other than booking sites, against 50.0% in English, because the Japanese web offered more of those sources. The Toronto team found the same pattern from another angle: changing the language moved cited sources more than rewording a question within one language.

The answers themselves differ too. In the European study, the median answer ran 289 words in English but 189 in Finnish. Finnish answers resembled answers in Germanic languages more than those in Estonian, its closest linguistic relative. The author’s plausible explanation is that Finnish business coverage leans heavily on English and Swedish sources. In other words, “language” really means the information ecosystem behind it. For how each engine localizes its sources, see [our guide to local-language citations](https://underneath.agency/resources/do-ai-engines-cite-local-language-sources).

## Does automatically translating your content keep your AI search visibility?

No study we reviewed tests this directly, and the indirect evidence says translation alone is not enough. Treat any claim otherwise as unproven.

Three findings point the same way. First, the Toronto authors write about engines that switch to local sources. Their conclusion: “simply translating your own brand content is not enough; you need earned coverage in the local language.” Second, in Tokyo, content depth by language mattered. The authors report that Japanese queries produce 2.27 times more hotel-direct citations than English ones. They link this to Japanese hotel sites carrying neighborhood guides and transit details, while English versions focused on booking. A translated booking page adds no such depth.

Third, a [survey of AI visibility measurement](https://arxiv.org/abs/2609.06811) warns that a translated question may not preserve local availability, vocabulary or regulation. The same logic applies to pages: a literal translation can miss the words local buyers use. The Tokyo team wrote its Japanese questions as natural-sounding queries rather than literal translations for exactly this reason. This is our inference from the research, not a tested result.

## Is AI search even the same product in every market?

No: whether AI answers appear at all, and how often, depends heavily on the country. Your exposure is uneven before any content work begins.

[Aral, Li and Zuo](https://arxiv.org/abs/2602.13415) ran the same searches on Google in 2024 and 2025 across 243 countries. AI Overviews, the AI summary at the top of Google’s results, went from 7 countries to 229. In most countries that had them in 2025, they appeared on 55% to 70% of searches, but in Iceland only 7.4%. France, Turkey, Iran, China and Cuba were excluded altogether, even as 222 new countries were added. Among early markets, Japan saw a 78% rise in AI answers in a year. The country explained more of whether a search showed AI than its topic or wording.

The timing gap has real costs. [Khosravi and Yoganarasimhan](https://arxiv.org/abs/2602.18455) used Google’s staggered rollout, US default in May 2024 and the EU only in early 2025, as a natural test. Monthly search traffic to English Wikipedia articles fell by approximately 5.45% relative to the same articles in German. The same page can feel AI’s effect a year earlier in one language than in another.

## Which markets and categories need the most local attention?

Markets where you are a local champion, and categories where products, prices or rules are national. Global consumer products move least.

In [our country study](https://underneath.agency/research/ai-recommendations-by-country-study), we asked ChatGPT and Gemini the same 40 English buyer questions five times each from the US, UK, Canada and Australia in September 2026. ChatGPT’s brand picks moved between countries far more for insurance, lending and tax questions (a gap of 0.180 in brand overlap) than for global products such as headphones and CRM software (0.034). For small business accounting software, it led with QuickBooks in the US and Canada and with Xero in the UK and Australia.

The assistant also adapted the answer itself. ChatGPT mentioned the country, its currency or its regulators in 68.8% of UK answers, though the question named no country. If your category depends on local rules or prices, the AI answer in each market is effectively a different answer. China adds a further split, which our guide on [brand visibility in Chinese versus Western AI models](https://underneath.agency/resources/chinese-vs-western-ai-brand-visibility) covers.

## Do the risks to your brand cross languages too?

Yes: in one lab test, fake pages misled AI assistants about as often in English as in Chinese. Brand protection has to run in every language.

[Luo and Chen](https://arxiv.org/abs/2606.13610) swapped real products for invented ones in the top three web pages an AI assistant read, then counted how often twelve AI models recommended the fake. They ran it mainly in Chinese, then repeated it in English with 360 trials. In English, the fake was recommended 87% of the time for San Francisco restaurants and 43% for smartphones; the matched Chinese figures were 82% and 23%. Categories where the AI knew less about real brands were the most exposed.

This was a controlled setup, not the live web. Still, a market where your brand has thin local coverage is the kind of market where false pages have the least to compete with.

## What should you do about it?

Run one global GEO program with a local evidence base and local measurement in every priority market. In practice:

1. **Classify each market.** Decide whether you are a local champion or a global brand there; the European study shows that this sets how much local work you need.
2. **Check the AI layer first.** Confirm whether and how often AI answers appear in each country before setting targets.
3. **Map the local sources.** List the publishers, review sites and directories that AI answers cite in each language, and earn coverage there, not only in English trade press.
4. **Write natively, not literally.** Give local pages real local substance: prices, rules, places, the words buyers use. Have a reviewer in each market check them.
5. **Measure each market separately.** Run your buyer questions in each language and location, more than once, and report them separately rather than as one global number.
6. **Watch for impersonation in every language.** Fake pages worked in both languages tested, so monitor what AI says about you in each market.

If you want help planning a multi-market program, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

Most of the strategy questions are untested: no study measures whether translated or localized content changes AI visibility. The main gaps:

- No controlled test compares machine-translated pages with natively written ones for AI citations.
- Brand-level evidence covers twelve European languages, five major world languages, Japanese hotels and four English-speaking countries. Most other markets have only data on whether Google shows AI answers.
- The largest multilingual study was written by an author affiliated with a brand-monitoring company and covers one window in April and May 2026.
- The link between local sources and local brands is an association; no study has changed a market’s sources to test cause.
- AI products and rollouts change quickly; country-level availability in 2025 may already differ today.

## Frequently asked questions

### Should we translate our website for AI search in other markets?

Translate, but do not stop there. No study shows translation alone keeps AI visibility, and the research points to local earned coverage and locally useful content as what AI assistants cite.

### Do AI assistants recommend different brands in different languages?

Yes. In 35,640 answers about 66 European brands, asking in a local champion’s home language raised how often it was named by 0.80 on a 0-to-1 scale, against 0.15 for global brands.

### Do Google AI Overviews appear in every country?

No. A study of 243 countries found them in 229 in 2025, with France, Turkey, Iran, China and Cuba excluded. They appeared on 55% to 70% of searches in most countries, but 7.4% in Iceland.

### Is our English content enough for AI visibility in Europe?

For a global brand it may carry some weight; for a local leader it is not enough. The European study found local champions’ visibility depends mainly on answers in their home language.

## Sources

- Aral, Li and Zuo (2026), [The Rise of AI Search: Implications for Information Markets and Human Judgement at Scale](https://arxiv.org/abs/2602.13415), arXiv:2602.13415.
- Żatuchin (2026), [The Language Blind Spot: How Query Language and Brand Recognition Tier Shape AI-Constructed Brand Reputation Across Twelve European Languages](https://arxiv.org/abs/2606.23165), arXiv:2606.23165.
- Chen, Wang, Chen and Koudas (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919), arXiv:2509.08919.
- Zhu and Chang (2026), [The End of Rented Discovery: How AI Search Redistributes Power Between Hotels and Intermediaries](https://arxiv.org/abs/2603.20062), arXiv:2603.20062.
- Khosravi and Yoganarasimhan (2026), [Impact of AI Search Summaries on Website Traffic: Evidence from Google AI Overviews and Wikipedia](https://arxiv.org/abs/2602.18455), arXiv:2602.18455.
- Luo and Chen (2026), [One Polluted Page Is Enough: Evaluating Web Content Pollution in LLM Recommenders](https://arxiv.org/abs/2606.13610), arXiv:2606.13610.
- Martinez (2026), [Measuring GEO Visibility: Prompt Corpora Define the Answer Market](https://arxiv.org/abs/2609.06811), arXiv:2609.06811.
- Underneath (2026), [Same question, four countries: do AI recommendations change?](https://underneath.agency/research/ai-recommendations-by-country-study)

---

This is the Markdown twin of https://underneath.agency/resources/geo-strategy-for-global-brands-across-languages. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "When Google changes its results, how much control do we have?"
description: "Very little: Google can shift traffic with design changes publishers cannot control. AI Overviews cut English Wikipedia’s search visits by about 5%."
canonical: "https://underneath.agency/resources/google-search-design-changes-traffic-risk"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How much control do we have when Google changes how search results look?

Very little over the change itself, and some over how exposed you are to it. When Google made AI Overviews the default in the US, search visits to English Wikipedia fell about 5% relative to the same articles in other languages, with nothing changing on Wikipedia’s side. Your control lies in how much of your business depends on one layout, and how quickly you notice when it moves.

## The short version

1. Google’s switch to AI Overviews by default cut English Wikipedia’s monthly search traffic by 5.45% relative to German and 4.82% relative to French, about 100.27 million visits a month ([Khosravi and Yoganarasimhan](https://arxiv.org/abs/2602.18455)).
2. Google’s AI layer shifts week to week: re-searching 96 keywords two days later, 16 changed whether they showed an AI Overview at all ([our AI Overview frequency study](https://underneath.agency/research/ai-overviews-frequency-study)).
3. Google’s own rankings move too: measured against results from two days later, ChatGPT’s match with Google’s top 10 fell from 8.3% to 5.2% ([our AI citations and Google rankings study](https://underneath.agency/research/ai-citations-google-rankings-study)).
4. Each new Google surface rewards pages differently: AI Overviews cited 29.7% of top-10 pages and AI Mode 16.3% on the same searches ([our AI Mode vs AI Overviews study](https://underneath.agency/research/ai-mode-vs-ai-overviews-study)).

## Can a business control how Google designs its results?

No: Google sets the layout, and publishers can only respond to it. [Khosravi and Yoganarasimhan](https://arxiv.org/abs/2602.18455), researchers at the University of Washington, put it plainly: search platforms “can reallocate traffic through product-design changes that publishers cannot directly control.”

That is a policy argument rather than a measured result, but the stakes behind it are measured. They cite Google’s share of the global search market at approximately 90%. When one company controls the front door for that much demand, a design decision on its side becomes a traffic event on yours.

The timeline shows how fast this moves. Google made AI Overviews, the AI summary at the top of its results, part of the default US experience in May 2024. By early 2025 the feature had reached more than 100 countries.

## How much traffic can one design change move?

Enough to matter even for a site that AI summaries cite often. The researchers compared English Wikipedia articles with the same articles in German and French. Readers of those editions were mostly in Europe, where AI Overviews were not yet the default.

After the US rollout, monthly search traffic to English articles fell 5.45% relative to German and 4.82% relative to French. The authors estimate that as about 100.27 million fewer search visits to English Wikipedia per month. A third comparison, with Japanese, pointed the same way at 16.53%, over a shorter window.

The striking part is that Wikipedia is among the sources AI summaries cite most. Being cited did not offset the loss. And these figures likely understate the effect for searches that actually showed a summary, because only about 40% of English Wikipedia traffic comes from the US. Our guide on [AI Overview citations and lost clicks](https://underneath.agency/resources/ai-overview-citations-vs-lost-clicks) weighs the wider evidence.

## How often does Google’s AI layer change what it shows?

Often enough that a single check tells you little. In [our study of US searches](https://underneath.agency/research/ai-overviews-frequency-study), we re-ran 96 keywords two days apart. 83.3% had the same outcome, but 16 flipped between showing an AI Overview and not.

The sources shifted more than the outcome. Where both days showed an AI Overview, fewer than half the cited pages were the same: the typical overlap was 0.43 on a scale where 1 means identical.

Google’s regular rankings moved as well. Between 26 and 28 September, Google’s top 10 for the same 80 questions changed substantially ([our AI citations and Google rankings study](https://underneath.agency/research/ai-citations-google-rankings-study)). Measured against the later results, ChatGPT’s match with Google’s top 10 fell from 8.3% to 5.2%.

## Does each new Google surface reward the same pages?

No, which means each new layout can reshuffle who gets seen. In [our comparison of AI Mode and AI Overviews on the same 400 searches](https://underneath.agency/research/ai-mode-vs-ai-overviews-study), AI Overviews cited 29.7% of the top-10 pages and AI Mode 16.3%.

AI Mode also leans on Google’s own pages in some searches. On searches with no AI Overview, 69.3% of AI Mode’s citations went to Google pages, mostly listing pages for local businesses and products. That is a design choice on Google’s side, and it leaves less room for outside sites in those answers. We cover that scenario in [our article on AI Mode as Google’s default](https://underneath.agency/resources/ai-mode-default-traffic-loss).

AI answers are also sensitive to small changes on the user’s side. In a small test by [Wen and colleagues](https://arxiv.org/abs/2606.12439), rewording questions by only 13% on average changed which domains Gemini cited for every one of 30 question pairs.

## What part of the risk can you actually manage?

You can manage your dependence on any one layout and how fast you react. Ranking still helps inside AI Overviews: [the first organic result was cited in 49.5% of AI Overviews](https://underneath.agency/research/ai-overview-citations-study), the ninth in 15.5%. So strong rankings remain a hedge, just not a guarantee.

Spreading demand across channels is the other lever. On one site studied by [Watanabe and Nakayashiki](https://arxiv.org/abs/2606.04362), Google clicks fell about 20% site-wide while ChatGPT referrals grew. Even pages nobody optimized grew 3.5 times in ChatGPT referrals over about five months. Our guide on [whether ChatGPT citations lift Google rankings](https://underneath.agency/resources/do-chatgpt-citations-boost-google-rankings) covers what that study means for Google traffic.

Position inside AI answers is partly outside your hands too. In a simulated hotel test, being placed higher in the list an AI assistant saw had the same pull as a price cut of about $11.7 per night ([Baig and colleagues](https://arxiv.org/abs/2606.16344)). The authors advise “continuous measurement rather than a one-time fix.”

## What should you do about it?

Treat Google’s layout as a risk you monitor and hedge, not a fixed condition. In practice:

1. Map your exposure. List which revenue pages depend on informational searches, the kind AI summaries answer directly.
2. Monitor weekly, not once. Track a fixed set of searches for whether an AI Overview appears and which pages it cites.
3. Compare against a baseline. Judge traffic drops against similar pages or sections, so you can tell a Google change from your own mistake.
4. Keep rankings strong. Top positions are still cited far more often inside AI Overviews.
5. Build other routes to demand: AI assistants, email, direct and partner channels, so one layout change cannot cut off most of it.
6. Plan for scenarios. Model what a further drop in search visits, at least the size Wikipedia saw, would do to your pipeline.

If you want a monitoring setup built around this, see [our generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The research measures past changes, not the next one, so forecasts remain guesses.

- The traffic evidence comes from Wikipedia, a non-commercial reference site. Commercial publishers may lose more or less.
- The Wikipedia study cannot see visits lost inside the site after a search visit disappears, so the total effect could be larger.
- Our frequency and ranking checks cover two dates in September 2026 and US searches only.
- The rewording and hotel findings come from small or simulated tests.
- No study yet measures how much diversifying channels actually protects revenue when Google changes its layout.

## Frequently asked questions

### Did Google AI Overviews reduce website traffic?

For English Wikipedia, yes. Its monthly search traffic fell about 5.45% relative to the German edition after AI Overviews became the US default.

### Can we opt out of Google’s AI Overviews without losing rankings?

The research reviewed here does not answer that. It shows only that being cited, as Wikipedia often is, did not prevent a traffic decline.

### How often do AI Overviews change?

Within days. In our check, 16 of 96 keywords changed whether they showed an AI Overview two days later, and the cited pages changed more.

### Should we still invest in SEO if Google keeps changing the layout?

Yes, as one hedge among several. The first organic result was cited in 49.5% of AI Overviews in our study, against 15.5% for the ninth.

## Sources

- Khosravi and Yoganarasimhan (2026), [Impact of AI Search Summaries on Website Traffic: Evidence from Google AI Overviews and Wikipedia](https://arxiv.org/abs/2602.18455), arXiv:2602.18455.
- Watanabe and Nakayashiki (2026), [Disentangling Answer Engine Optimization from Platform Growth: A Log-Based Natural Experiment on ChatGPT Referral Traffic](https://arxiv.org/abs/2606.04362), arXiv:2606.04362.
- Wen, Zhang, Yuan, Chen, Zhang and Guo (2026), [Position: Generative Engine Optimization Creates Underexamined Risks, Governance Must Target Concentration, Disclosure, and Academic Blind Spots](https://arxiv.org/abs/2606.12439), arXiv:2606.12439.
- Baig, Gillani and Ali (2026), [Whose hotel does the AI recommend? An algorithm audit of reputation signals in LLM-assisted hotel selection](https://arxiv.org/abs/2606.16344), arXiv:2606.16344.
- Underneath (2026), [When does Google show an AI Overview? 1,248 US searches](https://underneath.agency/research/ai-overviews-frequency-study)
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude cite pages that rank?](https://underneath.agency/research/ai-citations-google-rankings-study)
- Underneath (2026), [AI Mode vs AI Overviews: how different are the sources?](https://underneath.agency/research/ai-mode-vs-ai-overviews-study)
- Underneath (2026), [Do AI Overviews cite the pages that rank? 4,051 citations](https://underneath.agency/research/ai-overview-citations-study)

---

This is the Markdown twin of https://underneath.agency/resources/google-search-design-changes-traffic-risk. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How can health insurers win members when shoppers ask AI?"
description: "By making plan facts, networks and ratings accurate and public before enrollment opens, because shoppers now ask AI tools to explain and compare coverage."
canonical: "https://underneath.agency/resources/health-insurers-members-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How can a health insurer win new members when shoppers ask AI to explain their plan choices?

By making sure AI answers find accurate, current and consistent facts about your plans, networks and ratings before each enrollment window opens. Plan shoppers are confused, rarely compare, and increasingly ask chatbots about health and coverage. No study yet links AI answers to enrollment numbers, so treat this as a strong bet, not a measured channel. Nothing here is advice on choosing a plan.

## The short version

1. The individual market is large and partly new each year: 23.0 million people selected 2026 Marketplace plans, including 3.4 million new consumers, according to [CMS](https://www.cms.gov/newsroom/fact-sheets/marketplace-2026-open-enrollment-period-report-national-snapshot-2).
2. Medicare buyers face many options and mostly do not shop: the average beneficiary could choose among 42 plans in 2025, [KFF](https://www.kff.org/medicare/issue-brief/medicare-advantage-2025-spotlight-a-first-look-at-plan-offerings/) reports, and in an earlier KFF analysis 69% did not compare their coverage with other options during open enrollment.
3. Shoppers want help: in an [eHealth survey](https://s204.q4cdn.com/837903328/files/doc_news/Survey-75-of-Medicare-Beneficiaries-Say-Selecting-a-Plan-Is-Confusing-2025.pdf), 75% of Medicare beneficiaries called choosing a plan confusing, and an EBRI survey reported by [Health Populi](https://www.healthpopuli.com/2026/03/24/consumer-adoption-of-ai-for-health-and-self-care-doubling-to-36-in-a-year-via-rock-healths-latest-snapshot) found 42% of privately insured adults want AI tools to help them choose a health plan but do not know where to start.
4. OpenAI documents the use case: [ChatGPT Health](https://openai.com/index/introducing-chatgpt-health/) is described as helping people “understand the tradeoffs of different insurance options,” and OpenAI says over 230 million people a week ask ChatGPT health and wellness questions.
5. Google’s AI answers are hard to avoid in this category: in [our study of 1,248 US searches](https://underneath.agency/research/ai-overviews-frequency-study), financial services and insurance keywords triggered an AI Overview 88.0% of the time.

## Who buys health insurance, and what is a new member worth to an insurer?

Three different buyers: individuals, Medicare beneficiaries and employers. Each won member tends to stay for years.

**Individual market.** CMS counts 23.0 million 2026 plan selections on the Marketplaces. Of those, 19.6 million had 2025 coverage and selected a plan or were automatically re-enrolled, so the fight for new members centers on the 3.4 million newcomers and on returning members who actively shop.

**Medicare.** Medicare Advantage enrolled 34.1 million of about 62.8 million eligible beneficiaries in 2025, or 54%, according to KFF data [summarized by TechTarget](https://www.techtarget.com/healthcarepayers/news/366628088/Top-stats-on-Medicare-Advantage-enrollment-costs-in-2025). The market is concentrated: UnitedHealth held 29% of enrollees and Humana 17%. Growth slowed to 1.3 million beneficiaries, or 4%, in 2025, so plans increasingly win members from each other rather than from new eligibles.

**Employers.** The [KFF 2025 Employer Health Benefits Survey](https://www.kff.org/health-costs/2025-employer-health-benefits-survey/) puts the average annual premium at $9,325 for single coverage and $26,993 for family coverage. Employers usually choose through brokers and consultants, and their reasons for switching include their workers’ experience: in the [J.D. Power 2025 commercial member study](https://chaindrugreview.com/j-d-power-gap-widens-between-highest-and-lowest-performing-employer-sponsored-health-plans), 20% of employers cited low employee satisfaction as a top reason for switching health plans. Health startups selling to employers and plans face a related test, covered in [how healthcare startups compete with established brands](https://underneath.agency/resources/healthcare-startups-demand-ai-search).

What makes a member valuable is how rarely people move. KFF found that 82% of Medicare Advantage drug plan enrollees did not compare their plan’s drug coverage with other plans in their area. The flip side: a shopper who leaves you during an enrollment window may be gone for many years. Carriers selling auto, home and other lines can see [how insurance carriers win quotes from AI](https://underneath.agency/resources/insurance-companies-customers-ai-search).

## Are plan shoppers already asking AI assistants about coverage?

Yes for health questions in general; for plan choice specifically, the evidence shows interest more than measured use.

The general shift is well documented. A [KFF poll](https://www.kff.org/health-information-and-trust/poll-finding/kff-tracking-poll-on-health-information-and-trust-use-of-ai-for-health-information-and-advice/) found 32% of adults turned to AI tools for health information and advice in the past year. Rock Health’s consumer survey, as reported by Health Populi, found use of AI chatbots for health information doubled from 16% to 32% between 2024 and 2025; ChatGPT was used by 23% of seekers and Gemini by 15%.

On insurance itself, three signals point the same way:

- **The platform documents it.** OpenAI lists insurance among ChatGPT Health’s uses, and one of its example prompts is “Based on my medical history, which of these insurance plans might be best for me?” OpenAI updated the page in July 2026 to say Health in ChatGPT is launching to US users 18 and older.
- **Shoppers say they want it.** EBRI’s 2025 survey of 2,001 privately insured Americans aged 21 to 64 is the source of the 42% who want AI help choosing a plan.
- **Brokers are building for it.** eHealth found 50% of Medicare beneficiaries would be interested in working with an AI agent by phone if it made shopping more efficient. eHealth sells plans and has an interest in that finding.

None of these surveys measures how many enrollees asked ChatGPT or Gemini to compare carriers before enrolling. That figure does not exist yet.

## What do plan shoppers ask AI assistants?

Mostly questions about plan types, networks, drugs, costs and ratings. We wrote these examples ourselves; none were collected from real shoppers.

| Shopping concern | Illustrative question |
|---|---|
| Plan type | “What is the difference between an HMO and a PPO on the Marketplace?” |
| Medicare choice | “What should I know when comparing Medicare Advantage and Medigap?” |
| Network | “Which 2026 plans in Maricopa County include the Banner Health network?” |
| Drugs | “How do Part D plans cover insulin in 2026?” |
| Ratings | “Which Medicare Advantage plans near Tampa have high star ratings?” |
| Extras | “Which plans in my area include dental and vision benefits?” |
| Employer | “What should a 50-person company compare when switching group health carriers?” |

Behind each question there may be several searches. Google says AI Overviews and AI Mode [may use a “query fan-out” technique](https://developers.google.com/search/docs/appearance/ai-features), issuing related searches across subtopics, and OpenAI says [ChatGPT search rewrites a question](https://help.openai.com/en/articles/9237897-chatgpt-search) into one or more targeted queries. A question about a county’s plans could become separate searches for networks, star ratings and drug coverage. When [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study) logged ChatGPT’s searches, it looked for reviews or ratings in 46.2% of answers.

A caution for insurers: AI answers on these topics are general information, not plan advice, and an assistant can get plan details wrong. That is why the accuracy of your public plan information matters so much.

## How does an AI answer turn into an enrolled member?

By shaping which carriers and plan types a shopper checks before the enrollment window closes.

As we understand the evidence, the path looks like this:

1. Before or during an enrollment window, a shopper or family member asks an assistant to explain options, check a doctor or drug, or compare carriers.
2. The answer names plan types, sometimes carriers, and cites sources such as government pages, insurer pages and comparison sites.
3. The shopper checks the insurer’s website, the official plan finder or a broker. In [Media Logic’s 2025 survey](https://www.medialogic.com/blog/healthcare-marketing/2025-medicare-aep-shopping-experience/) of 450 Medicare enrollees, insurer websites were the primary resource and brokers facilitated 58% of switches.
4. The shopper enrolls before the deadline; Medicare’s annual window runs October 15 to December 7, and HealthCare.gov’s 2026 window ran through January 15.
5. Retention follows. Because most people do not compare again, a member won this year is likely to renew.

The windows make timing unusual. A software vendor can build visibility at any time; an insurer needs its plan facts to be right in AI answers in the weeks before and during enrollment. Media Logic found 20% of shoppers switched plans, and 64% of switchers changed insurers entirely. Those are the members up for grabs.

Measurement is hard. A member who first heard your name in an AI answer may enroll through a broker or call center with no trace of that answer. We cover attribution in [how to prove GEO caused sales](https://underneath.agency/resources/prove-geo-caused-sales).

## What decides whether an AI assistant mentions your plans accurately?

The platforms keep selection rules private; studies favor authoritative, consistent sources, and stale pages breed wrong plan facts.

**Documented by the platforms.** Google and OpenAI confirm their AI answers search the web and link to sources. Neither explains how a carrier or plan is chosen.

**Observed in studies.**

- In ChatGPT’s answers to consumer health questions, [Jacques and colleagues](https://arxiv.org/abs/2601.17109) found 75.7% of 615 cited sources came from institutions such as medical centers, government agencies and Wikipedia. Commercial sites that were cited showed visible trust signals; our summary is in [what commercial health sites cited by ChatGPT have in common](https://underneath.agency/resources/chatgpt-cited-commercial-health-sites).
- [Our pricing study](https://underneath.agency/research/ai-pricing-accuracy-study) found 61.9% of plan prices quoted by four assistants for software products were fully faithful to the official pricing page. Of 64 differing prices, 39 appeared on another page of the vendor’s own site. That study covered software, not insurance, but the lesson transfers, we infer: last year’s plan pages and outdated benefit summaries are likely sources of wrong answers.

**What shoppers trust.** KFF found 77% of the public is concerned about the privacy of personal medical information given to AI tools. J.D. Power found members who understand their out-of-pocket and out-of-network coverage report higher satisfaction. Clear, accurate public explanations serve both the shopper and the assistant.

**Our inference for insurers.** The facts shoppers verify (networks, formularies, costs, star ratings, extras) are mostly public, but they are often spread across PDFs, county-specific pages and old plan years. A benefit change buried in a PDF the assistant never reaches will not make it into the answer. A carrier whose current facts are easy to find, consistently named and supported by third-party ratings gives the assistant less room for error. No study has yet tested this with health plans.

## What does it cost an insurer to be missing or misdescribed in AI answers?

Mostly lost consideration among the minority who actively shop, plus misinformation that sales teams must correct.

- **The shoppers who matter are few and decisive.** If 69% of Medicare beneficiaries do not compare, the members you can win sit in the remaining 31%. eHealth found 51% intended to review coverage for 2026, down from 63% who said they did so the year before.
- **Confusion is the norm.** In eHealth’s survey, 33% of beneficiaries said they did not have a good understanding of how Medicare Advantage, Medicare Supplement and Part D plans differ, and 33% wrongly believed Medicare covers GLP-1 drugs for weight loss. An AI answer that repeats a wrong benefit claim about your plan can cost a sale or create a complaint, we infer.
- **Brand gaps are visible.** J.D. Power found regional satisfaction scores ranging from 594 to 523 on a 1,000-point scale. Ratings like these are public and can be cited, for good or ill.
- **Answers shift.** The same question can return different carriers on different days; see [why AI answers about your brand change](https://underneath.agency/resources/why-ai-answers-about-your-brand-change).

## What GEO work fits a regulated health insurer?

Accurate, current, consistent public plan information, backed by independent ratings, and reviewed by compliance.

1. **One clear identity per plan.** Use the same plan names, plan years and service areas on your site, in directories and in partner listings. When assistants mix up plan names or years, [fixing wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) sets out the steps.
2. **Plan-year pages that can be read.** Publish benefits, networks and drug coverage as current web pages, not only PDFs, and mark or retire last year’s pages so they stop competing with this year’s.
3. **Plain explanations.** Write neutral explainers on plan types, enrollment windows and how to check a doctor or drug. Keep them educational; no personalized recommendations, and every claim checked against your filed plan documents.
4. **Independent proof.** Star ratings, accreditation and J.D. Power results are third-party sources assistants can cite. Report them accurately and with their year. The broader method is in [how brands build authority for AI search](https://underneath.agency/resources/how-brands-build-authority-for-ai-search).
5. **Reputation.** Reviews, complaint resolution and coverage in local news feed what assistants say about a carrier.
6. **Timed tracking.** Ask shoppers’ questions across ChatGPT, Gemini, Perplexity, Copilot, Claude and Google’s AI features before and during each enrollment window, more than once; [how many prompts to track](https://underneath.agency/resources/how-many-prompts-to-track-ai-visibility) explains the sample size.

Medicare and Marketplace marketing are regulated, so treat AI-facing content as marketing material subject to your usual review. Do not seed misleading content or fake reviews; see [legitimate GEO versus manipulation](https://underneath.agency/resources/legitimate-geo-vs-manipulation).

## Which questions about AI and plan shopping are still unanswered?

Whether AI answers change enrollment choices is unmeasured; most data show interest in AI help, not its effect.

- **No enrollment data.** No survey counts members who used an AI assistant before choosing a plan.
- **Some figures are older or partial.** KFF’s 69% comes from the 2021 open enrollment period. The eHealth and Media Logic surveys come from firms that sell plans or insurer marketing.
- **Health citation research covers medical questions.** Coverage questions may draw on different sources than symptom questions.
- **The software pricing study is not about insurance.** We use it as a pattern, not as a measure of insurer accuracy.
- **Revenue links are thin in every industry, not only insurance.** The evidence that does exist is gathered in [does AI visibility drive business results](https://underneath.agency/resources/does-ai-visibility-drive-business-results).

## Where should a health insurer start before the next enrollment window?

Check what the main assistants say about your plans in your top markets, then fix the public facts behind errors.

Pick your largest counties and products. Write the questions an individual shopper, a Medicare beneficiary, an adult child helping a parent, and an employer’s broker would ask. Ask each one in several assistants, then ask again a few days later. Record which carriers are named, which sources are cited, and whether your networks, drug coverage, ratings and costs come back right for the correct plan year. Most errors in this industry trace back to outdated plan documents or inconsistent plan names.

If you would like a second pair of eyes, [request an enrollment-season AI visibility review](https://underneath.agency/contact). It shows how AI answers describe your plans in the markets that matter most, where they cite outdated or wrong information, and which fixes are most likely to protect new member enrollment and renewals before the window opens. Our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) maps that work onto a regulated insurer’s calendar, with readable plan-year pages, compliance review and tracking timed around each enrollment window.

## Frequently asked questions

### Do Medicare beneficiaries use ChatGPT to pick plans?

No study measures it directly. eHealth found 50% of beneficiaries would be interested in an AI agent by phone, and OpenAI lists insurance comparisons as a ChatGPT Health use.

### Can an AI assistant recommend our plan to a specific person?

Assistants can describe plans, but they are not licensed agents. Insurers should keep their own content educational and avoid personalized recommendations.

### Why do AI answers show last year’s benefits?

Often because old plan-year pages remain online. In our software pricing study, 39 of 64 differing prices appeared on another page of the vendor’s own site.

### Do star ratings and J.D. Power awards help?

They are independent sources assistants can cite, and ChatGPT searched for reviews or ratings in 46.2% of answers in our study. Their effect on mentions is untested.

### When should an insurer check its AI visibility?

Before each enrollment window, and again during it. Medicare’s annual window runs October 15 to December 7.

## Sources

- Centers for Medicare & Medicaid Services (2026-01-28), [Marketplace 2026 Open Enrollment Period Report: National Snapshot](https://www.cms.gov/newsroom/fact-sheets/marketplace-2026-open-enrollment-period-report-national-snapshot-2)
- KFF (2024-11), [Medicare Advantage 2025 Spotlight: A First Look at Plan Offerings](https://www.kff.org/medicare/issue-brief/medicare-advantage-2025-spotlight-a-first-look-at-plan-offerings/)
- KFF (2024-09-26), [Nearly 7 in 10 Medicare Beneficiaries Did Not Compare Plans During Medicare’s Open Enrollment Period](https://www.kff.org/medicare/issue-brief/nearly-7-in-10-medicare-beneficiaries-did-not-compare-plans-during-medicares-open-enrollment-period)
- KFF (2025), [2025 Employer Health Benefits Survey](https://www.kff.org/health-costs/2025-employer-health-benefits-survey/)
- KFF (2026), [KFF Tracking Poll on Health Information and Trust: Use of AI for Health Information and Advice](https://www.kff.org/health-information-and-trust/poll-finding/kff-tracking-poll-on-health-information-and-trust-use-of-ai-for-health-information-and-advice/)
- TechTarget (2025), [Top stats on Medicare Advantage enrollment, costs in 2025](https://www.techtarget.com/healthcarepayers/news/366628088/Top-stats-on-Medicare-Advantage-enrollment-costs-in-2025)
- Chain Drug Review (2025-05-28), [J.D. Power: Gap widens between highest- and lowest-performing employer-sponsored health plans](https://chaindrugreview.com/j-d-power-gap-widens-between-highest-and-lowest-performing-employer-sponsored-health-plans)
- eHealth (2025-10-01), [Survey: 75% of Medicare Beneficiaries Say Selecting a Plan Is Confusing](https://s204.q4cdn.com/837903328/files/doc_news/Survey-75-of-Medicare-Beneficiaries-Say-Selecting-a-Plan-Is-Confusing-2025.pdf)
- Media Logic (2025), [What We Learned from the 2025 Medicare AEP Shopping Experience Survey](https://www.medialogic.com/blog/healthcare-marketing/2025-medicare-aep-shopping-experience/)
- Health Populi (2026-03-24), [Consumer adoption of AI for health and self-care doubling, via Rock Health’s latest snapshot](https://www.healthpopuli.com/2026/03/24/consumer-adoption-of-ai-for-health-and-self-care-doubling-to-36-in-a-year-via-rock-healths-latest-snapshot)
- OpenAI (2026-01-07, updated 2026-07-23), [Introducing ChatGPT Health](https://openai.com/index/introducing-chatgpt-health/)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- OpenAI Help Center (2025), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Jacques and colleagues (2026), [Authority Signals in AI Cited Health Sources: A Framework for Evaluating Source Credibility in ChatGPT Responses](https://arxiv.org/abs/2601.17109), arXiv:2601.17109.
- Underneath (2026), [When does Google show an AI Overview? 1,248 US searches](https://underneath.agency/research/ai-overviews-frequency-study)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [How faithfully do AI assistants quote software prices?](https://underneath.agency/research/ai-pricing-accuracy-study)

---

This is the Markdown twin of https://underneath.agency/resources/health-insurers-members-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How health providers can win new patients from AI search"
description: "By making providers easy for AI to name and describe correctly: full physician profiles, consistent facts, strong reviews. Patients now ask AI who to see."
canonical: "https://underneath.agency/resources/healthcare-providers-patients-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How can a hospital, clinic or practice win new patients from AI search?

By making sure AI assistants can find, name and correctly describe your physicians, locations and services when a patient asks who to see. Patients increasingly ask an assistant before they book, and one recent survey found AI tools now influence the choice of doctor about as much as Google results and physician referrals. Early research also suggests the answers favor hospital systems whose physician pages are complete and current, which leaves many independent practices unseen.

## The short version

1. Patients are asking AI who to see: in [rater8’s 2026 survey](https://rater8.com/2026-patient-choice-report/) of 992 US adults, 47% had used AI to research healthcare providers, up from 31% in 2025.
2. AI now rivals referrals: among people who looked for a new doctor, 36% named AI tools as a top influence, against 34% for Google search results and 32% for referrals from another doctor.
3. Independents can vanish: in a [study reported by Medical Economics](https://www.medicaleconomics.com/view/why-chatgpt-favors-hospitals-over-independent-practices), ChatGPT cited and mentioned none of 200 randomly chosen independent practices across about 4,950 patient-style questions.
4. Hospital rosters feed the answers: in the same study, hospital staff rosters were the largest source of ChatGPT’s citations, at 27.5%.
5. Wrong details are common: 66% of patients who had used AI to research providers said they had seen incorrect provider information, such as wrong addresses, phone numbers, insurance details or hours.

A note before you read: this article is about how provider organizations are found in AI answers. It is not medical advice, and it does not suggest that AI assistants should replace a clinician’s judgment.

## Who chooses a provider today, and what is a new patient worth?

Patients choose more actively than before, often starting with a local search, and each one can bring years of care.

Switching is common. In rater8’s survey, 72% of patients were in the market for a new provider over the past year: 47% chose a new doctor and 25% searched without switching. When they searched, 55% started with some version of “[specialty] near me.” rater8 sells reputation software to medical practices, so treat its numbers as vendor research; the survey was fielded in April 2026, according to [Medical Economics](https://www.medicaleconomics.com/view/patients-now-shop-for-doctors-like-consumers-and-the-bar-just-got-higher).

Most providers now sit inside larger organizations. The [American Hospital Association](https://www.aha.org/statistics/fast-facts-us-hospitals) counts 6,100 hospitals in the United States. According to Avalere Health’s analysis for the Physicians Advocacy Institute, [reported by Medical Economics](https://www.medicaleconomics.com/view/physician-independence-vanishes-as-corporate-medicine-swallows-up-u-s-health-care), 82% of practicing doctors were employed by hospitals or corporate entities at the start of 2026, and those entities owned 63.9% of physician practices, against 29.8% in 2018.

A new patient is worth far more than one visit. The best-known measure is old but telling: in a 2019 Merritt Hawkins survey of 62 hospital finance chiefs, [reported by Fierce Healthcare](https://www.fiercehealthcare.com/practices/physicians-generate-2-4m-each-year-for-hospitals-survey), physicians generated an average of $2.38 million a year in net revenue for their affiliated hospitals, through admissions, tests, treatments and procedures. For a health system, the patient who picks one of its physicians brings that downstream care too. For an independent practice, the same patient is the business.

## How far has AI entered patients’ search for care?

Far enough that a third of US adults use it for health information, and many ask it about providers.

The [KFF Tracking Poll on Health Information and Trust](https://www.kff.org/health-information-and-trust/poll-finding/kff-tracking-poll-on-health-information-and-trust-use-of-ai-for-health-information-and-advice/) found about a third (32%) of adults turning to AI for health information and advice in the past year. Among those users, 41% said a major reason was to look up information before deciding whether to see a provider, and 18% said they did not have a regular provider or could not get an appointment. Health care professionals (80%) and search engines (68%) remain far more common sources, but many search results now carry AI summaries of their own. Our guide on [how telehealth companies get chosen by AI](https://underneath.agency/resources/telehealth-patients-from-ai-search) covers patients who turn to virtual care instead.

The platforms describe the scale themselves. OpenAI says that, based on its de-identified analysis, [over 230 million people globally](https://openai.com/index/introducing-chatgpt-health/) ask health and wellness questions on ChatGPT every week. That is documented by the platform, not measured by us.

Provider choice is following. In rater8’s survey, AI tools were a top influence for 36% of people who searched for a new doctor, and for 39% of those who actually switched. Asked which part of a Google results page they trust most when researching a provider, 37% chose the AI Overview, ahead of the regular links and the map.

## Where is the line between finding a provider and medical advice?

A provider’s job in AI search is to be found and described correctly, not to diagnose anyone.

The platforms draw the same line. OpenAI says its health experience “is not intended for diagnosis or treatment.” Google has already pulled back where answers went wrong: after a Guardian investigation, [TechCrunch reported](https://techcrunch.com/2026/01/11/google-removes-ai-overviews-for-certain-medical-queries) that AI Overviews were removed for queries such as “what is the normal range for liver blood tests.” Drugmakers work under the same line plus FDA rules, as our guide on [keeping drug brands accurate in AI answers](https://underneath.agency/resources/pharma-brand-visibility-ai-search) explains.

The two kinds of question also look different on Google. In [our study of when Google shows an AI Overview](https://underneath.agency/research/ai-overviews-frequency-study), healthcare and dental keywords showed one 43.0% of the time, and 60.0% of them showed a local map pack. Healthcare and dental searches with a map showed an AI Overview 13.3% of the time; those without a map, 87.5%. In plain terms, AI summaries mostly answer the “what is this?” questions, while the map still answers “who near me?”

Our inference for provider marketing: keep two kinds of page separate. Patient education should be written or reviewed by clinicians, sourced and dated. Pages meant to win the booking should state facts a patient can act on: who the physicians are, what they treat, where, which insurance plans are accepted, and how to get an appointment. What tends to appear on cited commercial health pages is covered in [what commercial health sites cited by ChatGPT have in common](https://underneath.agency/resources/chatgpt-cited-commercial-health-sites).

## What do patients ask AI when choosing care?

Questions that join a specialty, a place, an insurance plan and a reputation check. We wrote these example prompts to show the pattern; they are not recorded patient searches.

| Stage | Illustrative prompt |
|---|---|
| Specialty and place | “Best orthopedic surgeon for knee replacement in Columbus” |
| Access | “Pediatric dentist near me accepting new patients on Medicaid” |
| Insurance | “Which dermatologists in Charlotte take Blue Cross?” |
| Condition to specialist | “What kind of doctor treats long-term dizziness, and who is good for it nearby?” |
| Reputation | “Is [hospital] good for heart surgery? What do patients say?” |
| Comparison | “[Health system A] vs [health system B] for having a baby” |

Only the condition question touches medical information, and even there the useful answer for a provider is a referral to the right specialty, not advice. Every other question is about access, fit and trust, which is exactly the information a provider organization controls.

## How does an AI answer become a booked appointment, and what does a miss cost?

The answer names the provider, the patient checks reviews and access, then books. A missing or wrong name breaks the chain.

The path is short: an AI answer → the physician or location page → a rating and review check → insurance and availability → a call or online booking. Each step can end it.

- **The rating filter.** In rater8’s survey, 75% of patients would not book with a provider rated below 4.0 stars, and 55% had canceled or avoided an appointment because of online reviews, up from 40% in 2025.
- **Wrong details.** Of the 465 respondents who had used AI to research providers, 66% said they had met incorrect provider information, and 60% of AI users said they trusted the summary without checking it. The same report cites a 2024 study of insurer directories in which physician addresses were consistent only 17% to 28% of the time.
- **Our own check.** When we asked four assistants for the address, phone, website and hours of real local businesses, [answers differed from the Google profile](https://underneath.agency/research/ai-business-facts-accuracy-study) for 17.2% of dentists and 21.1% of physiotherapists.

What it costs is not yet measured in dollars. Our inference: a patient who is given another practice’s name, or your old phone number, rarely comes back to look for you, and each lost new patient takes their future visits with them. If assistants already describe your organization wrongly, [here is how to fix wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers).

## Why do AI assistants name some providers and skip others?

Because they can only name providers they find described in sources, and hospital systems have far more of those sources.

**Documented by Google.** For local results on Google, the company says ranking is [mainly based on relevance, distance and popularity](https://support.google.com/business/answer/7091?hl=en), and it asks businesses to provide complete and detailed information. That covers the map, not how assistants choose.

**Observed in a study.** The Medical Economics study, run in May 2026 by a company that sells AI marketing software to independent practices, tested ChatGPT only. Across its larger set of questions, ChatGPT’s citations came from:

| Source | Share of citations |
|---|---|
| Hospital staff rosters | 27.5% |
| Practice-owned websites, mostly large groups | 23.0% |
| Specialty-association and credentialing pages | 22.1% |
| National directories | 8.3% |

Employed physicians appear on hospital rosters automatically, and system web teams keep them current. The pattern varied by specialty and city: cardiology and orthopedics tilted hardest toward hospitals, dermatology and plastic surgery gave independents more room, and independents’ share ranged from about 7% in Boston to more than 20% in Charlotte.

**Observed in our studies.** When we asked ChatGPT for the best local providers in services that included dentists and physical therapists, businesses with more Google reviews than the local median were [19.5 points more likely to be listed](https://underneath.agency/research/chatgpt-local-picks-google-profile-study), after adjusting for map rank and other signals.

**Our inference.** Assistants appear to favor providers whose facts are published in many consistent, credible places: roster pages, association listings, review profiles and the provider’s own pages. Size helps because it produces those pages, not because assistants prefer hospitals for their own sake. The wider pattern is in [do AI assistants favor big brands over smaller competitors?](https://underneath.agency/resources/do-ai-assistants-favor-big-brands)

## How does GEO work for a provider organization?

Generative engine optimization (GEO) for providers means publishing accurate, complete facts about physicians, locations and services wherever assistants look.

1. **A full profile for every clinician.** Name, specialty, board certifications, training, conditions treated, procedures performed, locations, languages and insurance accepted, on one page per physician. For independents, this is the equivalent of the hospital roster.
2. **Consistent facts everywhere.** The same addresses, phone numbers, hours and plan lists on the website, Google Business Profile, insurer directories and listing sites. Inconsistent sources produce inconsistent answers.
3. **Association and credentialing pages.** Check what specialty societies and credentialing listings say about each physician, and correct them. These made up 22.1% of citations in the Medical Economics study, and most practices never look.
4. **Reviews and replies.** Ask patients for reviews and reply to them. In rater8’s survey, 66% said a provider’s response to reviews influences their trust. Privacy rules limit what a reply can say about any patient, so keep replies general.
5. **Service-line pages built for access.** For each service, state who it is for, who provides it, where, and how to book. Keep clinical education on separate, clinician-reviewed pages.
6. **Local coverage.** News about new physicians, locations and services in local and specialty press places names next to places and conditions.
7. **Regular checks.** Ask assistants the specialty, access and reputation questions above every quarter, and track which physicians and sources appear.

Health system IT buying is a separate market; for vendors selling to hospitals, see [our healthcare software article](https://underneath.agency/resources/healthcare-software-ai-search). Vendors selling to independent practices can see [how practice software vendors win demos](https://underneath.agency/resources/medical-practice-software-ai-search). GEO cannot guarantee that an assistant names any provider; it makes the facts it finds complete, correct and easy to verify.

## What is still unknown about AI and patient choice?

Who gets named is starting to be measured; how many appointments AI answers produce is not.

- **Vendor research.** The patient survey and the independent-practice study both come from companies that sell to practices.
- **One assistant, one month.** The citation study tested ChatGPT through its developer interface in May 2026; Gemini, Perplexity and Google’s AI features may differ.
- **No booking data.** We found no public data linking AI visibility to scheduled appointments or patient revenue.
- **Old value figures.** The physician revenue survey dates from 2019.
- **Policies are moving.** Google and OpenAI are both changing how they handle health questions, so today’s answers may not hold.

## Where should a provider organization start?

Start by asking assistants the specialty, insurance and reputation questions your patients ask, for your top service lines.

Pick the three service lines and markets that matter most for new patients. Ask ChatGPT, Gemini, Perplexity and Google AI Mode who to see, and note which physicians and locations are named, which sources are cited, and whether your addresses, phone numbers and plans are right.

If new-patient growth in a service line depends on being chosen before anyone calls, [ask us to check how AI assistants present your physicians and locations](https://underneath.agency/contact). We will test the questions patients ask in your specialties and cities, find the wrong or missing facts, and plan the profile, review and coverage work that helps patients find and book with you. Our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page describes how that work is run for provider organizations, from finding the physician facts assistants miss to measuring whether the fixes held.

## Frequently asked questions

### Do patients use ChatGPT to find a doctor?

Increasingly. In one 2026 survey of 992 US adults, 47% had used AI tools to research providers, and 36% of people seeking a new doctor named AI as a top influence.

### Why does ChatGPT recommend hospital doctors over independent practices?

One study found hospital staff rosters were its largest citation source, at 27.5%. Hospital physicians have maintained profile pages; many independents have none.

### Should a practice write health content to appear in AI answers?

Only clinician-reviewed, sourced content. Pages meant to win bookings should state facts patients act on: physicians, services, locations, insurance and how to book.

### Do online reviews affect whether AI recommends a provider?

They appear to. In our local study, businesses with more reviews than the local median were 19.5 points more likely to be listed by ChatGPT.

## Sources

- rater8 (2026), [2026 Patient Choice Report](https://rater8.com/2026-patient-choice-report/)
- Medical Economics (2026), [Patients now shop for doctors like consumers, and the bar just got higher](https://www.medicaleconomics.com/view/patients-now-shop-for-doctors-like-consumers-and-the-bar-just-got-higher)
- Medical Economics (2026-07-20), [Why ChatGPT favors hospitals over independent practices](https://www.medicaleconomics.com/view/why-chatgpt-favors-hospitals-over-independent-practices)
- Medical Economics (2026-05-13), [Physician independence vanishes as corporate medicine swallows up U.S. health care](https://www.medicaleconomics.com/view/physician-independence-vanishes-as-corporate-medicine-swallows-up-u-s-health-care)
- KFF (2026-03), [KFF Tracking Poll on Health Information and Trust: Use of AI for Health Information and Advice](https://www.kff.org/health-information-and-trust/poll-finding/kff-tracking-poll-on-health-information-and-trust-use-of-ai-for-health-information-and-advice/)
- OpenAI (2026-01-07), [Introducing ChatGPT Health](https://openai.com/index/introducing-chatgpt-health/)
- TechCrunch (2026-01-11), [Google removes AI Overviews for certain medical queries](https://techcrunch.com/2026/01/11/google-removes-ai-overviews-for-certain-medical-queries)
- American Hospital Association (2026), [Fast Facts on U.S. Hospitals](https://www.aha.org/statistics/fast-facts-us-hospitals)
- Fierce Healthcare (2019), [Survey: Physicians net $2.4M in revenue for hospitals each year](https://www.fiercehealthcare.com/practices/physicians-generate-2-4m-each-year-for-hospitals-survey)
- Google Business Profile Help, [Tips to improve your local ranking on Google](https://support.google.com/business/answer/7091?hl=en)
- Underneath (2026), [When does Google show an AI Overview? 1,248 US searches](https://underneath.agency/research/ai-overviews-frequency-study)
- Underneath (2026), [Do AI answers match a business’s Google profile?](https://underneath.agency/research/ai-business-facts-accuracy-study)
- Underneath (2026), [Which Google Maps businesses does ChatGPT recommend?](https://underneath.agency/research/chatgpt-local-picks-google-profile-study)

---

This is the Markdown twin of https://underneath.agency/resources/healthcare-providers-patients-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Can AI search help healthcare software firms win hospital deals?"
description: "Yes, if AI answers can find proof of EHR integration, HIPAA security and peer results, because hospitals now narrow long vendor lists before any demo."
canonical: "https://underneath.agency/resources/healthcare-software-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Can AI search help a healthcare software company win health system deals?

It can help you reach the shortlist, which is where most health system deals are decided, but only if AI answers can find independent proof that you integrate with the hospital’s EHR, protect patient data and deliver results for peers. Health systems are buying faster and evaluating more vendors than before, so the first narrowing happens earlier and with less help from your sales team. No study has yet measured how often health IT buyers use AI assistants, so treat what follows as a well-supported bet, not a proven channel.

## The short version

1. Healthcare is buying software fast: [Menlo Ventures](https://menlovc.com/perspective/2025-the-state-of-ai-in-healthcare/), surveying more than 700 healthcare executives, estimates $1.4 billion of healthcare AI spending in 2025, with health systems’ buying cycles down from 8.0 months to 6.6 months.
2. The field of vendors is crowded: one health system, Advocate Health, evaluated over 225 AI solutions to select 40 use cases, Menlo reports, and 85% of generative AI spend went to startups rather than incumbents.
3. The EHR shapes every purchase. Epic held 43.7% of the US acute care EHR market in 2025, per [KLAS data reported by Fierce Healthcare](https://www.fiercehealthcare.com/health-tech/epic-continues-grow-ehr-market-share-it-makes-gains-small-health-systems), and in a [Bain and KLAS survey](https://www.biospace.com/bain-and-company-and-klas-study-finds-80-percent-of-us-healthcare-providers-are-accelerating-spending-on-it-and-software-with-ai-top-of-mind) nearly two-thirds of providers looked first to existing vendors before evaluating new ones, though 94% were open to looking elsewhere.
4. Security is a gate, not a feature: [HIPAA Journal](https://www.hipaajournal.com/healthcare-data-breach-statistics/) counts a record 772 large healthcare data breaches in 2025, and [IBM](https://newsroom.ibm.com/2025-07-30-ibm-report-13-of-organizations-reported-breaches-of-ai-models-or-applications,-97-of-which-reported-lacking-proper-ai-access-controls) puts the average healthcare breach at $7.42 million.
5. Clinicians are already AI users and want a vote: 81% of physicians use AI professionally, and 85% want a say on AI adoption in their practice, according to the [American Medical Association’s 2026 survey](https://www.fiercehealthcare.com/ai-and-machine-learning/ama-physicians-use-ai-doubled-2023-2026).

## Who signs a health system software contract, and what is the account worth?

A committee buys it, with IT, finance, clinicians and security all holding a veto. A won health system pays for years.

No single executive signs a health system deal alone. A chief information officer or chief medical information officer may lead, but finance wants a return, security wants a risk review, and clinicians want to know the tool fits their day. The [AMA’s 2026 survey](https://ascopost.com/news/march-2026/ama-survey-finds-rapid-growth-in-physician-ai-adoption/) of nearly 1,700 doctors found 55% want to be involved in AI implementation decisions to evaluate the clinical evidence, and Fierce reports that 85% want a say on AI adoption in their practice.

Budgets are moving toward software. In the 2023 Bain and KLAS survey of more than 200 provider executives, nearly 80% said they had increased IT spending materially over the past year. [G2’s healthcare survey](https://research.g2.com/insights/healthcare-roi-survey) of 208 buyers at health systems, hospitals and practices found 80% planned to raise their technology budget within 12 months.

The money is concentrated. Menlo estimates US healthcare administration spending at $740 billion a year, of which healthcare IT is only $63 billion. Health systems supplied $1 billion of the $1.4 billion flowing into healthcare AI in 2025. In practice, that means a small number of large systems decide which vendors grow. Vendors selling to independent practices face a different buyer, covered in [how practice software reaches specialty clinics](https://underneath.agency/resources/medical-practice-software-ai-search).

How much a health system account is worth depends on what you sell:

- **Systems of record are sticky.** Oracle Health held 21.9% of the acute care EHR market in 2025, and KLAS found 35% of sampled customers say they might leave or “want to leave but can’t.” Once a core platform is installed, it tends to stay for years.
- **Point AI tools are not yet sticky.** Menlo found large health systems using ambient scribes were just as likely to switch vendors as to stay, and among outpatient providers the likelihood to switch rose to 67%. For these vendors, a customer’s value depends on being named again at every review.

Virtual care platforms sell into the same health systems; see [how virtual care platforms win contracts](https://underneath.agency/resources/virtual-care-platforms-ai-search).

## How far have AI assistants entered health IT vendor research?

Mostly as a research and narrowing tool around the first shortlist; no published study isolates health IT buyers yet.

The people who sign health IT contracts already use AI daily. The AMA found 81% of physicians use AI in their practices, up from 38% in 2023, and the share using it to summarize medical research rose to 39% in 2026 from 13% in 2024. Among health systems, an [Eliciting Insights survey reported by Fierce Healthcare](https://fiercehealthcare.com/ai-and-machine-learning/75-us-healthcare-systems-use-plan-use-ai-platform-2026) found 75% now use at least one AI application, up from 59% in 2025. And [KLAS](https://engage.klasresearch.com/blog/where-healthcare-is-investing-and-betting-on-ai-in-2025/8955/), drawing on 228 executives, reports that 70% of providers and 80% of payers have AI strategies underway.

None of these surveys asks whether those executives use ChatGPT or Gemini to research vendors. For vendor research itself, the nearest data point comes from outside healthcare: [G2’s March 2026 survey](https://company.g2.com/news/g2-research-the-answer-economy) of 1,076 software buyers found that 51% started research with an AI chatbot more often than with Google. G2 runs a review marketplace and has an interest in this finding, and its sample is not healthcare-specific.

What is documented is the scale of the narrowing job. Menlo reports that Advocate Health evaluated over 225 AI solutions to select 40 use cases, and that SimonMed, a radiology group, went from co-building with fewer than 10 vendors to piloting solutions from more than 50. A team screening hundreds of vendors needs a fast first pass. We infer that AI assistants increasingly do some of that first pass, alongside peers, KLAS reports and conferences.

A health IT leader who sticks to Google will probably meet an AI answer anyway. Across the 1,248 US searches in [our AI Overview study](https://underneath.agency/research/ai-overviews-frequency-study), keywords in B2B software and technology triggered an AI Overview on 96.0% of searches, the highest rate among the eight industries we sampled.

## Which questions do health IT buyers ask AI assistants?

Questions about fit with their EHR, HIPAA and security, peer results and return on investment. We wrote the prompts in this table ourselves to show how a CIO or CMIO might phrase these concerns; none was collected from an actual hospital buyer.

| Buying concern | Illustrative prompt |
|---|---|
| Category | “What are the leading ambient documentation tools for a 12-hospital health system?” |
| EHR fit | “Which prior authorization tools integrate natively with Epic?” or “…with Oracle Health?” |
| Alternatives | “What are the alternatives to Nuance DAX Copilot for an academic medical center?” |
| Security and compliance | “Which patient engagement platforms will sign a business associate agreement and have HITRUST certification?” |
| Peer proof | “How do these revenue cycle vendors score in KLAS?” |
| Return | “What denial reduction have health systems reported from AI coding tools?” |
| Interoperability | “Which care coordination platforms support FHIR APIs and TEFCA exchange?” |

A single question from a hospital IT team can trigger several searches behind the scenes. Google’s documentation says AI Overviews and AI Mode [may use a “query fan-out” technique](https://developers.google.com/search/docs/appearance/ai-features), running multiple related searches on subtopics; for a health IT question, those could cover EHR fit and certifications separately. OpenAI, for its part, explains that [ChatGPT search typically rewrites a question](https://help.openai.com/en/articles/9237897-chatgpt-search) as one or more targeted queries that go to its search partners.

When [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study) logged ChatGPT’s queries, it went looking for reviews or ratings in 46.2% of its answers and went after a named publication, ranking or award in 43.8%. In health IT, a reasonable expectation is that the named sources would include KLAS, analyst coverage and trade publications.

## What path runs from an AI answer to a signed health system contract?

By getting you onto the first vendor screen, which is where hospital pilots and system-wide contracts begin.

As we understand the evidence, the path runs like this:

1. A CIO, CMIO, revenue cycle leader or innovation team asks an assistant a category, EHR-fit or alternatives question.
2. The assistant returns a short list of health IT vendors, with links to some of the pages it drew on.
3. The team checks those names against peers and KLAS. KLAS says 95% of its data comes from in-depth phone conversations with providers and payers.
4. Survivors face an integration and security review: EHR connection, business associate agreement, risk questionnaire.
5. A pilot follows, then a system-wide contract and, often, expansion to more sites.

Two features make healthcare different from general software. First, the incumbent EHR competes at every step: Menlo found most customers prefer to buy AI from their incumbent EHR for everything except ambient scribes and chart review. Second, the pace has changed. Menlo says buying cycles have compressed from 12 to 18 months to under six for many providers, which leaves less time to get noticed once a search starts. Payers are the exception: their cycles lengthened from 9.4 months to 11.3 months.

A hospital contract that began with an AI answer almost never carries a referral tag in your analytics. A buyer who met you in an AI answer may appear months later through a conference meeting, a peer referral or an EHR app marketplace. We cover the measurement problem in [how to prove GEO caused sales](https://underneath.agency/resources/prove-geo-caused-sales).

## Why does an AI assistant list one health IT vendor and skip another?

The assistants keep their selection rules private; research favors independent sources, and hospitals favor proof they can check.

**What Google and OpenAI disclose.** Both companies confirm that their AI answers run web searches and show links to what they drew on. Neither explains why a particular EHR add-on or revenue cycle vendor gets named.

**What researchers have measured.** On US software questions, AI search drew 72.7% of its sources from earned sites such as reviews and independent publications, [Chen and colleagues](https://arxiv.org/abs/2509.08919) report, compared with 45.4% for Google; in health IT, KLAS and the trade press are the obvious earned sites. In [our four-assistant study](https://underneath.agency/research/ai-assistants-brand-agreement-study), the overlap between assistants’ brand picks was highest for B2B software, at 0.543 on a scale from 0 to 1, so being named by one assistant is no guarantee of the others.

**What health buyers check.** Several trust factors are specific to this industry:

- **EHR integration.** The Bain and KLAS survey called seamless EHR integration a key purchasing criterion for all providers. With Epic at 56.9% of US hospital beds, per KLAS, “works with Epic” is often the first filter.
- **Security and privacy.** Hacking and IT incidents caused more than 80% of large breaches in 2025, HIPAA Journal reports, and the largest that year was at a business associate, Conduent, exposing the health information of more than 62 million Americans. Vendors are part of the risk hospitals screen for. Physicians feel it too: on patient privacy, 41% expected AI to cause harm against 13% expecting help.
- **Interoperability.** In 2023, 70% of US hospitals at least sometimes exchanged data across all four domains of interoperability (send, find, receive, integrate), according to federal data [reported by TechTarget](https://techtarget.com/searchhealthit/news/366586275/70-of-hospitals-participate-in-healthcare-interoperability). Buyers expect vendors to fit that exchange.
- **Peer evidence.** Health executives buy on outcomes they can verify with a peer. In G2’s survey, 83% said it is important that software they buy incorporates AI, but KLAS reports leaders “increasingly seeking investments that provide measurable ROI,” with providers refocused on revenue cycle tools.

**Our inference.** Most of what a health system verifies is public or could be: integration listings, certifications, security documentation, peer ratings, published outcomes. An assistant can read and cite the same pages a hospital’s security reviewer opens. We would expect a health IT vendor whose integration, security and outcome proof is public, up to date and consistent to give both the buying committee and the assistant more to go on. No study has yet checked it with EHR add-ons, scribes or revenue cycle tools.

## How much can a missing AI mention cost an EHR add-on or scribe vendor?

Mainly seats in health system evaluations rather than website visits, though nobody has yet priced that loss.

- **Shorter cycles shrink the window.** If a health system decides in 6.6 months instead of 8.0, a vendor that is not in the first screen has less time to be added later.
- **The incumbent fills the gap.** When buyers look first to existing vendors, as nearly two-thirds did in the Bain and KLAS survey, a startup missing from AI answers loses by default to the EHR’s own module, we infer.
- **Weak stickiness raises the cost.** In categories where customers are as likely to switch as stay, being absent at renewal reviews can mean losing existing revenue, not only new deals.
- **One good answer is not a fixed place.** When [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study) put the same question to ChatGPT five times, only 25.2% of the brands ChatGPT named came back in every run.

## Which GEO work matters most when your buyers are hospitals?

Making EHR, security and outcome proof easy for assistants to find and repeat, though no mention is guaranteed.

1. **One clear identity.** Say what you do, for which care setting and with which EHRs, the same way on your site, app marketplace listings, KLAS profile and company databases. When an assistant gets your EHR support or care setting wrong, our guide to [fixing wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) sets out the corrections.
2. **Public integration and security pages.** Publish which EHRs you connect to and how, which standards you support, your certifications, and whether you sign business associate agreements, on plain web pages rather than only in sales decks.
3. **Independent validation.** Take part in KLAS interviews, publish outcomes with named health system customers who agree to it, and earn coverage in trade publications. KLAS profiles and health IT trade press fit the wider method we describe in [how brands build authority for AI search](https://underneath.agency/resources/how-brands-build-authority-for-ai-search).
4. **Content for real buyer questions.** Write honest comparison, alternatives and return-on-investment pages for specific care settings. No medical advice and no unsupported clinical claims. Third-party rankings of health IT vendors tend to carry more weight than your own site; [which pages to target](https://underneath.agency/resources/best-of-lists-ai-recommendations) explains why.
5. **Newness is a handicap to plan for.** With 85% of AI spend going to startups, many vendors are young. Assistants often miss recent launches, as we explain in [why ChatGPT misses new products](https://underneath.agency/resources/why-chatgpt-misses-new-products).
6. **Tracking every assistant a hospital buyer might open.** Put your health system buyers’ questions repeatedly to ChatGPT, Google AI Overviews and AI Mode, Gemini, Perplexity, Copilot and Claude; [how many prompts to track](https://underneath.agency/resources/how-many-prompts-to-track-ai-visibility) covers the sample size.

Stay honest. Planted reviews or invented outcomes are a regulatory and reputational risk in healthcare; see [legitimate GEO versus manipulation](https://underneath.agency/resources/legitimate-geo-vs-manipulation). Much of this adapts the approach in [B2B SaaS revenue from AI search](https://underneath.agency/resources/b2b-saas-revenue-from-ai-search) to hospital procurement. School districts buy in a similar way, through committees, evidence reviews and privacy checks, as [how EdTech companies reach schools](https://underneath.agency/resources/edtech-ai-search) shows.

## Which questions about AI search and hospital purchasing remain open?

Nobody has yet shown that appearing in AI answers grows a health IT vendor’s pipeline.

- **No survey of health IT buyers’ AI search use.** The AI adoption surveys above measure use of AI products, not use of assistants to choose vendors.
- **The sources have interests.** Menlo Ventures invests in healthcare AI companies, G2 sells review visibility, and Sage Growth Partners, which found 57% of executives rank AI clinical tools as their [top technology initiative](https://www.healthcareittoday.com/2026/03/29/bonus-features-march-29-2026-57-of-execs-say-ai-based-clinical-tools-are-their-top-tech-initiative-57-of-patients-say-ai-isnt-mature-enough-for-docs-to-trust-it-plus-32-more-s/), advises health tech marketers.
- **Some data are dated.** The Bain and KLAS finding on looking first to existing vendors comes from 2023.
- **The link to revenue is the thinnest evidence of all.** The general case is reviewed in [does AI visibility drive business results](https://underneath.agency/resources/does-ai-visibility-drive-business-results); health IT has no data of its own yet.

## How can a health IT vendor tell whether AI answers are costing it evaluations?

Ask what health system buyers ask, see which vendors come back, then publish the missing proof.

Write down what a CIO, CMIO or revenue cycle director would ask at each stage and in each care setting: category, EHR fit, alternatives, security and return. Repeat every question several times in ChatGPT, Gemini, Perplexity and Google’s AI answers, because one run can mislead. Note which vendors are named, which sources are cited, and whether your EHR integrations, certifications and outcomes come back accurately. In health IT, the gaps usually trace to proof kept in sales decks rather than on public pages.

We can run that check with you: [ask us for a review of your health system shortlist visibility](https://underneath.agency/contact). It shows where AI answers place you on, or leave you off, those shortlists, and which gaps in your integration, security and outcome evidence are most likely costing you pilots, contracts and renewals. The [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page lays out how we diagnose those gaps, publish the missing EHR and security proof, and track what assistants then tell hospital buyers.

## Frequently asked questions

### Do hospital CIOs use ChatGPT to find vendors?

No published study measures it. What is known: 75% of US health systems use at least one AI application, and 81% of physicians use AI professionally.

### Does a KLAS rating help AI visibility?

Plausibly, but untested. ChatGPT targeted a named publication, ranking or award in 43.8% of answers in our study, and KLAS is health IT’s best-known rating source.

### Should we publish our security and HIPAA documentation?

Publish the facts buyers verify first, such as certifications and whether you sign business associate agreements. Keep sensitive detail behind a review process.

### Can we compete with the EHR vendor’s own AI module?

Yes, where you are clearly better: Menlo found buyers favor startups for ambient scribes and chart review, but prefer their EHR for most other uses.

### How fast can GEO show up in health system pipeline?

Expect quarters, not weeks. Health systems’ buying cycles average about 6.6 months for AI, and payers’ about 11.3 months.

## Sources

- Menlo Ventures (2025-10-21), [2025: The State of AI in Healthcare](https://menlovc.com/perspective/2025-the-state-of-ai-in-healthcare/)
- Bain & Company and KLAS Research, via BioSpace (2023-09-12), [Bain & Company and KLAS study finds 80% of US healthcare providers are accelerating spending on IT and software](https://www.biospace.com/bain-and-company-and-klas-study-finds-80-percent-of-us-healthcare-providers-are-accelerating-spending-on-it-and-software-with-ai-top-of-mind)
- KLAS Research (2025-10-16), [Where Healthcare Is Investing and Betting on AI in 2025](https://engage.klasresearch.com/blog/where-healthcare-is-investing-and-betting-on-ai-in-2025/8955/)
- KLAS Research (n.d.), [Uncover the Truth Behind the Hype](https://engage.klasresearch.com/healthcare-it-insights/)
- Fierce Healthcare (2026), [Epic grows EHR footprint among small health systems even as overall market sales decline in 2025](https://www.fiercehealthcare.com/health-tech/epic-continues-grow-ehr-market-share-it-makes-gains-small-health-systems)
- Fierce Healthcare (2026), [Health system AI adoption surges in 2026 with execs reporting increased ROI: survey](https://fiercehealthcare.com/ai-and-machine-learning/75-us-healthcare-systems-use-plan-use-ai-platform-2026)
- Fierce Healthcare (2026), [AMA: Physicians’ use of AI doubled from 2023 to 2026](https://www.fiercehealthcare.com/ai-and-machine-learning/ama-physicians-use-ai-doubled-2023-2026)
- The ASCO Post (2026-03), [AMA Survey Finds Rapid Growth in Physician AI Adoption](https://ascopost.com/news/march-2026/ama-survey-finds-rapid-growth-in-physician-ai-adoption/)
- G2 (2024), [Key Insights from G2’s 2024 Healthcare ROI Survey](https://research.g2.com/insights/healthcare-roi-survey)
- G2 (2026-04-15), [In the Answer Economy, Don’t Win the Click — Win the Answer](https://company.g2.com/news/g2-research-the-answer-economy)
- HIPAA Journal (2026), [Healthcare Data Breach Statistics](https://www.hipaajournal.com/healthcare-data-breach-statistics/)
- IBM (2025-07-30), [IBM Report: 13% of Organizations Reported Breaches of AI Models or Applications](https://newsroom.ibm.com/2025-07-30-ibm-report-13-of-organizations-reported-breaches-of-ai-models-or-applications,-97-of-which-reported-lacking-proper-ai-access-controls)
- TechTarget (2024-05), [70% of hospitals participate in healthcare interoperability](https://techtarget.com/searchhealthit/news/366586275/70-of-hospitals-participate-in-healthcare-interoperability)
- Healthcare IT Today (2026-03-29), [Bonus Features, March 29, 2026](https://www.healthcareittoday.com/2026/03/29/bonus-features-march-29-2026-57-of-execs-say-ai-based-clinical-tools-are-their-top-tech-initiative-57-of-patients-say-ai-isnt-mature-enough-for-docs-to-trust-it-plus-32-more-s/)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- OpenAI Help Center (2025), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Chen, Wang, Chen and Koudas (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919), arXiv:2509.08919.
- Underneath (2026), [When does Google show an AI Overview? 1,248 US searches](https://underneath.agency/research/ai-overviews-frequency-study)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude agree on brands?](https://underneath.agency/research/ai-assistants-brand-agreement-study)
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)

---

This is the Markdown twin of https://underneath.agency/resources/healthcare-software-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Can AI search help healthcare startups compete with big brands?"
description: "It can, if AI answers find independent proof of outcomes; buyers now pay for results and patients ask AI first, but new companies are rarely named unprompted."
canonical: "https://underneath.agency/resources/healthcare-startups-demand-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Can AI search help a healthcare startup compete with established health brands?

It can, but only for startups whose clinical and financial results are public, independently confirmed and easy for an assistant to find. Employers and health plans now tie fees to outcomes, and patients increasingly ask AI tools about their conditions, so proof decides both who gets bought and who gets used. AI assistants rarely name young companies in open-ended questions, so a startup has to earn its place; no study yet shows this lifts a healthcare startup’s sales.

## The short version

1. Purchasers pay for results now: in the [Peterson Health Technology Institute’s 2026 survey](https://www.phti.org/research-insight/2026-state-of-digital-health-purchasing/) of 321 digital health buyers, 68% of employers and 55% of health plans use performance-based contracts, and 85% of those contracts tie at least a quarter of fees to performance.
2. Winning a contract is only half the sale: 47% of purchasers told PHTI that fewer than a quarter of eligible members enroll, and poor member engagement is the leading reason they switch vendors.
3. Money is flowing, but unevenly: US digital health startups raised $14.2 billion in 2025 across 482 deals, according to Rock Health data [reported by Healthcare Dive](https://www.healthcaredive.com/news/digital-health-funding-2025-boosted-ai-rock-health/809449/), while 35% of deals were flat or down rounds, [Fierce Healthcare](https://www.fiercehealthcare.com/digital-health/jpm26-digital-health-funding-hit-142b-2025-ai-companies-taking-lions-share-dollars) reports.
4. Patients have moved to AI: Rock Health found use of AI chatbots for health information doubled from 16% to 32% between 2024 and 2025, as [Health Populi](https://www.healthpopuli.com/2026/03/24/consumer-adoption-of-ai-for-health-and-self-care-doubling-to-36-in-a-year-via-rock-healths-latest-snapshot) reports, and [KFF](https://www.kff.org/health-information-and-trust/poll-finding/kff-tracking-poll-on-health-information-and-trust-use-of-ai-for-health-information-and-advice/) also puts adult use of AI for health information at 32%.
5. Newness is a handicap: in a [study of 112 Product Hunt startups](https://arxiv.org/abs/2601.00912), ChatGPT recognized them by name 99.4% of the time but surfaced them in only 3.32% of discovery-style questions.

## Who does a healthcare startup have to convince, and what is a customer worth?

Usually three groups: a purchaser who signs, the people who must enroll, and investors who fund the gap.

**Purchasers.** Most digital health startups sell through employers, health plans or health systems. Spending is holding up: in PHTI’s 2026 survey, 56% of purchasers plan to keep digital health spending at current levels over the next 12 months and 38% plan to increase it. But buyers are cutting the number of vendors. A Solera Health survey cited by [D Magazine](https://www.dmagazine.com/healthcare-business/2026/07/baylor-scott-white-is-betting-employers-want-fewer-digital-health-vendors/) found 42% of employers manage eight or more digital health vendors, and 90% of those spend more than $1 million a year administering them. Startups compete for fewer slots. Our guide on [how healthtech vendors get named by employers and plans](https://underneath.agency/resources/healthtech-employers-payers-ai-search) covers that purchaser sale in depth.

**Members and patients.** A contract pays off only when people use the service. A Baylor Scott & White executive told D Magazine that for many employer-purchased products, “the average usage rate is about two or three percent.” With performance-based fees, low enrollment can mean lost revenue as well as a lost renewal. Health plans face their own enrollment challenge, covered in [how health insurers win members through AI](https://underneath.agency/resources/health-insurers-members-ai-search).

**Investors.** Rock Health’s 2025 data show a split market. AI-enabled companies took 54% of funding, megadeals over $100 million took 42%, and average deal size rose to $29.3 million. Meanwhile, digital health M&A rose to 195 deals, some of them survival moves: Fierce reports that Thirty Madison’s valuation reportedly fell from $1 billion to $500 million in its sale. Biotechs courting partners and investors face a related search, covered in [how biotech firms find partners through AI](https://underneath.agency/resources/biotech-partnering-ai-search).

What a won customer can be worth shows in the sector’s public companies. Hinge Health, one of five digital health companies that went public in 2025, reported [full-year 2025 revenue](https://s205.q4cdn.com/978984498/files/doc_news/Hinge-Health-reports-fourth-quarter-and-full-year-2025-financial-results-2026.pdf) of $587.9 million, with 2,830 clients and 25 million contracted lives. Its value lies in contracts that put a service in front of millions of eligible people.

## How far has AI entered the way patients and health buyers research options?

Deeply for patients, documented by surveys; for purchasers, AI use for vendor research is not yet measured.

**Patients and members.** Beyond the 32% figures, Health Populi reports that ChatGPT was used by 23% of Rock Health’s health information seekers and Gemini by 15%. Among people using AI for health, 59% explored treatment options based on a diagnosis and 55% researched prescription drugs or side effects. These users are also the digital health market: 84% of AI users had used an app or virtual care program in the past year, against 42% of non-users. OpenAI says [over 230 million people a week](https://openai.com/index/introducing-chatgpt-health/) ask ChatGPT health and wellness questions.

**Purchasers.** PHTI found AI widely adopted inside health plans and health systems, mostly for administrative work such as clinical documentation (71% of health systems report some deployment). That measures AI in operations, not in vendor research. No survey yet tells us how often a benefits leader asks ChatGPT which virtual physical therapy or diabetes program to shortlist.

The live article on [healthcare software and health system deals](https://underneath.agency/resources/healthcare-software-ai-search) covers vendors selling IT into hospitals. This article is about the startup that has to win a purchaser and then the patient. Platforms selling virtual care to health systems and plans have their own guide on [winning virtual care contracts and visits](https://underneath.agency/resources/virtual-care-platforms-ai-search).

## What do purchasers and patients ask AI about healthcare startups?

Category, alternatives, evidence and legitimacy questions. We wrote these examples ourselves; none were collected from real users.

| Who asks | Illustrative question |
|---|---|
| Benefits leader | “Which virtual physical therapy programs have independent evidence of cost savings?” |
| Health plan | “What are the alternatives to Omada for diabetes management?” |
| Benefits consultant | “Which digital mental health vendors offer performance guarantees?” |
| Eligible member | “Is the back pain program my employer offers any good?” |
| Patient | “Are online menopause clinics legitimate, and how do they compare?” |
| Investor | “Which women’s health startups have published clinical outcomes?” |

A single question can trigger several searches. Google says AI Overviews and AI Mode [may use a “query fan-out” technique](https://developers.google.com/search/docs/appearance/ai-features), and OpenAI says [ChatGPT search rewrites a prompt](https://help.openai.com/en/articles/9237897-chatgpt-search) into targeted queries. For “is this program any good,” we would expect those searches to reach for reviews, evaluations and news. Our study of [how AI judges whether a business is legitimate](https://underneath.agency/research/is-it-legit-ai-reputation-study) looks at that kind of question directly.

Startups must keep their own content within limits: general information, no diagnosis, no personal medical advice, and claims that match the evidence.

## How does an AI mention become a contract or an enrollment?

Through two paths: purchaser shortlists that lead to contracts, and member searches that lead to enrollment.

As we read the evidence, the two paths look like this:

**Purchaser path.**
1. A benefits leader, consultant or plan executive asks about a category or alternatives to an incumbent.
2. The answer names vendors and cites evaluations, news and company pages.
3. The buyer requests evidence; independent reviews like PHTI’s matter here, because PHTI says digital health companies’ own return-on-investment estimates “vary in methodological rigor and reliability.”
4. A contract follows, increasingly with fees at risk on outcomes.

**Member path.**
1. An eligible employee or plan member has a condition and asks an assistant about options, or about the program in their benefits.
2. If the answer describes your service accurately and points to it, enrollment becomes more likely, we infer.
3. Enrollment and engagement drive the outcomes that performance-based fees and renewals depend on.

The second path is the one most startups overlook. A contract covering thousands of eligible employees is worth little if an AI answer tells them a competitor is the better-known option, or gets your eligibility rules wrong. Neither path is easy to trace in analytics; we cover the measurement problem in [how to prove GEO caused sales](https://underneath.agency/resources/prove-geo-caused-sales).

## Why do AI assistants name the established health brand and skip the startup?

Selection rules are undisclosed; studies show new companies are rarely surfaced and independent, institutional sources dominate health answers.

**Documented by the platforms.** Google and OpenAI confirm their AI answers search the web and link sources. Neither discloses why one health company is named over another.

**Observed in studies.**

- The Product Hunt study found the gap between being known and being found: Perplexity recognized startups by name 94.3% of the time but surfaced them in 8.29% of discovery questions. The study found referring links and community presence related to Perplexity visibility.
- In ChatGPT’s answers to consumer health questions, [Jacques and colleagues](https://arxiv.org/abs/2601.17109) found 75.7% of cited sources came from institutions such as medical centers, government agencies and Wikipedia. Commercial platforms that were cited usually stated a medical review (71.1%). We summarize that study in [what commercial health sites cited by ChatGPT have in common](https://underneath.agency/resources/chatgpt-cited-commercial-health-sites).
- In [our study of brand entities](https://underneath.agency/research/brand-entity-ai-recommendations-study), independent coverage was the strongest predictor we measured: each tenfold increase in independent sites naming a brand in the cited pages went with 4.7 times the odds of being recommended.

**What health buyers check.** Evidence first. PHTI reported in 2023 that 80% of digital health products lack clinical evidence. Its later work, reviewed by the [Society of Actuaries](https://phti.org/?p=2234), found digital diabetes solutions generally increase spending while virtual musculoskeletal solutions can replace in-person care at lower cost. Independent verdicts like these are public, specific and citable.

**Our inference.** Established brands enjoy years of coverage, reviews and citations. A startup closes that gap by giving the assistant what a skeptical buyer wants: published outcomes, independent evaluations, named clinicians, and coverage by outlets with no stake in the sale. No study has yet tested this for healthcare startups specifically.

## What does it cost a healthcare startup to be invisible in AI answers?

Mostly lost shortlists and weak enrollment, with funding consequences; no study has put a number on it.

- **Shortlists are shrinking.** With employers consolidating vendors, a startup not named in early research may not make the reduced list.
- **Enrollment pays the bills.** If fewer than a quarter of eligible members enroll for nearly half of purchasers, a startup that members cannot find or understand in AI answers loses performance fees and renewals, we infer.
- **Investors notice traction.** In a market where 35% of deals were flat or down rounds, weak demand shows up in the next raise.
- **Misinformation spreads.** Wrong pricing, coverage or eligibility details can appear in answers; see [fixing wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers).

## Which GEO work helps a young health company most?

Public evidence, independent validation and accurate, consistent facts that both purchasers and patients can find.

1. **One clear identity.** State what you treat, for whom, through which purchasers, and in which states, consistently across your site, partner pages, benefits marketplaces and company databases.
2. **An evidence page.** Publish peer-reviewed results, study designs and independent evaluations in plain language, with links. Avoid unsupported savings claims; purchasers and regulators read them too.
3. **Visible clinical oversight.** Name your medical leadership and state how content is reviewed; the cited commercial health sites in the Jacques study showed this signal.
4. **Independent coverage.** Trade press, peer-reviewed papers, evaluations and conference talks give assistants third-party confirmation. The method is in [how brands build authority for AI search](https://underneath.agency/resources/how-brands-build-authority-for-ai-search); [how small brands get recommended by AI](https://underneath.agency/resources/how-small-brands-get-recommended-by-ai) covers the startup case.
5. **Member-facing explainers.** Write accurate pages on the conditions you treat, how to enroll through an employer or plan, and what it costs, without medical advice.
6. **Plan for the newness gap.** Assistants often miss recent launches; see [why ChatGPT misses new products](https://underneath.agency/resources/why-chatgpt-misses-new-products). Then track purchaser and patient questions across ChatGPT, Gemini, Perplexity, Claude, Copilot and Google’s AI features, repeatedly.

Planted reviews and inflated outcomes are a legal and reputational hazard in health; see [legitimate GEO versus manipulation](https://underneath.agency/resources/legitimate-geo-vs-manipulation). Broader startup lessons are in [how AI startups win customers from AI search](https://underneath.agency/resources/ai-startups-customers-from-ai-search).

## How solid is the case that AI visibility builds healthcare startup demand?

The supporting trends are well documented; the direct link from AI visibility to contracts or enrollment is not.

- **No purchaser data.** No survey measures how often benefits leaders use AI assistants to shortlist digital health vendors.
- **The startup study is general.** The Product Hunt sample covers technology products, not healthcare companies.
- **Health citation research covers medical questions.** Purchaser questions may draw on different sources.
- **Several sources have interests.** Solera sells navigation services; Hinge reports its own metrics; Rock Health invests in digital health.
- **No one has tied AI answers to signed health contracts.** The wider evidence on revenue is thin as well; [our review of whether AI visibility drives business results](https://underneath.agency/resources/does-ai-visibility-drive-business-results) sets out what is known.

## What should a healthcare startup do first to show up in AI answers?

Ask what your purchasers and patients ask, see who is named, then publish the evidence they cannot find.

List the questions a benefits leader, a plan executive, an eligible member and a patient would ask about your category, your competitors and your company by name. Pose each one to more than one assistant, and on more than one day, because a single answer can flatter or overlook a young company. Record which companies appear, which evaluations and articles are cited, and whether your outcomes, coverage and eligibility come back correctly. For most young health companies, the gap is evidence that exists but is not public, or is public only in a sales deck.

When you are ready, [ask us for a startup AI visibility check](https://underneath.agency/contact). It shows where AI answers place you against established brands for purchaser and patient questions, which sources they trust, and which missing evidence is most likely costing you shortlists, member enrollment and renewals. How a young health company’s evidence, coverage and listings are then built up, step by step, is set out on our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page.

## Frequently asked questions

### Do employers use ChatGPT to choose digital health vendors?

No survey measures it yet. PHTI’s 2026 survey shows purchasers focus on outcomes, with most employers using performance-based contracts.

### Why does ChatGPT recommend the big brand instead of us?

Big brands have more independent coverage, which our brand study found was the strongest predictor of being recommended. New products are also often missing from answers.

### Should we publish our clinical outcomes?

Yes, where you can do so accurately. Independent and peer-reviewed results are what both purchasers and AI answers can cite.

### Can AI visibility improve member enrollment?

Plausibly, by helping eligible members understand and trust the program, but it is untested. Enrollment below a quarter of eligible members is common, PHTI found.

### Is GEO worth it before product-market fit?

Usually not as a priority. It pays most once you have evidence to show and a purchaser channel to support.

## Sources

- Peterson Health Technology Institute (2026-09-29), [2026 State of Digital Health Purchasing](https://www.phti.org/research-insight/2026-state-of-digital-health-purchasing/)
- Peterson Health Technology Institute (2023), [Digital Health Tools for Diabetes Management and Virtual Musculoskeletal Care to Undergo Independent Evaluation](https://www.phti.org/?p=762)
- Peterson Health Technology Institute (2024), [Society of Actuaries Report Validates PHTI Economic Impact Analysis](https://phti.org/?p=2234)
- Healthcare Dive (2026-01-13), [Digital health funding increases in 2025, spurred by AI: report](https://www.healthcaredive.com/news/digital-health-funding-2025-boosted-ai-rock-health/809449/)
- Fierce Healthcare (2026-01-12), [JPM26: Digital health funding hit $14.2B in 2025, with AI companies taking the lion’s share of dollars](https://www.fiercehealthcare.com/digital-health/jpm26-digital-health-funding-hit-142b-2025-ai-companies-taking-lions-share-dollars)
- D Magazine (2026-07-30), [Baylor Scott & White is betting employers want fewer digital health vendors](https://www.dmagazine.com/healthcare-business/2026/07/baylor-scott-white-is-betting-employers-want-fewer-digital-health-vendors/)
- Hinge Health (2026-02-10), [Hinge Health reports fourth quarter and full year 2025 financial results](https://s205.q4cdn.com/978984498/files/doc_news/Hinge-Health-reports-fourth-quarter-and-full-year-2025-financial-results-2026.pdf)
- Health Populi (2026-03-24), [Consumer adoption of AI for health and self-care doubling, via Rock Health’s latest snapshot](https://www.healthpopuli.com/2026/03/24/consumer-adoption-of-ai-for-health-and-self-care-doubling-to-36-in-a-year-via-rock-healths-latest-snapshot)
- KFF (2026), [KFF Tracking Poll on Health Information and Trust: Use of AI for Health Information and Advice](https://www.kff.org/health-information-and-trust/poll-finding/kff-tracking-poll-on-health-information-and-trust-use-of-ai-for-health-information-and-advice/)
- OpenAI (2026-01-07), [Introducing ChatGPT Health](https://openai.com/index/introducing-chatgpt-health/)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- OpenAI Help Center (2025), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Sharma (2025), [The Discovery Gap: How Product Hunt Startups Vanish in LLM Organic Discovery Queries](https://arxiv.org/abs/2601.00912), arXiv:2601.00912.
- Jacques and colleagues (2026), [Authority Signals in AI Cited Health Sources: A Framework for Evaluating Source Credibility in ChatGPT Responses](https://arxiv.org/abs/2601.17109), arXiv:2601.17109.
- Underneath (2026), [Do Wikipedia and schema make AI assistants recommend a brand?](https://underneath.agency/research/brand-entity-ai-recommendations-study)
- Underneath (2026), [“Is this brand legit?” How AI assistants build a reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study)

---

This is the Markdown twin of https://underneath.agency/resources/healthcare-startups-demand-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How healthtech firms win employer and health plan buyers via AI"
description: "With independent evidence, plain outcome and cost facts, and coverage buyers trust. Employers are cutting vendors, and AI is an early place they look."
canonical: "https://underneath.agency/resources/healthtech-employers-payers-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How does a healthtech company get named when employers and health plans ask AI for vendors?

By publishing evidence that independent sources can confirm, stating outcomes, pricing terms and compliance plainly, and earning coverage where benefits and health plan buyers look. Employers facing their steepest cost increases in years are reviewing every health vendor, and buyers in general now use AI assistants to gather information on vendors. For a company that sells virtual care, chronic-condition or navigation programs to employers and health plans, the shortlist increasingly starts with an answer it does not write.

## The short version

1. Employers are re-tendering: in [Business Group on Health’s 2026 survey](https://www.businessgrouphealth.org/newsroom/news-and-press-releases/press-releases/2026-employer-health-care-strategy-survey) of 121 large employers, 51% were changing or running a request for proposal (RFP) for health and well-being vendors, as they projected a median 9% cost increase.
2. Purchasers keep buying but demand proof: in a [Peterson Health Technology Institute survey](https://employercoverage.substack.com/p/survey-shows-growing-interest-in) of 332 health plans, employers and provider organizations, 75% were spending more on digital health than two years earlier, and employers evaluated vendors on health outcomes (99%) and costs (93%).
3. Contracts come up for renewal often: most digital health contracts in that survey ran two years or less, and 59% of employers contracted directly with vendors.
4. Independent evaluation shapes categories: the institute found that physical therapist-guided virtual programs for back pain could [save an estimated $4.4 million per 1 million](https://techtarget.com/virtualhealthcare/answer/PHTI-validates-virtual-MSK-solutions-after-taking-diabetes-tools-to-task) commercially insured people, after finding earlier that many diabetes tools did not deliver clinical benefits.
5. The winners are growing fast: [Hinge Health](https://s205.q4cdn.com/978984498/files/doc_news/Hinge-Health-reports-record-second-quarter-2026-financial-results-signs-definitive-agreement-to-acquire-Cylinder-Health-2026.pdf) reported 2,929 clients in mid-2026, up 24% in a year, and second-quarter revenue up 53%.

A note before you read: this article is about how healthtech companies are found and described by AI assistants. It makes no health or clinical claims, and nothing here is medical, legal or regulatory advice. Consumer health apps sold directly to people are a different market; hospital IT is covered in [our healthcare software article](https://underneath.agency/resources/healthcare-software-ai-search).

## Who buys healthtech, and what is a customer worth?

Employers, health plans and provider organizations buy the contract; their employees or members then have to enroll.

The purchaser survey from the Peterson Health Technology Institute (PHTI), summarized by Jeff Levin-Scherz in his Employer Coverage newsletter, covered 115 health plans, 117 employers and 100 health care delivery organizations. Among employers, more than half bought digital programs for diabetes (78%), obesity (63%) and mental health (57%). Direct contracting was the most common route for health plans (63%) and employers (59%), followed by contracting through a pharmacy benefit manager (PBM) at 30%.

That gives most healthtech companies two sales. The first is the contract with the employer or plan, often through a benefits consultant or a channel partner. The second is each employee or member who signs up, because revenue usually follows enrollment.

Public companies show what a customer is worth when both sales work:

- **Hinge Health**, a virtual musculoskeletal care company, had 2,929 clients on June 30, 2026, against 2,359 a year earlier, and second-quarter revenue of $212.8 million. It counts organizations that buy through its partners as individual clients, a sign of how much reach channel partners provide.
- **Omada Health**, a cardiometabolic care company, reported [second-quarter revenue of $88 million](https://www.biopharmawatch.com/news/OMDA/omada-health-q2-2026-results-ceo-transition), up 43%, total members up 45%, and trailing twelve-month revenue per member of $284.

One employer contract can therefore bring thousands of members, each worth a recurring fee for as long as the program is renewed. Platforms that also sell virtual care to health systems are covered in [how virtual care platforms win contracts](https://underneath.agency/resources/virtual-care-platforms-ai-search).

## Why are employers and health plans re-examining vendors now?

Because costs are rising fast, and buyers are cutting programs that cannot prove their value.

Business Group on Health’s survey covered employers with 11.6 million people on their plans. They projected a median 9% cost increase for 2026, cut to 7.6% after plan design changes, and the group’s chief executive said employers would be “rigorously evaluating benefit offerings, vendor performance and patient outcomes.” Besides the 51% reviewing health and well-being vendors, 41% were changing pharmacy benefit managers or running an RFP.

Evidence now separates categories. PHTI reported that virtual musculoskeletal programs improved pain and function, with physical therapist-guided programs saving money. Its earlier review of digital diabetes tools found that many did not provide clinical benefits. In the newsletter’s summary of the purchaser survey, only 47% of buyers who increased spending said digital health had decreased their costs, and its author warned that vendor claims of savings might not hold up in practice.

Our inference: in a market where buyers check claims against independent evaluations, what assistants say about a category and its evidence will shape who makes the longlist. Younger companies face that check with less coverage behind them, as our guide on [how healthcare startups compete with big brands](https://underneath.agency/resources/healthcare-startups-demand-ai-search) explains.

## Where does AI search enter a healthtech purchase?

Early, at the research and longlist stage, though no public survey yet measures benefits buyers specifically.

Across business purchases, a [Gartner survey of 645 B2B buyers](https://www.gartner.com/en/newsroom/press-releases/2026-05-20-gartner-survey-finds-sixty-nine-percent-of-b-two-b-buyers-turn-to-sales-reps-to-validate-ai-generated-insights) found 45% had used generative AI during a recent purchase, primarily to gather information on vendors and products, and 69% preferred to validate AI-generated insights with sales reps. We found no comparable figure for benefits leaders, consultants or health plan buyers, and we do not assume one.

A reasonable expectation is that AI enters at three points:

1. **Category research.** A benefits manager asked to “look at virtual MSK” or “find a GLP-1 support program” starts by asking what the options are.
2. **Longlist and RFP drafting.** Consultants and procurement teams compare vendors on evidence, guarantees and integration.
3. **Member questions.** After launch, employees ask whether a program is legitimate, private and free through their employer. In [our study of “is it legit?” questions](https://underneath.agency/research/is-it-legit-ai-reputation-study), 88.0% of AI answers cited a review or complaint platform.

The third point is easy to miss. A contract won in the RFP still depends on enrollment, and enrollment depends partly on what members hear when they check.

## What do employers, plans and members ask AI about healthtech vendors?

Questions about evidence, savings, guarantees, privacy and fit. We drafted the prompts below to show typical questions from benefits teams, plans and members; none was collected from a real buyer.

| Buyer | Illustrative prompt |
|---|---|
| Benefits leader | “Which virtual physical therapy programs have independent evidence of lower costs for self-insured employers?” |
| Benefits consultant | “Compare Hinge Health and Sword Health on outcomes, pricing model and guarantees” |
| Health plan | “Digital diabetes programs with peer-reviewed results and fees tied to outcomes” |
| Employer, GLP-1 costs | “Weight management programs that work alongside GLP-1 coverage, with outcome guarantees” |
| Provider organization | “Remote monitoring vendors for hypertension that integrate with Epic” |
| Member | “Is [program] really free through my employer, and who sees my health data?” |

Notice how many ask for proof: independent evidence, peer-reviewed results, guarantees. A vendor whose evidence exists only in a sales deck gives an assistant nothing to repeat.

## How does an AI mention turn into contract value and enrollment?

By putting the vendor on the longlist, then supporting the evidence review, launch and every renewal after it.

The path runs: an AI answer during category research → the vendor appears on the consultant’s or employer’s longlist → evidence and savings review → a pilot or contract with risk-based terms → launch → employee enrollment → renewal. Because most contracts run two years or less, the evaluation repeats often, and each renewal is another moment when a buyer may ask what the market looks like now.

Being missing has a cost, though no one has measured it in dollars. Our inference: a vendor left out of category answers loses a place on longlists during a year when 51% of large employers are reviewing well-being vendors, and a vendor described with outdated or unsupported claims risks failing the evidence review that follows. The broader pattern of how early AI answers shape shortlists is covered in [are AI assistants now shaping which enterprise software gets shortlisted?](https://underneath.agency/resources/enterprise-software-shortlists-ai-search)

## What decides whether an assistant names a healthtech vendor?

Platforms do not document how vendors are chosen; studies point to independent coverage and evidence.

**Documented by platforms.** No assistant publishes how it selects vendors in a category. We found nothing specific to healthtech.

**Observed in a study.** In [our brand study](https://underneath.agency/research/brand-entity-ai-recommendations-study), independent coverage was the strongest predictor we measured: each tenfold increase in independent sites naming a brand in the cited pages went with 4.7 times the odds of being recommended. That study covered consumer brands, not health services.

**Our inference for healthtech.** Buyers in this market already trust independent evaluators, peer-reviewed studies, consultants and trade coverage more than vendor claims. We expect assistants to reflect the same sources when they describe a category, which means:

- **Independent evaluations travel.** A PHTI assessment, a peer-reviewed trial or a published validation is the kind of source an assistant can cite when asked which programs work.
- **Category position matters.** PHTI evaluated eight virtual musculoskeletal solutions in one report. If an evaluator groups you in a category, that grouping may shape how you are described.
- **Unsupported claims are a risk.** Our article on [how often AI answers say things their sources do not support](https://underneath.agency/resources/ai-answers-unsupported-claims) shows how easily a claim drifts from its evidence; precise, sourced statements are harder to distort.

## How does GEO work for a healthtech company?

Generative engine optimization (GEO) for healthtech means making your evidence, terms and fit easy to find, verify and quote.

1. **Evidence pages for buyers.** Summaries of each study with a link, who ran it, the population, the measure and its limits. Name independent evaluations, including unfavorable ones you have answered.
2. **Savings and guarantee terms in plain words.** How fees work, what is guaranteed, how savings are calculated. Buyers ask for this; vague claims invite doubt.
3. **Category explainers.** An honest guide to the category, how to compare vendors and what evidence to ask for. These pages can be cited even when the question does not name you.
4. **Security and privacy facts.** Certifications, data handling and what members’ employers can and cannot see, on public pages.
5. **Channel and partner listings.** Health plan, PBM and benefits platform listings that describe you accurately and consistently.
6. **Coverage where buyers read.** Benefits and health plan trade press, consultant briefings, conference talks and peer-reviewed publications, so independent sources describe you in your category.
7. **Member-facing pages and reviews.** Eligibility, cost to the member and privacy, written for employees, plus reviews and complaint responses handled well.
8. **Regular checks.** Ask employer, consultant, plan and member questions across assistants before RFP season and open enrollment, and track who is named and which sources are cited.

The approach overlaps with [how brands build authority for AI search](https://underneath.agency/resources/how-brands-build-authority-for-ai-search), with one difference: every claim must survive a clinical and actuarial review. GEO cannot guarantee that an assistant names any vendor; it makes your evidence easier to find and harder to misstate.

## What does no one know yet about AI in healthtech buying?

Whether AI answers change which vendors win contracts; the evidence so far is indirect.

- **No benefits-buyer data.** The AI-use figures come from cross-industry B2B surveys, not from employers, consultants or health plans.
- **Older purchaser data.** The PHTI purchaser survey was published in October 2024.
- **Consumer-brand studies.** Our brand study tested consumer products, not health services sold to employers.
- **No contract link.** We found no public data connecting AI visibility to RFP invitations, contracts or enrollment.
- **Company figures are self-reported.** Client and member counts come from company releases.

## Where should a healthtech company start?

Start by asking assistants the category, comparison and member questions your buyers ask before the next RFP cycle.

List the questions a benefits leader, a consultant, a health plan and a member would ask about your category. Put each question to ChatGPT, Gemini, Perplexity and Google AI Mode, the way a benefits consultant building a vendor shortlist might. Note which vendors are named, which evidence is cited, whether your outcomes and terms are described correctly, and what members hear when they ask if you are legitimate.

If your growth depends on employer contracts, plan partnerships and enrollment, [ask us to test what AI tells your buyers and members](https://underneath.agency/contact). We will map how assistants describe your category and your evidence, find the gaps and errors, and plan the evidence, coverage and listing work that supports your next RFP and renewal season. For the full scope, our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page explains how that program is diagnosed, carried out and measured for vendors selling to employers and health plans.

## Frequently asked questions

### Do employers use ChatGPT to choose health benefit vendors?

No public survey measures this yet. Across B2B purchases, 45% of buyers in a Gartner survey used generative AI, mainly to gather information on vendors and products.

### Does an independent evaluation help a healthtech company in AI answers?

Probably, though it is not proven. Independent sources are what buyers trust, and in our brand study independent coverage was the strongest predictor of being recommended.

### Should we publish our pricing and guarantees?

Publish how fees and guarantees work, at least in outline. Buyers ask assistants about them, and 93% of employers in one survey evaluated vendors on costs.

### Do member reviews matter if we sell to employers?

Yes, because enrollment drives revenue. When people ask AI whether a company is legitimate, 88.0% of answers in our study cited a review or complaint platform.

## Sources

- Business Group on Health (2025-08-19), [Business Group on Health Survey: 9% Health Care Cost Increase for 2026](https://www.businessgrouphealth.org/newsroom/news-and-press-releases/press-releases/2026-employer-health-care-strategy-survey)
- Employer Coverage, Jeff Levin-Scherz (2024-10), [Survey shows growing interest in digital health](https://employercoverage.substack.com/p/survey-shows-growing-interest-in)
- TechTarget (2024-06), [PHTI validates virtual MSK solutions after taking diabetes tools to task](https://techtarget.com/virtualhealthcare/answer/PHTI-validates-virtual-MSK-solutions-after-taking-diabetes-tools-to-task)
- Hinge Health (2026-08-04), [Hinge Health reports record second quarter 2026 financial results](https://s205.q4cdn.com/978984498/files/doc_news/Hinge-Health-reports-record-second-quarter-2026-financial-results-signs-definitive-agreement-to-acquire-Cylinder-Health-2026.pdf)
- Omada Health, via BioPharmaWatch (2026-08), [Omada Health Q2 2026 results](https://www.biopharmawatch.com/news/OMDA/omada-health-q2-2026-results-ceo-transition)
- Gartner (2026-05-20), [Gartner Survey Finds 69% of B2B Buyers Turn to Sales Reps to Validate AI-Generated Insights](https://www.gartner.com/en/newsroom/press-releases/2026-05-20-gartner-survey-finds-sixty-nine-percent-of-b-two-b-buyers-turn-to-sales-reps-to-validate-ai-generated-insights)
- Underneath (2026), [Do Wikipedia and schema make AI assistants recommend a brand?](https://underneath.agency/research/brand-entity-ai-recommendations-study)
- Underneath (2026), [“Is this brand legit?” How AI assistants build a reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study)

---

This is the Markdown twin of https://underneath.agency/resources/healthtech-employers-payers-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How home decor brands get found through AI product discovery"
description: "By showing up in the visual product sets AI tools build from style questions, with images, colors, materials and reviews an assistant can read."
canonical: "https://underneath.agency/resources/home-decor-brands-ai-product-discovery"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How do home decor brands get found when shoppers ask AI for a look, not a product?

By being one of the products an AI tool shows when a shopper describes a style, a room or a color, and by giving that tool the images, colors, materials and reviews it needs to match your product to the request. Decor shoppers rarely ask for a brand. They ask for a “vibe,” and AI search is now built to turn that vibe into a set of products to buy.

## The short version

1. Decor shoppers find new sites through search and other retailers. In a March 2025 survey of 300 US home and decor shoppers by Wunderkind and MX8 Labs, search engines drove 39% of first visits to a decor website, retail stores like IKEA 26% and marketplaces like Wayfair 23%.
2. Decor is bought often. In the same survey, 15% of shoppers bought from home and decor brand websites weekly, including 28% of Gen Z.
3. The discovery platforms decor depends on are adding AI. Pinterest, which counted 640 million monthly users in mid-2026, launched an AI assistant built for requests like throw pillows that match a living room.
4. Google built visual shopping into AI Mode around examples like maximalist bedroom design, drawing on more than 50 billion product listings.
5. Checkout moved toward the answer: Etsy, whose Home and Living category grew in late 2025, announced an integration with ChatGPT Instant Checkout and agentic shopping partnerships with Google and Microsoft. OpenAI scaled Instant Checkout back in March 2026, according to [FashionUnited](https://fashionunited.com/news/retail/us-consumers-would-let-ai-agents-buy-clothes-but-not-without-a-say/2026092974919), though its help page still describes it for some eligible merchants.

## Who buys home decor, and how often?

Style-led shoppers who browse for inspiration and buy small items often, especially younger ones.

Decor is not furniture. The ticket is smaller, the purchase is more frequent and the decision is about taste more than specifications. Furniture is a larger, more considered sale, covered in [how furniture brands win high-ticket orders](https://underneath.agency/resources/furniture-brands-sales-from-ai-search). The [Wunderkind Home and Décor Consumer Insights Report](https://hello.wunderkind.co/hubfs/PDFS%20-%20Content/2025-Home-and-Decor-Consumer-Insights-Report.pdf?hsLang=en), based on a survey of 300 US shoppers fielded with MX8 Labs, shows the pattern. Weekly purchases from home and decor brand websites were highest among Gen Z (28%) and millennials (25%), against 1% of boomers, and nearly half of boomers (49%) bought only once or twice a year.

Browsing without buying is part of the category. In the survey, 31% of shoppers browsed decor sites looking for design inspiration, and 23% bookmarked or saved items for later. The report, written by a marketing vendor, says only 3% of visitors complete a purchase in a single session. For a decor brand, the customer’s value comes from repeat baskets of pillows, throws, candles and frames, not one large order.

## Where do decor shoppers discover new brands today?

Mostly through search engines, other retailers, marketplaces and social platforms, with AI now built into each.

The Wunderkind survey asked how shoppers first find a home and decor website. Search engines led at 39%, followed by links from online retail stores like IKEA or Home Depot (26%), ads (24%) and marketplaces like Amazon or Wayfair (23%). Social media posts and influencer recommendations brought 16% overall, but 29% of Gen Z. Reading about a brand in an article, blog or review brought 12%. Every one of those doors now has an AI layer.

**Search.** Google rebuilt shopping in AI Mode around visual, style-led questions. Its [announcement of visual search in AI Mode](https://blog.google/products/search/search-ai-updates-september-2025/) uses a decor example: “Let’s say you’re searching for maximalist design inspiration for your bedroom.” The shopper can follow up by asking for “dark tones and bold prints,” start from a photo, and click through to buy from listings drawn from Google’s Shopping Graph of more than 50 billion product listings, with more than 2 billion refreshed every hour.

**Social discovery.** Pinterest, a primary home for room inspiration, launched [Pinterest Assistant](https://channelx.world/2025/10/new-pinterest-assistant-ai-feature-for-enhanced-discovery-and-shopping/) in late 2025. Its launch example: “I need new throw pillows that match my living room decor,” answered from the shopper’s saves and boards and from people with similar taste. Pinterest’s [second quarter 2026 results](https://s204.q4cdn.com/369458543/files/doc_earnings/2026/q2/earnings-result/Q2-2026-Press-Release.pdf) reported 640 million monthly active users and revenue of $1,180 million, up 18%, and its CEO said “AI is at the heart of our momentum.”

**Marketplaces.** Etsy, where handmade and vintage decor is a core category, reported in its [fourth quarter 2025 results](https://investors.etsy.com/_assets/_fea57334fe81a735b34cf3bf4bdcb553/etsy/db/938/10062/earnings_release/Exhibit+99.1+12.31.2025.pdf) that its Home and Living category grew year over year, with 86.5 million active buyers on the Etsy marketplace.

## Which questions lead a decor shopper to a product?

Questions about a style, a room, a color match or a gift, rarely a brand name.

The examples below are ours, written to show how decor questions are usually framed; they are illustrative, not logged queries:

- Style: “Warm minimalist living room decor ideas under $300.”
- Color match: “Throw pillows that go with a rust velvet sofa and a cream rug.”
- Constraint: “Renter-friendly wall decor that doesn’t need nails.”
- Gift and season: “Unique handmade housewarming gifts” or “Japandi holiday decor.”
- Comparison and trust: “Best washable rugs for a home with pets” or “Is this brand’s bedding good quality?”

Decor questions differ from most product questions in one way: the shopper often cannot name what they want until they see it. Pinterest calls this the “I’ll know it when I see it” problem. That makes images and descriptive attributes, not product names, the way a decor item gets matched to a request. An [academic study of brand recommendations](https://arxiv.org/abs/2609.16304) by Northwestern and Boston University researchers found that when a request spelled out the shopper’s goals and constraints, the brands an AI model retrieved changed, and brands left out of ordinary answers could surface when distinctive cues matched their positioning. Fashion brands meet the same look-first questions, covered in [how fashion brands get recommended by AI](https://underneath.agency/resources/fashion-ecommerce-ai-recommendations).

## How does AI discovery turn into revenue for a decor brand?

Through inclusion in a visual set of options, then a click or an in-chat checkout, then repeat baskets.

**Inclusion.** A decor answer is usually a set of images and listings, not a single pick. Being one of eight pillows shown for “rust velvet sofa” is the equivalent of a shelf placement.

**The click or the checkout.** When it launched Instant Checkout in September 2025, OpenAI said [more than 700 million people use ChatGPT each week](https://openai.com/index/buy-it-in-chatgpt/) and that US users could buy directly from US Etsy sellers in chat. FashionUnited reports OpenAI scaled the feature back in March 2026, though OpenAI’s help page still describes it for some eligible merchants. Etsy added partnerships with Google and Microsoft in January 2026 so that signed-in US users can buy select Etsy items inside AI Mode, the Gemini app and Copilot Checkout, “at the moment shoppers move from inspiration to intent.”

**The repeat basket.** Because younger decor shoppers buy weekly or monthly, a brand that is found once through an AI answer can be bought from again directly. A reasonable expectation is that the first AI-led discovery matters more than its first order value.

The evidence on conversion is mixed, and decor executives should know it. When [Adobe](https://blog.adobe.com/en/publish/2025/03/17/adobe-analytics-traffic-to-us-retail-websites-from-generative-ai-sources-jumps-1200-percent) first measured visits from AI tools in early 2025, conversion rates from that traffic were highest in electronics and jewelry and lowest in apparel, home goods and grocery. In the [2025 holiday season](https://news.adobe.com/news/2026/01/adobe-holiday-shopping-season), AI traffic to US retail sites rose 693.4%, but AI was used most for video games, toys, appliances, electronics and personal care. Holiday decor sales rose 285% against pre-season levels, a seasonal surge that Adobe did not attribute to AI. We infer that AI already shapes decor discovery more than it closes decor sales.

## What decides whether an AI tool shows your decor products?

Product data and images, independent mentions and reviews; platforms document some of it, studies show the rest.

**Documented by the platforms.** OpenAI says [ChatGPT selects shopping results](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search) using structured metadata from first-party and third-party providers and other third-party content, and may show summaries of reviews from public websites. Its [product feed specification](https://developers.openai.com/commerce/specs/feed) includes fields for color (“consistent with the product image”), pattern and material, the attributes a decor match depends on. Its Instant Checkout announcement added that items sold through that checkout were not preferred in product results. Google’s [guidance on AI features](https://developers.google.com/search/docs/appearance/ai-features) says there are no additional requirements or special optimizations needed to appear in AI Overviews or AI Mode, and recommends supporting text with high-quality images, keeping structured data consistent with the visible page and keeping Merchant Center information up to date.

**Observed in studies.** The Northwestern study found that how prominently a brand was recommended was associated with broader marketplace visibility, particularly search interest and online brand conversation, more than with conventional brand popularity. In [our study of “best of” lists](https://underneath.agency/research/self-promoting-best-lists-study) cited by AI engines, 65 of 269 lists (24.2%) ranked their own publisher first, so retailer-written roundups of “best throw blankets” often favor the retailer’s own goods. And answers move: in [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), a single ChatGPT answer showed 57.8% of the brands its five answers named between them.

**Trust factors specific to decor.** We infer that what helps an AI tool match a decor item is what helps a shopper picture it in their room: accurate color names, material and texture, real dimensions, several photos in a styled room, and reviews that mention color accuracy and quality. Editorial coverage in gift guides and design publications matters because decor shoppers already find brands through articles and reviews. Sportswear brands make the same bet on exact product attributes, matched to an activity rather than a room, as our guide on [how sportswear brands win AI shoppers](https://underneath.agency/resources/sportswear-brands-ai-search) shows.

## What does a decor brand lose if AI tools skip it?

The early, style-led moment when a shopper decides what the room should look like.

The cost is not easy to measure, and we found no study that counts decor sales lost to AI answers. What we can say is where the exposure lies. Search engines, marketplaces and social platforms brought most first visits in the Wunderkind survey, and each is now inserting an AI layer between the shopper’s question and the products shown. A brand whose products lack clear color, material and style data gives that layer less to match. A brand sold mainly through marketplaces may be shown under the marketplace’s name, not its own. Independent decor makers face a further handicap, covered in [whether AI assistants favor big brands](https://underneath.agency/resources/do-ai-assistants-favor-big-brands): well-known names tend to win when products look alike.

## How does GEO work for a home decor brand?

By making each product easy to match to a style request and easy to trust, without promising placement.

1. **Attribute-rich product data.** Color names a shopper would use, pattern, material, texture, dimensions and care, in product pages, structured data and feeds to Google, OpenAI and the marketplaces you sell on.
2. **Images with words.** Several photos per item, including styled rooms, with text on the page that describes what the image shows, the kind of copy our review of [the product content AI shopping assistants prefer](https://underneath.agency/resources/product-content-ai-shopping-assistants-prefer) favors.
3. **Style and room content.** Guides that answer style questions in your own voice: how to layer textiles, which pillows go with a green sofa, small-space decor.
4. **Independent coverage.** Gift guides, design editors and creators who show your products in real rooms. “Best of” roundups such as holiday gift lists are cited often, so it pays to know [which best-of lists AI answers draw on](https://underneath.agency/resources/best-of-lists-ai-recommendations).
5. **Presence where AI shopping runs.** Accurate listings on Etsy, Pinterest, Google Merchant Center and the retailers that carry you, with consistent names and facts.
6. **Measurement over many runs.** A tracked set of style, room and gift prompts asked repeatedly, because [AI answers about a brand change from run to run](https://underneath.agency/resources/why-ai-answers-about-your-brand-change).

## What is still unknown about AI and decor sales?

How often AI answers lead to decor purchases, and which decor brands they show.

The decor survey we cite is small (300 shoppers) and comes from a marketing vendor. Adobe’s AI traffic data covers all retail, and its early conversion figures put home goods near the bottom. Pinterest’s and Google’s launch posts describe what their tools are designed to do, not how often shoppers use them for decor. IKEA put an AI assistant in the OpenAI GPT Store in [February 2024](https://www.retailtouchpoints.com/news/ikea-launches-gen-ai-powered-design-and-shopping-tool/140545), and we found no public data on what it sold. No study we found measures which decor products AI tools show for style questions or how stable those sets are. Treat any claim of guaranteed AI placement for decor with suspicion.

## Where should a home decor brand start?

With an audit of the style, room and gift questions your customers ask, across the AI tools they use.

Pick your best-selling categories and the style words your customers use. Check whether your products appear in AI Mode, ChatGPT and the marketplaces’ AI tools for those requests, which competitors and retailers appear instead, and whether your color, material and image data is complete enough to be matched. Then fix the data and coverage gaps that keep you out of the set. If you would like us to look with you, [start with a short note about your range](https://underneath.agency/contact) and we will show where your products stand when shoppers ask AI for a look, and what would bring more of that discovery back to your own store. Our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page walks through how that product data, imagery and coverage work is planned and checked for decor brands.

## Frequently asked questions

### Do decor shoppers use AI tools or social platforms for inspiration?

Both, and the two are merging. Pinterest added an AI assistant, and Google’s AI Mode now answers style questions with images and products.

### Does selling on Etsy or marketplaces help or hurt AI visibility?

It can help discovery, since Etsy announced ways to buy its listings inside ChatGPT, Google AI Mode and Copilot. OpenAI scaled back ChatGPT’s Instant Checkout in March 2026, according to FashionUnited, though its help page still describes it for some eligible merchants. The risk is that shoppers credit the marketplace, not your brand.

### What product data matters most for decor in AI answers?

Color, pattern, material, size and good images. OpenAI’s product feed has fields for color, pattern and material, and Google recommends high-quality images backed by text.

### Is AI traffic for home goods worth chasing yet?

It is growing but converts less well than in electronics, according to Adobe’s early data. The stronger case is discovery: being in the set shoppers see when they decide on a look.

## Sources

- Wunderkind and MX8 Labs (2025), [The Home and Décor Consumer Insights Report 2025](https://hello.wunderkind.co/hubfs/PDFS%20-%20Content/2025-Home-and-Decor-Consumer-Insights-Report.pdf?hsLang=en)
- Pinterest (2026-08-04), [Pinterest Announces Second Quarter 2026 Results](https://s204.q4cdn.com/369458543/files/doc_earnings/2026/q2/earnings-result/Q2-2026-Press-Release.pdf)
- ChannelX (2025-10-30), [New Pinterest Assistant AI feature for enhanced discovery and shopping](https://channelx.world/2025/10/new-pinterest-assistant-ai-feature-for-enhanced-discovery-and-shopping/)
- Google (2025-09-30), [AI Mode can now help you search and explore visually](https://blog.google/products/search/search-ai-updates-september-2025/)
- Google Search Central (2026), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- Etsy (2026-02-19), [Etsy, Inc. Reports Fourth Quarter and Full Year 2025 Results](https://investors.etsy.com/_assets/_fea57334fe81a735b34cf3bf4bdcb553/etsy/db/938/10062/earnings_release/Exhibit+99.1+12.31.2025.pdf)
- OpenAI (2025-09-29), [Buy it in ChatGPT: Instant Checkout and the Agentic Commerce Protocol](https://openai.com/index/buy-it-in-chatgpt/)
- FashionUnited (2026-09-29), [US consumers would let AI agents buy clothes, but not without a say](https://fashionunited.com/news/retail/us-consumers-would-let-ai-agents-buy-clothes-but-not-without-a-say/2026092974919)
- OpenAI (2026), [Shopping with ChatGPT Search](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search)
- OpenAI (2026), [Product Feed Spec](https://developers.openai.com/commerce/specs/feed)
- Adobe (2025-03-17), [Traffic to U.S. retail websites from generative AI sources jumps 1,200 percent](https://blog.adobe.com/en/publish/2025/03/17/adobe-analytics-traffic-to-us-retail-websites-from-generative-ai-sources-jumps-1200-percent)
- Adobe (2026-01-07), [Holiday Shopping Season Drove a Record $257.8 Billion Online](https://news.adobe.com/news/2026/01/adobe-holiday-shopping-season)
- Retail TouchPoints (2024-02-05), [Ikea Launches Gen AI-Powered Design and Shopping Tool](https://www.retailtouchpoints.com/news/ikea-launches-gen-ai-powered-design-and-shopping-tool/140545)
- Northwestern University and Boston University researchers (2026-09), [Evaluating Brand Retrieval and Ranking in Large Language Model Recommendations](https://arxiv.org/abs/2609.16304)
- Underneath (2026), [Self-promoting “best of” lists study](https://underneath.agency/research/self-promoting-best-lists-study)
- Underneath (2026), [AI recommendation consistency study](https://underneath.agency/research/ai-recommendation-consistency-study)

---

This is the Markdown twin of https://underneath.agency/resources/home-decor-brands-ai-product-discovery. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How do brands build authority that AI search recognizes?"
description: "By being written about by sources AI engines already trust, in the outlets each engine reads, and by keeping a site AI agents can read."
canonical: "https://underneath.agency/resources/how-brands-build-authority-for-ai-search"
published: 2026-10-07
updated: 2026-10-07
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How do brands build authority that AI search recognizes?

Brands build AI search authority mostly by being covered by independent sources the engines already trust, in the outlets each engine actually reads. Your own site still matters, but as the place where facts are checked, not as the voice that vouches for you. Measure it by where you appear in AI answers, not by backlinks alone.

## The short version

1. In tests of niche brands, 95.1% of the sources ChatGPT drew on were earned media, such as reviews and independent publications ([Chen and colleagues](https://arxiv.org/abs/2509.08919)).
2. In [our brand entity study](https://underneath.agency/research/brand-entity-ai-recommendations-study), each tenfold increase in independent sites naming a brand went with 4.7 times the odds of being recommended.
3. Products AI assistants recognized by name 99.4% of the time surfaced in only 3.32% of open “best tools” questions on ChatGPT ([Sharma](https://arxiv.org/abs/2601.00912), a small study).
4. In lab tests, fake authority claims worked, but when all rivals used them, the famous brand’s recommendation rate recovered to 93.8% ([Chu and Hou](https://arxiv.org/abs/2606.17443)).

## What does “authority” mean to an AI search engine?

To an AI engine, authority is mostly what credible others publish about you, not what you say about yourself.

[Chen and colleagues](https://arxiv.org/abs/2509.08919) compared AI search engines with Google across product categories. For niche brands, ChatGPT drew 95.1% of its sources from earned media and only 4.9% from brand sites. Perplexity was more mixed, with 17.5% from social platforms such as YouTube and Reddit. Google’s results kept a far larger share of brand-owned and social pages.

A health study shows what kind of authority counts. [Jacques and colleagues](https://arxiv.org/abs/2601.17109) coded 615 sources ChatGPT cited for 100 consumer health questions in January 2026. Institutional sources such as hospitals, government sites and encyclopedias made up 75.7%. Yet 64.7% of the cited pages had no author attribution at all. The authority was the institution’s, not the byline’s.

## Why does authority matter differently in generative search?

Because an AI answer names only a few brands, and a short list of trusted sources decides which ones.

A ranked page of results lets a reader browse past the leaders. An AI answer does the browsing and hands back a verdict. In an analysis of a public dataset of 602 prompts, official, news and specialist sources made up 79.12% to 87.52% of citations on each platform ([Zhang, He and Yao](https://arxiv.org/abs/2604.25707)). Being one of the recognized sources is the entry ticket.

Google rankings carry only part of that weight. In [our AI citations and Google rankings study](https://underneath.agency/research/ai-citations-google-rankings-study), only 8.3% of the pages ChatGPT cited ranked in Google’s top 10 for the question. Across six AI search engines and 55,936 queries, 37% of domains cited were unique to AI search and never appeared in Google or Bing results ([Zhang and colleagues](https://arxiv.org/abs/2512.09483)). For more on rankings, see [is traditional SEO still important for AI search?](https://underneath.agency/resources/is-seo-still-important-for-ai-search)

## Will authority matter more than keyword rankings?

For being recommended, the evidence points that way: being known by name is not the same as being recommended.

In [Sharma’s](https://arxiv.org/abs/2601.00912) test of 112 Product Hunt startups, AI assistants recognized products asked about by name 99.4% of the time on ChatGPT. Asked open questions such as “What are the best AI tools launched this year?”, ChatGPT surfaced them in only 3.32% of cases. Referring domains, a classic sign of others linking to you, predicted visibility on Perplexity. This is a small, single-author study.

A vendor’s tracking data shows the same ladder. Household-name brands appeared in 73% of relevant AI answers on their first tracking run, established mid-market brands in 44% and small brands in 11% ([Kumar](https://arxiv.org/abs/2606.20065), 2026). Those tiers reflect fame built over years, which no keyword ranking supplies on its own.

## Should companies invest more in third-party coverage?

Yes; independent coverage is the strongest signal of being recommended that we have measured.

In [our brand entity study](https://underneath.agency/research/brand-entity-ai-recommendations-study), each tenfold increase in the number of independent sites naming a brand in the cited pages went with 4.7 times the odds of being recommended. Having a Wikipedia article mattered mainly because it went with being named in the evidence. Brands with an article were named in the cited pages 74.3% of the time, against 59.5% without. Once named there, they were recommended at about the same rate.

So coverage works by getting you into the pages AI engines read before they choose. The live guides on [small brands](https://underneath.agency/resources/how-small-brands-get-recommended-by-ai) and [ad spend](https://underneath.agency/resources/does-ad-spend-help-ai-brand-recommendations) cover who benefits and what does not substitute for it.

## Is PR becoming part of search strategy in the age of AI search?

Yes, because some AI assistants now search specifically for publications, rankings and awards before answering.

In [our hidden searches study](https://underneath.agency/research/ai-hidden-searches-study), ChatGPT ran a search aimed at a named publication, ranking or award in 43.8% of its answers. Gemini did so in 26.2% and Claude in 1.2%. Coverage in those outlets is what such searches find.

PR coverage gets you selected, but other pages may shape the words. In the public-dataset analysis, encyclopedia pages scored 0.2144 on the authors’ 0-to-1 measure of influence on the answer text, against 0.0726 for news pages. News gets you into the pool; clear explanatory pages supply more of what is said. In one vendor’s data, ranked “best of” lists were the most cited format, at about 21% of all citations ([Kumar](https://arxiv.org/abs/2606.20065)).

## Does AI search change the value of third-party reviews and publications?

It raises their value, and concentrates it in a small number of outlets per engine.

In [our reputation study](https://underneath.agency/research/is-it-legit-ai-reputation-study), answers about whether brands were legitimate cited 484 domains. The evidence was concentrated enough to behave like 30.5 equally used sources. A few review sites and publications carry most of the weight, so knowing which ones your category’s answers cite matters more than total coverage volume.

Engines also read different outlets. In [our brand agreement study](https://underneath.agency/research/ai-assistants-brand-agreement-study), 66.3% of the options recommended for a question came from one assistant only. Coverage that one engine reads may be invisible to another. The guide [do AI engines cite your own website or third-party reviews?](https://underneath.agency/resources/do-ai-engines-cite-your-own-website) covers the split by engine.

## Can a brand manufacture its own authority?

Not reliably, and the attempts that work in tests carry real risks.

Self-ranked lists are common but small. In [our best-lists study](https://underneath.agency/research/self-promoting-best-lists-study), 24.2% of cited numbered “best X” lists ranked their own publisher first. Such lists were 1.1% of all citations, and we found no detectable difference in how often their top pick was named.

Invented authority is another matter. In lab tests by [Chu and Hou](https://arxiv.org/abs/2606.17443), claims that mimic evidence, including made-up clinical trials, beat a famous brand 50–73% of the time. But when all nine rival brands used the same language, the famous brand’s recommendation rate recovered to 93.8%. The authors call fabricated claims potential false advertising. The advantage disappears once everyone copies it.

## Should you optimize your entire digital presence, not just your website?

Yes; AI answers draw on peers, reviews and publications, and your site must still be readable when agents check it.

In one vendor’s data, own and third-party company pages together made up about 78% of 149,912 citations, often competitors’ pages in “alternatives” answers ([Kumar](https://arxiv.org/abs/2606.20065)). Your presence on comparison, review and partner pages is part of your authority.

Your own site still decides the final check. In a vendor study that held third-party mentions equal, AI agents clearly recommended businesses with agent-readable sites 20% of the time, against 11% for sites agents struggled to read ([Finder and colleagues](https://arxiv.org/abs/2609.34951)). See [where AI agents get their answer when they can’t read your site](https://underneath.agency/resources/when-ai-agents-cant-read-your-site).

## What should you do about it?

Treat authority as a coverage program aimed at the specific sources AI engines read, and measure it there.

1. **Map the cited sources in your category.** Ask each engine your buyers’ questions and list the outlets, reviews and lists they cite.
2. **Pitch those outlets first.** Prioritize publications, rankings and awards that the assistants search for, over volume of placements.
3. **Earn real, checkable proof.** Use genuine certifications, studies and expert reviews; avoid invented claims.
4. **Keep review profiles healthy.** Review platforms are a major source for reputation answers.
5. **Keep one clear reference page about the company.** Explanatory, factual pages shape what answers say.
6. **Measure authority in AI answers.** Track, per engine, how often you are named, which sources name you, and how you are described.

Ownership across SEO, PR and content is covered in [who should own AI search visibility](https://underneath.agency/resources/who-should-own-ai-search-visibility). For outside help, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The research links coverage to recommendations, but no study has yet proved that new coverage causes them.

- The coverage findings, ours included, compare brands at one moment; none tracked a brand before and after a PR push.
- Several key figures come from vendors who sell AI visibility tools or readiness scores.
- The startup study is small and single-authored, and tested older assistant versions.
- The fabricated-authority tests used invented products, mostly in skincare, in a lab setup.
- How long earned coverage takes to show up in AI answers is unmeasured.
- Results vary by engine, language and industry, so one category’s cited outlets may not transfer.

## Frequently asked questions

### Is PR becoming part of SEO in the AI search era?

In practice, yes. ChatGPT ran a search aimed at a named publication, ranking or award in 43.8% of its answers in our study, so press coverage is now direct input to AI answers.

### Should CMOs measure earned-media authority differently for AI search?

Yes. Count how often each engine names you and which outlets it cites, because only 8.3% of the pages ChatGPT cited ranked in Google’s top 10 for the question.

### Do third-party reviews matter more for AI search than for Google?

They appear to. For niche brands, ChatGPT drew 95.1% of its sources from earned media such as reviews, while Google showed far more brand-owned pages.

### Does a Wikipedia page give my brand authority in AI search?

Mostly as a sign of wider coverage. In our study, the Wikipedia advantage largely disappeared once brand prominence and independent coverage were taken into account.

## Sources

- Chen and colleagues (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919), arXiv:2509.08919.
- Jacques and colleagues (2026), [Authority Signals in AI Cited Health Sources: A Framework for Evaluating Source Credibility in ChatGPT Responses](https://arxiv.org/abs/2601.17109), arXiv:2601.17109.
- Zhang, He and Yao (2026), [From Citation Selection to Citation Absorption: A Measurement Framework for Generative Engine Optimization Across AI Search Platforms](https://arxiv.org/abs/2604.25707), arXiv:2604.25707.
- Zhang and colleagues (2025), [Source Coverage and Citation Bias in LLM-based vs. Traditional Search Engines](https://arxiv.org/abs/2512.09483), arXiv:2512.09483.
- Sharma (2025), [The Discovery Gap: How Product Hunt Startups Vanish in LLM Organic Discovery Queries](https://arxiv.org/abs/2601.00912), arXiv:2601.00912.
- Kumar (2026), [Generative Engine Optimization at Scale: Measuring Brand Visibility Across AI Search Engines](https://arxiv.org/abs/2606.20065), arXiv:2606.20065.
- Chu and Hou (2026), [Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems](https://arxiv.org/abs/2606.17443), arXiv:2606.17443.
- Finder, Elovic, Shalev and Yosef (2026), [AX is the New AEO](https://arxiv.org/abs/2609.34951), arXiv:2609.34951.
- Underneath (2026), [Do Wikipedia and schema make AI assistants recommend a brand?](https://underneath.agency/research/brand-entity-ai-recommendations-study)
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude cite pages that rank?](https://underneath.agency/research/ai-citations-google-rankings-study)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [“Is this brand legit?” How AI assistants build a reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study)
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude agree on brands?](https://underneath.agency/research/ai-assistants-brand-agreement-study)
- Underneath (2026), [How many “best of” lists cited by AI rank their own brand first?](https://underneath.agency/research/self-promoting-best-lists-study)

---

This is the Markdown twin of https://underneath.agency/resources/how-brands-build-authority-for-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How many prompts do we need to track to measure AI visibility?"
description: "There is no fixed number. Studies suggest about 40 to 150 or more prompts per topic per AI engine, each asked several times, before a ranking is trustworthy."
canonical: "https://underneath.agency/resources/how-many-prompts-to-track-ai-visibility"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How many prompts do we need to track to measure AI visibility?

There is no single right number, and anyone who quotes one is guessing. The best current research points to roughly 40 to 150 or more prompts per topic for each AI engine, each asked several times. The exact figure depends on the engine, the topic and how fine a difference you need to see. Which prompts you choose matters as much as how many.

## The short version

1. Across three consumer topics, Gemini needed about 40 to 50 prompts per topic for a tight reading, Perplexity about 100 and OpenAI’s search model 150 or more (Sielinski, 2026).
2. In a second dataset of 30 engine-and-topic pairs, no budget below 94 answers would have been enough for every pair, and three pairs were still unsettled at 125 (Sielinski, 2026).
3. A Swiss study across four AI engines recommends asking each prompt at least 7 times a day to estimate whether a brand appears (Schulte and colleagues, 2026).
4. In our own test, a single ChatGPT answer showed only 57.8% of the brands that five answers to the same question named between them ([our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study)).
5. The mix of prompts can swing a score on its own: reweighting the same published data produced AI answer rates of 39.7% or 70.5% (Martinez, 2026).

## Is there a standard number of prompts?

No. The research shows the right number changes with the engine and the topic, so no fixed count holds everywhere.

The clearest test comes from [Sielinski](https://arxiv.org/abs/2607.10341), who tracked 10 topics on Gemini, Perplexity and OpenAI’s search model. He asked when each ranking of cited websites stopped moving and became precise enough to act on. The method settled for 27 of the 30 engine-and-topic pairs within 125 answers. In that data, no budget below 94 answers would have covered every pair that settled, and three pairs had still not settled by 125.

Early answers are especially misleading. In the same paper, the ranking of the most-cited websites kept changing substantially through the first 20 to 60 queries, depending on topic and engine. The author works for a company that sells AI visibility measurement, and the queries were written by an AI tool rather than taken from real users.

## How many prompts does each AI engine need?

Fewer for Gemini, more for Perplexity, and the most for ChatGPT’s search model. The gap tracks how many sources each engine cites per answer, and how erratically.

In an earlier paper, [Sielinski](https://arxiv.org/abs/2603.08924) sent the same 200 queries per topic to the three engines every day for nine days, across bird feeders, multivitamins and running gear. He then asked how many queries it takes to narrow a website’s share of citations to a range about five percentage points wide.

| AI engine | Prompts per topic for a five-point reading |
|---|---|
| Gemini | About 40 to 50 |
| Perplexity | About 100 |
| OpenAI search model (called SearchGPT in the study) | 150 or more, and sometimes no fixed number worked |

There is a hidden cost on top. OpenAI’s model sometimes answers without citing anything: on one topic in the later paper, it returned citations for only 104 of 125 questions, a zero-citation rate of about 17%. To end up with 100 usable answers, the author says you would need to send about 121 questions rather than 100. Our guide on [measuring share of citations](https://underneath.agency/resources/measuring-share-of-citations-in-ai-search) explains how to count those empty answers.

## How many times should each prompt be asked?

More than once, and probably at least seven times. The same prompt gives a different answer from one run to the next.

[Schulte and colleagues](https://arxiv.org/abs/2604.07585) ran eight prompts per industry through ChatGPT, Gemini, Google AI Mode and Perplexity in four Swiss industries over 45 days. From repeated same-day runs, they concluded that brand tracking needs at least 7 runs per prompt per day, and at least 8 if you care which web pages are cited. One author is affiliated with the company that supplied the data, and the thresholds assume brands that appear some of the time, not always or never.

Our own test points the same way. We asked ChatGPT, Gemini and Perplexity the same 20 buyer questions five times each. One ChatGPT answer showed 57.8% of the brands that the five answers named between them. Even five runs is coarse: pinning a brand that appears about half the time to within 10 points would take about 97 runs ([our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study)). That is why [a one-time AI visibility report](https://underneath.agency/resources/one-time-ai-visibility-report) can mislead.

## Does the mix of prompts matter more than the count?

Often, yes. Which questions you ask, in what words, language and country, can move a score more than adding runs.

[Martinez](https://arxiv.org/abs/2609.06811) argues that a prompt set defines its own “answer market”, which need not match what real customers ask. Saying a source is cited in 40% of answers means little, he notes, until you know what was asked. Open questions, questions naming the brand and comparisons give very different rates. Using published figures for Google, he showed that two weightings of the same three groups of searches produced AI answer rates of 39.7% and 70.5%, with nothing else changed. That calculation concerns whether an AI answer appears, not brand visibility, but the lesson carries over.

Other studies show the same pattern for brands:

- **Question type.** On one tracking platform, broad discovery questions named a tracked brand 23% of the time on average. Specific problem and use-case questions did so about 11% of the time ([Kumar](https://arxiv.org/abs/2606.20065), who co-founded the platform).
- **Wording.** In [our wording study](https://underneath.agency/research/ai-prompt-phrasing-study), asking the same question again kept the same first brand 68.0% of the time. Adding “on a tight budget” kept it only 15.3% of the time.
- **Country.** On a scale where 1 means identical brand lists, two ChatGPT answers from the same country scored 0.594 and answers from different countries 0.429 ([our country study](https://underneath.agency/research/ai-recommendations-by-country-study)).
- **Language.** For 66 European brands, switching from English to the brand’s home language raised how often local champions were recommended by 0.80 on a 0-to-1 scale. For global brands the rise was 0.15 ([Żatuchin](https://arxiv.org/abs/2606.23165), who also works for a measurement company).
- **The question itself.** In our consistency test, the question accounted for 30.0% of the variation in how stable answers were, more than the choice of assistant.

## Is it better to add new prompts or repeat old ones?

Adding different prompts usually buys more precision than repeating the same ones. Repeats reduce only one kind of noise.

A study of 12,933 answers about 20 Central and Eastern European brands split the variation in the tone of AI answers into its sources ([Żatuchin](https://arxiv.org/abs/2607.13304)). A design that spread 360 questions across 15 wordings measured brands more reliably than a 600-question design that asked 5 wordings 5 times each. The outcome was the tone of each answer, not whether a brand was named, and the author sells brand measurement, so treat the exact sizes with care.

Schulte and colleagues reached a similar view by another route. Within one industry, some prompts gave nearly the same sources every run, with overlap scores above 0.8, while others stayed below 0.2. Tracking one or two prompts would mostly measure the quirks of those prompts.

## How long should you track before trusting a trend?

Weeks, not days. Daily snapshots of a single brand stay noisy for several weeks of collection.

In the Swiss data, the estimate for one brand’s appearance rate only became reasonably steady after about 24 days of daily tracking. The authors recommend rolling averages over two to four weeks. Sielinski found a related trap: rankings can look steady from one day to the next while drifting over nine days. The same caution applies when you [judge whether a GEO campaign worked](https://underneath.agency/resources/did-geo-improve-ai-citations).

## What should you do about it?

Start with the decision the numbers must support, then size the prompt set to it. A practical plan:

1. **Set the precision you need first.** Telling a brand at 12% from one at 10% needs far more prompts than spotting that a competitor appears twice as often as you.
2. **Size per engine.** Plan for about 40 to 50 prompts per topic on Gemini and 100 or more on Perplexity and ChatGPT’s search, then check whether your rankings have settled.
3. **Send extra prompts where answers lack citations.** If about one answer in six comes back with no sources, as in one test, send about a fifth more prompts.
4. **Repeat each prompt several times.** Seven runs is a reasonable starting point for brand mentions.
5. **Build the set to match real demand.** Mix broad and specific questions, budget and premium wording, and every country and language you sell in. Report each group separately.
6. **Report a range, not a single figure.** “Named in 4 of 7 runs” is honest; “visible in ChatGPT” is not.
7. **Judge trends on two-to-four-week averages,** not day-to-day moves.

If you want help designing a prompt set like this, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The research has not produced a tested, general rule for how many prompts a brand needs. Specific gaps:

- **Real user questions.** The main studies used prompts written by an AI tool or taken from search suggestions, not logs of what customers actually ask.
- **Business-to-business topics.** The sample-size work covered consumer topics such as running gear and smoke detectors.
- **Brand mentions on every engine.** The per-engine figures above measure how often websites are cited, not how often a brand is named.
- **Long-term drift.** Most datasets cover days or weeks, so model updates over months are not captured.
- **Independent replication.** Several key studies come from companies that sell measurement, and none has been repeated by an outside team.

## Frequently asked questions

### Is 50 prompts enough to measure AI visibility?

It can be for one topic on Gemini, but probably not on ChatGPT’s search or Perplexity. In Sielinski’s test, those engines needed about 100 and 150 or more prompts per topic for a similar precision.

### Should I track the same prompts every day?

Yes, keep a fixed core set so you can compare over time, and repeat each prompt several times. Schulte and colleagues recommend at least 7 runs per prompt per day for brand tracking.

### Why does my AI visibility score jump around from week to week?

Mostly because each answer is a sample from a changing distribution. A single ChatGPT answer showed only 57.8% of the brands five answers named in our test, so small prompt sets swing a lot.

### Do I need prompts in other languages?

Yes, if you sell in other languages. For local European brands, asking in the home language raised how often they were recommended by 0.80 on a 0-to-1 scale.

## Sources

- Sielinski, R. (2026), [Quantifying Uncertainty in AI Visibility: A Statistical Framework for Generative Search Measurement](https://arxiv.org/abs/2603.08924), arXiv:2603.08924.
- Sielinski, R. (2026), [From Stochastic to Stable: Rank Stability and Structural Sufficiency in AI Visibility Measurement](https://arxiv.org/abs/2607.10341), arXiv:2607.10341.
- Schulte, J., Bleeker, M. and Kaufmann, P. (2026), [Don’t Measure Once: Measuring Visibility in AI Search (GEO)](https://arxiv.org/abs/2604.07585), arXiv:2604.07585.
- Martinez, O. (2026), [Measuring GEO Visibility: Prompt Corpora Define the Answer Market](https://arxiv.org/abs/2609.06811), arXiv:2609.06811.
- Kumar, P. (2026), [Generative Engine Optimization at Scale: Measuring Brand Visibility Across AI Search Engines](https://arxiv.org/abs/2606.20065), arXiv:2606.20065.
- Żatuchin, D. (2026), [The Language Blind Spot: How Query Language and Brand Recognition Tier Shape AI-Constructed Brand Reputation Across Twelve European Languages](https://arxiv.org/abs/2606.23165), arXiv:2606.23165.
- Żatuchin, D. (2026), [Where Does the Noise Come From? A Variance-Components Decomposition of Non-Determinism in LLM Brand Answers](https://arxiv.org/abs/2607.13304), arXiv:2607.13304.
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)
- Underneath (2026), [Does rewording a question change AI brand recommendations?](https://underneath.agency/research/ai-prompt-phrasing-study)
- Underneath (2026), [Same question, four countries: do AI recommendations change?](https://underneath.agency/research/ai-recommendations-by-country-study)

---

This is the Markdown twin of https://underneath.agency/resources/how-many-prompts-to-track-ai-visibility. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How many sources does each AI search engine cite per answer?"
description: "Anywhere from about 3 to 40 sources per answer, depending on the engine, the topic and how it was measured. The engine ranking flips between studies."
canonical: "https://underneath.agency/resources/how-many-sources-ai-search-engines-cite"
published: 2026-10-11
updated: 2026-10-11
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How many sources does each AI search engine cite per answer?

There is no fixed number: published studies find from about 3 to about 40 cited sources per answer, depending on the engine, the topic and how the engine was reached. Which engine cites the most also changes from study to study. A citation benchmark therefore only means something for the engine, topic and setting it was measured in.

## The short version

1. In a June 2026 audit of 15 US commercial prompts, ChatGPT cited a mean of 17.8 web addresses per prompt, Google 12.1, Perplexity 10.0 and Copilot 4.5; the author sells AI visibility software.
2. In a university audit of 2,848 answers on politics, health and the environment, ChatGPT cited 14.7 sources per answer, Copilot 8.4, Perplexity 7.9 and Gemini 7.3.
3. Through developer access the order can reverse: one study counted 36 to 40 citations per answer on Gemini and 5 to 7 on SearchGPT.
4. In our own test of 80 US buyer questions, Perplexity listed a median of 20 sources per answer and the other three assistants 3 to 5.
5. Engines read far more than they show: ChatGPT’s developer version returned 37.38 pages per question from its searches but cited 3.22.

## How many sources does each engine cite?

Most engines cite somewhere between a handful and a few dozen sources, and the counts vary widely by study. The table brings together the largest recent measurements.

| Study | ChatGPT | Google AI | Perplexity | Copilot | Gemini |
|---|---|---|---|---|---|
| Tannenbaum, June 2026, 15 US commercial prompts | 17.8 | 12.1 | 10.0 | 4.5 | not tested |
| Allaham and Diakopoulos, 2026, 712 queries on politics, health, environment | 14.7 | not tested | 7.9 | 8.4 | 7.3 |
| Zhang, He and Yao, 2026, 602 controlled prompts | 6.88 | 12.06 | 16.35 | not tested | see Google |

[Tannenbaum](https://arxiv.org/abs/2609.22655) ran one day of prompts about AI visibility software in a US setting. The paper discloses that the author founded the company whose tool collected the data, so treat it as a vendor study. [Allaham and Diakopoulos](https://arxiv.org/abs/2605.23684) used the public web versions of each engine and found ChatGPT cited the most. [Zhang, He and Yao](https://arxiv.org/abs/2604.25707), three independent researchers, re-analyzed a public dataset and found the opposite order, with Perplexity highest; in their data “Google” means AI Overviews and Gemini together.

Google’s own surfaces differ too. In [our AI Overview study](https://underneath.agency/research/ai-overview-citations-study) of 481 US AI Overviews, the median was 8 citations per AI Overview, the AI summary at the top of Google’s results. On the searches where both appeared, [our comparison of AI Mode and AI Overviews](https://underneath.agency/research/ai-mode-vs-ai-overviews-study) found a median of 8 against 4.5 for AI Mode, Google’s separate chat-style search tab.

## Why do studies get such different numbers?

Topic, date and the way an engine is reached all change the count. None of the studies above is wrong; they measured different things.

Topic matters. In the Allaham and Diakopoulos audit, answers on the environment carried 10.9 citations on average, against 8.4 for health and 7.7 for politics. Access route matters even more. [Sielinski](https://arxiv.org/abs/2603.08924), who works for an AI visibility company, queried engines through their developer interfaces over nine days in February 2026, on three consumer product topics. It found median citations of 36 to 40 on Gemini, 19 to 22 on Perplexity and 5 to 7 on SearchGPT. Our guide on [tracking AI visibility through developer access](https://underneath.agency/resources/api-ai-visibility-monitoring) explains why that gap matters for monitoring.

The question itself also matters. In the Zhang, He and Yao dataset, harder questions with several conditions pushed ChatGPT down to 3.4 citations, while Google rose to 12.6 and Perplexity to 17.7. Our own test of 80 US buyer questions on 26 September 2026 found yet another pattern, described in [our Google rankings study](https://underneath.agency/research/ai-citations-google-rankings-study): Perplexity listed a median of 20 sources per answer and ChatGPT, Gemini and Claude 3 to 5.

## Do some AI answers cite nothing at all?

Yes, and how often depends heavily on the engine. Some engines answer from memory for a large share of questions.

In a [2025 comparison of six AI search engines with Google and Bing](https://arxiv.org/abs/2512.09483), run in incognito browser sessions on trending US topics, the AI engines cited 4.3 web addresses per answer on average, against 10.3 for the classic search engines. 82% of Grok’s answers and 38% of Gemini’s contained no cited website.

Other studies find fewer empty answers. In the politics, health and environment audit, 11.6% of health questions got an answer with no citations. [Another 2026 study](https://arxiv.org/abs/2603.16138) of 11,000 real search questions found that SearchGPT cited a median of 3 sources, with 97% of answers citing at least one. In [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study), none of the 42 answers that ran no web search cited anything.

## Do AI engines read more pages than they cite?

Yes, many more: the citations you see are a small slice of what the engine retrieved.

An [audit of AI product recommendations](https://arxiv.org/abs/2609.18729) could see both layers in ChatGPT’s developer version. Its searches returned a mean of 37.38 pages per question, but the final answer cited 3.22. Only 9.5% of the returned pages made it into the citations.

Our hidden-searches study points the same way from another angle. ChatGPT ran the most searches, yet cited only 2.38 websites per answer, against 6.28 for Gemini and 3.14 for Claude. How many results an assistant keeps matters as much as how much it searches.

## Why does the number of citation slots matter for a brand?

Fewer slots mean fewer chances to be cited, so raw citation counts cannot be compared across engines. A share that looks small on one engine may be strong on another.

Tannenbaum puts it simply: a page that is never eligible for a four-source answer faces different odds from one judged in an 18-source answer. The same study found that 96.4% of all cited web addresses appeared on only one of the four engines. The engines are not just citing different amounts; they are citing different pages. Our guide on [how concentrated AI citations are across websites](https://underneath.agency/resources/ai-search-citation-concentration) looks at who fills those slots.

Counting also distorts totals. Sielinski concluded that raw citation counts are not comparable across platforms and recommended share of citations within each engine instead. If you add up mentions across engines, the engine that cites the most will dominate the total. Our guide on [combining engines into one visibility score](https://underneath.agency/resources/combine-ai-engines-visibility-score) shows how to weight them instead.

## What should you do about it?

Measure each engine separately and judge your share against that engine’s own number of slots. Then:

1. Ask your team or agency for citation share per engine, never one blended number across engines.
2. Record how many sources each engine cited for your key questions, so a “low” share can be read against a small denominator.
3. Track the questions your buyers actually ask; counts change with topic and question complexity.
4. Check whether your tracking uses the consumer app or developer access, because the two can differ widely.
5. Treat any benchmark as dated; every study here is a snapshot.

If you want help setting up per-engine tracking, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The research does not give a stable, current number for any engine. The main gaps:

- Most studies are single snapshots, from one day to a few weeks, and engines change often.
- Topics are narrow: commercial software prompts, politics and health, or consumer products. Other industries are untested.
- Several key figures come from developer interfaces, which do not match what consumers see in the apps.
- No study we reviewed measures how the number of citations affects clicks or sales for a brand.
- The largest four-engine count comes from a vendor with a disclosed conflict of interest and has not been independently repeated.

## Frequently asked questions

### Which AI search engine cites the most sources?

It depends on the study. ChatGPT cited the most in two 2026 audits of public web versions (17.8 and 14.7 per answer), while Perplexity and Gemini cited more in studies using developer access.

### How many sources does a Google AI Overview cite?

About 8 in our data. Our study of 481 US AI Overviews in September 2026 found a median of 8 citations each.

### Does ChatGPT cite every page it reads?

No. In one audit its developer version retrieved 37.38 pages per question on average and cited 3.22.

### Is a higher citation count better for my brand?

Not by itself. More slots give more chances to appear, but each citation then carries less weight, so compare your share within each engine.

## Sources

- Tannenbaum (2026), [Scoring With the Engine: Retrieval Exposure, Cross-Engine Divergence, and the Limits of Engine-Agnostic GEO Scores](https://arxiv.org/abs/2609.22655), arXiv:2609.22655.
- Allaham and Diakopoulos (2026), [Synthetic Sources?: Auditing Generative Search Engine Citations for Evidence of AI-Generated Sources](https://arxiv.org/abs/2605.23684), arXiv:2605.23684.
- Zhang, He and Yao (2026), [From Citation Selection to Citation Absorption: A Measurement Framework for Generative Engine Optimization Across AI Search Platforms](https://arxiv.org/abs/2604.25707), arXiv:2604.25707.
- Zhang and colleagues (2025), [Source Coverage and Citation Bias in LLM-based vs. Traditional Search Engines](https://arxiv.org/abs/2512.09483), arXiv:2512.09483.
- Huang and colleagues (2026), [Answer Bubbles: Information Exposure in AI-Mediated Search](https://arxiv.org/abs/2603.16138), arXiv:2603.16138.
- Sielinski (2026), [Quantifying Uncertainty in AI Visibility: A Statistical Framework for Generative Search Measurement](https://arxiv.org/abs/2603.08924), arXiv:2603.08924.
- Uberti-Bona Marin and colleagues (2026), ["If I Had to Buy Just ONE: Galaxy S26 Ultra": Auditing AI-Generated Product Recommendations](https://arxiv.org/abs/2609.18729), arXiv:2609.18729.
- Underneath (2026), [Do AI Overviews cite the pages that rank?](https://underneath.agency/research/ai-overview-citations-study)
- Underneath (2026), [AI Mode vs AI Overviews: how different are the sources?](https://underneath.agency/research/ai-mode-vs-ai-overviews-study)
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude cite pages that rank?](https://underneath.agency/research/ai-citations-google-rankings-study)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)

---

This is the Markdown twin of https://underneath.agency/resources/how-many-sources-ai-search-engines-cite. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How much web content is already optimized for AI search?"
description: "About 9% of pages Google and Gemini returned for 1,000 real searches showed signs of AI-search optimization, rising to 16.36% for pages updated in 2026."
canonical: "https://underneath.agency/resources/how-much-web-content-is-optimized-for-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How much web content is already optimized for AI search?

About one page in eleven, according to the only large measurement so far: 8.90% of pages returned by Google Search and Gemini for 1,000 real searches showed signs of being optimized for AI engines. The share is higher among recently updated pages, reaching 16.36% for pages last modified in 2026. These are detector estimates, not confirmed counts, but they show that the field is no longer empty.

## The short version

1. A detector flagged 8.90% of 10,095 pages returned for 1,000 real searches as optimized for AI search: 8.14% from Google Search and 9.09% from Gemini ([Chu and colleagues](https://arxiv.org/abs/2608.16824)).
2. Among pages with a readable update date, the flagged share rose from 7.02% for 2024 to 16.36% for 2026, though only 19.57% of pages carried such a date (Chu and colleagues).
3. The share varies by site: 20.37% of Amazon pages in the sample were flagged, 4.38% of YouTube pages and none of 613 Wikipedia pages (Chu and colleagues).
4. Other AI-facing signals are still a minority: 11.5% of top websites publish a valid llms.txt file and 3.2% serve Markdown to AI agents ([our llms.txt study](https://underneath.agency/research/llms-txt-adoption-study), [our agent-readable web study](https://underneath.agency/research/agent-readable-web-study)).
5. In a simulated test, one brand optimizing cut the market leader’s share of recommendations from 100% to 19.8%, but when all nine rivals did the same, the leader recovered to 93.8% ([Chu and Hou](https://arxiv.org/abs/2606.17443)).

## How much of what Google and Gemini show looks optimized for AI?

About 9% of pages, in the one large audit published so far. [Chu and colleagues](https://arxiv.org/abs/2608.16824), researchers at CISPA Helmholtz Center for Information Security, HPE and the University of Waterloo, built a detector for GEO: generative engine optimization, meaning edits that make a page more likely to be picked and cited by AI search engines. Our [plain-language AI search glossary](https://underneath.agency/resources/glossary) defines GEO and the related terms.

They ran it on the pages that Google Search and Gemini returned for 1,000 real user searches, fetched between 28 and 31 July 2026. Of 10,095 usable pages, the detector flagged 898, an estimated 8.90%. The rate was 8.14% for pages from regular Google Search and 9.09% for pages Gemini used to ground its answers.

The authors call these estimates, not measurements. Live pages carry no label saying they were optimized, and the detector makes mistakes. It also struggled more with light-touch and human-written optimization, so subtle work may be undercounted.

## Is the share of AI-optimized content growing?

It appears to be, though the evidence is a trend, not proof. Among pages that declared when they were last modified, the flagged share rose from 7.02% in 2024 to 12.80% in 2025 and 16.36% in 2026. For 2026 pages it reached 13.52% on Google Search and 18.20% on Gemini.

There is a large caveat. Only 19.57% of pages had a readable modification date, and that date says nothing about when or whether optimization happened. The authors say the rise “should therefore be treated as a descriptive trend rather than evidence of increasing GEO adoption.”

Even so, the direction matters for planning. Pages that were recently touched are more likely to look optimized, so the baseline you compete against is rising, not static.

## Where is AI-optimized content most common?

On commercial and creator platforms more than on reference sites, in this sample. 20.37% of the Amazon pages in the audit were flagged, against 4.38% of YouTube pages. None of the 613 Wikipedia pages was flagged, and several large health sites also had none.

The samples per site are small: 54 Amazon pages and 434 YouTube pages. The authors suggest commercial platforms have stronger reasons to optimize for AI shopping and discovery tools, but say their data “cannot establish the intent behind individual pages.”

Quality is a concern on the flagged pages. Of 6,663 citations inside them, 69.34% pointed to sources the authors rated low in verifiability, meaning weak editorial accountability or hard to check. That share was 74.15% on pages Gemini used and 45.88% on pages from Google Search. A low rating does not mean the claim is false. A related question is [whether AI engines cite AI-written pages](https://underneath.agency/resources/are-ai-search-engines-citing-ai-content).

## What other signs show websites adapting to AI?

Technical signals for AI agents remain a minority practice among large websites. In [our study of 5,902 top websites](https://underneath.agency/research/llms-txt-adoption-study), 11.5% served a valid llms.txt file, a proposed summary file for AI tools. Only 3.2% returned a Markdown version of a page when an AI agent asked for one ([our agent-readable web study](https://underneath.agency/research/agent-readable-web-study)). Structured data is more common: 45.8% of homepages carried it.

Some content is shaped to steer AI answers more directly. In [our study of AI-cited “best of” lists](https://underneath.agency/research/self-promoting-best-lists-study), 24.2% of numbered lists with an identifiable publisher ranked their own publisher first.

At the far end are hidden instructions aimed at AI tools. Scanning 1.2 billion web addresses, [Khodayari and colleagues](https://arxiv.org/abs/2604.27202) found 15.3 thousand confirmed instances, including roughly 1.5 thousand meant to manipulate reputation. Many were long-lived: 65% of archived pages already carried them 12 months earlier. These are tiny numbers against the whole web, and in the authors’ tests AI tools followed them at most 8% of the time.

## What happens when everyone in a category optimizes?

The advantage shrinks, but sitting out looks worse, at least in one simulated market. [Chu and Hou](https://arxiv.org/abs/2606.17443) tested skincare recommendations from three AI assistants using fictional challenger brands and one real market leader.

When one challenger used authority-style marketing language, the leader’s survival as the recommended pick fell from 100% to 19.8%. When all nine challengers used it, the leader recovered to 93.8%. Each brand’s gain on the authors’ scale fell from 0.802 for the first mover to 0.007 once everyone optimized, close to nothing, and brands that did not optimize received zero recommendations.

Two cautions apply. The language tested included made-up clinical claims, which no brand should copy; our guide on [how GEO can push false claims](https://underneath.agency/resources/can-geo-push-false-information-into-ai-answers) covers that risk. And this was a controlled test of three assistants in one product category, not a live market. Still, it describes the pressure the audit above suggests is building.

## What should you do about it?

Assume competitors are already optimizing, and compete on verifiable substance rather than tricks. In practice:

1. Check your category. Ask AI assistants your buyers’ questions and note which pages and lists they cite.
2. Look at your recently updated pages first. That is where optimized competitors cluster, by the audit’s date trend.
3. Back claims with checkable sources. Flagged pages leaned heavily on low-accountability citations, so verifiable evidence is a way to stand apart.
4. Do not rely on self-ranking lists or hidden instructions. Self-ranking lists were 1.1% of all AI citations in our study, and hidden instructions rarely worked in testing.
5. Track whether your share of AI recommendations holds as competitors optimize, not just whether you appear once.

For help running that kind of category review, see [our generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

Nobody knows the true share of AI-optimized content, because no page carries a reliable label.

- The 8.90% figure comes from one detector, one set of 1,000 searches and one week of fetching. The detector can miss subtle, human-written optimization.
- The rise over time rests on the 19.57% of pages with a readable date and is not evidence of adoption.
- Per-site figures rest on small samples, such as 54 Amazon pages.
- The competition study is simulated, with fictional brands in one product category.
- No study yet links a page being flagged as optimized to whether it actually wins more AI citations in live search.

## Frequently asked questions

### Are my competitors already doing GEO?

Some probably are. An audit of 10,095 pages from Google Search and Gemini flagged 8.90% as optimized for AI, and 16.36% of pages updated in 2026.

### Is Gemini more exposed to AI-optimized content than Google Search?

Slightly, in the one audit available. The detector flagged 9.09% of pages Gemini used against 8.14% of Google Search pages, and 18.20% against 13.52% for pages updated in 2026.

### How many websites have an llms.txt file?

A minority of large sites do. In our check of 5,902 top websites, 11.5% served a valid llms.txt.

### If everyone optimizes for AI, does it stop working?

In one simulated market, mostly yes for the gain, but not for the cost of opting out. Each brand’s gain fell to near zero once all nine challengers optimized, and brands that did not optimize got zero recommendations.

## Sources

- Chu, Leng, Li, Shen, Shen and Zhang (2026), [GEO-Flag: Detecting and Measuring GEO-Optimized Web Content](https://arxiv.org/abs/2608.16824), arXiv:2608.16824.
- Chu and Hou (2026), [Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems](https://arxiv.org/abs/2606.17443), arXiv:2606.17443.
- Khodayari, Zhang, Acharya and Pellegrino (2026), [Indirect Prompt Injection in the Wild: An Empirical Study of Prevalence, Techniques, and Objectives](https://arxiv.org/abs/2604.27202), arXiv:2604.27202.
- Underneath (2026), [How many websites have an llms.txt file? 2026 adoption data](https://underneath.agency/research/llms-txt-adoption-study)
- Underneath (2026), [Do websites serve Markdown to AI agents? 2026 data](https://underneath.agency/research/agent-readable-web-study)
- Underneath (2026), [How many “best of” lists cited by AI rank their own brand first?](https://underneath.agency/research/self-promoting-best-lists-study)

---

This is the Markdown twin of https://underneath.agency/resources/how-much-web-content-is-optimized-for-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How can a small brand get recommended by AI assistants?"
description: "Small brands start far behind in AI answers. The best evidence says to get written about by independent sites and make your advantages easy to verify."
canonical: "https://underneath.agency/resources/how-small-brands-get-recommended-by-ai"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How can a small brand get recommended by AI assistants?

Mainly by being written about on independent sites and by stating a clear, checkable advantage, not by tweaking pages for each assistant. Small brands start far behind: in one vendor’s data they appeared in about one unbranded AI answer in ten. The research suggests building that wider presence first, then fine-tuning.

## The short version

1. In one vendor’s tracking data, small brands appeared in 9.9% of unbranded ChatGPT answers on day one, 8.9% on Gemini and 6.8% on Perplexity ([Kumar](https://arxiv.org/abs/2606.20065), 2026).
2. The same study’s fastest-rising small brands gained only 10 to 20 points between March and May 2026, so its author advises building broad presence before per-assistant tactics.
3. In [our brand entity study](https://underneath.agency/research/brand-entity-ai-recommendations-study), the number of independent websites naming a brand was the strongest signal: 4.7 times the odds of a recommendation per tenfold increase.
4. In controlled tests, an unknown brand beat a famous one half the time with a price just 7.3% lower ([Chu and Hou](https://arxiv.org/abs/2606.17443), 2026).
5. Asking in a brand’s home language raised how often AI assistants named local champions by 0.80 on a 0 to 1 scale, against 0.15 for multinationals ([Żatuchin](https://arxiv.org/abs/2606.23165), 2026).

## How far behind do small brands start?

Far behind, on every major assistant measured. [Kumar](https://arxiv.org/abs/2606.20065), a co-founder of the AI visibility company Ranqo, grouped brands tracked on its platform from March to May 2026 into three tiers by stature. On each brand’s first tracking run, niche and small brands appeared in 11% of unbranded category answers.

Outside the two top tiers, the picture was similar everywhere: 9.9% of unbranded answers on ChatGPT, 8.9% on Gemini and 6.8% on Perplexity. Small brands are hard to surface on every assistant, not just on one.

Nor did they climb much on their own. The fastest-rising small brands moved only 10 to 20 percentage points across the observation window. This is a vendor’s customer data, with tiers assigned by hand, so treat the exact figures as indicative.

## Why does being written about matter more than your own website?

Because AI assistants mostly build answers from other people’s pages. In Kumar’s data, “best of” lists were the most cited kind of content, at about 21% of all citations. A single list that includes your brand can be reused across many different questions. Among sources other than company websites, YouTube was cited most, at 4.2% of citations.

[Our brand entity study](https://underneath.agency/research/brand-entity-ai-recommendations-study) found the same thing from another angle. Across ChatGPT, Gemini, Perplexity and Claude, the signal most closely tied to being recommended was independent coverage: how many websites, other than the brand’s own, named it in the pages the assistants cited. Each tenfold increase went with 4.7 times the odds of being recommended.

The other route in is genuine community discussion. In [Sharma’s study](https://arxiv.org/abs/2601.00912) of 112 Product Hunt startups, links from other websites and real Reddit discussion went with being surfaced by Perplexity, which searches the web. A score for on-page AI optimization showed no link at all.

## Do you need a Wikipedia page first?

No, and on its own it seems to add little once a brand’s wider prominence is counted. Kumar suggests small brands invest in Wikipedia, mainstream press and sustained YouTube presence, though he frames this as a hypothesis. Our data tests the Wikipedia part.

Brands with a Wikipedia article were more often named in the pages assistants cited, 74.3% against 59.5%. But once named there, they were recommended at the same rate, 50.2% against 48.9%. Most of the Wikipedia advantage reflected how well known those brands already were.

Consensus without an article also happens. Of the 110 options all four assistants named for a question, 21 had no Wikipedia article for themselves or a parent brand. In Kumar’s data, Wikipedia was also cited less often than video, media or forums, at 2.6% of citations. Coverage, not the encyclopedia entry, appears to do the work. Our guide on [why Wikipedia matters for AI search](https://underneath.agency/resources/why-wikipedia-matters-for-ai-search) covers where it does carry weight.

## Can clear facts help a small brand beat a bigger one?

Yes, when the assistant can see them side by side. [Chu and Hou](https://arxiv.org/abs/2606.17443) gave three AI models lists of skincare products with one real brand and nine invented ones. When everything was identical, the real brand always won. Our guide on [whether AI assistants favor big brands](https://underneath.agency/resources/do-ai-assistants-favor-big-brands) gathers the wider evidence.

The moment an invented brand had a visible edge, the picture flipped. It won half the time with a rating 0.075 stars higher, 1.6 times as many reviews, or a price 7.3% lower. In a small follow-up where the products were first found by a simple search step, the real brand ranked near the bottom on relevance and was recommended 0% of the time.

The lesson for a small brand is to make real advantages explicit and comparable: ratings, review counts, prices, specifications and certifications. The caveat is that these tests handed the facts to the assistant directly. In real answers, your facts must first be on pages it finds.

## Should a small brand win its home market and niche first?

The evidence points that way. A study of 66 European brands across twelve languages by [Żatuchin](https://arxiv.org/abs/2606.23165), who is affiliated with the AI brand-intelligence company Rankfor.AI, found query language changed which brands were recommended far more than how they were described. Asking in a brand’s home language raised how often local champions were named by 0.80 on a 0 to 1 scale, against 0.15 for global multinationals.

In other words, a local champion that seems invisible in English answers may be the default recommendation at home. Check AI answers in the languages and markets your buyers actually use before concluding you are absent.

Specific buyer needs work the same way. [Malthouse and colleagues](https://arxiv.org/abs/2609.16304) found that detailed questions about a buyer’s goals brought previously omitted brands into AI answers, and advise brands to stand for a few clear points of difference consistently everywhere they are described. The same team also looked at advertising, as our guide on [whether ad spend helps AI recommendations](https://underneath.agency/resources/does-ad-spend-help-ai-brand-recommendations) explains.

## Do shortcuts like bold claims or self-ranked lists work?

Briefly at best, and they carry real risks. Chu and Hou found that invented “clinical” claims let an unknown brand break through. But when every competitor used the same language, the advantage nearly vanished and the famous brand won again in 93.8% of trials.

Brands that did not compete that way fared worst of all: across 4,745 trials, they received zero recommendations. The authors used fabricated claims only to find the upper limit, and class them as potential false advertising. Real, verifiable evidence is the defensible version.

Self-ranked lists are common too. In [our study of self-promoting lists](https://underneath.agency/research/self-promoting-best-lists-study), 24.2% of the numbered “best” lists that AI engines cited ranked their own publisher first. We found no detectable difference in how often their top pick was named compared with independent lists, so they are not a proven shortcut.

## What should you do about it?

Build presence where AI assistants look, and make your advantage easy to check.

1. **Earn independent coverage.** Pursue reviews, comparisons and inclusion in credible “best of” lists in your category.
2. **Show up where buyers talk.** Genuine community discussion went with visibility in one study, and video was the top non-corporate source in another; fake posts do not count.
3. **Publish verifiable facts.** Put prices, ratings, review counts, specifications and real certifications in plain, comparable form.
4. **Own a niche and a market.** Measure AI answers for the specific needs, languages and places you serve before chasing the whole category.
5. **Keep the entity record tidy.** Consistent names, websites and descriptions prevent confusion, even if they do not win recommendations alone.
6. **Measure repeatedly.** Track unbranded questions across several runs and assistants, and expect progress in months, not days.

If you want help planning this, see [our approach to generative engine optimization](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The research describes where small brands stand far better than it proves how to move them.

- **No controlled test shows a lasting lift.** The tier advice is the vendor’s own hypothesis, and its planned experiment has not reported results.
- **Coverage findings are associations.** Our data cannot show that new coverage causes new recommendations.
- **Lab tests simplify reality.** The skincare experiments put the products in the question rather than letting assistants search.
- **Several sources have commercial ties.** The tier and language studies come from authors tied to companies that sell AI visibility tools.

## Frequently asked questions

### Can a small business get recommended by ChatGPT?

Yes, but it starts behind. In one vendor’s data, small brands appeared in 9.9% of unbranded ChatGPT answers, so independent coverage and clear facts matter more than for big brands.

### What is the fastest way for a small brand to appear in AI answers?

There is no proven fast route. The strongest signal in our data was independent coverage, and the fastest-rising small brands in one vendor’s data gained only 10 to 20 points across a three-month window.

### Does a small brand need a Wikipedia page for AI visibility?

No. In our data, 21 of the 110 options that all four assistants named had no Wikipedia article for themselves or a parent brand.

### Do fake reviews or bold claims help with AI recommendations?

Made-up claims swayed AI models in lab tests, but the edge vanished once rivals copied them, and invented evidence is potential false advertising.

## Sources

- Kumar (2026), [Generative Engine Optimization at Scale: Measuring Brand Visibility Across AI Search Engines](https://arxiv.org/abs/2606.20065), arXiv:2606.20065.
- Chu and Hou (2026), [Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems](https://arxiv.org/abs/2606.17443), arXiv:2606.17443.
- Sharma (2026), [The Discovery Gap: How Product Hunt Startups Vanish in LLM Organic Discovery Queries](https://arxiv.org/abs/2601.00912), arXiv:2601.00912.
- Malthouse, Lee, Yang, Pal and Feng (2026), [Evaluating Brand Retrieval and Ranking in Large Language Model Recommendations](https://arxiv.org/abs/2609.16304), arXiv:2609.16304.
- Żatuchin (2026), [The Language Blind Spot: How Query Language and Brand Recognition Tier Shape AI-Constructed Brand Reputation Across Twelve European Languages](https://arxiv.org/abs/2606.23165), arXiv:2606.23165.
- Underneath (2026), [Do Wikipedia and schema make AI assistants recommend a brand?](https://underneath.agency/research/brand-entity-ai-recommendations-study)
- Underneath (2026), [How many “best of” lists cited by AI rank their own brand first?](https://underneath.agency/research/self-promoting-best-lists-study)

---

This is the Markdown twin of https://underneath.agency/resources/how-small-brands-get-recommended-by-ai. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How to Choose a GEO Agency: 8 Criteria and Red Flags | Underneath"
description: "How to choose a generative engine optimization (GEO) agency in 2026: eight criteria, a scorecard, the red flags, and the questions to ask before you sign."
canonical: "https://underneath.agency/resources/how-to-choose-a-geo-agency"
published: 2026-09-25
updated: 2026-09-28
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How to choose a GEO agency

To choose a generative engine optimization (GEO) agency, pick the one that can show you repeated, per-engine measurements of a brand’s AI visibility, explains a method that is more than SEO with a new label, and ties its reporting to leads rather than screenshots. Below are eight criteria, a scorecard, the red flags and the questions to ask, checked against what the AI engines themselves tell buyers.

Underneath teamGuide

On this page

1. [Start with the question the engines ask](#start-with-the-question-the-engines-ask)
2. [Eight criteria for choosing a GEO agency](#eight-criteria-for-choosing-a-geo-agency)
3. [A scorecard you can use on a call](#a-scorecard-you-can-use-on-a-call)
4. [Red flags](#red-flags)
5. [Questions to ask before you sign](#questions-to-ask-before-you-sign)
6. [Agency, in-house or tool](#agency-in-house-or-tool)
7. [Frequently asked questions](#frequently-asked-questions)

Related service

Generative Engine Optimization

Get mentioned, cited and recommended by ChatGPT, Gemini, Perplexity and Copilot, and described correctly when you are.

[How we do it](https://underneath.agency/services/generative-engine-optimization)

## The short version

1. Judge a GEO agency on measurement first: it must run your buyer questions repeatedly on each engine and show who is cited, not one screenshot.
2. Its method must cover the pages, the structured data, the brand’s entity footprint and the third-party sources the engines cite; an SEO plan with “AI” added to the title is not that.
3. Ask to see the per-engine reporting, the crawler and schema checklist, and the list of third-party pages the agency intends to get you into. How it answers will sort the field quickly.

Plenty of SEO agencies now sell GEO too, and their proposals tend to read alike. Hold each one to three plain tests: is it specific, can it be measured, and does it hold up over time?

## Start with the question the engines ask

We asked ChatGPT, Google AI Overviews, Google AI Mode, Gemini and Perplexity “how to choose a GEO agency” 22 times between 15 and 25 September 2026. Read together, the answers make a decent buyer’s brief: they show which criteria the engines put first.

- Every platform answered with a numbered list of criteria, and the same three led each list: measurement and tracking of AI visibility, a GEO-specific methodology rather than repackaged SEO, and proof of results tied to business outcomes.
- ChatGPT’s answer ran to eight criteria and ended with a scorecard; Gemini’s opened with “evaluate their measurement and tracking capabilities”. Google’s AI Overview added a “red flags” section every time.
- Perplexity cited 13 sources per answer here, more than for any other question tested, and its stable sources were long checklists and guides. Google AI Mode cited a Reddit thread in 18 of 22 runs, so buyers who ask are shown peer experience alongside agency claims.
- The People Also Ask box asked “how to pick an SEO agency?” (21 times), “is SEO dead now with AI?” (21) and “what is the best GEO strategy?” (18). The SEO question is never far from the GEO one.

## Eight criteria for choosing a GEO agency

1. Measurement you can repeat. The agency should track a fixed set of your buyer questions on each engine daily, or at least across many days, and report who is cited and named. A single check is noise: Underneath’s own data shows sources cited once in twelve runs that never appeared again, and sources cited in ten of twelve that were positions worth defending.
2. A method that is more than SEO. Ask what the agency does that an SEO agency does not. The answer should include definition-first page structure, structured data for AI engines, entity and knowledge-graph work, and building citations on the third-party pages the engines use. If the answer is “content and links”, it is an SEO agency.
3. Per-engine thinking. ChatGPT, Google’s AI surfaces and Perplexity do not share sources. In Underneath’s study of a service keyword, ChatGPT took 26 of 28 citations from outside Google’s top 10 while Google’s AI Overview reused the top 10 heavily. Across 80 buyer questions, four assistants [recommended the same first pick for only 10.0% of questions](https://underneath.agency/research/ai-assistants-brand-agreement-study). An agency that talks about “AI search” as one thing has not measured it.
4. Technical competence. Crawler access for OAI-SearchBot, PerplexityBot and Google-Extended, server-rendered text, validated schema, one canonical URL per page. Ask to see the agency’s technical GEO checklist. If the engines’ crawlers cannot fetch and read a page, none of the content work on it will be retrieved.
5. Proof, without guarantees. Ask for a client’s before-and-after measurement: questions, engines, dates, citations. Be wary of any guarantee of being cited; no agency controls the engines, and the honest ones say so.
6. A third-party plan. Ask which roundups, directories, forum threads and publications the agency intends to get you into, and how it chose them. The right answer names the pages the engines actually cite for your category, which the audit should have found.
7. Reporting tied to outcomes. Mentions and citations by engine, share of voice against named competitors, and sessions and leads from AI referrals, monthly, against a baseline. Not “AI mentions up 300%” with no denominator.
8. Fit with your stack and your team. Implementation that works on the platform you already run, a plan for training your own people, and a clear line between what is in scope and what costs extra.

## A scorecard you can use on a call

Score each agency on the eight criteria from 0 to 3. An agency scoring under 2 on measurement or method should be dropped regardless of the total.

| Criterion | 0 | 1 | 2 | 3 |
| --- | --- | --- | --- | --- |
| Measurement | Screenshots | One-off audit | Repeated audit, one engine | Daily tracking, every engine, baseline |
| Method beyond SEO | “Content and links” | Adds schema | Adds entity work | Adds entity, schema, structure and third-party citations |
| Per-engine thinking | “AI search” as one thing | Names the engines | Explains differences | Plans and reports per engine |
| Technical GEO | None | Mentions robots.txt | Has a checklist | Checklist plus log-file evidence of AI crawlers |
| Proof | Claims | Testimonials | One dated case | Several dated, measured cases |
| Third-party plan | “Digital PR” | Generic outreach | Named target pages | Named pages per engine, from your audit |
| Reporting | Activity reports | Mentions | Mentions and citations by engine | Adds share of voice and AI referral leads |
| Fit | Fixed package | Some flexibility | Platform-agnostic | Platform-agnostic, trains your team, clear scope |

## Red flags

- A guarantee of citations, or a promise of “number one in ChatGPT”. The engines change their sources without notice; in our tracking, Perplexity replaced its entire source set for one query in a single day.
- One screenshot as evidence. One answer proves nothing about the next.
- No mention of third-party sources. On Perplexity and ChatGPT, third-party pages are most of the answer.
- No technical checklist. An agency that cannot say which AI crawlers your site allows has not checked whether its content work can be retrieved.
- SEO reports with “AI” added. Rankings and traffic are SEO metrics; GEO needs mentions, citations and share of voice.
- Talk of “gaming” or “hacking” the engines. Google’s guidance for its AI features says no special optimizations are needed to appear in them, which leaves little for a hack to do.

## Questions to ask before you sign

1. Which questions will you track, on which engines, how often, and can I see the baseline before we start?
2. What do you do that an SEO agency does not, and which of it happens off my site?
3. Which third-party pages will you try to get us into for our category, and how did you identify them?
4. Which AI crawlers do you allow, and how do you check that our pages render without JavaScript?
5. Show me one client’s measurement six months apart: questions, engines, dates, citations, and the leads that followed.
6. What is in the monthly report, and what is the first number you would expect to move?
7. What is in scope, what is an add-on, and what happens to the work when the engagement ends?

## Agency, in-house or tool

An agency makes sense when you need the method, the third-party outreach and the measurement at once, and do not yet have someone who has run GEO end to end. In-house makes sense once you have the tracker and the writing standards and need volume rather than method. A tool alone measures visibility; it does not change it. There is also a middle route: buy the audit from an agency, run the tracker yourself, and buy the third-party work as a project. Underneath’s programs are built to hand the tracker and the standards to your team at the end.

## Frequently asked questions

### How do I choose a GEO agency?

Choose the agency that measures your AI visibility repeatedly on each engine, explains a method that includes page structure, structured data, entity work and third-party citations, shows dated proof, and reports mentions, citations and share of voice against a baseline. Drop any agency that guarantees citations or shows one screenshot as evidence.

### How is choosing a GEO agency different from choosing an SEO agency?

The usual due diligence still applies, plus three tests: per-engine measurement, a plan for the third-party pages the engines cite, and technical checks for AI crawlers and rendering. An SEO agency can pass the usual checks and still fail all three.

### What is the best GEO strategy?

There is no single tactic. What the engines reward, in our measurements, is a system: definition-first pages with question headings, complete structured data, a consistent brand entity, presence on the third-party pages each engine cites, and daily tracking to see what changed, all run and measured per engine.

### How much does a GEO agency cost?

Published 2026 price guides range from about $1,500 to $50,000 a month, and [WebFX](https://www.webfx.com/blog/ai/generative-engine-optimization-cost/) puts mid-sized companies at $5,000 to $25,000. Underneath’s programs are priced in three bands, $5,000 to $10,000, $11,000 to $20,000 and $20,000+ a month, each beginning with an audit whose fee is credited if you continue. When you compare quotes, compare the audits too: a one-off check and a repeated, per-engine audit are not the same product at any price.

### How long before a GEO agency shows results?

Changes to pages that are already indexed can show up in Google AI Overviews and AI Mode within weeks, and in Perplexity once it recrawls them. Third-party citations take longer to land, usually months, and ChatGPT recommendations take longest. Ask for the first comparison against the baseline at 30 days.

### Is SEO still worth it in 2026?

Yes. Ranking is still the main way into Google’s AI surfaces: in our study of 486 searches, [pages ranking in the top three were cited in the AI Overview about twice as often](https://underneath.agency/research/ai-overview-cited-pages-study) as pages at positions 7 to 10. Ask any GEO agency how its work fits alongside your SEO, and be wary of one that proposes to replace it.

## Keep *reading.*

[All resources](https://underneath.agency/resources)

- Guide · AI search

  ### [What does a GEO agency do?](https://underneath.agency/resources/what-does-a-geo-agency-do)

  The nine things a generative engine optimization agency does, what it costs, and how the AI engines describe the job.

  Underneath team
- Guide · AI search

  ### [Is a GEO agency worth it?](https://underneath.agency/resources/is-a-geo-agency-worth-it)

  When hiring one pays back, when it does not, and how to decide with numbers rather than fear of missing out.

  Underneath team
- Guide · AI search

  ### [GEO agency vs SEO agency](https://underneath.agency/resources/geo-agency-vs-seo-agency)

  What each one does, where the work overlaps, and whether you need one, the other or both.

  Underneath team

Free strategy call

## Put us through the scorecard.

On a free 30-minute call we answer the seven questions above about our own work, and run your ten most important buyer questions through the engines.

[Book a strategy call](https://underneath.agency/contact)[How we work](https://underneath.agency/approach)

---

This is the Markdown twin of https://underneath.agency/resources/how-to-choose-a-geo-agency. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How should you design prompts and runs to track AI visibility?"
description: "Spread your budget across languages, assistants and wordings before buying more repeats of one prompt, and fix the question set and run count in advance."
canonical: "https://underneath.agency/resources/how-to-design-ai-visibility-tracking"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How should you design prompts and runs to track AI visibility?

Spend your tracking budget on variety first: more languages, more AI assistants and more ways of asking, before more repeats of the same prompt. One detailed study found that adding languages cut measurement error about fifteen times as much as adding five more repeats. The question set you choose also defines what the score means, so it deserves as much care as the tool that runs it.

## The short version

1. In a test of 20 brands, a single AI answer could barely rank brands at all (about 0.01 on a 0 to 1 reliability scale), and even 8 languages, 3 models and 15 phrasings reached only about 0.36 ([Żatuchin, 2026](https://arxiv.org/abs/2607.13304)).
2. In the same study, 15 phrasings asked once each (360 queries) ranked brands better than 5 phrasings asked 5 times each (600 queries).
3. A Swiss study of four AI search engines recommends at least 7 runs per prompt per day for brand tracking, pooled over two to four weeks ([Schulte and colleagues, 2026](https://arxiv.org/abs/2604.07585)).
4. Weighting the same published data two ways gave AI answer rates of 39.7% and 70.5%, which shows how much the choice of questions drives a headline score ([Martinez, 2026](https://arxiv.org/abs/2609.06811)).

## What are you really measuring when you track AI visibility?

You are measuring your brand against a market you defined yourself: your list of questions and how much each counts. [Martinez](https://arxiv.org/abs/2609.06811) calls this the “answer market”. A score of “cited in 40% of answers” means little until you know which questions, which assistant, which period and which weights produced it.

The paper shows how much this matters with published data on Google’s AI Overviews. It took activation rates for three groups of questions and weighted them two different ways, without changing any rate inside a group. The result was an AI answer rate of 39.7% under one weighting and 70.5% under the other. That example concerns whether an AI answer appears, not brand visibility, but the lesson carries over.

Two practical points follow. A question set that leans heavily on easy, branded questions will flatter you. And adding many rewordings of one question quietly gives that question more weight unless you correct for it.

## Where does the noise in AI visibility scores come from?

At least four places: chance, wording, which assistant you ask and which language you ask in. [Żatuchin](https://arxiv.org/abs/2607.13304) measured all four on 12,933 answers about 20 Central and Eastern European brands, in eight languages, from three AI models. The point of the design was to see which source of noise a tracking budget should target.

Most tracking tools spend their budget on repeats of the same prompt, typically five. The paper argues that repeats tackle only one source of noise, and the cheapest one: averaging over every query already shrinks it. Language and assistant differences shrink only when you add more languages and assistants.

The author is affiliated with Rankfor.AI, which sells brand-visibility measurement, and says the method is meant for sizing its own measurements. The study ran in one window in spring 2026 and measured the tone of answers rather than whether brands were named.

## Is it better to add more runs or more variety?

Variety wins, by a wide margin in this test. Starting from a common design of one language, one model, five phrasings and five repeats, the study compared where the next queries should go. Three more languages cut measurement error about fifteen times as much as five more repeats. The language option used three times as many extra queries, yet it still came first per query. More models and more phrasings came next; repeats came last.

Piling on repeats barely helped. Twenty repeats of one prompt in one language lifted the brand-ranking reliability score only to 0.020. A design of 15 phrasings asked once each, 360 queries in total, scored higher than one of 5 phrasings asked 5 times each, at 600 queries. Breadth beat depth even while costing less.

The ceiling was low everywhere. On a 0 to 1 scale, a single answer scored about 0.01, and the full design of 8 languages, 3 models and 15 phrasings reached about 0.36. The models were run at a low randomness setting, so real chat apps may vary more between repeats than this study saw. That is why [a single AI visibility snapshot](https://underneath.agency/resources/one-time-ai-visibility-report) is a weak basis for decisions.

## How many runs and questions do you need?

There is no universal number; it depends on the assistant and the topic. Still, several studies give useful floors.

| Study | Setting | Suggested minimum |
|---|---|---|
| [Schulte and colleagues](https://arxiv.org/abs/2604.07585) | 4 engines, Swiss-German prompts, early 2026 | At least 7 runs per prompt per day for brands, 8 for sources, pooled over two to four weeks |
| [Sielinski](https://arxiv.org/abs/2603.08924) | 3 engines, 3 consumer topics, 9 days | About 40 to 50 questions on Gemini, about 100 on Perplexity, 150 or more on ChatGPT search, to pin citation shares to a range about five points wide |
| [Sielinski](https://arxiv.org/abs/2607.10341) | 3 engines, 10 topics, 125 questions each | No budget below 94 responses was enough for every topic; three topics were not settled within 125 |
| [Our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study) | 3 assistants, 20 questions, 5 runs | About 97 runs to pin a brand’s naming rate near 50% to within 10 points |

Sielinski is affiliated with the company IQRush and Schulte with Aurora Intelligence, so treat these as practitioner studies. The first Sielinski paper warns against stopping as soon as the numbers look steady, because the error can narrow and then widen again. Its advice is to fix the number of questions in advance, based on earlier measurements of that assistant and topic.

## Which wordings belong in your prompt set?

Wordings that real buyers use, chosen on purpose, because a different wording is often a different question. In our [rewording study](https://underneath.agency/research/ai-prompt-phrasing-study), asking the identical question again kept the same first brand 68.0% of the time. Adding “on a tight budget” kept it only 15.3% of the time. That change is not noise; the buyer asked for something else.

So separate two kinds of variation. Rewordings that keep the need the same, such as “best” versus “top-rated”, measure how robust your visibility is. Rewordings that add a need, such as a budget, a company size or a region, are new questions that deserve their own line in the report.

Individual prompts also behave very differently. In the Swiss study, some prompts returned almost the same sources run after run, with overlap above 0.8, while others stayed below 0.2. A score built on one or two prompts mostly reflects the quirks of those prompts.

## Should you track the consumer app or the developer API?

Track what buyers see, which usually means the consumer app, or label clearly that you measured the API. In a [Dutch audit of product questions](https://arxiv.org/abs/2609.18729), ChatGPT’s app and its API, asked the same question moments apart, shared only 12.0% of their cited domains on average. Our guide on [monitoring visibility through the API](https://underneath.agency/resources/api-ai-visibility-monitoring) covers that gap in detail.

Assistants also differ in how often they search at all. In the Swiss study, ChatGPT left 57.8% of its runs with no citations because it only searched for some questions. If you track citations, budget for those empty runs. [Sielinski](https://arxiv.org/abs/2607.10341) gives a worked example: needing 100 answers with citations when 83% of queries return them means submitting about 121 queries. Our guide on [measuring share of citations fairly](https://underneath.agency/resources/measuring-share-of-citations-in-ai-search) explains how to count those empty runs.

## What should you do about it?

Design the tracking program before you buy the tool, and write the design down. Steps:

1. Define the answer market: list the buyer questions, the markets and the weight each should carry, and keep branded and unbranded questions separate.
2. Cover the languages and locations your buyers use, then the assistants they use, before adding repeats.
3. Use several real-buyer wordings per need, and report need-changing wordings such as budget or company size as separate questions.
4. Fix the run count in advance. As a floor, use the published minimums above, and pool results over two to four weeks rather than reading daily swings.
5. Measure the consumer app where you can, and record the date, location, assistant and settings with every run.
6. Report each result as a rate with its range, such as “named in 6 of 10 runs”, per assistant and per market.

For help designing a tracking program like this, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The research agrees on the direction but cannot yet give you an exact recipe. Open questions:

- The clearest budget comparison comes from one study of 20 brands in one region and one window, measuring tone rather than whether brands were named. Its author plans a version that measures naming.
- Most studies here were written by people tied to AI-visibility vendors. Their methods are transparent, but independent replication is limited.
- Suggested run counts come from specific topics and markets: Swiss-German consumer queries, US consumer products, Central European brands. B2B, local and regulated categories are largely untested.
- No study has yet tied a prompt set to real buyer question volumes, because AI companies do not publish them. Every prompt set is still an informed guess about demand.

## Frequently asked questions

### How many times should I run each prompt to track AI visibility?

One Swiss study recommends at least 7 runs per prompt per day for brand tracking, pooled over two to four weeks. Another found repeats past five add little once you cover more languages, assistants and wordings.

### How many prompts do I need for AI visibility tracking?

It depends on the assistant. One study needed about 40 to 50 questions on Gemini, about 100 on Perplexity and 150 or more on ChatGPT search before citation shares settled to a range about five points wide.

### Should AI visibility prompts include different phrasings?

Yes. In one test, 15 phrasings asked once each ranked brands better than 5 phrasings asked 5 times each, while using fewer queries.

### Can I track AI visibility through the API instead of the app?

You can, but label it as API data. In one audit, ChatGPT’s app and API shared only 12.0% of their cited domains for the same question asked moments apart.

## Sources

- Żatuchin (2026), [Where Does the Noise Come From? A Variance-Components Decomposition of Non-Determinism in LLM Brand Answers](https://arxiv.org/abs/2607.13304), arXiv:2607.13304.
- Schulte, Bleeker and Kaufmann (2026), [Don’t Measure Once: Measuring Visibility in AI Search (GEO)](https://arxiv.org/abs/2604.07585), arXiv:2604.07585.
- Sielinski (2026), [Quantifying Uncertainty in AI Visibility: A Statistical Framework for Generative Search Measurement](https://arxiv.org/abs/2603.08924), arXiv:2603.08924.
- Sielinski (2026), [From Stochastic to Stable: Rank Stability and Structural Sufficiency in AI Visibility Measurement](https://arxiv.org/abs/2607.10341), arXiv:2607.10341.
- Martinez (2026), [Measuring GEO Visibility: Prompt Corpora Define the Answer Market](https://arxiv.org/abs/2609.06811), arXiv:2609.06811.
- Uberti-Bona Marin and colleagues (2026), ["If I Had to Buy Just ONE: Galaxy S26 Ultra": Auditing AI-Generated Product Recommendations](https://arxiv.org/abs/2609.18729), arXiv:2609.18729.
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)
- Underneath (2026), [Does rewording a question change AI brand recommendations?](https://underneath.agency/research/ai-prompt-phrasing-study)

---

This is the Markdown twin of https://underneath.agency/resources/how-to-design-ai-visibility-tracking. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How can an HR consulting firm win clients through AI search?"
description: "By being the firm AI answers connect to a named people problem, such as pay transparency or work redesign, backed by research and coverage others cite."
canonical: "https://underneath.agency/resources/hr-consulting-firms-clients-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How can an HR consulting firm win clients through AI search?

By becoming the firm that AI answers connect to a specific people problem, such as pay transparency, job architecture or redesigning work around AI, and backing that link with research and coverage that other sites cite. HR leaders now bring these problems to AI assistants before they call anyone, and the firms that publish the clearest, most quoted thinking on a problem are the ones a reasonable buyer would expect to see named.

## The short version

1. The deadline is a buying trigger: EU member states had to put the Pay Transparency Directive into national law by [7 June 2026](https://commission.europa.eu/strategy-and-policy/policies/justice-and-fundamental-rights/gender-equality/equal-pay/eu-action-equal-pay_en), and in an IBEC survey only [1% of organizations felt fully prepared](https://www.lawsociety.ie/gazette/top-stories/2025/october/20-of-firms-not-ready-for-key-eu-pay-directive/).
2. The people problems are expensive: Gallup puts global employee engagement at 20%, with low engagement costing about $10 trillion in lost productivity.
3. Pay money is tight, so advice on spending it matters: WTW found US salary budgets for 2026 holding at 3.4%, with 24% of organizations reporting trouble attracting or retaining employees.
4. Advice is high-value work: Korn Ferry’s Consulting business billed $691.7 million in fiscal 2026 at an average bill rate of $458 an hour.
5. HR teams already work with AI: in SHRM’s survey of 1,908 HR professionals, 39% had AI adopted in their HR function.

This article is for firms that sell HR advice: compensation, organization design, workforce planning, leadership and change. If you sell HR software, our article on [whether AI search decides which HR software gets the demo](https://underneath.agency/resources/hr-software-ai-search) covers that market. Recruiters can see [how agencies win employer clients through AI](https://underneath.agency/resources/recruiting-agencies-employer-clients-ai-search).

## Who hires an HR consulting firm, and what is one client worth?

A CHRO, head of rewards, CFO or CEO with a deadline or a problem their own team cannot staff. One engagement often leads to years of follow-on work.

The buyers are senior. Gartner’s advice to chief HR officers for 2026, reported by [HRD America](https://www.hcamag.com/us/specialization/leadership/what-should-hrs-priorities-be-for-2026/551961), put an HR-focused AI strategy among the top priorities, alongside helping leaders make change routine and helping managers hold better productivity conversations. Mercer’s [Global Talent Trends 2026](https://www.mercer.com/insights/people-strategy/future-of-work/global-talent-trends/) drew on nearly 12,000 executives, HR leaders, employees and investors across 16 geographies, and its four trends read like a list of consulting briefs: redesigning work, talent intelligence, the employee value proposition and a new HR operating model.

The value of a client shows in public filings. Korn Ferry, which calls itself a global consulting firm, reported in its [fiscal 2026 results](https://ir.kornferry.com/sec-filings/all-sec-filings/content/0000056679-26-000017/kfy-20260430xex991q4fy26.htm) that its Consulting business earned $691.7 million in fee revenue from 1,522 consultants and execution staff, at an average bill rate of $458 an hour. It also held $390.1 million of estimated remaining fees under signed consulting contracts. Smaller firms charge less, but the shape is the same: a scoped project, then annual pay cycles, surveys and refreshes. Firms that also run leadership searches should read [how boards now check executive search firms](https://underneath.agency/resources/executive-search-firms-clients-ai-search).

## What is pushing employers to hire HR advisers now?

Regulation, AI-driven work redesign, weak engagement and flat pay budgets, each of which creates a defined project.

**Pay transparency.** The European Commission says member states must transpose the directive by 7 June 2026, and it has since issued EU-wide guidelines on gender-neutral job evaluation. In IBEC’s survey of more than 391 HR professionals in Ireland, 22% said preparations had yet to begin, up from 12% in 2024. The hardest tasks were grouping roles of equal value (25%) and making pay criteria accessible to employees (21%). Those are job architecture and pay equity projects. The Irish recruiter Sigmar reports [the sharpest increase in demand](https://www.sigmarrecruitment.com/resources/download/human-resources-job-market-outlook-2026--talent-retention--hybrid-working-and-skills-demand/) for workforce planning and HR project management expertise, “especially for projects related to the upcoming EU Pay Transparency Directive,” and higher salary bands for consulting work.

**Engagement.** Gallup’s [State of the Global Workplace](https://www.gallup.com/workplace/349484/state-of-the-global-workplace.aspx) found engagement fell to 20% in 2025, its lowest level since 2020, and manager engagement dropped from 27% to 22% in a year. Gallup estimates the cost at 9% of global GDP.

**Pay strategy.** In [WTW’s Salary Budget Planning Survey](https://www.wtwco.com/en-us/news/2026/01/salary-budgets-have-stabilised-as-employers-focus-on-pay-strategy-for-2026), with 1,876 US organizations responding, 21% of employers planned to cut pay budgets, and WTW described a move from spreading budget across everyone to “strategic use of each dollar.” Deciding where each dollar goes is compensation consulting.

## Where does AI search sit in how HR leaders find advisers?

Mostly at the start, where a leader works out what the problem is and who handles it. Direct evidence for HR consulting is still thin.

We found no published study of how HR leaders use AI assistants to choose consultants. What exists is adjacent:

- **HR teams use AI daily.** SHRM’s [State of AI in HR 2026](https://www.shrm.org/content/dam/en/shrm/research/ai-in-hr_executive-summary.pdf) found 62% of respondents’ organizations using AI somewhere, and 60% of extra-large organizations had AI in their HR function, against 33% of small ones.
- **Business buyers research vendors with AI.** In a Gartner survey of 645 B2B buyers, [45% said they used generative AI](https://www.gartner.com/en/newsroom/press-releases/2026-05-20-gartner-survey-finds-sixty-nine-percent-of-b-two-b-buyers-turn-to-sales-reps-to-validate-ai-generated-insights) in a recent purchase, mainly to gather information on vendors, and 69% preferred to check AI-generated insights with a salesperson. These buyers were mostly buying products, not advice.
- **Service directories are moving into the chat.** In May 2026, Clutch, a review platform for B2B service providers, [launched an app inside ChatGPT](https://www.demandgenreport.com/?p=52857) so buyers can compare firms, reviews and pricing signals without leaving the conversation.

We infer that the HR buyer’s path looks like other professional services: AI helps define the problem and names a first set of firms, then peers and a scoping call decide. In Responsive’s survey of buyers, [industry expertise carried significant weight for 52%](https://www.responsive.io/news/buyer-intelligence-2025), ahead of price at 49%. For an HR consultancy, expertise on the specific problem is the product.

## What do HR leaders ask AI before they call a consultancy?

Questions that pair a problem with their situation: country, size, industry or deadline. The examples below are our own wording, written to show the pattern; none comes from logged searches.

| Situation | Example question |
|---|---|
| EU staff, US headquarters | “What does the EU pay transparency directive require of a US company with 300 employees in Germany and Ireland, and who helps with job leveling?” |
| Fast-growing company | “Compensation consultants for a Series C software company setting pay bands for the first time” |
| After an AI rollout | “Firms that help redesign roles and org structure after automating customer service” |
| Engagement slump | “Best consultancies for manager effectiveness programs at a 5,000-person manufacturer” |
| Mid-size employer | “Mercer vs WTW vs a boutique firm for a pay equity audit: pros and cons” |
| Due diligence | “Is [firm name] good at change management? What do clients say?” |

Two things follow. The questions are about problems, not about “HR consulting” in general, so a firm that is described only as “full-service HR consulting” gives an assistant little to match. And the comparison question puts global firms beside smaller ones, which is where a specialist with a sharp point of view can be named. Our guide on [how specialist consulting firms get shortlisted](https://underneath.agency/resources/consulting-firms-clients-ai-search) covers the same pattern outside HR.

## How does an AI answer turn into a consulting engagement?

Through a short chain: an answer names the firm, the leader reads its view, then books a scoping call.

The path is an AI answer → the firm’s practice page, report or article → a scoping call → a proposal, often against two or three other firms → a project → annual work. That last step is why one client matters: pay cycles, engagement surveys and job architecture all need refreshing.

Two features make HR consulting different from buying software. First, there is no free trial. A buyer judges a consultancy by its thinking, so the content an assistant cites, such as a readiness checklist or a survey of employers, does the job a demo does for software. Second, the people matter. A CHRO hires a team, and named consultants with credentials and a record on the problem are what the buyer checks after the AI answer. A reasonable expectation is that firms whose practice pages name their consultants and show results with client permission convert more of that interest.

## What decides whether an assistant names an HR consulting firm?

Mostly what others write about the firm, and whether its own pages answer the exact question. Platforms say little; most of what we know is observed.

**Documented by the platform.** OpenAI’s help page on [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search) says ChatGPT may search the web automatically, that answers may include citations, and that it typically rewrites a question into “one or more targeted queries” for its search partners. So an answer is built from pages that matched those rewritten searches.

**Observed in our studies.** These are cross-industry; none was about HR consulting.

- **Rankings and named publications get searched.** In our [hidden searches study](https://underneath.agency/research/ai-hidden-searches-study), ChatGPT ran a mean of 3.7 searches per answer, and in 43.8% of its answers one search was aimed at a named publication, ranking or award. Industry rankings, trade press lists and awards are therefore part of the evidence an answer may draw on.
- **Independent coverage predicts recommendations.** In our [brand entity study](https://underneath.agency/research/brand-entity-ai-recommendations-study), each tenfold increase in the number of independent sites naming a brand in the cited pages went with 4.7 times the odds of being recommended.
- **Lists can be self-serving.** Of the numbered “best” lists AI answers cited, [24.2% ranked their own publisher first](https://underneath.agency/research/self-promoting-best-lists-study). Buyers may meet consultancy lists written by consultancies.
- **Reputation answers cite review sites.** When asked whether a brand was legitimate, [88.0% of answers cited a review or complaint platform](https://underneath.agency/research/is-it-legit-ai-reputation-study).

**Our inference for HR consulting.** The large firms publish the data others quote. WTW’s salary budgets, Mercer’s talent trends and Gallup’s engagement figures appear across trade press, which gives assistants many independent pages naming them. A boutique cannot match that volume, but it can own a narrower problem: one survey of employers on pay transparency readiness in a specific country, quoted by trade press, ties its name to that question. [Supply chain consultancies face the same contest](https://underneath.agency/resources/supply-chain-consulting-firms-ai-search) with larger rivals.

## What does it cost an HR consultancy to be missing from AI answers?

Fewer invitations to scope the work, especially on new problems where buyers have no adviser yet. No study has measured the loss.

The risk is greatest on fresh problems. Pay transparency and AI work redesign are new to most HR teams, so buyers lack a trusted name and are more likely to ask a general tool for one. If the answer names three global firms and two software vendors, a specialist that has done the work may never be asked. That is our inference. For firms advising on the technology side of AI rollouts, see [how technology consultancies reach enterprise shortlists](https://underneath.agency/resources/technology-consulting-firms-ai-search).

The second cost is being described wrongly. Assistants do not always agree on who to name: across four assistants, [the options recommended for the same question overlapped by only 0.327 on average](https://underneath.agency/research/ai-assistants-brand-agreement-study). A firm checked on one assistant has not been checked on the others. If an assistant describes the firm as a staffing agency or a payroll provider, the buyer with a pay equity question moves on. Our guide on [fixing wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) covers the repair.

## How does GEO work for an HR consulting firm?

Generative engine optimization (GEO) makes a firm’s expertise on specific people problems easy for assistants to find, confirm and repeat. It cannot guarantee a mention.

1. **A page per problem, not per service line.** “EU pay transparency readiness for US employers” is a question buyers ask; “total rewards solutions” is not. State who the work is for, what it involves, how long it takes and which consultants lead it.
2. **Original data that others cite.** A short annual survey of employers on one topic, published with its method, is the boutique version of WTW’s salary budgets. It gives trade press a reason to name the firm. Our guide on [how brands build authority that AI search recognizes](https://underneath.agency/resources/how-brands-build-authority-for-ai-search) explains why independent coverage counts most.
3. **Named experts.** Consultant pages with credentials, published work and the problems each one handles, the same across the website, LinkedIn and conference bios.
4. **Third-party profiles and rankings.** Directory profiles such as Clutch, association listings and any independent rankings the firm qualifies for, with consistent descriptions. See [which pages to target for AI recommendations](https://underneath.agency/resources/best-of-lists-ai-recommendations).
5. **Client evidence with permission.** Case studies that name the problem, the size of the employer and the outcome in plain terms, without confidential detail.
6. **Regular checks.** Ask the problem questions above across ChatGPT, Gemini, Perplexity, Claude and Google’s AI features each quarter, several times each, and record who is named and which pages are cited.

## What can’t the current research tell an HR consultancy?

Whether AI visibility turns into signed engagements, and how HR leaders specifically use AI to pick advisers.

- **No HR-consulting study of AI answers.** We found none that measures which HR consultancies assistants name.
- **Buyer surveys are general.** The Gartner and Responsive findings cover B2B purchases broadly, mostly of products.
- **The directive is moving.** Countries are transposing it on different timelines, so readiness questions and answers will keep changing.
- **Answers vary.** Our studies show lists change between assistants and between runs, so a single check is a snapshot.

## Where should an HR consulting firm start?

Pick the two or three problems you most want to be hired for, and see who AI names today.

Write the questions a CHRO or CFO would ask about each problem, with country, size and deadline. Ask them in several assistants, more than once, and note which firms appear, which pages are cited and how your firm is described. The gap between that picture and the clients you want is the work plan.

If you want more scoping calls on pay transparency, rewards or work redesign from buyers who are researching with AI, [talk to us about an AI visibility review for your practice areas](https://underneath.agency/contact). We will test how assistants answer your buyers’ questions, show which sources they rely on, and plan the pages, research and coverage that connect your firm to the problems you solve. You can see how such a program runs for a consultancy, from the first diagnosis of buyer questions to ongoing measurement, on our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page.

## Frequently asked questions

### Do AI assistants recommend boutique HR consultancies or only Mercer and WTW?

There is no HR-specific study yet. Across industries, independent coverage was the strongest signal we measured, so a boutique that is widely cited on one narrow problem has a realistic opening.

### Does publishing salary or workforce surveys help an HR firm appear in AI answers?

It is one of the most direct routes. Surveys give trade press and other sites a reason to name the firm, and independent mentions went with higher odds of being recommended in our study.

### Should an HR consultancy create pages about the EU Pay Transparency Directive?

If it does that work, yes. The transposition deadline was 7 June 2026 and few organizations felt ready, so buyers are asking. Keep pages factual; they are not legal advice.

### How often should a firm check what AI says about it?

Quarterly at least, with several runs per question. Answers differ between assistants and between runs, so one check can mislead.

## Sources

- European Commission (n.d.), [EU action for equal pay](https://commission.europa.eu/strategy-and-policy/policies/justice-and-fundamental-rights/gender-equality/equal-pay/eu-action-equal-pay_en)
- Law Society Gazette, Ireland (2025-10-23), [20% of firms not ready for key EU pay directive](https://www.lawsociety.ie/gazette/top-stories/2025/october/20-of-firms-not-ready-for-key-eu-pay-directive/)
- Sigmar Recruitment (2025-12-22), [Human Resources Job Market Outlook 2026](https://www.sigmarrecruitment.com/resources/download/human-resources-job-market-outlook-2026--talent-retention--hybrid-working-and-skills-demand/)
- Gallup (2026), [State of the Global Workplace](https://www.gallup.com/workplace/349484/state-of-the-global-workplace.aspx)
- WTW (2026-01-21), [Salary budgets have stabilized as employers focus on pay strategy for 2026](https://www.wtwco.com/en-us/news/2026/01/salary-budgets-have-stabilised-as-employers-focus-on-pay-strategy-for-2026)
- Korn Ferry (2026-06-23), [Fourth quarter and full year FY’26 results](https://ir.kornferry.com/sec-filings/all-sec-filings/content/0000056679-26-000017/kfy-20260430xex991q4fy26.htm)
- HRD America (2025-10-05), [What should HR’s priorities be for 2026?](https://www.hcamag.com/us/specialization/leadership/what-should-hrs-priorities-be-for-2026/551961)
- Mercer (2026), [Global Talent Trends 2026](https://www.mercer.com/insights/people-strategy/future-of-work/global-talent-trends/)
- SHRM (2026), [The State of AI in HR 2026: Executive Summary](https://www.shrm.org/content/dam/en/shrm/research/ai-in-hr_executive-summary.pdf)
- Gartner (2026-05-20), [Gartner Survey Finds 69% of B2B Buyers Turn to Sales Reps to Validate AI-Generated Insights](https://www.gartner.com/en/newsroom/press-releases/2026-05-20-gartner-survey-finds-sixty-nine-percent-of-b-two-b-buyers-turn-to-sales-reps-to-validate-ai-generated-insights)
- Responsive (2025), [GenAI overtakes search for a quarter of B2B buyers](https://www.responsive.io/news/buyer-intelligence-2025)
- Demand Gen Report (2026-05-12), [Clutch Launches First B2B Services Marketplace App on ChatGPT](https://www.demandgenreport.com/?p=52857)
- OpenAI (2026), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [Do Wikipedia and schema make AI assistants recommend a brand?](https://underneath.agency/research/brand-entity-ai-recommendations-study)
- Underneath (2026), [How many “best of” lists cited by AI rank their own brand first?](https://underneath.agency/research/self-promoting-best-lists-study)
- Underneath (2026), [“Is this brand legit?” How AI assistants build a reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study)
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude agree on brands?](https://underneath.agency/research/ai-assistants-brand-agreement-study)

---

This is the Markdown twin of https://underneath.agency/resources/hr-consulting-firms-clients-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Is AI search deciding which HR software gets the demo?"
description: "Increasingly it shapes who gets considered. HR software vendors win demos when AI answers can verify their fit, security and compliance."
canonical: "https://underneath.agency/resources/hr-software-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Is AI search deciding which HR software gets the demo?

Not yet on its own, but it increasingly shapes the short list of vendors an HR leader invites to demo. HR teams already use AI in their own work, the first list of vendors usually decides the purchase, and the pages AI answers draw on are often written by vendors with a stake in the ranking. For an HR software company, being named and described correctly in those answers is now part of winning the demo, and the multi-year contract that follows.

## The short version

1. HR software is bought with long, high-value contracts: Workday reported more than 11,500 customers, fiscal 2026 subscription revenue of $8.833 billion and $28.101 billion of total subscription revenue backlog.
2. HR teams are adopting AI: in SHRM’s 2026 survey of 1,908 HR professionals, 39% had AI in their HR functions, and recruiting was the most common use, at 27%.
3. HR buyers screen hard for risk: in Capterra’s survey of 3,256 HR professionals in 11 countries, 67% call security a critical feature when researching HR software, and 46% are concerned with assessing AI’s value and risks.
4. The sources AI answers cite are not neutral: in our study of 269 “best” lists cited by AI search, 24.2% ranked their own publisher first, and Rippling was among the most frequent self-rankers, with three lists.
5. In our study of Google’s AI Overviews, searches about B2B software cited a YouTube video on 91.0% of searches, with a channel named HR Toolbox among the most-cited channels.

## Who picks an HR platform, and how much revenue does one contract carry?

HR leaders choose it with IT, finance and procurement, and a won customer usually signs a multi-year subscription.

The market runs from small businesses buying an all-in-one HR and payroll tool to global employers replacing a core human capital management (HCM) system. At the top, [Workday’s fiscal 2026 results](https://newsroom.workday.com/2026-02-24-Workday-Announces-Fiscal-2026-Fourth-Quarter-and-Full-Year-Financial-Results) show more than 11,500 customers, including more than 7,000 core Workday Financial Management and HCM customers, and total revenues of $9.552 billion. Its total subscription revenue backlog, contracted revenue not yet recognized, was $28.101 billion. That is what a long HR contract looks like: revenue booked years ahead.

Below the enterprise tier, newer platforms compete on breadth and speed. [Deel](https://www.deel.com/about/) says it is trusted by 40,000+ companies; [Rippling](https://www.rippling.com/company/about) describes itself as integrated with 600+ apps. Their pitch, one system for HR, payroll, IT and global hiring, means buyers compare platforms that bundle very different things under the same label. Payroll specialists compete in that overlap too; see [how payroll providers get named by AI](https://underneath.agency/resources/payroll-software-ai-search).

What HR leaders buy most is the core record. In [Capterra’s 2025 HR Software Trends Survey](https://www.capterra.com/resources/hr-technology-trends/) of 3,256 HR professionals, 50% rated an HR information system (HRIS) as critical to their HR operations. The same survey found 47% saying effective implementation is a challenge, and 48% naming training new users on HR software as their main software-related challenge. Implementation pain is why HR systems are replaced rarely, and why the first choice carries so much value.

## How far into HR buying have AI tools reached?

Inside HR teams’ daily work first; direct evidence that HR leaders shortlist vendors with chatbots is still thin.

AI adoption in HR is real but uneven. [SHRM’s State of AI in HR 2026](https://www.shrm.org/content/dam/en/shrm/research/ai-in-hr_executive-summary.pdf), based on 1,908 HR professionals, found 39% with AI adopted in their HR functions and 62% using AI somewhere in their organizations. Most extra-large organizations (60%) had AI in HR, against 33% of small and 35% of midsize ones. Capterra found that over half of surveyed organizations already use AI features in their HR software.

How software buyers in general build lists matters here. Gartner Digital Markets’ [2025 survey of 3,500 software buyers](https://cdn-static.bizzabo.com/bizzabo.file.upload/yVkIqwmWRDKbEhqeE3aQ_Gartner-Digital-Markets_2025-Report_Making-The-List.pdf) across industries found buyers start with an average of 4.4 options on their initial list, and 81% end up buying from that list most or all of the time. Buying teams are mixed: 39% use a formal team from several departments, which in HR usually means HR, IT and finance. Our construction software guide applies the same Gartner data to small contractors in [how contractors build a software shortlist](https://underneath.agency/resources/construction-software-customers-ai-search).

Video is part of Google’s layer: across the B2B software searches in [our YouTube study](https://underneath.agency/research/ai-overview-youtube-videos-study), Google’s AI Overviews cited a YouTube video on 91.0% of searches. The channels cited most often across the study were specialists, and a channel named HR Toolbox was among them, with four cited videos. No published survey yet measures how many HR leaders start a vendor search in ChatGPT or Gemini; we say so plainly rather than borrow a figure from another industry. For the cross-category survey evidence, see [our B2B SaaS article](https://underneath.agency/resources/b2b-saas-revenue-from-ai-search).

## What do HR leaders ask AI assistants before booking a demo?

Fit questions by company size and country, alternatives to an incumbent, side-by-side comparisons, and risk and compliance questions.

We wrote the prompts below ourselves, modeled on how an HR director or people-operations lead might frame a search; none were observed:

- **Fit:** “Best HRIS for a 400-person company with employees in 12 US states and Canada.”
- **Alternatives:** “Alternatives to Workday for a mid-market company that needs to go live in under six months.”
- **Comparison:** “BambooHR vs HiBob vs Rippling for a fast-growing tech company.”
- **Global hiring:** “Which platform can hire and pay employees in Germany without a local entity?”
- **Compliance:** “Which applicant tracking systems have completed a bias audit for New York City’s AI hiring law?”
- **Integration:** “Which HR platforms integrate with NetSuite, Greenhouse and Okta?”

The compliance questions are specific to this category. New York City’s [Local Law 144 of 2021](https://www.nyc.gov/site/dca/about/automated-employment-decision-tools.page) bars employers from using an automated employment decision tool unless it has had a bias audit within one year of use. The EU AI Act lists AI systems used for recruitment, filtering job applications and evaluating candidates among its [high-risk uses](https://artificialintelligenceact.eu/annex/3/). An HR vendor with AI features will be asked about both, and an assistant can only answer from what is published.

## How does an AI answer turn into an HR software demo and contract?

Through the shortlist: the answer names you, the buying team checks you, you get the demo, then the security review.

**Shortlist.** If the initial list decides the purchase most of the time, an AI answer that names three vendors for a 400-person, multi-state company is effectively choosing who gets considered. We infer that for HR software this matters most for mid-market buyers without a consultant running the selection. Finance tools follow a similar path, as our look at [AI and accounting software choices](https://underneath.agency/resources/accounting-software-ai-search) describes.

**Demo and validation.** HR buyers check risk before they sign. In Capterra’s survey, 67% consider security a critical feature when researching HR software, and 43% said security concerns triggered HR purchases in the previous year. A reasonable expectation is that an assistant’s answer to “is this platform compliant” pulls from your public security, privacy and AI pages, or from someone else’s description of them.

**Contract.** The prize is a multi-year subscription that is painful to replace, which is why Workday can report subscription backlog of $28.101 billion. For smaller vendors the same logic holds at a smaller scale: a demo won through an AI shortlist can turn into years of per-employee revenue, plus payroll, benefits and other modules sold later.

## Why does an assistant mention one HR platform and leave out another?

Platforms document how they search; studies show what they cite; most HR-specific trust factors are still inference.

**Documented by the platforms.** [Google’s guidance](https://developers.google.com/search/docs/appearance/ai-features) explains that AI Overviews and AI Mode may use “query fan-out”, issuing multiple related searches across subtopics, and that appearing needs nothing technical beyond what normal search requires. A question about multi-state HRIS can therefore pull in pages on pricing, payroll, compliance and reviews at once.

**Observed in our studies.** [Our self-ranking lists study](https://underneath.agency/research/self-promoting-best-lists-study) found that of 269 numbered “best” lists with an identifiable publisher cited by six AI surfaces, 24.2% ranked their own publisher first. Rippling and Zendesk published the most such lists, three each, and 92.9% of lists that included their publisher put it at number one. In HR software, where vendors publish many “best HR software” pages, buyers and assistants are reading lists written by competitors. In [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), B2B software had the most stable brands across repeated runs (0.708 on a 0-to-1 scale), so whoever is named tends to stay named. Our guide on [why AI keeps naming the same CRMs](https://underneath.agency/resources/crm-software-ai-search) shows how challengers break in.

**Our inference about HR-specific trust factors.** We expect assistants to rely on what HR buyers themselves trust: independent review platforms, analyst coverage, HR practitioner communities, published security certifications, bias audit summaries for AI features, and clear statements of supported countries, company sizes and integrations. No AI platform has named any of them as something it weighs when choosing HR vendors.

## How much does an HR platform lose when an AI shortlist leaves it off?

It usually loses the whole evaluation, and with it years of contract value.

When most purchases come from the first list of vendors, a platform left out of the AI answer rarely gets a second chance in that cycle. Because HR systems are hard to implement and replace, the next chance may be years away. The cost also shows up as wrong information: an answer that says your platform lacks a country, an integration or an AI audit can remove you from a list without anyone telling you.

The self-ranking lists finding adds a specific risk for HR. If a competitor’s “best HR software” page ranks itself first and gets cited, its framing of the category reaches the buyer before yours does. Our article on [which pages to target](https://underneath.agency/resources/best-of-lists-ai-recommendations) explains why independent lists matter more than your own.

## What does GEO involve for an HRIS or HCM vendor?

Generative engine optimization (GEO) means making your HR platform’s fit, compliance and proof easy for AI search to verify. No one can guarantee a mention.

1. **A clear entity.** Describe who you serve, by company size, industry and country, in the same words on your site, review profiles, partner marketplaces and press pages.
2. **Public compliance and security pages.** Publish security certifications, data residency, AI governance and bias audit summaries in plain, ungated pages, since 67% of HR buyers treat security as critical and assistants can only cite what is public.
3. **Independent coverage.** Earn mentions in practitioner communities, HR analyst research, podcasts and independent comparison lists, not only on your own blog.
4. **Fair comparison content.** Publish honest “alternatives” and comparison pages that state where you fit and where you do not; see [our summary of comparison-page research](https://underneath.agency/resources/do-comparison-pages-help-b2b-ai-citations).
5. **Video and reviews.** Product walkthroughs and customer reviews are sources AI search already cites in this category.
6. **Measurement tied to demos.** Track a fixed set of HR buyer questions across ChatGPT, Gemini, Perplexity, Copilot and Google, and compare it with demo requests and how new customers say they found you. Our guide on [what to measure](https://underneath.agency/resources/what-to-measure-ai-visibility) sets out the options.

## Which gaps in the evidence matter most for HR software vendors?

The biggest: no published study yet links AI visibility to HR software demos or contracts.

- We found no survey measuring how many HR leaders use AI assistants to find or compare HR vendors. The adoption figures above describe AI inside HR work, not vendor research.
- The cross-industry buying data from Gartner Digital Markets covers all software, not HR alone.
- Vendor figures such as customer counts are self-reported.
- Our self-ranking and consistency findings come from US data collected on one day in September 2026, across categories rather than HR alone.
- Whether a compliance page changes what an assistant says about an HR platform has not been tested.

## How can an HR software vendor find out whether AI answers are costing it demos?

Check which AI answers include you for the questions that come before your demo requests.

Run your buyers’ fit, alternatives, comparison and compliance questions across the main assistants, note who is named and whether your countries, integrations and security facts are right, and rank the gaps by the deal size of the segment that asks them. If you would rather have that done from outside, with a plan measured in demos and pipeline, [talk to us about an HR software visibility audit](https://underneath.agency/contact). Our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page covers what that ongoing work includes, from public compliance pages and independent coverage to tracking buyer questions against demo requests.

## Frequently asked questions

### Do AI assistants recommend Workday, Rippling or BambooHR?

They often name well-known platforms for general HR questions, but which ones depends on the question’s size, country and needs. We have not published HR-specific recommendation counts, so test your own buyer questions rather than relying on general claims.

### Does a “best HR software” page on our own site help?

It may get cited, but buyers and assistants increasingly see self-ranking lists. In our study, 24.2% of cited numbered lists ranked their own publisher first. Independent lists and reviews are stronger evidence.

### Should HR vendors publish bias audit results for AI features?

Where the law requires it, yes, and New York City’s Local Law 144 already does. Beyond compliance, a public summary gives assistants and buyers a fact to cite instead of a guess.

### How long does it take for GEO to affect HR software demos?

There is no published benchmark for this category. Track your buyer questions monthly and compare changes with demo requests over several quarters, because HR buying cycles are long.

## Sources

- Workday (2026), [Workday Announces Fiscal 2026 Fourth Quarter and Full Year Financial Results](https://newsroom.workday.com/2026-02-24-Workday-Announces-Fiscal-2026-Fourth-Quarter-and-Full-Year-Financial-Results)
- SHRM (2026), [The State of AI in HR 2026: Executive Summary](https://www.shrm.org/content/dam/en/shrm/research/ai-in-hr_executive-summary.pdf)
- Capterra (2025), [Capterra’s 2025 HR Software Trends](https://www.capterra.com/resources/hr-technology-trends/)
- Gartner Digital Markets (2025), [Making the List: How Software Buyers Pare Down Their Options](https://cdn-static.bizzabo.com/bizzabo.file.upload/yVkIqwmWRDKbEhqeE3aQ_Gartner-Digital-Markets_2025-Report_Making-The-List.pdf)
- Deel (2026), [About Deel](https://www.deel.com/about/)
- Rippling (2026), [About Rippling](https://www.rippling.com/company/about)
- NYC Department of Consumer and Worker Protection, [Automated Employment Decision Tools](https://www.nyc.gov/site/dca/about/automated-employment-decision-tools.page)
- EU Artificial Intelligence Act, [Annex III: High-Risk AI Systems](https://artificialintelligenceact.eu/annex/3/)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- Underneath (2026), [Self-ranking “best” lists](https://underneath.agency/research/self-promoting-best-lists-study), [YouTube videos in AI Overviews](https://underneath.agency/research/ai-overview-youtube-videos-study) and [recommendation consistency](https://underneath.agency/research/ai-recommendation-consistency-study)

---

This is the Markdown twin of https://underneath.agency/resources/hr-software-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Will AI name your identity platform when buyers start looking?"
description: "Identity vendors get named in AI answers when their standards support, security record and facts are public, consistent and confirmed by independent sources."
canonical: "https://underneath.agency/resources/iam-enterprise-demand-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Will AI assistants name your identity platform when a credential breach sends buyers looking?

Only if the facts buyers check, standards support, security record, integrations and price, are public, consistent and confirmed by independent sources. Stolen credentials are now a leading way into companies, so identity purchases often start with an incident and a deadline. The first places a buyer looks, increasingly AI answers, shape a shortlist that is then tested hard by security reviews and procurement.

## The short version

1. Identity is the attacker’s favorite entry point: [Verizon’s 2025 breach report](https://verizon.com/about/news/2025-data-breach-investigations-report) found credential abuse was the leading initial attack vector, at 22%, and [CrowdStrike](https://www.crowdstrike.com/en-us/press-releases/crowdstrike-releases-2025-global-threat-report/) reported that 79% of attacks to gain initial access were malware-free.
2. Breaches that start with compromised credentials cost an average of $4.67 million, and phishing, often aimed at logins, was the most common initial vector at 16%, according to [IBM’s 2025 report](https://www.bakerdonelson.com/webfiles/Publications/20250822_Cost-of-a-Data-Breach-Report-2025.pdf).
3. Identity and access management spending grows from $20.7 billion in 2025 to $23.4 billion in 2026, an 11.8% rise, in Gartner’s forecast as summarized by [Louis Columbus](https://softwarestrategiesblog.com/2026/04/01/top-10-fastest-growing-security-categories-gartner-2026-forecast/).
4. Identity platforms are prized: [Palo Alto Networks](https://www.paloaltonetworks.com/company/press/2025/palo-alto-networks-announces-agreement-to-acquire-cyberark--the-identity-security-leader) agreed to buy CyberArk at an equity value of approximately $25 billion, a 26% premium.
5. A new buying trigger is arriving: in [Okta’s Businesses at Work report](https://www.okta.com/reports/businesses-at-work/), 78% cited controlling access for non-human identities as a top concern and only 10% had a strategy for governing them.

## Who buys identity and access management, and what is a customer worth?

A CISO or CIO signs, IAM architects run the evaluation, and procurement tests the vendor’s own security.

Identity and access management (IAM) covers several products that enterprises increasingly buy together: workforce single sign-on and multi-factor authentication (MFA), identity governance and administration (IGA) for joiners, movers and leavers, privileged access management (PAM) for administrator accounts, customer identity for consumer logins, and newer tools that detect identity attacks. Gartner’s 1Q26 forecast, as summarized by Columbus, puts the IAM subsegment at $20.7 billion in 2025 and $23.4 billion in 2026. Identity governance grows faster than the subsegment, at 14.7%, and access management at 14.5%.

The buying group is wide because identity touches everything:

- **The CISO**, who owns the risk after a credential-based incident.
- **The CIO and IT**, who run the directory and every application connected to it.
- **IAM architects**, who check standards, integrations and migration effort.
- **Audit and compliance**, who need access reviews and evidence for regulators.
- **Procurement and vendor risk**, who send the identity vendor a security questionnaire, since a compromised identity provider exposes every connected system.

A won customer tends to be large and durable. Once an identity platform is wired into a company’s applications, replacing it is a major project, so contracts tend to renew and expand into governance and privileged access, we infer. The market values that position highly. Palo Alto Networks said the CyberArk deal would establish “Identity Security as a new core platform,” and its CEO described identity security as a category at its “inflection point.”

## What sends enterprises shopping for identity security?

Credential-based breaches, regulatory pressure for stronger authentication, and the arrival of AI agents with their own access.

**Breaches that begin with a login.** Verizon analyzed 12,195 confirmed data breaches; credential abuse (22%) and exploitation of vulnerabilities (20%) were the leading ways in. CrowdStrike’s 2025 threat report found a 442% increase in voice phishing between the first and second halves of 2024, and said valid account abuse accounted for 35% of cloud incidents in the first half of 2024. Vendors selling the cloud side of that defense can read [how cloud security vendors reach enterprise shortlists](https://underneath.agency/resources/cloud-security-enterprise-buyers-ai-search).

The 2024 attacks on Snowflake customers showed the pattern plainly. [Mandiant](https://cloud.google.com/blog/topics/threat-intelligence/unc5537-snowflake-data-theft-extortion) reported that the attackers used stolen customer credentials, mostly from infostealer malware, and that the affected accounts “were not configured with multi-factor authentication enabled.” Mandiant and Snowflake notified approximately 165 potentially exposed organizations. Mandiant found no evidence of a breach of Snowflake’s own enterprise environment.

**Pressure for stronger authentication.** [CISA](https://www.cisa.gov/sites/default/files/publications/fact-sheet-implementing-phishing-resistant-mfa-508c.pdf) “strongly urges all organizations to implement phishing-resistant MFA,” and notes that the Office of Management and Budget requires federal agencies to adopt it. In Okta’s customer data, high-assurance MFA adoption grew from 41% to 58%. For the network side of these projects, see [how network security vendors win zero trust buyers](https://underneath.agency/resources/network-security-zero-trust-customers-ai-search).

**AI agents.** AI agents need accounts and permissions too. Okta’s report found 78% of organizations cite controlling non-human identity access as a top concern, and Palo Alto Networks framed its CyberArk deal partly around securing “autonomous AI agents.” Okta and Palo Alto Networks both sell identity products, so treat their framing as interested.

Each trigger arrives with urgency: an audit finding, a board question or an incident review. That urgency, we infer, pushes buyers to fast research tools such as AI assistants.

## Where does AI search sit in an identity purchase?

Early, during research and shortlisting; no public study isolates identity buyers.

The broadest recent evidence covers technology buyers. [TrustRadius’s 2026 B2B Buying Disconnect Report](https://www.demandgenreport.com/industry-news/news-brief/trustradius-ai-has-changed-how-buyers-research-but-not-what-they-trust/53853/) found that 63% of buyers used AI during their purchase journey, but 94% of them fact-checked its answers at least some of the time and 74% used reviews to inform their decisions. Analyst reports were used by only 13%. As a seller of review visibility, TrustRadius has an interest in that finding.

For identity, the fact-checking step is unusually formal. A shortlisted vendor will face a security questionnaire, a review of its certifications and incident history, and an architecture check against the buyer’s directory and applications. An AI answer can get a vendor into that process; it cannot carry the vendor through it. Endpoint vendors meet a similar test, described in [how EDR vendors win when CISOs ask AI](https://underneath.agency/resources/endpoint-security-enterprise-customers-ai-search).

## Which questions do identity buyers ask AI assistants?

Questions about platforms, standards, migration, governance, AI agents and the vendor’s own security. We wrote these prompts to illustrate identity buying; they were not observed from real IAM teams.

| Buying need | Illustrative prompt |
|---|---|
| Platform choice | “Okta or Microsoft Entra ID for 6,000 employees who mostly use Google Workspace?” |
| Phishing-resistant MFA | “Which identity providers support passkeys and FIDO2 for every user, including contractors?” |
| Privileged access | “What are the alternatives to CyberArk for privileged access in a hybrid Windows and Linux estate?” |
| Governance | “Do we need a separate identity governance tool, or can our identity provider handle access reviews for SOX?” |
| AI agents | “How should we give AI agents their own identities and limit what they can access?” |
| Migration | “How long does it take to migrate 400 apps from one single sign-on provider to another?” |
| Vendor risk | “Has this identity vendor had security incidents, and how did it disclose them?” |

Each of these can trigger many searches. Google documents that AI Overviews and AI Mode [may use a “query fan-out” technique](https://developers.google.com/search/docs/appearance/ai-features), and OpenAI documents that [ChatGPT search rewrites a question](https://help.openai.com/en/articles/9237897-chatgpt-search) into “one or more targeted queries.” A question about passkeys for contractors, we infer, sends those searches to documentation and standards pages as well as to review sites.

## What path leads from an AI answer to a signed identity contract?

Through an urgent shortlist: trigger, AI-assisted research, security review, pilot, then a multi-year platform contract.

1. **Trigger.** An incident, audit finding or AI-agent rollout makes identity a priority.
2. **Research.** The IAM lead or CISO asks assistants for options that fit the company’s directory, applications and regulators.
3. **Shortlist and review.** A few vendors receive security questionnaires and architecture questions; the vendor’s own security record is examined.
4. **Pilot.** The chosen vendor connects a group of applications and users.
5. **Contract and expansion.** The platform is priced by users or identities and grows into governance, privileged access and non-human identities, the bundle Palo Alto Networks is assembling.

The AI answer matters at step 2, and the urgency of step 1 shortens the time buyers spend there, we infer. A vendor not named early may never reach step 3.

## Why do assistants name some identity vendors and skip others?

The platforms do not document vendor selection; studies show answers repeat what a company’s own sources say, including contradictions.

**Documented by the platforms.** Both Google and OpenAI state that their AI answers go out to the web and link the pages behind them. Neither says what makes one SSO, MFA or privileged access vendor appear and another not.

**Observed in studies.** In [our business-facts study](https://underneath.agency/research/ai-business-facts-accuracy-study), consistency across a company’s own sources made a large difference: where a business’s profile phone number did not appear on its own website, 30.6% of the numbers AI engines gave differed from the profile, against 1.6% where it did. The study covered local businesses, but the lesson carries over, we infer: an identity vendor whose integration list, standards support or certifications differ between its site, documentation and review profiles invites answers that differ too.

Pricing is another fact buyers ask about, and [our pricing study](https://underneath.agency/research/ai-pricing-accuracy-study) found only 61.9% of software plan prices quoted by four AI engines were fully faithful to the official pricing page. Identity pricing, with per-user tiers and add-ons, gives answers plenty of room to drift.

**Identity-specific trust signals, our inference.** Most of what an IAM team checks before a pilot sits on public pages an assistant can also read:

- standards support: SAML, OpenID Connect, SCIM provisioning, and FIDO2 or passkey authentication;
- the vendor’s own security record: certifications such as SOC 2 and ISO 27001, FedRAMP status, a public trust center, and dated incident disclosures;
- the integration catalog, with clear statements of depth, not just logos;
- migration guides and reference architectures;
- independent coverage, analyst placements and practitioner reviews.

Our working assumption is that identity vendors who keep these answers public, current and consistent hand assistants better material to work with. That assumption has not been tested on SSO, governance or privileged access vendors.

## What is at stake for an IAM vendor that assistants overlook?

A missed urgent decision and a long lock-in afterward; no study has measured the cost in dollars.

- **Urgent cycles are short.** When an incident drives the purchase, buyers decide quickly, so a vendor not named at the start has little time to enter, we infer.
- **Lock-in cuts both ways.** The switching cost that protects an incumbent also keeps a missing vendor out for years.
- **Bundles expand the loss.** Losing the access decision can mean losing governance, privileged access and agent identities later, as platforms like Palo Alto Networks bundle them.
- **Wrong facts can disqualify.** An AI answer that misstates your passkey support, certifications or a past incident can remove you before the security review. Correcting those errors is covered in [how to fix wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers).

## Which GEO work helps an IAM vendor get verified and named?

Publish standards support, security record and fit in one consistent form; nobody can force a mention.

1. **One consistent set of facts.** Keep product names, supported standards, integrations, certifications and pricing identical across your site, documentation, marketplace listings and review profiles.
2. **A public trust center.** Publish certifications, audit dates, data residency, subprocessors and incident history on readable pages. Your own security is part of the product.
3. **Standards and migration content.** Explain passkey rollout, provisioning, and migration from named competitors in practical detail.
4. **Answers for new triggers.** Publish clear guidance on AI-agent and non-human identities, tied to what your product actually does.
5. **Independent coverage.** Contribute identity attack research and expert comment after major incidents; brief analysts. This is the identity version of [how brands build authority for AI search](https://underneath.agency/resources/how-brands-build-authority-for-ai-search).
6. **Reviews from practitioners.** Encourage detailed reviews from IAM teams, within review platforms’ rules.
7. **Honest comparisons.** Fair pages comparing you with Okta, Entra ID or CyberArk help buyers, though third-party sources weigh more; see [whether comparison pages help B2B citations](https://underneath.agency/resources/do-comparison-pages-help-b2b-ai-citations).
8. **Measurement.** Track platform, standards, governance, agent and vendor-risk questions across assistants and repeated runs; see [how many prompts to track](https://underneath.agency/resources/how-many-prompts-to-track-ai-visibility).

For the wider security buying picture, see [cybersecurity software and AI search](https://underneath.agency/resources/cybersecurity-software-revenue-from-ai-search); for subscription economics, see [B2B SaaS and AI search](https://underneath.agency/resources/b2b-saas-revenue-from-ai-search).

## What remains unmeasured about AI in identity purchases?

Nobody has yet measured how IAM buyers use assistants or whether being named changes identity vendors’ win rates.

- **No survey of IAM buyers.** The TrustRadius figures describe technology buyers as a whole, not identity teams.
- **Vendor sources.** Threat and adoption figures here come largely from security vendors, including CrowdStrike, Okta and Palo Alto Networks.
- **Transfer from our studies.** Our facts and pricing studies did not test identity vendors; we apply them by inference.
- **Ties to closed identity deals.** None are measured here; the general evidence is in [does AI visibility drive business results](https://underneath.agency/resources/does-ai-visibility-drive-business-results).

## What should an identity vendor check before the next breach-driven buying cycle?

What assistants say about your passkey and SSO support, your security record, and your fit for each buying trigger.

List questions for each trigger: breach response, phishing-resistant MFA, governance audits, AI-agent identities, migration and vendor risk. Repeat every question on each main assistant, and on more than one day. Record whether you are named, whether your facts are right, and which sources are cited. Fix inconsistencies in your own sources first; they are the cheapest gap to close.

If you would rather have us run it, [ask us to review your identity visibility](https://underneath.agency/contact). We will check where assistants include or omit you when enterprises shortlist SSO, MFA, governance and privileged access vendors, and which gaps most likely keep you out of security reviews, pilots and per-user platform contracts. The way we then align an identity vendor’s standards, trust center and integration facts across every source is described on our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page.

## Frequently asked questions

### Do identity breaches really drive IAM purchases?

They are a leading trigger. Credential abuse was the top initial attack vector in Verizon’s 2025 report, at 22%, and incidents put identity on the board agenda.

### Should identity vendors publish their own incident history?

Yes. Buyers ask about past incidents, and an assistant that finds only third-party accounts of them will repeat those. A dated, factual record gives answers your version.

### Does supporting passkeys help AI visibility?

Only if it is clearly documented. Buyers increasingly ask about phishing-resistant MFA, which CISA strongly urges, and answers can only repeat what they can find.

### How do we keep AI answers from quoting the wrong price?

Make one clear public pricing page and keep it consistent elsewhere. Only 61.9% of software prices quoted by AI engines in our study were fully faithful.

## Sources

- Verizon (2025-04-23), [2025 Data Breach Investigations Report news release](https://verizon.com/about/news/2025-data-breach-investigations-report)
- CrowdStrike (2025-02-27), [CrowdStrike Releases 2025 Global Threat Report](https://www.crowdstrike.com/en-us/press-releases/crowdstrike-releases-2025-global-threat-report/)
- IBM and Ponemon Institute (2025-07), [Cost of a Data Breach Report 2025](https://www.bakerdonelson.com/webfiles/Publications/20250822_Cost-of-a-Data-Breach-Report-2025.pdf)
- Louis Columbus, Software Strategies Blog (2026-04-01), [Gartner’s $246.2B Security Forecast shows 10 categories growing 2x to 3x the market](https://softwarestrategiesblog.com/2026/04/01/top-10-fastest-growing-security-categories-gartner-2026-forecast/)
- Palo Alto Networks (2025-07-30), [Palo Alto Networks Announces Agreement to Acquire CyberArk, the Identity Security Leader](https://www.paloaltonetworks.com/company/press/2025/palo-alto-networks-announces-agreement-to-acquire-cyberark--the-identity-security-leader)
- Okta (2026), [Businesses at Work Report 2026](https://www.okta.com/reports/businesses-at-work/)
- Mandiant, Google Cloud (2024-06-10), [UNC5537 Targets Snowflake Customer Instances for Data Theft and Extortion](https://cloud.google.com/blog/topics/threat-intelligence/unc5537-snowflake-data-theft-extortion)
- CISA (2022-10), [Implementing Phishing-Resistant MFA](https://www.cisa.gov/sites/default/files/publications/fact-sheet-implementing-phishing-resistant-mfa-508c.pdf)
- Demand Gen Report, James Hickey (2026-07-30), [TrustRadius: AI Has Changed How Buyers Research, But Not What They Trust](https://www.demandgenreport.com/industry-news/news-brief/trustradius-ai-has-changed-how-buyers-research-but-not-what-they-trust/53853/)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- OpenAI Help Center (2025), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Underneath (2026), [Do AI answers match a business’s Google profile?](https://underneath.agency/research/ai-business-facts-accuracy-study)
- Underneath (2026), [How faithfully do AI assistants quote software prices?](https://underneath.agency/research/ai-pricing-accuracy-study)

---

This is the Markdown twin of https://underneath.agency/resources/iam-enterprise-demand-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How industrial equipment makers win B2B leads through AI search"
description: "Equipment makers reach more bid lists when AI can verify their duty ranges, efficiency data, standards and local service, the facts engineers specify on."
canonical: "https://underneath.agency/resources/industrial-equipment-leads-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Will AI help put our pumps, compressors and valves on the bid list?

It can, when an assistant can match your equipment to a duty, confirm its efficiency and see who services it nearby. Engineers now use AI early in research, but they still choose on lifecycle cost and trust in the supplier, so the job is to make those facts easy to find and check.

This article is for makers of process and plant equipment that engineers specify: pumps, compressors, blowers, valves, dryers, boilers and similar packages bought through requests for quotation (RFQs), often via a consulting engineer or an engineering, procurement and construction (EPC) firm. If you make components that engineers design in, see how part makers get named by AI.

## The short version

1. The installed base is the prize: a guide from the [US Department of Energy and the Hydraulic Institute](https://www.energy.gov/sites/prod/files/2014/05/f16/pumplcc_1001.pdf) notes that there are at least 20 times as many pump systems installed as are built each year, and that pumping systems often last 15 to 20 years.
2. Energy decides the economics: the same guide says pumping systems account for nearly 20% of the world’s electricity demand, and the [Compressed Air and Gas Institute](https://www.cagi.org/performance-verification/) works through an example in which one 100 hp compressor uses $41,968 of electricity a year.
3. Service is where much of the money is: Ingersoll Rand’s CEO told analysts that [40% of its revenue is aftermarket](https://finance.yahoo.com/news/ingersoll-rand-ir-q4-2025-144202373.html), and its recurring revenue passed $450 million in 2025.
4. Project demand is shifting: the [US Census Bureau](https://www.census.gov/construction/c30/pdf/release.pdf) put power construction at an annual rate of $185,973 million in August 2026, up 8.5% in a year, while manufacturing construction fell 19.2%.
5. Engineers start with AI but verify with people: in the 2026 [State of Marketing to Engineers](https://advertising.globalspec.com/state-of-marketing-to-engineers-research-report) research by TREW Marketing and GlobalSpec, only 6% of technical buyers who notice AI summaries in search say the summary is usually enough, and 41% routinely consult sales or application engineers.

## Who specifies plant equipment, and what is one award worth?

Plant, process and consulting engineers write the specification; one award can bring decades of parts and service.

Most process equipment is chosen against duty conditions: flow, pressure or head, fluid, temperature, materials and the standard the plant requires. A plant engineer, reliability manager or energy manager defines the need. On larger projects, a consulting engineer or EPC firm writes the specification and the bid list, and procurement runs the RFQ. Final selection usually weighs lifecycle cost, delivery, service coverage and the supplier’s record. Catalog components sold through distributors are chosen differently, as [how part makers get named by AI](https://underneath.agency/resources/industrial-manufacturers-ai-search) explains.

What one customer is worth depends on the package, but the shape is clear in public data. The DOE and Hydraulic Institute guide says the initial purchase price is a small part of the lifecycle cost of a high-use pump; energy and maintenance usually dominate. That makes the aftermarket large. Ingersoll Rand’s chief executive, Vicente Reynal, said on the company’s fourth-quarter 2025 call that 40% of revenue is aftermarket, and that recurring revenue exceeded $450 million in 2025, with about $1.1 billion more already contracted for future years.

Our inference: a single award opens a 15- to 20-year relationship for spare parts, service contracts and the replacement that eventually follows, so the moment a supplier makes the bid list carries far more value than the first invoice. Production machinery rests on a similar capital case, covered in [how AI steers a plant’s machine purchases](https://underneath.agency/resources/machinery-companies-buyers-ai-search).

## How far has AI reached into equipment research?

Most engineers now use it somewhere in a purchase, but they treat it as a first pass, not a decision.

The best industry-specific evidence is the 2026 survey of more than 1,000 technical buyers by TREW Marketing and GlobalSpec, with Elektor:

| Technical buyers, 2026 | Share |
|---|---|
| Use generative AI at some point when buying | 69% |
| Have noticed AI-generated summaries at the top of search results | 75% |
| Of those, say the summary is usually enough | 6% |
| Routinely consult sales or application engineers when researching a purchase | 41% |
| Contact a salesperson first to validate information gathered online | 23% |
| Contact a salesperson first because of a solution’s technical complexity | 21% |

For equipment makers, the last three lines matter as much as the first. Engineers bring AI-assisted research to an application engineer and ask them to confirm it. A supplier whose published data disagrees with what an assistant said starts that conversation on the back foot.

## Which questions do engineers ask AI about process equipment?

Questions shaped like a datasheet: duty point, fluid, standard, efficiency, region and the brands being compared.

The examples below are written by us to show the pattern. They are not logged from real engineers or captured from an assistant:

| Need | Example question |
|---|---|
| Duty point | “Centrifugal pump for 400 gpm at 150 ft of head with 30% glycol, which models fit?” |
| Efficiency | “Oil-free rotary screw compressors around 75 hp with the lowest specific power” |
| Standard | “API 610 pump suppliers with service centers on the Gulf Coast” |
| Hard service | “Which manufacturers make control valves for flashing and cavitating service?” |
| Lifecycle cost | “Compare blowers for wastewater aeration by energy use over ten years” |
| Replacement | “Efficient replacement for a 1990s split-case pump that meets current DOE rules” |

Notice how many carry a place or a standard. Those are the details an assistant must find stated somewhere in text. Performance curves locked in PDFs, sizing tools behind logins and service maps drawn as images give it little to work with.

## How does an AI answer become an RFQ and an award?

Through the bid list: an assistant shapes who is considered, then data, references and service decide who wins.

1. **Considered.** An engineer or specifier asks an assistant or search engine which suppliers fit the duty and standard, alongside trade publications and colleagues.
2. **Checked.** They read performance data, efficiency ratings and references, then call an application engineer or local representative.
3. **Invited.** The supplier is named on the bid list or in the specification, and receives the RFQ.
4. **Evaluated.** Bids are compared on lifecycle cost, compliance, delivery and service.
5. **Supported.** The winner supplies parts, service and upgrades for years, and is first in line for the replacement.

AI visibility acts on the first two steps. Our inference: being left off a bid list cannot be fixed by a better price later, because the RFQ never arrives. That makes the early research stage the one to watch, as our article on [what lost clicks to AI answers mean for pipeline](https://underneath.agency/resources/ai-answers-pipeline-revenue) argues more generally.

## What makes an assistant name one equipment supplier over another?

Facts it can confirm in outside sources, with verified performance data and local service carrying special weight.

What the platforms document: [Google says](https://blog.google/products/search/ai-mode-search/) AI Mode uses a “query fan-out” technique, “issuing multiple related searches concurrently across subtopics and multiple data sources.” [OpenAI says](https://help.openai.com/en/articles/9237897-chatgpt-search) ChatGPT search rewrites a question into one or more targeted queries and that a site must allow its OAI-SearchBot crawler to be eligible. Neither explains how an equipment brand is chosen.

What our studies observed, across consumer and business questions rather than equipment:

- **Assistants look for rankings and publications.** In [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study), ChatGPT ran a mean of 3.7 searches per buyer question, and 43.8% of its answers included a search aimed at a named publication, ranking or award.
- **Place-specific questions split the assistants.** In [our brand agreement study](https://underneath.agency/research/ai-assistants-brand-agreement-study), two assistants’ picks overlapped 0.160 on questions naming a place, against 0.390 on national questions. For equipment, “who services this in my region” is that kind of question.
- **Answers move between runs.** In [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), only 25.2% of the brands ChatGPT named for a question appeared in all five repeats.
- **Some rankings are written by vendors.** In [our study of self-ranking lists](https://underneath.agency/research/self-promoting-best-lists-study), 24.2% of AI-cited numbered “best” lists with an identifiable publisher put that publisher first. The pump industry has such pages: one pump maker’s own “[10 top pump manufacturers of the world](https://jeepumps.com/10-top-pump-manufacturers-of-the-world)” lists itself fourth, after Grundfos, Xylem and KSB. We have not checked whether assistants cite it.

What we infer for equipment makers: the trust signals are the ones engineers already rely on. Third-party verified ratings are the clearest. The Compressed Air and Gas Institute runs a [performance verification program](https://www.cagi.org/performance-verification/) in which an independent lab tests rotary compressors from 5 to 200 hp against the data sheets manufacturers publish; the Hydraulic Institute runs an [Energy Rating label and database](https://www.pumps.org/what-we-do/energy-rating/) aligned with the DOE pump standard. Add named reference installations, coverage in technical publications, and service locations stated plainly.

## What does it cost to be missing from AI answers?

RFQs that never arrive, in a market where lifecycle savings and the installed base reward whoever is considered first.

We found no public measurement of equipment bids lost to AI absence, so the reasoning is labeled:

- **The savings case is large.** The DOE guide says studies have shown that 30% to 50% of the energy used by pump systems could be saved through equipment or control changes, and that pumping can account for 25% to 50% of energy use in some plants. Retrofit projects follow energy audits, and the supplier an engineer finds while researching is the one that gets asked.
- **Demand is moving between sectors.** The Census figures show power construction rising while manufacturing construction fell 19.2% year over year. Water supply ran at $36,804 million and sewage and waste disposal at $53,791 million. Suppliers following the work into utilities and water need to be findable for those applications. This is our inference from the spending data.
- **Familiarity tips close calls.** In the engineers’ survey, 70% were likely to choose the better-known brand when two solutions were technically similar. Being named in research builds that familiarity.
- **Wrong facts travel.** An outdated efficiency figure or a service center that closed years ago can rule you out. Our guide to [correcting wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) shows how to find the page an assistant relied on.

## How does GEO work for a process equipment maker?

Generative engine optimization (GEO) makes your equipment easy for AI assistants to match to a duty and verify.

For a pump, compressor or valve company, the work usually includes:

1. **Application pages in text.** One page per application and industry, with duty ranges, materials, standards met and efficiency data written out, not only in curves, PDFs or a sizing tool.
2. **Verified ratings up front.** Program participation and results (CAGI data sheets, Hydraulic Institute Energy Rating, DOE compliance) stated on the product page and linked to the program’s own listing.
3. **Lifecycle-cost content.** Worked energy and maintenance cost examples with the assumptions published, so an engineer, and an assistant, can check them.
4. **Service and representation by region.** Service centers, authorized representatives and response commitments listed as text by state or region, consistent with your representatives’ own sites.
5. **Independent proof.** Application articles in technical publications, conference papers, case studies with named plants where customers allow it, and listings in the vendor directories EPC firms use. For why outside sources matter, see [how brands build authority for AI search](https://underneath.agency/resources/how-brands-build-authority-for-ai-search).
6. **Fair comparisons.** Honest pages on how your technology compares with alternatives for a given duty. Our review of [whether comparison pages help B2B brands get cited](https://underneath.agency/resources/do-comparison-pages-help-b2b-ai-citations) covers what works.
7. **Measurement.** Track a stable list of duty, standard, replacement and regional-service questions in ChatGPT, Gemini, Perplexity, Copilot and Google’s AI features over several runs, and line the results up against the RFQ invitations you receive.

None of this promises a place in any answer. It makes your equipment the easiest to verify for both the assistant and the engineer who checks it.

## What remains unknown about AI and equipment specification?

Nobody has yet measured how often an AI answer adds a supplier to a bid list or changes an award.

- **Usage is not influence.** The TREW and GlobalSpec survey measures research habits, not which RFQs came from AI answers.
- **Our studies are not about equipment.** They covered buyer questions in other industries; applying them to process equipment is our inference.
- **Some figures are old or specific.** The DOE lifecycle guide dates from 2001, and the CAGI cost example rests on stated assumptions about hours and electricity prices.
- **Company figures describe companies.** Ingersoll Rand’s aftermarket share is one firm’s mix, not an industry average.

## Where should an equipment company start?

Start with the duty conditions and regions where you win most often, and see whether assistants name you there.

A first review shows whether you appear for your core applications and standards, whether assistants describe your efficiency and service coverage correctly, which publications, directories and rankings they draw on, and which competitors are named in your place.

If your growth depends on being invited to more bids, [ask us to look at your equipment’s AI visibility](https://underneath.agency/contact). We will show where assistants name your products for the duties you serve, why rivals appear instead, and which changes are most likely to put you on more bid lists and RFQs. For a pump, compressor or valve maker, our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page explains how application pages, verified ratings and regional service facts are rebuilt and then measured.

## Frequently asked questions

### Do consulting engineers and EPC firms use AI to build bid lists?

There is no published survey of EPC bid-list practice and AI. Broader surveys show most technical buyers use generative AI somewhere in purchasing, then verify with application engineers.

### Do verified performance programs help with AI visibility?

We have not measured that. They give an assistant an independent source that confirms your claims, which is the kind of evidence engineers trust, so we treat them as worth stating prominently.

### Should we publish prices for engineered equipment?

Often that is impractical. Publish what you can verify instead: duty ranges, efficiency data, lead-time ranges and typical lifecycle costs with assumptions shown.

### Does this matter for replacement business, not just new projects?

Yes. With far more systems installed than built each year, many RFQs are replacements and retrofits after an energy audit or a failure.

## Sources

- US Department of Energy, Hydraulic Institute and Europump (2001), [Pump Life Cycle Costs: A Guide to LCC Analysis for Pumping Systems, Executive Summary](https://www.energy.gov/sites/prod/files/2014/05/f16/pumplcc_1001.pdf)
- Compressed Air and Gas Institute (n.d.), [Performance Verification Program](https://www.cagi.org/performance-verification/)
- Hydraulic Institute (n.d.), [Energy Rating](https://www.pumps.org/what-we-do/energy-rating/)
- Yahoo Finance (2026-02), [Ingersoll Rand (IR) Q4 2025 earnings call transcript](https://finance.yahoo.com/news/ingersoll-rand-ir-q4-2025-144202373.html)
- US Census Bureau (2026-10-01), [Monthly Construction Spending, August 2026](https://www.census.gov/construction/c30/pdf/release.pdf)
- TREW Marketing and GlobalSpec (2026), [State of Marketing to Engineers research report](https://advertising.globalspec.com/state-of-marketing-to-engineers-research-report)
- Google (2025-03-05), [Expanding AI Overviews and introducing AI Mode](https://blog.google/products/search/ai-mode-search/)
- OpenAI Help Center (2026), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- JEE Pumps (n.d.), [10 top pump manufacturers of the world](https://jeepumps.com/10-top-pump-manufacturers-of-the-world)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude agree on brands?](https://underneath.agency/research/ai-assistants-brand-agreement-study)
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)
- Underneath (2026), [How many “best of” lists cited by AI rank their own brand first?](https://underneath.agency/research/self-promoting-best-lists-study)

---

This is the Markdown twin of https://underneath.agency/resources/industrial-equipment-leads-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How industrial manufacturers win buyers through AI search"
description: "By making every part number, spec and equivalent easy for AI to verify, on your site and in distributor catalogs, before engineers design it in."
canonical: "https://underneath.agency/resources/industrial-manufacturers-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# When an engineer asks AI for a part, will it name ours?

It will only if an assistant can find your part number, read its specifications and confirm where to buy it. For a maker of catalog components, AI search matters at two moments: when an engineer chooses parts for a new design, and when a maintenance team needs a replacement today.

This article is for manufacturers of standard industrial products sold from a catalog, mostly through distributors: bearings, motion and fluid power components, connectors, sensors, fasteners, enclosures, lubricants and similar lines. If you build parts to a customer’s drawing, our article on how contract manufacturers win RFQs through AI search fits better. Makers of pumps, compressors and valves can read how equipment makers reach more bid lists.

## The short version

1. Engineers are letting AI into component choice: in a [Newark survey](https://news.newark.com/newark-research-shows-engineers-now-trust-ai-to-select-components/), 86% of engineers trusted AI to play at least some role in selecting components, and 23% of those would trust it completely.
2. The choice is made early: in a 2026 survey of 400 North American engineers [commissioned by Weidmuller](https://tedmag.com/weidmuller-usa-releases-new-survey/), 43% said most component selections are finalized in the concept or schematic phase.
3. Your part competes inside huge catalogs: [Grainger’s 2025 annual report](https://s1.q4cdn.com/422144722/files/doc_financials/2025/ar/2025-GWW-Annual-Report.pdf) lists more than 5,000 primary suppliers, about 2 million products in its high-touch business, about 13 million at Zoro and about 29 million at MonotaRO.
4. Ordering is already digital: [Fastenal’s digital footprint](https://www.digitalcommerce360.com/article/fastenal-digital-sales/) was 61.6% of its sales in the second quarter of 2026, and Grainger’s CEO said electronic procurement is “closer to 40%” of its business.
5. Engineers still verify: in the 2026 [State of Marketing to Engineers](https://advertising.globalspec.com/state-of-marketing-to-engineers-research-report) research by TREW Marketing and GlobalSpec, technical buyers rated their trust in AI answers 4.7 out of 10, and 77% still use a search engine more often than AI.

## Who buys catalog industrial products, and what is one design-in worth?

Two buyers: the engineer who specifies a part into a design, and the maintenance or purchasing team that reorders it.

The first buyer is a design engineer choosing components for a new machine, product or line. Once a part is on the drawing and the bill of materials, it is bought for as long as the product is made, usually through a distributor. The second is a maintenance, repair and operations (MRO) buyer who needs a replacement fast, often by part number, and orders from whichever distributor has it in stock.

Most of that volume flows through distribution. The [US Census Bureau](https://www.census.gov/wholesale/pdf/mwts/currentwhl.pdf) estimates that merchant wholesalers of machinery, equipment and supplies sold $59,883 million of goods in July 2026 alone, seasonally adjusted, 12.1% more than a year earlier. Electrical wholesalers, which carry connectors, controls and sensors, sold $111,485 million.

There is no public benchmark for what one design-in is worth; it depends on the part’s price and the product’s volume and life. The mechanism is the point. Our inference: a single specification decision can produce years of repeat distributor orders that no salesperson ever sees, so the moment a part is chosen is worth far more than one sale. Parts built to a customer’s drawing follow a quote-led path instead, covered in [how contract manufacturers win RFQs](https://underneath.agency/resources/contract-manufacturers-rfqs-from-ai-search).

## Where does AI already sit when engineers choose components?

At the start of research and in early component selection, used widely but checked against datasheets and publications.

Industry surveys show engineers bringing AI into the parts decision, carefully:

| Finding | Source |
|---|---|
| 69% of technical buyers use generative AI during the purchasing process | TREW Marketing and GlobalSpec, 2026 |
| 21% routinely research purchases on generative AI platforms, up 8 points in a year | TREW Marketing and GlobalSpec, 2026 |
| 31% routinely use industry directory websites | TREW Marketing and GlobalSpec, 2026 |
| 86% trust AI to play at least some role in component selection | Newark survey |
| 91% have used AI-based tools in their PCB design workflow | Weidmuller and EETech Research, 2026 |
| 79% say reliability and durability matter most in connectivity components, against 22% for price | Weidmuller and EETech Research, 2026 |

Newark’s respondents also described the limits: several wanted AI as “an enhanced search engine of sorts,” with every selection reviewed by an engineer, especially for safety-critical designs.

The design files themselves are moving online too. [TraceParts](https://info.traceparts.com/pcs/publishing-3d-catalogs/understanding-industrial-engineers-designers-survey-report/), a platform for downloadable 3D models, says it hosts nearly 120 million searchable part numbers used by more than 6 million registered engineers and designers. An engineer who drops your model into a design has, in effect, chosen your part.

Purchasing is changing faster still. Gartner forecasts, as [reported by Digital Commerce 360](https://www.digitalcommerce360.com/2025/11/28/gartner-ai-agents-15-trillion-in-b2b-purchases-by-2028/), that AI agents will intermediate more than $15 trillion of business spending by 2028, relying on “verifiable data feeds.” That is a forecast, not a measurement, but it points at product data, the thing a catalog manufacturer controls.

## What do engineers and MRO buyers ask AI about parts?

Questions built from specifications, standards, equivalents and availability, often naming a part number or a rival brand.

We wrote the examples below to show the shape of these questions; they are not captured from real buyers or assistants:

| Moment | Example question |
|---|---|
| Specification | “Inductive proximity sensor, 30 mm sensing range, IP69K, for a washdown food line” |
| Standard | “Pneumatic cylinders to ISO 15552 with 50 mm bore, stainless, which brands?” |
| Equivalent | “Drop-in replacement for a discontinued solenoid valve, part number …” |
| Compliance | “Which manufacturers make NSF H1 food-grade gear oil?” |
| Comparison | “Compare sealed deep-groove bearings from three brands for high-speed motors” |
| Availability | “Where can I buy this M12 connector in stock for next-day delivery?” |

Each question is a filter on attributes. If your sensing range, rating, standard, material or cross-reference only lives inside a PDF datasheet or a configurator, an assistant may not be able to match it. The equivalence question deserves special attention: when a rival’s part is discontinued or out of stock, the brand an assistant offers as the substitute picks up the business.

## How does an AI answer turn into distributor orders?

Through a design-in or a replacement order, then repeat purchases through whichever distributor the buyer uses.

For a new design:

1. **Named.** An engineer asks an assistant, a search engine or a distributor’s search box which parts meet a requirement.
2. **Checked.** The engineer reads the datasheet, downloads the 3D model and often orders samples.
3. **Specified.** The part goes onto the drawing and bill of materials. In the Weidmuller survey, 43% said most selections are finalized early, and another 26% said selection is iterative across stages.
4. **Bought for years.** Production purchasing orders the specified part number, usually from distribution.

For a replacement, the path is shorter: an MRO buyer looks up the failed part or an equivalent, checks availability, and orders through a distributor website or an electronic procurement link. Grainger’s CEO told analysts, as [reported by Digital Commerce 360](https://www.digitalcommerce360.com/2026/02/03/grainger-ai-data-digital-sales-q4-2025/), that electronic procurement “is the biggest share we have at this point. Closer to 40%.” Fastenal said its digital footprint, its vending and inventory systems plus electronic business, reached 61.6% of sales, with electronic business daily sales up 12.6%.

The money comes from sell-through. Our inference: AI visibility matters most at the naming step for new designs and at the equivalence step for replacements; after that, availability and price at the distributor decide the order. Pumps, compressors and valves reach buyers through bid lists instead, as [how equipment makers reach more bid lists](https://underneath.agency/resources/industrial-equipment-leads-ai-search) explains.

## What decides whether an assistant names your part?

Product facts it can read and confirm, plus who sells the part and whether it is in stock.

What the platforms document:

- **Structured product data counts.** OpenAI’s [shopping help page](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search) says ChatGPT considers “structured metadata from first-party and third-party providers (e.g., price, product description)” when choosing products, and ranks the merchants offering a product by “availability, price, quality, and whether they are the maker or primary seller of that item.” Merchants can apply to send a direct product feed.
- **Google’s shopping answers draw on a product database.** Google says AI Mode shopping uses its [Shopping Graph](https://blog.google/products/shopping/google-shopping-ai-mode-virtual-try-on-update/), with more than 50 billion product listings, more than 2 billion of them refreshed every hour.
- **Assistants search, then answer.** OpenAI says [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search) rewrites a question into one or more targeted queries, and a site must allow its OAI-SearchBot crawler to be eligible.

These pages describe consumer shopping, and none of them explains how an industrial part is chosen for a design question. Applying them to components is our inference.

What our studies observed, across other industries:

- In [our study of hidden searches](https://underneath.agency/research/ai-hidden-searches-study), ChatGPT ran a mean of 3.7 searches per buyer question and looked for prices in 23.8% of its answers.
- In [our pricing study](https://underneath.agency/research/ai-pricing-accuracy-study), 61.9% of the software plan prices assistants quoted were fully faithful to the official page; most of the rest were real variants found elsewhere on the vendor’s own site or on third-party pages. Inconsistent prices across your own pages and your distributors’ pages give an assistant more than one answer to repeat.
- In [our study of agent-readable websites](https://underneath.agency/research/agent-readable-web-study), 45.8% of top homepages carried JSON-LD structured data, so many sites still give machines no structured description at all.

What we infer for catalog manufacturers: the trust factors are the ones an engineer already checks. Complete specifications in text, standards and approvals (UL, CE, ATEX, NSF, IP ratings) named precisely, cross-references to competitor and legacy part numbers, application notes, published reliability data and in-stock availability at named distributors.

## What happens to sales when your parts are left out?

You lose design-ins you never hear about and replacement orders that go to the brand an assistant offered instead.

We found no public measurement of orders lost to AI absence, so here is the labeled reasoning:

- **Selection happens before contact.** The TREW and GlobalSpec research found that 62% of the technical buying journey happens online before a vendor is contacted, and the top reason buyers reach out at all is pricing or inventory questions (26%). A part missing at the research stage is rarely added later.
- **Catalogs are crowded.** With about 2 million products at Grainger alone and 85,000 net additions in 2025, a distributor’s own search and content decide which brand appears first. Grainger said its category reviews focus on “improving product search, organization, and content.”
- **Familiarity tips close calls.** In the same engineers’ survey, 53% said brand familiarity influenced their most recent purchase. We infer that being named in AI answers and technical publications builds the familiarity that later decides a tie.
- **Wrong data does damage.** An assistant that repeats an outdated rating or a discontinued part number can steer an engineer away. Finding the page an assistant took the outdated rating from comes first; our guide to [fixing wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) shows how.

## How does GEO work for a catalog manufacturer?

Generative engine optimization (GEO) makes your parts easy for AI assistants to find, match to a requirement and verify.

For a component maker, the work usually includes:

1. **One page per product family, with specs in text.** Dimensions, ratings, materials, standards and operating ranges written as text and tables on the page, not only in PDF datasheets, configurators or images.
2. **Cross-reference and equivalence pages.** Honest tables mapping your part numbers to legacy, discontinued and competitor numbers, with the differences stated.
3. **Consistent data in every catalog.** The same attributes, names and part numbers on your site and at each distributor, from national houses to regional specialists, so an assistant does not find conflicting facts.
4. **Product feeds and structured data.** Where platforms accept feeds or read structured product data, supply them accurately and keep them current.
5. **Design resources that can be found.** 3D models, application notes and selection guides published where engineers and crawlers can reach them.
6. **Independent coverage.** Application stories and technical articles in trade publications, which 76% of technical buyers routinely read. For why third-party coverage carries weight with assistants, see [how brands build authority for AI search](https://underneath.agency/resources/how-brands-build-authority-for-ai-search).
7. **Measurement.** Ask a fixed set of specification, equivalence and where-to-buy questions across ChatGPT, Gemini, Perplexity, Copilot and Google’s AI features, repeat them, and record which parts and sources are named. Our article on [designing AI visibility tracking](https://underneath.agency/resources/how-to-design-ai-visibility-tracking) explains why one answer is not enough.

None of this guarantees a mention. It makes your part the easiest one for an assistant to match and for an engineer to confirm. For the research on which product facts move AI choices, see [what kind of product content AI shopping assistants prefer](https://underneath.agency/resources/product-content-ai-shopping-assistants-prefer).

## What can’t the data tell a component maker yet?

It shows engineers using AI in research, not how many design-ins or distributor orders AI answers cause.

- **Trust is not attribution.** The Newark and Weidmuller surveys measure attitudes and tool use. Neither links an AI answer to a specified part.
- **Sponsors have interests.** Newark is a distributor, Weidmuller sells components, TREW and GlobalSpec sell marketing to engineering firms, and TraceParts hosts CAD content for manufacturers.
- **Platform pages describe consumer shopping.** OpenAI and Google document how product results work for shoppers. How assistants handle technical part questions is not documented, and we have not measured it for industrial parts.
- **Forecasts are forecasts.** Gartner’s $15 trillion figure is a prediction about AI agents in purchasing, not today’s behavior.

## Where should an industrial manufacturer start?

Start with your best-selling part families and ask assistants the specification and replacement questions your customers ask.

That first check shows whether your parts are named for the requirements they meet, whether assistants describe your ratings and part numbers correctly, which distributor and publication pages they rely on, and which brand they offer as the equivalent when yours is the one being replaced.

If your growth depends on being designed in and reordered through distribution, [talk to us about a review of your product visibility](https://underneath.agency/contact). We will show which of your parts assistants name, where they get the facts, and which changes are most likely to put your part numbers on more drawings and distributor orders. What the follow-on work involves, such as cross-reference pages, consistent distributor data and repeated part-question checks, is outlined on our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page.

## Frequently asked questions

### Should we prioritize our own site or our distributors’ product pages?

Both, with one set of facts. Assistants may cite either, and OpenAI says it ranks merchants partly on whether they are the maker or primary seller, so your own page should be the most complete version.

### Do engineers trust AI to pick components?

Partly. 86% in Newark’s survey would let AI play some role, but trust in AI answers averaged 4.7 out of 10 in the TREW and GlobalSpec research, so engineers check datasheets before committing.

### Are PDF datasheets enough?

Keep them, but put the key specifications on the page as text too. A requirement can only be matched to your part if the attribute can be read and compared.

### Does this help with replacement and MRO orders?

Yes, especially for equivalence questions. Clear cross-reference pages and accurate stock information at distributors help an assistant offer your part when another brand’s part fails or is discontinued.

## Sources

- Newark (2025), [Newark research shows engineers now trust AI to select components](https://news.newark.com/newark-research-shows-engineers-now-trust-ai-to-select-components/)
- tED magazine (2026), [Weidmuller USA releases new survey](https://tedmag.com/weidmuller-usa-releases-new-survey/)
- W.W. Grainger (2026), [2025 Annual Report](https://s1.q4cdn.com/422144722/files/doc_financials/2025/ar/2025-GWW-Annual-Report.pdf)
- Digital Commerce 360 (2026-02-03), [Grainger leans on AI, data and digital tools as Q4 sales rise and profits fall](https://www.digitalcommerce360.com/2026/02/03/grainger-ai-data-digital-sales-q4-2025/)
- Digital Commerce 360 (2026), [Fastenal digital sales](https://www.digitalcommerce360.com/article/fastenal-digital-sales/)
- TREW Marketing and GlobalSpec (2026), [State of Marketing to Engineers research report](https://advertising.globalspec.com/state-of-marketing-to-engineers-research-report)
- US Census Bureau (2026-09-10), [Monthly Wholesale Trade: Sales and Inventories, July 2026](https://www.census.gov/wholesale/pdf/mwts/currentwhl.pdf)
- TraceParts (n.d.), [Understanding Industrial Engineers and Designers (Survey Report)](https://info.traceparts.com/pcs/publishing-3d-catalogs/understanding-industrial-engineers-designers-survey-report/)
- Digital Commerce 360 (2025-11-28), [Gartner: AI agents will command $15 trillion in B2B purchases by 2028](https://www.digitalcommerce360.com/2025/11/28/gartner-ai-agents-15-trillion-in-b2b-purchases-by-2028/)
- OpenAI Help Center (2026), [Shopping with ChatGPT Search](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search)
- OpenAI Help Center (2026), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Google (2025-05-20), [Shop with AI Mode, use AI to buy and try clothes on yourself virtually](https://blog.google/products/shopping/google-shopping-ai-mode-virtual-try-on-update/)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [How faithfully do AI assistants quote software prices?](https://underneath.agency/research/ai-pricing-accuracy-study)
- Underneath (2026), [Do websites serve Markdown to AI agents?](https://underneath.agency/research/agent-readable-web-study)

---

This is the Markdown twin of https://underneath.agency/resources/industrial-manufacturers-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How industrial software vendors win plant pilots via AI search"
description: "When plant teams research MES, maintenance or industrial data platforms with AI, the vendors named are those whose plant-level facts are easy to verify."
canonical: "https://underneath.agency/resources/industrial-software-plants-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Will plant leaders find our MES or IIoT platform when they ask AI for options?

They will if an assistant can find and confirm what your platform does on a real shop floor: which processes, which equipment, which industries and how fast a pilot pays back. Industrial software is still a fragmented, under-adopted market, so many plants are choosing a first system now, and B2B buyers increasingly start that research with generative AI.

This guide is for companies that sell industrial software and industrial IoT (IIoT) platforms to plants: manufacturing execution systems (MES), predictive maintenance and asset performance tools, industrial data platforms and digital twin software. It is not about robots or the integrators who install automation. The prize is a pilot at one plant that can grow into a contract across many.

## The short version

1. Most plants have no commercial system yet: [IoT Analytics](https://iot-analytics.com/mes-vendors-replace-pen-paper-spreadsheets/) estimates 54% of plants worldwide ran operations on pen, paper or spreadsheets in 2024, and just 8% used a commercial MES.
2. The market is crowded and split by vertical: more than 300 vendors serve a $5.5 billion MES market, the leader holds less than 10% share, and nine different vendors lead the 13 industries IoT Analytics analyzed.
3. Budgets are there: in [Rockwell Automation’s 2025 State of Smart Manufacturing Report](https://rockwellautomation.com/en-us/company/news/press-releases/Ninety-Five-Percent-of-Manufacturers-Are-Investing-in-AI-to-Navigate-Uncertainty-and-Accelerate-Smart-Manufacturing.html), 95% of 1,560 manufacturers had invested or planned to invest in AI and machine learning within five years.
4. Large plants commit real money: in [Deloitte’s 2025 Smart Manufacturing Survey](https://www.deloitte.com/us/en/insights/industry/manufacturing/2025-smart-manufacturing-survey.html) of 600 executives, 78% put more than 20% of their improvement budget into smart manufacturing, and execution systems were a top-two investment priority for 33%.
5. Decisions run through pilots: 41% of organizations in one 2025 survey were piloting digital twins, against 20% with full integration, so the vendor that gets on the longlist gets the chance to prove value on one line.

## Who buys industrial software for a plant, and what is a customer worth?

Operations leaders usually own the decision, with IT, OT engineers, integrators and procurement shaping the shortlist.

Deloitte found that 51% of large manufacturers say smart manufacturing initiatives are owned by operations leaders such as the chief operating officer, and 38% by technology owners such as the chief technology officer. Around them sit plant managers, controls and OT engineers, reliability and quality leaders, and procurement. System integrators matter too: IoT Analytics notes that many MES projects include a large share of integrator-led customization. How plants find those integrators is covered in [how automation integrators win industrial buyers](https://underneath.agency/resources/automation-integrators-industrial-buyers-ai-search).

The market these buyers face is large and unsettled. IoT Analytics puts MES at about 6.5% of an $85 billion industrial software market, and counts roughly 5 million factories worldwide. Its [digital twin research](https://iot-analytics.com/our-coverage/iot-platforms-software/) sizes the standalone digital twin market at $1.3 billion in 2025, forecast to reach $4.2 billion by 2030. With only 8% of plants on a commercial MES, the competition for most plants is not a rival vendor. It is a spreadsheet, a homegrown tool or an add-on to the existing ERP.

There is no public benchmark for the value of a typical MES or IIoT contract, so we do not quote one. The public data shows the shape of the opportunity instead:

- Deloitte’s respondents were companies with at least $500 million in revenue, and nearly nine in ten expected their smart manufacturing investment to continue or increase in the next fiscal year.
- Research from Siemens and S&P Global, [reported by Process Excellence Network](https://www.processexcellencenetwork.com/tools-technologies/news/manufacturing-downtime-operational-costs-digital-twins), found almost a third (30%) of organizations spending over $10 million on digital twin technology, in a study of 907 businesses.
- Deloitte’s respondents reported, on average, a 10% to 20% improvement in production output from smart manufacturing. That is the business case your champion has to defend internally.

Our inference: the value of a new customer sits less in the first license and more in the rollout. A platform proven on one line can spread to every line, then every plant.

## Where do AI assistants enter a plant software decision?

Early, when teams frame the problem and build a longlist, and then again when they check what they were told.

We found no survey that isolates how plant teams use AI to choose industrial software. The closest evidence is cross-industry. In [Gartner’s survey of 645 B2B buyers](https://www.gartner.com/en/newsroom/press-releases/2026-05-20-gartner-survey-finds-sixty-nine-percent-of-b-two-b-buyers-turn-to-sales-reps-to-validate-ai-generated-insights), 45% said they used generative AI in a recent purchase, primarily to gather information on vendors and products, and buyers used an average of seven information sources. The same survey found 69% prefer to validate AI-generated insights with sales reps. AI starts the research. It does not finish it.

Manufacturers are also adopting AI inside the plant, which makes them comfortable using it for research. Rockwell found 50% of manufacturers planned to apply AI to product quality in 2025, the top use case for the second year running. Deloitte found 24% had deployed generative AI at facility or network scale, and 23% were piloting AI and machine learning.

The traditional channels still matter. IoT Analytics sent analysts on more than 400 booth visits at [Hannover Messe 2025](https://iot-analytics.com/?p=161547), a reminder that trade shows, analyst reports and integrator advice remain part of how plants find vendors. We infer that AI answers increasingly sit in front of all of these, because an assistant summarizes trade press, analyst coverage and vendor pages in a single reply.

## Which plant questions do buyers put to AI?

Questions that describe a production problem, a vertical and an existing system, rather than a product name.

The plant questions below are our own illustrations of that pattern, not queries logged from real buyers.

| Buyer’s situation | Example question |
|---|---|
| Replacing paper | “MES for a mid-size food plant that still tracks batches on spreadsheets” |
| Downtime | “Predictive maintenance software for pumps and motors that works with our existing vibration sensors” |
| Data foundation | “Industrial data platforms that support a unified namespace and connect older PLCs” |
| Regulated production | “MES vendors with electronic batch records for pharmaceutical plants” |
| Digital twin | “Digital twin software for a packaging line where we can run a pilot in under six months” |
| Replacing a vendor | “Alternatives to our current MES for discrete assembly with better cloud options” |

Two things stand out. Vertical fit is the first filter: IoT Analytics argues that “one size fits all” is unrealistic in MES and that vendors should focus on specific verticals. And integration is the second: the buyer already owns controllers, historians and an ERP, so the question is whether your platform connects to them.

Deloitte’s priority list shows where demand is concentrated. Over the next two years, respondents ranked advanced production scheduling (35%), execution systems (33%) and quality management (28%) as their first or second system investment priorities, with 40% naming data analytics among their top solution priorities.

## How does an AI answer become a pilot and then a multi-site contract?

Through a longlist, a demo, a pilot on one line, a measured result and a rollout across plants.

1. **Framing.** An operations or engineering lead asks what kind of system fixes the problem, and which vendors do it in their industry. The answer shapes the category and the longlist.
2. **Checking.** The team reads vendor sites, analyst coverage, integrator recommendations and peer references. Few teams rely on a single source.
3. **Piloting.** One line or one plant tests the platform. Pilots are normal here: the Manufacturing IT/OT Trend Report 2025 found 41% of organizations in the pilot phase with digital twins and 20% fully integrated.
4. **Proving value.** The pilot is judged on downtime, quality, throughput or labor. Of organizations that had used digital twins, 65% reported reduced downtime and operational costs and 55% improved predictive maintenance.
5. **Rolling out.** A successful pilot becomes a program across lines and sites, with recurring subscription revenue.

AI visibility influences the first two steps. Your product, implementation team and integrators win the last three. A platform that is missing from the longlist never gets the pilot.

## What gets an MES or analytics platform named by an assistant?

Facts it can find and confirm in trusted sources; platforms document how they search, not how they rank.

**Documented by the platforms.** [OpenAI says](https://help.openai.com/en/articles/9237897-chatgpt-search) ChatGPT search rewrites a question into one or more targeted queries sent to search providers, and that a site must allow its crawler, OAI-SearchBot, to be eligible for inclusion. [Google says](https://blog.google/products/search/ai-mode-search/) AI Mode uses a “query fan-out” technique that issues multiple related searches across subtopics, and its [guidance for site owners](https://developers.google.com/search/docs/appearance/ai-features) says there are no additional requirements to appear in AI Overviews or AI Mode beyond being eligible for Search. Neither company publishes how it chooses which vendors to name.

**Observed in our studies.**

- Behind each buyer question sit several searches: [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study) counted a mean of 3.7 per question before ChatGPT answered. A question about MES for food plants may be answered from searches about traceability, batch records and cloud MES vendors.
- Google rank is a weak proxy here: [our study of AI citations and Google rankings](https://underneath.agency/research/ai-citations-google-rankings-study) found just 8.3% of the pages ChatGPT cited sat in Google’s top 10 for the question. Ranking for your category term is not the same as being cited.
- In [our crawler study](https://underneath.agency/research/ai-crawler-blocking-study), 7.4% of top sites blocked OpenAI’s search crawler. A security-minded IT team can block it without realizing what that does to visibility.

**Our inference for industrial software.** The trust factors are the ones a plant team already checks, written where machines can read them: industries and processes served, equipment and protocols supported, deployment options (cloud, edge or on-premises), OT security posture, typical pilot scope and time to value, named reference plants used with permission, integrator partnerships, analyst coverage and trade press. If those facts live only in gated PDFs and sales decks, an assistant has little to quote.

## What does an industrial software vendor lose when it is left out?

Pilots it never hears about, in a market where most plants are choosing their first system.

We have no measurement of pilots lost to AI absence, so we label the reasoning as ours:

- **The buying window is open now.** With 54% of plants on paper or spreadsheets and Rockwell reporting that 81% of manufacturers say outside and internal pressures are speeding up digital transformation, many plants are making a first selection rather than renewing an incumbent.
- **The field is wide.** Three hundred MES vendors cannot all make a longlist. When an assistant names four or five, the rest are not compared.
- **Losses are invisible.** A plant that piloted a rival never appears in your pipeline, so no report records the miss.
- **Wrong facts filter you out.** An answer that says you lack an on-premises option, or do not serve regulated industries, removes you from a pharma or defense evaluation. To find which page planted the error and get it changed, use our guide on [correcting wrong brand facts in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers).

## How does GEO work for an industrial software or IIoT company?

Generative engine optimization (GEO) makes your platform easy for AI assistants to find, describe accurately and verify.

For an industrial software vendor, the work usually covers:

1. **Vertical pages written as plain text.** One page per industry and use case, stating processes, regulations, equipment and outcomes, not just a brochure download.
2. **Integration facts.** Supported controllers, historians, ERP connectors and protocols such as OPC UA and MQTT, listed in text a crawler can read.
3. **Deployment and security facts.** Cloud, edge and on-premises options, data residency and OT security practices, because IT and security teams join the evaluation.
4. **Proof from plants.** Reference sites, measured results and case studies published with customer permission, with the plant’s industry and scope stated.
5. **Independent coverage.** Analyst reports, trade publications, conference talks and integrator partner directories. Our article on [how brands build authority for AI search](https://underneath.agency/resources/how-brands-build-authority-for-ai-search) explains why outside sources carry weight.
6. **Consistent identity.** The same product names, capabilities and partner lists on your site, marketplaces, partner pages and review platforms, especially after acquisitions and rebrands.
7. **Crawl access.** Allow the search crawlers the assistants document, and keep key pages out from behind forms.
8. **Measurement.** Ask a fixed set of plant-problem questions across ChatGPT, Gemini, Perplexity, Copilot and Google’s AI features, repeatedly, and record who is named and which sources are cited. Then compare that with where pilots actually come from.

No vendor can be promised a place on a plant’s longlist, and anyone who offers that guarantee is overselling. What this work does is make your platform the easiest one for an assistant, and then an engineer, to check. Software buyers in other categories follow a similar path; see [how ERP vendors get on AI-built long lists](https://underneath.agency/resources/erp-software-ai-search) and [whether AI assistants shape enterprise software shortlists](https://underneath.agency/resources/enterprise-software-shortlists-ai-search). Engineering teams choosing design tools face a similar test, covered in [how CAD and simulation vendors get recommended](https://underneath.agency/resources/engineering-software-ai-search).

## Which questions can’t the data answer yet for industrial software?

It shows manufacturers investing and buyers using AI, not how many pilots AI answers create.

- **No plant-specific AI buying data.** The AI usage figures here come from cross-industry B2B surveys. We found no public survey of how plant teams use AI to choose MES or IIoT software.
- **Sponsors have interests.** Rockwell sells automation and software, Siemens sells digital twin software, and IoT Analytics sells market reports to vendors.
- **No ranking studies for this category.** Our studies covered buyer questions across several industries. Applying them to industrial software is our inference.
- **Contract values are private.** Public figures show budgets and market sizes, not the typical value of a plant subscription or a multi-site program.

## Where should an industrial software company start?

Ask assistants the plant problems your best customers had before their first pilot, and see whether you are named.

That first check usually shows whether your platform appears for its core use cases and verticals, whether integrations and deployment options are described correctly, which analyst, trade and integrator sources the answers rely on, and which rivals are named instead.

If your growth depends on turning a handful of pilots into multi-site programs each year, [talk to us about a visibility review](https://underneath.agency/contact). We will show where your platform appears when plant teams ask AI for options, why rivals are named instead, and which fixes are most likely to put you on more pilot shortlists. Our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page sets out how vertical pages, integration facts and plant proof are built and measured for MES and IIoT vendors.

## Frequently asked questions

### Do plant managers really use ChatGPT to research MES or IIoT vendors?

There is no plant-specific survey yet. Cross-industry surveys show many business buyers use generative AI to gather vendor information, then check it with people and other sources.

### Should we publish integration and pricing details publicly?

Publish integration, deployment and industry facts in plain text. Pricing is a business decision; at minimum, say how pilots are scoped and what drives cost, so an assistant does not guess.

### Do analyst reports on MES still count when plant teams ask AI?

Yes. Assistants cite outside sources, and analyst coverage, trade press and integrator directories are among the sources plant teams already trust. We infer they also shape what assistants say.

### Can a niche vendor compete with the largest automation companies?

In vertical questions, often yes. The MES leader holds less than 10% share, and nine different vendors lead the 13 industries IoT Analytics studied. Specific facts help a specialist win the specific question.

## Sources

- IoT Analytics (2025-12-15), [Manufacturing Execution Systems: The 300+ vendors looking to displace pen, paper, and spreadsheets in the factory](https://iot-analytics.com/mes-vendors-replace-pen-paper-spreadsheets/)
- IoT Analytics (2026-08-26), [Inside the $1.3B digital twin market (IoT platforms and software coverage)](https://iot-analytics.com/our-coverage/iot-platforms-software/)
- IoT Analytics (2025-05), [Hannover Messe 2025: the latest Industrial IoT/Industry 4.0 Trends](https://iot-analytics.com/?p=161547)
- Rockwell Automation (2025-06-03), [Ninety-Five Percent of Manufacturers Are Investing in AI to Navigate Uncertainty and Accelerate Smart Manufacturing](https://rockwellautomation.com/en-us/company/news/press-releases/Ninety-Five-Percent-of-Manufacturers-Are-Investing-in-AI-to-Navigate-Uncertainty-and-Accelerate-Smart-Manufacturing.html)
- Deloitte (2025), [2025 Smart Manufacturing and Operations Survey](https://www.deloitte.com/us/en/insights/industry/manufacturing/2025-smart-manufacturing-survey.html)
- Process Excellence Network (2025-04-23), [Manufacturing firms reduce downtime and operational costs with digital twins](https://www.processexcellencenetwork.com/tools-technologies/news/manufacturing-downtime-operational-costs-digital-twins)
- Gartner (2026-05-20), [Gartner Survey Finds 69% of B2B Buyers Turn to Sales Reps to Validate AI-Generated Insights](https://www.gartner.com/en/newsroom/press-releases/2026-05-20-gartner-survey-finds-sixty-nine-percent-of-b-two-b-buyers-turn-to-sales-reps-to-validate-ai-generated-insights)
- OpenAI Help Center (2026), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Google (2025-03-05), [Expanding AI Overviews and introducing AI Mode](https://blog.google/products/search/ai-mode-search/)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude cite pages that rank?](https://underneath.agency/research/ai-citations-google-rankings-study)
- Underneath (2026), [Which AI crawlers do top websites block?](https://underneath.agency/research/ai-crawler-blocking-study)

---

This is the Markdown twin of https://underneath.agency/resources/industrial-software-plants-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How insurers win more quotes when shoppers ask AI"
description: "By being named, correctly and state by state, when shoppers ask AI to explain and compare coverage. 32% of auto shoppers already use AI tools."
canonical: "https://underneath.agency/resources/insurance-companies-customers-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How can an insurance carrier or agency win more quotes when shoppers ask AI?

By being named, and described correctly for each state and line of coverage, when shoppers ask AI assistants to explain and compare policies. Nearly a third of auto insurance shoppers already use AI tools, and those who do switch insurers more often. For carriers and agencies, that makes AI answers a new place where quote requests are won or lost before anyone visits a comparison site.

## The short version

1. AI is already in the shopping journey: [J.D. Power](https://www.cbtnews.com/auto-insurers-customer-gaps/) reports 32% of auto insurance shoppers used AI tools, and they were more than 1.3 times as likely to switch insurers.
2. Shoppers compare more than ever: J.D. Power’s [2026 shopping study](https://www.cbtnews.com/digital-is-the-new-front-door-for-auto-insurance/) found 53% of customers shopped, collecting 3.5 quotes on average, and 48% of new auto policies were bought digitally.
3. Each switch carries real premium: among switchers, the median premium moving between carriers is more than $3,200, according to J.D. Power’s [quarterly shopping report](https://hub.jdpower.com/hubfs/FSAAS/Insurance/26-JDP_LIST%20Report%20Q2%202026_FINAL.pdf).
4. Life insurance needs explaining: [LIMRA and Life Happens](https://www.limra.com/siteassets/newsroom/liam/2026/2026-barometer-member-pr-toolkit.pdf) estimate 92 million US adults need coverage or more of it, and only 32% call themselves very or extremely knowledgeable about it.
5. Google answers most insurance searches with AI: in [our study of 1,248 searches](https://underneath.agency/research/ai-overviews-frequency-study), 88.0% of financial services and insurance keywords showed an AI Overview.

A note before you read: this article is about how insurers appear in AI answers. It recommends no policy or carrier, and it is not legal or regulatory advice.

## Who shops for insurance now, and what is a new policyholder worth?

Households at renewal or after a life event, comparing several quotes, and increasingly buying online or after online research.

Auto insurance is the front door. In the J.D. Power 2026 U.S. Insurance Shopping Study of 12,437 customers, the share shopping fell from 57% to 53% as rate increases cooled, but customers collected an average of 3.5 quotes, the most in the study’s history. Nearly half (48%) of new auto policies were bought digitally, up from 36% five years earlier.

A new policyholder is worth more than one policy. J.D. Power’s quarterly report puts the median premium moving between carriers among switchers at more than $3,200. The bigger prize is the household: 45% of active auto shoppers said they have a homeowners policy, but only 20% received a homeowners quote while shopping. As J.D. Power’s Stephen Crewdson put it, if the auto quote isn’t competitive, customers “don’t stick around to discuss home, life or other financial products.”

Life insurance follows a different path. The [2026 Insurance Barometer](https://www.limra.com/siteassets/newsroom/liam/2026/2026-barometer-member-pr-toolkit.pdf), a survey of 5,237 adults, found 52% own at least one policy and 38% have a need-gap. 43% would research online but buy in person, 25% would research and buy entirely online, and 45% work with a financial professional. Demand is growing: [LIMRA](https://rethinking65.com/life-insurance-growth-opportunities-shifting-in-2026-limra/) reports new annualized individual life premium reached $12.7 billion in the first nine months of 2025, a 12% increase.

For agencies, this means the research often happens before the first call. For carriers, it means the comparison starts before the quote form. Health plans run on their own enrollment calendar; see [how health insurers win members through AI](https://underneath.agency/resources/health-insurers-members-ai-search).

## Where is AI already part of insurance shopping?

At the explanation and comparison stage: shoppers use AI to understand coverage, compare policies and decide whether to switch.

J.D. Power’s data is the most direct evidence:

- **Use.** 32% of auto insurance shoppers used AI tools during their search, most often for general questions, quotes, policy comparisons and decisions.
- **Quality.** A similar share, 33%, found the content unhelpful.
- **Switching.** Shoppers who used AI were more than 1.3 times as likely to switch insurers as those who did not. Newer digital insurers court these switchers, as [how insurtechs win switching shoppers](https://underneath.agency/resources/insurtech-customers-ai-search) explains.
- **Understanding.** Only 58% of customers said they completely understand their auto policy, down 4 points from 2025. J.D. Power’s reading is that when insurers fail to explain coverage, customers turn to AI to fill the gap.

J.D. Power’s quarterly report adds an early look at a new study: two-thirds of consumers use AI as part of researching insurance coverage, and more than one-third made a policy change based on the advice.

Google adds AI to many of these searches by default. Across the eight industries in our AI Overviews study, financial services and insurance had one of the highest rates: 88.0% of keywords showed an AI Overview, or 77.7% after adjusting for the mix of searches.

Our inference: in insurance, AI is acting as the explainer that many carriers’ own materials are not. Whoever supplies the clearest explanation has a good chance of shaping the shortlist.

## Which questions do insurance shoppers ask AI?

Questions about coverage, cost, comparisons and claims, almost always tied to a state. We wrote the shopper prompts below as examples; they were not collected from real policyholders.

| Line | Illustrative prompt |
|---|---|
| Auto | “Cheapest full coverage for a 17-year-old driver in Texas?” |
| Auto | “Is usage-based insurance worth it if I drive under 8,000 miles a year?” |
| Home | “Does homeowners insurance cover water backup in Florida?” |
| Home and auto | “Which insurers give the best bundle discount in Ohio?” |
| Life | “How much term life insurance do I need with two kids and a mortgage?” |
| Agency | “Independent insurance agent near me who writes flood insurance” |
| Claims | “Which insurers handle hail claims fastest in Colorado?” |

Assistants also run searches of their own to answer these. In [our study of the searches AI assistants run behind the scenes](https://underneath.agency/research/ai-hidden-searches-study), one example was “J.D. Power homeowners insurance customer satisfaction Florida 2025”, an observed search for a third-party ranking in a specific state. Ratings and rankings are part of how assistants research insurers, and the state is part of the question.

## How does AI visibility turn into quotes and policies?

Through the shortlist: an answer names carriers or agencies, the shopper requests quotes, and one becomes a policy.

The path, as we infer it from the documented shopping process:

1. **Trigger.** A renewal increase, a new car or home, or a life event.
2. **Explanation.** The shopper asks an assistant what coverage they need and who offers it in their state.
3. **Shortlist.** The answer names a few carriers, comparison sites or agencies.
4. **Quotes.** The shopper requests quotes, on average 3.5, through apps, websites or an agent.
5. **Policy and household.** One quote becomes a policy, renewals follow, and a competitive auto quote opens the door to home and life.

Carriers already pay heavily for the quote-request stage. [EverQuote](https://s25.q4cdn.com/599700560/files/doc_news/EverQuote-Announces-Second-Quarter-2026-Financial-Results-2026.pdf), an online insurance marketplace, reported second-quarter 2026 revenue of $195.1 million, up 25%, including $172.1 million from auto insurance and $23.0 million from home and renters, and said carriers “continue to target growth across digital channels.” That revenue comes from insurers and agents buying shopping traffic.

An AI answer sits one step earlier. Our inference: a carrier that assistants name and explain well can start conversations it would otherwise pay to buy, while a carrier that is missing must compete for the same shopper later, at a price.

## What decides which insurers an AI answer names?

Independent ratings and comparison sites, the insurer’s own clear explanations, and the shopper’s location. Only some of this is documented.

**Documented by platforms.** Google says its systems give even more weight to strong expertise and trust signals on [“Your Money or Your Life” topics](https://developers.google.com/search/docs/fundamentals/creating-helpful-content), which include financial stability, and that such content “must be highly accurate.”

**Observed in studies.**

- **Comparison sites and insurers’ own pages are both cited.** In [our study of 4,051 AI Overview citations](https://underneath.agency/research/ai-overview-citations-study), nerdwallet.com was cited in 10.0% of the 481 AI Overviews, libertymutual.com in 7.5%, allstate.com in 6.9% and thezebra.com in 6.0%. Insurers’ own sites were among the most-cited sources in their category.
- **Ranking is not enough.** In finance and insurance, 34.8% of AI Overview citations were page-one organic results for the same search, so most cited pages came from elsewhere.
- **Editorial sources lead in money questions.** A [study of banking questions](https://arxiv.org/abs/2509.08919) by Mahe Chen and colleagues found 64.6% of cited sources were editorial or third-party sites. Fund questions show the same pull toward outside sources, with assistants finding fund families in trusted fund research, as [how asset managers get funds considered](https://underneath.agency/resources/asset-managers-fund-demand-ai-search) shows.
- **Forums matter less here.** Reddit was cited in 5.7% of financial services and insurance AI Overviews in [our Reddit study](https://underneath.agency/research/ai-reddit-citations-study), among the lowest of the eight industries.
- **Assistants disagree.** In [our four-assistant study](https://underneath.agency/research/ai-assistants-brand-agreement-study), recommended options for financial services and insurance questions overlapped by 0.332, against 0.543 for business software.
- **Location changes the answer.** In [our country study](https://underneath.agency/research/ai-recommendations-by-country-study), the country gap in ChatGPT’s recommendations was 0.180 for insurance, lending and tax questions, against 0.034 for global products.

**Our inference for insurers.** The country study tested countries, not states, but insurance is sold and priced state by state. An insurer that is clear about where it writes each line, and what its coverage includes there, gives assistants what they need to place it in the right answer.

## How do state rules shape what insurers publish for AI?

They set the limits: content that assistants repeat must be accurate for each state where it applies.

Insurance is regulated by the states, and products, rates and availability differ between them. The NAIC’s [Unfair Trade Practices Act model law](https://content.naic.org/sites/default/files/model-law-880.pdf), the template for many state laws, prohibits misrepresentations in insurance advertising, including “any intentional misquote of premium rate,” and covers advertising on the internet.

That has two practical consequences, which we infer:

1. **Clarity by state is a compliance task and a visibility task.** A page that says which lines you write in which states, and what each policy covers there, helps shoppers, regulators and assistants alike.
2. **Stale pages are a risk.** Old rate examples, retired discounts or outdated availability can be picked up and repeated. Our guide to [fixing wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) explains how to trace an error back to its source.

States are also writing rules for insurers’ own use of AI. The NAIC’s [map of states adopting its AI model bulletin](https://content.naic.org/sites/default/files/legal-adoption-map-ai-model-bulletin.pdf) shows how quickly that is spreading. That bulletin governs decisions such as underwriting and claims rather than marketing, but it signals how closely regulators watch AI in this industry.

## What does GEO look like for a carrier or an agency?

For insurers, generative engine optimization (GEO) means clear coverage explanations, consistent facts and independent evidence assistants can trust.

1. **Explain coverage in plain language, by state.** What a policy covers, what it does not, and which discounts apply where. J.D. Power’s finding that only 58% fully understand their policy shows the gap.
2. **Keep facts consistent.** States served, lines written, discounts, claims steps and contact details should match across your site, agency profiles and comparison listings.
3. **Earn third-party evidence.** Customer satisfaction rankings, financial strength ratings, editorial reviews and comparison sites are what assistants cite. Our guide on [which pages to target for AI recommendations](https://underneath.agency/resources/best-of-lists-ai-recommendations) covers ranked lists.
4. **Give agencies a local presence.** For agents, accurate local profiles and reviews matter, since local answers vary more than national ones.
5. **Use video for explanation.** LIMRA found YouTube (66%) has overtaken Facebook as the top platform for financial information.
6. **Review with compliance.** Every page assistants might repeat is advertising under state rules.
7. **Measure by state and by engine.** Answers change from run to run and between assistants; see [why AI answers about your brand change](https://underneath.agency/resources/why-ai-answers-about-your-brand-change).

No one can promise an assistant will recommend a particular insurer. GEO improves what assistants find, so that when they do mention you, the description is correct.

## What can’t the research tell insurers yet?

How assistants choose among insurers state by state, and how many policies AI answers produce.

- **Surveys, not logs.** The AI usage figures come from J.D. Power surveys, which show behavior shoppers report, not what assistants said to them.
- **Countries, not states.** Our location study compared countries. We found no public study of how AI answers differ between US states.
- **No link to bound policies.** We found no public data connecting an insurer’s AI visibility with quotes or policies written.
- **Fast change.** Assistants and Google’s AI features change often, so any snapshot ages quickly.

## Where should an insurance company start?

Ask assistants your customers’ coverage and comparison questions, state by state, and record who is named and why.

That check shows whether assistants know where you write each line, whether your coverage and discounts are described correctly, and which ratings, comparison sites or competitors shape the answers instead of you. For agencies, it shows whether you appear for local questions at all.

If more quote starts, bundles or life applications are on your plan, [talk to us about an AI visibility review for your lines and states](https://underneath.agency/contact). We will test the questions behind your most valuable lines in the states that matter, trace the sources behind each answer, and plan the content, coverage and data fixes, reviewed with your compliance team, that give you a fair chance of being named. How carriers and agencies keep coverage explanations and state-by-state facts consistent over time is described on our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page.

## Frequently asked questions

### Do AI assistants give insurance quotes?

Some shoppers use AI to compare quotes, and J.D. Power counts quotes among common uses. A binding price still comes from the insurer or agent, so assistants mostly shape who gets asked.

### Why does an assistant name comparison sites instead of insurers?

Comparison sites publish the rankings and explanations assistants draw on. In our AI Overview study, NerdWallet and The Zebra were cited often, but so were Liberty Mutual and Allstate’s own sites.

### Does AI search matter for independent agencies?

Yes, especially for local questions. Shoppers ask for agents near them, and answers that name a place agree far less across assistants, so accurate local profiles and reviews matter.

### Can an insurer correct a wrong AI answer about its coverage?

Usually by correcting the source. Find the page the answer relies on, often an old page of your own or a third-party listing, and update it.

## Sources

- J.D. Power, via CBT News (2026-06-09), [2026 U.S. Auto Insurance Study: auto insurers struggle to maintain seamless interactions across channels](https://www.cbtnews.com/auto-insurers-customer-gaps/)
- J.D. Power, via CBT News (2026-06-04), [2026 U.S. Insurance Shopping Study: digital becomes the new front door for auto insurance shopping](https://www.cbtnews.com/digital-is-the-new-front-door-for-auto-insurance/)
- J.D. Power (2026), [Loyalty Indicator and Shopping Trends Report, Q2 2026](https://hub.jdpower.com/hubfs/FSAAS/Insurance/26-JDP_LIST%20Report%20Q2%202026_FINAL.pdf)
- LIMRA and Life Happens (2026), [2026 Insurance Barometer Study toolkit](https://www.limra.com/siteassets/newsroom/liam/2026/2026-barometer-member-pr-toolkit.pdf)
- Rethinking65 (2026), [Life-Insurance Growth Opportunities Shifting in 2026: LIMRA](https://rethinking65.com/life-insurance-growth-opportunities-shifting-in-2026-limra/)
- EverQuote (2026-08-03), [EverQuote Announces Second Quarter 2026 Financial Results](https://s25.q4cdn.com/599700560/files/doc_news/EverQuote-Announces-Second-Quarter-2026-Financial-Results-2026.pdf)
- NAIC, [Unfair Trade Practices Act (Model 880)](https://content.naic.org/sites/default/files/model-law-880.pdf)
- NAIC (2026-08-31), [Implementation of NAIC Model Bulletin: Use of Artificial Intelligence Systems by Insurers](https://content.naic.org/sites/default/files/legal-adoption-map-ai-model-bulletin.pdf)
- Google Search Central (2025), [Creating helpful, reliable, people-first content](https://developers.google.com/search/docs/fundamentals/creating-helpful-content)
- Chen, Wang, Chen and Koudas (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919)
- Underneath (2026), [When does Google show an AI Overview? 1,248 US searches](https://underneath.agency/research/ai-overviews-frequency-study)
- Underneath (2026), [Do AI Overviews cite the pages that rank? 4,051 citations](https://underneath.agency/research/ai-overview-citations-study)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [When does a Reddit thread become evidence in Google’s AI?](https://underneath.agency/research/ai-reddit-citations-study)
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude agree on brands?](https://underneath.agency/research/ai-assistants-brand-agreement-study)
- Underneath (2026), [Same question, four countries: do AI recommendations change?](https://underneath.agency/research/ai-recommendations-by-country-study)

---

This is the Markdown twin of https://underneath.agency/resources/insurance-companies-customers-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How insurtechs win customers from incumbents in AI search"
description: "By being the clear, well-reviewed alternative AI can explain when shoppers ask about switching. Insurify says AI-sourced new customers rose from 0.15% to 5.7%."
canonical: "https://underneath.agency/resources/insurtech-customers-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How can an insurtech win customers from incumbent insurers when shoppers ask AI?

By becoming the alternative an assistant can explain and vouch for: a clear category story, honest comparisons with the big carriers, and reviews and coverage that answer “is this legit?”. AI already sits in the insurance shopping journey, and shoppers who use it switch more often, which favors challengers. But assistants default to familiar brands, so an insurtech has to earn its place in the answer.

## The short version

1. AI is becoming a customer source for insurtechs: [Insurify](https://iireporter.com/insurify-expands-chatgpt-insurance-plugin/) says the share of its new customers who found it through ChatGPT and AI conversations rose from 0.15% in the first half of 2025 to 5.7% in the first half of 2026.
2. AI users are switchers: [J.D. Power](https://www.cbtnews.com/auto-insurers-customer-gaps/) found 32% of auto insurance shoppers used AI tools, and they were more than 1.3 times as likely to switch insurers.
3. Customers are expensive to win: [Lemonade](https://coverager.com/lemonade-reports-q2-2026-results/) spent $78 million on sales and marketing in the second quarter of 2026, a quarter in which it added about 166 thousand customers.
4. Distribution is shifting to partners: [Root](https://www.insurancejournal.com/news/national/2026/08/06/880452.htm) said partnership and independent agent channels produced about 51% of its new writings, up from about 44% a year earlier.
5. Early-stage money is tighter: [Gallagher Re](https://riskandinsurance.com/ai-dominates-insurtech-funding-in-q2-as-early-stage-deals-cool-sharply/) reports early-stage insurtech funding fell 51.8% in the second quarter of 2026, to $264.19 million.

A note before you read: this article is about how insurtech brands appear in AI answers. It does not recommend any policy, and none of it should be read as legal or regulatory guidance for a licensed carrier or agency.

## Why is AI search a different problem for an insurtech than for an incumbent?

Assistants lean on well-documented brands, and incumbents have decades of documentation while a startup has a few years.

State Farm, Geico and Progressive have years of ratings, press, comparison-site listings and customer reviews behind them. An insurtech may have a better product and a fraction of that record. Two studies show what that means in AI answers:

- **Known by name, missing from discovery.** In a [study of 112 Product Hunt startups](https://arxiv.org/abs/2601.00912) by Amit Prakash Sharma, ChatGPT recognized the products 99.4% of the time when asked about them by name, but named them in only 3.32% of open discovery questions.
- **Familiar brands win ties.** In a [controlled test](https://arxiv.org/abs/2606.17443) by Xi Chu and YuPeng Hou, assistants recommended the well-known brand 100% of the time when products had identical specifications. A competitor’s rating edge of less than 0.1 stars was enough to break that.

Neither study covered insurance. Our inference for insurtechs: when a shopper asks an open question, such as “who has the cheapest renters insurance?”, the assistant will reach for familiar carriers unless it finds specific, credible evidence that a newer option is different. The incumbents’ side of this contest is covered in [how carriers and agencies win quotes from AI](https://underneath.agency/resources/insurance-companies-customers-ai-search).

The funding climate raises the stakes. Gallagher Re counted $2.44 billion of insurtech funding in the second quarter of 2026, the most since 2022, but 99.1% went to AI-focused companies and early-stage funding fell sharply. For most startups, we infer, a channel built on evidence rather than ad spend is worth more when capital is concentrated.

## Where are insurtechs already meeting shoppers inside AI assistants?

Inside the assistants themselves, through in-chat apps, and in the answers shoppers read before they request a quote.

Insurify, a digital insurance agent, shows the in-chat route. In February 2026 it launched what [FinTech Global](https://fintech.global/2026/02/09/insurify-launches-industry-first-chatgpt-insurance-app/) describes as the industry’s first insurance app in ChatGPT, drawing on more than 196m auto insurance quotes and over 70,000 verified customer reviews. In August it [added real-time quotes](https://www.webull.com/news/15429157729682432) for shoppers in a first group of states, including Arizona, Ohio and Wisconsin. Insurify reports that in the six months after launch, ChatGPT-referred traffic grew 151% and revenue from those visitors grew 212%. These are the company’s own figures, but they are the clearest public numbers we found on AI as an insurance acquisition channel.

The broader shift shows up in J.D. Power’s data. Shoppers who used AI were more than 1.3 times as likely to switch insurers. J.D. Power reads this as customers turning to AI when insurers fail to explain coverage, which moves control of information away from carriers.

That matters more to a challenger than to an incumbent. Our inference: a shopper who is willing to switch, and is asking an assistant to explain the options, is exactly the shopper an insurtech needs to reach, provided the assistant knows the insurtech exists.

## Which questions decide whether an insurtech is considered?

Alternative, comparison and legitimacy questions, plus questions about new kinds of coverage. The examples below are our own drafts of typical questions, not prompts captured from real shoppers or partners.

| Stage | Illustrative prompt |
|---|---|
| Alternatives | “Alternatives to Geico for a new driver with a clean record” |
| Comparison | “Lemonade vs State Farm for renters insurance in Illinois” |
| Category | “Is pay-per-mile car insurance worth it if I work from home?” |
| Legitimacy | “Is Root insurance legit, and do they actually pay claims?” |
| Embedded | “Should I buy the insurance my car dealer offers at checkout?” |
| Partner search | “Embedded insurance providers for an online used-car marketplace” |

The legitimacy question deserves special attention. In [our study of “Is this brand legit?” answers](https://underneath.agency/research/is-it-legit-ai-reputation-study) across 79 brands, every complete answer said the brand was legitimate, but 99.7% raised at least one problem. 88.0% of answers cited a review or complaint platform, and Trustpilot and the BBB made up 61.7% of review-platform citations. For an insurtech, how complaints are handled on those platforms becomes part of the answer.

## How does AI visibility turn into customers for an insurtech?

It puts the insurtech on the shortlist of a shopper ready to switch, who can quote and buy in minutes.

For a direct insurtech, the path is short. A shopper asks about alternatives, the answer names the company, the shopper checks reviews, opens the app, gets a quote and binds. The economics show why that path is valuable. Lemonade ended the second quarter of 2026 with 3,308,666 customers, up 23%, and in-force premium of $1.43 billion, up 32%. Premium per customer was $433. It also spent $78 million on sales and marketing in the quarter, up from $60 million a year earlier.

Lemonade has grown from 1.00 million customers at the end of 2021, according to [The Motley Fool](https://www.fool.com/investing/2026/07/31/lemonade-cut-its-full-year-in-force-premium-outloo/), mostly by paying to acquire them. Our inference: every customer who arrives because an assistant named and explained the company is one the marketing budget did not have to buy. Insurify’s reported rise in AI-sourced customers, from 0.15% to 5.7% of new customers in a year, shows how fast that share can move once a company is present where shoppers ask.

## How does embedded insurance change the AI question?

It adds partners choosing an insurance provider, and shoppers deciding whether to trust the policy offered at checkout.

Distribution is moving toward partners. Root, which now operates in 37 states with 483,921 policies in force, said partnership and independent agent channels made up about 51% of new writings, while its direct channel “remained challenging.” Lemonade’s car product is now available in states representing nearly 50% of the US car insurance market, and offers Tesla owners 50% off every mile driven using Tesla’s Full Self-Driving technology.

Shoppers are open to it. J.D. Power’s [2026 shopping study](https://www.cbtnews.com/digital-is-the-new-front-door-for-auto-insurance/) found 36% of recent shoppers interested in buying embedded insurance sold through a car dealer or manufacturer. Usage-based insurance, a common insurtech product, was used by 34% of those buying from a new insurer.

Our inference for embedded providers: two kinds of AI question matter. Shoppers ask whether the policy offered at checkout is a good deal, so the provider needs a clear, independent reputation. And partners, such as marketplaces, dealers and lenders, research providers too, so business-facing pages and trade coverage matter as much as consumer reviews.

## What decides whether an insurtech is named?

Independent evidence of a real difference, a visible review record, and clear facts about what is sold where. Platforms document only part of this.

**Documented by platforms.** Google says its systems give even more weight to strong expertise and trust signals on [“Your Money or Your Life” topics](https://developers.google.com/search/docs/fundamentals/creating-helpful-content), which include financial stability.

**Observed in studies.**

- **Independent coverage.** In [our brand study](https://underneath.agency/research/brand-entity-ai-recommendations-study), each tenfold increase in independent sites naming a brand went with 4.7 times the odds of being recommended.
- **Search foundations.** In the Product Hunt study, referring domains, an ordinary search signal, predicted which startups Perplexity named.
- **A visible difference.** In the controlled brand test, a small, verifiable advantage broke the default toward the known brand.
- **Review platforms.** In our legitimacy study, review and complaint platforms carried most of the negative claims assistants made.

**Our inference for insurtechs.** The difference has to be specific and checkable: a price structure that is genuinely different, a claims process that reviewers confirm, a coverage feature incumbents lack, and plain facts about which states you write in. Inflated claims are a risk in a regulated industry, and [GEO can backfire](https://underneath.agency/resources/can-geo-backfire-on-your-brand) when it outruns the facts.

## What does GEO look like for an insurtech?

For an insurtech, generative engine optimization (GEO) means building, quickly and precisely, the evidence incumbents already have.

1. **Own the category explainer.** Explain usage-based, pay-per-mile or embedded insurance honestly: how it works, what it costs, who it suits and who it does not.
2. **Compare yourself with incumbents honestly.** Keep comparison pages current on price structure, coverage and claims. Our research on [whether comparison pages help brands get cited](https://underneath.agency/resources/do-comparison-pages-help-b2b-ai-citations) shows what tends to work.
3. **Answer the legitimacy question in public.** Maintain profiles on Trustpilot, the BBB and app stores, and respond to complaints, since these are what assistants cite.
4. **State availability plainly.** Which products you sell in which states, kept up to date as you expand.
5. **Earn independent coverage.** Insurance trade press, personal finance media and comparison platforms give assistants third-party evidence.
6. **Serve partners too.** Embedded providers need clear business-facing pages and trade coverage for the platforms evaluating them.
7. **Check launches.** New products are often missing at first; [why ChatGPT misses newly launched products](https://underneath.agency/resources/why-chatgpt-misses-new-products) explains why.

The broader playbook for challengers is in [how a small brand can get recommended by AI assistants](https://underneath.agency/resources/how-small-brands-get-recommended-by-ai), and the wider evidence on incumbents is in [do AI assistants favor big brands?](https://underneath.agency/resources/do-ai-assistants-favor-big-brands). No one can guarantee an assistant will recommend a startup; GEO makes the evidence it finds accurate and easy to verify.

## Where is the evidence on insurtechs still thin?

On how assistants treat insurance startups specifically, and on how many policies AI answers produce across the industry.

- **No insurance-specific test.** The studies of new and incumbent brands used startups across categories and consumer products, not insurers.
- **Company-reported figures.** Insurify’s growth numbers are its own, from a company with a product in ChatGPT, and not independently audited.
- **Surveys.** J.D. Power’s AI figures describe what shoppers report, not what assistants told them.
- **No industry-wide link to policies.** Beyond one company’s figures, we found no public data connecting insurtech AI visibility with policies written.

## Where should an insurtech start?

Ask assistants the alternative, comparison and legitimacy questions your target customers ask, and note which incumbents and sources win.

That first check shows whether assistants understand your category, whether they name you next to the carriers you compete with, how they answer “is it legit?”, and whether your state availability and pricing model are described correctly. It also shows which reviews, articles or comparison sites the incumbents have that you lack.

If your growth plan depends on lowering what you pay to win each customer, or on winning partners for embedded distribution, [ask us to map your AI visibility against the incumbents](https://underneath.agency/contact). We will test the questions that decide whether you are considered, trace the evidence behind each answer, and plan the coverage, content and review work, checked with your compliance team, that gives you a fair chance of being named. To see how a challenger’s category explainers, comparison pages and review profiles are planned and then tracked against incumbents, read our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page.

## Frequently asked questions

### Can a new insurer appear in ChatGPT answers at all?

Yes, but rarely at first. In one study, startups were recognized 99.4% of the time when named, but appeared in only 3.32% of open discovery questions.

### Do in-chat apps replace being named in AI answers?

No. An app helps shoppers who already chose it, while answers decide who gets considered. Insurify uses both: an app in ChatGPT and a brand that shoppers discover through AI conversations.

### How do bad reviews affect an insurtech in AI answers?

They tend to appear. In our legitimacy study, 99.7% of answers raised at least one problem, often drawn from review and complaint platforms, so visible, well-handled complaints matter.

### Should embedded insurance providers care about AI search?

Yes. Shoppers ask whether a checkout policy is a good deal, and partner platforms research providers before signing, so both consumer reputation and business-facing evidence matter.

## Sources

- Insurance Innovation Reporter (2026-08), [Insurify Expands ChatGPT Insurance Plugin](https://iireporter.com/insurify-expands-chatgpt-insurance-plugin/)
- Insurify, via PR Newswire and Webull (2026-08-19), [Insurify Expands ChatGPT Plugin with Real-Time Personalized Quotes and In-Chat Shopping](https://www.webull.com/news/15429157729682432)
- FinTech Global (2026-02-09), [Insurify launches industry-first ChatGPT insurance app](https://fintech.global/2026/02/09/insurify-launches-industry-first-chatgpt-insurance-app/)
- J.D. Power, via CBT News (2026-06-09), [2026 U.S. Auto Insurance Study: auto insurers struggle to maintain seamless interactions across channels](https://www.cbtnews.com/auto-insurers-customer-gaps/)
- J.D. Power, via CBT News (2026-06-04), [2026 U.S. Insurance Shopping Study: digital becomes the new front door for auto insurance shopping](https://www.cbtnews.com/digital-is-the-new-front-door-for-auto-insurance/)
- Coverager (2026-07-29), [Lemonade reports Q2 2026 results](https://coverager.com/lemonade-reports-q2-2026-results/)
- Reinsurance News (2026-07-29), [Lemonade’s revenue jumps 79% to $294m in Q2’26](https://www.reinsurancene.ws/lemonades-revenue-jumps-79-to-294m-in-q226/)
- The Motley Fool (2026-07-31), [Lemonade cut its full-year in-force premium outlook](https://www.fool.com/investing/2026/07/31/lemonade-cut-its-full-year-in-force-premium-outloo/)
- Insurance Journal (2026-08-06), [Auto Insurer Root Inc. Reports 12% Increase in Q2 Net Income](https://www.insurancejournal.com/news/national/2026/08/06/880452.htm)
- Risk & Insurance (2026-08), [AI Dominates Insurtech Funding in Q2 As Early-Stage Deals Cool Sharply](https://riskandinsurance.com/ai-dominates-insurtech-funding-in-q2-as-early-stage-deals-cool-sharply/)
- Amit Prakash Sharma (2025), [The Discovery Gap: How Product Hunt Startups Vanish in LLM Organic Discovery Queries](https://arxiv.org/abs/2601.00912)
- Xi Chu and YuPeng Hou (2026), [Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems](https://arxiv.org/abs/2606.17443)
- Google Search Central (2025), [Creating helpful, reliable, people-first content](https://developers.google.com/search/docs/fundamentals/creating-helpful-content)
- Underneath (2026), [“Is this brand legit?” How AI assistants build a reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study)
- Underneath (2026), [Do Wikipedia and schema make AI assistants recommend a brand?](https://underneath.agency/research/brand-entity-ai-recommendations-study)

---

This is the Markdown twin of https://underneath.agency/resources/insurtech-customers-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How investing apps and robo-advisors get chosen in AI answers"
description: "By being the clear answer for a specific use case, with fees and safety facts stated plainly and confirmed by the comparison sites AI cites."
canonical: "https://underneath.agency/resources/investment-platforms-customers-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How do investing apps and robo-advisors get chosen when people ask AI which platform to use?

They get chosen by being the obvious fit for a specific need, such as a first account, options trading or automated investing, in the comparison and review sources AI assistants cite. Most investors now use AI somewhere in their research, and the answers to “which app should I use?” lean on the review sites that compare brokers by use case. For an investment platform, AI visibility is won with plain fee, feature and safety facts that independent reviewers confirm.

## The short version

1. AI is part of retail investing: in an [Investing.com survey](https://hedgefundalpha.com/news/retail-investors-now-use-ai-to-inform-investment-decisions/) of 938 US investors in April 2026, 62% had used AI tools to help with investment decisions, and 54% had used chatbots such as ChatGPT for investing research.
2. Answers lean on comparison sites: in [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study), a source such as NerdWallet that ChatGPT’s own search had named was cited 44.0% of the time, against 8.1% otherwise.
3. A funded customer is valuable: [Robinhood](https://mondovisione.com/news/robinhood-reports-second-quarter-2026-results-2026729/) reported 28.4 million funded customers and average revenue per user of $187 a year in the second quarter of 2026.
4. Answers vary by assistant: in [our agreement study](https://underneath.agency/research/ai-assistants-brand-agreement-study), 66.3% of the options four assistants recommended were named by only one of them.

This guide covers how investment platforms get found and described in AI answers. Nothing here is investment, legal or regulatory advice, and nothing here suggests any platform or product suits any investor.

## Who chooses an investment platform, and what is a funded account worth?

Individuals choose, usually at a trigger moment; a funded account can pay back for many years once assets arrive.

The customer is a person opening a first account, adding a second platform for a different purpose, or moving assets from a provider they no longer like. Common triggers are a first salary, retirement-account season, a new product such as a retirement match, or frustration with an existing app. Platforms aim at a broad range of people: Wealthfront, which set terms for its stock market listing in December 2025, said in its [prospectus](https://www.iposcoop.com/the-ipo-buzz-wealthfront-wlth-proposed-sets-terms-for-450-million-ipo/) that its average individual funded client is 38 and that its funded clients earn about $165,000 a year.

The money comes after funding, not at sign-up:

| Platform | Latest public figures | What drives revenue |
|---|---|---|
| Robinhood (Q2 2026) | 28.4 million funded customers; average revenue per user of $187; 4.8 million Gold subscribers; net deposits of $21.7 billion | Trading, interest on balances, subscriptions, and new products such as its robo-advisor |
| Wealthfront (July 2025) | Over 1.3 million funded clients; $88.2 billion in platform assets | Advisory fees and cash management on growing balances |

Two details show where the value sits. About 40% of Robinhood’s new funded customers signed up for its Gold subscription in the quarter, and its automated investing product, Robinhood Strategies, grew to over 300 thousand funded customers. Robinhood also counts account transfer and retirement match incentives in its net deposits, a sign of how hard platforms compete for assets that are moving. We infer that the most valuable AI answer is not the one that produces an app download but the one that sends someone who will fund and stay.

## How much do investors already rely on AI?

Heavily for research, but with caution: most verify what AI tells them before acting.

- **Broad use.** In the Investing.com survey, 62% had used AI tools for investment decisions: 24% regularly, 27% occasionally and 11% once or twice.
- **Chatbots lead.** 54% had used chatbots such as ChatGPT for investing research. In an earlier [Motley Fool survey](https://www.fool.com/research/survey-how-investors-are-using-generative-ai) of 2,000 investors who use generative AI, conducted in September 2024, 66% said ChatGPT had helped them make investment decisions.
- **Verification is the norm.** 54% of Investing.com’s respondents said they trust AI analysis only somewhat and usually check it against other sources, and 39% worried about incorrect or misleading recommendations.

Regulators push the same habit. A joint [investor alert](https://www.investor.gov/introduction-investing/general-resources/news-alerts/alerts-bulletins/investor-alerts/artificial-intelligence-fraud) from the SEC’s investor education office, NASAA and FINRA warns that AI-generated information “might rely on data that is inaccurate, incomplete, or misleading,” and reminds investors that investment platforms generally must be registered. For a platform, that verification step is where review sites, registration records and your own disclosures do their work.

These surveys measure research about investments, not the choice of platform. No public survey we found measures how many people ask AI which brokerage or robo-advisor to open, so the scale of that specific behavior is unknown.

## What do people ask AI when choosing a brokerage or robo-advisor?

Questions about fit for a purpose, fees, safety and switching. The investor prompts in this table are our own examples, not questions collected from real users.

| Need | Illustrative prompt |
|---|---|
| First account | “Best investing app for a beginner with $500 to start” |
| Head-to-head | “Robinhood vs Fidelity for a Roth IRA” |
| Automated investing | “Cheapest robo-advisor with tax-loss harvesting” |
| Active trading | “Which broker has the lowest options fees and good charts?” |
| Safety | “Is my money protected if this investing app goes bankrupt?” |
| Switching | “How do I move my brokerage account, and which broker pays a transfer bonus?” |

Each of these needs is a separate question, and we expect each to produce its own shortlist: the app suited to a first $500 account may not be the one an active options trader is pointed to. A platform therefore has many AI reputations, one for each kind of question. Which funds people pick inside an account is a separate contest, covered in [how asset managers get funds considered](https://underneath.agency/resources/asset-managers-fund-demand-ai-search).

## Which platforms do AI answers name, and what decides it?

Mostly brokers that editorial comparison sites rate well; assistants publish little about how they choose.

**Documented by the platform.** Google says AI Overviews and AI Mode may use [“query fan-out”](https://developers.google.com/search/docs/appearance/ai-features), running several related searches across subtopics before answering. For money topics, Google’s helpful-content guidance says its systems [lean harder on signs of experience, expertise, authoritativeness and trust](https://developers.google.com/search/docs/fundamentals/creating-helpful-content), because the stakes for people’s financial stability are higher. Neither Google nor OpenAI documents how it chooses between brokers.

**Observed by others and by us.**

- **Awards are evidence, not a guarantee.** SoFi Invest ranked first among do-it-yourself platforms in [J.D. Power’s 2026 study](https://s27.q4cdn.com/749715820/files/doc_news/SoFi-Ranks-1-in-JD-Power-2026-U-S--Investor-Satisfaction-Study-for-DIY-Investors-2026.pdf), which surveyed 7,982 advised and 4,335 do-it-yourself investors. We found no public data showing whether an award like this changes how often assistants name a platform, so treat it as one more reviewer fact, not a shortcut.
- **Comparison sites pass into answers.** In [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study), when ChatGPT’s own search named a source such as NerdWallet, the answer cited it 44.0% of the time, against 8.1% when it did not. In [our citation study](https://underneath.agency/research/ai-overview-citations-study), nerdwallet.com appeared in 10.0% of all AI Overviews in the sample.
- **Each assistant has its own list.** In our agreement study, two-thirds of recommended options were named by only one of four assistants, so tracking one assistant gives a partial picture.

**Our inference.** Assistants appear to borrow the “best for” labels that reviewers assign: best for beginners, best for active traders, best for retirement. A platform that wants to be named for a use case needs that label in the reviews and comparisons assistants read, backed by facts on its own pages that the reviewers can check. Fees, minimums, account types, fractional shares, protections and the regulated entities behind the product are the facts reviewers line up in their comparison tables, so they are the facts most likely to be repeated.

## What path leads from an AI answer to a funded brokerage account?

Through four steps: named in the answer, confirmed on review sites, account opened, then assets funded or transferred.

1. **Named for a need.** The assistant lists two to four platforms for the person’s stated purpose.
2. **Checked.** The person compares fees and features, reads reviews, and checks safety and registration.
3. **Opened.** They download the app or open the account, often the same day.
4. **Funded.** They deposit cash or transfer an existing account. This is where revenue starts, and where premium tiers and other products follow.

Each step can fail on an out-of-date fact. If an assistant repeats an old fee, a retired promotion or a missing account type, the person may drop out before opening, and the platform never learns why. We suggest measuring AI visibility by use case and connecting it to funded accounts and net deposits, not downloads, and asking new customers where they first heard of you. Analytics often undercount AI-assisted visits, as covered in [why analytics miss AI visibility](https://underneath.agency/resources/why-analytics-miss-ai-visibility).

## Which marketing rules apply to what assistants repeat?

The same rules that govern your advertising. A fact an assistant repeats is only as safe as the page it came from.

Robo-advisers are usually registered investment advisers, and the SEC’s [marketing rule](https://www.ecfr.gov/api/renderer/v1/content/enhanced/current/title-17?part=275&section=275.206%284%29-1) bars advertisements that include an untrue statement of a material fact. It sets conditions on testimonials and endorsements, and it permits third-party ratings only with specific disclosures, such as whether the adviser paid in connection with obtaining or using the rating. Brokerages answer to FINRA’s communications rules, and many platforms run both kinds of entity. SoFi’s release, for example, names an SEC-registered investment adviser for its robo investing and a FINRA and SIPC member broker for self-directed trading.

Two practical points follow, which your compliance team should confirm. First, fee tables, promotions and award badges on your site are the facts assistants and reviewers pick up, so they need the same review and dating as any advertisement. Second, rankings you pay to appear in or license, such as award logos, carry disclosure duties; the J.D. Power award notice in SoFi’s release says use of award materials is subject to a license fee. Do not seed reviews or testimonials to influence AI answers; the rules on endorsements apply however the content is found. Lenders face a similar duty with rate claims, covered in [how mortgage platforms win borrowers through AI](https://underneath.agency/resources/mortgage-platforms-leads-ai-search).

## What does GEO look like for an investment platform?

Generative engine optimization (GEO) makes the use cases you serve, and their facts, easy for assistants to find and verify.

For a brokerage, investing app or robo-advisor, the work usually covers:

1. **Use-case pages.** One clear page for each audience you serve well (first-time investors, retirement savers, active traders, automated investing), stating what you offer them and what you do not.
2. **Fee and feature facts in text.** Commissions, account minimums, advisory fees, cash rates with dates, account types and transfer process, written so a reviewer or assistant can quote them correctly.
3. **Safety and entity facts.** Which registered entity provides each service, membership details, how customer assets are held, and what protection does and does not cover, written with compliance.
4. **Reviewer relationships.** Accurate data and product access for the editorial review sites assistants cite, plus prompt corrections when their tables go stale.
5. **Explainers people watch.** Short, accurate videos on account types and transfers; in [our YouTube study](https://underneath.agency/research/ai-overview-youtube-videos-study), AI Overviews for financial services searches cited a YouTube video on 49.0% of searches.
6. **Monitoring across assistants.** Track your use-case questions in several assistants and Google’s AI surfaces, since answers differ between them.

If an assistant quotes a stale commission or names the wrong custodian, [our guide to correcting wrong brand facts in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) walks through the repair. Smaller platforms facing incumbents can learn from [how a fintech startup gets recommended by AI when established brands own the answers](https://underneath.agency/resources/fintech-startups-customers-ai-search), and consumer money apps outside investing are covered in [how consumer fintech apps win customers when people ask AI about money](https://underneath.agency/resources/consumer-fintech-apps-customers-ai-search). Savings accounts are covered in [how digital banks win depositors through AI](https://underneath.agency/resources/digital-banks-depositors-ai-search). No one can guarantee that an assistant will name a platform; GEO makes the facts it finds accurate, current and confirmed by independent sources.

## Where does the evidence on investment platforms run out?

At the conversion step: we can see who gets named, but not how many accounts those answers open or fund.

- **Investor surveys measure research, not platform choice.** The Investing.com and Motley Fool surveys ask about investment decisions; neither measures how people pick a brokerage. Both come from publishers with a stake in investing content.
- **Our studies are snapshots.** They cover many industries at one point in time, not investing alone, and answers change between assistants and over time.
- **No public link to funded accounts.** We found no public data connecting AI visibility to account openings or net deposits for any platform. Our cross-industry review, [does AI visibility drive business results?](https://underneath.agency/resources/does-ai-visibility-drive-business-results), covers what little is known beyond investing.
- **Assistants inside brokerages are new.** Several platforms now offer their own AI assistants. How those change the choice of platform, rather than investing within one, is not yet measured.

## Where should an investment platform start?

Start by asking assistants the platform-choice questions for each use case you serve, then check every fee and safety fact.

A useful first review covers the “best app for” questions for your target customers, head-to-head comparisons with the platforms you lose to, and safety and “is it legit” questions about your brand. It shows where you are named, which reviewers the answers rely on, and which of your facts are missing, wrong or out of date.

If your growth depends on funded accounts and net deposits, [ask us to review how AI answers present your platform](https://underneath.agency/contact). We will compare your visibility with your closest competitors by use case, list the facts assistants and reviewers get wrong, and plan the pages, reviewer outreach and monitoring, checked by your compliance team, that give your platform a fair hearing. The [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page shows how use-case pages, fee and safety facts, and reviewer outreach run as one program for brokerages and robo-advisors.

## Frequently asked questions

### Do people really ask ChatGPT which investing app to use?

Many use AI for investing research: 54% of Investing.com’s respondents had used chatbots for it. How many ask specifically which platform to open is not publicly measured.

### Why do AI answers about brokers cite review sites instead of broker websites?

Assistants search for comparisons and reuse what they find: in our hidden-searches study, a source that ChatGPT’s own search had named was cited 44.0% of the time. That favors the brokers the comparison sites feature, which is why accurate reviewer data matters.

### Can a platform pay to be recommended in AI answers?

We know of no documented way to buy a place in an assistant’s recommendations; ads, where offered, are a separate question covered in [are ads coming to ChatGPT and other AI assistants?](https://underneath.agency/resources/are-ads-coming-to-ai-assistants). Paid rankings and licensed awards carry disclosure duties.

### Should a robo-advisor publish its fees in plain text?

Yes. Assistants and reviewers quote fees and minimums constantly. Clear, dated fee pages reduce the chance an answer repeats an old or wrong figure.

## Sources

- Investing.com, via Hedge Fund Alpha (2026-04-09), [Nearly Two-Thirds Of Retail Investors Now Use AI To Inform Investment Decisions](https://hedgefundalpha.com/news/retail-investors-now-use-ai-to-inform-investment-decisions/)
- The Motley Fool (2024), [Survey: How Investors Are Using Generative AI](https://www.fool.com/research/survey-how-investors-are-using-generative-ai)
- Robinhood Markets, via Mondo Visione (2026-07-29), [Robinhood Reports Second Quarter 2026 Results](https://mondovisione.com/news/robinhood-reports-second-quarter-2026-results-2026729/)
- IPO Scoop (2025-12-02), [The IPO Buzz: Wealthfront Sets Terms for $450 Million IPO](https://www.iposcoop.com/the-ipo-buzz-wealthfront-wlth-proposed-sets-terms-for-450-million-ipo/)
- Crowdfund Insider (2025-10), [Wealthfront Submits Registration Statement for Proposed IPO](https://www.crowdfundinsider.com/2025/10/253865-wealthfront-submits-registration-statement-for-proposed-ipo/)
- SoFi Technologies (2026-03-18), [SoFi Ranks #1 in J.D. Power 2026 U.S. Investor Satisfaction Study for DIY Investors](https://s27.q4cdn.com/749715820/files/doc_news/SoFi-Ranks-1-in-JD-Power-2026-U-S--Investor-Satisfaction-Study-for-DIY-Investors-2026.pdf)
- SEC, NASAA and FINRA, via Investor.gov (2024-01-25), [Artificial Intelligence (AI) and Investment Fraud: Investor Alert](https://www.investor.gov/introduction-investing/general-resources/news-alerts/alerts-bulletins/investor-alerts/artificial-intelligence-fraud)
- eCFR (2026), [17 CFR 275.206(4)-1, Investment adviser marketing](https://www.ecfr.gov/api/renderer/v1/content/enhanced/current/title-17?part=275&section=275.206%284%29-1)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- Google Search Central (2025), [Creating helpful, reliable, people-first content](https://developers.google.com/search/docs/fundamentals/creating-helpful-content)
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude agree on brands?](https://underneath.agency/research/ai-assistants-brand-agreement-study)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [Do AI Overviews cite the pages that rank?](https://underneath.agency/research/ai-overview-citations-study)
- Underneath (2026), [The YouTube videos Google’s AI cites are small](https://underneath.agency/research/ai-overview-youtube-videos-study)

---

This is the Markdown twin of https://underneath.agency/resources/investment-platforms-customers-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Is a GEO Agency Worth It? A Decision with Numbers | Underneath"
description: "Is a GEO agency worth it? When it pays back, when it does not, what it costs in 2026, and how the AI engines themselves answer the question."
canonical: "https://underneath.agency/resources/is-a-geo-agency-worth-it"
published: 2026-09-25
updated: 2026-09-27
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Is a GEO agency worth it?

A generative engine optimization (GEO) agency is worth it when your buyers already ask AI assistants for recommendations, when you have organic revenue to protect, and when you can sustain the work for months rather than buy it as a one-off. It is not worth it as a substitute for SEO, or for a business whose customers do not research with AI yet. This guide shows how to decide with three numbers, and what the engines themselves said when we asked them.

Underneath teamGuide

On this page

1. [What the AI engines answer](#what-the-ai-engines-answer)
2. [When a GEO agency is worth it](#when-a-geo-agency-is-worth-it)
3. [When it is not](#when-it-is-not)
4. [The cost side](#the-cost-side)
5. [The return side](#the-return-side)
6. [Agency, in-house or wait](#agency-in-house-or-wait)
7. [Frequently asked questions](#frequently-asked-questions)

Related service

Generative Engine Optimization

Get mentioned, cited and recommended by ChatGPT, Gemini, Perplexity and Copilot, and described correctly when you are.

[How we do it](https://underneath.agency/services/generative-engine-optimization)

## The short version

1. Worth it if buyers in your category ask AI assistants for recommendations and the assistants currently name competitors; not worth it if they do not yet, or if the agency treats AI visibility as separate from search.
2. Decide with three numbers: how often your buyer questions produce an AI answer that names a brand, how often that brand is you, and what a named recommendation is worth against a retainer, which published 2026 price guides put between about $1,500 and $50,000 a month.
3. The cheapest way to find out is a repeated audit of your own questions across the engines, which any credible agency will run before selling you a program.

“Is a GEO agency worth it” is one of the few marketing questions the AI engines answer with a hedge, and the hedge is informative. Google’s AI Overview opens with “usually not worth it if they treat AI optimization as a separate fad from traditional search”. ChatGPT says “sometimes, yes, but most businesses do not need a full GEO agency yet”. Both are right, and both point to the same test: is AI search already where your buyers make up their minds?

## What the AI engines answer

We put the question to ChatGPT, Google AI Overviews, Google AI Mode, Gemini and Perplexity 22 times between 15 and 25 September 2026.

- Every platform structured its answer as “when it is worth it” and “when it is not”, and every one made the verdict conditional on how much the audience uses AI search and on whether the agency builds on SEO or sells against it.
- Google AI Mode cited a Reddit thread in 22 of 22 runs, and Reddit was named inside 56 of the 110 answers, more than for any other GEO question tested. Buyers asking whether to pay for this are reading what other buyers say.
- ChatGPT’s most-cited source was academic research (14 of 22 runs), and it offered a rule of thumb: if organic-driven revenue is under about $20,000 a month, build the foundations before hiring a GEO agency.
- The People Also Ask box beneath the Google result was almost entirely about SEO anxiety: “is SEO dead now with AI?” (18 times), “can ChatGPT do SEO?” (12), “is GEO replacing SEO?” (12), “is AEO worth it?” (10). The real worry is whether money already spent on search has been wasted. It has not; the questions at the end of this guide explain why.

## When a GEO agency is worth it

- Your buyers ask assistants. Search your ten most important buyer questions on ChatGPT, Perplexity and Google. If the answers name brands, a recommendation is being given to someone. If a competitor is named and you are not, the cost of doing nothing is already being paid.
- You have organic revenue to protect. Ranking pages are the ones Google’s AI Overview is most likely to cite: across 486 searches, [41.7% of pages ranking 1 to 3 were cited](https://underneath.agency/research/ai-overview-cited-pages-study), against 20.1% at positions 7 to 10. A page that ranks but is not cited risks losing the click to the answer above it, and GEO work on that page has a return you can measure.
- Your category’s answers draw on a small set of pages. For the four GEO agency questions we tracked, a handful of pages were cited in 18 to 22 of 22 runs on Perplexity and Google. Where the source set is small and stable, entering it is worth more than where it churns.
- You can sustain it. The pages the engines cite are rarely thin: Perplexity’s had a median of about 3,700 words, against about 1,500 for pages it retrieved and skipped. Pages of that depth, and the third-party mentions around them, take months to build; a one-off project rarely gets there.

## When it is not

- Your customers do not yet research with AI. Some local and trade categories still convert through the map pack and referrals; the audit will show AI answers that name nobody, or that name marketplaces. Spend on the foundations first.
- The agency sells GEO against SEO. Google’s own guidance says its generative features rest on the same fundamentals as search. An agency that proposes to skip ranking and “optimize for AI” is proposing to skip the main route into Google’s AI surfaces.
- Your organic revenue is small. If it is, a $5,000-a-month retainer is a large share of it, and the foundational work (a crawlable site, clear pages, a consistent business profile) is cheaper to do in-house or with an SEO partner.
- You want a guarantee. No agency controls the engines, and their sources move: for one of the queries we tracked, Perplexity’s sources changed completely from one day to the next. An agency promising a place in ChatGPT’s answer is promising what it cannot deliver.

## The cost side

Published 2026 price guides range from about $1,500 to $50,000 a month. [WebFX](https://www.webfx.com/blog/ai/generative-engine-optimization-cost/) (May 2026) splits the range by company size: $1,500 to $5,000 a month for small businesses, $5,000 to $25,000+ for mid-sized companies and $25,000 to $50,000+ for enterprises, with one-off projects at $5,000 to $50,000 and consulting at $50 to $300 an hour. Underneath’s programs are priced in three bands, $5,000 to $10,000, $11,000 to $20,000 and $20,000+ a month, and start with a Marketing & AI Visibility Audit whose fee is credited if you continue.

The retainer is not the whole cost. Someone on your side has to approve page changes, supply proof for the third-party outreach and read the monthly report; budget two to four hours a week for it.

## The return side

The return is the value of being the named recommendation for the questions your buyers ask, multiplied by how often those questions are asked of an assistant, minus what you already earn from the same questions in ordinary search. Three numbers make it concrete.

| Number | How to get it | What it tells you |
| --- | --- | --- |
| Answer rate | Run your ten buyer questions on each engine over several days; count how often the answer names a brand | Whether a recommendation exists to be won |
| Your share | Count how often that brand is you | The gap between where you are and where the winner is |
| Value of a recommendation | Your close rate and average deal or basket value for inquiries that arrive already shortlisted | What each point of share is worth per month |

If the answer rate is high, your share is low and a recommendation is worth more than the retainer within the sales cycle, an agency is worth it. If the answer rate is low, wait and re-measure in a quarter; the engines change quickly, and the audit is cheap to repeat.

## Agency, in-house or wait

- Hire an agency when the audit shows competitors being named and nobody on your team has done this work before. Start with the audit, not a contract.
- Build in-house when you already have the tracker and the writing standards and need volume; Underneath hands these over at the end of every engagement.
- Wait when the audit shows AI answers that name nobody in your category. Fix the foundations and re-run the audit in a quarter.

## Frequently asked questions

### Is a GEO agency worth it?

Yes, when your buyers already ask AI assistants for recommendations, competitors are being named and you are not, and you can sustain the work for months. No, when your customers do not yet research with AI, when the agency sells GEO against SEO, or when your organic revenue is too small to carry a retainer of several thousand dollars a month.

### How much does a GEO agency cost?

Between about $1,500 and $50,000 a month, according to published 2026 price guides; [WebFX](https://www.webfx.com/blog/ai/generative-engine-optimization-cost/) puts mid-sized companies at $5,000 to $25,000. Underneath’s programs are priced in three bands, $5,000 to $10,000, $11,000 to $20,000 and $20,000+ a month, after an audit whose fee is credited if you continue. Add a few hours a week of your own team’s time.

### Is SEO dead now with AI?

No. Google’s AI features are built on Google’s index, and ranking still raises the odds of a citation: in our study of 486 searches, the AI Overview cited 41.7% of top-three pages against 20.1% at positions 7 to 10. SEO gets you into Google’s AI surfaces; GEO works on whether the ranking page is the one quoted, and on the engines where ranking counts for less.

### Is GEO replacing SEO?

No. They are complementary: SEO earns a position in the results, GEO earns a place in the answer, and the two should be measured together. Google’s guidance says the fundamentals are shared, and the AI answers to this very question warned against agencies that sell one against the other.

### Is AEO worth it?

Answer engine optimization, the work of being the extracted answer in featured snippets and AI Overviews, is part of the same program. On Google’s surfaces it and GEO overlap almost completely; you do not need to buy them separately.

### Can ChatGPT do SEO?

ChatGPT can draft content and suggest structure, but it cannot rank a page, earn a citation or fix crawler access. In Underneath’s tests ChatGPT cited academic research and official documentation for this question and rarely an agency; it is a source of advice, not of visibility.

## Keep *reading.*

[All resources](https://underneath.agency/resources)

- Guide · AI search

  ### [What does a GEO agency do?](https://underneath.agency/resources/what-does-a-geo-agency-do)

  The nine things a generative engine optimization agency does, what it costs, and how the AI engines describe the job.

  Underneath team
- Guide · AI search

  ### [How to choose a GEO agency](https://underneath.agency/resources/how-to-choose-a-geo-agency)

  Eight criteria, the red flags, and the questions to ask on the first call, checked against what the AI engines themselves tell buyers.

  Underneath team
- Guide · AI search

  ### [GEO agency vs SEO agency](https://underneath.agency/resources/geo-agency-vs-seo-agency)

  What each one does, where the work overlaps, and whether you need one, the other or both.

  Underneath team

Free strategy call

## Get the three numbers for your business.

On a free 30-minute call we run your ten most important buyer questions through the engines and tell you the answer rate, your share, and who is being named instead.

[Book a strategy call](https://underneath.agency/contact)[How we work](https://underneath.agency/approach)

---

This is the Markdown twin of https://underneath.agency/resources/is-a-geo-agency-worth-it. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Is ChatGPT replacing Google search? | Underneath"
description: "Not on current evidence. People use ChatGPT alongside Google: search appeared in 46.0% of assistant sessions in a 2026 US and UK browsing panel."
canonical: "https://underneath.agency/resources/is-chatgpt-replacing-google"
published: 2026-10-08
updated: 2026-10-10
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Is ChatGPT replacing Google search?

Not on current evidence. The best behavioral data shows people using AI assistants such as ChatGPT alongside search engines rather than instead of them: in a 2026 browsing panel, search still appeared in nearly half the sessions that included an assistant. But assistant sessions more often end without visiting any website, and no study can yet say whether AI is shrinking total search demand.

## The short version

1. A search engine appeared in 46.0% of browsing sessions that included an AI assistant, in a February 2026 panel of US and British users (Iannelli and Ai).
2. 34.1% of assistant sessions included no visit to any outside website, against 19.5% of search sessions by the same people.
3. People tend to search first and then open pages, but browse first and then turn to an assistant: a 20.6-point reversal in direction between the two.
4. ChatGPT runs its own web searches, a mean of 3.7 per answer in our study, yet only 8.3% of the pages it cited ranked in Google’s top 10 for the question.
5. Google’s own AI changes search more directly: AI Overviews appeared on 36% of searches in a 2026 US experiment, while people chose AI Mode for only 0.6%.

## What does the best evidence on behavior show?

People use AI assistants alongside search, not instead of it, in the one panel that tracks both.

[Iannelli and Ai](https://arxiv.org/abs/2607.04282) studied an opt-in panel of US and British users in February 2026. The panel’s software recorded conversations with ChatGPT, Gemini and Perplexity on the web, together with the same people’s searches and page visits. The authors then grouped each person’s activity into sessions and looked at where searching and browsing fell around assistant use.

A search engine appeared somewhere in 46.0% of sessions that included an assistant. It came before the first assistant prompt in 33.8% of them. The authors’ conclusion is careful: “containment and cross-surface coexistence occur at the same time.” Some assistant sessions stand alone, and many sit inside ordinary searching and browsing.

Read this study with its limits in mind. The authors work at a company that sells AI-visibility software, the paper was not yet peer reviewed, and they do not disclose the panel’s size. Use of assistants through phone apps was not observed. The main results did repeat in a second month, March 2026.

## How is using an assistant different from searching?

Assistant sessions more often stay self-contained, and they tend to come later in the journey.

The same people behaved differently around the two tools. The table compares their sessions that included an assistant with their sessions built around search alone.

| Measure (February 2026) | Sessions with an AI assistant | Search sessions, same people |
|---|---|---|
| No visit to any outside website | 34.1% | 19.5% |
| Typical session length | 22 minutes | 9 minutes |

Within the same person, the gap in self-contained sessions was 13.0 percentage points. In March 2026 it was 35.3% for assistant sessions, close to February’s figure. Our article on [how often assistant users skip websites](https://underneath.agency/resources/ai-assistant-sessions-without-website-visits) looks at this gap in detail.

Direction differed too. Around search, page visits mostly came after: they followed the last search in 60.2% of sessions and preceded the first in 46.7%. Around assistants this flipped: page visits came before the first prompt in 45.5% of sessions and after the last in 38.9%. In the authors’ words, search tends to open the journey, while assistants “sit deeper inside it”.

“Self-contained” does not mean the question was answered. The authors stress that the data cannot show whether a need was met. The levels also vary widely: self-contained assistant sessions were 50.9% in the United States and 20.7% in Great Britain.

## Is AI shrinking the number of Google searches?

No study we reviewed answers this cleanly; the evidence is mixed and indirect.

Iannelli and Ai deliberately make no claim about search volume. They say their design cannot show whether AI creates or destroys search demand, because people turn to assistants at moments when they are already busy online.

Other evidence points both ways. [Zhang, Cui and Zhang](https://arxiv.org/abs/2605.16428) summarize one study finding that adopting AI assistants cut total searches by roughly 20%, and another finding that it increased the number of websites people visit. We could not review either directly.

Single-site data shows how hard this is to read. [Watanabe and Nakayashiki](https://arxiv.org/abs/2606.04362), who work for the company that owns the site they studied, found that ChatGPT referrals to pages they had not changed grew 3.5 times between January and May 2026. Google organic clicks across the site fell about 20% from late 2025 into 2026. One site cannot show what caused either change.

## Does ChatGPT still depend on search engines?

Yes: ChatGPT runs its own web searches behind many answers, though it cites different pages from Google.

In [our study of hidden searches](https://underneath.agency/research/ai-hidden-searches-study), ChatGPT ran a mean of 3.7 searches per answer, Gemini 1.9 and Claude 0.76. In other words, the assistant often does the searching on the user’s behalf.

The results do not mirror Google. In [our study of AI citations and Google rankings](https://underneath.agency/research/ai-citations-google-rankings-study), only 8.3% of the pages ChatGPT cited ranked in Google’s top 10 for the question. 68.5% were not in Google’s top 100 for the question or for any of ChatGPT’s own searches.

[Huang and colleagues](https://arxiv.org/abs/2603.16138) found the same gap at the level of whole websites. OpenAI’s search-enabled GPT model cited a median of 3 sources per answer. Its 100 most-cited domains overlapped only 24% with those of Google’s AI Overviews and 25% with Google’s normal results.

## Is Google’s own AI changing search more than ChatGPT is?

For most businesses, probably yes, because Google’s AI answers sit inside the search people already use.

In the field experiment by [Wang and colleagues](https://arxiv.org/abs/2608.18352), 1,100 US volunteers saw AI Overviews on 36% of their searches before the study began. They chose Google’s chat-style AI Mode for only 0.6% of searches. On a fixed set of benchmark searches, [Aral, Li and Zuo](https://arxiv.org/abs/2602.13415) found 67% of US searches answered by AI in 2025, against 42% in 2024.

People resisted being pushed further. When every search was sent to AI Mode for a week, the share of people who tried Bing, DuckDuckGo or Yahoo rose 11.2 percentage points, and clicks to outside websites fell 18.8 points. Our article on [AI Mode as Google’s default](https://underneath.agency/resources/ai-mode-default-traffic-loss) covers that result.

## What should you do about it?

Plan for both channels, because the evidence shows people using them together rather than swapping one for the other.

1. Do not move your whole search budget to AI. Search still tends to open the journey toward websites.
2. Treat AI assistants as [a later step in the customer journey](https://underneath.agency/resources/chatgpt-vs-google-customer-journey), where people compare and decide. Content that helps a decision matters there.
3. Do not assume your Google rankings carry over. Only 8.3% of ChatGPT’s citations ranked in Google’s top 10 for the question in our study.
4. Measure AI referrals separately, and compare them with pages you did not change. One site’s untouched pages saw ChatGPT referrals grow 3.5 times in five months.
5. Watch Google’s own AI features, which reach far more searches than people choose to send to chat tools.

If you want help planning for both, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The research cannot yet say whether AI assistants reduce, add to or reshape total search demand.

- No causal study of search volume was available to us; the behavior panel describes patterns, not causes.
- That panel was run by an AI-visibility software company, is not yet peer reviewed, covers two months in two countries and misses phone apps.
- Nobody knows whether self-contained assistant sessions actually solved the user’s need.
- Differences between business buyers and consumers are not measured.
- Single-site referral data cannot be generalized to other websites.

## Frequently asked questions

### Are people using ChatGPT instead of Google?

Mostly alongside it. In a February 2026 US and UK panel, a search engine appeared in 46.0% of browsing sessions that included an AI assistant.

### Does ChatGPT use Google search results?

ChatGPT runs its own web searches, a mean of 3.7 per answer in our study, but only 8.3% of the pages it cited ranked in Google’s top 10 for the question.

### Do AI assistants send people to websites?

Less often than search does. 34.1% of assistant sessions in the 2026 panel included no visit to any outside website, against 19.5% of search sessions by the same people.

### Is Google losing users to AI chatbots?

The evidence does not show it yet. The panel study found people using both, and when a 2026 experiment forced Google’s AI Mode on people, 11.2 percentage points more of them tried another search engine.

## Sources

- Iannelli and Ai (2026), [The New Shape of Search: How Conversational AI Recomposes Information Seeking](https://arxiv.org/abs/2607.04282), arXiv:2607.04282.
- Zhang, Cui and Zhang (2026), [The Impact of AI Search on the Online Content Ecosystem: Evidence from Google and Reddit](https://arxiv.org/abs/2605.16428), arXiv:2605.16428.
- Watanabe and Nakayashiki (2026), [Disentangling Answer Engine Optimization from Platform Growth: A Log-Based Natural Experiment on ChatGPT Referral Traffic](https://arxiv.org/abs/2606.04362), arXiv:2606.04362.
- Huang and colleagues (2026), [Answer Bubbles: Information Exposure in AI-Mediated Search](https://arxiv.org/abs/2603.16138), arXiv:2603.16138.
- Wang, Gleason, Bart, Wilson and Metaxa (2026), [AI in Search Reduces Publisher Referrals Without Improving User Experience: Experimental Evidence](https://arxiv.org/abs/2608.18352), arXiv:2608.18352.
- Aral, Li and Zuo (2026), [The Rise of AI Search: Implications for Information Markets and Human Judgement at Scale](https://arxiv.org/abs/2602.13415), arXiv:2602.13415.
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude cite pages that rank?](https://underneath.agency/research/ai-citations-google-rankings-study)

---

This is the Markdown twin of https://underneath.agency/resources/is-chatgpt-replacing-google. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Is traditional SEO still important for visibility in AI search?"
description: "Yes. Pages that rank higher are cited more often, but Google rankings explain only part of AI citations, and less for ChatGPT than for Google’s own AI."
canonical: "https://underneath.agency/resources/is-seo-still-important-for-ai-search"
published: 2026-10-11
updated: 2026-10-11
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Is traditional SEO still important for visibility in AI search?

Yes. Search ranking is still the main door into AI answers: pages that rank higher are cited more often, and pages no engine retrieves are never cited. But rankings explain only part of what AI engines cite, and much less for ChatGPT than for Google’s own AI features.

## The short version

1. In Google’s AI Overviews, the first organic result was cited in 49.5% of answers and the ninth in 15.5%, in [our AI Overview citation study](https://underneath.agency/research/ai-overview-citations-study).
2. In a simulated test, moving a page to the top of what the AI reads beat every content rewrite tested, [Puerto and colleagues found](https://arxiv.org/abs/2506.11097).
3. Only 8.3% of the pages ChatGPT cited ranked in Google’s top 10 for the question, against 25.7% for Claude, in [our rankings study](https://underneath.agency/research/ai-citations-google-rankings-study).
4. ChatGPT ran a mean of 3.7 of its own searches per answer, so ranking for the buyer’s exact question is not enough, [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study) found.

## Does ranking on Google still matter for AI answers?

Yes, especially for Google’s own AI features. AI engines first search, then choose what to cite from what they found. A page that is not retrieved cannot be cited.

In [our AI Overview citation study](https://underneath.agency/research/ai-overview-citations-study) of 4,051 citations, 28.7% were page-one organic results for the same search. And 78.8% of AI Overviews cited at least one page-one result. Our guide on [whether Google rankings get you into AI Overviews](https://underneath.agency/resources/does-ranking-on-google-get-you-into-ai-overviews) breaks this down.

[Zhang and colleagues](https://arxiv.org/abs/2512.09483) saw the same pull in 2025 across six AI engines. Among domains that both AI and traditional engines returned, the most common was the domain ranked first in traditional search: 23.27% in Bing and 14.53% in Google. Those shared domains were a minority of everything the AI engines cited.

## How much does a page’s position matter?

A great deal: higher-ranking pages are cited far more often. In our AI Overview study, the first organic result was cited in 49.5% of AI Overviews, the third in 33.5%, the sixth in 23.2% and the ninth in 15.5%.

A second study of 3,096 top-10 pages confirmed it. In [our cited-pages study](https://underneath.agency/research/ai-overview-cited-pages-study), 41.7% of pages ranking 1 to 3 were cited, against 20.1% at positions 7 to 10. Position explained more than the 14 page features we tested, combined.

Controlled tests agree. [Puerto and colleagues](https://arxiv.org/abs/2506.11097) built a simulated AI search engine and tried published content rewrites meant to win citations. Most did little: 61.0% of retail product rankings were unchanged after one rewrite method. Placing a page first among the documents the AI read produced far larger gains than any rewrite.

That test simulated ranking by reordering documents, not by real SEO work. It still shows how much order matters once a page is retrieved.

A small 2024 study by [Narayanan Venkit and colleagues](https://arxiv.org/abs/2410.22349) saw the same pattern from the user’s side. Across sessions with 21 participants, people using Google often explored beyond the top five results. The answer engines they tested drew mainly on the top two or three.

## Why doesn’t a top ranking guarantee a citation?

Because AI engines run their own searches and choose differently from Google’s ranking. In [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study), ChatGPT ran a mean of 3.7 searches per answer. None of the 509 searches we saw repeated the user’s question word for word.

Many cited pages are outside Google’s rankings entirely. In our rankings study, 68.5% of ChatGPT’s citations were not in Google’s top 100 for the question or for any of the searches we tracked. For Claude the figure was 36.8%.

Google’s AI Overviews also look past page one. [Xu and colleagues](https://arxiv.org/abs/2605.14021) tracked 7,583 AI Overviews in spring 2026. Only 31.5% of cited domains overlapped with the top five results, rising to 70.2% across the full first page. The other 29.8% did not appear on the first page at all. We compare the overlap across studies in [whether AI engines cite the same sites as Google](https://underneath.agency/resources/do-ai-search-engines-cite-the-same-sites-as-google).

## Does SEO matter equally for every AI engine?

No. Google rankings carry over best to Google’s own AI and least to ChatGPT. In our rankings study, 8.3% of pages ChatGPT cited ranked in Google’s top 10 for the question, against 25.7% for Claude.

Even Google’s two AI products differ. In [our AI Mode comparison](https://underneath.agency/research/ai-mode-vs-ai-overviews-study), AI Overviews cited 29.7% of top-10 pages and AI Mode 16.3%, on the same 400 searches.

It varies by industry too. [Chen and colleagues](https://arxiv.org/abs/2509.08919) compared AI search with Google for local business questions in 2025. The share of sites both returned ranged from 20.6% for home cleaning to 2.5% for auto repair and 0.1% for IT support.

Engines without live search are harder still. [Sharma](https://arxiv.org/abs/2601.00912) tested 112 Product Hunt startups. A version of ChatGPT without web search surfaced them in only 3.32% of discovery questions. Perplexity, which searches the web, managed 8.29%, and its results tracked SEO signals such as the number of sites linking to a startup.

## Can content changes make up for a weak ranking?

Not on their own, according to the evidence so far. [Zhu and Chang](https://arxiv.org/abs/2603.20062) audited 14 Tokyo hotel websites against Gemini’s citations. Cited hotels scored 8.6 out of 15 on content depth, against 3.4 for hotels that were not cited.

The authors describe a two-stage process. A hotel site first needs content deep enough to rank in Google. Only then does it compete on how well it answers the question. The audit is small and exploratory, and the authors work at an AI company.

Another controlled test points the same way. [Vishwakarma and colleagues](https://arxiv.org/abs/2605.25517) ran 252,000 trials across six AI models. Topical relevance and list position were the biggest drivers of being cited first; formatting-only edits had little impact.

## What should you do about it?

Keep investing in SEO, then add the work that rankings alone do not cover.

1. Protect your rankings for the questions buyers ask. Higher positions are cited far more often.
2. Rank for the related searches AI engines run, not just the headline question, such as reviews, prices and comparisons.
3. Build links and mentions from other sites. They help ranking, and one study tied them to Perplexity visibility.
4. Check Google’s AI features and ChatGPT separately, since rankings carry over very differently.
5. Make each page answer its question clearly once it is retrieved.

For a combined approach, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

No study we reviewed tested whether improving a real page’s ranking changes its AI citations. The reverse question, [whether ChatGPT citations lift Google rankings](https://underneath.agency/resources/do-chatgpt-citations-boost-google-rankings), is also unproven.

- The strongest ranking evidence comes from simulations that reorder documents, not live SEO campaigns.
- Most live studies are observational snapshots on one or a few days.
- We cannot see the indexes ChatGPT and Perplexity use, so Google rankings are only a stand-in.
- Small samples, such as 14 hotels or 112 startups, limit how far some findings travel.
- Much remains unexplained. In our cited-pages study, 91.4% of the variation in which top-10 pages were cited was left unexplained by anything we measured.

## Frequently asked questions

### Do I still need SEO if I am optimizing for AI search?

Yes. Pages ranking 1 to 3 were cited in 41.7% of cases in our study of 3,096 top-10 pages, about twice the rate of positions 7 to 10.

### Does ranking first on Google get me into AI Overviews?

Often, not always. The first organic result was cited in 49.5% of AI Overviews in our study.

### Does ChatGPT use Google rankings?

Only loosely. Just 8.3% of ChatGPT’s cited pages ranked in Google’s top 10 for the question in our study of 80 buyer questions.

### Is GEO replacing SEO?

No. In a simulated test, a page’s position among the documents the AI read mattered more than any content rewrite tested.

## Sources

- Puerto et al. (2025), [C-SEO Bench: Does Conversational SEO Work?](https://arxiv.org/abs/2506.11097), arXiv:2506.11097.
- Zhang et al. (2025), [Source Coverage and Citation Bias in LLM-based vs. Traditional Search Engines](https://arxiv.org/abs/2512.09483), arXiv:2512.09483.
- Xu et al. (2026), [Measuring Google AI Overviews: Activation, Source Quality, Claim Fidelity, and Publisher Impact](https://arxiv.org/abs/2605.14021), arXiv:2605.14021.
- Zhu and Chang (2026), [The End of Rented Discovery: How AI Search Redistributes Power Between Hotels and Intermediaries](https://arxiv.org/abs/2603.20062), arXiv:2603.20062.
- Sharma (2026), [The Discovery Gap: How Product Hunt Startups Vanish in LLM Organic Discovery Queries](https://arxiv.org/abs/2601.00912), arXiv:2601.00912.
- Narayanan Venkit et al. (2024), [Search Engines in an AI Era: The False Promise of Factual and Verifiable Source-Cited Responses](https://arxiv.org/abs/2410.22349), arXiv:2410.22349.
- Chen et al. (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919), arXiv:2509.08919.
- Vishwakarma et al. (2026), [What Gets Cited: Competitive GEO in AI Answer Engines](https://arxiv.org/abs/2605.25517), arXiv:2605.25517.
- Underneath (2026), [Do AI Overviews cite the pages that rank? 4,051 citations](https://underneath.agency/research/ai-overview-citations-study)
- Underneath (2026), [What pages cited by AI Overviews have in common: 3,096 pages](https://underneath.agency/research/ai-overview-cited-pages-study)
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude cite pages that rank?](https://underneath.agency/research/ai-citations-google-rankings-study)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [AI Mode vs AI Overviews: how different are the sources?](https://underneath.agency/research/ai-mode-vs-ai-overviews-study)

---

This is the Markdown twin of https://underneath.agency/resources/is-seo-still-important-for-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Is Tracking ChatGPT Alone Enough to Measure AI Visibility?"
description: "No. AI engines cite largely different sources and pick different brands for the same question, so ChatGPT alone shows less than half the picture."
canonical: "https://underneath.agency/resources/is-tracking-chatgpt-enough"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Is tracking ChatGPT alone enough to measure our AI search visibility?

No: for the same question, AI engines cite mostly different sources and often recommend different brands, so one engine shows only part of your visibility. In one four-engine test, ChatGPT captured 42.6% of all the pages the engines cited, the most of any engine but still less than half. Tracking ChatGPT is a sensible start, not a measure of “AI search” as a whole.

## The short version

1. In a June 2026 test of 15 commercial prompts, no single engine captured half of the pages cited across ChatGPT, Google, Perplexity and Copilot; ChatGPT came closest at 42.6%.
2. In the same test, 96.4% of all cited pages appeared on only one engine.
3. On shopping questions asked from the Netherlands, ChatGPT and Gemini shared only 5.4% of the websites they displayed for the same question.
4. In our 80-question study, two assistants’ recommended brands overlapped by only about a third (0.327), and all four named the same first pick for just 10.0% of questions.
5. Tracking tools that query the developer version of ChatGPT see different sources from the consumer app: the two shared only 12.0% of displayed websites.

## How much of the AI source landscape does ChatGPT cover?

Less than half in the one study that measured it directly, and other engines each covered much less.

[Tannenbaum](https://arxiv.org/abs/2609.22655), founder of a company that sells AI visibility software, ran 15 US commercial prompts about that software category on four engines on 6 June 2026. For the ten prompts all four answered, he pooled every page cited by any engine. ChatGPT cited 42.6% of that pool, Google 24.3%, Perplexity 23.6% and Copilot 11.4%. The pool was 2.43 times as large as the broadest single engine’s list.

The overlap between engines was close to nothing. Of 528 distinct pages cited that day, 96.4% appeared on one engine only. In 84.9% of engine pairs, the two engines cited no page in common for the same prompt. ChatGPT and Google shared no cited page in any of their 11 comparisons. The author flags his conflict of interest and the narrow, self-referential topic.

| Engine | Share of all pages cited by the four engines |
|---|---|
| ChatGPT | 42.6% |
| Google | 24.3% |
| Perplexity | 23.6% |
| Copilot | 11.4% |

## Do other studies find the same split between engines?

Yes, the pattern of low overlap repeats across independent audits, countries and topics.

[Uberti-Bona Marin and colleagues](https://arxiv.org/abs/2609.18729), academics with no vendor tie, asked 117 real shopping questions in September 2026 from the Netherlands. For the same question at the same time, ChatGPT and Gemini shared only 5.4% of the websites they displayed. In 76.7% of comparisons they shared none.

[Martinez’s survey](https://arxiv.org/abs/2607.14035) of the research reports similar gaps elsewhere. One earlier audit found only 26% of domains were cited by both Bing Chat and Perplexity. Another found 53% of domains cited by Google’s AI Overviews did not appear in Google’s own top 10 results. Our guide on [whether AI engines cite the same sources](https://underneath.agency/resources/do-ai-search-engines-cite-the-same-sources) gathers these audits.

Google’s ranking does not stand in for ChatGPT either. In [our Google rankings study](https://underneath.agency/research/ai-citations-google-rankings-study), only 8.3% of the pages ChatGPT cited ranked in Google’s top 10 for the question, against 25.7% for Claude.

## Do engines also disagree on which brands they recommend?

Yes: different engines often name different brands for the same buyer question, not just different sources.

In [our four-assistant study](https://underneath.agency/research/ai-assistants-brand-agreement-study), we put 80 US buyer questions to ChatGPT, Gemini, Perplexity and Claude in September 2026. Two assistants’ recommended options overlapped by 0.327 on average, about a third. 66.3% of the options recommended for a question came from one assistant only. All four agreed on the first pick for 10.0% of questions.

This is not just run-to-run noise. Each assistant’s answers overlapped with its own repeat runs about twice as much as with another assistant’s. Of the brands one assistant named in at least three of five runs, 50.0% to 56.6% were never named by the other assistant in any of its five. Their taste in page quality is closer, as [whether AI engines prefer the same content](https://underneath.agency/resources/do-ai-engines-prefer-same-content) explains.

## Do engines differ in how they show sources and brands?

Yes, so the same metric can mean different things on each engine.

A [vendor study by Kumar at Ranqo](https://arxiv.org/abs/2606.20065) summarizes a January 2026 test of 50 questions about CRM software on five engines. The share of answers that included source links ranged from 95% on Perplexity to 15% on ChatGPT and 10% on Claude. A “citation share” on ChatGPT is therefore built from far fewer links than the same figure on Perplexity.

Language can create a similar blind spot. [Żatuchin](https://arxiv.org/abs/2606.23165), who works for a brand-monitoring firm, studied 66 European brands in twelve languages. Asking in a brand’s home language instead of English raised how often local brands were recommended by 0.80 on a 0 to 1 scale, against 0.15 for global brands. A single-engine, English-only tracker would miss both effects.

## Does it matter which version of ChatGPT a tool tracks?

Yes: the developer version that many tools query does not reproduce what consumers see in the app.

The Dutch audit paired each consumer-app question with the same question sent to OpenAI’s developer interface a fraction of a second later. The two shared an average of 12.0% of their displayed websites. Ask any tool which version of each engine it queries, from which country, and whether it is logged in.

## What should you do about it?

Track the engines your buyers actually use, and report each one separately before combining them.

1. Find out where your buyers ask. Include Google’s AI features, since they appear on the Google results page itself.
2. Track at least two or three engines, with the same fixed set of buyer questions on each.
3. Report every engine on its own. [Combine engines only with explicit weights](https://underneath.agency/resources/combine-ai-engines-visibility-score), never by adding raw citation counts.
4. Check whether a tool queries consumer apps or developer interfaces, and in which country and language.
5. When one engine moves, check the others before acting. A change on one may not appear anywhere else.

If you want help designing a multi-engine tracker, see [our generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

It does not show how much each engine is worth to a business, or whether one engine predicts the others.

- The four-engine coverage figures come from 15 prompts in one category on one day, by a vendor.
- No study here weights engines by how many buyers use them, so a coverage share is not a traffic share.
- Overlap at the level of brands, not just sources, has been measured on 80 questions or fewer.
- Claude, Copilot and Grok appear in few studies.
- Whether visibility on one engine leads to visibility on another over time has not been studied.

## Frequently asked questions

### Do ChatGPT and Google AI cite the same sources?

Rarely. In one test of 15 prompts, ChatGPT and Google shared no cited page in any of 11 comparisons, and 96.4% of all cited pages appeared on only one engine.

### Which AI engine should we track first?

The one your buyers use most, but not only that one. In one four-engine test, the broadest engine still covered just 42.6% of the pages cited across all four.

### Do Perplexity and ChatGPT recommend the same brands?

Often not. In our study, two assistants’ recommended options overlapped by about a third (0.327), and all four agreed on the first pick for 10.0% of questions.

### Is ranking on Google enough to show up in ChatGPT?

Not reliably. Only 8.3% of the pages ChatGPT cited in our study ranked in Google’s top 10 for the question.

## Sources

- Tannenbaum (2026), [Scoring With the Engine: Retrieval Exposure, Cross-Engine Divergence, and the Limits of Engine-Agnostic GEO Scores](https://arxiv.org/abs/2609.22655), arXiv:2609.22655.
- Uberti-Bona Marin, Bertaglia, Astante, Rijsbosch, van Dijck, Hannák, Spanakis and Kollnig (2026), ["If I Had to Buy Just ONE: Galaxy S26 Ultra": Auditing AI-Generated Product Recommendations](https://arxiv.org/abs/2609.18729), arXiv:2609.18729.
- Martinez (2026), [Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026)](https://arxiv.org/abs/2607.14035), arXiv:2607.14035.
- Kumar (2026), [Generative Engine Optimization at Scale: Measuring Brand Visibility Across AI Search Engines](https://arxiv.org/abs/2606.20065), arXiv:2606.20065.
- Żatuchin (2026), [The Language Blind Spot: How Query Language and Brand Recognition Tier Shape AI-Constructed Brand Reputation Across Twelve European Languages](https://arxiv.org/abs/2606.23165), arXiv:2606.23165.
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude agree on brands?](https://underneath.agency/research/ai-assistants-brand-agreement-study)
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude cite pages that rank?](https://underneath.agency/research/ai-citations-google-rankings-study)

---

This is the Markdown twin of https://underneath.agency/resources/is-tracking-chatgpt-enough. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How IT consulting firms win advisory clients from AI search"
description: "By being the independent adviser AI names for a company’s size, industry and decision, such as an ERP choice or a fractional CIO, with proof buyers can check."
canonical: "https://underneath.agency/resources/it-consulting-firms-clients-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How can an IT consulting firm win more advisory clients when mid-size companies ask AI first?

By being the independent adviser an AI assistant names when an executive describes their company’s size, industry and decision, and by publishing the proof that makes the recommendation hold up. Mid-size companies are under pressure to make big technology choices without a senior IT leader, and many now ask an assistant what to do before they call anyone. No study has yet measured how many advisory engagements begin that way.

This guide covers advisory IT consultancies: IT strategy and roadmaps, architecture reviews, software and vendor selection, fractional or virtual CIO services, and oversight of projects others deliver. Hands-on IT projects and support are covered in our guide to how IT services firms win projects from AI search. Large enterprise programs are covered in how technology consultancies reach enterprise shortlists.

## The short version

1. Mid-size companies need help with AI: in [RSM’s 2025 middle market survey](https://rsmus.com/insights/services/digital-transformation/rsm-middle-market-ai-survey-2025.html) of 966 technology decision-makers, 62% said generative AI had been harder to implement than expected, and 70% said they needed outside help.
2. The cost of a wrong system is high: [Gartner](https://www.theregister.com/2025/11/13/erp_disaster_gartner/) says 70 percent of ERP initiatives fail to fully meet their original business case goals, and 25 percent fail catastrophically.
3. Part-time leadership is growing from a small base: [Lightcast](https://lightcast.io/blog/rise-of-fractional-leadership) counted at least 34,000 US workers with “fractional” in their job title in 2025, up 265% since 2019, and 97% of fractional job postings in 2026 came from medium and small companies.
4. A full-time IT leader is expensive: the median wage for computer and information systems managers was $175,140 in 2025, according to [O*NET](https://www.onetonline.org/link/summary/11-3021.00) data from the Bureau of Labor Statistics.
5. Technology buyers now research with AI and use analysts less: in [TrustRadius’s 2026 report](https://www.demandgenreport.com/industry-news/news-brief/trustradius-ai-has-changed-how-buyers-research-but-not-what-they-trust/53853/), 63% of buyers used AI during their purchase, while analyst reports were used by only 13%.

## Who hires an IT consulting firm, and what is a client worth?

Usually a mid-size company’s CEO, CFO or COO with no CIO, or an IT manager needing strategic backing.

Three buyers account for most advisory work:

- **The executive without an IT leader.** A company with a few hundred employees, an IT manager or an outsourced provider, and no one at the leadership table who owns technology strategy. The CFO often holds IT by default.
- **The IT manager who needs cover.** Strong on operations, but facing a board question about AI, a major system replacement or a security audit, and wanting an outside view to support the business case.
- **The owner or investor preparing for change.** Growth, an acquisition or a sale exposes old systems, and someone must say what to fix first. The finance side of those moments is covered in [how financial consultancies win CFO and deal work](https://underneath.agency/resources/financial-consulting-firms-clients-ai-search).

The client is worth more than the first engagement. A paid assessment or roadmap often leads to a software selection project, then to oversight of the implementation, and in many cases to a monthly fractional CIO retainer. The comparison buyers make is with a hire: the median wage for an IT director-level manager was $175,140 in 2025, before benefits and recruiting. A part-time adviser is the cheaper way to get that judgment, which is why the fractional model has spread among smaller firms. Lightcast found only 0.3% of fractional leaders list a Fortune 1000 employer. Its data does not break out fractional CIOs; finance leads, at 46% of fractional postings. Large enterprise programs are a different sale, covered in [how technology consultancies reach enterprise shortlists](https://underneath.agency/resources/technology-consulting-firms-ai-search).

## Why do mid-size companies look for IT advice now?

AI pressure, risky system decisions and missing in-house expertise, often all at once.

**AI.** In RSM’s survey, 91% of respondents reported using generative AI, yet 92% had run into implementation challenges and 53% of those who had implemented it felt only “somewhat prepared.” Among those who felt unprepared, 39% named a lack of in-house expertise as their top issue. That is the gap an adviser fills.

**Big system choices.** ERP, CRM and data platform decisions shape a company for a decade, and they go wrong often. At Gartner’s Symposium, an analyst named “bad business cases, unrealistic scope, underestimated complexity, and scope creep” as common causes of ERP disasters. The Register cites Birmingham City Council, whose Oracle project costs rose from around £19 million to £170 million. A 2023 survey of 251 tech leaders found 73 percent felt their ERP strategy was not strongly aligned with their business strategy.

**Thin teams.** A mid-size company rarely employs an architect, a security strategist and a sourcing specialist. It rents that expertise when a decision demands it.

## Where does AI search fit in how executives find an IT adviser?

At the first two steps: understanding the decision, then finding someone independent to help make it.

A typical path:

1. **Trigger.** A board asks about AI, the ERP vendor ends support, the IT manager resigns, or an investor asks hard questions.
2. **Self-education.** The executive asks an assistant what the options are and what similar companies did.
3. **Finding help.** Referrals from the accountant, lawyer, investor or peers, plus searches and AI answers.
4. **Two or three conversations.** Fit, independence, fees and references.
5. **First engagement.** Usually a fixed-fee assessment or roadmap.
6. **Ongoing work.** Selection, oversight and a fractional retainer.

Step 2 is where AI changes this industry most. TrustRadius found 63% of technology buyers used AI during their purchase journey and 83% shortlisted three or fewer products. An assistant now drafts the software shortlist that a vendor-selection adviser used to build. That cuts both ways. Some executives will skip the adviser. Others will want an expert to check the assistant’s work: 94% of the AI users in the same survey fact-check its answers at least some of the time.

Services marketplaces are moving into the same conversations. Clutch, which lists IT consulting firms among many service categories, [launched an app inside ChatGPT](https://www.demandgenreport.com/?p=52857) in 2026 for comparing providers, reviews and “pricing signals.”

## What do executives ask AI when they need IT advice?

Questions about their own situation, framed by company size, industry and a decision they dread. Every example below is ours, written to show the phrasing; none is a captured buyer query.

| Situation | Example question |
|---|---|
| No IT leader | “Do I need a fractional CIO for a 200-person manufacturer, and how do I find a good one?” |
| ERP decision | “Should we replace our ERP or upgrade it? We are a $150 million distributor.” |
| Selection | “Independent ERP selection consultant that does not resell software, Midwest?” |
| Before a sale | “Who can assess our IT before a private equity sale?” |
| AI plan | “IT strategy consultant to build an AI roadmap for a regional bank?” |
| Price | “How much does an IT strategy assessment cost for a mid-size company?” |

The details in these questions matter. In [our prompt phrasing study](https://underneath.agency/research/ai-prompt-phrasing-study), adding “I run a small business with about 10 employees” to a question kept the original first brand only 40.6% of the time, against 68.0% when the same question was simply asked again. Who is asking changes who gets named. An adviser that only appears for generic questions may vanish when a buyer describes a 300-person company in a specific industry.

## How does an AI recommendation turn into an advisory engagement?

Through a discovery call. The assistant can put your name forward; the first conversation still has to earn the assessment.

1. **AI answer.** The executive sees a few advisers or types of adviser, sometimes with a reason for each.
2. **Check.** Your site, LinkedIn profiles, reviews and any articles quoting you. In [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study), ChatGPT itself looked for reviews in 46.2% of its answers before replying.
3. **Discovery call.** The buyer tests whether you understand companies like theirs.
4. **Fixed-fee assessment.** A defined first purchase with low risk.
5. **Selection, oversight or retainer.** The larger and longer revenue.

Independence is the deciding trust factor in this industry. A buyer choosing an ERP or a cloud platform wants to know whether the adviser earns referral fees or resells. A firm that says plainly what it does not sell, and shows decisions it has steered in both directions, gives both the buyer and the assistant a reason to recommend it, we infer. Firms that also deliver hands-on projects can read [how IT services firms win projects from AI search](https://underneath.agency/resources/it-service-firms-leads-ai-search).

## What decides whether an assistant names an IT adviser?

No platform documents how it picks consultants. Studies suggest answers vary widely, and that reviews and specific context matter.

**Documented by the platforms.** None of the major assistants publishes rules for recommending professional advisers.

**Observed in our studies, in other categories.** We have not studied IT consultancies directly:

- **Answers vary from run to run.** In [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), a single ChatGPT answer showed 57.8% of the brands its five answers named between them, and 25.2% of the brands appeared in all five runs. One check of one question proves little; we compare engines in [which AI engine is most consistent for brands](https://underneath.agency/resources/most-consistent-ai-engine-for-brands).
- **Context reshuffles the list.** The prompt phrasing study shows company size and budget wording change who is named.
- **Reviews are looked up.** ChatGPT searched for reviews in nearly half of its answers in the hidden-searches study.

**Our inference, specific to advisory IT.** Advisers that publish clear decision guides (replace or upgrade, how to run a selection, what a fractional CIO does in the first 90 days), state the company sizes and industries they serve, and are quoted in regional business and trade press give assistants more to cite. Smaller firms can win here; we explain why in [how small brands get recommended by AI](https://underneath.agency/resources/how-small-brands-get-recommended-by-ai).

## What does it cost an IT consultancy to be missing?

The executives who decide alone, or with whichever adviser the assistant named. No study has measured that loss.

- **The adviser’s role is being automated at the edges.** TrustRadius found analyst reports used by 13% of buyers, a 63% decrease since 2022. If buyers trust an assistant’s shortlist more than a research report, the adviser who is absent from the assistant’s answer loses the chance to challenge it.
- **The first engagement leads to the rest.** Missing the assessment means missing the selection, the oversight and the retainer.
- **Wrong descriptions send the wrong buyers.** An answer that calls you an IT support company, or a reseller of one vendor, repels the executives who want independent advice. Our guide to [correcting wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) shows how to fix that at the source.

## How does GEO work for an IT consulting firm?

It makes your independence, specialties and client fit clear to assistants and buyers. No one can guarantee a recommendation.

1. **Say what you are and are not.** Advisory, vendor-neutral, no resale, or the partnerships you do hold. State it the same way on your site, LinkedIn, Clutch and association profiles.
2. **Name your client profile.** Company sizes, industries and regions. Assistants answer size-specific questions, and our phrasing study shows those details change the list.
3. **Publish decision guides.** Replace or upgrade, build or buy, how to run a software selection, what a roadmap contains, how fractional CIO retainers work. Dated, specific and signed by a named adviser.
4. **Show outcomes and both-way decisions.** Case summaries where you recommended against a purchase are strong evidence of independence.
5. **Collect reviews that name the work.** “Ran our ERP selection and saved us from a bad fit” says more than “great to work with.”
6. **Explain fees.** Fixed-fee assessments, day rates or retainer ranges, at least in outline. Buyers ask, and assistants look for prices.
7. **Get quoted.** Regional business journals, industry associations and trade publications that cover your clients’ sectors.
8. **Test the questions in context.** Ask ChatGPT, Gemini, Perplexity, Claude, Copilot and Google’s AI Overviews and AI Mode your buyers’ questions, with their company size and industry, several times each. For the software those buyers end up choosing, see [how ERP vendors appear in AI search](https://underneath.agency/resources/erp-software-ai-search).

## What is still unproven about AI and IT advisory work?

Whether AI answers send executives to advisers or replace them. No study has tested either for IT consulting.

- **No survey of advisory IT buyers.** RSM surveys mid-size companies’ AI use, not how they hire advisers; TrustRadius surveys software buyers.
- **Fractional CIO data is thin.** Lightcast does not break out IT roles.
- **ERP failure rates are an analyst view.** Gartner’s 70 percent comes from a conference presentation, not a published study we could read in full.
- **Our studies covered other categories.** Their patterns may not hold for advisers.

## Where should an IT consulting firm start?

With the questions your best clients asked themselves before they called you, phrased with their size and industry.

Write 20 of them: fractional CIO, ERP decisions, AI roadmaps, pre-sale assessments, fees. Ask each major assistant several times. Record which firms or types of adviser are named, what sources are cited, and whether your independence and specialties are described correctly.

If you would like a second opinion on what you find, [ask us to review your advisory questions](https://underneath.agency/contact). We will show where your firm appears, where other advisers or software vendors fill the answer instead, and which gaps in your guides, reviews and profiles are most likely costing you discovery calls and first assessments. Beyond that review, our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page explains how an independent adviser’s decision guides, profiles and press mentions are built up and retested over time.

## Frequently asked questions

### Do executives really ask ChatGPT whether they need a fractional CIO?

No survey isolates that question. TrustRadius found 63% of technology buyers used AI during a recent purchase, and fractional leadership is growing fast among smaller firms.

### Will AI assistants replace vendor-selection consultants?

Not for complex decisions, we infer. Buyers shortlist with AI, but 94% of the AI users in TrustRadius’s survey fact-check its answers.

### Should an IT adviser publish fees?

At least explain the structure. Buyers ask about cost early, and fixed-fee assessments are easier to choose than open-ended work.

### Does company size in the question change the answer?

In our phrasing study, adding a small-business detail kept the original first brand only 40.6% of the time, against 68.0% for a plain rerun.

### How is this different from managed services?

Advisers sell judgment and independence; managed service providers sell ongoing operations. See [how MSPs win clients from AI search](https://underneath.agency/resources/msps-customers-from-ai-search).

## Sources

- RSM US (2025), [The middle market is embracing AI to drive innovation and efficiency: RSM Middle Market AI Survey 2025](https://rsmus.com/insights/services/digital-transformation/rsm-middle-market-ai-survey-2025.html)
- The Register (2025-11-13), [ERP carnage continues as orgs jump in unprepared](https://www.theregister.com/2025/11/13/erp_disaster_gartner/)
- CIO Dive (2025), [Companies mitigate ERP project risks](https://www.ciodive.com/news/companies-mitigate-ERP-project-risks/752801/)
- Lightcast (2026-08-20), [The rise of fractional leadership](https://lightcast.io/blog/rise-of-fractional-leadership)
- O*NET OnLine, U.S. Department of Labor (2025), [Computer and Information Systems Managers (11-3021.00)](https://www.onetonline.org/link/summary/11-3021.00)
- Demand Gen Report (2026-07-30), [TrustRadius: AI has changed how buyers research, but not what they trust](https://www.demandgenreport.com/industry-news/news-brief/trustradius-ai-has-changed-how-buyers-research-but-not-what-they-trust/53853/)
- Demand Gen Report (2026-05-12), [Clutch launches first B2B services marketplace app on ChatGPT](https://www.demandgenreport.com/?p=52857)
- Underneath (2026), [Does rewording a question change AI brand recommendations?](https://underneath.agency/research/ai-prompt-phrasing-study)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)

---

This is the Markdown twin of https://underneath.agency/resources/it-consulting-firms-clients-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How IT service firms can win projects and calls from AI search"
description: "By being the local or specialist firm AI names for a specific problem, place and project, backed by reviews, listings and clear proof of what you fix."
canonical: "https://underneath.agency/resources/it-service-firms-leads-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How can an IT services firm win more projects and support calls from AI search?

By being the firm an AI assistant names when a business owner describes a specific problem, in a specific place, for a specific project, and by keeping the reviews, listings and service pages that answer draws on accurate. IT support and project work is bought in a hurry or against a deadline, often by someone who is not an IT person. That buyer increasingly asks an assistant before calling anyone, but no study has yet measured how often that leads to a signed IT project.

This guide covers project-based IT consultancies and local IT support firms: break/fix work, on-site support, networking, migrations and one-off projects. Recurring managed-services contracts follow a different path and deserve their own treatment.

## The short version

1. The market is large and growing: [Gartner](https://www.gartner.com/en/newsroom/press-releases/2025-10-22-gartner-forecasts-worldwide-it-spending-to-grow-9-point-8-percent-in-2026-exceeding-6-trillion-dollars-for-the-first-time) forecasts worldwide IT services spending of $1,869,269 million in 2026, up 8.7%.
2. Outages make IT purchases urgent: [ITIC’s 2024 survey](https://calyptix.com/wp-content/uploads/ITIC-2024-Hourly-Cost-of-Downtime-Survey-Results-Part-2.pdf) of more than 1,000 firms found an hour of downtime costs over $300,000 for more than 90% of mid-size and large enterprises.
3. Small firms lean on their IT provider for advice: in the [UK government’s 2025 breaches survey](https://www.gov.uk/government/statistics/cyber-security-breaches-survey-2025/cyber-security-breaches-survey-2025), external cyber security consultants and IT providers were businesses’ most common source of security information (25%).
4. Local discovery is moving to AI: in [BrightLocal’s 2026 survey](https://www.brightlocal.com/research/local-consumer-review-survey/) of 1,002 US consumers, use of ChatGPT and other AI tools for local business recommendations rose from 6% to 45% in a year. These are consumers, not IT buyers, but small-business owners search the same way, we infer.
5. Deadlines create projects: [Microsoft](https://www.microsoft.com/en-us/windows/end-of-support) ended Windows 10 support on October 14, 2025, with paid security updates available until October 12, 2027.

## Who buys IT services, and what is a client worth?

Mostly small-business owners and office managers without IT staff, and IT managers at mid-size firms buying a project.

Three buyers matter most:

- **The small-business owner or office manager.** A 10- to 50-person law office, clinic, agency or manufacturer with no IT employee. They buy fixes, setups, office moves and security help, usually locally.
- **The mid-size IT manager.** A one- or two-person IT team that buys projects it cannot staff: a Microsoft 365 or cloud migration, a network refresh, a server replacement, a Windows 11 rollout.
- **The operations or finance leader after a scare.** An outage, a ransomware attempt or an auditor’s finding turns IT into an urgent purchase.

The client’s value is the relationship, not the first ticket. A firm that fixes a server today is the natural choice for the next office move, the next migration and, often, a later support agreement. That is why being the first name a buyer hears matters so much in this industry.

IT firms are mostly small too. In a [Techreviewer survey](https://techreviewer.co/research/main-marketing-channels-of-it-services-companies-in-2025) of 62 IT service providers, 71.5% had fewer than 100 employees. The sample leans toward software outsourcing firms, but the pattern holds across the field: many small firms competing for attention.

## What triggers an IT project or support call?

An outage, a deadline, a move or a security scare; rarely a calm, planned search.

**Outages.** ITIC found 84% of its respondents rank security attacks as the top cause of downtime, and 88% of organizations need at least 99.99% uptime. A small office does not lose $300,000 an hour, but ITIC notes that losses of $25,000 to $75,000 an hour “may be serious enough to put the SMB out of business.”

**Deadlines.** Windows 10 end of support is the clearest recent example. Microsoft says it “no longer provides software updates, security fixes, or technical assistance to Windows 10 PCs.” Every office with old machines faced a hardware and upgrade decision, and the paid extension runs only until October 2027.

**Advice gaps.** The UK survey describes boards placing “a great deal of trust” in “their external IT providers,” and smaller firms passing security responsibility to contractors. The buyer often wants someone to tell them what to do, not just to do it.

**Growth and moves.** New staff, a second site or a new office all create cabling, network and device projects.

## Where does AI search sit in the IT services buying journey?

At the first step, when a stressed owner or IT manager looks for someone who can fix this, here, soon.

The typical path:

1. **Trigger.** Something breaks, a deadline looms or a project is approved.
2. **First search.** Peers, Google, Google Maps and, increasingly, an AI assistant.
3. **Check.** Reviews, Google Business Profile, the firm’s website, perhaps a Clutch profile.
4. **Call or quote request.** Often to two or three firms.
5. **Project or ticket.** Then repeat work if the experience is good.

Two documented changes put AI at step 2. First, OpenAI says [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search) “can also use location information to find local results” and may rewrite a “near me” question into a local search sent to partner search providers. Second, directories are moving into the assistants: Clutch, a B2B services marketplace, [launched an app inside ChatGPT](https://www.demandgenreport.com/?p=52857) in 2026 so buyers can compare service providers, reviews and “pricing signals” without leaving the conversation.

For bigger projects, buyer surveys point the same way. [Responsive’s 2025 survey](https://www.responsive.io/news/buyer-intelligence-2025) of 350 B2B buyers found 48% of US buyers use generative AI for vendor discovery, and 61% start with a preferred vendor in mind. Being named early, before that preference forms, is the opportunity.

## Which questions do IT service buyers ask AI assistants?

Concrete questions about a problem, a place, a system and a budget. We wrote the prompts below to show how an office manager or IT director might phrase a request; none of them was captured from a real buyer.

| Buyer | Illustrative prompt |
|---|---|
| Small law office | “IT support company in Columbus that works with law firms, about 15 people?” |
| Deadline project | “Who can upgrade our 40 PCs from Windows 10 to Windows 11 before our insurance renewal?” |
| Migration | “Microsoft 365 migration partner for a dental group with three locations?” |
| Office move | “Network cabling and Wi-Fi contractor for a new office in Austin?” |
| No contract wanted | “IT consultant for a one-off server replacement, no monthly contract?” |
| After a scare | “Our email was hacked. Who can help a small business in Leeds today?” |
| Price | “How much should a small office network upgrade cost?” |

Place names matter a great deal. In [our four-assistant study](https://underneath.agency/research/ai-assistants-brand-agreement-study), questions naming a place had an overlap of 0.160 between assistants, against 0.390 for national questions. The local list a buyer sees depends heavily on which assistant they use.

## How does AI visibility turn into IT projects and support calls?

When an assistant names your firm for the buyer’s problem and town, the next step is a phone call.

Local IT work is unusual among B2B services: the path from answer to call is short. The likely sequence:

1. **AI answer.** Two to five named firms, often with ratings, addresses and phone numbers.
2. **Quick check.** Reviews and a glance at the website. In BrightLocal’s survey, 31% of consumers will only use a business with 4.5 stars or more, up from 17% a year earlier.
3. **Call.** The buyer describes the problem; the firm scopes it on the phone or on site.
4. **Project.** Fixed-price for a migration or rollout, hourly for break/fix.
5. **Repeat work.** The next project, and possibly a support agreement.

Trust is the gate. BrightLocal found 40% of consumers trust AI platforms to provide business recommendations. A buyer who calls after an AI recommendation still checks you, so the reviews and website must confirm what the assistant said. Few of these leads show up as a website visit in analytics, a gap we explain in [what lost clicks mean for pipeline](https://underneath.agency/resources/ai-answers-pipeline-revenue).

## What decides whether an assistant names a local IT firm?

No platform publishes how it picks IT firms; for local questions, studies point to map listings and review counts.

**Documented by the platforms.** Google says [local results](https://support.google.com/business/answer/7091) are “mainly based on relevance, distance, and popularity.” OpenAI documents that ChatGPT uses approximate location and partner search providers for local questions. Neither says how an AI answer chooses among IT firms.

**Observed in our studies, in other local trades.** Our local studies asked about dentists, plumbers, lawyers, accountants and physical therapists, not IT firms, so treat them as a guide:

- In [our study of ChatGPT’s local answers](https://underneath.agency/research/chatgpt-local-recommendations-study), 53.2% of the businesses ChatGPT listed were in the same-day Google Maps top 20, and their details matched Maps almost field for field.
- In [our study of which Maps businesses ChatGPT picks](https://underneath.agency/research/chatgpt-local-picks-google-profile-study), businesses with more reviews than the local median were 19.5 points more likely to be listed at the same Maps rank. ChatGPT listed 67.7% of the businesses ranked 1 to 3 in Maps.
- In [our brand entity study](https://underneath.agency/research/brand-entity-ai-recommendations-study), on local questions “no entity signal made a difference in any model.” A Wikipedia page will not help a local IT shop; reviews and listings appear to matter more.

**Our inference.** For project work beyond the local area, a reasonable expectation is that assistants lean on directories such as Clutch, case studies and independent mentions, as they do for other B2B services. No study has tested that for IT firms.

## What does it cost an IT firm to be missing?

Mostly lost first calls and the repeat projects behind them. No study has measured that loss for IT firms.

- **The first call decides the relationship.** A buyer who never calls you does not come back for the next migration.
- **Deadlines bunch demand.** Events like Windows 10 end of support send many buyers searching at once; firms missing from the answer miss the wave.
- **Wrong facts turn buyers away.** An assistant that lists an old address, a closed phone line or “managed services only” when you also take one-off projects loses you exactly the buyers who would fit. Our guide to [correcting wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) walks through the fixes, starting with your listings.
- **Firms rely on channels that are changing.** In the Techreviewer survey, 69.4% of IT firms rated content marketing and SEO as an effective lead source, and 71% used referral platforms. If those platforms and searches move into assistants, the firms named there inherit the leads, we infer.

## How does GEO work for an IT services firm?

It makes your services, service area and proof easy for assistants and buyers to find and confirm. Nobody can promise a recommendation.

1. **Complete and correct your listings.** Google Business Profile categories, service area, hours and phone, consistent across directories. Our local studies found ChatGPT’s business details match Maps almost exactly, so errors there travel.
2. **Earn steady reviews that name the work.** “Migrated our 25 mailboxes to Microsoft 365” tells an assistant and a buyer more than “great service.”
3. **Write one page per service and per place.** Break/fix, Windows 11 upgrades, Microsoft 365 migrations, networking, cabling, security clean-up; each with systems supported, typical timelines and the towns you cover.
4. **Say how you price.** Hourly, fixed-price projects, or no-contract work. ChatGPT searched for prices in 23.8% of its answers in [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study).
5. **Publish real project stories.** Industry, size, systems and outcome, with the client’s permission.
6. **Build local and industry proof.** Chamber of commerce listings, local press, vendor partner directories (Microsoft, Cisco and others) and Clutch-style profiles. A local IT shop competing with national providers can borrow from [how small brands get recommended by AI](https://underneath.agency/resources/how-small-brands-get-recommended-by-ai) and [can small websites get cited by AI](https://underneath.agency/resources/can-small-websites-get-cited-by-ai).
7. **Test the questions your buyers ask.** Use your town, industries and projects across ChatGPT, Google AI Overviews and AI Mode, Gemini, Perplexity, Copilot and Claude, several times each.

## What can’t an IT firm yet prove about AI-sourced leads?

How many IT buyers use AI assistants to pick a provider, and how many projects result. No study answers either.

- **No survey of IT services buyers.** BrightLocal surveys consumers; Responsive and Gartner survey B2B buyers across industries.
- **Our local studies cover other trades.** IT firms may be listed differently.
- **Industry surveys are small or skewed.** Techreviewer’s 62 respondents lean toward software outsourcing firms.
- **The link to revenue is untested.** In [Martinez’s](https://arxiv.org/abs/2607.14035) review of research on AI search optimization, traffic and conversions had the weakest evidence.

## Where should an IT services firm start?

With the questions your best clients asked before they first called you, in their own words.

List 20 to 30 of them: urgent fixes, deadline projects, migrations, moves, price questions. Ask each major assistant several times, from your area. Record which firms are named, which sites are cited, and whether your services, area, hours and phone number are correct. Compare that with the projects you want more of.

If you would like an outside set of eyes on those answers, [ask us to run the check for your service area](https://underneath.agency/contact). We will show which local and project questions name your firm, which send callers to competitors, and which gaps in your listings, reviews and service pages are most likely costing you calls, quotes and projects. Listing cleanup, service-and-town pages and review work for local IT firms are all part of what our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page describes.

## Frequently asked questions

### Do business owners really ask ChatGPT to find an IT company?

No study isolates IT buyers. Among US consumers, AI use for local business recommendations rose from 6% to 45% in a year, in BrightLocal’s 2026 survey.

### Do reviews matter for AI recommendations?

In other local trades, yes. Businesses with more reviews than the local median were 19.5 points more likely to be listed by ChatGPT in our study.

### Should we publish prices?

Explain how you price, at least. ChatGPT searched for prices in 23.8% of answers in our hidden-searches study.

### Does Windows 10 end of support still create work?

Yes. Support ended on October 14, 2025, and paid security updates run only until October 12, 2027.

### Is this different for managed service providers?

Yes. Recurring managed contracts are bought through a longer comparison of pricing models and coverage. This guide covers projects and break/fix work.

## Sources

- Gartner (2025-10-22), [Gartner forecasts worldwide IT spending to grow 9.8% in 2026, exceeding $6 trillion for the first time](https://www.gartner.com/en/newsroom/press-releases/2025-10-22-gartner-forecasts-worldwide-it-spending-to-grow-9-point-8-percent-in-2026-exceeding-6-trillion-dollars-for-the-first-time)
- Information Technology Intelligence Consulting (ITIC) (2024), [ITIC 2024 Hourly Cost of Downtime Survey Results, Part 2](https://calyptix.com/wp-content/uploads/ITIC-2024-Hourly-Cost-of-Downtime-Survey-Results-Part-2.pdf)
- UK Department for Science, Innovation and Technology (2025-04-09), [Cyber security breaches survey 2025](https://www.gov.uk/government/statistics/cyber-security-breaches-survey-2025/cyber-security-breaches-survey-2025)
- BrightLocal (2026-02-11), [Local Consumer Review Survey 2026](https://www.brightlocal.com/research/local-consumer-review-survey/)
- Microsoft (2025), [End of support for Windows 10, Windows 8.1, and Windows 7](https://www.microsoft.com/en-us/windows/end-of-support)
- Techreviewer (2025), [Main marketing channels of IT services companies in 2025](https://techreviewer.co/research/main-marketing-channels-of-it-services-companies-in-2025)
- OpenAI Help Center (2026), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Demand Gen Report (2026-05-12), [Clutch launches first B2B services marketplace app on ChatGPT](https://www.demandgenreport.com/?p=52857)
- Responsive (2025-10-15), [GenAI overtakes search for a quarter of B2B buyers](https://www.responsive.io/news/buyer-intelligence-2025)
- Google Business Profile Help, [Tips to improve your local ranking on Google](https://support.google.com/business/answer/7091)
- Martinez (2026), [Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026)](https://arxiv.org/abs/2607.14035), arXiv:2607.14035.
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude agree on brands?](https://underneath.agency/research/ai-assistants-brand-agreement-study)
- Underneath (2026), [ChatGPT local recommendations: stable details, shifting lists](https://underneath.agency/research/chatgpt-local-recommendations-study)
- Underneath (2026), [Which Google Maps businesses does ChatGPT recommend?](https://underneath.agency/research/chatgpt-local-picks-google-profile-study)
- Underneath (2026), [Do Wikipedia and schema make AI assistants recommend a brand?](https://underneath.agency/research/brand-entity-ai-recommendations-study)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)

---

This is the Markdown twin of https://underneath.agency/resources/it-service-firms-leads-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How can a jewelry brand get recommended when shoppers ask AI?"
description: "By being the jeweler AI names for ring and gift questions, with current facts on certification, sourcing, price and where to buy, before peak seasons."
canonical: "https://underneath.agency/resources/jewelry-brands-ai-product-recommendations"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How can a jewelry brand get recommended when shoppers ask AI?

By giving AI assistants accurate, current and independently confirmed facts about your rings and gifts: certification, stone type, sourcing, price and where to see them in person. A monthly study of four AI assistants finds a few online jewelers named again and again for engagement ring questions, mall chains rarely named, and answers that still describe one brand as if a merger had not happened. For a jeweler, being named, and named accurately, decides who gets a high-value order.

## The short version

1. Jewelry purchases are large and planned: couples in [The Knot’s 2026 study](https://www.jckonline.com/editorial-article/lab-diamonds-engagement-market/) spent an average of $4,600 on an engagement ring, 61% chose a lab-grown center stone, and 57% of proposers began searching more than six months before proposing.
2. Gifting is a calendar business: the [National Retail Federation](https://nrf.com/media-center/press-releases/valentine-s-day-spending-expected-to-reach-new-records) expected Americans to spend $7 billion on jewelry for Valentine’s Day 2026, more than on any other gift.
3. AI answers concentrate on a few jewelers: in [Rings.com’s September 2026 study](https://rings.com/pages/ai-jewelry-recommendations-study) of 2,400 AI responses to 30 jewelry questions, Brilliant Earth and Blue Nile led, while four mall chains together drew 8.1% of brand mentions.
4. AI answers can be out of date: in the same study, James Allen appeared in 54.0% of responses, and 90.7% of those showed no acknowledgment of its consolidation by Signet.
5. Which assistant a shopper asks changes the answer: OpenAI’s model named Rare Carat 10 times in 600 responses, while Perplexity named it 332 times.

## Who buys jewelry, and what is one customer worth?

Proposers, gift buyers and self-purchasers, with orders that often run to thousands of dollars.

Three buyers matter most. The proposer buys an engagement ring, usually once, after months of research. The gift buyer shops around Valentine’s Day, Mother’s Day, anniversaries and the holidays. The self-purchaser buys everyday fine jewelry for themselves. Each asks different questions, but all three are spending enough to compare carefully.

The numbers are large. The Knot’s Real Weddings Study 2026, based on 10,474 US couples married in 2025, found an average engagement ring spend of $4,600 for an average 1.9-carat stone. [INSTORE’s summary](https://instoremag.com/what-10000-couples-who-just-got-married-can-tell-you-about-your-business/) adds that lab-grown buyers averaged $4,300 for a 2.0-carat center stone, and natural buyers $7,000 for 1.6 carats. At [Brilliant Earth](https://nationaljeweler.com/articles/15203-brilliant-earth-s-q2-sales-up-6-driven-by-high-income-shoppers), average order value rose 8% to $2,238 in the second quarter of 2026, with non-bridal fine jewelry driving growth. At the other end of the market, [Signet Jewelers](https://s26.q4cdn.com/755441662/files/doc_news/Signet-Jewelers-Reports-Preliminary-Results-for-Fourth-Quarter-and-Full-Year-Fiscal-2026-2026.pdf), which owns Kay, Zales, Jared and Blue Nile, reported preliminary fiscal 2026 sales of approximately $6.8 billion across about 2,600 stores. The top of the market is covered in [how luxury brands reach high spenders through AI](https://underneath.agency/resources/luxury-brands-ai-search).

A customer is often worth more than one order. The engagement ring leads to wedding bands, and the jeweler who sold it is in a good position for anniversary gifts. We infer that the first bridal purchase is where much of a jewelry customer’s long-term value is decided, though we found no public figure that measures it.

## How do people shop for jewelry now?

They research online for months, compare a handful of jewelers, and often still buy in person.

The Knot found that proposers visited an average of two retailers in person and looked at 10 rings before buying, that 64% bought in person and about a third bought online, and that 79% of ring recipients took part in choosing. Almost nine in ten made custom edits or designed the ring from scratch. Around 40% of proposals happen between Thanksgiving and Valentine’s Day, so ring research peaks months before that window.

Gift buying leans online. The NRF’s survey of 7,791 adults found that 25% planned to buy jewelry for Valentine’s Day 2026 and that online was the top shopping destination (38%). The Plumb Club’s 2025 survey of more than 2,000 US consumers, [reported by Rapaport](https://rapaport.com/news/online-jewelry-shopping-on-the-rise-plumb-club-study-shows/), found 62% had bought jewelry online, average online spend rose 22% to $1,652, and 34% preferred to browse and buy on brand websites.

AI now sits in front of much of this research. Across all US retail, [Adobe’s data](https://techcrunch.com/2026/04/16/ai-traffic-to-us-retailers-rose-393-in-q1-and-its-boosting-their-revenue-too/) shows AI traffic to retail sites rose 393% in the first quarter of 2026 compared with a year earlier, and AI visitors converted 42% better than other traffic in March. Those figures are not jewelry-specific. We found no public survey of how many jewelry buyers use AI assistants, so we do not quote one.

## Which questions do jewelry shoppers put to AI?

Where to buy, lab-grown or natural, ethical sourcing, custom designs, price and gift ideas.

The Rings.com study tests real-sounding questions in five groups: general purchase, lab-grown diamonds, ethical sourcing, custom and gifting, and price and value. Its question list includes “best place to buy an engagement ring online,” “best place to buy lab-grown diamond rings online,” “best alternative to Tiffany for engagement rings” and “best jewelry brands for anniversary gifts.” These are the study’s test questions, not a log of what shoppers typed.

We wrote the prompts below to illustrate other questions buyers bring; they are not observed queries:

- Bridal: “2-carat oval lab-grown diamond with an independent grading report, under $4,000, in yellow gold.”
- Gift: “Anniversary gift under $500 for a wife who only wears gold and dislikes big pendants.”
- Sourcing: “Which jewelry brands use recycled gold and can show where their diamonds come from?”
- Trust: “Is it safe to buy an engagement ring online, and which jewelers offer free resizing and returns?”
- Everyday: “Waterproof gold necklace I can wear every day that won’t tarnish.”

The pattern matters for revenue. Bridal questions are about a single large purchase. Gift questions spike before fixed dates. Both are high intent: the shopper is close to choosing where to spend. Adding a budget can reshuffle the names, as [our study of question wording](https://underneath.agency/research/ai-prompt-phrasing-study) found.

## How does an AI answer become a jewelry sale?

Through a named jeweler, a site visit or showroom appointment, then an order that can lead to more.

**The shortlist.** The assistant names a few jewelers. Because proposers visit only about two retailers, a reasonable expectation is that being on the AI shortlist strongly shapes which two they visit.

**The visit.** Most rings are still bought in person. An AI answer that names your brand may lead to a showroom appointment rather than a tracked click, so the revenue shows up in store sales, not in AI referral reports. Brilliant Earth, for example, sells online and through 43 showrooms.

**The online order.** For gifts and everyday pieces, the sale is more often online. OpenAI’s help page says ChatGPT may show an [Instant Checkout option](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search) for some eligible products and merchants, though [FashionUnited reports](https://fashionunited.com/news/retail/us-consumers-would-let-ai-agents-buy-clothes-but-not-without-a-say/2026092974919) OpenAI scaled the feature back in March 2026. The same page describes a “Try on” button for clothes and accessories, so a shopper can see how a pendant or pair of earrings could look on them.

**The next purchase.** A satisfied bridal customer returns for bands and anniversary gifts. We infer that AI visibility for bridal questions pays back over years, not one order.

## What decides which jeweler an AI assistant names?

Platforms document product data, reviews and price; studies show answers vary by assistant and lag behind the market.

**Documented by the platform.** OpenAI says ChatGPT considers structured metadata from first-party and third-party providers, such as price and product description, other third-party content and reviews when selecting products. If the same ring or pendant is sold by more than one store, ChatGPT orders those sellers by availability, price, quality and whether each is the maker or primary seller. On Google’s side, AI Mode pulls from [shopping data for billions of products](https://blog.google/products/search/ai-mode-search/), firing off several related searches in parallel and merging what comes back.

**Observed in a study.** The Rings.com study, run monthly since May 2026 across OpenAI, Gemini, Perplexity and Claude with web search turned on, plus Google AI Overviews, is the most detailed public record for this industry. Rings.com is a participant in the market it measures, running a ring-shopping guidance tool, and discloses that interest; its own brand recorded no mentions. Its September findings include:

- Brilliant Earth scored 15,860 weighted points with 1,910 mentions, ahead of Blue Nile at 15,448 and 1,809.
- Rare Carat’s score rose 102% in one month. The study says it measures what the models say, “not why.”
- Kay, Zales, Jared and Helzberg together drew 1,043 of 12,887 brand mentions.
- Forbes.com was the single largest cited source, with 2,677 captured citations, so independent publishers shape these answers.
- Google AI Overviews appeared on 564 of 600 searches for the study’s questions (94.0%).

**Trust factors specific to jewelry.** Buyers want proof. In the Plumb Club survey, [summarized by INSTORE](https://instoremag.com/new-research-by-the-plumb-club-reveals-consumers-more-cautious-and-want-assurances-when-buying-jewel/), 32% named a brand certificate of authenticity as the most important factor when buying jewelry and 29% named independent laboratory verification; 41% said they would be unlikely to buy if they could not verify that a piece had been responsibly sourced. Lab-grown or natural, grading reports, metal, recycled materials, return and resizing policies, and showroom locations are the facts a careful assistant needs to answer these questions well. Our inference is that jewelers who publish these facts clearly, and whose facts match what independent publishers say, give assistants more reason to name them. Makeup brands face a similar need for precise product facts, as [how beauty brands get recommended by AI](https://underneath.agency/resources/cosmetics-brands-product-discovery-ai-search) shows.

## What does a jeweler lose when AI leaves it out or gets it wrong?

Bridal shoppers who visit only two jewelers, gift buyers in peak weeks, and trust when facts are stale.

The cost of absence is clearest in bridal. If a proposer visits about two jewelers and an AI answer named three others, we infer the omitted jeweler has lost that sale without knowing it was in play. In gifting, the window is short: Valentine’s spending is concentrated in a few weeks, and an answer that is wrong in late January cannot be corrected in time.

The cost of being wrong is visible in the Rings.com data. Signet has consolidated James Allen, yet 1,295 of 2,400 responses still named it, and most of them showed no acknowledgment of the change. When Signet acquired The Clear Cut in late May 2026, the study found it named in only 5 responses in September, none mentioning Signet or Blue Nile. For any jeweler that merges brands, closes stores, changes its lab-grown policy or reprices, answers can lag. Two guides cover the repair work: [correcting wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) and [why ChatGPT misses new products](https://underneath.agency/resources/why-chatgpt-misses-new-products), which matters for new collections.

## How does GEO work for a jewelry brand?

It makes your jewelry facts easy for AI to find, trust and repeat, on your site and in independent sources.

For a jeweler, generative engine optimization (GEO) means making sure AI answers name the brand and describe its stones, certificates and stores correctly. It cannot guarantee a recommendation. The work usually covers:

1. **Brand and company facts.** Which brands you own, which sites and showrooms are open, what changed and when, stated plainly and consistently, so assistants stop repeating old structures.
2. **Product facts that answer trust questions.** Stone type, grading report and lab, carat, metal, recycled content, sourcing, price, resizing and returns, on every product page and in product feeds.
3. **Occasion and budget pages.** Clear gift guides by occasion and price, published well before Valentine’s Day, Mother’s Day and the holidays.
4. **Independent coverage.** Reviews and roundups on the publishers assistants cite, earned through digital PR and product loans for honest testing. Our article on [which pages to target](https://underneath.agency/resources/best-of-lists-ai-recommendations) explains why third-party lists carry weight.
5. **Reviews.** Honest customer reviews on your site and independent platforms, including bridal stories that mention service, custom work and resizing.
6. **Readable pages.** Product and showroom pages that AI systems can access and read.
7. **Measurement by assistant.** Tracking bridal and gift questions separately in ChatGPT, Gemini, Perplexity, Claude and Google, because the Rings.com data shows they disagree. Our article on [why tracking only ChatGPT is not enough](https://underneath.agency/resources/is-tracking-chatgpt-enough) covers this.

## What can’t jewelers yet measure about AI shoppers?

How many jewelry buyers use AI, and how much jewelry revenue AI answers drive.

We found no public survey measuring how many jewelry shoppers use AI assistants, and no jeweler has published revenue from AI referrals. The Rings.com study measures what models say for 30 questions, not what shoppers do, and its author is a participant in the ring market. Its month-to-month swings, including Rare Carat’s doubling and changes at the mall chains, show that one month’s ranking is a weak guide. Adobe’s conversion figures cover all of US retail. The Knot and Plumb Club figures describe buyers, not AI use. Treat any claim of a guaranteed AI ranking for jewelry with suspicion.

## Where should a jewelry brand start?

With your bridal and gift questions: see which jewelers AI names and what it gets wrong about you.

Choose 20 to 30 questions your buyers ask, split between bridal, gifting and everyday pieces, and run them across the main assistants and Google several times. Record who is named, how your brand is described (lab-grown or natural, certification, price, store status) and which sources are cited. Do it before your peak season, not during it. If you want help before the engagement season, [get in touch with us](https://underneath.agency/contact): we will map how AI assistants describe your brand today and set out the facts, coverage and product-data work most likely to bring more showroom appointments and online orders when gift and engagement demand peaks. Our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page explains how a jeweler’s certification, sourcing and store facts are corrected at the source and then tracked across each assistant.

## Frequently asked questions

### Do AI assistants recommend lab-grown or natural diamonds?

We found no public evidence that they lean either way; a reasonable expectation is that they follow the question asked. With 61% of engagement rings now lab-grown in The Knot’s data, a jeweler should state clearly what it sells and how each stone is graded.

### Can a small independent jeweler be named by AI?

Sometimes, for local and specific questions, but national “best place to buy” answers favor a few online brands in the Rings.com data. An independent jeweler can use the tactics in [how small brands get recommended by AI](https://underneath.agency/resources/how-small-brands-get-recommended-by-ai).

### Does ChatGPT let shoppers try on jewelry?

OpenAI documents a “Try on” button for clothes and accessories that shows how an item could look on the shopper, and warns that try-on images may not represent the product exactly.

### Why do AI answers about our brand change month to month?

Answers vary between runs and assistants, and sources change; our explainer on [why AI answers about a brand shift](https://underneath.agency/resources/why-ai-answers-about-your-brand-change) shows how to track that over a gifting season.

## Sources

- JCK, reporting The Knot Real Weddings Study 2026 (2026), [Lab diamonds and the engagement market](https://www.jckonline.com/editorial-article/lab-diamonds-engagement-market/)
- INSTORE, reporting The Knot Real Weddings Study 2026 (2026), [What 10,000 couples who just got married can tell you about your business](https://instoremag.com/what-10000-couples-who-just-got-married-can-tell-you-about-your-business/)
- National Retail Federation (2026), [Valentine’s Day spending expected to reach new records](https://nrf.com/media-center/press-releases/valentine-s-day-spending-expected-to-reach-new-records)
- Rings.com (September 30, 2026), [Which Jewelry Brands Does AI Recommend? A Monthly Study](https://rings.com/pages/ai-jewelry-recommendations-study)
- National Jeweler (August 2026), [Brilliant Earth’s Q2 sales up 6%, driven by high-income shoppers](https://nationaljeweler.com/articles/15203-brilliant-earth-s-q2-sales-up-6-driven-by-high-income-shoppers)
- Signet Jewelers (March 9, 2026), [Signet Jewelers Reports Preliminary Results for Fourth Quarter and Full Year Fiscal 2026](https://s26.q4cdn.com/755441662/files/doc_news/Signet-Jewelers-Reports-Preliminary-Results-for-Fourth-Quarter-and-Full-Year-Fiscal-2026-2026.pdf)
- Rapaport, reporting The Plumb Club 2025 study (2025), [Online Jewelry Shopping on the Rise, Plumb Club Study Shows](https://rapaport.com/news/online-jewelry-shopping-on-the-rise-plumb-club-study-shows/)
- INSTORE, reporting The Plumb Club 2025 study (2025), [New research by The Plumb Club reveals consumers more cautious and want assurances when buying jewelry](https://instoremag.com/new-research-by-the-plumb-club-reveals-consumers-more-cautious-and-want-assurances-when-buying-jewel/)
- TechCrunch, reporting Adobe data (April 16, 2026), [AI traffic to US retailers rose 393% in Q1, and it’s boosting their revenue too](https://techcrunch.com/2026/04/16/ai-traffic-to-us-retailers-rose-393-in-q1-and-its-boosting-their-revenue-too/)
- OpenAI (2026), [Shopping with ChatGPT search](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search)
- FashionUnited (2026-09-29), [US consumers would let AI agents buy clothes, but not without a say](https://fashionunited.com/news/retail/us-consumers-would-let-ai-agents-buy-clothes-but-not-without-a-say/2026092974919)
- Google (March 5, 2025), [Expanding AI Overviews and introducing AI Mode](https://blog.google/products/search/ai-mode-search/)

---

This is the Markdown twin of https://underneath.agency/resources/jewelry-brands-ai-product-recommendations. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How legal software companies win customers through AI search"
description: "By being named, accurately, when lawyers ask AI which tools fit their practice, with public proof of accuracy, confidentiality and ethics that buyers can check."
canonical: "https://underneath.agency/resources/legal-software-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How can legal software companies win law firms and in-house teams through AI search?

By being named, and described accurately, when lawyers and legal operations leaders ask AI assistants which practice management, contract, research or legal AI tool fits their work. Lawyers already use these assistants heavily, and they buy with an unusual set of fears: wrong answers, leaked client data and ethics complaints. The vendors that win are those whose accuracy, confidentiality and track record are documented in public, where both buyers and AI assistants can check them.

## The short version

1. Lawyers already work in AI: in [Wolters Kluwer’s 2024 survey](https://www.wolterskluwer.com/en/know/future-ready-lawyer-2024) of 712 lawyers, 76% of those in corporate legal departments and 68% in law firms used generative AI at least once a week.
2. ChatGPT leads their own tool lists: in the [American Bar Association’s 2024 survey](https://www.americanbar.org/groups/law_practice/resources/tech-report/2024/2024-artificial-intelligence-techreport/), 52.1% of attorneys whose firms had adopted or were considering AI research tools named ChatGPT, ahead of CoCounsel (26.0%) and Lexis+ AI (24.3%).
3. The prize is large: [Harvey](https://www.harvey.ai/blog/harvey-raises-dollar550m-at-a-dollar155b-valuation-to-help-legal-teams-own-their-intelligence), a legal AI company, raised $550 million at a $15.5 billion valuation in September 2026 and says 80% of Am Law 100 firms use it.
4. Accuracy is the deciding fear: a [Stanford study](https://hai.stanford.edu/news/ai-trial-legal-models-hallucinate-1-out-6-or-more-benchmarking-queries) found general chatbots hallucinated on 58% to 82% of legal queries, and a [public database](https://www.damiencharlotin.com/hallucinations/) lists 2149 court cases involving AI-hallucinated material.
5. Ethics rules shape every purchase: the ABA’s [Formal Opinion 512](https://www.americanbar.org/news/abanews/aba-news-archives/2024/07/aba-issues-first-ethics-guidance-ai-tools/) says lawyers using generative AI must weigh duties of competence, confidentiality, client communication and reasonable fees.

## Which law firms and legal teams buy software, and what is each one worth?

Three very different buyers: small firms, large firms and in-house legal departments, each with its own budget and process.

**Small and mid-sized firms** buy practice management, billing, document and intake tools, usually without an IT department. The [ABA’s 2024 practice management report](https://www.americanbar.org/groups/law_practice/resources/tech-report/2024/2024-practice-management-techreport/) found 65% of firms budget for technology, rising from 41% of solo respondents to 90% of firms with 100 or more attorneys, and an average annual technology spend of $13,991. Each customer is small, but there are many, and subscriptions renew.

**Large firms** buy legal research, document review, e-discovery and legal AI platforms through innovation teams, partners and security reviews. Their deals are larger and slower. Harvey’s claim that 80% of Am Law 100 law firms use it shows how concentrated and competitive the top of this market has become.

**In-house legal teams** buy contract management, matter management, e-billing and legal AI, often alongside procurement and IT. They are the heaviest AI users in Wolters Kluwer’s survey: 76% used generative AI weekly.

Investors are betting on the category. Harvey’s September 2026 round of $550 million at a $15.5 billion valuation is a self-reported figure, but it shows how much revenue the market expects legal AI to earn.

## Where do AI assistants sit in how lawyers find software?

Already inside lawyers’ daily work, though no study isolates AI use for choosing legal software.

Lawyers know these tools first-hand. In the [ABA’s 2024 AI survey](https://www.americanbar.org/groups/law_practice/resources/tech-report/2024/2024-artificial-intelligence-techreport/) of 512 attorneys, 30.2% said their offices were using AI-based tools. Among AI research tools adopted or seriously considered, ChatGPT led at 52.1%.

How lawyers learned about new technology in that survey is telling:

| Source for learning about new technology such as AI | Share of attorneys |
|---|---|
| CLE seminars or webinars | 60.9% |
| Publications | 36.7% |
| Legal news | 34.3% |
| Other law firms | 31.9% |
| Google | 25.2% |

That survey predates the spread of AI answers in search. Our inference is that the same sources, legal publications, legal news and peer firms, now also reach lawyers through AI answers that cite them.

Software buyers in general start more research with AI. [G2](https://company.g2.com/news/g2-research-the-answer-economy), a review platform, asked 1,076 software buyers in March 2026 where their research begins; 51% said they start with an AI chatbot more often than with Google. G2 sells visibility to software vendors, and its sample was not drawn from law firms or legal departments.

On Google, AI answers appear on most software searches. Across the 1,248 US searches in [our AI Overview study](https://underneath.agency/research/ai-overviews-frequency-study), the B2B software and technology group, where legal tools belong, drew an AI Overview (the AI summary above Google’s results) on 96.0% of searches.

## Which questions do legal buyers ask AI assistants?

Questions about fit, accuracy, confidentiality and ethics, as much as features. We wrote the example prompts below ourselves to show the pattern; none were collected from real lawyers.

| Buyer | Illustrative prompt |
|---|---|
| Small firm | “Best practice management software for a three-lawyer immigration firm that bills flat fees?” |
| Small firm | “Clio vs MyCase for a family law practice: which handles trust accounting better?” |
| Large firm | “Which legal AI tools are used by Am Law 100 firms for due diligence review?” |
| Large firm | “Which AI legal research tools have the lowest hallucination rates in independent tests?” |
| In-house | “What contract lifecycle management tools suit a 10-person legal team on Salesforce?” |
| Any | “Does this legal AI vendor train its models on client data?” |
| Any | “Can I use a generative AI tool with privileged documents under ABA Formal Opinion 512?” |

To answer a prompt like the trust-accounting comparison, assistants can search the web. Google says AI Overviews and AI Mode [may use a “query fan-out” technique](https://developers.google.com/search/docs/appearance/ai-features) for such questions, and OpenAI says [ChatGPT search rewrites a question](https://help.openai.com/en/articles/9237897-chatgpt-search) into targeted queries. [Our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study) found ChatGPT running a search aimed at a named publication, ranking or award in 43.8% of its answers. For legal software, we infer, those would include legal technology press, bar association resources and independent evaluations.

## How does an AI answer become a legal software subscription or contract?

Through trials for small firms and through pilots and security reviews for large firms and legal departments.

**Small firms: answer, trial, subscription.** A managing partner asks which practice management tool fits their practice area and size, gets a few names, and starts a trial. With no procurement team, the AI answer carries much of the shortlist, we infer.

**Large firms: answer, pilot, firm-wide rollout.** An innovation lead or partner researches legal AI tools, then runs a pilot that must pass security and ethics review. The answer decides who gets the pilot; the pilot decides the contract.

**In-house teams: answer, evaluation, enterprise contract.** A general counsel or legal operations lead frames options for contract or matter management, then runs a formal selection with procurement and IT.

Vendor diligence by lawyers is thinner than you might expect. In the [ABA’s 2024 cloud computing report](https://www.americanbar.org/groups/law_practice/resources/tech-report/2024/2024-cloud-computing-techreport/), only 23% of respondents said that, as a precaution when using cloud tools, they had evaluated the vendor company’s history, and 19% sought peer advice. If buyers check little themselves, what an AI answer says about you matters more, we infer.

The revenue will rarely appear as a tracked click. For ways to tie AI answers to signed firms and legal departments, see [what lost clicks mean for pipeline](https://underneath.agency/resources/ai-answers-pipeline-revenue).

## Why would an AI assistant name one legal tech vendor over another?

No platform says how it picks; studies point to independent sources, and lawyers add accuracy, confidentiality and ethics.

**Documented by the platforms.** Google and OpenAI both describe their AI answers searching the web and citing what they find. Neither explains how it chooses which legal software vendors to put in front of a lawyer.

**Observed in studies.** [Chen and colleagues](https://arxiv.org/abs/2509.08919) looked at US software questions, the nearest available proxy for legal tech, and found AI search drew 72.7% of its sources from earned media (independent reviews and publications), against 45.4% for Google, whose results leaned more on vendor sites.

**What legal buyers weigh.** Three trust factors are specific to this market:

- **Accuracy.** Stanford RegLab and HAI researchers tested leading legal research AI tools on over 200 legal queries and found that even they hallucinate in [1 out of 6 (or more) benchmarking queries](https://hai.stanford.edu/news/ai-trial-legal-models-hallucinate-1-out-6-or-more-benchmarking-queries): Lexis+ AI and Ask Practical Law AI gave incorrect information more than 17% of the time and Westlaw’s AI-Assisted Research more than 34%, still far less often than general chatbots. The study dates from 2024, and the vendors may have improved their tools since. In the ABA survey, accuracy was the top concern about AI tools, named by 74.7% of respondents.
- **Confidentiality.** In the ABA cloud report, confidentiality and security were the top concerns, cited by approximately 55% of respondents.
- **Ethics.** Formal Opinion 512 ties generative AI use to the model rules on competence, confidentiality, communication and fees. Stanford notes that by May 2024 more than 25 federal judges had issued standing orders on AI use in their courtrooms.

**Our inference.** Lawyers trust what other lawyers, courts and independent testers say. Assistants appear to lean on independent sources too. A reasonable expectation is that legal vendors with published security documentation, honest accuracy testing and coverage in legal publications give both readers more to work with. Nobody has yet tested this with practice management, research or contract tools.

## What does it cost a legal software vendor to be missing or misdescribed?

Lost trials and pilots, and in this market, reputational risk from wrong claims. Nobody has yet priced that loss for legal tech.

- **Small firms decide fast.** Without procurement teams, a firm that gets three names from an assistant may trial only those, we infer.
- **The top of the market is concentrated.** With one vendor reporting use at 80% of Am Law 100 firms, a challenger missing from AI answers starts further behind.
- **Wrong security or accuracy claims are costly.** If an assistant says your product trains on client data when it does not, or misstates your certifications, you may be ruled out before a demo. Our guide on [how to fix wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) covers corrections like these.
- **One answer is not a verdict.** Ask the same practice management question twice and the list may change: in [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), only 25.2% of the brands ChatGPT named for a question appeared in all five runs.

## What does GEO involve for a legal tech vendor?

It puts your accuracy, security and ethics evidence where AI assistants can find and quote it. No vendor can be promised a recommendation.

Generative engine optimization (GEO) means working to be mentioned, and described correctly, when AI assistants answer questions. For a legal software company it covers:

1. **Clear identity by buyer.** Separate pages for solo and small firms, large firms and in-house teams, stating practice areas, firm sizes and systems you integrate with.
2. **Public trust documentation.** Security certifications, data handling, whether client data trains any model, and data residency, on plain web pages, not only in sales decks. Confidentiality was the top cloud concern for lawyers.
3. **Honest accuracy evidence.** Publish how you test accuracy, take part in independent evaluations, and say plainly what the tool should not be used for. Overclaiming invites the scrutiny Stanford applied to “hallucination-free” marketing.
4. **Ethics alignment.** Explain how your product supports competence, confidentiality, supervision and billing under Formal Opinion 512 and state bar guidance.
5. **Legal press, CLE and peers.** Coverage in legal publications, CLE sessions and customer stories from named firms are the sources lawyers already learn from. Our piece on [how brands build authority for AI search](https://underneath.agency/resources/how-brands-build-authority-for-ai-search) explains why those third-party sources carry weight.
6. **Repeated checks for each buyer.** Ask the small-firm, large-firm and in-house questions several times each in ChatGPT, Google AI Overviews and AI Mode, Gemini, Perplexity, Copilot and Claude, and track the answers.

For the wider software picture, see [how B2B SaaS companies earn revenue from AI search](https://underneath.agency/resources/b2b-saas-revenue-from-ai-search), and for the limits of tactics, [legitimate GEO versus manipulation](https://underneath.agency/resources/legitimate-geo-vs-manipulation). Law firms are not the only buyers who check a vendor’s handling of sensitive data before they buy: school districts run privacy checks too, as [how EdTech companies win schools through AI](https://underneath.agency/resources/edtech-ai-search) shows.

## What don’t we know yet about AI search in legal software buying?

Nobody has yet shown that appearing in AI answers wins more law firm subscriptions or legal department contracts.

- **No study of legal buyers choosing software through AI.** The ABA and Wolters Kluwer surveys measure AI use in legal work, not in vendor selection.
- **Dated and interested sources.** The ABA figures are from 2024, before AI answers spread in search. Harvey’s figures are its own; Wolters Kluwer and G2 sell to this market.
- **Accuracy tests age quickly.** The Stanford study tested 2024 versions of legal research tools.
- **Revenue effects are the thinnest research.** Across the 45 studies of AI search optimization reviewed by [Martinez](https://arxiv.org/abs/2607.14035), traffic and conversions were the least supported area, the very outcome a vendor counting trials and contracts cares about.

## How should a legal tech vendor test what AI tells lawyers about it?

Ask the assistants what lawyers ask before a trial or pilot, then check how they describe your safeguards.

Draft questions for solo and small firms, large firms and in-house teams covering practice fit, alternatives, accuracy, confidentiality and ethics. Put each one to ChatGPT, Google, Gemini, Perplexity, Copilot and Claude more than once, since answers vary between runs. Note which vendors appear, which legal publications or bar resources are cited, and whether your security posture, model training policy and integrations come out right. Where an answer is wrong or thin, the cause is usually proof you have not yet published.

To have us run that test with you, [ask us for a review of your visibility to law firm and legal department buyers](https://underneath.agency/contact). We will show where AI answers place you for small firms, large firms and legal departments, and which missing public proof is most likely costing you trials, pilots and contracts. Publishing that trust, accuracy and ethics evidence, then checking how assistants repeat it, is the work our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page lays out for legal tech vendors.

## Frequently asked questions

### Do lawyers use ChatGPT to choose legal software?

No study measures that directly. But ChatGPT was the top AI research tool firms had adopted or considered in the ABA’s 2024 survey, at 52.1%.

### Does legal AI accuracy affect whether assistants recommend a tool?

Unknown, but it affects buyers. Stanford found even leading legal research tools hallucinated in at least 1 out of 6 benchmark queries in 2024.

### Should legal software vendors publish their security details?

Yes, in plain web pages. Confidentiality and security were the top cloud concerns for about 55% of attorneys in the ABA survey.

### Do ethics rules matter for legal software marketing?

Yes. ABA Formal Opinion 512 links generative AI use to competence, confidentiality, communication and fees, so buyers will ask how your tool supports them.

### Is AI visibility more important for small-firm software?

Probably. Small firms rarely run formal selections, and only 19% of attorneys in the ABA cloud survey sought peer advice as a precaution with cloud tools.

## Sources

- Wolters Kluwer (2024), [2024 Future Ready Lawyer Survey Report](https://www.wolterskluwer.com/en/know/future-ready-lawyer-2024)
- American Bar Association (2025), [2024 Artificial Intelligence TechReport](https://www.americanbar.org/groups/law_practice/resources/tech-report/2024/2024-artificial-intelligence-techreport/)
- American Bar Association (2025), [2024 Practice Management TechReport](https://www.americanbar.org/groups/law_practice/resources/tech-report/2024/2024-practice-management-techreport/)
- American Bar Association (2025), [2024 Cloud Computing TechReport](https://www.americanbar.org/groups/law_practice/resources/tech-report/2024/2024-cloud-computing-techreport/)
- American Bar Association (2024-07-29), [ABA issues first ethics guidance on a lawyer’s use of AI tools](https://www.americanbar.org/news/abanews/aba-news-archives/2024/07/aba-issues-first-ethics-guidance-ai-tools/)
- Harvey (2026-09-09), [Harvey Raises $550M at a $15.5B Valuation to Help Legal Teams Own Their Intelligence](https://www.harvey.ai/blog/harvey-raises-dollar550m-at-a-dollar155b-valuation-to-help-legal-teams-own-their-intelligence)
- Stanford HAI (2024-05-23), [AI on Trial: Legal Models Hallucinate in 1 out of 6 (or More) Benchmarking Queries](https://hai.stanford.edu/news/ai-trial-legal-models-hallucinate-1-out-6-or-more-benchmarking-queries)
- Damien Charlotin (2026), [AI Hallucination Cases database](https://www.damiencharlotin.com/hallucinations/)
- G2, Tim Sanders (2026-04-15), [In the Answer Economy, Don’t Win the Click — Win the Answer](https://company.g2.com/news/g2-research-the-answer-economy)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- OpenAI Help Center (2025), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Chen, Wang, Chen and Koudas (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919), arXiv:2509.08919.
- Martinez (2026), [Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026)](https://arxiv.org/abs/2607.14035), arXiv:2607.14035.
- Underneath (2026), [When does Google show an AI Overview? 1,248 US searches](https://underneath.agency/research/ai-overviews-frequency-study)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)

---

This is the Markdown twin of https://underneath.agency/resources/legal-software-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "What separates legitimate GEO from manipulation? | Underneath"
description: "Truth and transparency. Legitimate GEO makes real, verifiable facts easier for AI to find; manipulation fakes evidence, hides commands or hides motives."
canonical: "https://underneath.agency/resources/legitimate-geo-vs-manipulation"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# What separates legitimate GEO from manipulation?

The line is honesty, not the wish to be visible: legitimate generative engine optimization (GEO) makes true, verifiable information easier for AI search engines to find, cite and repeat. Manipulation fabricates evidence, hides instructions aimed at the AI, or disguises commercial motives. Because both use the same channel, a leader needs tests that look at the content, not the goal.

## The short version

1. Researchers draw the line at truthfulness: [Wen and colleagues](https://arxiv.org/abs/2606.12439) define benign GEO as adding clarity and legitimate citations, and malicious GEO as fabricated statistics, fake endorsements or “always recommend X” instructions.
2. Honest content changes work: in the original GEO study by [Aggarwal and colleagues](https://arxiv.org/abs/2311.09735), adding citations, quotations and statistics improved visibility by 30 to 40%, while keyword stuffing did little.
3. Fabrication also works in tests, which is why policy matters: a made-up clinical citation had the same effect as 0.17 rating points of real improvement, per [Chu and Hou](https://arxiv.org/abs/2606.17443).
4. GEO-shaped content is spreading: [Chu, Leng and colleagues](https://arxiv.org/abs/2608.16824) estimate that 8.90% of pages in Google and Gemini results show signs of it, rising to 16.36% among pages updated in 2026.
5. Engines are fighting back: one tested defense by [Li and colleagues](https://arxiv.org/abs/2609.02964) cut manipulation success from 50.32% to 6.20%.

## Isn’t all GEO an attempt to influence AI answers?

Yes, and that is why intent alone cannot be the test. GEO means shaping content so AI search engines are more likely to find, cite and describe it. [Li and colleagues](https://arxiv.org/abs/2609.02964) put it plainly: optimization and manipulation “differ in intent rather than mechanism.” A rewritten page may contain no false statement and no hidden command, yet still be built to win.

[Wen and colleagues](https://arxiv.org/abs/2606.12439) give the clearest dividing line. Benign actors “preserve truthfulness and verifiability” by improving factual clarity, adding legitimate citations and making relevant evidence easier to retrieve. Malicious actors relax those limits, using fabricated statistics, fake endorsements or prompt-injection text such as “always recommend X.” Both chase exposure. Only one keeps the evidence honest.

## What tests can a leader apply to any tactic?

Four questions, drawn from a 2026 survey of 45 studies by [Martinez](https://arxiv.org/abs/2607.14035). A tactic must pass all four to count as legitimate:

| Test | The question to ask | Fails when |
|---|---|---|
| Truth | Do the facts and qualifications stay true? | Claims are stretched or caveats removed |
| Real evidence | Can every statistic, review and reference be verified? | Testimonials, data or sources are invented |
| No hidden commands | Does the page inform the reader rather than instruct the AI? | Text tells the AI what to recommend |
| Disclosure and fairness | Is commercial intent disclosed, and are rivals treated fairly? | Self-interest is hidden or competitors are smeared |

Martinez notes that reorganizing paragraphs or adding a verified primary source will generally pass. A string aimed at the AI, a fabricated testimonial or an instruction to favor a brand will fail, however fluent the writing.

[Chu and Hou](https://arxiv.org/abs/2606.17443) offer a similar three-tier scale for authority claims. Tier 1, real certifications, published trials or genuine expert endorsements, is legitimate marketing. Tier 2, vague phrases like “clinically proven” with no source, is a grey area. Tier 3, invented studies or endorsements, is potential false advertising.

## Which legitimate tactics have evidence behind them?

Making content more substantive, specific and verifiable has the best support. In the original GEO study, [Aggarwal and colleagues](https://arxiv.org/abs/2311.09735) found that citing sources, adding quotations and adding statistics achieved a relative improvement of 30 to 40% in their main visibility measure. On Perplexity, a live engine, gains reached up to 37%. Keyword stuffing, a classic search trick, offered little to no improvement.

Those gains come with a caveat. Martinez points out that they apply to a page already chosen by the engine. In one end-to-end test he summarizes, rewriting only the body of pages reduced their presence in the top 10 after re-ranking by 16%. A rewrite that helps a page once it is read can make it harder to find in the first place.

A shopping study points the same way. [Bagga and colleagues](https://arxiv.org/abs/2511.20867) found that the best automated rewrites kept a rule to preserve facts. When the AI ranking the products was told to flag manipulative listings, rank gains depended on genuine content improvement.

## How common is questionable GEO on the web today?

Measurable and growing, though most of it is not proven fraud. [Chu, Leng and colleagues](https://arxiv.org/abs/2608.16824) built a detector for GEO-optimized pages and ran it on Google Search and Gemini results for 1,000 real user queries. Of 10,095 pages, they estimate 8.90% were GEO-optimized, rising to 16.36% among pages last modified in 2026. On Amazon, the rate reached 20.37%.

Detection is not the same as wrongdoing. But the citations inside those detected pages were often weak: 69.34% of citation occurrences got a low verifiability label from the authors’ automated check.

Openly self-serving content is common too. In [our self-ranking lists study](https://underneath.agency/research/self-promoting-best-lists-study), 24.2% of AI-cited “best X” lists with an identifiable publisher ranked that publisher first. That is not manipulation by itself, but it fails the disclosure test when the self-interest is not stated.

Hidden commands are rarer. [Khodayari and colleagues](https://arxiv.org/abs/2604.27202) scanned 1.2 billion web addresses and confirmed 1,521 hidden instructions aimed at reputation, such as demands for positive reviews or forced citations. Our guide on [websites hiding instructions for AI](https://underneath.agency/resources/websites-hiding-instructions-for-ai-search) covers what that scan found.

## Does manipulation pay off even when it works?

Rarely for long, because engines, rivals and regulators all push back. Engines are building defenses. In tests by [Li and colleagues](https://arxiv.org/abs/2609.02964), a two-stage defense cut manipulation success from 50.32% to 6.20% while keeping 94.12% of the honest evidence the AI used. Whether live engines use such filters yet is covered in [whether AI search can filter manipulative GEO](https://underneath.agency/resources/can-ai-search-filter-manipulative-geo).

Simple warnings also change the game. In the shopping study by [Bagga and colleagues](https://arxiv.org/abs/2511.20867), an automated attacker was told to rank high without being flagged. Its mean flag rate fell by 60 percentage points, and its rewrites drifted to careful, fact-grounded prose. Under that defense, honesty was the winning strategy.

Copying wears gains away. In tests by [Chu and Hou](https://arxiv.org/abs/2606.17443), the first brand to use authority-style copy gained a payoff of +0.802 in their measure. When every brand did the same, it fell to +0.007. Brands that did nothing got zero recommendations in those tests, which is why doing nothing is not safe either. For the risk that a rival games the answers, see [how competitors try to game AI recommendations](https://underneath.agency/resources/can-competitors-game-ai-recommendations).

## What should you do about it?

Write a short GEO policy and hold agencies and in-house teams to it. Practical steps:

1. Adopt the four tests above as your rule: true, verifiable, no hidden commands, disclosed.
2. Ban specific tactics in writing: invented statistics or reviews, fake endorsements, hidden text, instructions aimed at AI, and posting fake user content.
3. Require a source for every number and claim in content written for AI visibility, and keep a record.
4. Disclose self-interest in comparison content, such as a “best X” list that includes your product.
5. Ask any agency to show which tactics it uses and how it measures results across repeated AI answers.
6. Review a sample of published content each quarter against the policy.

If you want a partner that works within these limits, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The research gives clear principles but no settled legal or platform rulebook for AI answers. Specifically:

- The benign and malicious distinction comes from position papers and surveys, not from regulation.
- Most evidence on what works comes from controlled tests, not long-term results on live engines.
- Martinez found no reviewed technique with a stable, long-term, cross-platform effect on being discovered organically.
- Detectors of GEO content are new; detection does not prove intent or harm.
- Tier 2 claims, such as “clinically proven” with no source, remain a grey area that no study resolves.

## Frequently asked questions

### Is generative engine optimization ethical?

It can be. Researchers class GEO as benign when it preserves truth and verifiability, and as malicious when it uses fabricated statistics, fake endorsements or hidden instructions.

### Is adding statistics and quotes to content manipulation?

Not if they are real and sourced. In the original GEO study, adding statistics and quotations improved visibility by 30 to 40% on the main measure, but invented numbers would fail the real-evidence test.

### Is it manipulation to put hidden text on a page for AI assistants?

Yes, if the text tells the AI what to recommend. One scan found 1,521 hidden instructions aimed at reputation, and researchers treat such commands as a clear sign of manipulation.

### How do I know if my GEO agency is using manipulative tactics?

Ask for its tactics in writing and check sample content against four tests: true, verifiable, no hidden commands and disclosed interest. Any invented statistic, review or endorsement is a red flag.

## Sources

- Wen, Y., Zhang, N., Yuan, H., Chen, X., Zhang, H. and Guo, H. (2026), [Position: Generative Engine Optimization Creates Underexamined Risks, Governance Must Target Concentration, Disclosure, and Academic Blind Spots](https://arxiv.org/abs/2606.12439), arXiv:2606.12439.
- Martinez, O. (2026), [Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026)](https://arxiv.org/abs/2607.14035), arXiv:2607.14035.
- Aggarwal, P., Murahari, V., Rajpurohit, T., Kalyan, A., Narasimhan, K. and Deshpande, A. (2024), [GEO: Generative Engine Optimization](https://arxiv.org/abs/2311.09735), arXiv:2311.09735.
- Chu, X. and Hou, Y. (2026), [Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems](https://arxiv.org/abs/2606.17443), arXiv:2606.17443.
- Chu, J., Leng, Y., Li, M., Shen, Y., Shen, X. and Zhang, Y. (2026), [GEO-Flag: Detecting and Measuring GEO-Optimized Web Content](https://arxiv.org/abs/2608.16824), arXiv:2608.16824.
- Li, H., Shao, Y., Lin, X., Guan, Z., Zhou, M. and Shi, J. (2026), [When Optimization Becomes Manipulation: Defending Generative Search against Malicious Generative Engine Optimization](https://arxiv.org/abs/2609.02964), arXiv:2609.02964.
- Bagga, P. S., Farias, V. F., Korkotashvili, T., Peng, T. and Wu, Y. (2025), [E-GEO: A Testbed for Generative Engine Optimization in E-Commerce](https://arxiv.org/abs/2511.20867), arXiv:2511.20867.
- Khodayari, S., Zhang, X., Acharya, B. and Pellegrino, G. (2026), [Indirect Prompt Injection in the Wild: An Empirical Study of Prevalence, Techniques, and Objectives](https://arxiv.org/abs/2604.27202), arXiv:2604.27202.
- Underneath (2026), [How many “best of” lists cited by AI rank their own brand first?](https://underneath.agency/research/self-promoting-best-lists-study)

---

This is the Markdown twin of https://underneath.agency/resources/legitimate-geo-vs-manipulation. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How lending platforms reach borrowers who ask AI about loans"
description: "By being named, with accurate rates, terms and reputation, when borrowers ask AI where to get a loan. Visibility brings applications, never approvals."
canonical: "https://underneath.agency/resources/lending-platforms-borrowers-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How do lending platforms reach borrowers who ask AI where to get a loan?

By making sure AI assistants describe their loans, costs, eligibility and reputation accurately when people and small businesses ask where to borrow. Online lenders are taking a growing share of personal and small-business lending, and the first comparison often happens in a conversation, not on a rate table. Visibility can bring qualified applications; it can never promise anyone a loan.

## The short version

1. Online lenders lead personal-loan growth: [TransUnion](https://media.transunion.com/content/dam/transunion/us/business/collateral/report/1-fs/q2-2026-ciir-consumer-lending-report.pdf) reports fintech lenders’ share of unsecured personal loan originations rose from 39.8% to 45.0% in a year, with balances at a record $281 billion.
2. Small businesses are moving online too: in the Federal Reserve Banks’ [Small Business Credit Survey](https://www.fedsmallbusiness.org/-/media/project/clevelandfedtenant/fsbsite/reports/2026/2026-report-on-employer-firms/2026-report-on-employer-firms.pdf), the share of applicants that sought financing at online lenders rose from 17% to 29% over five years.
3. Cost surprises are the trust gap: 60% of firms that borrowed from online lenders said costs were higher than expected, against 37% at small banks and 32% at large banks.
4. Approval is never certain: the [Federal Reserve’s household survey](https://www.federalreserve.gov/publications/files/2024-report-economic-well-being-us-households-202505.pdf) found one-third of 2024 credit applicants were denied or approved for less than they requested.
5. Answers depend on the market: in [our country study](https://underneath.agency/research/ai-recommendations-by-country-study), questions such as lending, insurance and tax showed the largest country effect on the brands ChatGPT named.

A note before you read: this article is about how lenders appear in AI answers. Nothing here is financial, legal or compliance advice.

## Who borrows from lending platforms, and what is a funded loan worth?

Consumers consolidating debt or funding a purchase, and small businesses seeking working capital; value comes from funded loans, not clicks.

Lending platforms serve two very different borrowers. Consumers take unsecured personal loans, often to pay down credit cards. TransUnion counts personal loan balances at $281 billion, up 9.6% in a year, and originations up 19.5%. Fintech lenders drove that growth, with their originations up 35.0%, while banks’ share fell from 13.1% to 10.8%. Subprime borrowers made up about 38% of new loans, a reminder that many of these customers are stretched. Home loans follow their own path, covered in [how mortgage lenders win borrowers through AI](https://underneath.agency/resources/mortgage-platforms-leads-ai-search).

Small businesses borrow for cash flow, equipment and expansion. The [Small Business Credit Survey](https://www.fedsmallbusiness.org/-/media/project/clevelandfedtenant/fsbsite/reports/2026/2026-report-on-employer-firms/2026-report-on-employer-firms.pdf), with 6,525 responses from employer firms, found that large banks remain the top place firms apply, but that applicants at online lenders, such as OnDeck in the survey’s own examples, rose to 29%. Those applicants prioritized speed and their expected chance of being funded, while bank applicants chose on existing relationships. For the payments side of small-business finance, see [how payment processors reach merchants through AI](https://underneath.agency/resources/payment-processors-merchant-demand-ai-search).

Comparison marketplaces such as LendingTree, NerdWallet and Credit Karma sit between the two: they earn from referring borrowers to lenders. For a lender, the commercial value of a borrower is a funded loan that performs; for a marketplace, it is a qualified referral.

## Where does AI already sit in a borrower’s decision?

In early research, where many consumers and small businesses already use AI, though usually as a starting point.

An [Intuit Credit Karma survey](https://stacker.com/stories/personal-finance-investing/rise-fin-ai-why-americans-are-trusting-generative-ai-their) found that 66% of Americans who had used generative AI had used it to seek financial advice, and 85% of those had acted on what it said. The same survey adds a warning: 52% of those who acted said they had made a poor financial decision or mistake based on it, and 80% said they still research and validate the advice first. Credit Karma runs a lending marketplace, so read this as vendor research.

Small businesses use AI widely too. The Small Business Credit Survey found 46% of firms said the business or its employees currently use AI, most often for writing or marketing. We found no public survey measuring how many borrowers ask an AI assistant which lender to choose.

The Federal Reserve’s household survey frames the stakes. Thirty-four percent of adults applied for some type of credit in 2024, and twenty-one percent reported experiencing financial fraud or scams involving their money. Borrowers asking AI “is this lender legit?” have good reason to ask.

## What do borrowers ask AI before applying for a loan?

What they can get, what it will cost, whether a lender is trustworthy and whether checking will hurt their credit. We wrote the borrower prompts in the table ourselves as examples; none comes from a real applicant.

| Need | Illustrative prompt |
|---|---|
| Debt consolidation | “Best way to consolidate $15,000 of credit card debt with a 680 credit score” |
| Cost | “What APR should I expect on a personal loan with fair credit?” |
| Process | “Does checking my rate with an online lender hurt my credit?” |
| Trust | “Is [lender] legit, and what do borrowers complain about?” |
| Small business | “Online lender or bank for a $50,000 line of credit?” |
| Speed | “How fast can a restaurant get a working capital loan without collateral?” |
| Marketplace | “Is it better to use a loan comparison site or go straight to a lender?” |

The small-business questions track the survey closely. Online-lender applicants said speed and the chance of being funded drove their choice, and high interest rates and unfavorable repayment terms were their most common complaints. An answer that explains speed without cost serves the borrower badly and, we infer, sets up the cost surprise the survey found.

## How does a borrower get from an AI answer to a funded loan?

Through a rate check and an application; the lender’s underwriting decides, so visibility shapes who applies, not approval.

The path, as we infer it from how online lending works:

1. **Question.** “Which lenders consolidate card debt for fair credit?”
2. **Answer.** The assistant names lenders or marketplaces and describes rates and requirements.
3. **Rate check.** The borrower checks a rate, often with a soft credit inquiry, on the lender’s site or a marketplace.
4. **Application.** Documents, verification and a credit decision.
5. **Decision.** Approved, approved for less, or declined; the household survey shows one-third of applicants get the second or third outcome.
6. **Funded loan.** Revenue for the lender over the loan’s life; a referral fee for a marketplace.

The useful measure is therefore not traffic but qualified applications: borrowers who arrive understanding the product, its cost and its eligibility. An AI answer that oversells approval odds or understates rates sends applicants who will be declined or disappointed. Our guide to [linking AI answers to pipeline and revenue](https://underneath.agency/resources/ai-answers-pipeline-revenue) explains measurement in general; for lenders, track applications and funded loans by source, not sessions.

## What decides whether an assistant names a lender?

Platforms say little; studies point to market-specific sources, independent coverage and reputation evidence, which matter more in lending.

**Documented by platforms.** Google says its AI features may use [“query fan-out”](https://developers.google.com/search/docs/appearance/ai-features), issuing related searches across subtopics. A question about consolidating debt with fair credit can pull in separate sources on rates, eligibility and individual lenders. We found no platform documentation on how assistants choose lenders.

**Observed in studies.**

- **The country matters most in lending.** In our country study, the gap between same-country and different-country answers on ChatGPT was 0.254 for high-dependence questions such as insurance, lending and tax, against 0.053 for global products.
- **Review sites settle the trust question.** Asked “is this brand legit?”, assistants in [our reputation study](https://underneath.agency/research/is-it-legit-ai-reputation-study) cited a review or complaint platform in 88.0% of answers; Trustpilot and the BBB together accounted for 61.7% of review-platform citations, and 99.7% of answers flagged at least one problem.
- **Independent mentions track with recommendations.** [Our brand study](https://underneath.agency/research/brand-entity-ai-recommendations-study) found that every tenfold rise in the number of independent sites naming a brand in the cited pages came with 4.7 times the odds of a recommendation, which matters for a lender that only its own site describes.

**Our inference for lenders.** Lending is often licensed state by state and priced by credit profile, so answers lean on sources that explain those details: comparison marketplaces, personal finance publishers and regulators. A lender whose rates, fees, eligibility and licensing are stated plainly gives those sources, and the assistants that read them, accurate material. Complaint patterns, especially about cost surprises, will surface in “is it legit?” answers.

## Which lending rules shape what a platform can say?

The same advertising and fair-lending rules that govern ads; content meant to shape AI answers needs the same compliance review.

We do not offer legal advice, but the rules any lender’s content team works within are public:

- **Truth in Lending advertising.** [Regulation Z](https://www.law.cornell.edu/cfr/text/12/1026.24) requires that an advertised rate of finance charge be stated as an “annual percentage rate,” using that term. Stating certain “triggering terms,” such as the number of payments or the amount of any payment, requires further disclosures.
- **Equal credit opportunity.** [Regulation B](https://www.law.cornell.edu/cfr/text/12/1002.4) bars statements “in advertising or otherwise” that would discourage a reasonable person from applying on a prohibited basis.
- **No approval promises.** Approval always depends on underwriting, and the household survey shows how often applicants are declined or offered less.

Our inference: a page written to be quoted by AI assistants is still marketing. Generative engine optimization (GEO) that inflates approval odds or hides costs is both a regulatory risk and the kind of tactic that can [backfire on a brand](https://underneath.agency/resources/can-geo-backfire-on-your-brand).

## What does GEO look like for a lending platform?

Making the facts borrowers and assistants check accurate, consistent, compliant and confirmed by independent sources.

1. **Rates and costs, plainly.** APR ranges, fees and representative examples on public pages, consistent everywhere, reviewed by compliance.
2. **Eligibility in the open.** Who the product suits and who it does not, including credit, income and state availability.
3. **Answers by need.** Pages for debt consolidation, home improvement, working capital or equipment, each explaining costs and alternatives honestly.
4. **Marketplace accuracy.** Correct, current listings on the comparison sites that assistants cite.
5. **Reputation work.** Answer complaints on Trustpilot and the BBB, especially about costs and repayment terms, and fix the causes.
6. **Licensing and identity facts.** State licenses, the legal entity, partner banks where relevant, and how to verify them, which helps answers to “is it legit?”
7. **Market-by-market checks.** Track answers in each state or country you lend in, because lending answers vary most by market.

Related guides cover [consumer fintech apps](https://underneath.agency/resources/consumer-fintech-apps-customers-ai-search), [fintech startups](https://underneath.agency/resources/fintech-startups-customers-ai-search) and [how to fix wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers). No one can guarantee that an assistant will recommend a lender, and no visibility work changes who qualifies for a loan.

## Which questions about AI and borrowing remain open?

How often borrowers choose a lender from an AI answer, and whether those borrowers apply and repay differently.

- **No direct measure.** We found no public data on how many borrowers pick a lender because an assistant named it.
- **Vendor surveys.** The consumer AI figures come from Credit Karma, which operates a lending marketplace.
- **Studies outside lending.** Our country study included lending questions, but our reputation and brand studies did not test lenders specifically.
- **Fair-lending questions.** We found no public research on whether AI answers describe loan options differently for different groups of borrowers, a question lenders’ compliance teams may want to watch.

## Where should a lending platform start?

Start by asking assistants your borrowers’ real questions, by product and market, and checking every rate and claim they repeat.

That first check shows which lenders and marketplaces are named for each need, whether your APR ranges, fees and eligibility come back correctly, how “is it legit?” answers describe you, and which sources they rely on. It also shows where answers imply approval odds you would never advertise.

If more qualified applications and funded loans are the goal, [ask us to review how assistants describe your loans](https://underneath.agency/contact). We will test the borrower questions that matter for your products and markets, find inaccurate or noncompliant descriptions, and plan, with your compliance team, the content, marketplace and reputation work that gives assistants accurate evidence. Lenders can read how that work is structured and checked market by market, without any promise about approvals, on our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page.

## Frequently asked questions

### Can GEO get more borrowers approved?

No. Visibility can influence who applies and how well they understand the product; approval depends entirely on the lender’s underwriting and the borrower’s situation.

### Do people really ask ChatGPT where to get a loan?

We found no direct measure. Credit Karma found 66% of generative AI users had used it for financial advice, and 80% of those who acted still validated it first.

### Why do AI answers about loans cite comparison sites?

They explain rates, eligibility and lenders side by side. Our country study also found lending answers vary by market, which favors sources specific to each market.

### Is content written for AI answers subject to lending advertising rules?

Treat it as marketing. Regulation Z’s disclosure rules and Regulation B’s ban on discouragement apply to advertising; your counsel should review any page meant to inform borrowers.

## Sources

- TransUnion (2026-08), [Q2 2026 Credit Industry Insights Report: Consumer Lending](https://media.transunion.com/content/dam/transunion/us/business/collateral/report/1-fs/q2-2026-ciir-consumer-lending-report.pdf)
- Federal Reserve Banks (2026), [Small Business Credit Survey: 2026 Report on Employer Firms](https://www.fedsmallbusiness.org/-/media/project/clevelandfedtenant/fsbsite/reports/2026/2026-report-on-employer-firms/2026-report-on-employer-firms.pdf)
- Board of Governors of the Federal Reserve System (2025-05), [Economic Well-Being of U.S. Households in 2024](https://www.federalreserve.gov/publications/files/2024-report-economic-well-being-us-households-202505.pdf)
- Intuit Credit Karma, via Stacker (2025), [The rise of fin-AI: Why Americans are trusting generative AI with their wallets](https://stacker.com/stories/personal-finance-investing/rise-fin-ai-why-americans-are-trusting-generative-ai-their)
- Legal Information Institute, Cornell Law School, [12 CFR § 1026.24 – Advertising (Regulation Z)](https://www.law.cornell.edu/cfr/text/12/1026.24)
- Legal Information Institute, Cornell Law School, [12 CFR § 1002.4 – General rules (Regulation B)](https://www.law.cornell.edu/cfr/text/12/1002.4)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- Underneath (2026), [Same question, four countries: do AI recommendations change?](https://underneath.agency/research/ai-recommendations-by-country-study)
- Underneath (2026), [“Is this brand legit?” How AI assistants build a reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study)
- Underneath (2026), [Do Wikipedia and schema make AI assistants recommend a brand?](https://underneath.agency/research/brand-entity-ai-recommendations-study)

---

This is the Markdown twin of https://underneath.agency/resources/lending-platforms-borrowers-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How 3PL and fulfillment providers win customers in AI search"
description: "3PL and fulfillment providers reach more shippers in AI search when locations, capabilities, integrations and outside proof are easy to verify."
canonical: "https://underneath.agency/resources/logistics-providers-shipper-contracts-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# When a shipper asks AI for a 3PL, does our warehouse network make the list?

It can, if an assistant can find and confirm where your buildings are, what you handle, which systems you connect to and who vouches for you. Shippers and ecommerce brands choose logistics partners rarely and keep them for years, so being on the first list of candidates matters more than any single click.

This guide is for third-party logistics providers (3PLs) that sell warehousing, fulfillment and contract logistics: ecommerce fulfillment centers, regional warehouse operators, cold storage providers and large contract logistics firms. It does not cover the software sold to logistics companies, which we discuss in our guide to logistics software in AI search. Carriers and brokers are covered in how freight providers win shippers through AI.

## The short version

1. The market is large and growing again: US 3PL gross revenue rose 5.0% to $323.4 billion in 2025, including $72.7 billion for value-added warehousing and distribution, according to [Armstrong & Associates](https://www.logisticsmgmt.com/article/u.s_3pl_revenues_see_strong_annual_gains_reports_armstrong_associates).
2. Winning customers is getting harder: in the [2025 Inbound Logistics 3PL Perspectives report](https://www.inboundlogistics.com/articles/2025-inbound-logistics-perspectives-3pl-market-research-report/), finding and retaining customers was 3PLs’ third-biggest challenge, cited by 46%, up 13 points in two years.
3. Contracts stick: in the [2026 Third-Party Logistics Study](https://www.mychesco.com/?p=642622), more than half of shippers said they do not rebid at contract end, and technology capability had become a decisive selection factor.
4. Ecommerce brands outsource as they grow: in a Radial survey of 200 retail decision-makers, [reported by CXM](https://cxm.world/?p=130037), 72% of brands with $100 million to $150 million in revenue relied on 3PLs, rising to 76% at $150 million to $200 million.
5. Logistics buyers already use generative AI at work: 96% of 616 shippers and logistics providers surveyed by Descartes said they use it for transportation management, [DC Velocity reported](https://www.dcvelocity.com/technology/descartes-transportation-management-benchmark-survey).

## Who hires a 3PL, and what is a contract worth?

Shippers and ecommerce brands, usually through a structured search, and a won account can last many years.

Two kinds of buyers dominate. **Shippers**, meaning manufacturers, retailers and distributors, hire 3PLs for warehousing, distribution, transportation management and value-added services such as kitting or returns. Their buyers are supply chain and logistics leaders with procurement alongside; across business purchases generally, [Forrester](https://www.forrester.com/press-newsroom/forrester-2026-the-state-of-business-buying/) finds procurement professionals are decision-makers in 53% of buying cycles. **Ecommerce brands** hire fulfillment providers to store inventory, pick, pack and ship orders, and connect to their store and marketplaces. In smaller brands, the founder or head of operations often runs the search. Those brands run a similar search for packaging, described in [how packaging suppliers win quote requests through AI](https://underneath.agency/resources/packaging-suppliers-quote-requests-ai-search).

The point at which brands outsource is visible. In Radial’s survey, 70% of fast-growing retail brands still handled fulfillment in-house, and 59% operated from a single facility, but outsourcing became the norm once brands passed $50 million in revenue. Sellers that still ship their own orders are covered in [how parcel carriers win small businesses through AI](https://underneath.agency/resources/shipping-companies-customers-ai-search). Ecommerce keeps growing underneath: the [US Census Bureau](https://www.census.gov/retail/mrts/www/data/pdf/ec_current.pdf) estimated second-quarter 2026 ecommerce sales at $340.2 billion, 17.1% of total retail sales and up 12.2% from a year earlier.

What a customer is worth depends on volume and service scope, and there is no public average. Large providers show the shape of the prize. [GXO](https://container-news.com/gxo-posts-record-revenue-in-q4-and-full-year-2025/), a contract logistics provider, reported 2025 revenue of $13.2 billion and more than $1 billion in new contracts for the third consecutive year, which it expects to add $774 million in incremental revenue in 2026. Because many shippers do not rebid at contract end, our inference is that each new account can be worth several years of revenue, not one season’s.

Capacity also shapes the market. National industrial vacancy fell to 6.9% in the second quarter of 2026, according to Cushman & Wakefield data [reported by IREI](https://irei.com/news/modest-new-supply-and-intensifying-demand-push-vacancy-back-below-7/), and Prologis said [roughly one-third of its leasing](https://www.freightwaves.com/news/prologis-says-warehouse-demand-is-piling-up) in the second quarter of 2025 came from 3PLs. Providers that add space need customers to fill it.

## Where does AI search sit in choosing a logistics partner?

Mostly in early research; buyers use AI widely at work but still choose partners through RFPs and site visits.

The evidence is stronger on AI use in logistics operations than on AI use in choosing a provider, so we keep the two apart.

- **Logistics teams use generative AI.** In the Descartes and SAPIO Research survey of shippers and logistics providers in North America and Europe, the top uses were data entry (41%), route or load optimization (39%), freight forecasting (35%) and load matching or capacity sourcing (35%). These are operational uses, not partner searches. Carriers and brokers face their own partner search, covered in [how freight providers win shippers through AI](https://underneath.agency/resources/freight-companies-shippers-ai-search).
- **3PLs see AI as the defining technology.** In the Inbound Logistics survey, AI was the clear choice when respondents were asked about the most impactful technologies in the field, with 94% of replies.
- **Technology now decides selection.** The 2026 Third-Party Logistics Study, led by Penn State’s Dr. C. John Langley with [NTT DATA](https://www.nttdata.com/en-us/insights/reports/transforming-the-supply-chain) and Penske Logistics, found 88% of shippers and 100% of 3PLs reporting successful partnerships, and highlighted a persistent gap between shipper expectations and 3PL technology capabilities.

We found no public survey that measures how often shippers ask ChatGPT, Gemini or Perplexity to suggest a 3PL. A reasonable expectation, given how widely logistics teams already use generative AI, is that many start a provider search the same way they start other research: by asking an assistant for options, then checking them.

## What do shippers and brands ask when they look for a 3PL?

They combine a location, a product type, an order profile and a system, then ask who can handle it.

The questions below are our own examples of the shape these searches take. We did not observe them in real buyer logs.

| What the shipper needs | Sample question we drafted |
|---|---|
| Regional coverage | “3PL with warehouses on both coasts for two-day ground delivery to most of the US” |
| Product handling | “Temperature-controlled warehousing near Dallas for a frozen food brand” |
| Ecommerce fulfillment | “Fulfillment for a skin care brand shipping about 5,000 orders a month, with Shopify integration” |
| Retail compliance | “3PL experienced with Walmart and Target retail routing and labeling” |
| Regulated goods | “Contract logistics providers with FDA-registered warehouses in New Jersey” |
| Comparison | “Compare these three fulfillment providers on cost per order, returns handling and onboarding time” |

Location words matter. In [our country study](https://underneath.agency/research/ai-recommendations-by-country-study), adding “in the United Kingdom” to a question raised the share of local-market brands in ChatGPT’s answers from 25.4% to 49.1%. That study did not cover logistics, but a 3PL’s value is tied to place. We infer that providers who state their facility locations, delivery coverage and regional strengths clearly give assistants more to match when a question names a city, state or region.

## How does an AI answer become an RFP invitation and a contract?

Through a longlist, a check of capabilities, an RFP or quote, site visits and solution design, and a multi-year agreement.

1. **Longlist.** A logistics leader or brand founder asks an assistant, a peer, a directory or a search engine for providers that fit the job. Industry rankings, such as the Inbound Logistics Top 100 3PLs, often feed this list.
2. **Capability check.** The buyer visits provider websites to confirm locations, certifications, integrations, industries served and minimums.
3. **RFP or quote.** Shippers send a request for information or proposal; ecommerce brands share order data and ask for pricing per order, storage and receiving.
4. **Site visits and design.** Finalists host visits, propose staffing, technology and network design, and negotiate service level agreements.
5. **Contract and renewal.** The agreement often runs for years, and many shippers renew without a rebid.

Being named by an assistant helps a 3PL at the longlist and capability-check stages; service, reliability, price and the site visit settle the contract. Buyers say what matters: in a [Transport Intelligence survey](https://ti-insight.com/?p=248855) of logistics buyers in late 2023, reliability and accuracy made up 32.1% of responses on the most important selection criteria, price 20.5% and geographic coverage 10.8%. In the Inbound Logistics survey, shippers rated service more important than price by 72% to 28%.

## What decides whether an assistant names a logistics provider?

Facts and outside endorsements it can find; the platforms explain how they search, not how they choose providers.

**Documented by the platforms.** [OpenAI says](https://help.openai.com/en/articles/9237897-chatgpt-search) ChatGPT search typically rewrites a question into one or more targeted queries sent to search providers, and that a site should allow OpenAI’s search crawler, OAI-SearchBot, to be eligible for inclusion. [Google says](https://blog.google/products/search/ai-mode-search/) AI Mode runs multiple related searches across subtopics and data sources, which it calls “query fan-out.” Neither publishes how individual providers are picked.

**Observed in our studies**, across several industries but not logistics specifically:

- **Assistants look for rankings and publications.** In [our hidden searches study](https://underneath.agency/research/ai-hidden-searches-study), ChatGPT ran a search aimed at a named publication, ranking or award in 43.8% of its answers, and ran a mean of 3.7 searches per question.
- **Trade press and rankings carry weight.** [Our brand entity study](https://underneath.agency/research/brand-entity-ai-recommendations-study) found that a tenfold rise in the independent sites naming a provider in the cited pages came with 4.7 times the odds that it was recommended.
- **Lists vary from run to run.** In [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), only 25.2% of the brands ChatGPT named for a question appeared in all five repeats, so one test answer is a sample, not a verdict.

**Our inference for 3PLs.** The trust factors are what a logistics buyer checks on a site visit, written where a machine can read them: facility addresses and square footage, temperature zones, certifications such as food-grade, FDA registration, bonded or hazardous-materials handling, integrations with ecommerce platforms, marketplaces and EDI, order and pallet volumes you serve well, industries served, and service level commitments. Inclusion in trade rankings, coverage in logistics publications and customer case studies published with permission give an assistant outside confirmation.

## What does a 3PL lose when AI leaves it out?

RFP invitations and quote requests it never sees, in a market where customers rarely come back up for bid.

We found no public measurement of logistics contracts lost to AI absence, so we label our reasoning:

- **Fewer chances to compete.** With more than half of shippers not rebidding at contract end, each search a provider misses may be the only opening for years.
- **Customer acquisition is already the pressure point.** When 46% of 3PLs call finding and retaining customers a top challenge, an extra place on early lists matters.
- **Wrong facts disqualify quietly.** An assistant that says you have no cold storage, no West Coast building or no Shopify connection removes you from a list you never knew existed. Our guide to [fixing wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) shows how to trace an error to the page it came from.
- **Failed partnerships create new searches.** In the Inbound Logistics survey, poor customer service was shippers’ top reason for a failed 3PL partnership, at 34%. Every failure starts a search for a replacement.

## How does GEO work for a 3PL or fulfillment provider?

Generative engine optimization (GEO) helps AI assistants find, describe and verify your network, capabilities and reputation.

For a logistics provider, the work usually includes:

1. **One page per facility.** Address, size, temperature zones, certifications, equipment, carriers and delivery coverage, in text rather than only in a map image or brochure.
2. **Capability pages by buyer type.** Separate, plain descriptions for ecommerce fulfillment, retail compliance, B2B distribution, kitting, returns and cold chain, each stating order profiles and minimums.
3. **Integration facts.** A clear list of the store platforms, marketplaces, warehouse and transportation systems and EDI standards you connect to, with any limits stated honestly. How those system vendors get found is covered in [logistics software in AI search](https://underneath.agency/resources/logistics-software-ai-search).
4. **Consistent identity.** The same company name, facility list and capabilities across your site, directories, LinkedIn, association listings and trade rankings.
5. **Independent proof.** Rankings and awards you qualify for, trade press coverage, conference talks and customer case studies with permission. Our guide on [how brands build authority for AI search](https://underneath.agency/resources/how-brands-build-authority-for-ai-search) explains why outside sources carry weight.
6. **Comparison-ready content.** Straight answers to in-house versus outsourced fulfillment, one 3PL or several (27% of shippers in the Inbound Logistics survey preferred more than one), and how pricing is built.
7. **Crawl access and measurement.** Let the documented search crawlers reach your pages, then ask a fixed set of location, vertical and capability questions across ChatGPT, Gemini, Perplexity, Copilot and Google’s AI features, repeatedly, and compare the results with the RFPs and quote requests you receive.

Nobody can promise that an assistant will name your company. The work makes your network the easiest one for an assistant, and then a logistics buyer, to check.

## What remains unproven about AI search in 3PL selection?

The data shows logistics teams using AI and buyers choosing 3PLs on service, not how many contracts AI answers start.

- **No attribution data.** We found no public study linking AI answers to RFP invitations, quotes or 3PL contracts.
- **Operational AI is not search AI.** The Descartes survey measures generative AI in transportation management, not provider selection.
- **Some sources are older or partial.** The Transport Intelligence criteria come from a late-2023 survey, and much of the 2026 3PL study is behind a download.
- **Surveys come from interested parties.** Radial and GXO sell logistics services, and Descartes sells logistics software.
- **Our studies are cross-industry.** Applying them to logistics is our inference.

## Where should a logistics provider start?

Start by asking assistants the location, product and capability questions your best customers asked before they hired you.

That check shows whether your company is named for your strongest regions and verticals, whether your facilities, certifications and integrations are described correctly, which rankings and publications the answers rely on, and which providers appear in your place. The next step is to publish those facts where assistants read them and to earn the outside coverage that confirms them.

If your growth depends on a few new multi-year accounts a year, [get in touch for a review of your AI search visibility](https://underneath.agency/contact). We will show where your company appears when shippers and brands ask AI for a logistics partner, why other providers are named, and which changes are most likely to bring more RFP invitations and qualified quote requests. Facility pages, integration lists and trade coverage are the kind of 3PL work our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page describes, along with how results are measured against incoming RFPs.

## Frequently asked questions

### Do shippers really ask ChatGPT to recommend a 3PL?

There is no public measurement of that yet. Logistics teams already use generative AI widely at work, 96% in the Descartes survey, so it is reasonable to expect some provider searches to start there, followed by RFPs and site visits.

### Does a small regional 3PL have a chance against national providers?

Yes, for questions tied to a place, product type or service that fewer providers match. State those specifics plainly and earn local and trade coverage that confirms them.

### Are top 3PL lists worth the effort?

They can be. Our hidden searches study found ChatGPT often searching for named rankings and publications, and buyers read the same lists. Enter only those you qualify for honestly.

### Will AI search replace RFPs and site visits?

No. Contracts still depend on proposals, visits and pricing. AI can change which providers get invited.

## Sources

- Logistics Management (2026-06-16), [U.S. 3PL revenues see strong annual gains, reports Armstrong & Associates](https://www.logisticsmgmt.com/article/u.s_3pl_revenues_see_strong_annual_gains_reports_armstrong_associates)
- Inbound Logistics (2025-07), [2025 Inbound Logistics Perspectives: 3PL Market Research Report](https://www.inboundlogistics.com/articles/2025-inbound-logistics-perspectives-3pl-market-research-report/)
- MyChesCo (2025), [30th Annual 3PL Study Charts Shift to Strategic Partnerships and AI Adoption](https://www.mychesco.com/?p=642622)
- NTT DATA (2025), [2026 Third-Party Logistics Study: Transforming the supply chain](https://www.nttdata.com/en-us/insights/reports/transforming-the-supply-chain)
- CXM (2025-03-27), [Self-managed fulfilment is stunting retail growth (Radial survey)](https://cxm.world/?p=130037)
- DC Velocity (2025), [Survey: 96% of shippers using Gen AI in transportation management](https://www.dcvelocity.com/technology/descartes-transportation-management-benchmark-survey)
- US Census Bureau (2026), [Quarterly Retail E-Commerce Sales, 2nd Quarter 2026](https://www.census.gov/retail/mrts/www/data/pdf/ec_current.pdf)
- Container News (2026-02), [GXO posts record revenue in Q4 and full year 2025](https://container-news.com/gxo-posts-record-revenue-in-q4-and-full-year-2025/)
- IREI (2026-07), [Modest new supply and intensifying demand push vacancy back below 7%](https://irei.com/news/modest-new-supply-and-intensifying-demand-push-vacancy-back-below-7/)
- FreightWaves (2025-07), [Prologis says warehouse ‘demand is piling up’](https://www.freightwaves.com/news/prologis-says-warehouse-demand-is-piling-up)
- Transport Intelligence (2024-01-18), [Ti survey reveals key insight into 3PL procurement trends](https://ti-insight.com/?p=248855)
- Forrester (2026-01-21), [Forrester’s 2026 Buyer Insights: GenAI Is Upending B2B Buying](https://www.forrester.com/press-newsroom/forrester-2026-the-state-of-business-buying/)
- OpenAI Help Center (2026), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Google (2025-03-05), [Expanding AI Overviews and introducing AI Mode](https://blog.google/products/search/ai-mode-search/)
- Underneath (2026), [Same question, four countries: do AI recommendations change?](https://underneath.agency/research/ai-recommendations-by-country-study)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [Do Wikipedia and schema make AI assistants recommend a brand?](https://underneath.agency/research/brand-entity-ai-recommendations-study)
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)

---

This is the Markdown twin of https://underneath.agency/resources/logistics-providers-shipper-contracts-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How logistics software firms can win customers in AI search"
description: "By being named, accurately, when shippers, carriers and 3PLs ask AI assistants which TMS, WMS or fleet tool fits, then turning that into demos."
canonical: "https://underneath.agency/resources/logistics-software-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How can logistics software companies win more customers through AI search?

By being named, and described accurately, when shippers, carriers and third-party logistics providers (3PLs) ask AI assistants which transportation, warehouse or fleet software fits their operation. These three buyers search differently and buy differently, so one AI visibility plan will not serve all of them. The common thread is that assistants lean on independent proof, such as reviews, trade coverage and integration details, more than on vendor claims.

## The short version

1. Many logistics buyers are about to shop: in [Peerless Research Group’s 2026 survey](https://www.supplychain247.com/article/2026_software_survey_software_stays_at_the_center_of_the_automated_warehouse), 25% of companies planned to evaluate, purchase or upgrade a warehouse management system (WMS) within two years, and 12% a transportation management system (TMS).
2. Most carriers are small: of almost 580,000 active US motor carriers, 91.5% operate 10 or fewer trucks, per [American Trucking Associations](https://www.trucking.org/economics-and-industry-data) citing federal data. They research and buy without procurement teams.
3. Large fleet deals are big and growing: [Samsara](https://www.samsara.com/company/news/press-releases/q4-fiscal-year-2026-results) reported $1.2 billion of annual recurring revenue from customers paying $100,000 or more, up 37% in its 2026 fiscal year.
4. Technology now decides 3PL contracts: in the [2026 Third-Party Logistics Study](https://www.mychesco.com/?p=642622), “technology capability is now a decisive selection factor” for shippers choosing a 3PL, which makes 3PLs demanding software buyers.
5. Buyers worry most about cost and fit: total cost of ownership (32%) and compatibility with existing systems (31%) topped the adoption challenges in the Peerless survey.

## Which logistics companies buy this software, and what is each account worth?

Three different buyers: shippers, carriers and 3PLs. Each has its own budget, cycle and definition of value.

**Shippers** are manufacturers, retailers and distributors that move their own goods. They buy TMS, WMS, visibility and freight procurement tools. Adoption is uneven: a [FreightWaves analysis](https://www.freightwaves.com/news/the-evolution-of-tms-the-next-phase-of-logistics-technology-is-defined-by-democratization) put TMS adoption at about 50% of large shippers but only 25% of medium-size and 10% of small ones. The same piece cites industry research that companies with a TMS typically save 3% to 12% on freight expenses. The undecided majority is the growth market.

**Carriers** run the trucks. The market is fragmented: [American Trucking Associations](https://www.trucking.org/economics-and-industry-data) reports almost 580,000 active US motor carriers with at least one tractor, and 99.3% of them operate 100 or fewer trucks. They buy fleet management, telematics, dispatch, route optimization and safety tools, often after a short search and a demo. How carriers and brokers win freight is covered in [how freight providers reach shippers through AI](https://underneath.agency/resources/freight-companies-shippers-ai-search).

**3PLs** run warehouses and transport for shippers, and software is part of what they sell. The [2026 Third-Party Logistics Study](https://www.nttdata.com/en-us/insights/reports/transforming-the-supply-chain), led by Penn State’s Dr. C. John Langley with NTT DATA and Penske Logistics, found 88% of shippers and 100% of 3PLs called their relationships successful, and named “a persistent gap between shipper expectations and 3PL technology capabilities” as a challenge. How 3PLs win those shippers is covered in [how warehouse providers make shippers’ AI shortlists](https://underneath.agency/resources/logistics-providers-shipper-contracts-ai-search).

What a customer is worth varies by segment. At the top end, [Samsara](https://www.samsara.com/company/news/press-releases/q4-fiscal-year-2026-results), a fleet and operations platform, reported $1.9 billion in annual recurring revenue, up 30%, and a record 13 deals worth $1 million or more in new annual contract value in one quarter. [SaaStr’s analysis](https://www.saastr.com/5-interesting-learnings-from-samsara-at-1-9-billion-in-arr/) of the results notes Samsara ended the year with 3,194 customers paying $100,000 or more, at an average of $362K each. A small carrier pays far less, but there are hundreds of thousands of them.

## Where does AI search sit in logistics software buying?

At the early research stage for all three buyers, though no public study measures logistics buyers’ AI use.

The broad trend is documented for software buyers. When the review platform [G2](https://company.g2.com/news/g2-research-the-answer-economy) polled 1,076 software buyers in March 2026, 51% said an AI chatbot, more often than Google, is where their research begins. G2 sells visibility to software vendors and its sample is not specific to logistics, so treat that as context.

Google’s results put AI answers in front of logistics buyers as well. Freight and warehouse software sits inside the most exposed category we measured: across [our study of 1,248 US searches](https://underneath.agency/research/ai-overviews-frequency-study), B2B software and technology keywords brought up an AI Overview (Google’s AI summary above the results) on 96.0% of searches, the top rate among eight industries. In [our Reddit study](https://underneath.agency/research/ai-reddit-citations-study), 35.4% of AI Overviews on B2B software searches cited a Reddit thread. Logistics has active practitioner communities, so we infer that driver, dispatcher and warehouse discussions can feed those answers.

The buyers are busy and open to new tools. In the Peerless survey, 26% of respondents said they now use AI, up from 19% a year earlier, and 29% were evaluating it. Buyers who adopt AI in operations, we infer, are also likely to use AI assistants to research it.

## Which questions do shippers, carriers and 3PLs ask?

Different questions for each buyer: fit and integration for shippers, cost and compliance for carriers, client-facing capability for 3PLs. We wrote the example prompts below to show how each logistics buyer might phrase a question; they are not observed data.

| Buyer | Illustrative prompt |
|---|---|
| Shipper | “What’s the best TMS for a mid-sized manufacturer shipping LTL and truckload across the US?” |
| Shipper | “Which WMS integrates with NetSuite and handles ecommerce returns?” |
| Carrier | “Best fleet management software for a 25-truck company, with ELD and dashcams?” |
| Carrier | “Samsara vs Motive for a small fleet: which is cheaper over three years?” |
| 3PL | “Which WMS do multi-client 3PL warehouses use for billing by customer?” |
| 3PL | “What visibility tools can a 3PL offer shippers as a customer portal?” |
| Any | “Alternatives to our current route optimization software that work with Microsoft Dynamics?” |

A question about a TMS or a fleet tool does not stay a single search. Google says AI Overviews and AI Mode [may use a “query fan-out” technique](https://developers.google.com/search/docs/appearance/ai-features) that splits it into subtopics, and OpenAI says [ChatGPT search rewrites a question](https://help.openai.com/en/articles/9237897-chatgpt-search) into narrower, targeted queries. In [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study) of 80 buyer questions, ChatGPT searched for a specific feature in 47.5% of its answers and for reviews or ratings in 46.2%. For logistics software, features like electronic logging, carrier integrations and multi-client billing are exactly what those searches would look for, we infer.

## How does an AI answer become a signed logistics software deal?

By different routes: self-serve demos for carriers, formal evaluations for shippers, and partner-led deals for 3PLs.

**Carriers: answer, demo, contract.** A small fleet owner asks an assistant for the best tool for their size and budget, gets three or four names, and books demos. Without a procurement team, the path from answer to sale can be short. We infer that the AI answer carries more weight here than in enterprise deals, because fewer people check it.

**Shippers: answer, longlist, evaluation.** A logistics director uses AI research to frame the market before a formal selection. The Peerless survey shows how cautious that can be: the share of respondents holding off on software investments rose that year, up from 34% in 2025. The vendors named during research are the ones invited when the budget opens. Procurement professionals often sit in on selections like this, and how they vet vendors with AI is covered in [selling procurement software to RFP buyers](https://underneath.agency/resources/procurement-software-ai-search).

**3PLs: answer, pilot, expansion.** A 3PL picks a WMS or visibility tool it will put in front of its own clients. The 3PL study found more than half of shippers do not rebid at contract end, so a 3PL wants software it can rely on for years, and its software choices are sticky too.

**Expansion is where the value compounds.** SaaStr reports that 96% of Samsara’s $100,000-plus customers subscribe to two or more of its products. A customer who first met you in an AI answer is a starting point for later modules.

A fleet owner who books a demo after asking ChatGPT rarely leaves a referral trail, so freight and warehouse deals need other signals; [what lost clicks mean for pipeline](https://underneath.agency/resources/ai-answers-pipeline-revenue) explains which ones.

## Why does an assistant name one TMS or WMS vendor and not another?

No platform documents how it chooses; studies point to independent sources, and logistics buyers want proof of fit and cost.

**Documented by the platforms.** Both Google and OpenAI say their AI answers search the web and cite what they find, but neither explains how a TMS, WMS or fleet vendor gets picked over its rivals.

**Observed in studies.** [Chen and colleagues](https://arxiv.org/abs/2509.08919) looked at US software questions and found AI search drew 72.7% of its sources from earned media, such as reviews and independent publications, against 45.4% for Google. When [our reputation study](https://underneath.agency/research/is-it-legit-ai-reputation-study) asked assistants “Is this brand legit?”, 88.0% of the answers cited a review or complaint platform, the kind of site where a dispatcher or warehouse manager leaves a verdict.

**What logistics buyers weigh.** In the Peerless survey, the top adoption challenges were:

| Challenge with logistics software | Share of respondents |
|---|---|
| Total cost of ownership | 32% |
| Compatibility with existing systems | 31% |
| User acceptance | 30% |
| Integration with existing software applications | 29% |
| Lack of resources to implement and maintain | 28% |

The return on investment also varies widely: 32% said their WMS took more than 18 months to pay back, while 39% saw a TMS pay back in six to 12 months.

**Our inference.** Cost, integration and payback are facts an assistant can only repeat if they are published. A reasonable expectation is that vendors with public pricing ranges, integration lists and customer results with numbers give AI answers, and buyers, more to work with. Nobody has yet tested it on freight or warehouse software.

## What happens to a logistics vendor that assistants leave out?

Lost demos from small fleets and lost evaluations from shippers, though no study puts a figure on it.

- **Fragmented markets reward the named few.** With 91.5% of carriers running 10 or fewer trucks, most buyers will not run a formal search. If an assistant names three tools, we infer the fourth rarely gets a call.
- **Buying windows are narrow.** Only 12% of Peerless respondents planned to evaluate, buy or upgrade a TMS in the next two years. Missing that window can mean waiting years.
- **Shortlists shift between runs.** Ask ChatGPT the same question five times and the names change: in [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), only 25.2% of the brands it named appeared in all five runs. A WMS or TMS vendor that shows up only some of the time can lose a shipper’s evaluation without ever learning it was open.
- **Wrong facts cost deals.** If an assistant misstates your carrier integrations or pricing, a buyer may rule you out before a demo; our guide to [correcting wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) covers the fix.

## What does GEO involve for a TMS, WMS or fleet software vendor?

It puts the proof each logistics buyer needs where assistants can find, read and repeat it. No vendor, and no agency, can make an assistant recommend a particular fleet or warehouse tool.

Generative engine optimization (GEO) means earning accurate mentions in AI answers; for a logistics software vendor it covers:

1. **Clear positioning per buyer.** Separate pages for shippers, carriers and 3PLs, stating fleet sizes, freight modes and warehouse types you serve. Assistants answer specific questions; vague positioning gives them nothing to match.
2. **Integration facts in plain text.** List the ERP, ecommerce, carrier, load board and electronic logging systems you connect to. Compatibility was the second-ranked challenge in the Peerless survey.
3. **Cost and payback information.** Publish pricing ranges or a cost model and customer payback examples with numbers.
4. **Reviews and communities.** Encourage detailed reviews from real customers on the platforms buyers use, and take part honestly in trucking and warehouse communities. See [legitimate GEO versus manipulation](https://underneath.agency/resources/legitimate-geo-vs-manipulation).
5. **Trade press and associations.** Coverage in logistics publications, conference talks and industry studies are the named sources assistants search for; our piece on [building brand authority for AI search](https://underneath.agency/resources/how-brands-build-authority-for-ai-search) explains why they carry weight.
6. **Measurement across assistants.** Track ChatGPT, Google AI Overviews and AI Mode, Gemini, Perplexity, Copilot and Claude, with questions for each buyer type, asked several times.

Logistics is one corner of a larger software market; [how B2B SaaS companies earn revenue from AI search](https://underneath.agency/resources/b2b-saas-revenue-from-ai-search) covers the rest of it. Legal software shows the same reliance on independent proof, with lawyers trusting what courts and independent testers say, as [how legal software firms win law firms](https://underneath.agency/resources/legal-software-ai-search) shows.

## What don’t we know yet about AI search in freight and warehouse software buying?

Nobody has yet measured whether being named by assistants sells more TMS, WMS or fleet software.

- **No study of logistics buyers’ AI use.** The AI-in-buying surveys cover software buyers generally, not shippers, carriers or 3PLs.
- **Interested sources.** G2 sells review visibility; Samsara reports its own results; the 3PL study is sponsored by a consultancy and a 3PL.
- **Revenue effects have the thinnest support.** Of the 45 studies of AI search optimization that [Martinez](https://arxiv.org/abs/2607.14035) reviewed, the ones on traffic and conversions made up the least supported area.
- **Selection is opaque.** Outside the platforms, no one knows exactly how assistants pick a logistics vendor, and the picks vary between runs and between assistants.

## How should a logistics software vendor begin protecting its demo pipeline?

Ask assistants the questions your shipper, carrier and logistics-provider buyers ask, and note which vendors come back.

Write ten questions each for shippers, carriers and 3PLs, covering category, fit, cost, integrations and alternatives. Put each one through ChatGPT, Gemini, Perplexity, Copilot, Claude and Google’s AI features more than once, because the named vendors shift between runs. Record who is named, which sources are cited, and whether your integrations, pricing and customer results are described correctly. The gaps show where third-party proof is missing.

For an outside view, [ask us to review how assistants describe your logistics product](https://underneath.agency/contact). We will show where AI answers place you with each buyer type, and which gaps are most likely costing you fleet demos, shipper evaluations and multi-year contracts. Separate positioning for shippers, carriers and 3PLs, public integration and cost facts, and repeated checks across assistants are what our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page covers.

## Frequently asked questions

### Do trucking companies use AI assistants to choose fleet software?

No public study measures it. Software buyers broadly do, and with 91.5% of carriers running 10 or fewer trucks, most buy without a formal selection process.

### Is a TMS buyer different from a WMS buyer in AI search?

Often the same company, but different questions. In the Peerless survey, 25% planned to evaluate a WMS within two years against 12% for a TMS.

### Should logistics software vendors publish pricing?

Usually at least ranges. Total cost of ownership was the top adoption challenge, cited by 32% of respondents in the Peerless survey.

### Do review sites matter for logistics software in AI answers?

Likely yes. ChatGPT searched for reviews or ratings in 46.2% of answers in our study, and AI search favors independent sources for software questions.

### How fast can GEO produce demos?

Faster for small carriers than for enterprise shippers, we expect. Carriers often buy after a few demos; shippers usually run longer, formal evaluations.

## Sources

- Supply Chain 24/7, Peerless Research Group (2026), [2026 Software Survey: Software stays at the center of the automated warehouse](https://www.supplychain247.com/article/2026_software_survey_software_stays_at_the_center_of_the_automated_warehouse)
- American Trucking Associations (2026), [Economics and Industry Data](https://www.trucking.org/economics-and-industry-data)
- Samsara (2026-03-05), [Q4 Fiscal Year 2026 Results](https://www.samsara.com/company/news/press-releases/q4-fiscal-year-2026-results)
- SaaStr (2026), [5 Interesting Learnings from Samsara at $1.9 Billion in ARR](https://www.saastr.com/5-interesting-learnings-from-samsara-at-1-9-billion-in-arr/)
- NTT DATA (2025), [2026 Third-Party Logistics Study: Transforming the supply chain](https://www.nttdata.com/en-us/insights/reports/transforming-the-supply-chain)
- MyChesCo (2025), [30th Annual 3PL Study Charts Shift to Strategic Partnerships and AI Adoption](https://www.mychesco.com/?p=642622)
- FreightWaves (n.d.), [The evolution of TMS: The next phase of logistics technology is defined by democratization](https://www.freightwaves.com/news/the-evolution-of-tms-the-next-phase-of-logistics-technology-is-defined-by-democratization)
- G2, Tim Sanders (2026-04-15), [In the Answer Economy, Don’t Win the Click — Win the Answer](https://company.g2.com/news/g2-research-the-answer-economy)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- OpenAI Help Center (2025), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Chen, Wang, Chen and Koudas (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919), arXiv:2509.08919.
- Martinez (2026), [Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026)](https://arxiv.org/abs/2607.14035), arXiv:2607.14035.
- Underneath (2026), [When does Google show an AI Overview? 1,248 US searches](https://underneath.agency/research/ai-overviews-frequency-study)
- Underneath (2026), [When does a Reddit thread become evidence in Google’s AI?](https://underneath.agency/research/ai-reddit-citations-study)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [“Is this brand legit?” How AI assistants build a reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study)
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)

---

This is the Markdown twin of https://underneath.agency/resources/logistics-software-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How luxury brands win high spenders who ask AI what to buy"
description: "By shaping the sources AI cites when high spenders ask open questions: most luxury AI prompts name no brand, and official sites are a minority of citations."
canonical: "https://underneath.agency/resources/luxury-brands-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How do luxury brands win the high spenders who now ask AI what to buy?

By making sure the press, retailers and reference sources that AI assistants cite tell your story accurately, because most luxury questions put to AI name no brand at all. The biggest spenders adopt AI fastest, so the clients a house values most are the ones meeting it first. This guide covers where AI enters the luxury purchase, what decides which houses appear, and what a luxury ecommerce team can do about it without giving up control of its positioning.

## The short version

1. The biggest luxury buyers are the heaviest AI users: in [Bain & Company and Comité Colbert’s 2026 study](https://www.jckonline.com/editorial-article/ai-is-reshaping-luxury-shopping-faster-than-brands-are-adapting-bain-report-finds/), 82% of large luxury purchasers used AI during their latest purchase, as did 54% of US and 64% of Chinese luxury buyers.
2. The brand is often not chosen yet: about 70% of luxury searches on AI platforms mention no specific brand, and about 75% are about discovery or comparison.
3. Houses do not control most of what AI reads: [official brand sites made up only 10% of cited sources in watch searches and 45% in jewelry](https://www.modaes.com/global/back-stage/from-google-to-chatgpt-the-luxury-industry-faces-a-new-battle-for-visibility), yet only 26% of houses are working on external content, against about 60% working on their own sites.
4. Size does not buy presence: 70% of brands with more than €5 billion in revenue had a smaller share of AI visibility than of revenue, while every brand under €1 billion in the ranking had visibility three to eight times its market share.
5. The stakes per client are high: Mytheresa’s average order over the last 12 months reached a record €847, and LuxExperience aims for top customers who are about 4% of buyers to make 40% of sales ([LuxExperience earnings call, May 2026](https://finance.yahoo.com/markets/stocks/articles/luxexperience-luxe-q3-2026-earnings-191532465.html)).

## Who buys luxury online, and what is one client worth?

A small group of high spenders drives most luxury revenue, and each one is worth repeated, high-value orders.

The personal luxury goods market closed 2025 at about €358 billion, down 2% at current exchange rates and flat at constant rates, according to the [Altagamma–Bain Worldwide Luxury Market Monitor](https://altagamma.it/img/osservatorio-2025/Comunicato-Stampa-OSSERVATORIO-ALTAGAMMA-2025.pdf). The same study describes an active client base that is shrinking, with aspirational buyers pulling back while the wealthiest keep spending. Online is a real but minority channel: LuxExperience, owner of Mytheresa, NET-A-PORTER and MR PORTER, cited Bain and Altagamma’s estimate of a global online luxury market of €75 billion.

The economics are concentrated. Mytheresa’s average order value over the last 12 months rose 12.5% to €847; NET-A-PORTER and MR PORTER combined reached €865. Top customers were 9.7% of Mytheresa’s customers in the quarter, and the company described “the famous 4% making 40% ratio” as the target it wants for NET-A-PORTER and MR PORTER too. We infer that houses selling direct see a similar shape: a few clients who buy across categories, return each season and book private appointments matter more than a large volume of first orders.

That shapes how AI matters here. Losing a single high spender at the moment they form a shortlist is not the loss of one order; it is the loss of a client relationship that the whole luxury model is built to grow.

## Where does AI already sit in the luxury purchase?

At the start, before the client has chosen a brand, and increasingly even for purchases completed in a boutique.

Bain and Comité Colbert’s fifth annual *Luxe et Technologie* report found that, as of April 2026, 54% of US luxury buyers and 64% of Chinese luxury buyers used AI during their most recent luxury purchase, against 27% in France. Usage rises with spending: 82% of high spenders used some form of AI, [compared with 28% among lower-spending buyers](https://www.modaes.com/global/back-stage/from-google-to-chatgpt-the-luxury-industry-faces-a-new-battle-for-visibility). And 97% of luxury shoppers who had used AI said they intend to use it again.

AI is not only an online habit. Forty-seven percent of shoppers who ultimately bought in a store reported using AI during their journey. The report lists what they used it for: researching products and brands, styling advice, summarizing reviews, comparing prices and finding complementary pieces.

The tools themselves are built for this kind of browsing. [OpenAI documents](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search) that ChatGPT can show product options with images, details, review summaries and links to buy, and offers virtual try-on for clothing and accessories. [Google says](https://blog.google/products/shopping/google-shopping-ai-mode-virtual-try-on-update/) AI Mode draws on a Shopping Graph of more than 50 billion product listings and lets shoppers try clothes on a photo of themselves.

Across all US retail, [Adobe’s data](https://techcrunch.com/2026/04/16/ai-traffic-to-us-retailers-rose-393-in-q1-and-its-boosting-their-revenue-too/) shows AI traffic converted 42% better than other traffic in March 2026. That figure is not specific to luxury, but it suggests the visitors who arrive from AI answers come ready to buy.

## What do luxury clients ask AI assistants?

Mostly open questions about what to buy, what is worth the price and how pieces compare.

Modaes, reporting on the Bain study, gives the clearest real example: “Which luxury handbag to buy for less than 4,000 euros for everyday use” leaves the brand open, while a question about the price of a Chanel 2.55 starts from a choice already made. Most luxury AI questions are the first kind.

The prompts below are illustrative, written by us to show the shape of these questions; they are not observed data:

- Discovery: “Understated luxury handbags that work for the office and travel.”
- Value: “Is a heritage watch worth the premium over an independent maker at the same price?”
- Comparison: “How do these two houses compare on leather quality and repairs?”
- Authenticity: “Where can I buy this bag new from an authorized seller online?”
- Gifting: “A fine jewelry gift under $5,000 that will still feel special in ten years.”
- Styling: “What to pair with a camel coat for a winter wedding.”

Each of these is a moment where an assistant decides which names to put in front of a client who has not yet committed. That is the battleground Bain describes.

## How does an AI answer turn into luxury revenue?

Through the shortlist: the client takes the names an assistant gave to a website, retailer or boutique.

The path is longer and more varied than in mass ecommerce, so it is worth tracing step by step:

1. **The unbranded question.** A client asks about a category, a budget or an occasion. The assistant names a handful of houses and pieces.
2. **The shortlist.** Modaes puts it directly: the list of options the customer carries with them “may have been formed earlier,” before any visit to a brand’s site, a retailer or a store.
3. **Where the purchase happens.** The client may buy on the house’s own site, on a multibrand retailer such as Mytheresa, or in a boutique after booking an appointment. The in-store figure above shows that much of the influence will never appear as an AI referral in analytics.
4. **The relationship.** A first purchase at full price can lead to repeat orders, private appointments and, eventually, top-client status. This is where the revenue concentrates.

A reasonable expectation, which we infer from these figures rather than measure, is that luxury AI influence will be undercounted by traffic data and should be judged against new-client acquisition and full-price sales instead.

## What decides which luxury houses an AI assistant names?

Platforms document some product signals; the luxury findings are observed in Bain’s study; the rest is our inference.

**Documented by the platforms.** OpenAI states that ChatGPT considers structured product data from first-party and third-party providers, such as price and description, plus other third-party content. When it lists sellers for a product, “merchants are ranked based on factors like availability, price, quality, and whether they are the maker or primary seller.” That last factor matters to houses that sell direct. OpenAI also says review summaries are “based on reviews from public websites” and are not verified by OpenAI.

**Observed in the Bain and Comité Colbert study.** Third-party websites supply most citations in AI answers to unbranded luxury questions. Official sites were 10% of cited sources for watches and 45% for jewelry. Jewelers will find more detail in [how jewelry brands get recommended by AI](https://underneath.agency/resources/jewelry-brands-ai-product-recommendations). Revenue was a poor guide to visibility: smaller houses were far more visible than their size would suggest, while most of the largest were less visible than their market share. Bain’s researchers warn that brands “risk losing potential customers and narrative control.”

**Our inference.** For luxury, the sources that seem most likely to shape an answer are those a client would also trust: fashion and watch media, auction and resale references, authorized retailers, and long-form reviews by collectors. Our article on [whether AI assistants favor big brands](https://underneath.agency/resources/do-ai-assistants-favor-big-brands) shows that in controlled tests a familiar name wins when everything else is equal, but a clearly stated advantage can beat it. Heritage, craftsmanship and service only help if they are written down somewhere an assistant can read and quote.

**Trust factors specific to luxury.** Authenticity and authorized selling, price consistency across markets, repair and after-sales service, craftsmanship and provenance, and a brand staying true to its values. Savanta’s MillionaireVue data found that 52% of high-net-worth shoppers prefer deeper connections with brands, against 42% who prefer faster, more transactional engagement, which is a reminder that AI visibility has to serve the relationship, not replace it.

## What does a luxury house risk by staying absent from AI answers?

Clients, and control of how the house is described, including whether buyers are sent to authorized sellers.

The first risk is losing the client at the shortlist. With seven in ten luxury AI searches naming no brand, a house that is absent does not get the chance to be compared. The second is narrative. If most citations come from third parties, an assistant’s description of a house’s quality, price and heritage is assembled from what others wrote.

The third risk is specific to luxury: counterfeits. The OECD and the European Union Intellectual Property Office estimate that [trade in fake goods reached USD 467 billion, or 2.3% of total imports](https://www.oepm.es/cs/OEPMSite/contenidos/Revista_InfoPYM/2025/Mayo/en/noticia4.html), with counterfeit imports into the European Union worth EUR 117 billion, and they note that counterfeiters use online sales platforms. Savanta warns that AI platforms “can easily recommend lookalikes if protections aren’t in place.” We have not seen a study measuring how often assistants send luxury shoppers to unauthorized sellers, so treat this as a risk to check, not a measured rate. For the general case, see our article on [fake reviews and fake brands in AI recommendations](https://underneath.agency/resources/fake-reviews-fake-brands-ai-recommendations).

Most houses are not yet acting on this. Only 48% of groups and houses regularly monitor their performance in AI search, and only 10% consider their current position strong. Nearly a quarter of luxury companies now rank AI among their top three priorities, up from 5% in 2024, but Bain says most of that work is in the back office rather than with clients.

## How does GEO work for a luxury brand without diluting it?

It shapes the facts and sources assistants rely on, in the house’s own voice; it cannot guarantee a mention.

Generative engine optimization (GEO) for luxury is closer to press relations and brand protection than to performance marketing. It usually covers six areas:

1. **External coverage where AI looks.** Since official sites are a minority of citations, work with the fashion, watch and jewelry press, collector communities and reference sites to keep accurate, current descriptions of the house and its key lines. Our guide to [building authority for AI search](https://underneath.agency/resources/how-brands-build-authority-for-ai-search) explains how to choose those sources.
2. **Authorized-seller clarity.** Publish who sells your products, in which markets, and how to verify authenticity, so an assistant that is asked where to buy has a clear answer and the house stays the “maker or primary seller” OpenAI describes.
3. **Product facts written plainly.** Materials, origin, craft techniques, sizes, care and repair policies, stated as facts on product pages and in feeds, not only as imagery. Our article on [what drives AI product recommendations](https://underneath.agency/resources/what-drives-ai-product-recommendations) shows why concrete details matter.
4. **Consistent pricing and naming across markets.** An assistant that quotes last season’s price or misnames a line undercuts the house; our guide to [correcting wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) covers the repair.
5. **Multibrand partner alignment.** Retailers such as Mytheresa describe your pieces too; make sure their product copy matches yours.
6. **Market-by-market monitoring.** AI use differs sharply between the US, China and France, so track the unbranded discovery and comparison questions clients ask in each market, across ChatGPT, Gemini, Perplexity, Copilot and Google’s AI features. Our guide to [GEO across languages](https://underneath.agency/resources/geo-strategy-for-global-brands-across-languages) covers the multilingual side.

None of this asks a house to discount, chase volume or change its tone. It asks the house to make its own story easy to verify.

## What can’t a luxury house learn from current AI research?

How often an AI answer changes which house a client buys from.

The strongest figures here come from one consultancy study, reported through trade media; we could not access Bain’s full report directly. The adoption numbers are survey answers, and “used AI during the purchase” covers anything from a quick question to a long session. The citation shares cover watches and jewelry; other categories may differ. Mytheresa’s figures describe one multibrand retailer’s clients, not every house. Adobe’s conversion data covers all US retail. No public source yet links AI visibility to luxury client acquisition or lifetime value, and no platform documents how it chooses between luxury houses for an open question.

## Where should a luxury ecommerce team start?

Start by finding out what assistants say about your house when clients ask open questions in your key markets.

A useful first step is an audit of the unbranded discovery, comparison, gifting and “where to buy” questions your clients ask, across the main assistants and in your priority markets, with the sources each answer cites. That shows where your narrative is being written by others and where buyers may be sent to the wrong sellers. To have us run that audit with you and plan the work, tied to new-client acquisition and full-price sales, [arrange a private conversation with our team](https://underneath.agency/contact). Our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page outlines how that work proceeds for a house, from press and authorized-seller facts to monitoring each market, in its own voice.

## Frequently asked questions

### Do wealthy luxury clients really use AI to shop?

Yes, more than other buyers. In Bain and Comité Colbert’s study, 82% of high spenders used AI during their latest luxury purchase, against 28% of lower spenders.

### Can a luxury house control what AI says about it?

Not directly. It can make accurate facts easy to find on its own site and in the third-party sources assistants cite, which is where most citations come from.

### Will AI shopping push luxury toward discounting?

Not necessarily. Assistants weigh price when a budget is given, but clients also ask about quality, craftsmanship and value. Clear facts about those let a house compete without discounting.

### Should luxury brands block AI assistants from their websites?

Blocking may keep a house’s own pages out of answers while third-party descriptions remain. Our [crawler-blocking study](https://underneath.agency/research/ai-crawler-blocking-study) covers the trade-offs; decide deliberately rather than by default.

## Sources

- Bain & Company and Comité Colbert, reported by JCK (2026), [AI Is Reshaping Luxury Shopping Faster Than Brands Are Adapting, Bain Report Finds](https://www.jckonline.com/editorial-article/ai-is-reshaping-luxury-shopping-faster-than-brands-are-adapting-bain-report-finds/)
- Modaes (2026), [Luxury’s Next Frontier: Battling for Visibility in the Age of AI with ChatGPT](https://www.modaes.com/global/back-stage/from-google-to-chatgpt-the-luxury-industry-faces-a-new-battle-for-visibility)
- Altagamma and Bain & Company (2025), [Osservatorio Altagamma 2025](https://altagamma.it/img/osservatorio-2025/Comunicato-Stampa-OSSERVATORIO-ALTAGAMMA-2025.pdf)
- LuxExperience, transcript via Yahoo Finance (2026), [LuxExperience Q3 2026 earnings call](https://finance.yahoo.com/markets/stocks/articles/luxexperience-luxe-q3-2026-earnings-191532465.html)
- Savanta (2025), The new concierge: generative AI and the future of luxury shopping, savanta.com
- OECD and EUIPO, summarized by the Spanish Patent and Trademark Office (2025), [Global trade in counterfeit goods 2025](https://www.oepm.es/cs/OEPMSite/contenidos/Revista_InfoPYM/2025/Mayo/en/noticia4.html)
- OpenAI (2026), [Shopping with ChatGPT Search](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search)
- Google (2025), [Shop with AI Mode, use AI to buy and try clothes on yourself virtually](https://blog.google/products/shopping/google-shopping-ai-mode-virtual-try-on-update/)
- TechCrunch (2026), [AI traffic to US retailers rose 393% in Q1, and it’s boosting their revenue too](https://techcrunch.com/2026/04/16/ai-traffic-to-us-retailers-rose-393-in-q1-and-its-boosting-their-revenue-too/)
- Underneath (2026), [AI crawler blocking study](https://underneath.agency/research/ai-crawler-blocking-study)

---

This is the Markdown twin of https://underneath.agency/resources/luxury-brands-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How machine learning firms get found when buyers ask AI"
description: "By earning the independent proof AI assistants draw on, so they name you when enterprises ask who can build or supply a machine learning solution."
canonical: "https://underneath.agency/resources/machine-learning-companies-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How do machine learning companies get found when enterprise buyers ask AI who to hire?

By earning the independent, checkable proof that AI assistants draw on when a buyer asks who can build or supply a machine learning solution, then turning that first mention into a paid pilot that reaches production. Enterprises now buy most of their AI rather than build it, and they look hard for vendors who can show results. Whether an assistant names you depends far more on what others publish about your work than on what your own site claims.

## The short version

1. Enterprises spent $37 billion on generative AI in 2025, up from $11.5 billion in 2024, and 76% of AI use cases are now purchased rather than built internally, against 53% purchased a year earlier, according to [Menlo Ventures](https://menlovc.com/perspective/2025-the-state-of-generative-ai-in-the-enterprise/).
2. Buying from outside works better for buyers: in MIT research reported by [Fortune](https://fortune.com/2025/08/18/mit-report-95-percent-generative-ai-pilots-at-companies-failing-cfo/), purchasing AI tools from specialized vendors and building partnerships succeeded about 67% of the time, while internal builds succeeded only one-third as often.
3. Once an enterprise commits to exploring an AI solution, 47% of AI deals reach production, against 25% for traditional software, Menlo Ventures found.
4. Buyers have been burned: in [IBM’s survey](https://newsroom.ibm.com/2025-05-06-ibm-study-ceos-double-down-on-ai-while-navigating-enterprise-hurdles) of 2,000 CEOs, only 25% of AI initiatives had delivered the expected return and only 16% had scaled enterprise-wide.
5. Being known is not being found: in a [test of 112 Product Hunt startups](https://arxiv.org/abs/2601.00912), ChatGPT recognized 99.4% when asked by name but surfaced only 3.32% for discovery questions such as “What are the best AI tools launched this year?”

## Who pays for machine learning work, and what is a client worth?

Enterprise teams buy it, mostly from outside vendors now, starting with a pilot and growing into production contracts.

Spending on AI is huge and still climbing. [Gartner](https://www.gartner.com/en/newsroom/press-releases/2025-09-17-gartner-says-worldwide-ai-spending-will-total-1-point-5-trillion-in-2025) forecasts worldwide AI spending of nearly $1.5 trillion in 2025 and more than $2 trillion in 2026, though much of that is hardware and AI built into existing products. Business adoption is broad: the [Stanford AI Index](https://hai.stanford.edu/ai-index/2025-ai-index-report) reports that 78% of organizations used AI in 2024, up from 55% the year before.

Two kinds of machine learning company sell into that demand, and they monetize differently:

- **Machine learning consultancies and service firms** sell expertise: a discovery phase, a paid proof of concept, then a build, deployment and support engagement. The buyer is often a chief data or technology officer, or a business-unit leader with a specific problem such as fraud, forecasting or document processing.
- **Machine learning product companies** sell models, platforms or applications on subscription or usage pricing. By Menlo Ventures’ count, 27% of AI application spend comes through product-led motions, where individual users adopt first, nearly four times the traditional software rate. Platform vendors are covered in [how data science platforms win enterprise buyers](https://underneath.agency/resources/data-science-platforms-ai-search).

The shift toward buying is what makes discovery matter. In 2024, Menlo found 47% of AI solutions were built internally; in 2025, 76% of use cases were purchased. Startups captured 63% of the AI application market, up from 36% the year before, so new vendors can win against incumbents.

We found no reliable public benchmark for the size of a machine learning consulting engagement, so we do not quote one. What the evidence does show is the shape of the value: Menlo counts at least 10 AI products earning over $1 billion in annual recurring revenue and 50 earning over $100 million, and a pilot that reaches production usually becomes a multi-year relationship. We infer that the first pilot is where most of a machine learning vendor’s lifetime revenue with a client is decided.

## When do enterprise buyers ask assistants about machine learning vendors?

At the research and shortlist stages, where buyers now ask assistants to explain options and name vendors.

No public survey isolates machine learning buyers, so the best evidence covers software buyers in general. In [G2’s 2026 buyer behavior report](https://sell.g2.com/2026-buyer-behavior-report), more than 80% of buyers had sourced software recommendations from an AI chatbot in the last two years, and half of those said AI had the greatest influence during shortlisting and evaluation. G2 runs a review marketplace and has an interest in that finding.

Google’s own search is just as present. Of the 1,248 US searches in [our AI Overview study](https://underneath.agency/research/ai-overviews-frequency-study), B2B software and technology keywords carried an AI Overview 96.0% of the time, more than any of the other seven industries. Google’s announcement of AI Mode describes a [“query fan-out” technique](https://blog.google/products/search/ai-mode-search/) that splits a question into multiple related searches across subtopics and then combines what comes back. A single question about a fraud model vendor can therefore pull in case studies, reviews, pricing and compliance pages at once.

A reasonable expectation is that machine learning buyers use these tools more than most. They are technical, they work with AI daily, and the category changes so fast that last year’s comparison articles are already out of date.

## Which questions do machine learning buyers ask AI assistants?

Problem-first questions: who can solve this, build or buy, how vendors compare, and what it costs.

Buyers rarely start by naming a vendor. They start with a business problem and narrow from there. We wrote the example prompts below ourselves to show how a CTO or business-unit head might frame a machine learning need; none is a recorded query:

- Problem to vendor: “Which companies build demand forecasting models for mid-size grocery chains?”
- Build or buy: “Should a regional bank build its own fraud detection model or buy one?”
- Alternatives: “Alternatives to building on SageMaker for a team without platform engineers.”
- Comparison: “Compare machine learning consultancies with experience in medical imaging and FDA submissions.”
- Cost and timeline: “How long and how much does a computer vision proof of concept usually take?”
- Risk: “Which vendors can fine-tune a model on our data without it leaving our cloud account?”

Problem and build-or-buy questions decide who is considered. Comparison, cost and risk questions decide who gets the pilot. The answers to the cost questions are often generic, which is a reason to publish your own ranges and timelines where assistants can find them.

## How does an AI mention turn into machine learning revenue?

Through the pilot: the assistant names you, the buyer checks your proof, a paid pilot runs, and production follows.

**Shortlist.** A buyer who asks an assistant for vendors usually leaves with a short list of names. If you are not on it, the buyer may never visit your site or meet your team.

**Due diligence.** Technical buyers check what the assistant said. They read case studies, papers, code and reviews, and they talk to peers. For machine learning, the proof they want is specific: the problem, the data, the result and how it was measured. Proof matters even more when every rival claims AI, as [how AI software firms cut through the hype](https://underneath.agency/resources/ai-software-companies-in-ai-search) explains.

**Pilot.** This is where machine learning deals are most fragile. [Gartner predicted](https://www.gartner.com/en/newsroom/press-releases/2024-07-29-gartner-predicts-30-percent-of-generative-ai-projects-will-be-abandoned-after-proof-of-concept-by-end-of-2025) that at least 30% of generative AI projects would be abandoned after proof of concept by the end of 2025, citing poor data quality, inadequate risk controls, escalating costs or unclear business value. It also predicts that through 2026, organizations will [abandon 60% of AI projects](https://www.gartner.com/en/newsroom/press-releases/2025-02-26-lack-of-ai-ready-data-puts-ai-projects-at-risk) unsupported by AI-ready data.

**Production and expansion.** Pilots that survive tend to convert. Menlo’s 47% production rate for AI deals, against 25% for traditional software, suggests that the hard part is getting into consideration, not closing once there. That puts unusual weight on the first step, where AI assistants now sit.

## Why does an assistant name one machine learning firm for a problem and not another?

Assistants don’t explain their picks; research points to third-party coverage, links from other sites and practitioner discussion.

**Documented by the platform.** Beyond Google’s description of query fan-out, no assistant publishes how it chooses which vendors to name for a question like “who builds fraud models?” Anything more specific is inference.

**What researchers found.** In the US categories [Chen and colleagues](https://arxiv.org/abs/2509.08919) tested, 72.7% of AI search sources were independent “earned” sites; for Google the share was 45.4%. In the Product Hunt study, the signals that predicted visibility in Perplexity were the number of referring domains and community presence on Reddit; the products’ own optimization scores showed no correlation with discovery. For consultancies, one common shortcut does not help much: in [our study of AI-cited “best X” lists](https://underneath.agency/research/self-promoting-best-lists-study), 24.2% of numbered lists with an identifiable publisher ranked their own publisher first, and such lists made up just 1.1% of all citations.

**Trust factors specific to machine learning.** We infer that what persuades a skeptical technical buyer also gives an assistant something to cite:

- Case studies that state the problem, the data, the metric and the measured result, ideally confirmed by the client.
- Peer-reviewed papers, conference talks and open-source code that show real technical depth. The AI Index notes that nearly 90% of notable AI models in 2024 came from industry, so buyers expect vendors to publish, not just claim.
- Benchmarks reported honestly, with dates. The AI Index reports scores on one coding benchmark, SWE-bench, rising 67.3 percentage points in a single year; a benchmark claim from last year may already be stale.
- Clear statements on data handling, security and compliance, since the risk questions above are often where vendors are cut.

## What does a machine learning vendor lose by not being named?

The pilot, and with it most of the client relationship, because vendors rarely get added once a shortlist forms.

The Product Hunt study shows the size of the gap for young AI companies: near-perfect recognition by name, but a 3.32% chance of appearing when a buyer asks an open question. A consultancy that is well known to its existing clients can be equally invisible to a new buyer who starts with a problem rather than a name. Our article on [why well-known brands miss AI recommendations](https://underneath.agency/resources/why-well-known-brands-miss-ai-recommendations) covers the mechanism.

There is also an opportunity in the other direction. Menlo’s data shows startups taking most of the AI application market from incumbents in a single year. Buyers are open to new names; we infer that the vendors who make their proof easy to find are best placed to benefit. Companies selling AI agents face the same opening, covered in [how AI agent companies reach buyers’ shortlists](https://underneath.agency/resources/ai-agent-companies-customers-from-ai-search).

## How does GEO work for a machine learning company?

It makes your expertise and results easy for assistants to find, verify and cite; it cannot promise a mention.

For an ML consultancy or product company, generative engine optimization (GEO) tends to span six pieces of work:

1. **A clear entity.** Describe what you do, for whom and in which domains the same way on your site, LinkedIn, partner directories, cloud marketplaces and conference bios, so assistants have one consistent picture.
2. **Proof others publish.** Earn client-approved case studies, coverage in trade and technical publications, podcast and conference appearances, and mentions in independent vendor comparisons. How an ML firm can earn that kind of third-party attention is the subject of [our authority-building guide](https://underneath.agency/resources/how-brands-build-authority-for-ai-search).
3. **Technical depth in public.** Publish papers, open-source code, model cards and honest benchmark write-ups with dates and methods. These serve technical buyers and give assistants specific facts to cite.
4. **Answers to the buyer’s real questions.** Publish ungated pages on build-or-buy decisions, typical pilot timelines and cost ranges, data requirements and security practices, written for the industries you serve.
5. **Community presence.** Contribute where practitioners talk, such as GitHub, technical forums and relevant Reddit communities, since community presence predicted visibility in the Product Hunt study.
6. **Measurement tied to pipeline.** Track a fixed set of problem-first buyer questions across ChatGPT, Gemini, Perplexity, Copilot and Google’s AI features, repeated over time, and compare with discovery calls and pilots. For new companies, our article on [why ChatGPT misses new products](https://underneath.agency/resources/why-chatgpt-misses-new-products) explains the timing problem.

## Which gaps remain in the evidence on how ML buyers use AI assistants?

Nobody has measured how ML buyers use assistants, or whether a mention leads to a pilot.

The buyer data comes from general software surveys, several run by companies with an interest in the answer. Menlo Ventures invests in AI companies, and its market figures are its own estimates. The MIT findings come from interviews, a survey and public deployments, reported by Fortune. The Product Hunt study tested one cohort of startups with two assistants. No published study yet follows machine learning buyers from an AI answer to a signed pilot or a production contract, so any claim that AI visibility drives revenue in this industry should be tested against your own pipeline.

## How can a machine learning firm learn whether assistants put it in line for pilots?

Ask assistants about the business problems you solve, not your company name, and see whether you are named.

A useful first step is an audit of the problem-first, build-or-buy, comparison and risk questions your buyers ask, across the main assistants, checked against where your discovery calls and pilots actually come from. That shows which proof is missing and which gaps cost the most paid pilots. To run that audit with us and plan a route onto more enterprise shortlists, [ask us for a review of your machine learning visibility](https://underneath.agency/contact). How case studies, technical publishing and problem-first pages are then built and tracked against pilots is explained on our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page.

## Frequently asked questions

### Do AI assistants recommend machine learning consulting firms?

They can name vendors when asked, but no platform documents how it chooses them. Studies of AI search find it leans on independent sources more than Google does, so third-party coverage and client-approved case studies matter more than self-description.

### Do published papers and open-source code help an ML company appear in AI answers?

No study has measured that directly for machine learning firms. Papers and code are public, specific and often discussed by others, which makes them useful evidence for buyers and plausible material for assistants to draw on.

### Should a machine learning consultancy publish its own “top ML firms” list?

It is a weak bet. In our study, about a quarter of AI-cited numbered lists ranked their own publisher first, but such lists were a tiny share of citations. Independent comparisons carry more weight.

### How do we measure whether AI assistants send us pipeline?

Track how often assistants name you for a fixed set of buyer questions, ask every new lead how they found you, and follow those leads through to pilots and production contracts.

## Sources

- Menlo Ventures (2025), [2025: The State of Generative AI in the Enterprise](https://menlovc.com/perspective/2025-the-state-of-generative-ai-in-the-enterprise/)
- Fortune (2025), [MIT report: 95% of generative AI pilots at companies are failing](https://fortune.com/2025/08/18/mit-report-95-percent-generative-ai-pilots-at-companies-failing-cfo/)
- IBM (2025), [IBM Study: CEOs Double Down on AI While Navigating Enterprise Hurdles](https://newsroom.ibm.com/2025-05-06-ibm-study-ceos-double-down-on-ai-while-navigating-enterprise-hurdles)
- Gartner (2025), [Gartner Says Worldwide AI Spending Will Total $1.5 Trillion in 2025](https://www.gartner.com/en/newsroom/press-releases/2025-09-17-gartner-says-worldwide-ai-spending-will-total-1-point-5-trillion-in-2025)
- Gartner (2024), [Gartner Predicts 30% of Generative AI Projects Will Be Abandoned After Proof of Concept By End of 2025](https://www.gartner.com/en/newsroom/press-releases/2024-07-29-gartner-predicts-30-percent-of-generative-ai-projects-will-be-abandoned-after-proof-of-concept-by-end-of-2025)
- Gartner (2025), [Lack of AI-Ready Data Puts AI Projects at Risk](https://www.gartner.com/en/newsroom/press-releases/2025-02-26-lack-of-ai-ready-data-puts-ai-projects-at-risk)
- Stanford HAI (2025), [The 2025 AI Index Report](https://hai.stanford.edu/ai-index/2025-ai-index-report)
- G2 (2026), [2026 Buyer Behavior Report](https://sell.g2.com/2026-buyer-behavior-report)
- Google (2025), [Expanding AI Overviews and introducing AI Mode](https://blog.google/products/search/ai-mode-search/)
- Chen and colleagues (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919)
- Sharma (2025), [The Discovery Gap: How Product Hunt Startups Vanish in LLM Organic Discovery Queries](https://arxiv.org/abs/2601.00912)
- Underneath (2026), [When does Google show an AI Overview?](https://underneath.agency/research/ai-overviews-frequency-study) and [self-ranking “best X” lists](https://underneath.agency/research/self-promoting-best-lists-study)

---

This is the Markdown twin of https://underneath.agency/resources/machine-learning-companies-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How machinery companies win buyers through AI search"
description: "Machinery makers get shortlisted when AI can verify their applications, cycle times, prices, financing and dealer support, the facts a capital case rests on."
canonical: "https://underneath.agency/resources/machinery-companies-buyers-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Will AI steer a plant’s next machine purchase toward us or a rival?

It may, because buyers increasingly start capital research with AI, but the assistant can only steer them toward a builder whose applications, performance, price range and local support it can verify. The decision itself is still made by a buying group, on payback, in front of a running machine.

This article is for builders of production machinery sold as capital equipment: machine tools, fabrication and forming machines, and packaging and processing lines, sold direct or through dealers. For process equipment chosen against duty conditions, or for components sold through distributors, the buying journey is different.

## The short version

1. Capital demand is strong: according to [AMT’s order data](https://www.automation.com/article/583.4-million-usd-new-machinery-orders-us-economic-strengths), US manufacturing technology orders reached $2.77 billion in the first five months of 2026, up 31.9%, after a 2025 total of [$5.74 billion](https://www.mdm.com/news/research/economic-trends/december-metalworking-machinery-orders-set-new-monthly-record/), up 22.5%.
2. Packaging is a large, growing market: [PMMI](https://www.pmmi.org/news/pmmis-2026-state-of-the-industry-report-reveals-a-packaging-machinery-market-driven-by-flexibility-automation-and-new-north-american-opportunities) estimates US packaging machinery shipments at $11.7 billion in 2025, with food machinery projected to grow from $5.1 billion to about $7.1 billion by 2031.
3. Buyers plan to spend more: in the National Association of Manufacturers’ third-quarter 2026 survey, [reported by Industrial Info Resources](https://www.industrialinfo.com/news/article/us-manufacturing-boosts-capex-expectations--362933), respondents expected capital spending to rise 2.6% over the next 12 months, and nearly 50% expected to spend more.
4. Most machines are financed: in the [Equipment Leasing & Finance Foundation’s](https://www.leasefoundation.org/wp-content/uploads/2022/10/2022-Horizon-Report-Fact-Sheet.pdf) survey, 79.3% of businesses that acquired equipment or software in 2021 used at least one form of financing.
5. Decisions are collective and tested: [Forrester](https://www.forrester.com/press-newsroom/forrester-2026-the-state-of-business-buying/) says a typical business purchase now involves 13 internal stakeholders and nine external influencers, and 78% of buyers making purchases of $10 million or more use a trial first.

## Who buys production machinery, and what is one order worth?

Operations leaders, engineers and finance decide together; one order can lead to service revenue and repeat machines.

A machine purchase usually starts with an operations problem: not enough capacity, too few skilled operators, a new program to win, or a product mix that changes faster than the line can. A plant manager or vice president of operations owns the need; manufacturing engineers judge the machine; finance and procurement judge the payback and the terms. In a job shop, the owner may be all of these at once.

The market data shows where orders come from. According to AMT’s year-end report, as covered by Modern Distribution Management, orders from contract machine shops, the largest customer group, grew 19.1% in 2025, aerospace orders rose 45.1%, and auto manufacturers’ orders rose 22.2%. December 2025 alone brought $814.3 million in orders, a monthly record. How those shops win their own work is covered in [how CNC shops win RFQs through AI](https://underneath.agency/resources/cnc-machine-shops-rfqs-ai-search). In packaging, PMMI puts food at about 44% of US machinery shipments.

There is no public average order value for a machine builder, and prices range from a small machining center to a complete line. Our inference: each order also opens a stream of tooling, parts, service, automation add-ons and, if the machine performs, the next machine on the same floor. Process equipment such as pumps and compressors is chosen against duty conditions and reaches buyers through engineers’ bid lists, as [how equipment makers reach more bid lists](https://underneath.agency/resources/industrial-equipment-leads-ai-search) explains.

## How are capital equipment buyers using AI today?

As a starting point for research, alongside video, trade shows and peers, before a demo settles the choice.

Forrester’s 2026 buying study describes generative AI searches as “the starting point for B2B buyers,” followed by checks with colleagues and outside influencers. The 2026 [State of Marketing to Engineers](https://advertising.globalspec.com/state-of-marketing-to-engineers-research-report) research by TREW Marketing and GlobalSpec, which surveys technical buyers, adds detail on where they look:

| What technical buyers reported in 2026 | Share |
|---|---|
| Bring generative AI into the purchasing process | 69% |
| Routinely research purchases on YouTube | 30% |
| Routinely research at conferences and trade shows | 27% |
| First contacted a salesperson because they wanted a demo | 8% |

Manufacturers are also using AI inside their own buying process. PMMI’s 2026 report lists early uses in packaging and processing that include “developing specifications using previous projects,” alongside troubleshooting and capturing knowledge from retiring staff. Our inference: when a buyer’s own AI helps draft the specification, the assumptions it draws from public sources can shape which builders’ capabilities look like a fit.

Trade shows remain central. PMMI says PACK EXPO International brings together more than 2,600 exhibitors and 48,000 attendees. A buyer who meets you at a show will often research you again afterward, in a search box or an assistant.

## What do plant leaders ask AI before buying a machine?

Questions about the application, the payback, the budget, the alternatives and who supports the machine locally.

The examples below are our own, written to show the kinds of questions involved; no assistant or buyer produced them:

| Stage | Example question |
|---|---|
| Application | “Best 5-axis machining center for titanium aerospace parts, under $500,000” |
| Line change | “Case packers that handle frequent SKU changeovers on a mid-size food line” |
| Automation | “Payback period for a robotic palletizer on a three-shift beverage plant” |
| Comparison | “Compare three vertical machining centers for a job shop on spindle, support and resale value” |
| Financing | “Should a small shop lease or buy a CNC lathe?” |
| Support | “Machine tool dealers with applications engineers near Dayton, Ohio” |

Two kinds of wording deserve attention. Budget words change answers: in [our prompt-phrasing study](https://underneath.agency/research/ai-prompt-phrasing-study), asking the same question again kept the same first brand 68.0% of the time, but adding “on a tight budget” kept it only 15.3%. Place words matter too: in [our brand agreement study](https://underneath.agency/research/ai-assistants-brand-agreement-study), two assistants’ picks overlapped 0.160 on questions naming a place, against 0.390 on national ones, so dealer and service coverage needs to be stated where an assistant can read it.

## How does an AI answer turn into a machine order?

Through a shortlist, a demo or test, a payback case with financing, then a purchase order and years of service.

1. **Shortlisted.** A plant leader or engineer asks an assistant, a search engine or peers which builders fit the application. Several names go on a list.
2. **Shown.** The buyer watches videos, visits a showroom or show booth, and asks for a demonstration, test cut or factory acceptance test. Forrester’s finding that more than 60% of business buyers now use a trial fits this habit: for machinery, our inference is that the test cut is the trial.
3. **Justified.** Finance builds the payback case. Most buyers finance: the Foundation’s survey found leasing was the most common method in 2021, used by 26%. The Equipment Leasing and Finance Association expects the deal volume in its monthly index to reach $129 billion in 2026, according to [Supply Chain 24/7](https://www.supplychain247.com/article/equipment-deal-volume-to-reach-129-billion-in-2026), the highest since its survey began in 2006.
4. **Ordered and supported.** The purchase order goes to the builder or its dealer, followed by installation, training, parts and service.

AI visibility acts mainly on the first step and shapes the second, because the builders a buyer finds early are the ones invited to demonstrate. For the general link between early AI research and pipeline, see [what lost clicks to AI answers mean for pipeline and revenue](https://underneath.agency/resources/ai-answers-pipeline-revenue).

## What decides whether an assistant recommends your machines?

Verifiable facts about applications, performance, price and support, found on your site and in independent sources.

What the platforms document: [Google says](https://blog.google/products/search/ai-mode-search/) AI Mode uses a “query fan-out” technique that issues “multiple related searches concurrently across subtopics and multiple data sources.” [OpenAI says](https://help.openai.com/en/articles/9237897-chatgpt-search) ChatGPT search rewrites a question into one or more targeted searches and that sites must allow its OAI-SearchBot crawler to be eligible. Neither says how machine builders are chosen.

What our studies observed, across other categories:

- **Reviews and prices get searched.** In [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study), ChatGPT looked for reviews in 46.2% of its answers and for prices in 23.8%. A builder with no public price range leaves the assistant to quote someone else’s.
- **Video is read through its words.** In [our YouTube study](https://underneath.agency/research/ai-overview-youtube-videos-study), only 15.1% of the YouTube videos cited in Google’s AI Overviews were also shown on page one of the same search. For 97.9% of cited videos with a text excerpt, the excerpt was not in the video description; it came from what was said, and 99.4% of the videos we could check had captions. Our inference for machinery: a demo video that states the part, material, cycle time and machine model out loud is easier to cite than one set to music.
- **Lists vary by wording.** The budget effect above means a builder known only as the premium choice may vanish from answers that add a price constraint.

What we infer for machinery builders: the deciding facts are the ones in a capital request. Application examples with parts or packages named, cycle and changeover times, footprint and utilities, price ranges or “starting at” figures, financing options, warranty, and the dealer or service location nearest the plant. Independent coverage from trade publications and show reporting confirms them.

## What does a builder lose when AI leaves it out?

It loses the demo invitation, and with it the chance to compete on performance, price and service.

No public data measures machine orders lost to AI absence, so here is the labeled reasoning:

- **Budgets are moving now.** Orders up 31.9% in early 2026 and rising capital spending plans mean many buyers are researching this year. A builder absent from that research waits for the next cycle.
- **Shortlists are short.** With many stakeholders and a trial before purchase, buyers can demonstrate only a few machines. We infer that the builders named early fill those slots. Warehouse automation buyers build similar longlists, as [how warehouse automation vendors make shortlists](https://underneath.agency/resources/warehouse-automation-enterprise-leads-ai-search) shows.
- **Automation is a growth area.** PMMI and Interact Analysis estimate the US robotics market serving packaging and processing at more than [$440 million in 2025](https://www.pmmi.org/news/u-s-robotics-market-for-packaging-and-processing-poised-to-nearly-double-by-2031), reaching about $800 million by 2031, and 72% of surveyed end users already use robotics. Buyers who are new to a category lean on research more than on old relationships. Robot makers are covered in [how robotics companies reach manufacturers’ shortlists](https://underneath.agency/resources/robotics-companies-customers-ai-search).
- **Wrong facts cost deals.** An outdated spec, a discontinued model or a dealer that no longer represents you can send a buyer elsewhere. Our guide to [fixing wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) explains how to trace the source.

## How does GEO work for a machinery company?

Generative engine optimization (GEO) makes your machines easy for AI assistants to match to an application and verify.

For a machine builder, the work usually covers:

1. **Application pages with numbers.** For each machine family, the parts, materials or packages it handles, with cycle times, changeover times, tolerances, footprint and utilities written as text.
2. **Price and financing clarity.** Price ranges or starting prices where you can publish them, financing and leasing options, and what is included, kept consistent across your site and dealers.
3. **Payback content with assumptions shown.** Worked payback examples and calculators that state shift patterns, labor rates and utilization, so a buyer and an assistant can check them.
4. **Dealer and service coverage in text.** Each dealer, showroom and service center listed by region with what it supports, matching the dealers’ own pages.
5. **Demo videos that say the facts.** Spoken specifications, accurate captions and clear titles on run-off and application videos.
6. **Independent proof.** Customer stories with named plants where permitted, trade publication features and show coverage. Our guide on [why well-known brands miss AI recommendations](https://underneath.agency/resources/why-well-known-brands-miss-ai-recommendations) explains why recognition alone is not enough.
7. **Fair comparisons.** Pages on how your machine compares for a given application. See our review of [whether comparison pages help B2B brands get cited](https://underneath.agency/resources/do-comparison-pages-help-b2b-ai-citations).
8. **Measurement.** Run the same application, budget, comparison and dealer questions in ChatGPT, Gemini, Perplexity, Copilot and Google’s AI features on a regular schedule, and set the results beside your demo and quote requests.

No provider can promise a recommendation. This work makes your machines the ones an assistant can describe accurately and a buyer can verify quickly.

## What is still unproven about AI in machinery buying?

Whether, and how often, AI answers change which builders get the demo or the order.

- **No direct measurement.** The surveys above describe research habits and capital plans, not AI-driven machine orders.
- **Cross-industry evidence.** Forrester’s study covers business buying in general, and our studies covered other categories; applying them to machinery is our inference.
- **Old financing data.** The Foundation’s 79.3% figure describes 2021 acquisitions; current shares may differ.
- **Forecasts move.** AMT noted that forecasts had expected flat or slightly lower orders in 2026 before the strong first five months.

## Where should a machinery builder start?

Start with the applications that bring your best orders and ask assistants what they recommend for them.

A first review shows whether your machines are named for those applications and budgets, whether specs, prices and dealers are described correctly, which videos, publications and listings the answers use, and which builders appear in your place.

If your sales plan depends on more demos and quotes this year, [contact us to review your machines’ visibility in AI answers](https://underneath.agency/contact). We will show where assistants recommend your machines, why other builders are named instead, and which changes are most likely to bring more qualified demo and quote requests through your team and dealers. Our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page explains how we put application pages, dealer coverage and demo videos into a form both assistants and plant buyers can check.

## Frequently asked questions

### Do plant managers really use ChatGPT to choose machines?

Surveys show most technical buyers use generative AI somewhere in purchasing, often to start. The decision still comes after demos, test cuts and a payback case.

### Should we publish machine prices?

Where you can, publish ranges or starting prices with what is included. Assistants look for prices, and buyers use them to decide whether to ask for a demo.

### Do our dealers’ websites matter for AI visibility?

Yes. Assistants may cite dealer pages, so specs, models and coverage there should match yours, and outdated dealer listings should be fixed.

### Are trade shows still worth it if buyers use AI?

Yes. Shows like PACK EXPO and IMTS create the demonstrations and coverage that buyers, and assistants, later find.

## Sources

- Automation.com (2026-07-13), [$583.4M in new machinery orders highlight US economic strengths](https://www.automation.com/article/583.4-million-usd-new-machinery-orders-us-economic-strengths)
- Modern Distribution Management (2026-01), [December Metalworking Machinery Orders Set New Monthly Record](https://www.mdm.com/news/research/economic-trends/december-metalworking-machinery-orders-set-new-monthly-record/)
- PMMI (2026-09-14), [PMMI’s 2026 State of the Industry Report reveals a packaging machinery market driven by flexibility, automation and new North American opportunities](https://www.pmmi.org/news/pmmis-2026-state-of-the-industry-report-reveals-a-packaging-machinery-market-driven-by-flexibility-automation-and-new-north-american-opportunities)
- PMMI (2026), [US robotics market for packaging and processing poised to nearly double by 2031](https://www.pmmi.org/news/u-s-robotics-market-for-packaging-and-processing-poised-to-nearly-double-by-2031)
- Industrial Info Resources (2026-09), [US Manufacturing Boosts Capex Expectations](https://www.industrialinfo.com/news/article/us-manufacturing-boosts-capex-expectations--362933)
- Equipment Leasing & Finance Foundation (2022-10), [Equipment Leasing & Finance Industry Horizon Report fact sheet](https://www.leasefoundation.org/wp-content/uploads/2022/10/2022-Horizon-Report-Fact-Sheet.pdf)
- Supply Chain 24/7 (2026), [Equipment deal volume to reach $129 billion in 2026](https://www.supplychain247.com/article/equipment-deal-volume-to-reach-129-billion-in-2026)
- Forrester (2026-01-21), [Forrester’s 2026 Buyer Insights: GenAI Is Upending B2B Buying](https://www.forrester.com/press-newsroom/forrester-2026-the-state-of-business-buying/)
- TREW Marketing and GlobalSpec (2026), [State of Marketing to Engineers research report](https://advertising.globalspec.com/state-of-marketing-to-engineers-research-report)
- Google (2025-03-05), [Expanding AI Overviews and introducing AI Mode](https://blog.google/products/search/ai-mode-search/)
- OpenAI Help Center (2026), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Underneath (2026), [Does rewording a question change AI brand recommendations?](https://underneath.agency/research/ai-prompt-phrasing-study)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [The YouTube videos Google’s AI cites are small](https://underneath.agency/research/ai-overview-youtube-videos-study)
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude agree on brands?](https://underneath.agency/research/ai-assistants-brand-agreement-study)

---

This is the Markdown twin of https://underneath.agency/resources/machinery-companies-buyers-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How management consulting firms win clients in AI search"
description: "Only if its reputation is visible in sources AI reads: rankings, client perception studies, case evidence and independent coverage, not just its own claims."
canonical: "https://underneath.agency/resources/management-consulting-firms-clients-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# When a CEO asks AI which consulting firm to hire, will ours be named?

Only if your reputation is visible in the sources an assistant reads: rankings, client perception studies, published case evidence and independent coverage. Most CEOs now say they are comfortable acting on AI-generated input, and the assistants document that they search the web before answering, so a firm’s standing increasingly depends on what can be found and checked about it.

This article is for leaders of management consulting firms that sell strategy and operations work to senior executives, boards and private equity owners. It covers how reputation, rankings, case evidence and proposals translate into AI answers. Specialist boutiques in general and technology consulting are separate questions.

## The short version

1. CEOs are ready to use AI in big decisions: in [IBM’s 2026 CEO study](https://newsroom.ibm.com/2026-05-04-ibm-study-ceos-are-reshaping-c-suite-roles-for-the-ai-era) of 2,000 chief executives, 64% said they are comfortable making major strategic decisions based on AI-generated input.
2. The market rewards scale: in the UK, large firms earned 77% of consulting fee income in 2025, medium firms 20% and small firms 3%, according to the [Management Consultancies Association](https://www.consultancy.uk/news/44832/rising-exports-drove-uk-consulting-industry-growth-in-2025).
3. Rankings are built on client recommendations: [Forbes’ 2026 list](https://finance.yahoo.com/news/meet-america-best-management-consulting-134541359.html) of America’s best management consulting firms surveyed more than 1,250 clients across 33 categories and named 183 firms.
4. Demand has shifted toward AI work: [BCG](https://aijourn.com/bcg-reports-14-4-billion-in-revenue-marking-22nd-consecutive-year-of-growth/) reported $14.4 billion in 2025 revenue, with AI- and technology-focused services over 40% of the total.
5. Assistants look for rankings: in our own study, ChatGPT ran a search aimed at a named publication, ranking or award in 43.8% of its answers to buyer questions.

## Who hires a management consulting firm, and how concentrated is the market?

Chief executives, functional leaders, boards and investors, and most of the fees go to the largest firms.

Management consulting is bought at the top. The sponsor is usually a chief executive, chief financial officer, chief operating officer or business-unit head, often with a board or private equity owner behind the decision. [Forrester’s 2026 State of Business Buying](https://www.forrester.com/press-newsroom/forrester-2026-the-state-of-business-buying/) found the typical business purchase involves 13 internal stakeholders and nine external influencers, and procurement was a decision-maker in 53% of buying cycles.

The fees are concentrated. The Management Consultancies Association and Oxford Economics estimate the UK consulting market rose 3% to £21.8 billion in 2025, with large firms generating 77% of fee income. Member firms forecast 6% growth in 2026 and 8% in 2027. At the very top, BCG reported 7% revenue growth to $14.4 billion in 2025 and said AI services grew 25% year over year.

What is the client buying now? Mostly change it can measure. In Forbes’ 2026 coverage, BCG’s North America chair said AI now constitutes a quarter of BCG’s business, and leaders at Deloitte, EY, IBM and Alvarez & Marsal described clients asking for results and execution rather than presentations. That shift matters for AI visibility, because outcomes and named case evidence are things an assistant can find and repeat; reputation alone is not. Firms that lead with technology delivery are covered in [how technology consultancies win enterprise work](https://underneath.agency/resources/technology-consulting-firms-ai-search).

## Where does AI already sit when executives choose an adviser?

In the research and framing before a firm is called, by executives who increasingly trust AI input in decisions.

We found no published survey on how executives use AI to choose a management consulting firm specifically. The closest evidence:

| Finding | Share | Source |
|---|---|---|
| CEOs comfortable making major strategic decisions based on AI-generated input | 64% | IBM, 2026 |
| Organizations with a Chief AI Officer (26% a year earlier) | 76% | same |
| Business buyers who used generative AI in a recent purchase, mainly to gather information on vendors and products | 45% | [Gartner, 2026](https://www.gartner.com/en/newsroom/press-releases/2026-05-20-gartner-survey-finds-sixty-nine-percent-of-b-two-b-buyers-turn-to-sales-reps-to-validate-ai-generated-insights) |
| Business buyers who prefer to validate AI-generated insights with sales reps | 69% | same |

Forrester describes generative AI searches as the starting point for business buyers, followed by validation through colleagues and external influencers, and it advises providers to ensure their claims can be validated through trusted external voices. For a consulting firm, those voices are former clients, rankings, analysts and the business press. For a related board-level choice, see [how directors vet executive search firms](https://underneath.agency/resources/executive-search-firms-clients-ai-search).

Our inference: an executive who asks an assistant to frame a cost program or an AI operating model will also ask who does that work well. The answer shapes which firms are called for an introductory conversation, before any formal proposal process starts.

## What do executives ask AI about consulting firms?

Questions about who has solved this problem before, how firms compare, and whether a particular firm is credible.

The prompts below are our own illustrations of the kinds of questions a sponsor might type. They were not collected from real executives, and we do not claim any assistant answers them in a particular way.

| Intent | Illustrative prompt |
|---|---|
| Who does this work | “Which consulting firms are strongest in operations turnarounds for private equity portfolio companies?” |
| Comparison | “McKinsey, BCG or a mid-sized firm for a cost transformation at a regional bank?” |
| Track record | “Consulting firms that have led post-merger integration in specialty chemicals” |
| Credibility | “Is this firm’s AI strategy practice credible, and who are its clients?” |
| Rankings | “Which firms rank highest for pricing strategy?” |
| Fees and model | “Do management consultants offer outcome-based fees for procurement savings?” |

Two of these depend on rankings and named clients. That is where large firms hold an edge, and where a mid-sized firm with strong results in one sector can still be named if those results are published and confirmed elsewhere. Smaller specialist firms are covered in [how boutique consulting firms get shortlisted](https://underneath.agency/resources/consulting-firms-clients-ai-search). Supply chain specialists face this test too, as [how supply chain consultancies earn shortlist places](https://underneath.agency/resources/supply-chain-consulting-firms-ai-search) shows.

## How does an AI answer become a consulting mandate?

Through awareness, a shortlist, conversations, a proposal and procurement, then follow-on work.

1. **Awareness.** The sponsor knows some firms already. AI answers, rankings, peers and thought leadership add or remove names.
2. **Shortlist.** Source Global Research tracks this stage directly. Its 2025 study of client perceptions in financial services, [described by KPMG](https://kpmg.com/xx/en/about/industry-analyst-accolades/kpmg-most-well-known-firm-in-financial-services.html), surveyed 641 executives, directors and senior managers as they move from awareness to shortlisting a firm to their experience as a client.
3. **Conversations.** Partners meet the sponsor. Gartner found 69% of business buyers prefer to validate AI-generated insights with a person.
4. **Proposal and procurement.** Firms submit proposals; procurement compares scope, team and fees.
5. **Engagement and follow-on.** A first phase leads to implementation, new business units and repeat work.

AI visibility mainly affects steps 1 and 2. Partners, proposals and price decide the rest. But a firm that is never in the awareness set cannot be shortlisted, and Source’s own model of the funnel starts with awareness for that reason.

## What decides whether an assistant names your firm?

Prominence and independent evidence, as far as research shows; the platforms explain how they search, not how they choose.

The platforms describe their search, not their choices. [OpenAI explains](https://help.openai.com/en/articles/9237897-chatgpt-search) that ChatGPT search turns a sponsor’s question into one or more narrower queries for its search providers, and that a firm’s site can be included only if it lets OpenAI’s crawler, OAI-SearchBot, in. [Google describes](https://blog.google/products/search/ai-mode-search/) AI Mode splitting a question into subtopics and searching each, a method it calls “query fan-out.” Neither publishes how firms are selected for an answer.

What our studies observed, across buyer questions in several industries rather than consulting:

- **Assistants search for rankings and awards.** In [our study of hidden searches](https://underneath.agency/research/ai-hidden-searches-study), ChatGPT ran a search aimed at a named publication, ranking or award in 43.8% of its answers.
- **Prominence carries the Wikipedia effect.** In [our brand entity study](https://underneath.agency/research/brand-entity-ai-recommendations-study), 60.0% of the options all four assistants named had an English Wikipedia article, against 20.6% of options only one assistant named, but most of that gap went once brand prominence was accounted for. Outside coverage mattered most: a brand named on ten times as many independent sites within the cited pages had 4.7 times the odds of being recommended.
- **Assistants rarely agree on the top pick.** In [our brand agreement study](https://underneath.agency/research/ai-assistants-brand-agreement-study), all four assistants recommended the same first option for only 10.0% of questions.
- **Some lists rank their own author first.** In [our study of self-ranking lists](https://underneath.agency/research/self-promoting-best-lists-study), 24.2% of AI-cited numbered “best X” lists with an identifiable publisher put that publisher first.

Rankings built on client recommendations are among the independent sources available to consulting firms. Forbes and Statista combined a client survey and a peer survey, scored firms in 16 client industries and 17 functional areas, and handed out 930 awards. Five firms (Bain, BCG, Deloitte, McKinsey and PwC) received the maximum 32 star ratings. The [Financial Times ranking](https://www.consultancy.uk/news/43204/financial-times-ranking-names-best-uk-consulting-firms-for-2026), also compiled with Statista, gave Deloitte 25 gold ratings and listed more than 170 other firms in its silver and bronze tiers, including mid-market, boutique and specialist firms.

Our inference for management consulting: assistants have more to work with for a firm that appears in category rankings, client perception studies and business press, and that publishes case evidence with named sectors and measured results. A firm that relies on private relationships and confidential work leaves little for an assistant to find. Even famous names get skipped, as [why well-known brands miss AI recommendations](https://underneath.agency/resources/why-well-known-brands-miss-ai-recommendations) explains.

## What does a consulting firm lose when AI leaves it out?

A place in the awareness set that feeds every shortlist; the size of that loss has not been measured.

We found no study that measures consulting mandates lost to AI absence, so we set out the reasoning and label it:

- **Concentration can deepen.** With large UK firms already earning 77% of fee income, and our study showing prominence drives much of what assistants name, we infer that AI answers may reinforce the lead of the best-known firms unless smaller firms make their sector strength visible. Our article on [whether AI assistants favor big brands](https://underneath.agency/resources/do-ai-assistants-favor-big-brands) reviews the wider evidence.
- **The category gets framed without you.** If an assistant explains a problem using another firm’s method or research, that firm enters the conversation with an advantage.
- **Old facts persist.** A practice you closed, a partner who left or a sector you no longer serve can appear in answers. Our guide to [fixing wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) explains how to trace and correct the source.

## How does GEO work for a management consulting firm?

Generative engine optimization (GEO) makes your expertise easy for AI assistants to find, describe and confirm independently.

For a management consulting firm, the work usually covers:

1. **Case evidence in text.** Engagement summaries with sector, problem, approach and measured outcome, published with client permission or anonymized with enough detail to be useful.
2. **Partner and practice pages.** Who leads each practice, what they have done and where they have written or spoken, consistent with their public profiles.
3. **Rankings and perception results, stated plainly.** Where an independent ranking or client perception study names you, say which one, which year and which category, and link to it.
4. **Research others cite.** Thought leadership with original data that business publications and analysts reference, because citations by others are the independent coverage assistants appear to weigh.
5. **Independent coverage.** Business and trade press, conference platforms, industry association roles and, where genuinely notable, encyclopedia entries. Our guide on [why Wikipedia matters for AI search](https://underneath.agency/resources/why-wikipedia-matters-for-ai-search) explains what that can and cannot do.
6. **Consistent facts.** The same firm name, practices, offices and leadership across your site, rankings, directories and press.
7. **Crawl access and measurement.** Allow the crawlers the assistants document, then ask a fixed set of sponsor questions across ChatGPT, Gemini, Perplexity, Copilot and Google’s AI features, repeatedly, recording who is named and which rankings and pages are cited.

No adviser can promise that an assistant will put your firm in front of a sponsor, so be wary of anyone who does. Our review of [best-of lists and AI recommendations](https://underneath.agency/resources/best-of-lists-ai-recommendations) covers how ranked lists feed into AI answers more broadly.

## How many consulting mandates actually begin with an AI answer?

Nobody has measured that yet; the evidence shows executives open to AI input, not mandates won through it.

- **No consulting-specific study of AI in firm selection.** The IBM, Gartner and Forrester figures cover executives and business buyers in general. Applying them to consulting is our inference.
- **Interested parties.** IBM, BCG, KPMG, Deloitte and EY are consulting firms themselves; rankings are produced by publishers that also sell to the industry.
- **No attribution data.** We found no public evidence linking AI answers to proposals requested or fees won.
- **Our studies are cross-industry.** They covered buyer questions in several sectors, not consulting, and assistants change how they search over time.

## Where should a management consulting firm start?

Start by asking assistants the questions your sponsors ask before they call a firm.

Ask them by sector, problem and comparison, including the questions where your best competitors are strong. Then look at whether your firm is named, whether your practices and results are described correctly, which rankings, studies and articles the answers cite, and which firms appear instead. The work that follows is to publish the evidence of your results and earn the independent recognition that confirms it.

If your growth depends on being shortlisted for a few large mandates each year, [request a review of your AI visibility](https://underneath.agency/contact). We will show where your firm appears when executives research consulting firms with AI, why others are named instead, and which changes are most likely to put you on more shortlists. The page on our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page describes how that work runs for a firm, from publishing case evidence to tracking which sponsor questions name you.

## Frequently asked questions

### Do CEOs really use AI to choose consultants?

There is no direct measure yet. In IBM’s 2026 study, 64% of CEOs said they are comfortable making major strategic decisions based on AI-generated input, and 45% of business buyers in Gartner’s survey used generative AI to research vendors.

### Do rankings like Forbes or the Financial Times matter for AI answers?

They are independent sources built on client and peer recommendations, and in our study ChatGPT ran a search aimed at a named ranking, publication or award in 43.8% of its answers. Whether a ranking is cited for a given question varies.

### Can a mid-sized firm compete with the largest consultancies in AI answers?

It is harder on general questions, where prominence dominates. It is more achievable on narrow questions about a sector, a problem or a region, where published case evidence and specialist rankings can set a firm apart.

### Should we publish client results if our work is confidential?

Publish what clients allow, and anonymize the rest with sector, scale, problem and outcome. An assistant cannot repeat evidence that exists only in proposals.

## Sources

- IBM Institute for Business Value (2026-05-04), [IBM Study: CEOs are reshaping C-suite roles for the AI era](https://newsroom.ibm.com/2026-05-04-ibm-study-ceos-are-reshaping-c-suite-roles-for-the-ai-era)
- Consultancy.uk, Management Consultancies Association and Oxford Economics (2026-07), [Rising exports drove UK consulting industry growth in 2025](https://www.consultancy.uk/news/44832/rising-exports-drove-uk-consulting-industry-growth-in-2025)
- Forbes via Yahoo Finance (2026-03-17), [Meet America’s Best Management Consulting Firms 2026](https://finance.yahoo.com/news/meet-america-best-management-consulting-134541359.html)
- Consultancy.uk (2026-02-18), [Financial Times ranking names best UK consulting firms for 2026](https://www.consultancy.uk/news/43204/financial-times-ranking-names-best-uk-consulting-firms-for-2026)
- Boston Consulting Group via The AI Journal (2026-04-23), [BCG Reports $14.4 Billion in Revenue, Marking 22nd Consecutive Year of Growth](https://aijourn.com/bcg-reports-14-4-billion-in-revenue-marking-22nd-consecutive-year-of-growth/)
- KPMG (2025), [KPMG most well-known firm in financial services](https://kpmg.com/xx/en/about/industry-analyst-accolades/kpmg-most-well-known-firm-in-financial-services.html)
- Forrester (2026-01-21), [Forrester’s 2026 Buyer Insights: GenAI Is Upending B2B Buying](https://www.forrester.com/press-newsroom/forrester-2026-the-state-of-business-buying/)
- Gartner (2026-05-20), [Gartner survey finds 69% of B2B buyers turn to sales reps to validate AI-generated insights](https://www.gartner.com/en/newsroom/press-releases/2026-05-20-gartner-survey-finds-sixty-nine-percent-of-b-two-b-buyers-turn-to-sales-reps-to-validate-ai-generated-insights)
- OpenAI Help Center (2026), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Google (2025-03-05), [Expanding AI Overviews and introducing AI Mode](https://blog.google/products/search/ai-mode-search/)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [Do Wikipedia and schema make AI assistants recommend a brand?](https://underneath.agency/research/brand-entity-ai-recommendations-study)
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude agree on brands?](https://underneath.agency/research/ai-assistants-brand-agreement-study)
- Underneath (2026), [How many “best of” lists cited by AI rank their own brand first?](https://underneath.agency/research/self-promoting-best-lists-study)

---

This is the Markdown twin of https://underneath.agency/resources/management-consulting-firms-clients-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How do marketing software companies win buyers who ask AI first?"
description: "By being named when marketing leaders ask AI to map a category: most CMOs now start software searches in AI tools, and they shortlist about three vendors."
canonical: "https://underneath.agency/resources/marketing-software-ai-search-growth"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How do marketing software companies win buyers who ask AI first?

By earning a place in the answers marketing leaders get when they ask an AI assistant to map a category, then meeting those buyers with the pricing, demos and proof they expect. Marketing leaders are among the earliest adopters of AI at work, and in 2026 most of them start their software searches there. The channel is real and moving fast; what nobody has yet measured is how much closed revenue it causes.

## The short version

1. In [Wynter’s](https://wynter.com/cmo-b2b-saas-buyer-journey-report-2026) January 2026 survey of 101 CMOs at B2B software companies, 84% used AI tools to discover vendors, up from 24% a year earlier, and 68% started their search there before Google. Only 9% started by googling a software category.
2. The shortlist is tiny and decided early: 62% of those CMOs evaluate exactly three vendors, and 48% arrive at the first sales call very familiar with the vendor, up from 22%.
3. The category is crowded. [Chiefmartec’s 2025 landscape](https://chiefmartec.com/2025/05/2025-marketing-technology-landscape/) counted 15,384 marketing technology products in 49 categories, so an AI answer that names three of them is doing a lot of filtering.
4. Marketing software vendors report the channel converting: at [Ahrefs](https://ahrefs.com/blog/ai-search-traffic-conversions-ahrefs/), 0.5% of traffic came from AI search but produced 12.1% of signups, and [HubSpot](https://transcripts.platformaeronaut.com/transcripts/HUBS-2Q25-transcript) told investors that just 10% of its leads now come from blog traffic.
5. A won customer is worth years of revenue: [HubSpot](https://seekingalpha.com/pr/20396604) averaged $11,683 in subscription revenue per customer in late 2025, and [Klaviyo](https://s25.q4cdn.com/801988311/files/doc_news/Klaviyo-Delivers-Outstanding-2025-Results-32-Revenue-Growth-Record-Fourth-Quarter-and-Raised-Fiscal-Year-2026-Outlook-2026.pdf) reported net revenue retention of 110%.

## Who signs off on marketing technology, and how much is one customer worth?

The CMO leads strategic platform purchases, a committee of about five signs off, and a won account usually grows.

Wynter found the average buying committee in a marketing purchase now has 5.1 members: the CMO, finance, IT and security, revenue operations and the people who will use the tool. CMOs lead the research themselves for strategic platforms such as CRM and marketing automation, while their teams choose point tools. One CMO in the survey put it plainly: “If it touches revenue strategy or brand reputation, I’m doing the research myself.”

Those buyers work with tight budgets. [Gartner’s 2025 CMO Spend Survey](https://www.01net.it/gartner-2025-cmo-spend-survey-reveals-marketing-budgets-have-flatlined-at-7-7-of-overall-company-revenue) of 402 marketing leaders found budgets flat at 7.7% of company revenue, and most CMOs said they had too little budget to carry out their strategy. A new tool usually has to replace or consolidate an old one, which is why HubSpot’s results talk about customers who “consolidate tech stacks.”

What a customer is worth depends on the category, but public filings give a range. HubSpot ended 2025 with 288,706 customers. Klaviyo, the email and text messaging platform, had over 193,000 customers, of which 3,912 generated more than $50,000 a year in recurring revenue. [Semrush](https://s27.q4cdn.com/202248034/files/doc_financials/2025/q3/SEMR-Q3-2025-Earnings-Release.pdf) reported that its customers paying over $50,000 a year grew by over 72% in a year. Because these are subscriptions that expand, a buyer first met in an AI answer can be worth several years of renewals and add-ons, not one sale.

## At what stage do CMOs bring AI assistants into a martech search?

At the very start, alongside peer communities, with Google demoted to checking facts.

Wynter describes the 2026 pattern this way: AI tools and peer communities start the search together, Google verifies, and a shortlist comes out the other end. Google is still used by 72% of CMOs, but mostly to check reviews and complaints. “Google is my fact-checker now, not my idea generator,” one respondent said. When CMOs ranked what most drives a vendor into consideration, AI recommendations (10%) edged out Google Search (8%), though word of mouth still led at 42%, and 65% start in peer communities such as private Slack groups.

This is the irony of the category. The people who buy marketing software are the people who adopted AI first: in [Salesforce’s State of Marketing research](https://www.salesforce.com/au/marketing/streamlining-with-agentic-ai/), 75% of marketing organizations now use AI in at least one form. Marketing leaders who use these tools to write campaigns all day also use them to choose the tools.

Google’s own results are a second AI surface. Across the eight industries in [our study of 1,248 US searches](https://underneath.agency/research/ai-overviews-frequency-study), B2B software and technology keywords drew an AI Overview most often, on 96.0% of them. Marketing vendors are visible inside those answers: in [our AI Overview citations study](https://underneath.agency/research/ai-overview-citations-study), hubspot.com was cited in 29 AI Overviews, 6.0% of the 481 we captured, one of the few vendor sites among the most-cited domains.

Gartner Digital Markets adds one detail particular to this buyer. In its [2025 survey of 3,500 software buyers](https://cdn-static.bizzabo.com/bizzabo.file.upload/yVkIqwmWRDKbEhqeE3aQ_Gartner-Digital-Markets_2025-Report_Making-The-List.pdf), buyers from marketing and advertising businesses were the most likely to factor social media (47%) into which vendors make their first list. Marketing buyers listen to the conversation about a product, and AI answers draw on much of that same conversation.

## Which questions do marketing leaders ask AI about software?

Category maps, stack fit, alternatives, attribution and price, usually with their own constraints attached.

These sample prompts are ours, written to mirror how a CMO or marketing ops lead might ask about different martech categories; they are not logged queries:

- Marketing automation: “Best marketing automation platform for a small B2B team already on Salesforce.”
- Email and messaging: “Which email marketing platform is best for a mid-size Shopify store?”
- Attribution and analytics: “How do Dreamdata and HockeyStack compare for multi-touch attribution with a long sales cycle?”
- SEO and AI visibility: “What tools track whether my brand appears in ChatGPT and Google AI Overviews?”
- Personalization and testing: “Alternatives to Optimizely for website personalization that do not need a developer.”
- Campaign management: “What do mid-market teams use to plan campaigns across paid, email and events in one place?”

Small changes in the constraint change the answer. In [our prompt phrasing study](https://underneath.agency/research/ai-prompt-phrasing-study), Perplexity named Klaviyo every time it was asked which email marketing platform is best for ecommerce, and never once when the same question added a tight budget; MailerLite moved the other way. Across all the brands we tracked that were named in three or more answers, 88.6% had a visibility that differed by 40 points or more between their best and worst wording. A marketing software company should know which version of its buyers’ question it wins and which it loses.

## How does an AI answer turn into pipeline for a marketing software company?

Through a three-vendor shortlist, then a trial or demo that confirms a decision already half made.

**The shortlist.** Wynter found 62% of CMOs evaluate exactly three vendors. In Gartner Digital Markets’ data, buyers start with an average of 4.4 options on their initial list, and 81% of buyers end up purchasing from that initial list most or all of the time. If the assistant leaves you off when the list is drawn, there are few later chances.

**The trial or demo.** Marketing buyers want to see the product before they talk to anyone. Wynter found 58% try an interactive demo before talking to sales, up from 46%, and 38% rank the sales demo as the top factor in the final choice. Gartner Digital Markets found 62% of software buyers call the product trial their top factor in the final purchase decision. “If I have to book a demo to see either, I’m gone,” one CMO said of product and price.

**The self-serve signup.** For product-led marketing tools, the effect shows up in signups. Ahrefs, which sells SEO software, found that AI search visitors converted at a 23x higher rate than traditional organic search visitors. [Semrush’s traffic study](https://www.semrush.com/blog/ai-search-seo-traffic-study/) estimated that the average AI search visitor is 4.4 times as valuable as a traditional organic one, based on conversion rate. Both are vendor-reported figures from companies whose own business is search.

**The sales-led deal.** For enterprise platforms, the AI answer rarely appears in referral data. We infer it shows up as a buyer who arrives already informed: Wynter found 80% of CMOs come to the first call moderately or very familiar with the vendor. That is the moment the AI answer has already paid off, or already cost you the deal. Sales tool vendors face the same handoff, described in [how sales software turns AI answers into pipeline](https://underneath.agency/resources/sales-software-pipeline-from-ai-search).

## Why do assistants name some martech vendors and pass over others?

Little is documented; studies point to marketers’ communities, reviews, independent coverage and consistent product facts.

**Documented by the platform.** [Google says](https://developers.google.com/search/docs/appearance/ai-features) AI Overviews and AI Mode may use a “query fan-out” technique, issuing multiple related searches across subtopics and data sources to build one response, and that there are no additional requirements to appear beyond normal search eligibility. A single question about attribution software can therefore pull in review pages, community threads and vendor documentation at once. Our [hidden searches study](https://underneath.agency/research/ai-hidden-searches-study) shows what those background searches look like.

**Observed in studies.** In our citations research, practitioner communities and video carried weight for software questions; see [our Reddit study](https://underneath.agency/research/ai-reddit-citations-study). Self-promotion is present but small. In [our study of AI-cited “best of” lists](https://underneath.agency/research/self-promoting-best-lists-study), 65 of 269 numbered lists (24.2%) ranked their own publisher first, and HubSpot and Klaviyo were among the publishers with more than one such list. And a self-ranking list’s top pick was not named detectably more often than an independent list’s top pick.

**Claimed by vendors.** HubSpot’s CEO told investors that HubSpot is now cited by AI models more than any other CRM and is driving conversions from the channel. That is a company statement, not an independent measurement. Our guide on [why assistants favor a few leading CRMs](https://underneath.agency/resources/crm-software-ai-search) looks at that category.

**Our inference for this industry.** Marketing buyers trust peers first and vendors last, so a reasonable expectation is that what practitioners say in communities, reviews, newsletters and podcasts carries more weight in AI answers than vendor pages do. Accuracy also matters: in [our pricing study](https://underneath.agency/research/ai-pricing-accuracy-study), one answer gave Semrush’s $117.33 annual-billing price for one plan as the price of another. Marketing buyers who see a wrong price may simply move on.

## Does it matter that your buyers already know how GEO works?

Yes. Marketing leaders now buy GEO themselves, so they will recognize, and punish, tactics that look like gaming.

Wynter found 34% of CMOs invest in generative engine optimization (GEO) in 2026, a category that registered 0% in 2025. The vendors have noticed too. [Adobe agreed to acquire Semrush](https://news.adobe.com/news/2025/11/adobe-to-acquire-semrush) for about $1.9 billion, describing it as a brand visibility platform built on GEO and search optimization, and chiefmartec counted the SEO subcategory, which now includes AI visibility tools, growing 24% in a year, from 212 to 262 products.

That has two consequences for marketing software companies. First, your buyers can check your visibility in minutes, so being absent from the answers for your own category is a credibility problem as well as a pipeline one. Second, a self-ranked comparison page or a wave of thin listicles will be read as exactly what it is. See [our guide to legitimate GEO versus manipulation](https://underneath.agency/resources/legitimate-geo-vs-manipulation) and [how optimizing for AI can backfire](https://underneath.agency/resources/can-geo-backfire-on-your-brand).

## What does a martech vendor lose when AI answers skip it?

Mostly invisible losses: deals that never reach your pipeline because the shortlist formed without you.

The clearest public example is HubSpot. On its second-quarter 2025 call, its CEO said that “organic search traffic is declining globally,” that AI Overviews are giving answers with fewer clicks, and that just 10% of HubSpot’s leads now come from blog traffic after years of diversifying into video, newsletters and podcasts. HubSpot built that buffer over years; most marketing software companies still rely on search content for a much larger share of demand.

Semrush’s study projects that digital marketing and SEO topics may send more visitors from AI search than from traditional search by early 2028. Treat that as a forecast, not a measurement. Meanwhile the category is consolidating: chiefmartec removed 1,211 products from its landscape in a year, a churn rate of 8.6%. In a market where buyers evaluate three vendors and budgets are flat, a vendor that AI assistants do not name loses ground quietly, and its analytics may not show it. Our piece on [why analytics miss AI visibility](https://underneath.agency/resources/why-analytics-miss-ai-visibility) explains that blind spot.

## Which GEO steps fit a marketing software company selling to marketers?

It gives assistants accurate, checkable material about your martech product; nobody can guarantee where it appears.

1. **One clear story everywhere.** Describe the product, category, ideal customer and integrations the same way on your site, review profiles, app marketplaces (Salesforce AppExchange, the HubSpot and Shopify app stores), documentation and LinkedIn.
2. **Practitioner coverage.** Pursue the newsletters, podcasts, communities and independent comparison pages marketing leaders read. Our article on [which pages to target](https://underneath.agency/resources/best-of-lists-ai-recommendations) covers how to choose them.
3. **Reviews with detail.** Keep a flow of recent, specific reviews on the platforms your buyers check, naming use cases and team sizes.
4. **Decision content, ungated and honest.** Publish integration guides, implementation timelines, data and privacy pages and fair comparison pages. CMOs in Wynter’s survey asked vendors to “show at least a starting price.”
5. **One current pricing page.** Retire old price figures so assistants have one source to quote.
6. **Measurement by question and by stage.** Track a fixed set of category, alternatives and comparison questions in ChatGPT, Gemini, Perplexity, Copilot and Google, in several wordings, and connect it to demo requests and self-reported attribution. Our guide to [designing AI visibility tracking](https://underneath.agency/resources/how-to-design-ai-visibility-tracking) explains the method.

For the general SaaS playbook, see [how B2B SaaS companies earn revenue from AI search](https://underneath.agency/resources/b2b-saas-revenue-from-ai-search).

## What remains unproven about AI answers and martech revenue?

Whether AI visibility causes martech subscriptions and renewals, or just travels with strong brands, is unmeasured.

The strongest buyer data comes from small surveys by companies that sell research or software: Wynter’s 101 CMOs work at mid-market software companies, so they are not every marketing buyer. The conversion figures come from Ahrefs and Semrush, which sell search tools and have a stake in the topic. HubSpot’s statements are company claims made to investors. No published study follows marketing software buyers from an AI answer through to a signed contract and renewal, and none isolates the effect of GEO work from brand strength, pricing or product. A vendor that wants its own answer can follow [how to prove GEO caused a change in sales](https://underneath.agency/resources/prove-geo-caused-sales), using its trial and demo data.

## How can a martech vendor see whether AI answers are feeding its trials and demos?

Check which AI shortlists name you for the questions CMOs ask before starting a trial or booking a demo.

Map the category, alternatives, comparison and pricing questions your buyers ask in each of your segments, check them across the main assistants in several wordings, and rank the gaps by the trials and pipeline they touch. When you want a second pair of eyes, [talk to us about a martech visibility audit](https://underneath.agency/contact); we tie the findings to your demo and trial numbers rather than to visibility alone. For the ongoing program, the [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page sets out the category questions we track for CMOs and the review and pricing fixes that follow.

## Frequently asked questions

### Do marketing leaders really use ChatGPT to choose marketing software?

Most now use AI somewhere in discovery. Wynter found 84% of B2B software CMOs used AI tools to discover vendors in 2026, but they combine it with peer communities and reviews, and word of mouth still ranked first for consideration.

### Is AI search traffic worth more for marketing software companies?

In the published cases, yes. Ahrefs and Semrush both report AI search visitors converting far better than traditional search visitors. Both are vendor-reported, and part of the gap may come from fewer, more decided visitors.

### Should a marketing software company publish “best tools” lists that rank itself first?

For a martech vendor, probably not. Our study found such lists are a small share of what AI answers cite, and we found no detectable advantage for the self-ranked list’s top pick. Your buyers are marketers and will recognize the tactic.

### Which marketing software categories are most exposed?

We have no category-by-category measurement. A reasonable expectation is that crowded, self-serve categories such as email, SEO tools and social scheduling feel it first, while enterprise platforms feel it later through shortlists and sales calls.

## Sources

- Wynter (2026), [How B2B SaaS CMOs buy software](https://wynter.com/cmo-b2b-saas-buyer-journey-report-2026)
- Chiefmartec, Scott Brinker (2025-05), [2025 Marketing Technology Landscape Supergraphic](https://chiefmartec.com/2025/05/2025-marketing-technology-landscape/)
- Gartner, via 01net (2025-05-12), [Gartner 2025 CMO Spend Survey Reveals Marketing Budgets Have Flatlined at 7.7% of Overall Company Revenue](https://www.01net.it/gartner-2025-cmo-spend-survey-reveals-marketing-budgets-have-flatlined-at-7-7-of-overall-company-revenue)
- Gartner Digital Markets (2025), [Making the List: How software buyers pare down their options](https://cdn-static.bizzabo.com/bizzabo.file.upload/yVkIqwmWRDKbEhqeE3aQ_Gartner-Digital-Markets_2025-Report_Making-The-List.pdf)
- Salesforce (2026), [10 Ways Agentic AI is Supporting Marketing Teams](https://www.salesforce.com/au/marketing/streamlining-with-agentic-ai/)
- HubSpot, via Seeking Alpha (2026-02-11), [HubSpot Reports Strong Q4 and Full Year 2025 Results](https://seekingalpha.com/pr/20396604)
- HubSpot (2025-08), [Second quarter 2025 earnings call transcript](https://transcripts.platformaeronaut.com/transcripts/HUBS-2Q25-transcript)
- Klaviyo (2026-02-10), [Klaviyo Delivers Outstanding 2025 Results](https://s25.q4cdn.com/801988311/files/doc_news/Klaviyo-Delivers-Outstanding-2025-Results-32-Revenue-Growth-Record-Fourth-Quarter-and-Raised-Fiscal-Year-2026-Outlook-2026.pdf)
- Semrush (2025-11-05), [Semrush Announces Third Quarter 2025 Financial Results](https://s27.q4cdn.com/202248034/files/doc_financials/2025/q3/SEMR-Q3-2025-Earnings-Release.pdf)
- Semrush (2025), [AI Search and SEO Traffic Case Study](https://www.semrush.com/blog/ai-search-seo-traffic-study/)
- Ahrefs (2025), [Does AI Search Traffic Convert Better Than Traditional Search? For Ahrefs, Yes](https://ahrefs.com/blog/ai-search-traffic-conversions-ahrefs/)
- Adobe (2025-11-19), [Adobe to Acquire Semrush](https://news.adobe.com/news/2025/11/adobe-to-acquire-semrush)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- Underneath (2026), [When does Google show an AI Overview?](https://underneath.agency/research/ai-overviews-frequency-study), [AI Overview citations](https://underneath.agency/research/ai-overview-citations-study), [prompt phrasing](https://underneath.agency/research/ai-prompt-phrasing-study), [self-promoting “best of” lists](https://underneath.agency/research/self-promoting-best-lists-study), [software pricing accuracy](https://underneath.agency/research/ai-pricing-accuracy-study), [hidden searches](https://underneath.agency/research/ai-hidden-searches-study) and [Reddit citations](https://underneath.agency/research/ai-reddit-citations-study)

---

This is the Markdown twin of https://underneath.agency/resources/marketing-software-ai-search-growth. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How marketplaces win buyers and sellers through AI search"
description: "By being the answer on both sides: where buyers should shop and where sellers should list. AI traffic is still small for marketplaces, but it converts."
canonical: "https://underneath.agency/resources/marketplace-buyers-sellers-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How does an online marketplace win more buyers and sellers when people ask AI?

By being named, and described accurately, in two kinds of answers: where a shopper should buy, and where a seller, host or freelancer should list. AI assistants are still a small source of marketplace traffic, but the visits they send are unusually ready to buy. The marketplaces preparing now are treating AI answers as a channel for both sides of the network, not just for demand.

## The short version

1. AI is a small but high-intent channel for marketplaces today: Etsy, with about 87 million active buyers, said in August 2026 that traffic from AI shopping experiences was [below 1% of its total](https://www.thecerbatgem.com/?p=10343104) but showed higher intent and higher average order value.
2. AI shoppers convert: Adobe found AI traffic to US retailers rose 393% in the first quarter of 2026 and [converted 42% better](https://techcrunch.com/2026/04/16/ai-traffic-to-us-retailers-rose-393-in-q1-and-its-boosting-their-revenue-too/) than other traffic in March.
3. Marketplaces are moving inside the assistants: Etsy sellers were the first on [ChatGPT’s Instant Checkout](https://openai.com/index/buy-it-in-chatgpt/), which OpenAI scaled back in March 2026, according to [FashionUnited](https://fashionunited.com/news/retail/us-consumers-would-let-ai-agents-buy-clothes-but-not-without-a-say/2026092974919), and [Upwork’s app in ChatGPT](https://stocks.observer-reporter.com/observerreporter/article/gnwcq-2026-4-9-upworks-work-marketplace-comes-to-chatgpt) lets businesses search a pool of more than 18 million professionals.
4. Marketplace pages matter most when people are ready to buy: in [our AI Overview study](https://underneath.agency/research/ai-overviews-frequency-study), marketplaces and retailers were 1.5% of all citations but 7.9% on transactional searches.
5. Trust answers come from review sites: in [our reputation study](https://underneath.agency/research/is-it-legit-ai-reputation-study), 88.0% of answers to “is this brand legit?” cited a review or complaint platform.

## How does a marketplace make money, and why do both sides matter?

A marketplace earns a cut of every transaction, so it grows only when buyers and sellers keep finding each other.

Andreessen Horowitz’s [marketplace glossary](https://a16z.com/the-marketplace-glossary/) defines the core economics plainly. The take rate is “the percentage of the gross merchandise value (GMV) captured by the marketplace.” Liquidity, “the likelihood that a seller is able to find a buyer, or that a buyer is able to find the product or service they’re looking for,” is called the most critical aspect: “Most marketplaces fail because they never reach or maintain liquidity.”

Public results show what is at stake:

| Marketplace | Latest public figure | What it means for AI visibility |
|---|---|---|
| Etsy | $2.6 billion in quarterly sales, take rate of 25.9%, about 87 million active buyers | Each new buyer is worth roughly $124 a year in sales at the current average |
| Airbnb | $27.2B in quarterly gross booking value | Guests and hosts both research the platform before committing |
| Upwork | More than $30 billion in total transactions since founding | Clients and freelancers each choose a platform before work starts |

Etsy’s trailing 12-month sales per active buyer rose 2.8% to $124 in the second quarter of 2026. At a take rate near 26%, every buyer an AI answer sends who becomes a regular adds revenue for years.

The supply side matters as much. Airbnb’s [second-quarter 2026 letter](https://s26.q4cdn.com/656283129/files/doc_financials/2026/q2/Airbnb-Q2-2026-Shareholder-Letter-FINAL.pdf) says more than 150,000 homes across World Cup host cities were listed on Airbnb for the first time, and it grew its Experiences supply by nearly 80% year over year. Every marketplace runs a constant campaign to recruit sellers, hosts or freelancers, and those people research their options too.

## Where do AI assistants already sit in marketplace journeys?

At the start of both journeys, and increasingly inside the assistant, where the transaction can begin without a website visit.

On the buyer side, the shift is visible in retail data. In Adobe’s survey, reported by TechCrunch, 39% of US consumers said they used AI for online shopping. AI-driven revenue per visit was 37% higher than for other traffic in March 2026, a reversal from a year earlier.

The assistants are also becoming storefronts:

- **Product marketplaces.** OpenAI launched Instant Checkout so that US ChatGPT users “can now buy directly from U.S. Etsy sellers right in chat.” OpenAI says more than 700 million people use ChatGPT each week. Accounts of what happened next differ: [Reuters, via Zawya](https://www.zawya.com/en/business/americas/retailers-tap-ai-shopping-traffic-but-fight-to-keep-customer-data-424957), reported that OpenAI ended Instant Checkout in March 2026 and is helping marketplaces such as Etsy use their own checkouts, FashionUnited says the feature was scaled back, and OpenAI’s help page still describes it for some eligible merchants.
- **Travel and property.** Booking.com, Expedia and Zillow were among the first apps inside ChatGPT, according to [Tom’s Guide](https://www.tomsguide.com/ai/you-can-now-use-apps-inside-chatgpt-including-spotify-canva-and-zillow-heres-how), with Thumbtack, TripAdvisor and others announced to follow.
- **Talent and services.** Upwork’s ChatGPT app lets businesses describe a project, find talent across 130 categories of work and draft a job post, then finish hiring on Upwork itself.

Marketplace leaders describe the channel as small but promising. Etsy said its AI traffic converts with higher intent. eBay’s chief executive said in 2025 that agentic commerce was a small part of the business but “[it has a nice growth rate](https://www.retaildive.com/news/ebay-ceo-ai-change-customer-experience/756794).”

## What do buyers and sellers ask AI about marketplaces?

Buyers ask where to find something and whether a seller is safe; sellers ask where they will earn most. The prompts below are illustrative, written by us, not observed data.

| Side | Illustrative prompt | What the answer decides |
|---|---|---|
| Buyer | “Where can I buy a handmade wedding guest book with custom names?” | Which marketplace or shop gets the visit |
| Buyer | “Is it safe to buy a used designer bag online, and which sites authenticate?” | Which platforms are trusted for the category |
| Buyer | “Find me a plumber who can come tomorrow morning” | Which services marketplace or app is used |
| Seller | “Best place to sell vintage clothing online: Etsy, eBay or Depop?” | Which platform gets the new seller |
| Seller | “Etsy fees vs selling on my own Shopify store” | Whether the seller joins a marketplace at all |
| Host | “Is it worth listing my apartment on Airbnb during a big event?” | Whether new supply arrives |
| Freelancer | “Upwork or Fiverr for a new web developer?” | Which talent pool grows |

The seller questions are the ones most marketplaces overlook. A marketplace usually tracks its buyer search terms closely. Few, we infer, check what an assistant says about their fees, payout speed, seller protection or audience when a would-be seller compares platforms. A seller who opts for their own store instead moves on to a related question: [how a new business chooses a payment provider](https://underneath.agency/resources/payment-companies-merchants-ai-search).

## How does an AI answer become marketplace revenue?

Through a new buyer who keeps buying, or a new seller whose listings make the marketplace more useful for everyone.

**The buyer path.** A shopper asks a question, the answer names a marketplace or a listing, the shopper clicks or checks out in the assistant, and the marketplace takes its share. For a repeat buyer, the value compounds: at Etsy, $124 a year in sales per active buyer.

**The seller path.** A would-be seller asks where to list, the answer compares platforms, the seller signs up, lists inventory and, ideally, makes a first sale soon. Each new seller adds choice, which in a two-sided market makes the platform more valuable to buyers. The a16z glossary calls this the two-sided network effect: the network “becomes more valuable as the number of users on the other side of the marketplace increases.”

**The in-assistant path.** When the purchase or booking starts inside ChatGPT, as it could with Instant Checkout and can with the Upwork app, the marketplace may never get the first visit. The transaction still flows through it, but the assistant controls the first impression.

We suggest measuring AI visibility on both sides: buyer sessions and orders from AI sources, and seller or host sign-ups that name an AI assistant in an onboarding survey.

## What decides whether an assistant names your marketplace or a listing on it?

Partly documented rules, partly observed patterns, and partly inference; the platforms publish little about choosing between marketplaces.

**Documented by OpenAI.** In its [shopping help page](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search), OpenAI says merchants for a product “are ranked based on factors like availability, price, quality, and whether they are the maker or primary seller of that item.” That last factor matters for marketplaces: when a product is also sold on the maker’s own site, ChatGPT may favor the maker, not the platform listing it. Google says its AI features can run [several related searches](https://developers.google.com/search/docs/appearance/ai-features) before answering.

**Observed in studies.**

- Marketplaces earn more citations when the question is about buying. In our frequency study, their share rose from 1.5% of all AI Overview citations to 7.9% on transactional searches.
- In product questions, [Uberti-Bona Marin and colleagues](https://arxiv.org/abs/2609.18729) found retailers and marketplaces supplied 17.8% of the sources Google’s AI Overviews displayed. The same study found ChatGPT and Gemini interfaces shared only 5.4% of domains on average for the same question, so visibility in one assistant says little about another.
- Reputation answers lean on review sites. In our reputation study, Trustpilot and the BBB accounted for 61.7% of review-platform citations. A marketplace’s buyer-protection record and complaint handling are part of that public record.

**Our inference.** Assistants can only describe what they can read. A marketplace whose category pages, seller fee pages and trust policies are clear, crawlable and consistent with third-party coverage gives the assistant something accurate to repeat. Adobe also warned that around 34% of retail product pages cannot be properly accessed by AI; listing pages built only with scripts risk the same.

## What does GEO look like for a two-sided marketplace?

Generative engine optimization (GEO) for a marketplace runs two programs at once: one for buyers, one for supply.

For buyers:

1. **Category answers, not just category grids.** Pages that explain what to look for, price ranges and how your sellers differ, so an assistant can cite them for “where to buy” questions. Guidance on product content is in [what kind of product content AI shopping assistants prefer](https://underneath.agency/resources/product-content-ai-shopping-assistants-prefer).
2. **Readable listings.** Listing and category pages that load without scripts, with structured product, price and availability data.
3. **Trust in public.** Buyer protection, returns, authentication and dispute rules stated plainly, plus active work on Trustpilot and BBB profiles.

For sellers, hosts or freelancers:

1. **Plain seller economics.** Fees, payout timing, audience size and seller protection on public pages, so comparisons in AI answers are accurate.
2. **Third-party comparisons.** Coverage in the guides and communities sellers trust when they compare platforms, starting with the [best-of lists AI answers draw on](https://underneath.agency/resources/best-of-lists-ai-recommendations).
3. **Proof of earnings.** Named seller stories and published marketplace statistics that writers and assistants can quote.

Across both sides: consistent facts about your brand, presence in assistant app directories where your category is offered, and regular checks of how each assistant describes you. No one can guarantee an assistant will recommend a marketplace; GEO makes accurate, favorable evidence easier to find. The same shift looks different from the brand side; [will AI search reduce our dependence on marketplaces and aggregators?](https://underneath.agency/resources/ai-search-marketplace-dependence) covers sellers hoping to rely on you less.

## What does the evidence not tell marketplaces yet?

It shows AI shoppers convert well; it does not show AI answers recruit sellers, or how assistants choose between marketplaces.

- **Seller-side evidence is missing.** We found no public data on how many sellers, hosts or freelancers choose a platform after asking an assistant.
- **Retail data is not marketplace data.** Adobe’s figures cover US retail sites in general; Etsy’s and eBay’s comments are the only marketplace-specific signals we found, and both describe a small channel.
- **In-assistant sales may hide the source.** When checkout happens inside ChatGPT, standard analytics may not show where the buyer started.
- **Answers vary.** With only 5.4% of displayed domains shared between ChatGPT and Gemini, one snapshot from one assistant is not a measurement, for reasons laid out in [why AI answers about a brand keep changing](https://underneath.agency/resources/why-ai-answers-about-your-brand-change).

## Where should a marketplace start?

Start by asking the assistants your buyers’ and your sellers’ questions, and compare what they say with your own facts.

A useful first check covers both sides: where-to-buy questions in your strongest categories, trust questions about buying on your platform, and the seller questions a new host or freelancer would ask before choosing you or a rival. It shows whether you are named, which competitors and sources appear instead, and whether your fees and protections are described correctly.

If your growth depends on new buyers and on steady new supply, [talk to us about checking both sides of your marketplace](https://underneath.agency/contact). We will map how assistants describe your marketplace to buyers and to sellers, show where rivals or makers’ own sites take the answer, and plan the content, coverage and reputation work to close those gaps on both sides. How we run that work for shoppers and for sellers, hosts or freelancers is laid out on our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page.

## Frequently asked questions

### Do AI assistants send real buyers to marketplaces yet?

Yes, but few so far. Etsy said in August 2026 that AI shopping traffic was below 1% of its total, with higher intent and higher average order value than average.

### Can ChatGPT send a buyer to a maker instead of our marketplace?

It can. OpenAI says ChatGPT ranks merchants partly on “whether they are the maker or primary seller” of an item, so a maker’s own store can appear ahead of a marketplace listing.

### Should a marketplace build an app inside ChatGPT?

It depends on the category. Booking.com, Expedia, Zillow and Upwork have, but an app does not replace being named in ordinary answers, where most questions start.

### Why should sellers’ questions matter to a marketplace’s growth team?

Because supply drives liquidity. If assistants misstate your fees or audience when a seller compares platforms, you lose supply that would have made the marketplace better for buyers.

## Sources

- Andreessen Horowitz (2020-02-18), [The Marketplace Glossary](https://a16z.com/the-marketplace-glossary/)
- The Cerbat Gem (2026-08-09), [Etsy Q2 Earnings Call Highlights](https://www.thecerbatgem.com/?p=10343104)
- Airbnb (2026-08-06), [Q2 2026 Shareholder Letter](https://s26.q4cdn.com/656283129/files/doc_financials/2026/q2/Airbnb-Q2-2026-Shareholder-Letter-FINAL.pdf)
- Upwork, via GlobeNewswire (2026-04-09), [Upwork’s Work Marketplace Comes to ChatGPT](https://stocks.observer-reporter.com/observerreporter/article/gnwcq-2026-4-9-upworks-work-marketplace-comes-to-chatgpt)
- TechCrunch (2026-04-16), [AI traffic to US retailers rose 393% in Q1, and it’s boosting their revenue too](https://techcrunch.com/2026/04/16/ai-traffic-to-us-retailers-rose-393-in-q1-and-its-boosting-their-revenue-too/)
- OpenAI (2025-09-29), [Buy it in ChatGPT: Instant Checkout and the Agentic Commerce Protocol](https://openai.com/index/buy-it-in-chatgpt/)
- Reuters, via Zawya (2026), [Retailers tap AI shopping traffic but fight to keep customer data](https://www.zawya.com/en/business/americas/retailers-tap-ai-shopping-traffic-but-fight-to-keep-customer-data-424957)
- FashionUnited (2026-09-29), [US consumers would let AI agents buy clothes, but not without a say](https://fashionunited.com/news/retail/us-consumers-would-let-ai-agents-buy-clothes-but-not-without-a-say/2026092974919)
- OpenAI Help Center (2026), [Shopping with ChatGPT Search](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search)
- Tom’s Guide (2025-10), [You can now use apps inside ChatGPT, including Spotify, Canva and Zillow](https://www.tomsguide.com/ai/you-can-now-use-apps-inside-chatgpt-including-spotify-canva-and-zillow-heres-how)
- Retail Dive (2025-08-06), [eBay CEO: AI “continues to fundamentally change” customer experience](https://www.retaildive.com/news/ebay-ceo-ai-change-customer-experience/756794)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- Uberti-Bona Marin and colleagues (2026), ["If I Had to Buy Just ONE: Galaxy S26 Ultra": Auditing AI-Generated Product Recommendations](https://arxiv.org/abs/2609.18729)
- Underneath (2026), [When does Google show an AI Overview?](https://underneath.agency/research/ai-overviews-frequency-study)
- Underneath (2026), [“Is this brand legit?” How AI assistants build a reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study)

---

This is the Markdown twin of https://underneath.agency/resources/marketplace-buyers-sellers-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How mattress brands win customers through AI search"
description: "By being on the short list when shoppers ask AI to compare mattresses, with verifiable facts, real reviews, clear trial terms and no unsupported sleep claims."
canonical: "https://underneath.agency/resources/mattress-brands-customers-from-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How can a mattress brand win customers who ask AI which mattress to buy?

By being one of the few mattresses an AI assistant names when a shopper describes how they sleep and what they can spend, and by giving that assistant facts it can check: materials, firmness, trial and return terms, and reviews from real owners. Mattress shoppers face too many look-alike choices, so an answer that narrows the field carries weight. That weight also makes accuracy and honest claims more important than in most categories.

## The short version

1. AI is already part of mattress research. In a survey of 1,006 US mattress shoppers run in August 2026 by the mattress maker Amerisleep, 37% used AI tools in their search and 49% started on Google.
2. Shoppers are overwhelmed: in the same survey, 33% had delayed or abandoned a mattress purchase because the options were too confusing, and price drove the decision for 73%.
3. The market is shrinking, so every sale is contested. The International Sleep Products Association reports US mattress wholesale value fell to $9.25 billion in 2025, with unit shipments down 13.2%.
4. Mattress shoppers research and read reviews. In Better Sleep Council research, consumers who researched a mattress were most interested in comparing prices (64%) and consumer reviews (58%).
5. Ratings move AI picks. In a test of AI recommendations for skincare, another product people cannot judge before using, a well-known brand’s lead disappeared when a rival had less than a 0.1-star rating advantage.

## Who buys mattresses, and what is a customer worth?

Adults replacing a mattress years late, spending around a thousand dollars, and adding bedding afterward.

The [Amerisleep survey](https://amerisleep.com/blog/state-of-mattress-buying/) is a mattress brand’s own research, so read it as one source, but it maps the journey in detail. Respondents waited an average of 5 years after realizing their mattress needed replacing. They set an average spending ceiling of $982 for the mattress itself, from $606 for Gen Z to $1,109 for Gen X. After buying, 63% kept spending on toppers, pillows and bedding, about $74 on average. Memory foam led at 37% of purchases or planned purchases, followed by hybrids at 29%.

The industry is under pressure. The [International Sleep Products Association](https://sleepproducts.org/2026/07/mattress-industry-trends-report-2026-overview/) (ISPA) found total US mattress market value declined 6.5% in 2025 while unit shipments fell 13.2%; [BedTimes](https://bedtimesmagazine.com/2026/05/the-mattress-industry-trends-report-sharpens-the-industry-picture-for-2025/), ISPA’s magazine, puts the 2025 total at $9.25 billion in wholesale value, down from $9.9 billion. The largest company in the category feels it too. [Somnigroup](https://www.finanznachrichten.de/nachrichten-2026-08/69236400-somnigroup-international-inc-reports-second-quarter-2026-results-008.htm), owner of Tempur-Pedic, Sealy, Stearns & Foster and Mattress Firm, reported second quarter 2026 net sales of $1,823.5 million, down 3.0%, with Mattress Firm at $922.2 million, down 2.8%, mostly from store closures. In a market this flat, growth comes from taking customers from competitors.

A sale is only worth something if the mattress stays. Online brands built the category on long trials: Casper promised a 100-night “risk-free” trial, and [NPR reported](https://www.wlrn.org/2020-01-17/the-cost-of-free-casper-pays-a-price-for-generous-mattress-returns) in 2020 that its “refunds, returns and discounts” were worth almost $81 million in 2018. A shopper who picks the wrong mattress costs the brand twice.

## Where does AI already sit in the mattress buying journey?

At the research step, where shoppers compare types, brands and prices before choosing where to buy.

Mattress shoppers do their homework. In consumer research for the Better Sleep Council, ISPA’s consumer education arm, [most consumers said they research](https://bedtimesmagazine.com/2023/03/mattress-shoppers-most-interested-in-price-comparisons/) before buying a mattress, a larger share than in 2020. They were most interested in comparing prices (64%), consumer reviews (58%) and promotions (55%), and 48% wanted to compare product features. An earlier Better Sleep Council survey found shoppers used [three to four information sources](https://bedtimesmagazine.com/?p=116308) on average, and almost 75% shopped both online and in store.

AI tools now sit among those sources. In the Amerisleep survey, AI use was fairly even across generations, from 32% of baby boomers to 39% of Gen Z. Across all US retail, AI visits are worth more than they were. Adobe data reported by [TechCrunch](https://techcrunch.com/2026/04/16/ai-traffic-to-us-retailers-rose-393-in-q1-and-its-boosting-their-revenue-too/) shows AI traffic to retailers converted 42% better than other traffic in March 2026, after converting worse a year earlier. Adobe does not break out mattresses.

The appeal is obvious. A shopper facing dozens of near-identical boxed mattresses wants someone to narrow the field. A third of shoppers in the Amerisleep survey gave up or delayed because the options were too confusing. An AI answer that names three options for a side sleeper on a budget solves that problem, for better or worse.

## Which questions lead shoppers to a mattress brand?

Questions about how they sleep, what they can spend and which brands to trust.

The examples below are ours, written to show the shape of mattress questions; they are illustrative, not logged queries:

- Sleeper fit: “Best mattress for a side sleeper who runs hot, queen, under $1,000.”
- Type: “Hybrid or memory foam for a couple where one person moves a lot?”
- Comparison: “Saatva vs Helix: which is better value?”
- Trust: “Is this online mattress brand legit?” or “What do owners complain about?”
- Terms: “Mattresses with the longest trial and free returns.”

Some questions touch health, such as back pain. Assistants may answer them, but a mattress brand should not try to win them with medical claims. Amerisleep itself argues that an AI model cannot see the variables that decide whether a mattress works, such as body weight, sleep position and temperature sensitivity. A reasonable expectation is that answers to health-tinged questions lean on medical and review publishers more than on brands; see [what commercial health sites cited by ChatGPT have in common](https://underneath.agency/resources/chatgpt-cited-commercial-health-sites).

## How does an AI answer turn into a mattress sale?

Through a short list that sends the shopper to a brand site or store, ready to buy.

**The short list.** An answer names a few mattresses. Given that a third of shoppers stall on choice, being named may be the difference between a purchase and another year on the old mattress.

**The purchase.** The shopper visits your site, a retailer or a showroom. Many mattress shoppers move between channels, so some AI-influenced sales will close in a store where no tracking links them to the answer.

**The keep.** If the answer matched the shopper’s needs, the mattress is more likely to stay. In the Amerisleep survey, 58% of buyers had at least one regret, and 18% said they chose the wrong firmness or comfort. We infer that clear, accurate firmness and fit information, both on your site and in what AI tools say about you, protects revenue after the sale as well as winning it.

**The add-ons.** Toppers, pillows and bedding follow a kept mattress. Brands that also sell bed frames or bedroom sets can borrow from [how furniture brands win high-ticket AI orders](https://underneath.agency/resources/furniture-brands-sales-from-ai-search).

## What decides whether an AI assistant names your mattress?

Product data, reviews and reputation; the platform documents some of it, studies show the rest.

**Documented by the platform.** OpenAI says that [when ChatGPT selects products](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search), it considers structured metadata such as price and product description from first-party and third-party providers, plus other third-party content, and may weigh price, reviews and ease of use. It also says labels such as “Budget-friendly” are generated by ChatGPT and are not guarantees, and that reviews and ratings are not verified by OpenAI.

**Observed in studies.** An [academic study of brand bias in AI recommendations](https://arxiv.org/abs/2606.17443) tested skincare, where buyers rely on reputation because they cannot judge quality before use, much as with a mattress. When products had the same specifications, well-known brands were recommended 100% of the time, but that dominance disappeared with less than a 0.1-star rating advantage for a competitor. The same study found that authority-style language, including fabricated clinical-evidence claims, broke the leader’s hold 50% to 73% of the time. That is a finding about how easily models can be swayed, not a tactic. Under the [FTC’s health products guidance](https://www.ftc.gov/business-guidance/resources/health-products-compliance-guidance), claims about health benefits require competent and reliable scientific evidence. Our guide to [skincare products in AI answers](https://underneath.agency/resources/skincare-brands-ai-search) covers the same study and the same FTC rule from the skincare side.

Reputation questions lean on review sites. In [our study of “is this brand legit?” answers](https://underneath.agency/research/is-it-legit-ai-reputation-study), 88.0% of answers cited a review or complaint platform, and claims attached only to review platforms were negative 56.5% of the time. For online mattress brands, which shoppers often ask about by name, that makes complaint handling and owner reviews part of AI visibility.

**Trust factors specific to mattresses.** We infer that what helps an assistant name a mattress is what a careful shopper checks: a firmness rating on a stated scale, materials, height, trial length, return fees, warranty terms, delivery and old-mattress removal, and independent testing by reviewers who sleep on the product. For how ratings and prices compare with brand names in AI product picks, see [what drives AI product recommendations](https://underneath.agency/resources/what-drives-ai-product-recommendations).

## What does it cost a mattress brand to be missing, or misrepresented?

Shoppers who never consider you, and shoppers who buy the wrong mattress because of a wrong fact.

We found no study that counts mattress sales lost to AI answers. The exposure is clear, though. With 37% of shoppers using AI tools in their research and a third stalling on choice, a brand absent from those answers is absent from a growing share of short lists. Direct-to-consumer brands built their reach on paid marketing; NPR reported that Casper’s founders spent hundreds of millions of dollars on it. Being named in a free answer is one way to depend less on that spend, though we cannot yet measure how much less.

Misrepresentation has its own cost. An assistant that states an outdated trial length, the wrong firmness or a complaint pattern drawn from old reviews can send away a buyer who would have kept the mattress, or attract one who will return it.

## How does GEO work for a mattress brand?

By making your mattresses easy to understand, compare and verify, without promising placement or making health claims.

1. **Verifiable product facts.** Firmness on a stated scale, materials and layers, height, weight limits, edge support, sizes, trial length, return fees, warranty and delivery terms, consistent across your site, product feeds and every retailer that sells you.
2. **Honest fit guidance.** Pages that explain which sleepers each model suits and which it does not. Admitting a trade-off is more believable than claiming to suit everyone.
3. **Substantiated claims only.** No claims about pain, health or sleep quality without competent and reliable scientific evidence. Unsupported claims are a legal risk and can surface in AI answers next to the evidence against them.
4. **Real reviews, gathered fairly.** The [FTC’s rule on fake reviews](https://www.ftc.gov/news-events/news/press-releases/2024/08/federal-trade-commission-announces-final-rule-banning-fake-reviews-testimonials) bans fake reviews and incentives conditioned on a particular sentiment. Planted reviews can also distort AI answers; see [how fake reviews affect AI recommendations](https://underneath.agency/resources/fake-reviews-fake-brands-ai-recommendations).
5. **Independent testing coverage.** Reviewers and publications that test mattresses in person and explain their methods.
6. **Reputation monitoring.** Track what assistants say when shoppers ask whether your brand is legit, and fix the underlying problems they cite.

For where useful GEO ends and manipulation begins, see [legitimate GEO versus manipulation](https://underneath.agency/resources/legitimate-geo-vs-manipulation).

## What can’t the data yet say about selling mattresses through AI?

Which mattress brands AI assistants name today, how stable those picks are, and how many sales they drive.

The best mattress-specific data on AI use comes from a survey run by a mattress company. The Better Sleep Council figures predate the spread of AI assistants. Adobe’s conversion data covers all retail. The brand-bias study tested skincare in a controlled setting, not mattresses on live assistants. We found no independent, published measurement of which mattresses assistants recommend for common sleeper questions, and nothing that separates AI-influenced sales from returns. Any promise that a brand will become “the AI’s pick” for mattresses goes beyond the evidence.

## Where should a mattress brand start?

With an audit of what AI assistants say about your mattresses, your trial terms and your reputation.

List the questions your buyers ask: sleeper type, budget, mattress type, brand comparisons and “is it legit.” Check which brands are named, which review sites are cited, and whether the facts about your firmness, trial and returns are right. Then fix the facts, the review gaps and the claims you cannot support. If you would rather have us run it, [ask us for a mattress visibility review](https://underneath.agency/contact) and we will show where your brand stands in AI answers for the questions that lead to mattress orders, and what would help more of those orders stay sold. For the ongoing work, our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page covers keeping firmness, trial and review facts consistent wherever sleepers and assistants look.

## Frequently asked questions

### Can an AI assistant tell a shopper which mattress is best for back pain?

It may try, but a mattress brand should not rely on it or encourage it. Health claims need scientific evidence, and AI tools cannot assess a shopper’s body or condition. Shoppers with pain should ask a clinician.

### Do mattress shoppers still start on Google?

Most do. In the Amerisleep survey, 49% started on Google, while 37% used AI tools at some point in their research.

### Can we offer discounts in exchange for reviews?

Not discounts tied to a positive review. The FTC’s rule bans incentives conditioned on reviews expressing a particular sentiment, positive or negative.

### Do longer sleep trials help with AI visibility?

We do not know of evidence that they do directly. Clear trial and return terms help shoppers and assistants compare offers, and fewer mismatched purchases mean fewer returns.

## Sources

- Amerisleep (2026-09-02), [State of Mattress Buying 2026: How U.S. Shoppers Research, Choose and Regret Mattresses](https://amerisleep.com/blog/state-of-mattress-buying/)
- International Sleep Products Association (2026-07), [ISPA Report Shows Mattress Manufacturers Maintaining Value Despite Shipment Declines](https://sleepproducts.org/2026/07/mattress-industry-trends-report-2026-overview/)
- BedTimes (2026-05), [The Mattress Industry Trends Report sharpens the industry picture for 2025](https://bedtimesmagazine.com/2026/05/the-mattress-industry-trends-report-sharpens-the-industry-picture-for-2025/)
- Somnigroup (2026-08-06), [Somnigroup International Inc. Reports Second Quarter 2026 Results](https://www.finanznachrichten.de/nachrichten-2026-08/69236400-somnigroup-international-inc-reports-second-quarter-2026-results-008.htm)
- BedTimes (2023-03-17), [Mattress shoppers most interested in price comparisons](https://bedtimesmagazine.com/2023/03/mattress-shoppers-most-interested-in-price-comparisons/)
- BedTimes (2021-01-12), [Mattress buyers are less reliant on such reviews than other shoppers but still value them](https://bedtimesmagazine.com/?p=116308)
- NPR (2020-01-17), [The Cost Of Free: Casper Pays A Price For Generous Mattress Returns](https://www.wlrn.org/2020-01-17/the-cost-of-free-casper-pays-a-price-for-generous-mattress-returns)
- TechCrunch (2026-04-16), [AI traffic to US retailers rose 393% in Q1](https://techcrunch.com/2026/04/16/ai-traffic-to-us-retailers-rose-393-in-q1-and-its-boosting-their-revenue-too/)
- OpenAI (2026), [Shopping with ChatGPT Search](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search)
- Federal Trade Commission (2022-12), [Health Products Compliance Guidance](https://www.ftc.gov/business-guidance/resources/health-products-compliance-guidance)
- Federal Trade Commission (2024-08-14), [Federal Trade Commission Announces Final Rule Banning Fake Reviews and Testimonials](https://www.ftc.gov/news-events/news/press-releases/2024/08/federal-trade-commission-announces-final-rule-banning-fake-reviews-testimonials)
- Chu et al. (2026), [Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems](https://arxiv.org/abs/2606.17443)
- Underneath (2026), [“Is it legit?” AI reputation study](https://underneath.agency/research/is-it-legit-ai-reputation-study)

---

This is the Markdown twin of https://underneath.agency/resources/mattress-brands-customers-from-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How MDR and managed SOC providers win leads from AI search"
description: "By being named, with defined response metrics and independent proof, when stretched IT and security leaders ask AI which MDR or managed SOC fits them."
canonical: "https://underneath.agency/resources/mdr-providers-qualified-leads-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How can MDR and managed SOC providers win qualified leads from AI search?

By being named, with defined response metrics and independent proof, when stretched IT and security leaders ask an AI assistant which managed detection and response (MDR) or managed SOC service fits their size, tools and insurer. A managed service is bought on trust in people you cannot see, so the buyer checks hard before calling. The AI answer can do part of that qualifying before the first call, though no study has yet measured how often MDR buyers use it.

## The short version

1. Many organizations cannot staff security operations themselves: in [ISC2’s 2025 study](https://www.isc2.org/Insights/2025/12/2025-ISC2-Cybersecurity-Workforce-Study) of 16,029 professionals, 33% said they lack the budget to staff their teams adequately, and 19% bring in third-party service providers to fill skills gaps.
2. Contract sizes span a wide range: [Vendr’s purchase data](https://www.vendr.com/marketplace/huntress) put the median Huntress buyer at $11,000 a year, [Red Canary](https://www.vendr.com/marketplace/red-canary) at $79,881 and [eSentire](https://www.vendr.com/marketplace/esentire) at $120,813. Qualification by size matters.
3. Incidents drive demand: in a 2026 survey of 1,350 security and IT decision-makers for [Arctic Wolf](https://arcticwolf.com/resources/press-releases/arctic-wolf-2026-trends-report-reveals-ai-trust-gap-as-organizations-race-to-modernize-security-operations-for-the-age-of-ai/), 63% reported a significant incident in the past year, and 48% of those affected lost productivity for two weeks or longer.
4. Insurers are in the loop: [Marsh](https://www.marsh.com/en-gb/about/media/incident-response-planning-emerges-key-cybersecurity-control-reducing-cyber-risk.html) found each 25% increase in endpoint detection and response coverage went with a 10% lower likelihood of a breach.
5. AI assistants research this category through analyst reports. In our hidden-searches study, ChatGPT ran the search “Gartner Magic Quadrant 2024 managed detection and response” while answering a buyer question.

## Who buys MDR and managed SOC services, and what is a contract worth?

Mostly mid-market IT leaders without a full SOC, enterprise CISOs extending their own, and MSPs serving small clients.

The terms overlap, so buyers often ask an assistant to explain them first. MDR is a service in which a provider’s analysts watch an organization’s security tools around the clock and act on threats, often by isolating a device or disabling an account. A managed SOC, or SOC as a service, usually runs a fuller security operations function, often around the customer’s own SIEM, the system that collects security logs. A managed security service provider (MSSP) may manage devices and alerts with less hands-on response.

Three buyer types stand out:

- **Mid-market organizations** with a small IT team and no one watching alerts at night.
- **Enterprises** with a SOC that need round-the-clock coverage, specialist threat hunting or relief for overloaded analysts. SIEM vendors court the same SOC leaders; see [how SIEM vendors reach enterprise shortlists](https://underneath.agency/resources/siem-enterprise-pipeline-ai-search).
- **Managed service providers (MSPs)** that resell or bundle MDR for many small clients. They face their own AI shortlist, covered in [how MSPs win clients from AI search](https://underneath.agency/resources/msps-customers-from-ai-search).

The staffing gap is the root of demand. ISC2’s 2025 study found 29% of organizations cannot afford to hire staff with the skills they need. To fill gaps, 20% outsource work and 19% bring in third-party providers. 72% agreed that reducing cybersecurity personnel significantly increases breach risk.

Contract values follow the buyer type. Vendr’s medians, from purchases on one procurement platform:

| Provider | Median annual spend | Range shown by Vendr |
|---|---|---|
| Huntress | $11,000 | $6,768 to $39,372 |
| Red Canary | $79,881 | $26,980 to $154,687 |
| eSentire | $120,813 | $37,539 to $232,789 |

The wider services market is large. Gartner forecasts worldwide spending on security services, a category that includes consulting and other services as well as managed ones, at $92,780 million in 2026, up from $77,130 million in 2024, as reported by [CRN Asia](https://www.crnasia.com/india/news-network/news/gartner-forecasts-worldwide-end-user-spending-on-information-security-to-total-213-bn-in-2025).

## Why do companies outsource security operations now?

Because incidents keep happening, detection is slow without round-the-clock eyes, and insurers reward strong controls.

**Incidents.** Arctic Wolf’s survey found 63% of organizations had a significant incident in the past year, while 96% of leaders said they were confident their teams could keep pace. Arctic Wolf sells MDR, so read its survey as one input. Its [2026 threat report](https://arcticwolf.com/resources/press-releases/arctic-wolf-threat-report-highlights-11x-growth-in-data-extortion-incidents-and-continued-dominance-of-ransomware/) found ransomware, business email compromise and data incidents made up 92% of its incident response cases.

**Slow detection is expensive.** A breach took 241 days on average to identify and contain, according to [IBM’s 2025 breach report](https://newsroom.ibm.com/2025-07-30-ibm-report-13-of-organizations-reported-breaches-of-ai-models-or-applications,-97-of-which-reported-lacking-proper-ai-access-controls). Where an organization’s own team caught the intrusion, it saved $900,000 compared with learning of it from the attacker, the gap round-the-clock MDR monitoring is pitched against.

**Insurance.** Marsh’s analysis of claims ranked endpoint detection and response, and logging and monitoring, among the controls most linked to fewer breach-based claims. [Coalition](https://www.coalitioninc.com/announcements/2025-cyber-claims-report), an insurer that also sells security services, traced 60% of its 2024 claims to business email compromise and funds transfer fraud. MDR providers court this channel: [Arctic Wolf](https://arcticwolf.com/company/) runs an insurance partner program for brokers and carriers and says its service can “increase the likelihood of insurability.”

## Where does AI search enter the MDR buying journey?

At the research step, where a buyer narrows dozens of providers to a few, then checks them with people.

Gartner’s 2026 survey of 645 B2B buyers, across industries, describes the pattern. 45% used generative AI in a recent purchase, mainly to gather information on vendors and products. [69% prefer to validate AI-generated insights with sales reps](https://www.gartner.com/en/newsroom/press-releases/2026-05-20-gartner-survey-finds-sixty-nine-percent-of-b-two-b-buyers-turn-to-sales-reps-to-validate-ai-generated-insights). For a service bought on trust, we infer that validation happens in scoping calls, reference checks and trials.

There is direct evidence that assistants research this category through analyst reports. In [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study), one of ChatGPT’s own searches was “Gartner Magic Quadrant 2024 managed detection and response.” ChatGPT ran a search aimed at a named publication, ranking or award in 43.8% of its answers. Analyst market guides, peer reviews and independent rankings are part of what assistants look for, we observe.

## Which questions do MDR and managed SOC buyers ask AI assistants?

Questions about fit, service model, response speed, tools, location and insurance. We drafted the prompts below to mirror how an IT director or CISO might ask; they are examples, not recorded buyer queries.

| Buyer | Illustrative prompt |
|---|---|
| Mid-market, first purchase | “What is the best MDR service for a 400-person manufacturer with two IT staff?” |
| Service model | “MDR, MSSP or SOC as a service: what is the difference and which do I need?” |
| Existing tools | “Which MDR providers work with Microsoft Defender and CrowdStrike, rather than their own agent?” |
| Speed | “Which MDR providers publish their mean time to respond, and how do they define it?” |
| Location | “MDR providers in Germany with an EU-based SOC and German-speaking analysts?” |
| Insurance | “Will my cyber insurer give a better rate if we use MDR?” |
| Enterprise | “Managed SOC to run our Splunk SIEM overnight and on weekends?” |
| MSP | “Best MDR to resell to 60 small business clients?” |

Location questions deserve special care. In [our country study](https://underneath.agency/research/ai-recommendations-by-country-study), naming the United Kingdom in the question raised the share of local-market brands in ChatGPT’s answers from 25.4% to 49.1%. In [our four-assistant study](https://underneath.agency/research/ai-assistants-brand-agreement-study), questions naming a place had far less agreement between assistants, an overlap of 0.160 against 0.390 for national questions. Regional providers may be named for local questions and missing from national ones, we infer.

## How does AI visibility become a qualified MDR lead?

When the AI answer has matched the provider to the buyer’s size, tools and region, the first call starts qualified.

A qualified MDR lead, in practice, has four facts in place: a size that fits the provider, compatible security tools, a budget in the provider’s range and a reason to buy now. The path, as we understand it:

1. **Trigger.** An incident, an insurance renewal, an analyst resignation or a board question.
2. **AI-assisted short list.** The buyer describes their situation, and the assistant names a few providers with reasons.
3. **Validation.** The buyer reads the providers’ sites, reviews and analyst mentions, then books scoping calls.
4. **Scoping and proposal.** Endpoints, identities, cloud accounts and log sources set the price.
5. **Contract and expansion.** An annual or multi-year agreement, often extended to incident response retainers or more coverage.

A specific question produces a pre-qualified buyer, we expect. A buyer who asked for a provider that supports their existing tools in their region has already filtered for fit. A vague or wrong description in the answer does the opposite, sending poorly matched buyers or none. The click itself is rarely visible in analytics; see [what lost clicks mean for pipeline](https://underneath.agency/resources/ai-answers-pipeline-revenue).

## Why does an assistant recommend one MDR service over another?

No platform documents how it picks providers; studies point to analyst reports and reviews, and buyers want verifiable proof.

**Documented by the platforms.** A question about overnight monitoring can set off several searches at once: Google says AI Overviews and AI Mode [may use a “query fan-out” technique](https://developers.google.com/search/docs/appearance/ai-features). None of the platforms explains how a managed security provider ends up named.

**Observed in our studies.** Besides the analyst-report searches above, [our reputation study](https://underneath.agency/research/is-it-legit-ai-reputation-study) found 88.0% of answers to “Is this brand legit?” cited a review or complaint platform. Ask ChatGPT the same question five times and the names move: in [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), just 25.2% of the brands it gave appeared in every run.

**What service buyers check.** Providers already compete on public service claims. [eSentire](https://www.esentire.com/what-we-do/managed-detection-and-response) advertises a “15-minute Mean Time to Contain” and says it protects more than 2,000 organizations. Arctic Wolf says its platform serves over 10,000 organizations. [Huntress](https://www.huntress.com/about) stresses its round-the-clock SOC. Such claims are measured in different ways, so buyers, and assistants, cannot compare them easily.

**Our inference.** A reasonable expectation is that providers whose metrics are defined, whose supported tools are listed and whose reviews and analyst mentions are current give assistants more to repeat accurately. No one has yet tested that idea on managed detection providers.

## What happens to an MDR provider that assistants leave out?

Mostly it loses first calls in a consolidating market, though no study has sized that loss for managed security.

- **Consolidation reshapes the list.** Red Canary’s product page now says [“Red Canary is now part of Zscaler.”](https://redcanary.com/products/managed-detection-and-response/) As product companies absorb services, buyers ask whether to choose an independent provider or one tied to a tool, we infer. Independents must be visible in that comparison.
- **Contracts recur.** With medians from $11,000 to $120,813 a year on Vendr’s data, plus renewals and expansion, each missed buyer is years of revenue.
- **Presence is partial.** A provider named in some runs and not others loses some buyers without knowing.
- **Wrong facts mislead.** An assistant that says a provider requires its own agent, or lacks a regional SOC, filters out buyers who would have fit. To correct an agent or SOC-location error, start with [how to fix wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers).

## How does GEO work for an MDR or managed SOC provider?

It puts your coverage hours, response metrics and supported tools where assistants and buyers can check them. Nobody can promise a recommendation.

1. **Define the service precisely.** Say whether you offer MDR, a managed SOC, MSSP services or several, and what each includes: hours, response actions, threat hunting, incident response. Buyers who need testing rather than monitoring ask differently, as [how penetration testing firms win scoped leads](https://underneath.agency/resources/pentest-firms-leads-ai-search) shows.
2. **Publish metrics with definitions.** If you cite response or containment times, define when the clock starts and stops, and how often you meet it.
3. **List supported tools and scope.** Endpoint, identity, cloud and SIEM integrations, plus how you price (endpoints, users, data).
4. **Show who stands behind the service.** Analyst team size, locations, certifications such as SOC 2, and data residency for each region.
5. **Earn independent proof.** Analyst market guides, peer review platforms and published threat research that the security press cites. See [how brands build authority for AI search](https://underneath.agency/resources/how-brands-build-authority-for-ai-search) and [how small brands get recommended by AI](https://underneath.agency/resources/how-small-brands-get-recommended-by-ai).
6. **Serve each buyer and region.** Separate pages for mid-market, enterprise and MSP buyers, and for each country you serve, in its language; see [GEO for global brands across languages](https://underneath.agency/resources/geo-strategy-for-global-brands-across-languages).
7. **Measure with real constraints.** Test size, tool, region and insurance wording across ChatGPT, Google AI Overviews and AI Mode, Gemini, Perplexity, Copilot and Claude, several runs each.

## Which questions about AI and managed detection buying remain open?

Two big ones: how many MDR buyers use assistants, and whether being named leads to signed contracts. No published study answers either.

- **No survey of managed security buyers.** Gartner’s AI-use figures describe B2B purchases across industries, not MDR selections.
- **Providers fund much of the data.** Arctic Wolf’s surveys, eSentire’s metrics and Coalition’s claims data come from companies that sell services.
- **Vendr is a sample.** Its medians reflect purchases on one platform.
- **Lead quality is unmeasured.** In [Martinez’s](https://arxiv.org/abs/2607.14035) review of 45 studies of AI search optimization, traffic and conversions had the weakest evidence, so nobody can yet say AI visibility yields qualified MDR leads.

## How can an MDR provider test whether assistants send it qualified buyers?

Ask assistants what your best-fit buyers would ask, including their size, tools and region, and note who is named.

Write 20 to 30 prompts for your three buyer types, each with tool, region, insurance and response-time wording. Put every prompt to each major assistant more than once, since a managed security shortlist can change between runs. Record who is named, which analyst reports and review sites are cited, and whether your service model, supported tools and regions are described correctly. Providers competing for a broader security budget can compare notes with [how cybersecurity software firms earn revenue from AI search](https://underneath.agency/resources/cybersecurity-software-revenue-from-ai-search). Providers that also sell staff phishing training can use [security awareness training in AI answers](https://underneath.agency/resources/security-awareness-training-customers-ai-search).

For an outside read on which of those questions you win, [talk to us about an MDR lead audit](https://underneath.agency/contact). We will show which buyer questions name you, which send qualified buyers to competitors, and which gaps in your public service proof are most likely costing you scoping calls and contracts. Our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page describes the work that follows, such as defining response metrics, listing supported tools and testing buyer questions over time.

## Frequently asked questions

### Do mid-market companies use AI to choose an MDR provider?

No study isolates them. Across B2B purchases, 45% of buyers in Gartner’s 2026 survey used generative AI, mainly to research vendors and products.

### Should we publish our response times?

Yes, with definitions. Providers already advertise figures such as a 15-minute mean time to contain, and undefined numbers are hard for buyers or assistants to compare.

### Does analyst recognition matter for AI visibility?

It appears to. In our hidden-searches study, 43.8% of ChatGPT’s answers involved a search for a named ranking, and one such search targeted the Gartner Magic Quadrant for MDR.

### Should regional MDR providers target local questions?

Yes. Naming the United Kingdom raised local-market brands in ChatGPT’s answers from 25.4% to 49.1% in our country study.

### Is MDR the same as a managed SOC?

Not quite. MDR focuses on detecting and acting on threats; a managed SOC usually runs a fuller operations function, often around your own SIEM.

## Sources

- ISC2 (2025-12), [2025 ISC2 Cybersecurity Workforce Study](https://www.isc2.org/Insights/2025/12/2025-ISC2-Cybersecurity-Workforce-Study)
- Vendr (2026), [Huntress software pricing and plans](https://www.vendr.com/marketplace/huntress)
- Vendr (2026), [Red Canary software pricing and plans](https://www.vendr.com/marketplace/red-canary)
- Vendr (2026), [eSentire software pricing and plans](https://www.vendr.com/marketplace/esentire)
- Arctic Wolf (2026-07-28), [Arctic Wolf 2026 Trends Report reveals AI trust gap](https://arcticwolf.com/resources/press-releases/arctic-wolf-2026-trends-report-reveals-ai-trust-gap-as-organizations-race-to-modernize-security-operations-for-the-age-of-ai/)
- Arctic Wolf (2026-02-17), [Arctic Wolf Threat Report highlights 11x growth in data extortion incidents](https://arcticwolf.com/resources/press-releases/arctic-wolf-threat-report-highlights-11x-growth-in-data-extortion-incidents-and-continued-dominance-of-ransomware/)
- Arctic Wolf (2026), [Company](https://arcticwolf.com/company/)
- Marsh (2025-08-27), [Incident response planning emerges as key cybersecurity control in reducing cyber risk](https://www.marsh.com/en-gb/about/media/incident-response-planning-emerges-key-cybersecurity-control-reducing-cyber-risk.html)
- Coalition (2025-05-07), [Coalition 2025 Cyber Claims Report](https://www.coalitioninc.com/announcements/2025-cyber-claims-report)
- IBM (2025-07-30), [IBM report: 13% of organizations reported breaches of AI models or applications](https://newsroom.ibm.com/2025-07-30-ibm-report-13-of-organizations-reported-breaches-of-ai-models-or-applications,-97-of-which-reported-lacking-proper-ai-access-controls)
- Gartner, via CRN Asia (2025-07-29), [Gartner forecasts worldwide end-user spending on information security to total $213 bn in 2025](https://www.crnasia.com/india/news-network/news/gartner-forecasts-worldwide-end-user-spending-on-information-security-to-total-213-bn-in-2025)
- Gartner (2026-05-20), [Gartner survey finds 69% of B2B buyers turn to sales reps to validate AI-generated insights](https://www.gartner.com/en/newsroom/press-releases/2026-05-20-gartner-survey-finds-sixty-nine-percent-of-b-two-b-buyers-turn-to-sales-reps-to-validate-ai-generated-insights)
- eSentire (2026), [Managed detection and response](https://www.esentire.com/what-we-do/managed-detection-and-response)
- Zscaler (Red Canary) (2026), [Zscaler Managed Detection & Response](https://redcanary.com/products/managed-detection-and-response/)
- Huntress (2026), [About Huntress](https://www.huntress.com/about)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- Martinez (2026), [Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026)](https://arxiv.org/abs/2607.14035), arXiv:2607.14035.
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [Same question, four countries: do AI recommendations change?](https://underneath.agency/research/ai-recommendations-by-country-study)
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude agree on brands?](https://underneath.agency/research/ai-assistants-brand-agreement-study)
- Underneath (2026), [“Is this brand legit?” How AI assistants build a reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study)
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)

---

This is the Markdown twin of https://underneath.agency/resources/mdr-providers-qualified-leads-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How should we measure share of citations in AI search tools?"
description: "Count every answer, including those that never searched the web, report each engine separately, and use enough repeated answers to tell real gaps from noise."
canonical: "https://underneath.agency/resources/measuring-share-of-citations-in-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How should we measure share of citations in AI search tools?

Measure it over all the answers an engine gives, not just the ones that cite something, and report each engine separately with a range around the figure. Many AI answers never search the web and cite nothing, so a share calculated only among cited answers overstates visibility. And because answers change from run to run, a share from a small sample can rank two sources the wrong way round.

## The short version

1. In one Swiss study, 57.8% of ChatGPT’s runs did not search the web and so cited nothing; a share counted only among the rest would overstate visibility.
2. In our 80-question study, none of the 42 answers that ran no web search cited anything.
3. The questions you choose and how you weight them change the result: reweighting the same published data gave rates of 39.7% or 70.5%.
4. Share estimates from 200 questions on OpenAI’s search still carried ranges of 3 to 6 percentage points, enough to make many apparent leads meaningless.
5. Engines cite very different numbers of sources, from about 40 to 43 per answer on Gemini to 6 or 7 on OpenAI’s search, so raw counts cannot simply be added up.

## Why does the denominator matter so much?

Because many AI answers cite no sources at all, and leaving them out inflates every share you calculate.

[Schulte and colleagues](https://arxiv.org/abs/2604.07585) ran German-language prompts in four consumer categories on four engines, from Swiss servers, in early 2026. ChatGPT searched the web only for some questions, leaving 57.8% of its runs with no citations. [Martinez’s survey](https://arxiv.org/abs/2607.14035) draws the lesson: a dashboard “cannot calculate a ‘share of citations’ only among responses that contain citations and then interpret it as overall visibility.” The 57.8% comes from one preprint with a small Swiss question set, so treat it as an example, not a norm.

Our own data shows the same link. In [our hidden searches study](https://underneath.agency/research/ai-hidden-searches-study), every answer that ran a search cited at least one page, and none of the 42 answers without a search cited anything. Claude averaged only 0.76 searches per answer, against 3.7 for ChatGPT.

The fix is simple arithmetic. Your share of all answers equals the share of answers that search, times your share of citations within those. Report both parts. A high share among cited answers can coexist with low visibility overall.

## Which questions should be in the denominator?

The ones your buyers actually ask, weighted by how often they ask them, and stated openly in the report.

A second paper by [Martinez](https://arxiv.org/abs/2609.06811) argues that a prompt list defines the market being measured. Stating that a source “is cited in 40% of answers is therefore insufficient” without saying which questions, which engine and which weights. To show the size of the effect, Martinez reweighted three published groups of Google queries. One weighting gave a rate of AI Overviews appearing of 39.7%, another 70.5%, with no change inside any group.

The same applies to citations. A tracker heavy on questions where you are strong will show a high share. Ask any vendor how its question list was built and whether it reflects real buyer demand.

## How many answers do we need before a share means anything?

More than most dashboards use, and the number differs by engine.

[Sielinski](https://arxiv.org/abs/2603.08924), from the measurement firm IQRush, ran 200 questions per topic for nine days on three engines. For most frequently cited sites on OpenAI’s search, the range around each share spanned 3 to 6 percentage points. In one example, one review site showed 9.5% and another 6.0%, yet the ranges overlapped so much that neither could be called the leader. The same problem is why [a one-off visibility report](https://underneath.agency/resources/one-time-ai-visibility-report) can mistake noise for a lead.

Reaching a range of about five points took roughly 40 to 50 questions on Gemini, about 100 on Perplexity and 150 or more on OpenAI’s search. A rise from 8% to 11% on OpenAI’s search, he notes, cannot be credited to your work with confidence.

| Engine (Sielinski, 2026) | Citations per answer | Questions needed for a range of about 5 points |
|---|---|---|
| Gemini | about 40 to 43 | about 40 to 50 |
| Perplexity | about 20 to 22 | about 100 |
| OpenAI search | about 6 to 7 | 150 or more |

A [follow-up by Sielinski](https://arxiv.org/abs/2607.10341) found no fixed budget works everywhere. Across 30 engine-and-topic combinations, no budget below 94 answers would have settled every ranking, and three never settled within 125. Brand counts behave the same way. In [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), pinning a frequency near 50% to within 10 points would take about 97 independent runs.

## Can we combine engines into one share of citations?

Only carefully: add raw counts and the engine that cites the most will drown out the others.

Engines cite very different numbers of sources. In Tannenbaum’s [four-engine test](https://arxiv.org/abs/2609.22655), ChatGPT cited 17.8 pages per prompt and Copilot 4.5. As he puts it, a page competing for a four-source answer faces different odds from one in an 18-source answer. Sielinski warns that pooling raw counts means “a dense platform overshadows sparse ones.”

The better method is to compute a share within each engine, with its own range, and then [combine engines using weights you state](https://underneath.agency/resources/combine-ai-engines-visibility-score), such as each engine’s estimated share of your buyers.

## Is a citation the same thing as visibility?

No: a citation does not show that anyone noticed it, or that it was accurate.

Citations can misrepresent their sources. In a 2023 audit summarized by Martinez, only 51.5% of sentences in four AI search engines’ answers were fully supported by their citations, and 74.5% of citations supported the claim they were attached to. [Kato and colleagues](https://arxiv.org/abs/2609.11915) add that how often a name appears is not yet exposure: it must be combined with how many people ask and the chance they notice it.

So report several measures side by side: whether the engine searched, whether you were cited, whether your brand was named, and what was said.

## What should you do about it?

Ask for a share that is complete, ranged and broken down by engine, before trusting or buying any figure.

1. Count every answer, including those with no search and no citations, and show the search rate alongside the share.
2. Report each engine separately, with its own question count and range.
3. Use [a question list built from real buyer questions](https://underneath.agency/resources/how-to-design-ai-visibility-tracking), and write down how it is weighted.
4. Run enough answers per engine: tens on some engines, well over a hundred on others.
5. Treat changes smaller than the range as noise, and fix the question set before comparing months.
6. Ask vendors these questions directly: what is the denominator, how many runs, which engine version, which country?

If you want help building a measurement that meets these tests, see [our generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

It does not yet give a standard sample size or tell us how citation share relates to sales.

- The 57.8% no-search figure comes from one small Swiss preprint; rates elsewhere are unknown.
- Sample-size guidance comes from one vendor’s data across three engines; its own author says thresholds may not transfer.
- Engines change their search behavior often, so today’s search rates may not hold.
- No study here links citation share to clicks, leads or revenue.
- How to weight engines by their real audience has not been studied.

## Frequently asked questions

### What is share of citations in AI search?

It is the fraction of the sources AI engines cite that point to your site. It means little unless you state the questions, the engine and whether answers with no citations were counted.

### Why do AI visibility tools show different numbers?

Because they use different question lists, engines, run counts and denominators. Reweighting the same published data alone moved one rate from 39.7% to 70.5%.

### How many prompts do we need to track AI citations reliably?

It depends on the engine. One study needed about 40 to 50 questions on Gemini but 150 or more on OpenAI’s search for a range of about five points.

### Should answers without citations count in our share?

Yes. In one study, 57.8% of ChatGPT runs searched nothing and cited nothing; dropping them overstates visibility.

## Sources

- Schulte, Bleeker and Kaufmann (2026), [Don’t Measure Once: Measuring Visibility in AI Search (GEO)](https://arxiv.org/abs/2604.07585), arXiv:2604.07585.
- Martinez (2026), [Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026)](https://arxiv.org/abs/2607.14035), arXiv:2607.14035.
- Martinez (2026), [Measuring GEO Visibility: Prompt Corpora Define the Answer Market](https://arxiv.org/abs/2609.06811), arXiv:2609.06811.
- Sielinski (2026), [Quantifying Uncertainty in AI Visibility: A Statistical Framework for Generative Search Measurement](https://arxiv.org/abs/2603.08924), arXiv:2603.08924.
- Sielinski (2026), [From Stochastic to Stable: Rank Stability and Structural Sufficiency in AI Visibility Measurement](https://arxiv.org/abs/2607.10341), arXiv:2607.10341.
- Tannenbaum (2026), [Scoring With the Engine: Retrieval Exposure, Cross-Engine Divergence, and the Limits of Engine-Agnostic GEO Scores](https://arxiv.org/abs/2609.22655), arXiv:2609.22655.
- Kato, Honma and Kato (2026), [Generative Marketing Mix Modeling: A Causal Inference Framework Linking GEO and GEM to Business Impact](https://arxiv.org/abs/2609.11915), arXiv:2609.11915.
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)

---

This is the Markdown twin of https://underneath.agency/resources/measuring-share-of-citations-in-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Can AI search build clinician demand for a medical device?"
description: "It can shape what clinicians and hospital committees read about your device, if peer-reviewed evidence and clear FDA facts are easy for AI tools to find."
canonical: "https://underneath.agency/resources/medical-device-demand-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Can AI search build demand for a medical device before the value analysis committee meets?

Yes, at the start of the process: AI tools now sit where surgeons and nurses first read about a technology, and the hospital’s value analysis committee then asks for the same evidence. What those tools can repeat depends on your peer-reviewed data, your clearance facts and how clearly you state your indications. No study yet measures how often AI answers lead to a device request, so treat what follows as a reasoned case, not a proven channel.

## The short version

1. The field is crowded: FDA’s device center authorized 124 novel devices in 2025, says its [2025 annual report](https://www.fda.gov/media/190779/download), and an [Emergo by UL count](https://www.emergobyul.com/news/us-fda-issues-2025-annual-report-medical-device-regulatory-activities) found 3,238 510(k) clearances that year.
2. Hospitals are under cost pressure: the [American Hospital Association](https://www.aha.org/press-releases/2026-03-11-new-aha-report-hospitals-face-increased-challenges-and-financial-pressures-they-care-patients) reports hospital spending on supplies rose 9.9% in 2025, faster than total expenses (7.5%).
3. Committees now gatekeep every new product: NCH Healthcare System in Florida told vendors that from August 3, 2026, no new products, trials or samples may enter its hospitals without [value analysis approval](https://nchmd.org/wp-content/uploads/2023/09/2026_7_31_NCH-Vendor-Letter.pdf), with clinical evidence first on its list of criteria.
4. Clinicians already use AI to read the evidence: OpenEvidence, an AI search tool for doctors, reported 760,000 registered US physicians and about 18 million consultations a month in December 2025, [according to Wikipedia’s summary of company figures](https://en.wikipedia.org/wiki/OpenEvidence).
5. One approval is worth years of revenue: Intuitive Surgical’s [2025 annual report](https://www.sec.gov/Archives/edgar/data/1035267/000103526726000010/isrg-20251231.htm) says a da Vinci system sells for $0.7 million to $3.1 million, and the company earns $900 to $3,700 in instruments and accessories per procedure.

## Who decides whether a hospital buys your device, and what is the account worth?

A clinician champion starts it, a value analysis committee vets it, and supply chain and senior management sign it.

The buying group for a medical device is wider than most marketers assume. A surgeon, interventional cardiologist or nurse leader usually asks for the product. A multidisciplinary value analysis committee then reviews it. NCH Healthcare System’s 2026 vendor letter lists what its committee weighs: clinical evidence and patient outcomes, patient and staff safety, regulatory compliance, operational impact, standardization, total cost of ownership, and contracting. Vendors may not hand products to physicians for evaluation without formal approval.

For capital equipment, more people join. Intuitive Surgical says the purchase of its systems “generally requires the approval of senior management of hospitals, their parent organizations, purchasing groups, and/or government bodies,” and that some sales go through competitive bidding or public tenders. It also notes that integrated delivery networks are “creating larger networks of system users with increasing purchasing power.”

What a won account is worth depends on the business model, and most device makers earn far more after the first sale than at it:

- **Capital plus recurring use.** A da Vinci system costs $0.7 million to $3.1 million. Service contracts run $95,000 to $225,000 a year, and each procedure brings $900 to $3,700 in instruments and accessories. Intuitive’s installed base reached about 11,106 systems at the end of 2025, up 12% in a year.
- **Consumables and implants.** For disposables and implants, committee approval opens the door to recurring purchases across every surgeon who adopts the product, which is why a single approval can matter for years.

Cost pressure raises the bar. With supply spending up 9.9% in 2025, committees have every reason to ask whether a new device is worth more than what is already on the shelf.

## Where do AI tools already appear in a clinician’s evaluation?

Mainly where clinicians read and summarize evidence; nobody has yet measured AI’s role in device requests.

Physicians are heavy AI users. The [American Medical Association’s 2026 survey](https://www.fiercehealthcare.com/ai-and-machine-learning/ama-physicians-use-ai-doubled-2023-2026) found 81% use AI professionally, and [The ASCO Post reports](https://ascopost.com/news/march-2026/ama-survey-finds-rapid-growth-in-physician-ai-adoption/) that 39% used it for summarizing medical research in 2026, up from 13% in 2024. That research step is where a clinician forms a view about a new device class.

Clinicians increasingly use AI tools built for them:

- **OpenEvidence.** Beyond the 760,000 registered physicians reported in December 2025, the company claims that more than 65% of US physicians use it monthly, [research firm Sacra reports](https://sacra.com/c/openevidence/). That figure is the company’s own and has not been independently checked. Sacra describes its search as running “exclusively through licensed medical content” such as NEJM, JAMA, specialty guidelines and drug labels.
- **General assistants for clinicians.** Sacra also notes that OpenAI launched ChatGPT for Clinicians, a free product for verified clinicians, in April 2026.
- **Established references.** UpToDate, which serves more than 3 million clinicians, launched an AI product in October 2025, according to the same Sacra profile.

Patients are part of demand too, especially for devices they can ask for by name. In a Rock Health survey [reported by HIT Consultant](https://hitconsultant.net/2026/03/23/rock-health-2025-survey-consumer-ai-adoption-chatgpt-healthcare/), 32% of US adults had used an AI chatbot for health information.

## What do clinicians and committees ask AI about a device?

Questions about evidence, comparisons, clearance, cost and alternatives. We wrote the prompts below as illustrations; none was collected from a real clinician or committee.

| Asker | What worries them | Example prompt (ours) |
|---|---|---|
| Surgeon | Evidence | “What do published trials show for robotic versus laparoscopic approaches in this procedure?” |
| Interventionalist | Comparison | “How do the main closure devices for this access site compare in published studies?” |
| Value analysis lead | Cost | “What is the total cost of ownership of a surgical robot over five years?” |
| Supply chain | Alternatives | “Which cleared alternatives exist for this single-use device?” |
| Any reviewer | Regulatory status | “Is this device FDA cleared, and for which indications?” |
| Patient | Awareness | “Is there a less invasive option for this procedure, and which hospitals near me offer it?” |

Two features make device questions different from software questions. Many have a clinical edge, so the answer an assistant gives draws on journals and guidelines rather than vendor marketing. And regulatory status is a hard filter: a device either is or is not cleared for a stated use, and committees check it.

## How does an AI answer turn into a purchase order?

Through the clinician who requests the device, the committee that checks the evidence, and the procedures that follow.

As we read the evidence, the path runs like this:

1. A clinician reads about a technology, often through an evidence tool or assistant that cites journals and guidelines.
2. The clinician asks the hospital to consider it, which triggers a value analysis request.
3. The committee checks clinical evidence, safety, regulatory status and total cost, and may approve a trial. At NCH, any request to evaluate a new product “must be coordinated through the Value Analysis committee.”
4. Contracting follows, often through a purchasing group or integrated delivery network.
5. Revenue grows with use: instruments per procedure, service contracts, more surgeons adopting.

AI can touch steps 1 and 3. At step 1 it shapes what the clinician believes before your representative calls. At step 3, committee members may use the same tools to check claims. We infer that a device whose evidence is published, indexed and clearly summarized is easier to champion and easier to approve; no study has tested this.

There is a limit you must respect. FDA’s final guidance of January 7, 2025 on [scientific information on unapproved uses](https://www.ropesgray.com/en/insights/alerts/2025/01/fda-finalizes-guidance-on-communication-of-scientific-information) sets out how firms may share such information with health care providers, and [Hall Render notes](https://hallrender.com/2025/01/30/fda-issues-final-guidance-on-communications-about-unapproved-uses-of-medical-products/) that it requires communications to be truthful and non-misleading. GEO work for a device maker has to stay within your cleared or approved indications and your regulatory team’s review. Drugmakers work under a similar line, covered in [keeping drug brands accurate in AI answers](https://underneath.agency/resources/pharma-brand-visibility-ai-search).

## What decides whether AI tools cite your device’s evidence?

The tools disclose little; analysts describe clinician tools reading licensed literature, and committees ask for independent proof.

**Described by the tools or analysts.** Sacra’s description of OpenEvidence says it answers from licensed journals, guidelines and labels, with inline citations. General assistants such as ChatGPT and Google’s AI features run web searches and show source links, but none explains why one device is named over another.

**Observed in studies.** When [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study) logged ChatGPT’s background queries, it looked for reviews in 46.2% of its answers and ran a search aimed at a named publication, ranking or award in 43.8%. In medical devices, the equivalent of a named publication is a peer-reviewed journal, a society guideline or a respected trade outlet. For consumer health questions, our summary of [what ChatGPT-cited health sites have in common](https://underneath.agency/resources/chatgpt-cited-commercial-health-sites) shows institutions dominate the citations.

**What committees check.** These trust factors are specific to devices:

- **Independent clinical evidence.** Intuitive says it seeks to show outcomes “validated by rigorous, independent, and peer-reviewed evidence.” Committees want the same.
- **Regulatory facts.** Clearance or approval type, indications and any recalls. FDA classifies devices by risk, and Class III devices generally need premarket approval, so the pathway itself signals the level of evidence behind a product.
- **Cost and operations.** Total cost of ownership, training, service and how the device fits existing workflows.
- **Peer adoption.** Which comparable hospitals use the device and what they report.

**Our inference.** An AI tool can only repeat the evidence it can find. If your trial results sit in a PDF behind a form, or your indications differ between your website, your instructions for use and distributor listings, assistants have little reliable material to work with. A reasonable expectation is that clear, consistent and independently published evidence gives both clinicians and AI tools more to cite.

## What does it cost a device maker to be absent from these answers?

Mostly lost requests and slower committee approvals; nobody has measured the loss directly.

- **No champion, no request.** A clinician who never meets your technology in the evidence they read is unlikely to start a committee request. We infer this from the request-led process hospitals like NCH describe.
- **A competitor’s evidence frames the category.** If AI answers summarize a rival’s trials and not yours, the committee may begin from their framing.
- **Errors spread.** An assistant that misstates an indication or cites an outdated study can create a compliance problem as well as a sales one; see [how to fix wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers).
- **Answers change.** In [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), only 25.2% of the brands ChatGPT named for a question appeared in all five runs, so one good answer proves little.

## How does GEO work for a medical device company?

By making accurate, on-label evidence and device facts easy for AI tools to find and cite. Nothing guarantees a mention.

1. **One consistent identity for each device.** Use the same device name, clearance type and indications on your website, product pages, distributor listings and press materials. Keep them aligned with your labeling.
2. **Evidence in the places AI tools read.** Publish trials and real-world studies in peer-reviewed journals, support registry participation, and help societies and guideline authors find your data. Clinician tools that answer from licensed literature cannot cite what was never published.
3. **Public, plain-language evidence summaries.** Give each study a readable page with methods, population, outcomes and limitations, linked to the journal version. Committees and assistants both benefit.
4. **Committee-ready facts.** Publish what value analysis teams check first: regulatory status, cost and training considerations, and service terms, where commercial policy allows.
5. **Earned coverage.** Clinical trade press, conference presentations and independent reviews are the kind of sources studies associate with AI citations; [how brands build authority for AI search](https://underneath.agency/resources/how-brands-build-authority-for-ai-search) covers the method.
6. **Keep paid and earned apart.** OpenEvidence is funded by pharmaceutical and device advertising, Sacra reports. Ads may reach clinicians, but they are not the same as being cited in an answer; our piece on [whether AI search ads should be labeled](https://underneath.agency/resources/should-ai-search-ads-be-labeled) explains why that line matters.
7. **Monitor general and clinician assistants.** Ask your buyers’ questions repeatedly in ChatGPT, Gemini, Perplexity, Copilot, Claude and Google’s AI answers, and in clinician tools your teams can access; [how many prompts to track](https://underneath.agency/resources/how-many-prompts-to-track-ai-visibility) covers the sample size.

Every piece should pass medical, legal and regulatory review, as your promotional materials already do. Shortcuts such as planted reviews are a serious risk in a regulated field; see [legitimate GEO versus manipulation](https://underneath.agency/resources/legitimate-geo-vs-manipulation). For the software side of hospital purchasing, see our guide for [healthcare software companies](https://underneath.agency/resources/healthcare-software-ai-search). Sellers of hospital equipment bought through quote requests can turn to [winning hospital RFQs through AI answers](https://underneath.agency/resources/medical-equipment-leads-ai-search).

## What is still unknown about AI search and device adoption?

The link from AI answers to device requests and approvals has not been measured anywhere.

- **No data on value analysis committees and AI.** We found no survey asking committee members whether they use AI tools to check evidence.
- **Clinician tool figures come from the vendors.** OpenEvidence’s usage numbers are company-reported, summarized by Sacra and Wikipedia.
- **Clinician tools are changing fast.** ChatGPT for Clinicians launched in April 2026, and how it chooses sources is not public.
- **One hospital’s letter is not a census.** NCH’s rules show how strict committees can be, not how every hospital works.
- **Revenue effects are unproven.** See [does AI visibility drive business results](https://underneath.agency/resources/does-ai-visibility-drive-business-results) for the general evidence.

## Where should a medical device company begin?

Check what clinicians and committees would read about your device today, then fix the gaps in evidence and facts.

Write down the questions a clinician champion, a value analysis lead and a supply chain manager would ask about your category, plus the questions patients ask about the procedure. Ask each in more than one assistant and repeat it on another day, because one answer is a sample, not a verdict on your device. Record whether your device is named, whether its clearance and indications are stated correctly, and which studies are cited, yours or a competitor’s.

To go further, [ask us to review how AI tools present your device to clinicians and committees](https://underneath.agency/contact). The review shows where answers misstate your clearance or skip your evidence, and which published proof would most help your clinician champions get product requests through value analysis and into procedures. Every piece of the ongoing evidence and device-fact work stays on-label and goes through your regulatory review, as our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page explains.

## Frequently asked questions

### Do surgeons use AI tools to research new devices?

No study measures device research specifically. Physicians broadly use AI: 81% do so professionally, and 39% use it to summarize medical research.

### Can we promote off-label uses through AI-friendly content?

No. FDA’s January 2025 guidance governs how firms share information on unapproved uses with providers. Keep content within your indications and regulatory review.

### Does advertising on clinician AI tools get our device cited?

There is no public evidence that it does. Ads and answer citations are separate, and clinician tools describe answering from licensed literature.

### Which matters more for AI visibility, our website or journal publications?

For clinician tools, published evidence, since they describe answering from licensed journals and guidelines. Your website still matters for clear device facts.

### How long before GEO affects device sales?

Expect a long horizon. Capital contracting cycles are lengthy, committees meet on schedules, and evidence takes time to publish.

## Sources

- U.S. Food and Drug Administration, CDRH (2026), [CDRH Annual Report 2025](https://www.fda.gov/media/190779/download)
- Emergo by UL (2026), [US FDA issues 2025 annual report on medical device regulatory activities](https://www.emergobyul.com/news/us-fda-issues-2025-annual-report-medical-device-regulatory-activities)
- U.S. Food and Drug Administration (n.d.), [Classify Your Medical Device](https://www.fda.gov/medical-devices/overview-device-regulation/classify-your-medical-device)
- American Hospital Association (2026-03-11), [New AHA Report: Hospitals Face Increased Challenges and Financial Pressures as They Care for Patients](https://www.aha.org/press-releases/2026-03-11-new-aha-report-hospitals-face-increased-challenges-and-financial-pressures-they-care-patients)
- NCH Healthcare System (2026), [Vendor letter on value analysis review](https://nchmd.org/wp-content/uploads/2023/09/2026_7_31_NCH-Vendor-Letter.pdf)
- Intuitive Surgical (2026-02-03), [Form 10-K for fiscal year 2025](https://www.sec.gov/Archives/edgar/data/1035267/000103526726000010/isrg-20251231.htm)
- Wikipedia (2026), [OpenEvidence](https://en.wikipedia.org/wiki/OpenEvidence)
- Sacra (2026), [OpenEvidence revenue, valuation and funding](https://sacra.com/c/openevidence/)
- Fierce Healthcare (2026), [AMA: Physicians’ use of AI doubled from 2023 to 2026](https://www.fiercehealthcare.com/ai-and-machine-learning/ama-physicians-use-ai-doubled-2023-2026)
- The ASCO Post (2026-03), [AMA Survey Finds Rapid Growth in Physician AI Adoption](https://ascopost.com/news/march-2026/ama-survey-finds-rapid-growth-in-physician-ai-adoption/)
- HIT Consultant (2026-03-23), [Rock Health 2025 survey: consumer AI adoption](https://hitconsultant.net/2026/03/23/rock-health-2025-survey-consumer-ai-adoption-chatgpt-healthcare/)
- Ropes & Gray (2025-01), [FDA Finalizes Guidance on Communication of Scientific Information](https://www.ropesgray.com/en/insights/alerts/2025/01/fda-finalizes-guidance-on-communication-of-scientific-information)
- Hall Render (2025-01-30), [FDA Issues Final Guidance on Communications About Unapproved Uses of Medical Products](https://hallrender.com/2025/01/30/fda-issues-final-guidance-on-communications-about-unapproved-uses-of-medical-products/)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)

---

This is the Markdown twin of https://underneath.agency/resources/medical-device-demand-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How medical equipment sellers can win RFQs from AI search"
description: "Hospital and clinic buyers are adopting AI fast. Equipment sellers win RFQs when assistants can verify their prices, contracts, service and evidence."
canonical: "https://underneath.agency/resources/medical-equipment-leads-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How can a medical equipment company win more hospital RFQs when buyers ask AI first?

By making the facts a capital committee checks (price range, GPO contract status, service coverage and independent evidence) easy for an assistant to find and repeat. Hospitals buy equipment through committees, contracts and long service agreements, so an AI answer rarely closes a sale. It can decide who receives the request for quote, and no study yet measures how often that happens.

## The short version

1. The buyer base is concentrated and contracted: the US has 6,100 hospitals ([AHA](https://www.aha.org/statistics/fast-facts-us-hospitals)), and the [Healthcare Supply Chain Association](https://www.supplychainassociation.org/) says more than 7000 hospitals use a group purchasing organization, with GPOs saving $55 billion a year.
2. Capital is tight and judged on return: 22% of health system executives plan to cut capital spending by 10% in 2026 and 19% by 20% or more, and 77% call anticipated ROI the most critical purchasing factor for digital health, in a [Sage Growth Partners survey reported by HFMA](https://www.hfma.org/fast-finance/health-system-capital-investment-strategy-2026/).
3. Equipment is aging: the average age of plant at Fitch-rated not-for-profit hospitals reached 12.7 years in FY24, the oldest in at least 13 years, according to [HFMA](https://www.hfma.org/fast-finance/hospital-capital-expenditures-aging-facilities/).
4. The independent safety body warns about AI answers: [ECRI](https://home.ecri.org/blogs/ecri-news/misuse-of-ai-chatbots-tops-annual-list-of-health-technology-hazards) ranked misuse of AI chatbots the top health technology hazard for 2026 and says chatbots have “promoted subpar medical supplies.”
5. Assistants look for prices and reviews: in [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study), ChatGPT searched for reviews in 46.2% of answers and for prices in 23.8%.

## Who buys hospital and clinic equipment, and how?

Committees do: supply chain, clinical engineering, department leaders and finance, usually working inside GPO contracts.

A capital purchase in a hospital involves more people than a medical device sale to a surgeon. The usual cast:

- **Supply chain and procurement leaders**, who run the sourcing event and negotiate within contract tiers.
- **Clinical or biomedical engineering**, who judge reliability, service history, parts and cybersecurity on older equipment.
- **Department leaders** in imaging, surgery, labs or nursing, who define requirements and run trials.
- **Finance and capital committees**, who rank requests against everything else the system wants to fund.
- **Value analysis committees**, which compare clinical evidence and total cost before a product enters the formulary. How device makers reach clinicians before that review is covered in [building device demand ahead of the committee](https://underneath.agency/resources/medical-device-demand-ai-search).

Most of them buy through group purchasing organizations such as Vizient, Premier and HealthTrust. HSCA counts more than 100 national, regional and local GPOs competing to serve hospitals. Systems matter too: the [AHA counts](https://www.aha.org/statistics/fast-facts-us-hospitals) 3,567 community hospitals in a system, where one decision can cover many sites, and 1,797 rural community hospitals, where budgets are thinner and refurbished equipment is often the realistic option.

Outside hospitals, clinics, imaging centers and ambulatory surgery centers buy the same categories with smaller teams and faster cycles. That is where refurbished dealers and independent servicers compete hardest. Sellers to dental offices, another small-team buyer, can compare notes with [dental technology on a dentist’s shortlist](https://underneath.agency/resources/dental-technology-ai-search).

## What is one customer worth to an equipment company?

The first sale is the smaller part. Service contracts, parts, upgrades and replacements carry the account for a decade or more.

Prices span orders of magnitude. [Block Imaging’s price guide](https://www.blockimaging.com/blog/1.5t-mri-machine-price-cost-guide), updated in January 2026, says refurbished 1.5T MRI machines sell for $150,000 to $500,000, and that 4 to 8-channel coils cost $12,000 to $80,000 new but $8,000 to $25,000 refurbished. New systems from the original manufacturers cost more, and the guide notes that upgrades, delivery, installation and first-year service often change the true price.

The account lasts because equipment is replaced slowly. HFMA reports that Fitch-rated systems spent a median 123.4% of depreciation on capital in FY24, the highest since 2013, as they caught up on aging plant. Northwell Health alone plans about $1.5 billion a year in capital spending. For an equipment seller, being on the shortlist when a system replaces a fleet of monitors or a set of C-arms can shape the next ten years of service revenue.

## How far has AI reached into hospital purchasing?

Into clinical and administrative work, quickly. Into vendor shortlists, likely but unmeasured.

The people who influence equipment choices now use AI routinely. The AMA’s 2026 survey of nearly 1,700 physicians found 81% use AI professionally, and 39% use it for summaries of research and standards of care, according to [Fierce Healthcare](https://www.fiercehealthcare.com/ai-and-machine-learning/ama-physicians-use-ai-doubled-2023-2026). A McKinsey survey of US healthcare leaders, [reported by DistilINFO](https://distilinfo.com/2026/04/20/half-of-us-hospitals-now-use-generative-ai/), found 50% of healthcare organizations actively use generative AI and more than 80% have deployed at least one use case.

In the Sage survey, AI-based clinical technology led planned technology initiatives, with 57% planning such spending in 2026 and 2027. Across all business purchases, [Gartner’s 2026 survey](https://www.gartner.com/en/newsroom/press-releases/2026-05-20-gartner-survey-finds-sixty-nine-percent-of-b-two-b-buyers-turn-to-sales-reps-to-validate-ai-generated-insights) of 645 buyers found 45% used generative AI in a recent purchase, mainly to gather information on vendors and products. No survey isolates hospital equipment buyers, so treat that as direction.

ECRI’s warning cuts both ways. It notes that more than 40 million people a day turn to ChatGPT for health information, and it found chatbots have suggested incorrect diagnoses and “promoted subpar medical supplies,” and it advises verifying chatbot information with a knowledgeable source. Buyers who follow that advice will check what an assistant tells them about your equipment. That favors companies whose facts are easy to confirm.

## Which questions do equipment buyers ask before they send an RFQ?

Questions about total cost, contract status, service, reliability and alternatives. We wrote these in buyers’ voices; they are illustrations, not logged queries.

| Buyer | Illustrative prompt |
|---|---|
| Imaging center owner | “Refurbished 1.5T MRI with a 70cm bore for under $400,000, including installation and first-year service” |
| Supply chain director | “Which C-arm makers are on our GPO contract, and how do their service terms compare?” |
| Clinical engineering manager | “Reliability and parts availability for patient monitors installed before 2018” |
| Rural hospital CFO | “Lease or buy a CT scanner for a 25-bed critical access hospital” |
| Surgery center administrator | “Independent servicers for endoscopy towers in Ohio with loaner equipment” |
| Value analysis lead | “Published evidence comparing infusion pump platforms on safety and total cost” |

Where a question names a place, answers diverge more. In [our brand agreement study](https://underneath.agency/research/ai-assistants-brand-agreement-study), two assistants’ recommendations overlapped 0.160 on questions that named a place, against 0.390 on national ones. For dealers and servicers that sell by region, a reasonable expectation is that regional questions are where visibility varies most.

## How does an AI answer become a qualified equipment lead?

By putting your company on the list that receives the RFQ, with the key facts already settled.

The path we see in this industry:

1. **Need.** Equipment reaches end of life, a service contract lapses, a new service line opens, or a recall forces replacement.
2. **Market scan.** The GPO contract catalog, distributor reps, ECRI evaluations, peers at other systems, web search and, increasingly, an assistant.
3. **RFI or RFQ.** Requirements go to a short list of manufacturers, distributors or refurbishers.
4. **Evaluation.** Demos, site visits, trials and a value analysis or capital committee review.
5. **Contract and install.** Pricing within the GPO tier, delivery, training.
6. **Service life.** Multi-year service, parts, upgrades and, eventually, the replacement decision.

A good lead arrives at step 3 already knowing your price band, your contract status and whether you service equipment in their region. A refurbished dealer that publishes clear price ranges, as Block Imaging does, gives an assistant something concrete to repeat when an imaging center asks what a used MRI costs. Most of this influence leaves no click to track; we explain why in [how AI answers affect pipeline](https://underneath.agency/resources/ai-answers-pipeline-revenue).

## What decides whether an assistant names your equipment or your company?

No platform documents how it picks equipment vendors; studies show assistants search for reviews, prices and rankings.

**Documented by the platforms.** According to Google, one question in AI Overviews or AI Mode [can be split into several related searches](https://developers.google.com/search/docs/appearance/ai-features), a technique it calls “query fan-out”. A question about refurbished CT scanners may fan out into searches about prices, warranties and service. Google does not explain how a vendor is chosen.

**Observed in studies.** In our hidden-searches study, ChatGPT ran a search aimed at a named publication, ranking or award in 43.8% of answers. In [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), only 25.2% of brands named appeared in all five runs of the same question. In a skincare experiment by [Chu and Hou](https://arxiv.org/abs/2606.17443), a fabricated clinical citation had the same effect as +0.17 rating points of real product improvement. That shows how much weight authority language can carry, and why invented evidence is a real risk in a field ECRI already watches for substandard products.

**What hospital buyers check.** ECRI describes itself as the only organization worldwide to conduct independent medical device evaluations, and its hazard reports are read by hospitals, health systems and manufacturers. Buyers also check GPO contract status, service response times, parts availability, cybersecurity on legacy devices and, for refurbished equipment, who did the work. The FDA’s May 10, 2024 [final guidance on remanufacturing](https://www.morganlewis.com/pubs/2024/05/fda-clarifies-distinction-between-device-remanufacturing-and-servicing-in-final-guidance), as summarized by Morgan Lewis, separates servicing, which faces limited oversight for independent third parties, from remanufacturing, which requires full compliance. A refurbisher that states which it does, plainly, answers a question buyers and assistants both ask.

**Our inference.** A reasonable expectation is that companies whose prices, contracts, service coverage and evidence are stated consistently across their own site, distributor listings and independent sources give assistants more to repeat accurately. No study has tested this for medical equipment.

## What does it cost an equipment company to be missing?

RFQs that go to competitors and service revenue that follows them, though no study has measured that loss.

- **The RFQ list is short.** A vendor that is not on it does not get to compete on price or terms.
- **Service follows the sale.** Losing an install often means losing years of service, parts and upgrades with it.
- **Capital is rationed.** With 19% of executives planning capital cuts of 20% or more, the purchases that do happen are scrutinized, and an assistant that cannot find your total cost may leave you out of the comparison.
- **Wrong facts filter buyers out.** An answer that says you do not service a region, or are not on a contract, sends the buyer elsewhere. The steps for setting the record straight are in [how to fix wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers).

## How does GEO work for a medical equipment company?

It makes your price logic, contract status, service and evidence easy to find and confirm. Nobody can promise a recommendation.

1. **Publish price ranges and total cost.** Purchase, installation, service, parts and expected life. Assistants search for prices, and capital committees ask for this first.
2. **State contract status consistently.** Which GPO agreements you hold and which product lines they cover, matching what distributors list.
3. **Make service coverage concrete.** Regions, response times, uptime commitments, loaner policy and parts availability, on pages a buyer can find.
4. **Explain refurbishment honestly.** Process, standards, warranty, and whether work is servicing or remanufacturing under the FDA’s definitions.
5. **Earn independent evidence.** ECRI evaluations where they exist, peer-reviewed studies, trade press and case studies that name the hospital and the result, with permission. [How brands build authority for AI search](https://underneath.agency/resources/how-brands-build-authority-for-ai-search) explains why this outside proof carries weight.
6. **Avoid unverifiable clinical claims.** Authority language works on assistants, which is exactly why invented or overstated evidence is a liability in a regulated market. What health pages that AI cites tend to show is covered in [what cited commercial health sites have in common](https://underneath.agency/resources/chatgpt-cited-commercial-health-sites).
7. **Align distributor and marketplace listings** so specifications, prices and model names match.
8. **Measure with real buyer wording.** Test procurement, clinical engineering, department and refurbished-buyer prompts, with region and budget variations, across ChatGPT, Google AI Overviews and AI Mode, Gemini, Perplexity, Copilot and Claude, several runs each.

Companies that also sell software into hospitals face a different evaluation; see [how healthcare software companies win health system deals](https://underneath.agency/resources/healthcare-software-ai-search).

## What remains unknown about AI and equipment buying?

Whether hospital equipment buyers use assistants to build RFQ lists, and whether being named wins contracts. No published study answers either.

- **No survey of equipment buyers.** The AMA, McKinsey and Gartner figures describe clinicians, healthcare organizations and business buyers in general.
- **The ROI figure covers digital health purchasing**, not capital equipment alone.
- **Price guides are a dealer’s own.** Block Imaging’s ranges describe its market view, not audited prices.
- **GPO and committee influence is unmeasured.** Nobody has studied how AI answers interact with contract catalogs and value analysis reviews.

## Where should a medical equipment company start?

With the questions buyers ask before they send an RFQ, and a record of what assistants answer today.

List 20 to 30 prompts across supply chain, clinical engineering, department leaders and clinic or imaging center buyers, including price, contract, service and region questions. Ask each major assistant several times. Note which companies are named, which sources are cited, and whether your contract status, service coverage and prices are described correctly. Compare notes with [how procurement teams buy software](https://underneath.agency/resources/procurement-software-ai-search), which covers buyers who run sourcing events for a living.

When you want a full picture, [request an RFQ-stage AI visibility audit](https://underneath.agency/contact). We will show which procurement and clinical engineering questions name your company, where assistants send buyers instead, and which missing or inconsistent facts are most likely costing you quote requests and service contracts. Our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page explains how we keep price ranges, contract status and service coverage consistent for capital committees and the assistants they consult.

## Frequently asked questions

### Do hospital buyers use ChatGPT to find equipment vendors?

Nobody has measured it for equipment. Clinicians and healthcare organizations use AI widely: 81% of physicians in the AMA survey and 50% of organizations in McKinsey’s.

### Should equipment companies publish prices?

Publish ranges and total cost at least. ChatGPT searched for prices in 23.8% of answers in our hidden-searches study, and capital committees judge on return.

### Does GPO contract status affect AI answers?

No study shows it does. It is a fact buyers check, so state it consistently on your site and distributor listings so assistants repeat it correctly.

### How should refurbished equipment dealers approach AI search?

Publish price ranges, warranties, service regions and whether you service or remanufacture. Regional questions vary most between assistants.

### Is it risky to promote equipment with clinical claims?

Yes, if the claims cannot be verified. Authority language can sway AI answers, and ECRI already flags substandard medical products as a 2026 hazard.

## Sources

- American Hospital Association (2026-02), [Fast facts on US hospitals, 2026](https://www.aha.org/statistics/fast-facts-us-hospitals)
- Healthcare Supply Chain Association, [HSCA home page](https://www.supplychainassociation.org/)
- HFMA (2026-04-14), [Health system capital investment strategy 2026](https://www.hfma.org/fast-finance/health-system-capital-investment-strategy-2026/)
- HFMA (2026-02-10), [Hospital capital expenditures and aging facilities](https://www.hfma.org/fast-finance/hospital-capital-expenditures-aging-facilities/)
- ECRI (2026-01-21), [Misuse of AI chatbots tops annual list of health technology hazards](https://home.ecri.org/blogs/ecri-news/misuse-of-ai-chatbots-tops-annual-list-of-health-technology-hazards)
- Block Imaging (2026-01-19), [1.5T MRI machine cost: price guide](https://www.blockimaging.com/blog/1.5t-mri-machine-price-cost-guide)
- Morgan Lewis (2024-05), [FDA clarifies distinction between device remanufacturing and servicing in final guidance](https://www.morganlewis.com/pubs/2024/05/fda-clarifies-distinction-between-device-remanufacturing-and-servicing-in-final-guidance)
- Fierce Healthcare (2026), [AMA: physicians’ use of AI doubled from 2023 to 2026](https://www.fiercehealthcare.com/ai-and-machine-learning/ama-physicians-use-ai-doubled-2023-2026)
- DistilINFO (2026-04-20), [Half of US hospitals now use generative AI](https://distilinfo.com/2026/04/20/half-of-us-hospitals-now-use-generative-ai/)
- Gartner (2026-05-20), [Gartner survey finds 69% of B2B buyers turn to sales reps to validate AI-generated insights](https://www.gartner.com/en/newsroom/press-releases/2026-05-20-gartner-survey-finds-sixty-nine-percent-of-b-two-b-buyers-turn-to-sales-reps-to-validate-ai-generated-insights)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- Xi Chu and YuPeng Hou (2026), [Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems](https://arxiv.org/abs/2606.17443)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude agree on brands?](https://underneath.agency/research/ai-assistants-brand-agreement-study)
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)

---

This is the Markdown twin of https://underneath.agency/resources/medical-equipment-leads-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Can AI search bring independent practices to your software demo?"
description: "Yes, if assistants find proof of specialty fit, certification, pricing and peer ratings; 23% of medical groups expect to switch or overhaul their EHR in a year."
canonical: "https://underneath.agency/resources/medical-practice-software-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Can AI search bring independent practices and specialty clinics to a medical software vendor’s demo?

It can help you make the short list, provided AI answers can find clear proof that your system fits the practice’s specialty, size, billing setup and certification needs. Nearly every office-based physician already uses an electronic health record, so most new customers are switchers, and they decide which vendors to demo before they talk to sales. No study yet measures how often practice owners ask AI assistants about software, so treat this as a strong bet backed by adjacent evidence, not a proven channel.

This article covers practice management systems, ambulatory and specialty EHRs and related tools sold to independent practices. For large hospital and health system buying, read our guide for healthcare software sold to health systems. Companies selling to dental offices should read how dental technology makes a dentist’s shortlist.

## The short version

1. The market is replacement-driven: 95% of US office-based physicians had adopted an EHR by 2024, and 91% a certified one, according to the [federal health IT office (ASTP/ONC)](https://www.healthit.gov/data/quickstats/office-based-physician-electronic-health-record-adoption).
2. Switching is common: in a March 2025 [MGMA Stat poll](https://www.mgma.com/mgma-stat/building-a-robust-rfi-process-and-rfp-for-a-new-ehr-system) of 455 practice leaders, 23% expected to switch or significantly update their EHR within 12 months.
3. The independent buyer is shrinking but still large: the [American Medical Association](https://www.medicaleconomics.com/view/ama-physician-private-practice-unraveling-due-to-low-payment-high-costs-administrative-burdens) found 42.2% of physicians worked in private practice in 2024, down from 60.1% in 2012, and 47.4% in practices of 10 or fewer physicians.
4. Physicians are heavy AI users: [81% used AI professionally](https://www.fiercehealthcare.com/ai-and-machine-learning/ama-physicians-use-ai-doubled-2023-2026) in the AMA’s 2026 survey, and 68% of medical groups added or expanded AI tools in 2025, per [MGMA](https://www.mgma.com/mgma-stat/document-schedule-communicate-ai-tools).
5. Trust must be earned with evidence: in [Black Book’s 2025 survey](https://www.accessnewswire.com/newsroom/en/healthcare-and-pharmaceutical/physician-practice-it-at-a-crossroads-black-book-survey-finds-progres-1077124) of 755 practice managers, only 19% trusted the AI prompts in their current systems, and 44% would be more likely to adopt AI decision support with transparent evidence trails.

## Who buys software for an independent practice, and what is a customer worth?

A physician owner and a practice administrator, usually together, with the billing lead holding an informal veto.

In a small or specialty practice, the person who signs is often a physician owner, but the person who runs the evaluation is the practice manager or administrator, and the person who lives with the result is the billing lead. In the AMA’s 2024 benchmark survey, 35.4% of physicians had an ownership stake in their practice, and single-specialty practices employed 37% of physicians, more than multi-specialty practices (27.8%). The same survey shows ownership is shifting: 38% of physicians in private equity owned practices said those practices were acquired in the past five years, which brings management service organizations into many buying decisions. Hospitals and health systems buy through a different process, covered in [our guide for healthcare software sold to health systems](https://underneath.agency/resources/healthcare-software-ai-search).

What a customer is worth depends on the sales model, and the two common models look very different:

| Model | Real example | How price is set |
|---|---|---|
| Self-serve trial | [SimplePractice](https://www.simplepractice.com/pricing/) for solo practitioners | Published plans at $49, $79 and $99 a month, with a free trial that can run to 30 days |
| Demo-led quote | [Tebra](https://www.tebra.com/pricing/) for independent practices | Licensed per prescribing provider; Tebra says “pricing is typically tailored during the demo process” |

In both models, the account is sticky. Moving clinical records, templates, claims history and staff habits is painful, so a won practice usually stays for years. We found no public benchmark for the lifetime value of a practice software customer and will not invent one. Our inference is that every won practice is worth several years of per-provider fees plus add-ons such as billing services, patient engagement and AI documentation.

## Where do AI assistants sit in a practice’s software search?

Probably near the start, since buyers already use AI daily, though direct evidence on vendor research is missing.

The adjacent evidence is strong. The AMA’s 2026 survey found 81% of physicians use AI in a professional context, more than double the 2023 rate, and 39% use it for summaries of research and standards of care. The [Doximity 2026 report cited by MGMA](https://www.mgma.com/mgma-stat/are-ai-tools-making-clinicians-more-productive-its-complicated) found ambient documentation adoption rose from 20% to 29% of physicians surveyed in a year. Practice leaders are buying AI, too: MGMA reports 68% of medical groups added or expanded AI tools in 2025.

Across all B2B software, [G2’s 2026 buyer research](https://company.g2.com/news/g2-research-the-answer-economy) found 51% of buyers now start their research with an AI chatbot more often than with Google. G2 did not break out healthcare practices, and we found no survey that asks practice managers how they research vendors. Our inference: a practice administrator comparing systems is likely to ask an assistant the same questions they would ask a peer, and the answer shapes which three vendors get a demo.

## What do practice owners and managers ask before they book a demo?

Specialty fit, billing and clearinghouse integration, certification, pricing, switching effort and what peers say.

We wrote the practice-manager and physician prompts below ourselves to show how buying questions tend to be phrased; none was captured from a real user.

| Stage | Illustrative prompt |
|---|---|
| Specialty fit | “Best dermatology EHR for a three-provider practice with a cosmetic side” |
| Switching | “Alternatives to our current EHR that can migrate charts and claims history” |
| Integration | “Practice management software that works with our clearinghouse and lab” |
| Certification | “Is this EHR on the ONC certified list for MIPS reporting?” |
| Price | “How much does an ophthalmology EHR cost per provider per month?” |
| AI features | “Which EHRs include an ambient AI scribe, and what does it cost extra?” |
| Peer proof | “What do billing managers say about this system’s denial management?” |

Some of these answers are easy to get wrong. Vendor names change after mergers: Tebra’s own pricing page notes that “Kareo is now part of Tebra.” An assistant that learned about the market before a rename may describe a product that no longer exists under that name. We explain why that happens in [why ChatGPT misses new products](https://underneath.agency/resources/why-chatgpt-misses-new-products).

## How does an AI answer become a signed practice?

Through a shortlist, a demo or trial, a quote and a migration plan, usually over weeks rather than months.

1. **Trigger.** A practice opens a new location, joins a management services organization, loses patience with a billing system or wants an AI scribe. MGMA’s respondents planning a change mostly said they were moving to a new vendor rather than upgrading.
2. **Shortlist.** The manager asks peers, checks KLAS or Black Book ratings, reads reviews and, we infer, increasingly asks an assistant. Vendors not named here rarely get a demo.
3. **Demo or trial.** Self-serve products convert through a trial; quote-based products through a demo where the price is set.
4. **Proof.** The practice checks references, certification, integrations and the contract. In Black Book’s survey, 17% of managers said they would replace their EHR or practice management platform with a vendor that guaranteed real-time FHIR-based data exchange.
5. **Contract and migration.** The practice signs per provider, and the revenue repeats for years.

Pricing is a weak point in step two. In [our pricing accuracy study](https://underneath.agency/research/ai-pricing-accuracy-study) across 45 software products, 61.9% of plan prices quoted by AI assistants were fully faithful to the official page. When prices differed, the figure often came from somewhere else on the vendor’s own site: for 39 of 64 differing prices, the same number appeared on another vendor page. For quote-based vendors, our inference is that a page explaining how pricing works (per provider, what is included, what costs extra) is the best defense against a wrong number.

## What decides which medical software vendor an assistant names?

Independent evidence of specialty fit and customer satisfaction; the platforms do not publish their selection rules.

What is documented: [Google says](https://developers.google.com/search/docs/appearance/ai-features) SEO best practices remain relevant for AI Overviews and AI Mode, with no additional requirements or special optimizations needed to appear.

What has been observed:

- **Answers vary.** In [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), only 25.2% of the brands ChatGPT named appeared in all five repeats of the same question. One check of “best EHR for my specialty” proves little.
- **Mention matters to B2B buyers.** In G2’s research, 85% of buyers said they think more highly of a vendor when AI includes it in an answer. That is a buyer survey across software, not a test of what assistants do.

What we infer, based on how practices already judge vendors: the evidence an assistant can find is the same evidence a practice manager trusts. That includes:

- **Independent ratings.** [KLAS](https://engage.klasresearch.com/press-release/epic-athenahealth-chartis-optimum-healthcare-it-and-impact-advisors-win-2026-overall-best-in-klasawards/) named athenahealth the Overall Independent Physician Practice Suite for a third consecutive year in 2026, and Greenway Health the most improved physician practice product after a 24% rise in satisfaction among its revenue cycle clients. Black Book’s 2025 ratings named ModMed for OB/GYN and NextGen Healthcare for practice management.
- **Certification and compliance.** Certified EHR status, a signed business associate agreement and interoperability readiness, stated plainly.
- **Specialty depth.** Templates, imaging, devices and workflows for the specialty, described on pages an assistant can read.
- **Transparent AI.** Black Book found 44% of managers would be more likely to adopt AI decision support if vendors showed the data source, logic and limits behind each recommendation. The same openness on your website gives an assistant something accurate to repeat.

## What does a medical software vendor lose when it is missing from AI answers?

A place on the shortlist during a switching window that may not reopen for years.

We have no direct measurement of deals lost to AI absence, so we label the reasoning. With 23% of medical groups planning to switch or overhaul their EHR in a year, the window is real but brief for any one practice. A vendor left off the shortlist loses that practice until the next switch, and a demo-led vendor loses the chance to set the price in a conversation. A vendor described with an old name, an old price or a feature it no longer offers may be shortlisted for the wrong reasons and lost in the demo. If an assistant still describes your old pricing or a retired feature, [the steps for fixing wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) apply.

## How does generative engine optimization work for practice software?

Generative engine optimization (GEO) makes your product easy for assistants to find, describe accurately and verify through outside evidence.

For a vendor selling to independent practices and specialty clinics, the work usually covers:

1. **One consistent identity.** The same product names, specialties served and practice sizes across your site, KLAS and Black Book profiles, review sites, marketplaces and partner pages, especially after a merger or rename.
2. **A page per specialty.** What the system does for dermatology, orthopedics, ophthalmology or behavioral health: templates, imaging, devices, billing codes and workflows, in text rather than only in videos. Dental offices are a separate market, covered in [how dental technology makes a dentist’s shortlist](https://underneath.agency/resources/dental-technology-ai-search).
3. **Pricing logic in writing.** Published plans, or a clear explanation of how quotes are built: per provider or per user, what is included, implementation fees and add-ons.
4. **Proof of certification and integration.** Certified status with a link to the [ONC Certified Health IT Product List](https://www.healthit.gov/topic/certification-ehrs/about-onc-health-it-certification-program), named clearinghouse, lab and imaging integrations, and a plain statement on business associate agreements.
5. **Switching guides.** Honest migration pages for practices leaving common systems: what moves, how long it takes and what it costs. Fair comparison pages also help; see [whether comparison pages help B2B AI citations](https://underneath.agency/resources/do-comparison-pages-help-b2b-ai-citations).
6. **Independent coverage and reviews.** KLAS and Black Book participation, specialty society and MGMA presence, case studies with named practices, and reviews from billing managers as well as physicians.
7. **Measurement.** Ask a fixed set of illustrative specialty, switching, pricing and certification questions across ChatGPT, Gemini, Perplexity and Google’s AI features, repeatedly, and track who is named and which sources are cited.

No vendor can be promised a place on an assistant’s shortlist, and anyone who promises one is overselling. What this work does is make your product the easiest one for an assistant to describe correctly to a practice that is ready to switch.

## What don’t we know yet about how practices choose software with AI?

How often practice managers ask assistants about vendors, and whether AI-sourced demos close at different rates.

- **No direct survey.** The AMA, Doximity and MGMA measure clinical and operational AI use, not vendor research. G2’s figures cover B2B software buyers in general.
- **Self-reported data.** MGMA Stat polls and Black Book surveys rely on what respondents say, and Black Book also publishes vendor ratings.
- **No tests on medical software queries.** Our consistency and pricing studies used other categories and products. Applying them here is our inference.
- **Platforms change.** Assistants update their search and answer features often, and none publishes how it chooses vendors.

## Where should a practice software company start?

Start with the questions a switching practice in your specialty would ask, then see who assistants name and why.

That first check usually shows whether your product is named for the specialties and practice sizes you serve, whether your name, pricing and certification are described correctly, which ratings and reviews shape the answer and which competitors appear instead. From there, the work is to publish the specialty, pricing and integration proof that a practice manager would want and to keep it consistent everywhere.

If your growth depends on demo requests and trials from independent practices, [ask us to map your visibility in AI answers](https://underneath.agency/contact). We will show where your product appears for specialty and switching questions, why competitors are named instead, and which changes are most likely to bring more qualified practices to your demo calendar. What the follow-on work involves, from specialty pages to pricing logic and switching guides, is described on our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page.

## Frequently asked questions

### Do practice managers use ChatGPT to choose an EHR?

No survey we found measures it directly. Physicians are heavy AI users, 81% in the AMA’s 2026 survey, and G2 found 51% of B2B software buyers start research with an AI chatbot.

### Should we publish our prices if we sell through demos?

Publishing at least the pricing logic helps. In our study, AI assistants quoted software prices fully faithfully 61.9% of the time, and many errors came from other pages on vendors’ own sites.

### Do Best in KLAS awards affect AI answers?

No study shows that directly. We infer that independent ratings help, because they are the evidence practices already trust and they appear on pages assistants can read.

### Is this different from selling to health systems?

Yes. Independent practices buy faster, with smaller committees and per-provider pricing, and often through trials. Health systems buy through long, committee-led evaluations tied to their main EHR.

## Sources

- ASTP/ONC (2024), [Office-based Physician Electronic Health Record Adoption](https://www.healthit.gov/data/quickstats/office-based-physician-electronic-health-record-adoption)
- MGMA (2025-03-19), [Building a robust RFI process and RFP for a new EHR system](https://www.mgma.com/mgma-stat/building-a-robust-rfi-process-and-rfp-for-a-new-ehr-system)
- Medical Economics (2025-05-29), [AMA: Physician private practice unraveling due to low payment, high costs, administrative burdens](https://www.medicaleconomics.com/view/ama-physician-private-practice-unraveling-due-to-low-payment-high-costs-administrative-burdens)
- Fierce Healthcare (2026), [AMA: Physicians’ use of AI doubled from 2023 to 2026](https://www.fiercehealthcare.com/ai-and-machine-learning/ama-physicians-use-ai-doubled-2023-2026)
- MGMA (2025), [Document, schedule, communicate: AI tools](https://www.mgma.com/mgma-stat/document-schedule-communicate-ai-tools)
- MGMA (2026-05), [Are AI tools making clinicians more productive? It’s complicated](https://www.mgma.com/mgma-stat/are-ai-tools-making-clinicians-more-productive-its-complicated)
- Black Book Research (2025-09-26), [Physician Practice IT at a Crossroads](https://www.accessnewswire.com/newsroom/en/healthcare-and-pharmaceutical/physician-practice-it-at-a-crossroads-black-book-survey-finds-progres-1077124)
- KLAS Research (2026-02-04), [Epic, athenahealth, Chartis, Optimum Healthcare IT and Impact Advisors win 2026 Overall Best in KLAS Awards](https://engage.klasresearch.com/press-release/epic-athenahealth-chartis-optimum-healthcare-it-and-impact-advisors-win-2026-overall-best-in-klasawards/)
- SimplePractice (2026), [SimplePractice EHR Pricing and Plans](https://www.simplepractice.com/pricing/)
- Tebra (2026), [Tebra Pricing: Software Plans for Private Practices](https://www.tebra.com/pricing/)
- G2 (2026-04-15), [In the Answer Economy, Don’t Win the Click — Win the Answer](https://company.g2.com/news/g2-research-the-answer-economy)
- ASTP/ONC (n.d.), [About the ONC Health IT Certification Program](https://www.healthit.gov/topic/certification-ehrs/about-onc-health-it-certification-program)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- Underneath (2026), [How faithfully do AI assistants quote software prices?](https://underneath.agency/research/ai-pricing-accuracy-study)
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)

---

This is the Markdown twin of https://underneath.agency/resources/medical-practice-software-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How mortgage platforms win qualified borrowers from AI answers"
description: "By being named, with accurate loan facts, when borrowers ask AI to estimate payments and compare lenders, and by keeping rate claims compliant."
canonical: "https://underneath.agency/resources/mortgage-platforms-leads-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How do mortgage lenders and marketplaces win qualified borrowers when buyers start with AI?

By being named, and described accurately, when a buyer asks an assistant what they can afford and which lender to use, and by making sure every rate and fee claim an assistant might repeat already follows mortgage advertising rules. A growing share of buyers now run payment estimates and lender comparisons in AI tools before they request a single quote. For a lender or marketplace, that is where a qualified lead is now won or lost.

## The short version

1. AI is now in the mortgage shopping step: in a [Veterans United survey](https://www.veteransunited.com/education/ai-homebuying-survey/) of 859 people in June 2026, 45% of prospective buyers had used AI tools in their home search, up from 37% a year earlier, and 30% used them to shop for the best interest rate.
2. Affordability is the entry point: among prospective buyers who used AI in a [Bank of America study](https://newsroom.bankofamerica.com/content/newsroom/press-releases/2026/06/bofa-study--more-americans-favor-buying-over-renting-for-the-fir.html), 57% used it to estimate affordability, mortgage payments or closing costs.
3. A funded loan barely clears its cost: lenders in the Mortgage Bankers Association’s second-quarter 2026 report, [as reported by HousingWire](https://housingwire.com/articles/imb-mortgage-profits-q2-2026), earned $11,909 per loan and spent $10,936, a pretax profit of $973.
4. Shopping changes what borrowers pay: [Freddie Mac](https://www.freddiemac.com/research/insight/20230216-when-rates-are-higher-borrowers-who-shop-around-save) estimated buyers could save $600 to $1,200 a year by applying with more than one lender, so the shortlist an assistant gives matters.
5. Google shows AI answers on most finance searches: in [our study](https://underneath.agency/research/ai-overviews-frequency-study), financial services and insurance searches triggered an AI Overview 88.0% of the time.

This guide is about how mortgage lenders and marketplaces get found and described in AI answers. It is not mortgage, legal or compliance advice; take rate, fee and referral questions to your counsel and compliance team.

## Who chooses a mortgage platform, and what is a funded loan worth?

A home buyer or homeowner chooses, usually under time pressure; each funded loan carries thin but real margin.

The buyer is a household, not a procurement team. They arrive at a few moments: they need a preapproval letter to make an offer, rates have dropped and a refinance looks worth it, or a life change forces a move. Younger borrowers carry growing weight. A [Cotality survey reported by HousingWire](https://housingwire.com/articles/homebuyers-want-ai-and-human-in-the-loop-cotality-2026-survey) notes that in the US, buyers under 35 account for 37% of originated loans.

The economics explain why lead quality matters more than lead volume. In the MBA’s second-quarter 2026 figures, lenders took in $11,909 in production revenue per loan against $10,936 in cost, for a pretax production profit of $973. The average first mortgage was $386,359, and purchase loans made up 80% of first-mortgage originations by dollar volume among the companies reporting. A lender that pays for leads that never close eats that thin margin quickly. A marketplace earns by passing borrowers to lenders, and its value to those lenders rests on how many of those borrowers fund. Platforms that also offer personal, auto or business loans can compare notes with [how lending platforms reach borrowers through AI](https://underneath.agency/resources/lending-platforms-borrowers-ai-search).

Two kinds of platform compete for the same borrower:

| Platform | How it earns from a borrower | What AI visibility has to deliver |
|---|---|---|
| Direct lender (bank unit, independent mortgage bank, digital lender) | Gain on sale and fees on a funded loan | Applications from borrowers who fit its loan programs |
| Marketplace or comparison platform | Payments from lenders for leads or click-throughs | Borrowers who complete a form and then fund with a participating lender |

## How many borrowers already use AI before they pick a lender?

Between one in five and almost half, depending on the survey; most use it for payment math and research.

- **Veterans United (June 2026).** 45% of prospective buyers had used AI tools. The leading uses were searching for homes (52%) and estimating monthly payments (43%). Shopping for the best interest rate (30%) and comparing lender reviews (28%) followed. ChatGPT was the most used tool (33%), then Gemini (20%). Veterans United is itself a lender, so read this as a lender’s survey.
- **Bank of America (June 2026).** 20% of prospective buyers and current homeowners used AI tools or chatbots for homebuying research in the past year, including 32% of Gen Z. Among AI users, estimating affordability, payments or closing costs led (57%).
- **Cotality (April 2026).** 80% of buyers surveyed assume lenders already use AI. Yet US trust in AI to help find a home fell to 16%, and 64% of buyers worry AI may repeat unverified information instead of relying on validated first-party data.

The surveys differ in who they asked, which explains the spread between 20% and 45%. They agree on the shape: buyers use AI to do the math and narrow options, and they still want people for the high-stakes steps. In Bank of America’s study, 54% preferred human expertise for legal or contractual advice.

The assistants are also becoming storefronts. Zillow launched an [app inside ChatGPT](https://www.housingwire.com/articles/zillow-chatgpt-launch-app-integration/) in October 2025 that guides users from listings back to Zillow, where they can explore financing with Zillow Home Loans. Redfin, now owned by Rocket Companies, [followed with its own ChatGPT app](https://www.housingwire.com/articles/redfin-launched-a-chatgpt-app-to-enable-conversational-home-searches-and-property-exploration-the-move-follows-similar-integrations-by-zillow-and-google-raising-questions-about-mls-data-licensing/). Both route home searchers toward businesses that also finance homes.

## What do borrowers ask AI on the way to a mortgage application?

Payment, eligibility, lender choice and trust questions. The borrower prompts below are our own examples of common questions, not prompts collected from real homebuyers.

| Stage | Illustrative prompt |
|---|---|
| Affordability | “How much house can I afford on $95,000 a year with $400 in car payments?” |
| Program fit | “FHA or conventional with 5% down and a 680 credit score: which costs less over seven years?” |
| Eligibility | “Which lenders do VA loans for a first-time buyer, and what fees should I expect?” |
| Lender choice | “Online lender or local credit union for a first mortgage?” |
| Comparison | “Lender A vs Lender B for a jumbo refinance” |
| Trust | “Is this mortgage company legit? What do borrowers complain about?” |

Each row is a different kind of lead. The affordability question is early and broad; the lender comparison and trust questions arrive close to an application. A reasonable expectation is that the later questions convert at much higher rates, so they deserve the most attention.

## How does an AI answer lead to a funded mortgage?

Through a short chain: the answer shapes the shortlist, and the shortlist decides who gets a quote request. Quotes then decide who funds.

1. **Payment math.** The buyer asks what they can afford. The assistant explains loan types and costs, sometimes citing lender or comparison pages.
2. **Shortlist.** The buyer asks which lenders fit. The assistant names a few, or points to comparison sites.
3. **Trust check.** The buyer asks about reviews, complaints or licensing.
4. **Quote request or form.** The buyer contacts one to three lenders, directly or through a marketplace.
5. **Application, lock and closing.** The lead becomes revenue only here.

Shopping is where the money moves. Freddie Mac found that in 2022, borrowers who applied with two lenders lowered their rate by an average of 20 basis points, and that two rate quotes could have saved as much as $600 a year. A lender left off the shortlist never gets the chance to compete on price. We suggest tracking AI visibility against funded loans and pull-through, not just form fills, and asking applicants where they first heard of you; analytics often miss AI-assisted visits, as explained in [why analytics miss AI visibility](https://underneath.agency/resources/why-analytics-miss-ai-visibility).

## What decides whether an assistant names a lender or marketplace?

In finance, sources that already rank, comparison sites and review platforms weigh heavily; platforms document little about choosing lenders.

**Documented by the platform.** A question such as “best FHA lender for first-time buyers” may, by Google’s own account, set off [“query fan-out”](https://developers.google.com/search/docs/appearance/ai-features) in AI Overviews and AI Mode: several related searches run before the answer is written. Because a mortgage can affect a household’s financial stability, Google’s guidance holds such topics to a higher bar, and its systems [weigh experience, expertise, authoritativeness and trust more heavily](https://developers.google.com/search/docs/fundamentals/creating-helpful-content) there. Neither Google nor OpenAI publishes how it picks one lender over another.

**Observed in our studies.**

- **Ranking pages feed mortgage answers.** Of the AI Overview citations for finance and insurance searches in [our citation study](https://underneath.agency/research/ai-overview-citations-study), 34.8% were page-one Google results; nerdwallet.com alone showed up in 10.0% of every AI Overview we sampled.
- **Comparison sites carry over into answers.** In [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study), when ChatGPT’s own search named a source such as NerdWallet, the answer cited it 44.0% of the time, against 8.1% when it did not.
- **Trust questions go to review platforms.** In [our reputation study](https://underneath.agency/research/is-it-legit-ai-reputation-study), 88.0% of answers to “is this brand legit?” cited a review or complaint platform.
- **Video explains the math.** In [our YouTube study](https://underneath.agency/research/ai-overview-youtube-videos-study), AI Overviews for financial services searches cited a YouTube video on 49.0% of searches, although Google showed one on page one for only 8.0%.

**Our inference.** For a lender, the facts that decide whether an assistant can recommend you are concrete: which loan programs you offer, in which states you are licensed, minimum credit and down payment requirements, how fast you close, and what borrowers say after closing. A lender whose program pages state these plainly gives an assistant something to repeat. One whose details sit behind a lead form gives it nothing, and the assistant may name a comparison site or a rival instead. Insurers face a similar state-by-state test, covered in [how insurance carriers win quotes from AI](https://underneath.agency/resources/insurance-companies-customers-ai-search).

## How do RESPA and TILA rules shape mortgage GEO?

They apply to your own content and referral arrangements; they make accuracy and neutral presentation the safest route to visibility.

**Advertising under Regulation Z.** The Truth in Lending Act’s [Regulation Z](https://www.ecfr.gov/api/renderer/v1/content/enhanced/current/title-12?part=1026&section=1026.24) says that if an advertisement states a rate, it must state it as an “annual percentage rate,” using that term. It lists “triggering terms,” such as a payment amount, that require further disclosures, and it bars misleading comparisons using rates or payments that apply for less than the full loan term. Rate tables, affordability calculators and “as low as” pages are exactly what an assistant may quote. We infer that a page written to these rules is also the page least likely to be misquoted.

**Referrals under RESPA section 8.** In a [2023 advisory opinion](https://www.govinfo.gov/content/pkg/FR-2023-02-13/pdf/2023-02910.pdf), the CFPB said a digital mortgage comparison-shopping platform violates section 8 if it gives enhanced placement or steers consumers to participants based on the payments it receives, rather than on neutral criteria. Federal guidance has since shifted: in [May 2025 the CFPB withdrew](https://www.govinfo.gov/content/pkg/FR-2025-05-12/html/2025-08286.htm) a long list of guidance documents, including its 2024 circular on preferencing and steering by digital intermediaries, and said the withdrawal “is not necessarily final.” The statute itself did not change, and state regulators enforce their own rules. Check the current status of each document with counsel.

For a marketplace, the practical point is that “best lender” rankings and comparison tables are both the content assistants like to cite and the content regulators have examined. Publishing the criteria behind any ranking, and keeping paid placement clearly separate, protects you on both fronts. For how ranked lists feed AI recommendations, see [which pages to target to show up in AI recommendations](https://underneath.agency/resources/best-of-lists-ai-recommendations).

## What does GEO work look like for a mortgage platform?

Generative engine optimization (GEO) makes your loan programs, costs and reputation easy for assistants to find and repeat accurately.

For a lender or marketplace, the work usually covers:

1. **Program pages in plain text.** One page per loan program (FHA, VA, conventional, jumbo, refinance) stating who qualifies, licensed states, typical timelines and documents needed, reviewed by compliance.
2. **Compliant rate and cost content.** APR shown wherever a rate appears, triggering-term disclosures in place, and a date on every rate example, so a quoted figure is not stale or misleading.
3. **Affordability explainers.** Clear, worked explanations of payments, closing costs and mortgage insurance, the questions buyers bring to AI first. Short explainer videos with accurate captions belong here too.
4. **Comparison and editorial coverage.** Accurate listings and reviews on the finance comparison and editorial sites assistants search and cite.
5. **Reputation work.** Monitor review and complaint platforms, answer complaints, and fix their causes; assistants summarize what those platforms say.
6. **Entity facts.** Consistent company name, licensing details, ownership and contact information across your site and directories, so assistants do not confuse you with a similarly named lender.
7. **Search foundations.** Because finance answers lean on page-one results, keep program and explainer pages ranking well in Google and Bing.

If an assistant misstates your loan limits, licensed states or program rules, [our guide to fixing wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) shows how to correct it. For marketplaces with two-sided dynamics, [how an online marketplace wins buyers and sellers when people ask AI](https://underneath.agency/resources/marketplace-buyers-sellers-ai-search) covers the supply side. No one can guarantee that an assistant will name a lender; GEO makes the evidence it finds accurate, current and verifiable.

## Where is the mortgage evidence still thin?

On conversion: surveys show borrowers use AI, but nothing public links AI visibility to applications or funded loans.

- **Surveys come from interested parties.** Veterans United and Bank of America are lenders, and Cotality sells data to the industry. Their samples and questions differ, which is why the AI-use figures range from 20% to 45%.
- **No public conversion data.** We found no public figures connecting AI answers to mortgage applications, lock rates or funded volume. The cross-industry evidence, thin as it is, is reviewed in [does AI visibility drive business results?](https://underneath.agency/resources/does-ai-visibility-drive-business-results).
- **Rules are in motion.** Federal guidance on comparison platforms changed in 2025 and may change again; how regulators view AI-generated rate quotes has not been tested in public.
- **Our studies are snapshots.** Our citation and frequency findings come from US searches in September 2026; answers change between assistants and over time.

## Where should a mortgage lender or marketplace start?

Start by asking assistants the questions your borrowers ask, then check every rate, program and licensing fact in the answers.

A useful first review covers affordability and program questions in your main states, lender-choice and comparison questions, and the reviews and “is it legit” questions about your company. It shows whether you are named, which lenders and comparison sites appear instead, and whether your programs, licensing and costs are described correctly.

If your growth depends on borrowers who apply and fund, [talk to us about a review of your mortgage platform in AI answers](https://underneath.agency/contact). We will map where assistants send borrowers in your markets, list the facts they get wrong or cannot find, and plan the program pages, coverage and reputation work, reviewed with your compliance team, that give you a fair chance at the shortlist. Lenders and marketplaces can see on our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page how that program and reputation work is run, with compliance review built into each step.

## Frequently asked questions

### Do home buyers really use ChatGPT to choose a mortgage lender?

Some do. In Veterans United’s June 2026 survey, 30% of prospective buyers used AI to shop for rates and 28% to compare lender reviews. Most AI use is still payment math and research.

### Can an AI assistant quote our mortgage rates?

It can repeat rates from pages it finds, including old ones. Date your rate examples, show the APR wherever a rate appears, and keep historical rate pages clearly labeled.

### Do mortgage comparison sites help or hurt a lender in AI answers?

Often they help. Finance answers cite comparison sites frequently, and in our hidden-searches study ChatGPT cited a named source 44.0% of the time. Accurate listings there matter.

### Does GEO replace paid mortgage leads?

No. It is a separate channel. Paid leads deliver contacts now; GEO improves how assistants describe you over time. Measure both against funded loans, not form fills.

## Sources

- Veterans United Home Loans (2026-07-02), [New Survey: More Homebuyers Turning to AI Tools in 2026](https://www.veteransunited.com/education/ai-homebuying-survey/)
- Bank of America (2026-06-23), [BofA Study: More Americans Favor Buying Over Renting for the First Time Since 2023](https://newsroom.bankofamerica.com/content/newsroom/press-releases/2026/06/bofa-study--more-americans-favor-buying-over-renting-for-the-fir.html)
- HousingWire (2026-04), [Homebuyers want AI and human in the loop: Cotality 2026 survey](https://housingwire.com/articles/homebuyers-want-ai-and-human-in-the-loop-cotality-2026-survey)
- HousingWire (2026-08), [IMB mortgage profits, Q2 2026](https://housingwire.com/articles/imb-mortgage-profits-q2-2026)
- Freddie Mac (2023-02-16), [When Rates Are Higher, Borrowers Who Shop Around Save More](https://www.freddiemac.com/research/insight/20230216-when-rates-are-higher-borrowers-who-shop-around-save)
- HousingWire (2025-10-06), [Zillow, ChatGPT launch app integration](https://www.housingwire.com/articles/zillow-chatgpt-launch-app-integration/)
- HousingWire (2026), [Redfin rolls out ChatGPT app for real estate searches](https://www.housingwire.com/articles/redfin-launched-a-chatgpt-app-to-enable-conversational-home-searches-and-property-exploration-the-move-follows-similar-integrations-by-zillow-and-google-raising-questions-about-mls-data-licensing/)
- eCFR (2026), [12 CFR 1026.24, Advertising (Regulation Z)](https://www.ecfr.gov/api/renderer/v1/content/enhanced/current/title-12?part=1026&section=1026.24)
- Federal Register, CFPB (2023-02-13), [Digital Mortgage Comparison-Shopping Platforms and Related Payments to Operators](https://www.govinfo.gov/content/pkg/FR-2023-02-13/pdf/2023-02910.pdf)
- Federal Register, CFPB (2025-05-12), [Interpretive Rules, Policy Statements, and Advisory Opinions; Withdrawal](https://www.govinfo.gov/content/pkg/FR-2025-05-12/html/2025-08286.htm)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- Google Search Central (2025), [Creating helpful, reliable, people-first content](https://developers.google.com/search/docs/fundamentals/creating-helpful-content)
- Underneath (2026), [When does Google show an AI Overview?](https://underneath.agency/research/ai-overviews-frequency-study)
- Underneath (2026), [Do AI Overviews cite the pages that rank?](https://underneath.agency/research/ai-overview-citations-study)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [“Is this brand legit?” How AI assistants build a reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study)
- Underneath (2026), [The YouTube videos Google’s AI cites are small](https://underneath.agency/research/ai-overview-youtube-videos-study)

---

This is the Markdown twin of https://underneath.agency/resources/mortgage-platforms-leads-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Which AI engine gives the most consistent answers about brands?"
description: "No AI engine is steadiest on every measure. Perplexity repeated brand lists most in two tests; Gemini and ChatGPT described brands most steadily in another."
canonical: "https://underneath.agency/resources/most-consistent-ai-engine-for-brands"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Which AI engine gives the most consistent answers about brands?

No single AI engine is the most consistent on every measure. In the two tests that compared which brands an engine names, Perplexity repeated itself most often; when researchers instead compared how engines describe a brand, Gemini and ChatGPT were steadier and Perplexity least steady. For anyone monitoring a brand, the engine you track shapes how reliable your numbers are, but so does the question you ask.

## The short version

1. In our test of 20 buyer questions asked five times each, 40.9% of the brands Perplexity named appeared in every run, against 25.2% for ChatGPT and 13.7% for Gemini.
2. A Swiss study of four engines found the same order for brand lists repeated within 24 hours: Perplexity overlapped most (0.492), Google’s AI Mode least (0.375).
3. On shopping questions asked from the Netherlands, Google’s AI Overviews repeated their product picks most (0.421 overlap) and ChatGPT least (0.178).
4. When a vendor compared the wording of repeated answers about 20 European brands, Gemini (0.952) and OpenAI’s model (0.950) were steadier than Perplexity (0.904).
5. In our test, the question explained more of the variation in consistency (30.0%) than the choice of assistant (25.5%).

## Which engine repeats the same brands most often?

Perplexity, in both independent tests that included it, with ChatGPT in the middle and Gemini or Google AI Mode last.

In [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), we asked ChatGPT, Gemini and Perplexity the same 20 US buyer questions five times each on one day in September 2026. Perplexity named 40.9% of its brands in all five runs. ChatGPT managed 25.2% and Gemini 13.7%. The first brand named changed at least once for 80.0% of questions on ChatGPT, 85.0% on Gemini and 40.0% on Perplexity.

[Schulte and colleagues](https://arxiv.org/abs/2604.07585) ran German-language prompts in four Swiss product categories up to ten times within a day, in March 2026. Measured as the share of brands two runs had in common, Perplexity scored 0.492, ChatGPT 0.437, Gemini 0.409 and Google AI Mode 0.375. One author also lists an industry affiliation, and all data came from Swiss servers.

| Measure of brand-list repeatability | Most repeatable | Least repeatable |
|---|---|---|
| Brands named in all 5 runs (our study, US) | Perplexity 40.9% | Gemini 13.7% |
| Brands shared between two runs (Swiss study) | Perplexity 0.492 | Google AI Mode 0.375 |
| Products shared between repeats (Dutch audit) | AI Overviews 0.421 | ChatGPT 0.178 |

## Does the ranking hold for product picks and cited sources?

Only partly: the order changes with the task, the country and whether you look at brands or sources.

[Uberti-Bona Marin and colleagues](https://arxiv.org/abs/2609.18729), independent academics, asked 117 real shopping questions three times each from the Netherlands in September 2026. They did not test Perplexity. Google’s AI Overviews repeated the most products between runs (0.421), Gemini 0.287 and ChatGPT only 0.178. The sources moved too: repeated answers shared 26.0% of cited websites on ChatGPT, 29.8% on Gemini and 45.9% on AI Overviews.

Sources and brands can also tell different stories. In the Swiss study, Gemini had the most stable sources, though not the most stable brands. In a June 2026 test of 15 prompts by [Tannenbaum](https://arxiv.org/abs/2609.22655), founder of a visibility software firm, 82.6% of ChatGPT’s cited pages turned over from one day to the next, against 45.5% on Perplexity. He notes that collection changes may explain part of that. [Sielinski](https://arxiv.org/abs/2603.08924), from another vendor, found OpenAI’s search often either repeated its sources exactly or changed them completely. For the reasons behind that churn, see [why cited sources change between checks](https://underneath.agency/resources/why-ai-search-citations-change).

## Which engine describes a brand most consistently?

Gemini and ChatGPT, in the one study that measured how similar the wording of repeated answers was.

[Żatuchin](https://arxiv.org/abs/2606.23165), who works for an AI brand-monitoring company, asked three engines about 20 Central and Eastern European brands, repeating each prompt five times in spring 2026. Repeated answers were very similar overall: average similarity was 0.935 on a scale up to 1, and 87.2% of comparisons scored above 0.90. By engine, Gemini (0.952) and OpenAI’s model (0.950) were steadier than Perplexity (0.904).

This does not contradict the brand-list results. Perplexity can name the same brands while phrasing its answer differently each time. Which engine looks “most consistent” depends on whether you track names or narrative.

## Does the engine matter more than the question or the language?

Not always: the question, and for some measures the language, matter as much as the engine.

In our study, the question accounted for 30.0% of the variation in stability and the assistant for 25.5%. A question that was stable on one assistant was not reliably stable on another. Żatuchin found little difference between languages in how steady the wording was: English scored 0.941 and Estonian 0.925.

A second paper by [Żatuchin](https://arxiv.org/abs/2607.13304), scoring how positive answers were about the same brands, found language mattered far more. Query language explained 26.5% of the variation in a single answer, the brand itself only 1.5%. That measure was weak, though: 91.9% of answers scored exactly neutral. Its practical point stands: a single answer says almost nothing about where a brand stands. The same caution applies to [a one-off AI visibility report](https://underneath.agency/resources/one-time-ai-visibility-report).

## Does a consistent engine give more trustworthy answers about you?

No: consistency only means the answer repeats, not that it is right or complete.

Perplexity’s steadiness shows the gap. In our study, 166 pairs of Perplexity runs cited exactly the same pages, yet the brand list still differed 91.6% of the time. Stable sources did not guarantee stable recommendations.

A vendor dataset from [Kumar at Ranqo](https://arxiv.org/abs/2606.20065), covering 102 brands, found brand mentions mostly fixed: only 6.8% of brand-question-engine combinations flipped between being named and not. But the tone of those mentions flipped 45.5% of the time. An engine can be reliable about whether it names you and unreliable about how it talks about you.

## What should you do about it?

Choose engines by where your buyers ask, then measure each one with enough repeats to see its real pattern.

1. Do not pick an engine to monitor because it looks stable. [Track the engines your customers use](https://underneath.agency/resources/is-tracking-chatgpt-enough), not just one.
2. Run each question several times per engine. One ChatGPT answer in our study showed only 57.8% of the brands its five answers named.
3. Track separately whether you are named, where you rank and how you are described. Each moves at a different rate.
4. Expect more noise from ChatGPT and Gemini brand lists, and set wider thresholds before reacting to a change there.
5. Test your questions in every language your buyers use, since language can matter as much as engine.

If you want this set up as an ongoing program, see [our generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

It does not tell us whether these rankings hold over weeks, in other categories or after engine updates.

- Every study is a snapshot of one to five days; engines change models often.
- Samples are small: 20 questions in our study, 15 prompts in Tannenbaum’s, 20 brands in Żatuchin’s.
- Three of the studies are by vendors that sell visibility measurement.
- Claude and Microsoft Copilot are missing from most comparisons.
- No study links an engine’s consistency to whether buyers trust or act on its answers.

## Frequently asked questions

### Is Perplexity more consistent than ChatGPT?

For brand lists, yes, in two tests. In ours, 40.9% of Perplexity’s brands appeared in all five runs against 25.2% for ChatGPT; for how answers are worded, one study found Perplexity less steady.

### Why does ChatGPT give different brand recommendations each time?

AI assistants generate each answer afresh and may search the web differently each run. In our test, ChatGPT’s first-named brand changed at least once for 80.0% of questions.

### How many times should we run a prompt to measure AI visibility?

More than once, and more on less stable engines. A single ChatGPT answer showed 57.8% of the brands its five runs named between them.

### Does the language of the question affect AI answers about brands?

It can matter more than the engine for some measures. In one study, query language explained 26.5% of the variation in tone of a single answer.

## Sources

- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)
- Schulte, Bleeker and Kaufmann (2026), [Don’t Measure Once: Measuring Visibility in AI Search (GEO)](https://arxiv.org/abs/2604.07585), arXiv:2604.07585.
- Uberti-Bona Marin, Bertaglia, Astante, Rijsbosch, van Dijck, Hannák, Spanakis and Kollnig (2026), ["If I Had to Buy Just ONE: Galaxy S26 Ultra": Auditing AI-Generated Product Recommendations](https://arxiv.org/abs/2609.18729), arXiv:2609.18729.
- Tannenbaum (2026), [Scoring With the Engine: Retrieval Exposure, Cross-Engine Divergence, and the Limits of Engine-Agnostic GEO Scores](https://arxiv.org/abs/2609.22655), arXiv:2609.22655.
- Sielinski (2026), [Quantifying Uncertainty in AI Visibility: A Statistical Framework for Generative Search Measurement](https://arxiv.org/abs/2603.08924), arXiv:2603.08924.
- Żatuchin (2026), [The Language Blind Spot: How Query Language and Brand Recognition Tier Shape AI-Constructed Brand Reputation Across Twelve European Languages](https://arxiv.org/abs/2606.23165), arXiv:2606.23165.
- Żatuchin (2026), [Where Does the Noise Come From? A Variance-Components Decomposition of Non-Determinism in LLM Brand Answers](https://arxiv.org/abs/2607.13304), arXiv:2607.13304.
- Kumar (2026), [Generative Engine Optimization at Scale: Measuring Brand Visibility Across AI Search Engines](https://arxiv.org/abs/2606.20065), arXiv:2606.20065.

---

This is the Markdown twin of https://underneath.agency/resources/most-consistent-ai-engine-for-brands. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How managed service providers can win clients from AI search"
description: "By being named when owners and IT leads describe their size, region and needs to AI, with pricing, coverage and security proof they can check."
canonical: "https://underneath.agency/resources/msps-customers-from-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How can a managed service provider win more clients from AI search?

By being the MSP an AI assistant names when a business owner or IT lead describes their size, location, industry and security needs, and by publishing the pricing model, coverage and proof that let them check the answer. A managed services contract is a long, recurring commitment, so buyers compare carefully and often switch only after a bad experience. The assistant increasingly shapes who makes that comparison, though no study has yet measured how many MSP contracts start there.

This guide covers recurring managed contracts: monthly support, monitoring, security and co-managed IT, priced per user, per device or by tier. One-off projects and break/fix work follow a different buying path. Advisory work such as an ERP choice is covered in how IT consulting firms win advisory clients.

## The short version

1. The market is large: Canalys estimates IT managed services will be worth [US$610 billion by the end of 2025](https://www.n-able.com/press/press-releases/n-ables-second-annual-msp-horizons-report-shows-significant-growth-opportunity-for-global-msps-with-cybersecurity-leading-the-way), with channel partners earning about 98% of it.
2. Winning clients is MSPs’ biggest worry: in [Kaseya’s 2025 State of the MSP survey](https://www.kaseya.com/press-release/kaseyas-state-of-msp-shows-what-makes-a-top-performing-msp/), new customer acquisition topped MSPs’ concerns (43%).
3. Buyers find MSPs through recommendations and search: [JumpCloud’s Q1 2025 survey](https://themspsummit.com/article/msps-gaining-steam-smes/) of 900 IT decision-makers found 45% found their MSP through recommendations and 37% through online search.
4. Most growing firms already use one: in [Barracuda’s 2025 survey](https://cxotoday.com/press-release/73-of-organizations-with-up-to-2000-employees-rely-on-msps-to-manage-the-security-challenges-of-growth-2/) of 2,000 decision-makers at firms with 50 to 2,000 employees, 73% already work with an MSP, and 45% would switch if theirs could not show round-the-clock security skills.
5. Company size changes AI answers: in [our phrasing study](https://underneath.agency/research/ai-prompt-phrasing-study), adding “I run a small business with about 10 employees” kept the original first brand only 40.6% of the time, against 68.0% for a plain rerun.

## Who buys managed services, and what is a client worth?

Owners of small businesses without IT staff, and mid-market IT leaders who want extra hands, mostly for security.

Three buyers stand out:

- **The small-business owner.** No IT employee, a growing headcount and a sense that IT is now risky. JumpCloud found 35% of SMEs (up to 2,500 employees) use MSPs to fully manage IT, up from 29% six months earlier.
- **The mid-market IT manager.** A small team that wants co-managed help with security, backup or after-hours cover. Kaseya found 83% of MSPs offer co-managed services.
- **The security-driven buyer.** Barracuda found 85% of organizations with 1,000 to 2,000 employees depend on MSPs for security support, compared with 61% of firms with 50 to 100 employees. MSP buyers are no longer only small businesses.

Outsourcing is common outside the US too. In the [UK government’s 2025 breaches survey](https://www.gov.uk/government/statistics/cyber-security-breaches-survey-2025/cyber-security-breaches-survey-2025), 44% of businesses had an external cyber security provider, rising to 62% of small and 68% of medium businesses.

A client’s value is the monthly fee times the seats, times the years, plus add-ons. Pricing models vary. In [Kaseya’s 2023 benchmark](https://www.kaseya.com/resource/msp-pricing-managed-it-services-pricing/) of 1,091 MSPs, 22% of respondents charged $50 to $100 per user per month, and per-device pricing had fallen to 13% of respondents from 17% in 2022. Security raises the value: Barracuda found customers prepared to pay MSPs up to 25% more for the services they need, and 92% willing to pay a premium for help integrating security tools.

## Why do businesses hire or switch MSPs now?

Because IT has become a security problem, small teams cannot cover it, and unhappy clients do switch.

**Effectiveness over savings.** In JumpCloud’s survey, 54% of IT admins said increased IT effectiveness was their top reason for using an MSP; 41% cited cost effectiveness, down from 58% two quarters earlier.

**Security demand.** [N-able’s 2025 report with Canalys](https://www.n-able.com/press/press-releases/n-ables-second-annual-msp-horizons-report-shows-significant-growth-opportunity-for-global-msps-with-cybersecurity-leading-the-way), based on 451 channel partners, found 90% expect growth in cybersecurity managed services sales. MSPs that resell staff training can see [how security awareness training vendors win buyers](https://underneath.agency/resources/security-awareness-training-customers-ai-search).

**Switching.** Barracuda found 45% of customers would switch if their MSP could not show the skills to deliver round-the-clock security support. An unhappy client of a competitor is often the best prospect an MSP has, and that client is asking someone, or something, who to call next.

**Growth.** Kaseya found 64% of MSPs reported revenue increases in the past year, so competition for each new client is real.

## Where does AI search sit in the MSP buying journey?

At the shortlist step, alongside the peer recommendation, before discovery calls and an IT assessment.

The path, as we understand it:

1. **Trigger.** A bad experience with the current provider, a security incident, an insurer’s questionnaire, a compliance audit or the departure of the one IT person.
2. **Shortlist.** Recommendations and search, per JumpCloud. AI assistants now sit inside that search for many buyers.
3. **Discovery calls and assessment.** Usually two or three MSPs, each reviewing the client’s environment.
4. **Proposal.** Seats, devices, service tier, security add-ons and response times.
5. **Contract and growth.** A multi-year agreement, then more seats and services.

Across B2B purchases, [Gartner’s 2026 survey](https://www.gartner.com/en/newsroom/press-releases/2026-05-20-gartner-survey-finds-sixty-nine-percent-of-b-two-b-buyers-turn-to-sales-reps-to-validate-ai-generated-insights) of 645 buyers found 45% used generative AI in a recent purchase, mainly to gather information on vendors, and 69% prefer to validate what AI tells them with sales reps. For an MSP, we infer the validation happens in the discovery call and assessment: the assistant suggests, the call confirms. Buyers facing a one-off decision such as an ERP choice often look for an adviser instead; see [how IT consulting firms win advisory clients](https://underneath.agency/resources/it-consulting-firms-clients-ai-search).

## Which questions do MSP buyers ask AI assistants?

Questions about size, location, industry, compliance, pricing and switching. We wrote the sample prompts below to sound like an office manager or IT lead shopping for support; none were captured from real users.

| Buyer | Illustrative prompt |
|---|---|
| Small professional firm | “Best managed IT provider for a 60-person accounting firm in Denver?” |
| Co-managed | “MSP that offers co-managed IT for a hospital with four IT staff?” |
| Pricing | “Per-user or per-device MSP pricing: what is fair for 40 users and 55 devices?” |
| Security | “Managed service provider with 24/7 security monitoring in Manchester?” |
| Compliance | “MSP that handles HIPAA for a three-location dental group?” |
| Switching | “Our MSP takes days to answer tickets. How do we switch, and to whom?” |
| Growth | “When should a 30-person company stop relying on one IT person and hire an MSP?” |

Each detail changes the answer. Our phrasing study shows company size alone moved the first brand. Location matters too: when [our country study](https://underneath.agency/research/ai-recommendations-by-country-study) added “in the United Kingdom” to a question, local-market brands went from 25.4% to 49.1% of the brands ChatGPT named. Regional MSPs compete inside these narrower answers, we infer, not in a national “best MSP” list.

## How does AI visibility turn into signed managed services contracts?

When an assistant matches your MSP to the buyer’s size, region and needs, the discovery call starts with a fit.

A qualified MSP lead has four facts settled: a seat count in your range, a location you serve, needs you cover (security, compliance, co-managed) and a reason to move now. A buyer who asked for an MSP with round-the-clock security for a 60-person firm in their city has filtered on most of them.

1. **AI answer.** A few named MSPs, with reasons drawn from their pages, reviews and rankings.
2. **Check.** Website, reviews, case studies, pricing model and service levels.
3. **Discovery call and assessment.** The buyer arrives with a size and a problem.
4. **Proposal and contract.** Seats times price, plus security and backup add-ons.
5. **Renewal and expansion.** Each year of good service compounds the value.

A vague or wrong answer works against you. An assistant that says you serve only enterprises, or only one city, sends away buyers who fit. Most of this happens before any visit you can track; see [what lost clicks mean for pipeline](https://underneath.agency/resources/ai-answers-pipeline-revenue).

## What decides whether an assistant names an MSP?

No platform documents how it picks MSPs; studies point to rankings, reviews and local listings.

**Documented by the platforms.** Google says AI Overviews and AI Mode [may use a “query fan-out” technique](https://developers.google.com/search/docs/appearance/ai-features), running several related searches for one question. No platform explains how an MSP gets named.

**Observed in our studies.** In [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study), ChatGPT searched for a named publication, ranking or award in 43.8% of its answers and for reviews in 46.2%. The MSP trade has its own rankings, such as Channel Futures’ MSP 501 list, which are the kind of named authority assistants look for. For local questions, [our study of ChatGPT’s local picks](https://underneath.agency/research/chatgpt-local-picks-google-profile-study) found businesses with more reviews than the local median were 19.5 points more likely to be listed at the same Google Maps rank; that study covered other local trades, not MSPs.

**What MSP buyers check.** Recommendations from peers, response times, security skills, compliance experience and whether the provider has served firms like theirs. Barracuda’s switching finding shows round-the-clock security is now a deciding factor.

**Our inference.** A reasonable expectation is that MSPs whose size range, regions, industries, pricing model and service levels are stated plainly, and confirmed by reviews and independent mentions, give assistants more to repeat accurately. No one has tested that for MSPs.

## What does it cost an MSP to be missing?

Lost discovery calls on contracts that would have paid every month for years, though no study has measured that loss.

- **Recurring revenue compounds.** A missed client is not one invoice; it is seats times months times years, plus add-ons.
- **Switchers decide fast.** A competitor’s unhappy client who asks an assistant and never hears your name is lost before your sales team knows they exist.
- **Partial presence is invisible.** In [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), only 25.2% of the brands ChatGPT named appeared in all five runs of the same question. An MSP named sometimes loses some buyers without knowing.
- **Wrong facts filter buyers out.** Old service areas, outdated pricing pages or a missing security offer send fits elsewhere. Our guide to [correcting wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) covers the repair work.

## How does GEO work for an MSP?

It makes your fit, pricing logic and proof easy to find and check, for buyers and assistants alike. Nobody can promise a recommendation.

1. **State who you serve.** Seat ranges, industries, regions and whether you offer fully managed or co-managed IT.
2. **Explain your pricing model.** Per user, per device or tiered, and what each tier includes, even without exact prices. Payroll providers, who also sell to small business owners, show why the page layout matters: assistants quoted prices less faithfully from pages with a monthly and annual toggle, as [how payroll providers get named by AI](https://underneath.agency/resources/payroll-software-ai-search) explains.
3. **Publish service levels.** Response and resolution times, after-hours cover and how you measure them.
4. **Show security proof.** Round-the-clock monitoring, partners, certifications and incident handling, since security drives switching. If you also offer testing, see [how penetration testing firms win scoped leads](https://underneath.agency/resources/pentest-firms-leads-ai-search).
5. **Make switching easy to understand.** An onboarding page that answers “how do we move from our current MSP?”
6. **Earn independent proof.** Reviews, industry rankings such as the MSP 501, local press and vendor partner listings. See [how small brands get recommended by AI](https://underneath.agency/resources/how-small-brands-get-recommended-by-ai), [which pages AI engines cite](https://underneath.agency/resources/best-of-lists-ai-recommendations) and, if you also sell security operations, [how MDR providers win leads](https://underneath.agency/resources/mdr-providers-qualified-leads-ai-search).
7. **Test real buyer wording.** Size, region, industry, compliance and switching prompts across ChatGPT, Google AI Overviews and AI Mode, Gemini, Perplexity, Copilot and Claude, several runs each. See [how question phrasing changes AI sources](https://underneath.agency/resources/does-question-phrasing-change-ai-sources).

## What can’t the research yet tell an MSP?

How many MSP buyers use AI assistants, and how often an AI mention becomes a signed contract. No study answers either.

- **Vendors fund most MSP data.** Kaseya, N-able, Barracuda and JumpCloud sell to MSPs or their clients; their surveys show direction, not neutral measurement.
- **Pricing data is dated.** Kaseya’s per-user figures come from its 2023 benchmark.
- **No survey isolates AI use by MSP buyers.** Gartner’s figures cover B2B buying across industries, and JumpCloud’s “online search” does not separate AI assistants from search engines.
- **The revenue link is untested.** In [Martinez’s](https://arxiv.org/abs/2607.14035) review of research on AI search optimization, traffic and conversions had the weakest evidence.

## Where should an MSP start?

With the questions your best-fit clients asked before they signed, including their size, city and reason for switching.

Write 20 to 30 of them across your buyer types: small firms with no IT, co-managed mid-market teams, regulated industries, switchers. Ask each major assistant several times. Record which MSPs are named, which rankings and review sites are cited, and whether your size range, regions, services and pricing model are described correctly.

If you want an independent read on your MSP, [get in touch with our team](https://underneath.agency/contact). We will show which buyer questions name your MSP, which send switchers and new buyers to competitors, and which gaps in your public proof are most likely costing you discovery calls and signed monthly contracts. The work after that review, from stating seat ranges and service levels to retesting switcher questions, is outlined on our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page.

## Frequently asked questions

### Do businesses use AI assistants to find an MSP?

No study isolates them. JumpCloud found 37% of SMEs found their MSP through online search, which increasingly includes AI answers, and 45% through recommendations.

### Should MSPs publish prices?

Publish the model at least. Buyers ask about per-user and per-device pricing, and assistants answer with or without your input.

### Does company size in the question matter?

Yes. Adding a 10-employee business context kept the original first brand only 40.6% of the time in our phrasing study.

### Do MSP rankings like the MSP 501 matter for AI?

Probably. ChatGPT searched for a named ranking or award in 43.8% of its answers in our hidden-searches study.

### Why do clients switch MSPs?

Security is a leading reason: 45% in Barracuda’s survey would switch if their MSP could not show round-the-clock security skills.

## Sources

- N-able with Canalys (2025-03-04), [N-able’s second annual MSP Horizons report](https://www.n-able.com/press/press-releases/n-ables-second-annual-msp-horizons-report-shows-significant-growth-opportunity-for-global-msps-with-cybersecurity-leading-the-way)
- Kaseya (2025-01-07), [Kaseya’s State of MSP shows what makes a top-performing MSP](https://www.kaseya.com/press-release/kaseyas-state-of-msp-shows-what-makes-a-top-performing-msp/)
- Kaseya (2023 survey data), [MSP pricing: a comprehensive guide to managed IT services pricing models](https://www.kaseya.com/resource/msp-pricing-managed-it-services-pricing/)
- JumpCloud, via The MSP Summit (2025), [MSPs gaining steam with SMEs](https://themspsummit.com/article/msps-gaining-steam-smes/)
- Barracuda Networks, via CXOToday (2025), [73% of organizations with up to 2,000 employees rely on MSPs](https://cxotoday.com/press-release/73-of-organizations-with-up-to-2000-employees-rely-on-msps-to-manage-the-security-challenges-of-growth-2/)
- UK Department for Science, Innovation and Technology (2025-04-09), [Cyber security breaches survey 2025](https://www.gov.uk/government/statistics/cyber-security-breaches-survey-2025/cyber-security-breaches-survey-2025)
- Gartner (2026-05-20), [Gartner survey finds 69% of B2B buyers turn to sales reps to validate AI-generated insights](https://www.gartner.com/en/newsroom/press-releases/2026-05-20-gartner-survey-finds-sixty-nine-percent-of-b-two-b-buyers-turn-to-sales-reps-to-validate-ai-generated-insights)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- Martinez (2026), [Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026)](https://arxiv.org/abs/2607.14035), arXiv:2607.14035.
- Underneath (2026), [Does rewording a question change AI brand recommendations?](https://underneath.agency/research/ai-prompt-phrasing-study)
- Underneath (2026), [Same question, four countries: do AI recommendations change?](https://underneath.agency/research/ai-recommendations-by-country-study)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [Which Google Maps businesses does ChatGPT recommend?](https://underneath.agency/research/chatgpt-local-picks-google-profile-study)
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)

---

This is the Markdown twin of https://underneath.agency/resources/msps-customers-from-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How neobanks turn AI answers into primary accounts"
description: "By answering the safety and fee questions AI users ask about app-only banks, so the account they open becomes the one their paycheck goes to."
canonical: "https://underneath.agency/resources/neobanks-account-openings-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How does a neobank turn AI answers into accounts that receive a paycheck?

By being named when young adults and gig workers ask AI which app to bank with, and by answering, plainly and checkably, the question every neobank faces next: is my money safe in an app that is not a bank? Neobanks already open accounts cheaply; the value arrives only when a customer moves their pay in. In AI answers, that step depends on partner-bank, deposit insurance, fee and reputation facts that the assistant can find and trust.

## The short version

1. Openings are cheap, primacy is not: [Dave](https://quartr.com/events/dave-inc-dave-q2-2026_oqMgziG5) added 951,000 new members in the second quarter of 2026 at a customer acquisition cost of $19, and [Chime](https://stocks.observer-reporter.com/observerreporter/article/bizwire-2026-8-5-chime-reports-second-quarter-2026-financial-results) now builds its premium tier around members with qualifying direct deposits of $3,000 or more a month.
2. Neobanks trail on satisfaction: in [J.D. Power’s 2026 study](https://www.01net.it/online-only-banking-providers-continue-to-win-over-customers-with-personalized-digital-experiences-but-some-struggle-with-customer-service-quality-jd-power-finds/), neobank checking scored 622 out of 1,000, against 674 for chartered online banks, with gaps driven by card and fraud problems and weaker support.
3. Trust was shaken: after the middleware firm Synapse failed, [more than 100,000 customers](https://www.bankingdive.com/news/synapse-trustee-85-million-shortfall-customer-funds/718796/) of its partner banks were locked out, and the [CFPB](https://www.crowdfundinsider.com/?p=256754) set aside $46.2 million to reimburse people affected.
4. The core customers ask AI: in [STRAT7’s April 2026 UK study](https://strat7.com/?p=48569), 45% of Millennials and Gen Z take most of their financial questions to AI, and AI finance users were more likely to bank with a neobank.
5. Gig workers are a real segment: the [Federal Reserve](https://www.federalreserve.gov/publications/files/2024-report-economic-well-being-us-households-202505.pdf) found 9 percent of US adults made money by doing short-term tasks such as giving rides, delivering takeout or doing odd jobs.

This guide is about how app-only banking brands are found and described in AI answers. It is not financial, legal or regulatory advice. Deposit insurance, fee and advance wording should be checked with your compliance team and your partner banks.

## Who chooses a neobank, and when does a member start to pay?

Younger adults, gig workers and people short between paydays choose them; members pay off once their pay arrives there.

A neobank, as we use the word, is an app-based banking brand without its own charter that offers accounts through partner banks. J.D. Power uses the same split: neobanks “do not have federal bank charters, but partner with federally chartered banks.” Chartered online banks are a different case; their economics run on deposits rather than paychecks. That case is covered in [how digital banks win depositors through AI](https://underneath.agency/resources/digital-banks-depositors-ai-search).

The customers have specific needs, and neobank marketing reflects them: no monthly fees, early access to direct deposits and small cash advances. That fits people with uneven income. Besides the 9 percent doing short-term gig tasks, the Federal Reserve found six percent of adults were unbanked in 2024. Earned wage access, which lets workers draw pay they have already earned, is a large market in its own right: a [Congressional Research Service report](https://www.everycrsreport.com/files/2024-07-31_IF12727_016d4fdcf42f18b428fa53fa04aa4279c6e61b87.html) cites a CFPB estimate that more than 10 million workers used such products in 2022, totaling $32 billion.

The economics turn on primacy:

| Neobank | Latest public figures | Where value comes from |
|---|---|---|
| Dave | 951,000 new members in the quarter, cost to acquire $19; 15.2 million members; quarterly revenue of $170.8 million | ExtraCash advances and subscriptions once members use the app regularly |
| Chime | 1.7 million net new active members over the past year; a premium tier for members depositing at least $3,000 a month | Card spending and liquidity products once pay is deposited; its fastest-growing segment is members earning $75,000 or more |

Chime also sells through employers: two new employer partners in the quarter together employ more than 350,000 people in the US. For a neobank, then, an AI answer is worth something only if it leads to an account that receives a paycheck. That is a narrower path than the general consumer finance app journey in [our guide to consumer fintech apps](https://underneath.agency/resources/consumer-fintech-apps-customers-ai-search).

## Where does AI enter a neobank customer’s decision?

Early, and among exactly the age groups neobanks serve.

- **Age.** In a J.D. Power survey [reported by the ABA](https://bankingjournal.aba.com/2025/09/survey-consumers-increasingly-turn-to-ai-for-financial-advice/), ChatGPT was the most popular AI tool for financial questions, particularly with respondents under 40.
- **Habit.** In STRAT7’s UK research, 45% of Millennials and Gen Z take most of their financial questions to an AI platform; among AI finance users, 35% are young families, against 7% of non-users.
- **Overlap.** STRAT7 found the people who use AI for money questions are more likely to bank with a neobank, invest in crypto and own stocks.

The UK figures are a signal, not a US measurement. A reasonable expectation is that neobank customers, younger and phone-first, meet AI answers at the moments that matter: comparing apps, checking an advance limit, or asking whether an app is safe after a bad headline.

Openings are already shifting toward digital players. According to a Cornerstone Advisors report [summarized by eMarketer](https://www.emarketer.com/content/fintechs-neobanks-outpace-big-incumbents-us-checking-account-openings), challengers and fintechs captured 47% of new checking accounts opened in the first half of 2023. That figure is disputed. An analysis in [Payments in Full](https://paymentsinfull.substack.com/p/are-fintechs-opening-44-of-new-checking) argues those counts include spending accounts that are not primary, citing a payroll survey in which only 1.7% of respondents sent direct deposits to payment apps such as PayPal and Cash App. Both views point to the same lesson: opening an account and becoming someone’s bank are different wins.

## What do people ask AI before choosing a neobank?

Questions about getting paid early, advances, fees, safety and alternatives. We wrote these examples to illustrate; they are not observed prompts.

| Need | Illustrative prompt |
|---|---|
| Cash gap | “Apps like Dave that give a cash advance without a credit check” |
| Gig income | “Best bank account for DoorDash and Uber drivers who get paid weekly” |
| Early pay | “Which banking apps let me get my paycheck two days early?” |
| Compatibility | “Cash advance apps that work with Chime” |
| Safety | “Is Chime a real bank? Is my money FDIC insured?” |
| After a scare | “What happened to Synapse customers, and could it happen with my app?” |
| Cost | “What fees do instant transfers and advances really cost on these apps?” |

The safety and cost prompts matter most for neobanks, because they test facts that are easy to get wrong: which bank holds the money, what pass-through insurance requires, and how advance fees work. According to the CRS report, the CFPB found 82% of employer-partnered earned wage access transactions had fees, with employer-integrated fees averaging $2.60 per transaction.

## How does an AI answer become a primary account?

Through five steps, and only the last two pay: named, checked, opened, paid in, used.

1. **Named.** A person asks for options and the answer lists a few apps.
2. **Checked.** They ask a follow-up about safety, fees or complaints. In [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study), ChatGPT looked for reviews in 46.2% of its answers.
3. **Opened.** They download the app and open an account, the step Dave can buy for $19.
4. **Paid in.** They switch their direct deposit, or take a first advance.
5. **Used.** Card spending, advances and premium tiers follow, which is where revenue builds.

An inaccurate or hedged answer at step two stops the chain before the cheap step even happens. We suggest measuring AI visibility against accounts that receive a first direct deposit, not installs, and asking new members where they first heard of the app. How to choose the right measures is in [what to measure in AI visibility](https://underneath.agency/resources/what-to-measure-ai-visibility).

## Why is “is my money safe?” the hardest question for a neobank?

Because the honest answer has conditions, and the industry’s worst failure is easy to find.

**The regulator’s view.** The FDIC wrote in its [2024 rule on deposit insurance misrepresentation](https://www.federalregister.gov/documents/2024/01/18/2023-28629/fdic-official-signs-and-advertising-requirements-false-advertising-misrepresentation-of-insured) that growth in fintech companies “has also blurred the distinction” between banks and non-banks “in the eyes of many consumers.” Chime’s own disclosure shows what a clear answer looks like: it says Chime is not FDIC-insured, names its partner banks, and states that “certain conditions must be satisfied for pass-through deposit insurance coverage to apply.”

**The failure people ask about.** When Synapse, a middleware company connecting fintech apps to banks, collapsed in 2024, its trustee found users were owed $265 million while partner banks held roughly $180 million for them. In November 2025 the CFPB approved $46.2 million from its Civil Penalty Fund to reimburse affected customers, about half the projected shortfall. Any neobank built on partner banks, we infer, will be asked by customers and assistants whether the same could happen to it. Vendors that sell the technology behind such partnerships face a different buyer, covered in [how banking technology vendors reach a bank’s shortlist](https://underneath.agency/resources/banking-technology-enterprise-deals-ai-search).

**What AI does with reputation.** In [our reputation study](https://underneath.agency/research/is-it-legit-ai-reputation-study), every complete answer to “is this brand legit?” called the brand legitimate, yet 99.7% made at least one negative claim. 88.0% of answers cited a review or complaint platform, and Trustpilot and the BBB accounted for 61.7% of review-platform citations. For a neobank, unresolved complaints about frozen accounts, card disputes or advance fees are, we infer, likely to surface in exactly these answers. J.D. Power’s finding that neobanks lag on debit card and fraud problems and on phone and chat support shows where those complaints come from.

## What decides whether a neobank is named?

Independent evidence and clear facts; platforms document how they search, not how they choose apps.

**Documented.** Google says AI Overviews and AI Mode [may use “query fan-out”](https://developers.google.com/search/docs/appearance/ai-features), issuing several related searches before answering. No assistant publishes how it ranks banking apps.

**Observed.** In our hidden-searches study, ChatGPT ran a search aimed at a named publication, ranking or award in 43.8% of its answers, and some of those looked for the latest edition of an annual ranking such as J.D. Power’s. In J.D. Power’s 2026 study, Chime ranked third among high-yield savings providers with a score of 714. The study also found a gap of 225 points between the best online banks and the lowest-performing neobanks in checking. Rankings like these are, we infer, part of what assistants read when they compare apps.

**Our inference.** Assistants can only repeat facts a neobank states in public and others confirm: the partner banks and what they hold, the insurance conditions, every fee and advance limit in text, and how complaints are handled. Regulation adds a wrinkle for advance products. On December 23, 2025, the CFPB issued an [advisory opinion](https://www.mofo.com/resources/insights/260107-cfpb-reestablishes-position-that-certain-earned-wage) that certain earned wage access programs repaid through payroll deduction are not credit under federal lending law; products outside those conditions remain, in the words of one law firm, “in a regulatory gray area.” How a product is described, then, is a compliance decision first and an AI visibility decision second.

## What does GEO look like for a neobank?

Generative engine optimization (GEO) makes your partner banks, protections, fees and reputation easy for assistants to find, check and repeat.

For a neobank, the work usually covers:

1. **A partner-bank and protection page.** Who holds the money, how pass-through insurance works and its conditions, what happens to funds if a partner fails, written with your partner banks and compliance team.
2. **Fee and advance pages in text.** Every fee, tip, transfer charge and advance limit, with plain examples, so cost comparisons in answers are right.
3. **Reputation work.** Monitor Trustpilot, the BBB and app store reviews, resolve complaints, and fix their causes. Whether assistants weigh reviews the way customers do is covered in [should reputation priorities differ for AI assistants?](https://underneath.agency/resources/ai-vs-human-reputation-priorities).
4. **Segment pages.** Honest pages for gig workers, first jobs and people between paydays that explain fit without promising outcomes.
5. **Independent proof.** Satisfaction rankings, press coverage and employer partnerships that a third party can confirm.
6. **Monitoring after every headline.** When a fintech failure or enforcement story breaks, ask assistants about your app and correct wrong or blended facts, following [how to fix wrong information about your brand in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers).

No one can guarantee that an assistant will recommend an app. GEO makes sure the evidence it finds about yours is complete and accurate. Early-stage brands facing established names can also read [how a fintech startup can get recommended by AI](https://underneath.agency/resources/fintech-startups-customers-ai-search).

## What remains uncertain for neobanks in AI search?

Much: there is no public link from AI answers to funded accounts, and US survey data is thin.

- **No US neobank-specific AI survey.** The strongest link between AI use and neobank customers comes from the UK.
- **Opening data is contested.** Account-opening shares differ by source and definition, as the Payments in Full critique shows.
- **No outcome data.** We found no public data connecting AI answers to direct deposits or advances. For the broader evidence beyond banking, read [does AI visibility drive business results?](https://underneath.agency/resources/does-ai-visibility-drive-business-results).
- **Our studies are snapshots.** They cover US answers collected in September 2026; assistants and their answers change.

## What should a neobank do first?

Ask assistants the questions your members ask, especially about safety and fees, and check every answer.

A useful first review covers “apps like” and comparison questions in your category, the safety questions about your partner banks, fee and advance questions, and how your app is described after the latest industry headline. It shows whether you are named, which rivals and review sites appear instead, and whether your protections and costs are stated correctly.

If your growth depends on members moving their paycheck to your app, [talk to us about how AI answers describe your neobank](https://underneath.agency/contact). We will compare how assistants present you and your competitors, list the safety, fee and reputation facts they miss or get wrong, and plan the content, coverage and review work, checked with your compliance team and partner banks, to correct them. How the partner-bank, fee and review work is carried out, and checked again after each industry headline, is explained on our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page.

## Frequently asked questions

### Do AI assistants treat neobanks differently from banks?

No assistant documents a different rule. In practice, safety questions about neobanks have conditional answers, so clear partner-bank and insurance facts matter more.

### Why would an AI answer bring up Synapse when people ask about our app?

Because it is the best-known failure of the partner-bank model: users were owed $265 million against roughly $180 million held. State plainly how your arrangement protects customers.

### Can a neobank with low satisfaction scores still be named by AI?

It can be named, but answers usually include negatives. In our reputation study, 99.7% of answers made at least one negative claim, often drawn from review platforms.

### Should we describe our cash advance as credit or not?

That is a legal question for your counsel. The CFPB’s December 2025 opinion covers only certain payroll-deduction products, so describe yours accurately and consistently everywhere.

## Sources

- Quartr (2026-08), [Dave (DAVE) Q2 2026 earnings summary](https://quartr.com/events/dave-inc-dave-q2-2026_oqMgziG5)
- Panabee (2026-08), [Dave Q2 2026 revenue grows 30% y/y amidst regulatory and accounting headwinds](https://www.panabee.com/news/dave-earnings-q2-2026)
- Chime, via Business Wire (2026-08-05), [Chime reports second quarter 2026 financial results](https://stocks.observer-reporter.com/observerreporter/article/bizwire-2026-8-5-chime-reports-second-quarter-2026-financial-results)
- J.D. Power, via 01net (2026-04-29), [Online-only banking providers continue to win over customers with personalized digital experiences](https://www.01net.it/online-only-banking-providers-continue-to-win-over-customers-with-personalized-digital-experiences-but-some-struggle-with-customer-service-quality-jd-power-finds/)
- Banking Dive (2024-06), [Synapse trustee McWilliams finds $85M gap in frozen funds](https://www.bankingdive.com/news/synapse-trustee-85-million-shortfall-customer-funds/718796/)
- Crowdfund Insider (2025-12), [CFPB allocates $46M to victims of Synapse fintech collapse](https://www.crowdfundinsider.com/?p=256754)
- STRAT7 (2026), [Who’s using AI for financial advice? New UK research](https://strat7.com/?p=48569)
- Board of Governors of the Federal Reserve System (2025-05), [Economic Well-Being of U.S. Households in 2024](https://www.federalreserve.gov/publications/files/2024-report-economic-well-being-us-households-202505.pdf)
- Congressional Research Service (2024-07-31), [Earned Wage Access](https://www.everycrsreport.com/files/2024-07-31_IF12727_016d4fdcf42f18b428fa53fa04aa4279c6e61b87.html)
- ABA Banking Journal (2025-09), [Survey: Consumers increasingly turn to AI for financial advice](https://bankingjournal.aba.com/2025/09/survey-consumers-increasingly-turn-to-ai-for-financial-advice/)
- eMarketer (2023), [Fintechs and neobanks outpace big incumbents in US checking account openings](https://www.emarketer.com/content/fintechs-neobanks-outpace-big-incumbents-us-checking-account-openings)
- Payments in Full (n.d.), [Are fintechs opening 44% of new checking accounts?](https://paymentsinfull.substack.com/p/are-fintechs-opening-44-of-new-checking)
- Federal Register, FDIC (2024-01-18), [FDIC Official Signs and Advertising Requirements, False Advertising, Misrepresentation of Insured Status](https://www.federalregister.gov/documents/2024/01/18/2023-28629/fdic-official-signs-and-advertising-requirements-false-advertising-misrepresentation-of-insured)
- Morrison Foerster (2026-01-07), [CFPB reestablishes position that certain earned wage access programs are not “credit”](https://www.mofo.com/resources/insights/260107-cfpb-reestablishes-position-that-certain-earned-wage)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [“Is this brand legit?” How AI assistants build a reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study)

---

This is the Markdown twin of https://underneath.agency/resources/neobanks-account-openings-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How network security vendors win buyers who ask AI"
description: "By being named, in NIST and CISA terms, in AI answers about firewall refreshes, SASE and zero trust access, with a public security record buyers can check."
canonical: "https://underneath.agency/resources/network-security-zero-trust-customers-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How do network security vendors win buyers who ask AI about zero trust?

By being named in the AI answers that turn “we need zero trust” into a shortlist, in the framework language buyers already use, and by keeping a public security record that survives checking. Network security is in the middle of a large replacement cycle: firewalls are being refreshed, and remote access is moving to zero trust. The questions buyers ask during that move are architectural, and the answers decide who gets the evaluation.

## The short version

1. The core market is still growing: [Gartner forecasts](https://softwarestrategiesblog.com/2026/04/01/top-10-fastest-growing-security-categories-gartner-2026-forecast/), as summarized by Louis Columbus, put firewall equipment at $16.8 billion of spending in 2025, growing 15.9% in 2026.
2. Zero trust access is replacing older tools: the same forecast has zero trust network access (ZTNA) growing 23.0% in 2026, while intrusion prevention falls 6.3% and network access control 7.7%.
3. [Dell’Oro Group](https://www.delloro.com/news/sase-revenue-grows-above-20-percent-as-sd-wan-reaccelerates-in-2q-2026/) counted $3.5 billion of secure access service edge (SASE) revenue in the second quarter of 2026, its fifth straight quarter of growth above 20%, and calls the single-vendor versus multi-vendor question “contested.”
4. Buyers have a public playbook: [NIST’s 2025 guidance](https://www.nist.gov/news-events/news/2025/06/nist-offers-19-ways-build-zero-trust-architectures) offers 19 example zero trust architectures built with commercial products from 24 industry collaborators, and [CISA’s maturity model](https://www.cisa.gov/zero-trust-maturity-model) sets out five pillars.
5. The perimeter itself is under attack: [Verizon’s 2025 breach report](https://verizon.com/about/news/2025-data-breach-investigations-report) found exploitation of vulnerabilities rose 34% as an initial attack vector, with a focus on perimeter devices and VPNs.

## Who signs off on a network security platform, and how big is the deal?

A network and security team buys it together, for the whole organization, usually as a multi-year platform decision.

Network security covers the controls between users, devices, applications and data: firewalls, secure web gateways, ZTNA, SD-WAN, and the cloud-delivered bundles sold as SASE or security service edge (SSE). Unlike many security tools, it sits in the path of every connection, so a bad choice breaks the business, not only a dashboard. That makes buyers careful and decisions slow.

The buying group is two teams with different priorities. Networking cares about performance, branch connectivity and operations. Security cares about policy, inspection and risk. Dell’Oro notes that SASE growth now comes from both: branch network refreshes and security services expanding into data protection and AI governance. A vendor has to win both teams.

Customers are large. Netskope said at its 2025 [initial public offering](https://www.netskope.com/press-releases/netskope-announces-pricing-of-initial-public-offering) that its customers include more than 30% of the Fortune 100. Platform deals bundle firewall, access and inspection, so one decision can cover every office and remote worker for years.

Budgets are being reshuffled rather than simply increased. In the Gartner forecast summarized by Columbus, overall security spending growth for 2026 is 12.2%, and the firewall equipment line alone is projected to reach $26.7 billion by 2030. ZTNA is forecast to grow from a $2.4 billion base to $6.4 billion by 2030, while older categories shrink. The money follows the move to zero trust.

## At what point do network security buyers turn to AI?

At the research stage, alongside sales calls, peers and frameworks; no public study isolates network security buyers.

The closest evidence covers business buyers generally. In a [Gartner survey](https://www.gartner.com/en/newsroom/press-releases/2026-05-20-gartner-survey-finds-sixty-nine-percent-of-b-two-b-buyers-turn-to-sales-reps-to-validate-ai-generated-insights) of 645 B2B buyers in 2025, 45% had used generative AI in a recent purchase, mainly to gather information on vendors and products. Gartner’s headline finding was that 69% turned to sales reps to validate AI-generated insights, and a slight majority said they were more likely to meet misleading information from generative AI than from a rep.

That pattern fits network security closely, we infer. A buyer can ask an assistant to explain SASE or compare ZTNA approaches in minutes, but no one replaces the firewalls at 400 branches on an AI answer alone. The answer shapes which vendors get the call; the proof of concept decides who wins.

Google’s AI features are also in the path, and Google says AI Overviews and AI Mode [may use a “query fan-out” technique](https://developers.google.com/search/docs/appearance/ai-features), turning one zero trust question into several related searches across subtopics. A question about zero trust for a hospital network, we infer, fans out into framework, vendor, compliance and review searches at once.

## What do firewall and SASE buyers type into AI assistants?

Architecture, framework, migration, comparison and vendor-risk questions. Each prompt below is our own illustration of how a CISO or network architect might phrase the question; none were collected from real buyers.

| Stage | Illustrative prompt |
|---|---|
| Strategy | “How do we move from VPN to zero trust access for 8,000 employees without breaking legacy apps?” |
| Framework | “Which vendors map to the CISA Zero Trust Maturity Model network pillar at the advanced level?” |
| Architecture | “Single-vendor SASE or best-of-breed SD-WAN plus SSE for a manufacturer with 120 sites?” |
| Refresh | “Our firewalls reach end of support next year: should we replace them or move to firewall as a service?” |
| Comparison | “Zscaler vs Netskope vs Palo Alto Prisma Access for a company on Microsoft 365” |
| Vendor risk | “Which firewall vendors have had actively exploited vulnerabilities in the last two years?” |
| Compliance | “Which ZTNA products are FedRAMP authorized for a federal contractor?” |

The framework questions matter because buyers borrow the government’s vocabulary. NIST’s Zero Trust Architecture (SP 800-207) describes the concept, and its 2025 implementation guide (SP 1800-35) shows 19 worked examples. CISA’s model gives five pillars, three cross-cutting capabilities and four maturity stages from traditional to optimal. A vendor whose public material maps its products to those terms gives an assistant an easy way to connect it to the question, we infer.

The vendor-risk questions are real, too. In September 2025 CISA issued [Emergency Directive 25-03](https://www.cisa.gov/news-events/directives/ed-25-03-identify-and-mitigate-potential-compromise-cisco-devices) ordering federal agencies to identify and mitigate potential compromise of certain Cisco firewall products, and updated it in April 2026. Every network security vendor faces questions about its own record, and buyers increasingly ask them of AI.

## How does a firewall or SASE vendor get from an AI mention to a signed platform deal?

Through the architecture decision: an AI answer frames the approach and names vendors, then a pilot decides.

1. A CISO, network architect or IT director asks how to replace VPNs, refresh firewalls or meet a zero trust goal.
2. The answer describes an approach (single-vendor SASE, SSE plus existing SD-WAN, firewall as a service) and names vendors that fit it.
3. The team checks those vendors against frameworks, analyst views, peers and vulnerability history.
4. Two or three vendors run a pilot on real users and sites, often through a reseller or managed service provider.
5. The winner signs a multi-year platform agreement that expands as sites and users move over.

Step 2 is where AI answers matter most, because the approach chosen narrows the vendor field before any vendor is called. Consolidation makes the stakes higher. In a [2022 Gartner survey](https://www.gartner.com/en/newsroom/press-releases/2022-09-12-gartner-survey-shows-seventy-five-percent-of-organizations-are-pursuing-security-vendor-consolidation-in-2022), 75% of organizations were pursuing security vendor consolidation, up from 29% in 2020, and SASE was named as a main area for it. When buyers consolidate, the vendor framed as the platform wins several product lines at once.

Tracing a SASE or firewall deal back to an AI answer is hard. The answer is rarely a click; it is a name on a whiteboard months before a reseller files the deal. How that hidden influence shows up later is the subject of [what lost clicks mean for pipeline](https://underneath.agency/resources/ai-answers-pipeline-revenue).

## Why does an assistant name one SASE or firewall vendor over another?

No platform explains its vendor choices; studies point to named sources and fresh pages, and security buyers check everything.

**Documented by the platforms.** When a buyer asks about ZTNA or a firewall refresh, Google’s and OpenAI’s AI answers run their own web searches and link the pages they used, as both companies document. Neither company says how a firewall or SASE vendor ends up in the answer.

**Observed in studies.** In [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study), when an assistant’s search named a specific source, the answer cited that source 44.0% of the time, against 8.1% when it did not. In security, named sources would include analyst reports, independent test labs and government guidance, we infer. When [our freshness study](https://underneath.agency/research/ai-source-freshness-study) dated the pages each assistant cited, those under 90 days old were 17.4% to 22.6% of the total, against 6.9% of Google’s top 10. In a field where vulnerabilities and products change monthly, stale pages are a real weakness. Practitioner communities matter as well: [our Reddit study](https://underneath.agency/research/ai-reddit-citations-study) found r/cybersecurity among the communities Google’s AI cited most.

**What network security buyers check.** Framework alignment, independent test results, certifications such as FedRAMP, vulnerability and patch history, and references from similar organizations. Verizon’s finding that attackers increasingly exploit perimeter devices and VPNs means a vendor’s own security record is part of the evaluation.

**Our inference.** The evidence that convinces a network architect is mostly public: NIST and CISA mappings, test reports, advisories, certifications and peer reviews. Those are the same pages an assistant can find. A reasonable expectation is that vendors whose proof is public, current and framed in the buyer’s framework language are named more often for the right questions. Nobody has yet tested that idea on firewall, ZTNA or SASE vendors.

## What happens to a firewall or SASE vendor that AI answers leave out?

Missed platform decisions that last years, though no study has measured the loss directly.

- **Refresh cycles are long.** A vendor left off the shortlist at a firewall refresh may wait for the next refresh, years later, we infer.
- **Consolidation multiplies the miss.** If the winning platform covers firewall, ZTNA and web gateway, missing the first decision means missing several product lines.
- **Old categories are shrinking.** With intrusion prevention and network access control forecast to decline, vendors framed only in those terms risk being described as legacy.
- **Inaccurate answers carry risk.** An AI answer that mixes up your products, or repeats an old vulnerability without the fix, can remove you from a list. If an assistant gets your products or patch history wrong, start with [how to fix wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers).

## What does GEO involve for a zero trust or SASE vendor?

It links your firewall, ZTNA and SASE products to the frameworks buyers already cite, in sources assistants trust. Being named is never guaranteed.

1. **Framework-mapped documentation.** Publish plain pages that map each product to NIST SP 800-207 components and CISA’s pillars and maturity stages, with honest limits.
2. **A public security record.** Keep advisories, patch timelines and certifications on readable pages. Transparent handling of your own vulnerabilities is evidence; hidden handling invites worse answers.
3. **Independent proof.** Independent test results, analyst coverage, participation in public projects such as NIST’s, and customer stories with named organizations give assistants sources other than you. For network security vendors, that outside proof is what [how brands build authority for AI search](https://underneath.agency/resources/how-brands-build-authority-for-ai-search) is about.
4. **Architecture guides.** Write honest guides to the real decisions: single-vendor or multi-vendor SASE, VPN replacement, firewall refresh. Third-party rankings of SASE and firewall vendors still carry weight, as our piece on [best-of lists](https://underneath.agency/resources/best-of-lists-ai-recommendations) shows.
5. **Fresh, consistent facts.** Keep product names, bundles and certifications current everywhere, including partner and reseller pages, since renamed products confuse buyers and assistants alike.
6. **Repeated checks on the zero trust questions.** Run the CISA-pillar, VPN-migration and vendor-risk questions your buyers ask through ChatGPT, Gemini, Perplexity, Claude, Copilot and Google’s AI features, more than once each. Our note on [how many prompts to track](https://underneath.agency/resources/how-many-prompts-to-track-ai-visibility) helps size that list.

For the wider picture of how security buyers use AI and what they trust, see our guide for [cybersecurity software companies](https://underneath.agency/resources/cybersecurity-software-revenue-from-ai-search). For the identity side of zero trust, see [how identity platforms get named by AI](https://underneath.agency/resources/iam-enterprise-demand-ai-search).

## What don’t we know yet about AI answers in zero trust buying?

Nobody has measured whether being named by AI changes who wins a firewall or SASE pilot.

- **No survey of network security buyers’ AI use.** Gartner’s figures cover B2B buyers in general.
- **Market figures are summarized forecasts.** The Gartner numbers here come from an analyst’s public summary of a paid forecast and will be revised.
- **The consolidation survey is dated.** Gartner’s vendor consolidation figures are from 2022.
- **The link to signed platform deals is the weakest part.** For what is known across categories, see [does AI visibility drive business results](https://underneath.agency/resources/does-ai-visibility-drive-business-results).

## Which questions should a firewall or SASE vendor test before the next refresh cycle?

The VPN replacement, SASE and firewall refresh questions that precede your pilots, plus how AI answers describe you.

List the decisions that lead to your pilots: VPN replacement, firewall refresh, SASE architecture, compliance needs. Put each one to the main assistants, and repeat it, so a single odd answer does not mislead you. Record whether you are named, how your products are framed against NIST and CISA terms, which sources are cited, and what is said about your security record. The gaps usually point to missing framework mappings, stale pages or thin independent proof.

To run that test with us, [ask us for a zero trust visibility review](https://underneath.agency/contact). We will look at where assistants place you in VPN replacement, SASE and firewall refresh answers, the decisions that feed pilots and multi-year platform agreements, and point out the gaps most likely to cost you those deals. The follow-up is described on our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page: framework mappings, a readable security record and repeated checks on the questions network and security teams ask.

## Frequently asked questions

### Do buyers really ask AI about zero trust vendors?

Many B2B buyers use AI in research: 45% in Gartner’s 2025 survey. No public study isolates network security buyers.

### Should we map our products to the CISA Zero Trust Maturity Model?

Yes, honestly. Buyers use its five pillars and maturity stages, and a clear mapping is easy for both people and assistants to repeat.

### Does a past vulnerability hurt AI visibility?

It can shape answers. Publish advisories, fixes and timelines plainly so assistants find the full story, not only the headline.

### Is single-vendor SASE always the winning position?

No. Dell’Oro calls the architecture contested; multi-vendor designs remain viable in complex enterprises, so describe where you fit.

### How fast do AI answers pick up new pages?

Often quickly. Pages under 90 days old were 17.4% to 22.6% of each assistant’s dated citations in our freshness study, so current advisories and architecture pages are worth the upkeep.

## Sources

- Louis Columbus, Software Strategies Blog (2026-04-01), [Gartner’s $246.2B Security Forecast shows 10 categories growing 2x to 3x the market](https://softwarestrategiesblog.com/2026/04/01/top-10-fastest-growing-security-categories-gartner-2026-forecast/)
- Dell’Oro Group (2026-09-15), [SASE Revenue Grows Above 20 Percent as SD-WAN Reaccelerates in 2Q 2026](https://www.delloro.com/news/sase-revenue-grows-above-20-percent-as-sd-wan-reaccelerates-in-2q-2026/)
- NIST (2025-06-11), [NIST Offers 19 Ways to Build Zero Trust Architectures](https://www.nist.gov/news-events/news/2025/06/nist-offers-19-ways-build-zero-trust-architectures)
- CISA (2023), [Zero Trust Maturity Model](https://www.cisa.gov/zero-trust-maturity-model)
- CISA (2025-09-25, updated 2026-04-23), [ED 25-03: Identify and Mitigate Potential Compromise of Cisco Devices](https://www.cisa.gov/news-events/directives/ed-25-03-identify-and-mitigate-potential-compromise-cisco-devices)
- Verizon (2025), [2025 Data Breach Investigations Report](https://verizon.com/about/news/2025-data-breach-investigations-report)
- Netskope (2025-09-17), [Netskope Announces Pricing of Initial Public Offering](https://www.netskope.com/press-releases/netskope-announces-pricing-of-initial-public-offering)
- Gartner (2026-05-20), [Gartner Survey Finds 69% of B2B Buyers Turn to Sales Reps to Validate AI-Generated Insights](https://www.gartner.com/en/newsroom/press-releases/2026-05-20-gartner-survey-finds-sixty-nine-percent-of-b-two-b-buyers-turn-to-sales-reps-to-validate-ai-generated-insights)
- Gartner (2022-09-12), [Gartner Survey Shows 75% of Organizations Are Pursuing Security Vendor Consolidation in 2022](https://www.gartner.com/en/newsroom/press-releases/2022-09-12-gartner-survey-shows-seventy-five-percent-of-organizations-are-pursuing-security-vendor-consolidation-in-2022)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [How fresh are the pages AI engines cite?](https://underneath.agency/research/ai-source-freshness-study)
- Underneath (2026), [When does a Reddit thread become evidence in Google’s AI?](https://underneath.agency/research/ai-reddit-citations-study)

---

This is the Markdown twin of https://underneath.agency/resources/network-security-zero-trust-customers-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Do off-site mentions help once an AI agent is reading about you?"
description: "Less than you might think. Once an AI agent researches a named business, whether it can read the site mattered far more than off-site mentions."
canonical: "https://underneath.agency/resources/off-site-mentions-vs-site-readability-ai-agents"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Do off-site mentions help once an AI agent is reading about you?

Less than you might expect: in the first large controlled test, off-site mentions had no measurable effect once an AI agent was researching a named business. What mattered was whether the agent could fetch and read the business’s own website. Off-site coverage still appears to help get a brand named in the first place, so the two jobs need different work.

## The short version

1. Across 37,927 agent research runs on 1,056 businesses, two stand-in measures of off-site optimization had no confirmed link to what agents said, while site readability did ([Finder and colleagues](https://arxiv.org/abs/2609.34951)).
2. Agents answered from the business’s own pages alone in 78% of runs when the site was readable, against 56% when it was not.
3. Readable businesses were clearly recommended 20% of the time, against 11% for the rest.
4. Answers built from the business’s site got 48.3% of the asked facts right, against 34.3% for answers built from elsewhere.
5. Off-site coverage still matters earlier: in [our brand study](https://underneath.agency/research/brand-entity-ai-recommendations-study), coverage on more independent sites went with 4.7 times the odds of being recommended for each tenfold increase.

## What did the first large test of this question find?

It found that being readable beat being talked about. Researchers at a company that sells agent-readiness scores for websites sent AI agents buyer questions about 1,056 real businesses ([Finder and colleagues](https://arxiv.org/abs/2609.34951)).

Each question named the business and asked about its pricing, features or setup. The agents could search the web and open pages before answering. In all, the team ran 37,927 of these research runs across four different agent setups.

They split businesses into two groups: sites an agent could read easily, and sites it could not. The groups were matched on fame, on how much the AI already knew about the brand, and on two stand-ins for off-site optimization. One counted mentions on third-party review and comparison sites, Wikipedia and Reddit; the other scored how findable the business was.

With readability held fixed, neither had a confirmed link to how the answer was built, whether the business was recommended, or what the run cost.

## What happens when an agent cannot read your site?

It answers anyway, using other people’s pages. In about 99% of runs that hit a dead end on a site, the agent still produced an answer from whatever it found elsewhere.

When the site was readable, the agent answered from the business’s own pages alone in 78% of runs, against 56% when it was not. As readability fell, agents ran more web searches, from 1.8 to 4.5 per run on average. Each extra search is another chance to pick up information the business never published.

The answers changed in tone too. When agents could not read the site, they were 4.4 times more likely to say they could not access the business. [Readable businesses were clearly recommended](https://underneath.agency/resources/do-ai-agents-recommend-readable-websites) in 20% of answers, against 11% for the rest.

Two AI judges from different companies had to agree before an answer counted as a clear recommendation.

## Does readability change what the agent gets right?

Yes, mainly by reducing what it leaves out. The researchers graded answers about 131 businesses against facts captured from each site. The table compares answers about the same business, agent and question.

| Where the answer came from | Asked facts right | Answers with none of the asked facts |
|---|---|---|
| The business’s own site | 48.3% | 6.7% |
| Other web pages | 34.3% | 25.0% |

Answers built from elsewhere were 3.7 times more likely to contain none of the facts the buyer asked for. The main failure was [omission, not invention](https://underneath.agency/resources/do-ai-agents-invent-or-omit-business-facts): facts never mentioned rose from 29% to 45%, while facts stated wrongly rose only from 4% to 6%.

One caveat matters. Compared by group rather than by source, accuracy differed by under two points, 51.7% against 49.8%, which was not a reliable difference. Reading the site helped; being in the readable group made reading more likely but did not guarantee it.

## So do off-site mentions not matter at all?

They matter, but at an earlier step. This study only asked about businesses by name, so it tested what happens after a brand is already on the agent’s radar.

Getting named in an open question is different. In [our brand study](https://underneath.agency/research/brand-entity-ai-recommendations-study), we counted the independent sites naming a brand in the pages an assistant cited. Each tenfold increase went with 4.7 times the odds of the brand being recommended.

A large vendor-authored study of 102 brands also found off-site pages dominate open questions. Only 2.9% of citations pointed at the tracked brand’s own website ([Kumar](https://arxiv.org/abs/2606.20065)). The agent study’s authors frame it the same way: being findable helps a business get named; a readable site decides what the answer says once the agent looks.

## How ready are most websites for agents today?

Not very. In [our study of the top 10,000 websites](https://underneath.agency/research/agent-readable-web-study), we asked each homepage for a Markdown version, a plain-text format agents read easily. Only 3.2% of live sites returned one.

Some sites also shut agents out on purpose. In [our crawler study](https://underneath.agency/research/ai-crawler-blocking-study), 15.2% of top sites with a readable robots.txt file, the file that tells crawlers what they may visit, blocked OpenAI’s training crawler. A smaller 7.4% blocked the crawler ChatGPT uses for search.

## Why does it matter that your site agrees with itself?

Because AI answers tend to repeat whatever your own sources say, including their contradictions.

In [our study of business facts in AI answers](https://underneath.agency/research/ai-business-facts-accuracy-study), we compared phone numbers in AI answers with each business’s Google profile. Where the profile number did not appear on the business’s own website, 30.6% of the numbers AI engines gave differed from the profile. Where it did appear, 1.6% did.

## What should you do about it?

Split the work into getting named and getting read, and fund the second as seriously as the first.

1. Check whether AI agents can read your key pages. Load them with scripts turned off; if pricing, features or setup steps disappear, agents may miss them too.
2. Review your robots.txt and bot protection. Make sure you are not blocking the search and user-requested agents you want to reach you. For Google’s case, see [whether blocking Google-Extended hurts AI Overviews](https://underneath.agency/resources/does-blocking-google-extended-hurt-ai-overviews).
3. Publish the facts buyers ask about in plain text: prices, plans, features and setup steps, on pages an agent can reach in one or two clicks.
4. Keep facts consistent across your own site and profiles, so an agent finds one answer.
5. Keep earning independent coverage, but do not expect it to fix what agents say once they read about you.

If you want help making your site readable to AI agents, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

No study has improved a site’s readability and measured the change in agent answers before and after.

- The main study is by a vendor that sells agent-readiness scoring, and it used its own scoring tool to define readable sites.
- Off-site work was measured by only two rough stand-ins. Richer measures of reputation might behave differently, and one showed a weak, unconfirmed hint of helping recommendation.
- The two groups were not matched on how much content their sites publish, which the authors flag as a gap.
- The businesses skewed toward software and commerce, in English, with fact-lookup questions only.
- How readable sites perform in open category questions, where no business is named, was not tested.

## Frequently asked questions

### Do mentions on other sites help AI agents recommend my business?

Not once the agent is researching you by name, in current evidence. In a study of 1,056 businesses, two stand-ins for off-site optimization had no confirmed effect, while site readability did.

### What happens if an AI agent cannot read my website?

It answers from other pages. In about 99% of runs that hit a dead end, the agent still answered, and businesses with hard-to-read sites were clearly recommended 11% of the time, against 20%.

### Are AI agents more accurate when they read my site?

Yes. Answers built from the business’s own site got 48.3% of the asked facts right, against 34.3% for answers built from other pages.

### Is PR still worth it for AI search?

Yes, for getting named. In our brand study, more independent coverage went with 4.7 times the odds of being recommended per tenfold increase.

## Sources

- Finder, Elovic, Shalev and Yosef (2026), [AX is the New AEO](https://arxiv.org/abs/2609.34951), arXiv:2609.34951.
- Kumar (2026), [Generative Engine Optimization at Scale: Measuring Brand Visibility Across AI Search Engines](https://arxiv.org/abs/2606.20065), arXiv:2606.20065.
- Underneath (2026), [Do Wikipedia and schema make AI assistants recommend a brand?](https://underneath.agency/research/brand-entity-ai-recommendations-study)
- Underneath (2026), [Do websites serve Markdown to AI agents? 2026 data](https://underneath.agency/research/agent-readable-web-study)
- Underneath (2026), [Which AI crawlers do top websites block? 10,000 sites, 2026](https://underneath.agency/research/ai-crawler-blocking-study)
- Underneath (2026), [Do AI answers match a business’s Google profile?](https://underneath.agency/research/ai-business-facts-accuracy-study)

---

This is the Markdown twin of https://underneath.agency/resources/off-site-mentions-vs-site-readability-ai-agents. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "On-page signals linked to AI Overview and Perplexity citations"
description: "Dates, clean HTML structure and structured data track with AI citations, but ranking matters more, and only a machine-readable date held up on Google."
canonical: "https://underneath.agency/resources/on-page-signals-linked-to-ai-citations"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# What on-page signals are linked to citations in Google AI Overviews and Perplexity?

Visible and machine-readable dates, a clean heading structure and structured data are the on-page signals most often linked to AI citations. The links are real but modest: on Google, ranking position matters more than any page feature, and only a machine-readable date held up in our own test. No study has yet shown that adding these signals causes more citations.

## The short version

1. Dates, clean HTML structure and structured data were the signals most linked to citation in [an audit of 1,100 B2B software pages](https://arxiv.org/abs/2509.10762) cited by Brave, Google and Perplexity.
2. In [our study of 3,096 ranking pages](https://underneath.agency/research/ai-overview-cited-pages-study), a machine-readable date was the only page feature still linked to AI Overview citation after a fair comparison: +7.9 points.
3. That date link was confined to informational searches: +17.8 points there, against +1.0 on commercial and transactional searches.
4. Perplexity cited much lower-scoring pages than Google: an average quality score of 0.300, against 0.687 for Google AI Overviews, in the same audit.

## Which on-page signals are most linked to AI citations?

Dates and freshness metadata, semantic HTML and structured data showed the strongest links. That is the main finding of the GEO-16 audit by Kumar and Palkhouski, at UC Berkeley and Wrodium Research.

They ran 70 prompts aimed at B2B software buyers and collected 1,702 citations from Brave, Google AI Overviews and Perplexity. AI Overviews are the AI summaries at the top of Google’s results. Other terms used here are defined in [our plain-language AI search glossary](https://underneath.agency/resources/glossary). The team then scored 1,100 unique cited pages on 16 quality “pillars”, from metadata to readability.

Three pillars stood out.

On a scale where 0 means no link and 1 a perfect one, Metadata and Freshness scored 0.68, Semantic HTML 0.65 and Structured Data 0.63. Semantic HTML means one main title and a logical order of headings, so a machine can tell what each part of the page is. Evidence and citations, authority and trust, and internal linking came next.

## How strong is that evidence?

It is suggestive, not proof, because the audit compared pages as they were and changed nothing. The authors say so themselves.

They report that pages scoring at least 0.70 overall, with at least 12 of the 16 pillars rated well, reached a 78% cross-engine citation rate. Pages cited by more than one engine (134 URLs) scored 71% higher on quality than pages cited by one engine only.

The limits matter. The pages were English B2B software content, collected at a single point in time.

Brand reputation and backlinks, links from other sites, were not controlled. The authors note they “may influence both scores and citations.” They list schema experiments as future work.

## Does schema markup get pages cited in AI Overviews?

The best evidence says schema alone does not. Once pages are compared with competitors on the same search, most schema links shrink toward zero.

In our study of 3,096 top-10 pages on 486 US searches, Organization schema added +0.5 points to the chance of citation and BreadcrumbList schema +1.3, both meaningless. FAQPage schema showed +5.9 points but did not survive a correction for testing 14 features at once.

The strongest design we know of points the same way. Our study summarizes a before-and-after test by the SEO tool maker Ahrefs, of 1,885 pages that added schema against 4,000 controls.

That test found adding schema “produced no major uplift in citations on any platform”. Schema is also common among cited pages anyway. In [a study of 615 ChatGPT health citations](https://arxiv.org/abs/2601.17109), 74.3% of cited sources used schema markup.

## Do dates and freshness help pages get cited?

A machine-readable date helps on informational searches, but a recent date does not add much on Google. On Google, being dated seems to matter more than being new.

In our study, a machine-readable date was the only one of 14 features still clear after the correction, at +7.9 points. On informational searches it rose to +17.8 points. Pages dated within the last 90 days were, if anything, 4.0 points less likely to be cited than other dated pages.

[Our freshness study](https://underneath.agency/research/ai-source-freshness-study) adds a warning about what dates mean. Among pages in Google’s top 10 carrying a recent date, 66.4% were older pages with a new modified date. A lab test by [Vishwakarma and colleagues](https://arxiv.org/abs/2605.25517) did find a strong pull toward recent dates, but it compared content dated 2026 versus 2019 in a simulation.

| Signal on Google AI Overviews | Linked to citation? | Source |
|---|---|---|
| Machine-readable date | Yes, +7.9 points, mainly informational searches | Our study of 3,096 pages |
| Date within 90 days | No (−4.0 points) | Our study of 3,096 pages |
| Organization or BreadcrumbList schema | No (+0.5 and +1.3) | Our study of 3,096 pages |
| Question-style headings | Not distinguishable from zero (+3.5) | Our study of 3,096 pages |
| HTML table | No for citation (+0.6), yes for wording used | Our study of 3,096 pages |

## Is Perplexity different from Google?

Yes: Perplexity cites more sources, and the pages it cites score lower on quality audits. It also cites newer pages than Google ranks.

In the GEO-16 audit, Perplexity’s cited pages averaged a quality score of 0.300, against 0.687 for Google AI Overviews and 0.727 for Brave. In [a study of 602 prompts](https://arxiv.org/abs/2604.25707), Perplexity cited 16.35 sources per prompt on average, against 12.06 for Google. Each of its cited pages shaped less of the answer than ChatGPT’s sources did.

In our freshness study, 19.2% of Perplexity’s dated citations were pages published in the last 90 days, against 9.0% of Google’s top 10 for the same questions. Access rules also work differently. [Our crawler study](https://underneath.agency/research/ai-crawler-blocking-study) found 40.4% of the pages Perplexity cited on top sites were closed to its crawler by the site’s own robots.txt file.

## Do headings, tables and page length matter?

Not much for being cited, though tables seem to help once a page is cited. Most layout signals vanish when pages are compared fairly.

Question-style headings showed +3.5 points in our study, which is not distinguishable from zero. Word count did not matter once position and search were held constant. Tests that changed layout alone are weighed in [whether restructuring content lifts AI citations](https://underneath.agency/resources/does-content-structure-increase-ai-citations).

But cited pages with an HTML table were 14.7 points more likely to have the answer’s wording traced back to them. Tables likely hold the specific figures an answer repeats.

Ranking dwarfed all of this. Of top-10 pages, 41.7% ranking 1 to 3 were cited, against 20.1% at positions 7 to 10. Most of the variation, 91.4%, was explained by nothing we measured.

## What should you do about it?

Treat these signals as basic hygiene, and put most effort into ranking and answering the question well. Concretely:

1. Show a visible date and a matching machine-readable date on informational content, and only update it when the content changes.
2. Keep one main title and a logical heading order on every important page.
3. Keep structured data valid and matching the visible page, but do not expect it to lift citations alone.
4. Put key figures in simple tables, so answers that cite you can quote you accurately.
5. Check your robots.txt deliberately for each AI crawler instead of relying on defaults.
6. Measure Perplexity and Google separately, because they reward different things.

If you want help prioritizing these across a large site, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

No published study has changed these signals on real pages and measured citations before and after across engines.

- **Cause and effect.** The GEO-16 audit and our study both compare existing pages. Better-run sites may simply have better pages and better rankings.
- **Other sectors and languages.** GEO-16 covered English B2B software pages; our study covered US searches on one day.
- **Brand strength.** Backlinks and reputation were not controlled in the audit, and the GEO-16 scores were produced by the authors’ own framework.
- **Stability.** Both studies are snapshots, and engines change often.
- **Perplexity in depth.** Fewer studies cover Perplexity than Google, and none tested schema there directly.

## Frequently asked questions

### Does schema markup help with Google AI Overviews?

Not on its own, according to the best current evidence. In our study of 3,096 pages, most schema types showed no reliable link once pages were compared on the same search, and a before-and-after test of 1,885 pages found no major uplift.

### Should I add dates to my pages for AI search?

Yes, on informational content. A machine-readable date was the one feature linked to AI Overview citation in our study, worth +17.8 points on informational searches, but a newer date added nothing extra.

### Why does Perplexity cite different pages than Google?

Perplexity casts a wider net, and its cited pages score lower on quality audits. It cited 16.35 sources per prompt in one study and cited pages with an average quality score of 0.300, far below Google’s 0.687.

### What is semantic HTML?

It is page code that labels each part by its role: one main title, ordered subheadings, lists and tables. It was one of the three signals most linked to AI citation in the GEO-16 audit.

### Do author bios help AI citations?

The evidence does not show it. In our Google study, an author signal showed no reliable link, and 64.7% of ChatGPT-cited health sources had [no author attribution at all](https://underneath.agency/resources/do-ai-cited-pages-need-author-bylines).

## Sources

- Kumar and Palkhouski (2025), [AI Answer Engine Citation Behavior An Empirical Analysis of the GEO16 Framework](https://arxiv.org/abs/2509.10762), arXiv:2509.10762.
- Vishwakarma, Kumar and Jamidar (2026), [What Gets Cited: Competitive GEO in AI Answer Engines](https://arxiv.org/abs/2605.25517), arXiv:2605.25517.
- Zhang, He and Yao (2026), [From Citation Selection to Citation Absorption: A Measurement Framework for Generative Engine Optimization Across AI Search Platforms](https://arxiv.org/abs/2604.25707), arXiv:2604.25707.
- Jacques and colleagues (2026), [Authority Signals in AI Cited Health Sources: A Framework for Evaluating Source Credibility in ChatGPT Responses](https://arxiv.org/abs/2601.17109), arXiv:2601.17109.
- Underneath (2026), [What pages cited by AI Overviews have in common: 3,096 pages](https://underneath.agency/research/ai-overview-cited-pages-study)
- Underneath (2026), [How fresh are the pages AI engines cite?](https://underneath.agency/research/ai-source-freshness-study)
- Underneath (2026), [Which AI crawlers do top websites block?](https://underneath.agency/research/ai-crawler-blocking-study)

---

This is the Markdown twin of https://underneath.agency/resources/on-page-signals-linked-to-ai-citations. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Can I trust a one-time AI visibility report for my brand?"
description: "Not on its own. AI answers vary between runs, so a one-off report can mistake noise for a lead. Reliable figures need repeated runs over two to four weeks."
canonical: "https://underneath.agency/resources/one-time-ai-visibility-report"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Can I trust a one-time AI visibility report for my brand?

Not on its own. AI assistants give a different answer each time, so a report built from one round of questions can rank you ahead of a competitor, or behind, purely by chance. Researchers who have measured the noise recommend repeated runs, several wordings and several engines, tracked over two to four weeks, before a number is solid enough to drive a budget decision.

## The short version

1. In our test of three assistants, a single ChatGPT answer showed only 57.8% of the brands that five answers to the same question named between them.
2. A visibility lead of 9.5% against 6.0% turned out to be statistically indistinguishable from a tie once the measurement’s margin of error was added (Sielinski, who works for an AI visibility company).
3. Swiss researchers recommend at least 7 runs per question per day and rolling results over two to four weeks before trusting a brand-level number.
4. Repeating one question helps less than asking it in more languages, on more engines and in more wordings, according to a study of 20 brands in 8 languages.
5. One vendor’s data suggest a single check is right more often for clear cases: 63.2% of brand-and-question pairs were never mentioned in any run.

## Why can’t one check tell you where your brand stands?

Because the same question asked twice gets a different answer, so one answer is only a sample. In [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), we asked ChatGPT, Gemini and Perplexity the same 20 buyer questions five times each. Of the brands ChatGPT named for a question across five runs, 25.2% appeared in all five and 36.6% in only one.

A single ChatGPT answer showed 57.8% of the brands its five answers named between them; for Gemini, 48.4%. ChatGPT’s first pick changed at least once for 80.0% of questions. How steady each engine is on its own is compared in [which AI engine answers most consistently](https://underneath.agency/resources/most-consistent-ai-engine-for-brands). Timing mattered too: an answer to the same question about 4.4 hours later overlapped less with the earlier runs than they did with each other.

[Schulte and colleagues](https://arxiv.org/abs/2604.07585) put it plainly after tracking four engines for 45 days. A snapshot “may differ substantially from a second query executed minutes later under identical conditions.” [Malthouse and colleagues](https://arxiv.org/abs/2609.16304) at Northwestern University reach the same view: a recommendation should be treated as a set of probabilities, not a single list.

## How big is the noise compared with the differences you care about?

Often bigger than the gaps reports highlight. [Sielinski](https://arxiv.org/abs/2603.08924) gives a typical case from ChatGPT search. In a sample of 200 running-gear questions, tomsguide.com held a 9.5% share of citations and runnersworld.com 6.0%. The first looks like the clear leader.

Add the margin of error and the picture changes. The plausible range ran from 5.5% to 12.5% for the first site and from 4.0% to 8.0% for the second, so the ranges overlap and the lead may be noise. On ChatGPT search, those ranges were typically 3 to 6 points wide. By the same logic, a rise from 8% to 11% after a content change cannot be credited to the change. Telling a real gain from noise is covered in [judging whether a GEO campaign worked](https://underneath.agency/resources/did-geo-improve-ai-citations).

Single samples can also mislead badly. One site took a share of 0.032 of Gemini’s citations in the first daily sample, against a long-run average of 0.005 across the following days. Anyone reading only the first day would have ranked it among the top sources.

## How many checks does a reliable number need?

More than most one-off reports use, and the right number depends on the engine and topic. Schulte’s team found brand-level figures needed at least 7 runs per question per day, and source-level figures at least 8. Over time, a brand’s figure became reasonably steady after about 10 days and tight after about 24 days.

Sielinski estimated how many questions it takes to pin a domain’s citation share within a five-point range. Gemini needed roughly 40 to 50 questions, Perplexity about 100 and ChatGPT search 150 or more.

| What you want to know | What the research suggests |
|---|---|
| Whether a brand appears for one question | At least 7 runs per day (Swiss study) |
| A share accurate to about 10 points | About 97 runs (our consistency study) |
| A citation share accurate to 5 points | 40 to 150 or more questions, by engine (Sielinski) |
| A stable trend for one brand | Two to four weeks of rolling results (Swiss study) |

Five runs leave wide margins. In our study, a brand named in 3 of 5 runs could plausibly appear anywhere from 23.1% to 88.2% of the time. In a later paper, Sielinski found that the data needed before rankings settled, under one common test, ranged from 40 responses to never across 30 engine and topic combinations.

## Is asking the same question more times enough?

No: the questions you choose and the engines you ask matter as much as repeat runs. [Żatuchin](https://arxiv.org/abs/2607.13304), who is affiliated with the AI visibility company Rankfor.AI, analyzed answers about 20 Central and Eastern European brands in 8 languages on 3 engines. The language of the question explained 26.5% of the variation in a single answer; the brand itself explained 1.5%.

On a 0-to-1 scale of how reliably brands could be ranked, one answer scored near 0.01. Even 8 languages, 3 engines and 15 wordings together reached only about 0.36. The outcome measured was the tone of answers, and 91.9% of answers were neutral, so the figures may not carry over to recommendations.

Wording matters on its own. In [our phrasing study](https://underneath.agency/research/ai-prompt-phrasing-study), simply asking again kept the same first brand 68.0% of the time; adding “on a tight budget” kept it 15.3%. [Martinez](https://arxiv.org/abs/2609.06811) shows how much the choice of question set can move a headline figure. Reweighting the same published data gave AI answers appearing for 39.7% or 70.5% of searches, depending only on how the question groups were weighted.

## Can a single check ever be trusted?

Sometimes, for clear-cut cases: a brand that is never or always named is easy to spot. [Kumar](https://arxiv.org/abs/2606.20065), a co-founder of the tracking company Ranqo, analyzed more than 100 brands tracked between March and May 2026. Of brand, question and engine combinations, 77.5% were always or never mentioned across runs, and only 6.8% flipped often.

The most common result was absence: 63.2% of combinations were never mentioned at all. For the rest, Kumar recommends at least 3 runs before trusting a mention rate. A one-off check can therefore tell you that you are missing, but it is a poor guide to how often you appear when you sometimes do.

## Why does tracking need to be continuous?

Because the engines change, and a brand’s position can drift without anyone noticing. Kumar’s data show that brands which did nothing slowly lost visibility on ChatGPT and Perplexity, by an average of 1.34% per tracking run on ChatGPT, while holding steady on the other engines. Kumar notes that part of this may be measurement drift.

[Baig and colleagues](https://arxiv.org/abs/2606.16344), who audited how AI models choose hotels, end with a warning for managers. The weights they measured describe the model versions of the day, and “the appropriate managerial posture is continuous measurement rather than a one-time fix.”

## What should you do about it?

Treat a one-off AI visibility report as a starting hypothesis, and ask how it was measured before acting on it.

1. Ask how many times each question was run, on how many days, and on which engines. One run per question is a sample of one.
2. Ask for margins of error. A lead that sits inside the margin is a tie.
3. Use [a broad, fixed set of questions](https://underneath.agency/resources/how-to-design-ai-visibility-tracking) in several wordings, and in each language your buyers use.
4. Track over at least two to four weeks before comparing yourself with competitors or judging a change.
5. Act quickly on clear absences. If you are never named across repeated runs, that finding is reliable.
6. Keep measuring after any change, because engines update and early gains can fade.

For help building tracking that holds up, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

There is no agreed standard for how much AI visibility data is enough.

- The recommended run counts come from small settings: German questions on Swiss servers, three consumer topics, or 20 regional brands.
- Several of the key studies come from authors who sell tracking tools, which gives them a stake in the case for repeated measurement.
- Żatuchin measured the tone of answers, not recommendations; a version using recommendations is still to come.
- No study has yet shown how much a measurement error costs a brand in lost sales.

## Frequently asked questions

### How many times should I run a prompt to measure AI visibility?

At least 7 times per day per question, according to a 45-day Swiss study. Precise shares take far more: pinning one to within about 10 points took about 97 runs in our estimate.

### How long should I track AI visibility before judging results?

Two to four weeks of rolling results, according to the Swiss study. A brand’s figure became reasonably steady after about 10 days and tight after about 24.

### Is an AI visibility score from a single audit useful?

Yes for spotting clear absences, much less for rankings. In one vendor dataset, 63.2% of brand and question combinations were never mentioned in any run, while close rankings often fell within the margin of error.

### Why did my brand’s AI visibility change between two reports?

Possibly only chance. In one example, a site at 9.5% and another at 6.0% had overlapping margins of error, so a change of a few points can be noise.

## Sources

- Ronald Sielinski (2026), [Quantifying Uncertainty in AI Visibility: A Statistical Framework for Generative Search Measurement](https://arxiv.org/abs/2603.08924), arXiv:2603.08924.
- Ronald Sielinski (2026), [From Stochastic to Stable: Rank Stability and Structural Sufficiency in AI Visibility Measurement](https://arxiv.org/abs/2607.10341), arXiv:2607.10341.
- Julius Schulte, Malte Bleeker and Philipp Kaufmann (2026), [Don’t Measure Once: Measuring Visibility in AI Search (GEO)](https://arxiv.org/abs/2604.07585), arXiv:2604.07585.
- Dmitrij Żatuchin (2026), [Where Does the Noise Come From? A Variance-Components Decomposition of Non-Determinism in LLM Brand Answers](https://arxiv.org/abs/2607.13304), arXiv:2607.13304.
- Pratyush Kumar (2026), [Generative Engine Optimization at Scale: Measuring Brand Visibility Across AI Search Engines](https://arxiv.org/abs/2606.20065), arXiv:2606.20065.
- Olivier Martinez (2026), [Measuring GEO Visibility: Prompt Corpora Define the Answer Market](https://arxiv.org/abs/2609.06811), arXiv:2609.06811.
- Edward Malthouse and colleagues (2026), [Evaluating Brand Retrieval and Ranking in Large Language Model Recommendations](https://arxiv.org/abs/2609.16304), arXiv:2609.16304.
- Mirza Samad Ahmed Baig, Syeda Anshrah Gillani and Asher Ali (2026), [Whose hotel does the AI recommend? An algorithm audit of reputation signals in LLM-assisted hotel selection](https://arxiv.org/abs/2606.16344), arXiv:2606.16344.
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)
- Underneath (2026), [Does rewording a question change AI brand recommendations?](https://underneath.agency/research/ai-prompt-phrasing-study)

---

This is the Markdown twin of https://underneath.agency/resources/one-time-ai-visibility-report. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How can outdoor brands get onto AI gear shortlists?"
description: "By matching gear to specific activities and conditions, and earning the expert reviews and retailer guides AI draws on when it builds an outdoor shortlist."
canonical: "https://underneath.agency/resources/outdoor-brands-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How does an outdoor brand get onto the shortlist when people ask AI what gear to buy?

By publishing the specifications that matter for a given activity and set of conditions, keeping them consistent at every retailer, and earning a place in the expert reviews and buying guides that AI answers draw on. Outdoor gear is bought for a trip, a season or a climate, often by people who are still new to it. Those are the shoppers most likely to ask an assistant to explain the choice, and the shortlist it gives them can shape a whole kit.

## The short version

1. The market is big and full of newcomers: [the Outdoor Industry Association](https://oia.outdoorindustry.org/2026-participation-trends-report-summary) counted a record 183.2 million US participants in 2025, and more than half of today’s participants have less than 10 years of outdoor experience ([Shop Eat Surf Outdoor](https://shop-eat-surf-outdoor.com/?p=620899)).
2. The money behind it is national-scale: the outdoor recreation economy added $639.5 billion, 2.3% of US GDP, in 2023, with retail trade contributing $156.3 billion ([Bureau of Economic Analysis](https://www.bea.gov/news/2024/outdoor-recreation-satellite-account-us-and-states-2023)).
3. Growth is concentrated in technical brands: [Amer Sports](https://s21.q4cdn.com/769814866/files/doc_news/Amer-Sports-Reports-Second-Quarter-2026-Financial-Results-Raises-Full-Year-Revenue-Margin-and-EPS-Guidance-2026.pdf) grew its Arc’teryx-led Technical Apparel segment 32% to $674 million in the second quarter of 2026.
4. AI is already moving orders for outdoor names: [Shopify](https://www.shopify.com/news/agentic-holiday-2026) reports that Stanley 1913, which it describes as growing from an iconic outdoor brand, had five times as many AI-referred US orders as in the same window a year earlier.
5. Google opens its own [AI Mode shopping announcement](https://blog.google/products/shopping/google-shopping-ai-mode-virtual-try-on-update/) with an outdoor question, “How do I choose a pair of hiking boots?”, and describes running several searches at once to work out what makes gear right for a trip’s conditions.

## Who buys outdoor gear today, and what is a customer worth?

A growing, less experienced crowd of hikers, campers, anglers and families, who often expand into several activities.

The OIA’s 2026 report shows participation reaching 59% of Americans aged six and over, even as growth slowed to 1.1%. The base is changing. Americans aged 65 and older are a fast-growing group, at 23.9 million participants with a participation rate of 41.6%. Kids aged 6 to 12 grew 5% to 22.6 million. Meanwhile the average participant gets out less often: 65.2 outings a year, against 87 in 2012, according to the OIA’s research director.

What a customer is worth comes from breadth more than frequency. The same report found that more than 9 in 10 campers also hike, fish, paddle or bike, which is why coverage of the report calls camping the on-ramp. A family that buys a tent this year may buy packs, rain shells, fishing gear and a paddleboard over the next few. Our inference is that the first trip’s shopping list is the start of a multi-year relationship, and the brand on that list has the first chance at the rest.

Where the sale lands varies. REI, the largest consumer co-op in the US, reported [$3.54 billion in net sales](https://gearjunkie.com/outdoor-business/rei-narrows-2025-losses-union-boycotts) in 2025 and more than 26 million members. Specialty retailers like it remain central for many brands. At the same time, Amer Sports’ direct-to-consumer revenue rose 39.9% to $896.8 million in the quarter, out of total revenue of $1,633 million. A brand can win the recommendation and see the sale go through its own site or a retail partner.

## How do outdoor shoppers research gear, and where does AI fit?

Through expert reviews, retailer guides, forums and video, and now AI that summarizes all of them for a specific trip.

Outdoor shopping has always been research-heavy. Buyers want to know whether a tent survives wind, whether a jacket breathes on a climb, how much a pack weighs, and whether a sleeping bag is warm enough in October. Much of that guidance comes from independent testers, specialty retailers’ advice pages and community discussion.

AI search now pulls those threads together. [Google says](https://blog.google/products/shopping/google-shopping-ai-mode-virtual-try-on-update/) AI Mode draws on a Shopping Graph of more than 50 billion product listings, and gives an example of the method: asked for a travel bag for a rainy trip to Portland, Oregon, it runs several simultaneous searches “to figure out what makes a bag good for rainy weather and long journeys,” then suggests waterproof options. That is exactly how outdoor gear is chosen: conditions first, then products that meet them.

Newcomers are the shoppers most likely to need that kind of help. With more than half of participants under 10 years of experience, many shoppers do not know what “hydrostatic head,” “fill power” or “R-value” mean, but they can describe the trip. Shopify’s holiday report makes the same point about AI in general: shoppers can describe a problem in their own words and an assistant can match it to a product that solves it. Shopify also reports that 65% of shoppers plan to use AI for at least one shopping task this season.

## Which outdoor gear questions do people ask AI assistants?

Questions about a trip, its conditions and a budget, plus head-to-head comparisons and first-time kit lists.

We wrote the examples below to illustrate outdoor shopping questions; they are not observed prompts:

- Conditions: “Rain jacket for Pacific Northwest hiking that still breathes on steep climbs.”
- Weight and budget: “Two-person backpacking tent under 4 pounds for windy Colorado trips, under $400.”
- First trip: “What do I need for our first car camping trip with two young kids?”
- Comparison: “Arc’teryx Beta vs Patagonia Torrentshell for everyday hiking and city rain.”
- Season: “Warmest sleeping bag under $300 for late-fall camping in the Smokies.”
- New activity: “Good beginner binoculars for birding, light enough for long walks.”
- Durability: “Which hiking boots can be resoled, and how long do they last?”

The first-trip question matters most commercially, in our view, because one answer can list ten or more items. A brand named for the tent may also be considered for the sleeping bags and the stove. Safety-critical gear, such as avalanche equipment or climbing protection, deserves extra care: describe certifications and intended use precisely, and do not imply that a product makes a dangerous activity safe.

## How does an AI shortlist become an outdoor sale?

Through a named product, a check against trusted reviews, and a purchase at a retailer or the brand’s site.

**Shortlist.** The assistant names a few products for the trip and conditions described. For many newcomers, this may be the first time they hear the brand’s name.

**Verification.** Outdoor buyers check before spending a few hundred dollars. They read expert reviews, compare weights and ratings, and look for people who have used the gear in similar conditions. If the assistant’s description of your product disagrees with what reviewers and retailers say, the shopper notices.

**Purchase.** [OpenAI documents](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search) that when a shopper opens a product in ChatGPT, it may list several merchants ranked “based on factors like availability, price, quality, and whether they are the maker or primary seller of that item.” A specialty retailer with full stock data might be listed alongside, or above, the brand.

**Expansion.** The shopper returns for the next activity. Stanley 1913’s senior director of commerce technology told Shopify that its AI work now spans social media, PR and data analytics, “to increase our share of the agentic conversation in awareness, consideration and decision making.” That breadth fits a category where the next purchase follows the next adventure.

## What decides whether an assistant names your gear?

Platforms document product data and outside content; studies show reviews, lists and fresh pages matter; the rest is inference.

**Documented by the platforms.** OpenAI says ChatGPT considers “structured metadata from first-party and third-party providers (e.g., price, product description) and other third-party content,” and builds review summaries from public websites. Google describes using its Shopping Graph’s reviews, prices, color options and availability, and refreshing more than 2 billion listings every hour.

**Observed in our studies.** Outdoor buying guides are often numbered “best” lists, and some are published by sellers. In [our study of “best of” lists cited by AI](https://underneath.agency/research/self-promoting-best-lists-study), 24.2% of the cited numbered lists with an identifiable publisher ranked that publisher first. Freshness counts too. In [our freshness study](https://underneath.agency/research/ai-source-freshness-study), the four assistants cited pages first published about half as long ago as Google’s top results for the same questions (a ratio of 0.50), and pages from the last 90 days made up 17.4% to 22.6% of their dated citations against 6.9% of Google’s top 10. Annual gear guides fit that pattern. Community discussion is a smaller factor than people assume for retail: in [our Reddit study](https://underneath.agency/research/ai-reddit-citations-study), 10.3% of retail and ecommerce AI Overviews cited a Reddit thread, and ChatGPT and Claude cited none across 80 buyer questions. Where Google’s AI did cite Reddit, the cited threads had more discussion, a median of 40 comments against 20.

**Our inference for outdoor gear.** The trust factors are measurable: weight and packed size, waterproof and breathability figures, temperature ratings and the standard used, materials, warranty and repair options, and real-world testing in named conditions. Independent testers and specialty retailers publish those facts in comparable form. A reasonable expectation is that a brand whose specifications match across its own site, retailer listings and test reviews gives an assistant consistent evidence. [Our article on “best of” lists](https://underneath.agency/resources/best-of-lists-ai-recommendations) and [our guide to how brands build authority for AI search](https://underneath.agency/resources/how-brands-build-authority-for-ai-search) cover the mechanics.

## What does an outdoor brand risk by staying invisible to AI?

The newcomer’s first kit, and the adjacent activities that follow it.

We found no published figure for outdoor sales lost to AI answers, and we will not estimate one. The risk is structural. The fastest-growing groups, newcomers, older adults and families, need the most guidance. Camping leads into hiking, fishing and paddling. A brand missing from the answer for “what do I need for our first camping trip” loses not one item but a place in that family’s gear closet. Stanley 1913’s fivefold growth in AI-referred orders, a self-reported result, shows what the other side of that risk looks like for one outdoor brand.

## How does generative engine optimization work for outdoor brands?

By making each product’s performance in real conditions clear, current and confirmed by the testers and retailers AI reads.

- **Conditions-first product pages.** Say what each product is for: terrain, weather, season, trip length. Give weight, packed size, ratings and the test standards behind them in plain text, not only in images. Training apparel faces similar activity questions, covered in [how sportswear brands get named by AI](https://underneath.agency/resources/sportswear-brands-ai-search).
- **Retailer and feed consistency.** Make specialty retailer listings, marketplace pages and product feeds match your specifications, prices and stock, so assistants can name you as the maker with accurate availability.
- **Independent testing.** Get gear into expert review programs and keep reviewers informed of updates, so current-year guides include current models.
- **Beginner and kit content.** Publish first-trip checklists and “how to choose” guides that answer newcomers’ questions honestly, including when they do not need the premium option.
- **Durability and repair facts.** Warranty terms, repair services and resoling are real buying criteria in this category; state them clearly. Boot makers can also use [how footwear brands get shoes found in AI answers](https://underneath.agency/resources/footwear-brands-ai-search).
- **Visibility tracking.** Check which products are named for which trips and conditions across ChatGPT, Google AI Mode, AI Overviews, Gemini and Perplexity, repeated over time.

No one can guarantee that an assistant will recommend a product. What a brand controls is whether accurate, current, verifiable facts exist where assistants look.

## What is still unknown about AI and outdoor gear sales?

How many outdoor purchases start with AI, and how assistants weigh expert tests against retailer data.

The participation and economic data here are strong; the AI-specific data are thin. We found no outdoor-only study of AI-referred sales, no published share of outdoor shoppers using AI for gear, and no platform documentation on how outdoor gear in particular is ranked. Stanley 1913’s result is a merchant story reported by Shopify. Our studies cover many industries, not outdoor gear alone. The BEA’s most recent detailed release covers 2023. Use it to choose where to start, not as proof of payback.

## Where should an outdoor brand start?

With an audit of the trip and conditions questions that lead to your hero products and first-trip kit lists.

Choose the products that anchor your range, a shell, a tent, a boot, a pack, and write down the trip and conditions questions their buyers would ask. Then check what ChatGPT, Google’s AI features, Gemini and Perplexity say, run several times: which products are named, whether specifications are right, which reviewers and retailers are cited, and where the shopper is sent to buy. If you would like us to run that check, [contact our team](https://underneath.agency/contact): we will map where your gear appears and where it is missing, and the work most likely to earn more places on the shortlists that turn into first kits and repeat customers. From conditions-first product pages to consistent specifications at every retailer that sells your gear, our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page covers how that work is carried out.

## Frequently asked questions

### Do expert gear reviews still matter if shoppers use AI?

Yes. OpenAI says ChatGPT uses third-party content and public reviews, and outdoor buyers still verify before buying. Expert tests give assistants comparable facts to work from.

### Should we publish our own “best of” gear lists?

You can, but be honest. In our study, about a quarter of AI-cited numbered lists with an identifiable publisher ranked that publisher first; shoppers and platforms may discount lists that look self-serving.

### Will AI send gear shoppers to REI instead of our own store?

Sometimes. ChatGPT ranks merchants on availability, price, quality and whether the seller is the maker. Accurate stock and price data on your own store helps, and retailer sales still count.

### How should we handle safety-critical gear in AI answers?

State certifications, ratings and intended use precisely, and never imply that a product makes a hazardous activity safe. Accurate, limited claims are less likely to be repeated wrongly.

### How quickly do new models show up in AI answers?

It varies. Assistants tended to cite recently published pages in our freshness study, but new products still need reviews and listings to exist before an answer can name them.

## Sources

- Outdoor Industry Association (2026), [2026 Outdoor Participation Trends Report: Executive Summary](https://oia.outdoorindustry.org/2026-participation-trends-report-summary)
- Shop Eat Surf Outdoor (July 2026), [183 Million Outdoor Participants and a Retail Playbook Hiding in the Numbers](https://shop-eat-surf-outdoor.com/?p=620899)
- US Bureau of Economic Analysis (November 2024), [Outdoor Recreation Satellite Account, U.S. and States, 2023](https://www.bea.gov/news/2024/outdoor-recreation-satellite-account-us-and-states-2023)
- Amer Sports (August 2026), [Amer Sports reports second quarter 2026 financial results](https://s21.q4cdn.com/769814866/files/doc_news/Amer-Sports-Reports-Second-Quarter-2026-Financial-Results-Raises-Full-Year-Revenue-Margin-and-EPS-Guidance-2026.pdf)
- GearJunkie (2026), [REI narrows 2025 losses by $102 million as union boycotts anniversary sale](https://gearjunkie.com/outdoor-business/rei-narrows-2025-losses-union-boycotts)
- Shopify (October 2026), [Welcome to the first holiday season of the agentic era](https://www.shopify.com/news/agentic-holiday-2026)
- Google (May 2025), [Shop with AI Mode, use AI to buy and try clothes on yourself virtually](https://blog.google/products/shopping/google-shopping-ai-mode-virtual-try-on-update/)
- OpenAI (2026), [Shopping with ChatGPT Search](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search)
- Underneath (2026), [How many “best of” lists cited by AI rank their own brand first?](https://underneath.agency/research/self-promoting-best-lists-study)
- Underneath (2026), [How fresh are the pages AI engines cite?](https://underneath.agency/research/ai-source-freshness-study)
- Underneath (2026), [When does a Reddit thread become evidence in Google’s AI?](https://underneath.agency/research/ai-reddit-citations-study)

---

This is the Markdown twin of https://underneath.agency/resources/outdoor-brands-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How packaging suppliers win quote requests through AI search"
description: "Brand owners now use AI in packaging sourcing. Packaging suppliers win quote requests when assistants can verify their materials, compliance and minimums."
canonical: "https://underneath.agency/resources/packaging-suppliers-quote-requests-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# When a brand asks AI for a packaging supplier, will our name come up?

It can, if an assistant can find and confirm what you make, from which materials, for which uses, at what minimums and to which rules. Brand owners already use AI in packaging sourcing, and most of them keep more than one supplier per format, so a place on the AI-built longlist is a real door to a quote request.

This guide is for custom and stock packaging suppliers: corrugated and folding carton converters, flexible packaging and label printers, rigid plastic and molded fiber makers, and distributors of stock packaging. The goal is not website traffic. It is a sample request, a quote and, in time, a share of a brand’s packaging spend.

## The short version

1. AI is already in packaging procurement: in [L.E.K. Consulting’s 2026 study](https://www.lek.com/en-gcc/node/14796) of 450 US brand owners, around 69% had adopted AI in procurement and sourcing, second only to product development.
2. Brands keep a second supplier: roughly 91% multisource their packaging, about 64% use a primary and a secondary supplier for a format, and around 27% use three or more, per [L.E.K.’s sourcing findings](https://www.lek.com/insights/industrials/brand-owners-have-shock-proofed-their-sourcing-strategies).
3. Almost every brand is about to change its packaging: roughly 99% expect changes over the next three years, led by sustainability, aesthetics and shelf life, according to the [same L.E.K. study](https://www.lek.com/insights/industrials/lek-consulting-2026-cpg-and-foodservice-brand-owner-packaging-study).
4. Rules are forcing redesigns: the EU’s packaging regulation [applies from 12 August 2026](https://circulareconomy.europa.eu/platform/en/news-and-events/all-news/eu-packaging-and-packaging-waste-regulation-when-will-it-come-effect-and-what-does-it-cover), and seven US states have packaging extended producer responsibility (EPR) laws, with Oregon’s fees ranging from $0.01 to $2.73 per pound across 60 materials, per [L.E.K.’s EPR analysis](https://www.lek.com/insights/paper-packaging/epr-inflection-point-understanding-impact-epr-regulations).
5. The market is large and shifting: [Smithers](https://www.smithers.com/Resources/2025/July/Tariffs-to-Reshape-Packaging-Market) projects US packaging consumption of $255.4 billion to $279.6 billion in 2030, depending on tariffs, while L.E.K. forecasts the share of US packaging sourced from abroad to fall to 10% by 2028.

## Who buys packaging, and what is a customer worth?

Brand managers, packaging engineers and procurement teams at consumer brands, and each win is a share of a format’s spend.

Packaging is bought by the companies that put products on shelves and doorsteps: food, beverage, beauty, household, health and electronics brands, plus restaurants and foodservice chains. Inside those companies, a brand manager cares about how the pack looks and sells, a packaging engineer cares about barrier, strength and line speed, and a procurement manager cares about cost, lead time and supply risk. Across business buying generally, [Forrester](https://www.forrester.com/press-newsroom/forrester-2026-the-state-of-business-buying/) finds procurement professionals are decision-makers in 53% of buying cycles.

The buyers care a great deal. In L.E.K.’s 2026 US study, 98% of respondents rated packaging as highly important to brand success, and micro and challenger brands put the most weight on it. That matters for suppliers that serve small brands with low minimums, because those buyers are often the ones searching without a procurement department.

What a customer is worth depends on the format and the volume, and no public benchmark gives an average account value. The shape of the prize is clearer. Among brands that use a primary and a secondary supplier, L.E.K. found about half give the primary supplier 65% to 75% of the spend for that format. Our inference: a supplier that wins a secondary slot has a realistic path to a larger share later, and one well-matched quote request can turn into years of repeat print and production runs.

The market behind those accounts is moving. Smithers expects global packaging demand to reach $1.52 trillion by 2030, and warns that rigid plastics and flexible packaging, which lean heavily on imports, face the sharpest cost increases from tariffs. Brand owners are responding by buying closer to home: L.E.K. forecasts that by 2028 the share of packaging sourced from outside the US will be just 10%, half of 2019 levels. Every one of those moves starts with a search for a new supplier.

## Where does AI already sit in packaging sourcing?

In supplier research and early comparison, where most brand owners already use it and plan to use more.

The packaging-specific evidence comes from L.E.K.’s eighth annual brand owner survey, run in late 2025 and early 2026:

| US brand owners, 2026 (L.E.K. Consulting) | Share |
|---|---|
| Aware of AI use in procurement and sourcing | about 82% |
| Have adopted AI in procurement and sourcing | about 69% |
| Have adopted AI in product development | about 74% |
| Use B2B customer portals with suppliers today | about 79% |
| Expect to use B2B customer portals within three years | about 94% |
| Use AI in customer-focused marketing, today and planned | about 52%, rising to 88% |

The survey does not say which AI tools brand owners use for sourcing, or whether they ask public assistants such as ChatGPT or Gemini for supplier names. It does show that AI is now part of how packaging buyers research and compare, and that large brands are investing in it: L.E.K. cites a generative AI tool that Nestlé R&D built with IBM Research to identify new high-barrier packaging materials.

Packaging engineers behave like other technical buyers. In the [2026 State of Marketing to Engineers research](https://advertising.globalspec.com/state-of-marketing-to-engineers-research-report) by GlobalSpec and TREW Marketing, which covers engineers across industries rather than packaging alone, 69% of technical buyers used generative AI during purchasing, they rated their trust in its answers 4.7 out of 10, and 62% of the buying process happened online before they contacted a vendor.

Industrial directories still matter too. Thomas’ 2025 sourcing report, [summarized by Distribution Strategy Group](https://distributionstrategy.com/2026/01/industrial-sourcing-behavior-shifts-in-2025-signaling-strategic-imperatives-for-distributors/), found food and beverage buyers concentrated on food ingredients, shrink and stretch packaging and contract manufacturing. The platform reports more than 1.5 million monthly sourcing sessions. Directory profiles are also pages an assistant can find and quote.

## What do packaging buyers ask AI assistants?

They describe a product, a material, a use and a rule, then ask who can make it at their volume.

None of the questions below were captured from real buyers; we wrote them to show the shape of a packaging sourcing request.

| What the brand needs | Sample sourcing question (our wording) |
|---|---|
| Material and application | “Recyclable mono-material pouch suppliers for whole-bean coffee in the US” |
| Low minimums | “Custom printed mailer boxes for a skin care startup, under 2,000 units” |
| Compliance | “PFAS-free molded fiber takeout containers for restaurants in Washington state” |
| Regional supply | “Corrugated box makers near Chicago with digital printing for short runs” |
| Redesign | “Alternatives to multilayer plastic film that still keep snacks fresh for six months” |
| Regulation | “Which packaging formats are likely to meet the EU recyclability rules?” |

Each question is a filter. If your site, directory listings and trade coverage never state your substrates, barrier options, food-contact status, print methods, minimum order quantities or plant locations, an assistant has nothing to match against.

Wording also changes the answer. In [our prompt phrasing study](https://underneath.agency/research/ai-prompt-phrasing-study), asking the same question again kept the same first-named brand 68.0% of the time, but adding “on a tight budget” kept it only 15.3%. Packaging buyers often add constraints like that, such as a low minimum, a deadline or a price ceiling. We infer that suppliers who state their minimums, lead times and price-relevant facts plainly give assistants more to work with when a question narrows.

## How does an AI answer become a quote request and an award?

Through a longlist, a website check, samples and quotes, line trials, and a share of the format’s spend.

1. **Longlisted.** A packaging engineer or buyer asks an assistant, a directory or a search engine which suppliers can make a format in a material at a volume. The answer is a starting list.
2. **Checked.** The buyer visits supplier websites to confirm materials, certifications, printing, capacity and location. Most of this happens before anyone calls.
3. **Sampled and quoted.** The supplier gets a request for samples and pricing, usually alongside rivals. Dielines, artwork and specifications move back and forth.
4. **Trialed.** The pack runs on the brand’s filling or packing line and goes through shipping, shelf-life or retailer tests.
5. **Awarded.** The supplier becomes primary or secondary for that format, with a share of the spend and repeat orders as long as quality, cost and supply hold.

An assistant’s answer can only touch the longlist and the website check. Price, quality, lead time and line performance decide the rest. Because roughly 91% of brand owners multisource, being named as a credible second option is often the realistic first win.

## What decides whether an assistant names a packaging supplier?

Facts it can find and confirm in other sources; the platforms document how they search, not how they pick suppliers.

What the platforms document: [OpenAI says](https://help.openai.com/en/articles/9237897-chatgpt-search) ChatGPT search typically rewrites a question into one or more targeted queries sent to search providers, and that a site must allow OpenAI’s search crawler, OAI-SearchBot, to be eligible for inclusion. [Google says](https://blog.google/products/search/ai-mode-search/) AI Mode uses a “query fan-out” technique that runs multiple related searches across subtopics and data sources. Neither explains how specific suppliers are chosen.

What has been observed in our studies, across several industries rather than packaging alone:

- **Several searches sit behind each answer.** In [our study of hidden searches](https://underneath.agency/research/ai-hidden-searches-study), ChatGPT ran a mean of 3.7 searches per buyer question before answering.
- **Outside mentions carry weight.** [Our brand entity study](https://underneath.agency/research/brand-entity-ai-recommendations-study) found that a brand named on ten times as many independent sites within the cited pages had 4.7 times the odds of being recommended.
- **Answers change between runs.** In [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), only 25.2% of the brands ChatGPT named for a question appeared in all five repeats.

Our inference for packaging: the trust factors are the ones a packaging engineer already checks, written where a machine can read them. That means materials and structures by name, barrier and food-contact information, recyclability and recycled-content claims with the evidence behind them, PFAS status, certifications such as FSC or BRCGS with the issuing body named, print methods, minimums, lead times and plant locations. Trade press features, packaging awards and case studies published with a customer’s permission give an assistant outside sources to confirm those claims.

## Why do sustainability rules make AI visibility more urgent?

Because new rules are pushing brands to redesign packs and, in many cases, to look for new suppliers now.

In Europe, the Packaging and Packaging Waste Regulation entered into force on 11 February 2025 and generally applies from 12 August 2026. The European Commission summary says all packaging must be recyclable and that recyclability performance grades apply from 2030. Brand owners are already reacting: in L.E.K.’s [European brand owner survey](https://www.lek.com/insights/paper-packaging/european-brand-owner-packaging-survey-2026-sustainability-balancing) of about 400 companies, the desire to move to sustainable packaging was the most-cited reason for changing primary packaging materials, at 43%, and 39% expected the sustainable share of their packaging to rise by 2030.

In the US, packaging EPR laws shift the cost of recycling from local governments to producers. [Mayer Brown](https://www.mayerbrown.com/en/insights/publications/2026/02/epr-packaging-laws-moving-from-concept-to-compliance) lists California, Colorado, Maine, Maryland, Minnesota, Oregon and Washington as states with EPR statutes, and notes that in Washington nonmembers cannot sell their products after March 2029. L.E.K. expects Oregon’s program to deploy more than $700 million in its first 2.5 years, with fees set by material so that harder-to-recycle materials pay more. The law firm also points out that packaging makers should expect EPR to change the types of packaging their customers buy.

For a supplier, every one of these rules creates questions a buyer may put to an assistant: which formats are recyclable, which materials carry lower fees, which suppliers can provide the data a producer must report. This article is not compliance advice, and buyers should confirm obligations with counsel. Our point is narrower: a supplier whose site explains, accurately and with sources, how its materials relate to these rules gives an assistant something to cite when the question comes up.

## What does a packaging supplier lose when AI leaves it out?

It loses quote requests it never sees, during a period when many brands are redesigning packs and adding suppliers.

We found no public measurement of quote requests lost to AI absence, so the reasoning is ours:

- **Redesign is near-universal.** With roughly 99% of US brand owners expecting packaging changes within three years, many will look beyond their current suppliers for materials or formats they do not yet buy.
- **Supply is moving home.** As imported packaging falls toward 10% of the total, brands need domestic suppliers they may never have used. A supplier absent from the first list is absent from the comparison.
- **Wrong facts filter you out.** An assistant that says you lack food-contact approval, cannot print digitally or have a minimum of 50,000 units, when you do not, removes you from a shortlist you never knew about. Our guide to [fixing wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) explains how to trace an error to the page it came from.

## How does GEO work for a packaging company?

Generative engine optimization (GEO) helps AI assistants find, describe and verify your materials, capabilities and compliance facts.

For a packaging supplier, the work usually includes:

1. **Material and application pages in plain text.** One page per structure or format, stating substrates, barrier levels, food-contact status, recyclability evidence, print methods, sizes, minimums and lead times. Spec sheets locked in PDFs or images are harder for assistants and buyers to use. Chemical makers face the same spec-sheet problem, covered in [how chemical suppliers reach formulators through AI](https://underneath.agency/resources/chemical-suppliers-b2b-buyers-ai-search).
2. **Compliance facts with sources.** Plain explanations of how your formats relate to state EPR programs and the EU regulation, with dates and links to the official texts, written as information rather than legal advice.
3. **Certifications that match the register.** Each certification named with its scope and issuing body, matching what the certifier publishes.
4. **One consistent identity.** The same company name, plant addresses, capabilities and minimums on your site, Thomasnet and other directories, LinkedIn and trade association listings.
5. **Independent coverage.** Packaging trade press features, award entries, conference talks and customer case studies published with permission. Our article on [how small brands get recommended by AI](https://underneath.agency/resources/how-small-brands-get-recommended-by-ai) explains why those outside sources can outweigh a supplier’s own pages.
6. **Honest comparison content.** Clear answers to “paper or plastic for this use,” “mono-material or multilayer,” and “domestic or imported,” without attacking rivals. Our review of [whether comparison pages help B2B brands get cited](https://underneath.agency/resources/do-comparison-pages-help-b2b-ai-citations) covers what such pages can and cannot do.
7. **Crawl access and measurement.** Allow the search crawlers the assistants document, then ask a fixed set of material, application and compliance questions across ChatGPT, Gemini, Perplexity, Copilot and Google’s AI features, repeatedly, and compare who is named with the quote requests you receive.

No one can promise that an assistant will name a supplier. The work makes your company the easiest one for an assistant, and then a packaging engineer, to check.

## Which questions about packaging and AI search remain open?

The data shows brand owners using AI in sourcing, not how many quote requests or awards AI answers produce.

- **Adoption is not attribution.** L.E.K. measures AI use in procurement and sourcing, not whether public assistants named the suppliers that won. We found no public data linking AI answers to packaging quote volume.
- **Tools are unspecified.** “AI in sourcing” may mean internal analytics, supplier platforms or public assistants. The surveys do not separate them.
- **Cross-industry studies.** The engineering survey and our own studies cover many industries. Applying them to packaging is our inference.
- **Rules are still moving.** EPR timelines can slip, as Mayer Brown notes for California, and Oregon’s program faces federal litigation, according to L.E.K. Claims about compliance need regular review.

Contract manufacturers face the same longlist problem from the production side, covered in our guide on [how contract manufacturers win RFQs through AI search](https://underneath.agency/resources/contract-manufacturers-rfqs-from-ai-search). For the warehousing and fulfillment side, see [how 3PLs make a shipper’s AI shortlist](https://underneath.agency/resources/logistics-providers-shipper-contracts-ai-search).

## Where should a packaging supplier start?

Start by asking assistants the material, application and compliance questions your best customers would ask.

That first check shows whether your company is named for your core formats and materials, whether your minimums, plants and certifications are described correctly, which directories and articles the answers rely on, and which suppliers appear in your place. The work then is to publish those facts where assistants read them and to earn the outside coverage that confirms them.

If your growth depends on winning new formats and secondary-supplier slots with brand owners, [talk to us about a packaging visibility review](https://underneath.agency/contact). We will show how assistants answer the sourcing questions your buyers ask, why other suppliers are named, and which changes are most likely to bring more qualified quote requests. How we handle the ongoing material, compliance and directory work for converters and distributors is set out on our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page.

## Frequently asked questions

### Do brand owners really use AI to choose packaging suppliers?

They use AI in sourcing: about 69% of US brand owners in L.E.K.’s 2026 study had adopted AI in procurement and sourcing. The study does not show how often public assistants name the supplier that wins.

### Is a stock packaging distributor in the same position as a custom converter?

Not quite. Stock buyers often ask about price, availability and delivery speed, while custom buyers ask about materials, printing and minimums. Both need those facts stated in plain text, consistently, across their own site and directories.

### Should we publish our sustainability claims for AI to read?

Publish only claims you can support, with the evidence and the standard named. Assistants can repeat what they find, so an unsupported recyclability claim can spread errors and create regulatory risk.

### Will AI search replace packaging sales reps?

No. Samples, artwork, line trials and pricing still need people. AI can change which suppliers a buyer contacts first.

## Sources

- L.E.K. Consulting (2026-04-13), [L.E.K. Consulting 2026 CPG and Foodservice Brand Owner Packaging Study](https://www.lek.com/insights/industrials/lek-consulting-2026-cpg-and-foodservice-brand-owner-packaging-study)
- L.E.K. Consulting (2026-04-13), [Brand Owners Are Embracing Digital and AI in Packaging](https://www.lek.com/en-gcc/node/14796)
- L.E.K. Consulting (2026-04-13), [Brand Owners Have Shock-Proofed Their Sourcing Strategies](https://www.lek.com/insights/industrials/brand-owners-have-shock-proofed-their-sourcing-strategies)
- L.E.K. Consulting (2026-06-08), [European Brand Owner Packaging Survey 2026: Sustainability](https://www.lek.com/insights/paper-packaging/european-brand-owner-packaging-survey-2026-sustainability-balancing)
- L.E.K. Consulting (2026-07-02), [The EPR Inflection Point: Understanding the Impact of EPR Regulations](https://www.lek.com/insights/paper-packaging/epr-inflection-point-understanding-impact-epr-regulations)
- European Circular Economy Stakeholder Platform (2026-04-06), [The EU Packaging and Packaging Waste Regulation: when will it come into effect and what does it cover?](https://circulareconomy.europa.eu/platform/en/news-and-events/all-news/eu-packaging-and-packaging-waste-regulation-when-will-it-come-effect-and-what-does-it-cover)
- Mayer Brown (2026-02-03), [EPR Packaging Laws Moving from Concept to Compliance](https://www.mayerbrown.com/en/insights/publications/2026/02/epr-packaging-laws-moving-from-concept-to-compliance)
- Smithers (2025-07), [Smithers Forecasts Tariffs to Reshape $1.5 Trillion Global Packaging Market by 2030](https://www.smithers.com/Resources/2025/July/Tariffs-to-Reshape-Packaging-Market)
- Distribution Strategy Group (2026-01-30), [Industrial Sourcing Behavior Shifts in 2025 (Thomas 2025 Annual Sourcing Activity Report)](https://distributionstrategy.com/2026/01/industrial-sourcing-behavior-shifts-in-2025-signaling-strategic-imperatives-for-distributors/)
- GlobalSpec and TREW Marketing (2026), [State of Marketing to Engineers research report](https://advertising.globalspec.com/state-of-marketing-to-engineers-research-report)
- Forrester (2026-01-21), [Forrester’s 2026 Buyer Insights: GenAI Is Upending B2B Buying](https://www.forrester.com/press-newsroom/forrester-2026-the-state-of-business-buying/)
- OpenAI Help Center (2026), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Google (2025-03-05), [Expanding AI Overviews and introducing AI Mode](https://blog.google/products/search/ai-mode-search/)
- Underneath (2026), [Does rewording a question change AI brand recommendations?](https://underneath.agency/research/ai-prompt-phrasing-study)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [Do Wikipedia and schema make AI assistants recommend a brand?](https://underneath.agency/research/brand-entity-ai-recommendations-study)
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)

---

This is the Markdown twin of https://underneath.agency/resources/packaging-suppliers-quote-requests-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How payment companies win small merchants in AI search"
description: "By being named, with the right fees, when a new business asks AI how to take payments. Owners choose on cost and simplicity, and rarely switch later."
canonical: "https://underneath.agency/resources/payment-companies-merchants-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# When a new business asks AI how to take payments, how does a payment company get chosen?

By being the provider the answer names, with fees and features stated correctly, at the moment the business opens its doors. Small merchants choose a payment company mostly on cost and ease, often at founding, and tend to stay. An AI answer at that moment can decide years of card volume.

## The short version

1. The founding moment is the market: in a [2026 Sonary survey](https://sonary.com/content/what-do-small-businesses-actually-want-from-payment-processors/) of 3,603 small businesses shopping for payment processing, 44.1% were just starting their business.
2. Cost decides the rest: among businesses with an existing setup, 35.2% named high fees as their biggest problem, and [Javelin](https://javelinstrategy.com/node/34611) found that lowering total cost remained the top reason for choosing a primary card processor.
3. Volume compounds: Square processed $72.8 billion in the second quarter of 2026, up 13%, according to [SiliconANGLE](https://siliconangle.com/2026/08/05/block-shares-slip-despite-second-quarter-beat-raised-2026-guidance/), and businesses on Stripe generated $1.9 trillion in 2025, up 34%.
4. Fees are easy to misquote: in [our pricing study](https://underneath.agency/research/ai-pricing-accuracy-study), 61.9% of 840 software plan prices quoted by four AI assistants were fully faithful to the official page.
5. Owners already use the tools: the [U.S. Chamber of Commerce](https://www.uschamber.com/technology/empowering-small-business-the-impact-of-technology-on-u-s-small-business) reports that 58% of small businesses say they use generative AI.

## Who chooses a payment company, and what is one merchant worth?

Usually the owner, often in the first weeks of the business, and the merchant is worth years of card volume.

A bakery, a mobile dog groomer, a yoga studio, a two-person online shop: the buyer for an all-in-one payment company is rarely a procurement team. It is the owner, deciding quickly, often alone, and often before the first sale. Sonary’s survey found that 44.1% of payment shoppers were just starting out, the single most common answer. Sonary is a comparison site that earns commissions from featured providers, so treat its survey as industry research with a commercial interest, not an independent census.

What that merchant is worth comes from processing volume, not a contract. Payment companies earn a share of every card payment, plus software, hardware and add-ons such as invoicing or loans. The public fee schedules show the scale of the take:

- [Square](https://squareup.com/us/en/payments/our-fees) lists in-person card payments at 2.4% to 2.6% + 15¢ depending on plan, and online and invoice payments at 2.9% or 3.3% + 30¢.
- [Stripe](https://stripe.com/pricing) lists 2.9% + 30¢ per successful transaction for domestic cards.
- [PayPal](https://www.paypal.com/us/business/paypal-business-fees) lists 3.49% + a fixed fee for PayPal Checkout and 2.99% + a fixed fee for standard card payments.

Small percentages on a growing business add up over years, and the choice tends to stick. Javelin reports that small businesses “widely express satisfaction” with their primary card provider, and that merchants selling both in stores and online were less likely to switch than a year earlier. Our inference: the cheapest merchant to win is the one choosing for the first time, and the hardest to win is one already settled with a rival.

The leaders show what winning early looks like. Stripe says businesses on its platform generated $1.9 trillion in total volume in 2025, that it serves more than 5 million businesses directly or through platforms, and that 25% of all Delaware corporations are now created with Stripe Atlas, its incorporation service. In Block’s second quarter, Square’s volume growth in the US accelerated to 10% while international volume grew 28%.

## Where does AI already sit in a small merchant’s decision?

In the research before sign-up, where owners already use AI for daily work, and increasingly in how they sell.

There is no public survey we could find on how many owners ask an AI assistant to pick a payment provider. What is documented is that the tools are part of daily work. The U.S. Chamber of Commerce’s 2025 report found that 58% of small businesses say they use generative AI, and it describes AI use as more than double the 2023 level.

AI also touches payments from the selling side. In September 2025, [OpenAI](https://openai.com/index/buy-it-in-chatgpt/) launched Instant Checkout in ChatGPT, “powered by the Agentic Commerce Protocol, built with Stripe,” and said more than 700 million people used ChatGPT each week. Merchants paid a small fee on completed purchases. The feature’s status has since changed: [Reuters reported](https://www.zawya.com/en/business/americas/retailers-tap-ai-shopping-traffic-but-fight-to-keep-customer-data-424957) that OpenAI ended Instant Checkout in March 2026 to focus on product discovery and merchants’ own checkouts, [FashionUnited](https://fashionunited.com/news/retail/us-consumers-would-let-ai-agents-buy-clothes-but-not-without-a-say/2026092974919) described it as scaled back, and [OpenAI’s help page](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search) still says ChatGPT may show an Instant Checkout option “for some eligible products and merchants.”

The detail matters less than the direction. [Stripe’s 2025 letter](https://stripe.com/newsroom/news/stripe-2025-update) describes an Agentic Commerce Suite for selling across several AI interfaces and work with Microsoft on Copilot. A reasonable expectation is that owners will add a new question to their checklist: “can I sell through AI assistants with this provider?” Sellers weighing a marketplace instead of their own store face a related choice, covered in [how marketplaces win sellers through AI](https://underneath.agency/resources/marketplace-buyers-sellers-ai-search).

## What do small business owners ask AI about payments?

How to start, what it costs, what fits their trade, and whether a provider can be trusted. We wrote the sample prompts below to show how merchants tend to phrase these questions; they are not recorded queries.

| Stage | Illustrative prompt |
|---|---|
| Getting started | “How do I take card payments at a farmers market stall?” |
| Cost | “Cheapest way to accept payments if I sell about $4,000 a month” |
| Comparison | “Square vs Stripe vs PayPal for a mobile dog groomer” |
| Fit | “Best payment app for a hair salon with bookings and tips” |
| Integration | “Which payment provider works with QuickBooks and my booking tool?” |
| Trust | “Why do payment apps freeze accounts, and which one is safest?” |
| Selling channels | “Which payment provider lets me sell through ChatGPT?” |

These questions mirror what owners say they need. In Sonary’s data, established businesses named high fees (35.2%) more than slow settlement (12.7%) and outdated equipment (8.1%) combined. Javelin also tracks which payment types small businesses accept, and found in-store buy now, pay later acceptance fell from 46% to 32% of respondents in a year. Owners ask about the methods their customers use, so an answer about “which provider accepts Apple Pay and pay-later” is a selection question too.

## How does an AI answer turn into processed volume?

Through self-serve sign-up: the answer names a provider, the owner signs up, and card volume follows for years.

Most small merchants never speak to a salesperson. The path, as we infer it from how these products are sold:

1. **Question.** “How should a new food truck take cards?”
2. **Answer.** The assistant names two or three providers and summarizes fees and hardware.
3. **Check.** The owner opens a pricing page, perhaps a review site.
4. **Sign-up.** Self-serve onboarding, often within minutes.
5. **Volume.** Every card payment for as long as the business stays.
6. **Attach.** Point-of-sale software, invoicing, payroll or financing.

Two facts make step 2 unusually valuable. The merchant often has no incumbent to compare against, and Javelin’s data suggests that once satisfied, small businesses rarely move. A provider missing from the answer at the founding moment may not get a second look until something goes wrong.

Our [guide to linking AI answers to pipeline and revenue](https://underneath.agency/resources/ai-answers-pipeline-revenue) covers how to measure that path; for payment companies the useful unit is new accounts and their first-year volume, not traffic.

## What decides whether an assistant names your payment company?

Platforms document little about provider recommendations; studies point to independent coverage, consistent facts and reputation evidence.

**Documented by platforms.** According to Google, its AI features can handle one merchant question with [“query fan-out”](https://developers.google.com/search/docs/appearance/ai-features), running several related searches on subtopics before writing the answer. A question such as “Square vs Stripe for a salon” can therefore pull in pages about fees, salon software and reviews separately. For its shopping results, OpenAI’s help page says Instant Checkout items “are not preferred in product results.” OpenAI scaled the feature back in March 2026, according to FashionUnited, though that help page still describes it for some eligible merchants. We found no platform documentation on how assistants choose payment providers.

**Observed in studies.**

- **Independent coverage.** Trade press and comparison sites count, not only your own pages: in [our brand study](https://underneath.agency/research/brand-entity-ai-recommendations-study), every tenfold rise in outside sites naming a brand within the cited pages came with 4.7 times the odds of a recommendation.
- **Reputation sources.** An owner asking whether a provider is safe will be shown review sites: in [our “is this brand legit?” study](https://underneath.agency/research/is-it-legit-ai-reputation-study), 88.0% of answers cited a review or complaint platform, 61.7% of review-platform citations went to Trustpilot and the BBB, and 99.7% of answers raised at least one problem.
- **Price fidelity.** In our pricing study, when an assistant’s price differed from the official page, 39 of 64 differing figures appeared on another page of the vendor’s own site.
- **Self-serving lists.** [Our study of “best of” lists](https://underneath.agency/research/self-promoting-best-lists-study) found that 24.2% of 269 AI-cited numbered lists ranked their own publisher first. It is worth checking who wrote the lists cited for your own category.

**Our inference for payment companies.** Fee schedules with plan tiers, card-present and online rates, and fixed fees are exactly the kind of detail an answer can flatten. Old pricing pages, help articles and partner pages that state outdated rates are a likely source of wrong quotes. And the reputation questions that matter most in payments, frozen funds, account holds and reserves, are the ones review platforms carry. Sonary’s survey also advised owners to ask about “holds or rolling reserves on new accounts” before signing.

The pricing study covered software subscriptions, not payment fees, and the reputation study did not include payment companies, so these are expectations to test, not measured facts about this industry.

## What does a payment company lose when AI gets its fees or fit wrong?

The founding-moment merchant, who signs up elsewhere the same day and may never revisit the choice.

We found no public data on how many merchants a payment company loses to inaccurate AI answers. The exposure follows from the facts above: owners choose on cost, many choose once, and assistants faithfully quoted 61.9% of plan prices in our software sample. An answer that quotes a provider’s highest tier as its standard rate, or omits a no-monthly-fee plan, removes it from a cost-driven shortlist without anyone noticing. Our guide on [how to fix wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) explains the correction process.

## What does GEO look like for a payment company?

Generative engine optimization (GEO) here means making every fact an owner checks easy to find, consistent and confirmed.

1. **One fee story everywhere.** Pricing page, help center, partner listings and old announcements should state the same current rates and conditions. Retire or update stale pages.
2. **Answers by trade.** Plain pages for the businesses you serve (salons, food trucks, contractors, online shops): what they need, what it costs, which hardware and which integrations.
3. **Integration facts.** Which accounting, booking and ecommerce tools you work with, stated on your site and in those partners’ directories, since an owner’s question often starts from a tool they already use. Online sellers ask a parallel question about delivery, covered in [how shipping platforms win small sellers](https://underneath.agency/resources/shipping-companies-customers-ai-search).
4. **Comparison and review presence.** Accurate listings on independent comparison sites and small-business publications; for how “best of” lists feed answers, see [what best-of lists mean for AI recommendations](https://underneath.agency/resources/best-of-lists-ai-recommendations).
5. **Reputation in the open.** Explain holds, reserves and payout timing plainly, and answer complaints on Trustpilot and the BBB, where assistants find them.
6. **Selling-channel readiness.** If you support selling through AI assistants, say exactly what, for whom and where, without overstating features that are changing.
7. **Measured visibility.** Track the founding-moment questions by trade and by market; [how many prompts to track](https://underneath.agency/resources/how-many-prompts-to-track-ai-visibility) explains how to size the set.

The neighboring guides cover adjacent buyers: [ecommerce platforms](https://underneath.agency/resources/ecommerce-platform-ai-search) choosing partners, and [fintech software sold to banks and finance teams](https://underneath.agency/resources/fintech-software-customers-from-ai-search). Larger merchants with their own payments teams are covered in [how processors reach payments teams through AI](https://underneath.agency/resources/payment-processors-merchant-demand-ai-search). No one can guarantee that an assistant will recommend a payment company; GEO makes the evidence it finds accurate, current and easy to verify.

## What is still unknown about AI and payment-provider choice?

How often owners ask AI before signing up, and how much of a provider’s sign-ups those answers drive.

- **No direct survey.** We found no public data on the share of small businesses that use an AI assistant to choose a payment provider; the Chamber figure covers AI use in general.
- **Vendor and commercial research.** Sonary earns commissions and Javelin’s full report is paid; the figures here are from their public pages.
- **Studies outside payments.** Our pricing, reputation and brand studies did not test payment companies specifically.
- **Changing products.** Selling inside AI assistants is in flux, as the different accounts of Instant Checkout show.

## Where should a payment company start?

Start by asking assistants the founding-moment questions for each trade you serve, and checking the fees they quote.

That first check shows which providers are named for a salon, a food truck or an online shop, whether your rates and plan conditions come back correctly, and which comparison, review and partner pages the answers draw on. It also shows whether account holds and payout times are described fairly.

If new merchant sign-ups and first-year processing volume are what you need to grow, [talk to us about a merchant-question review](https://underneath.agency/contact). We will test how assistants answer the questions owners ask before they pick a provider, find where your fees or fit are described wrongly, and plan the content, coverage and reputation work that gives those answers better evidence. Keeping one consistent fee story and building pages for each trade you serve are part of the ongoing work described on our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page.

## Frequently asked questions

### Do small business owners really ask ChatGPT which payment provider to use?

We found no survey measuring that directly. What is documented is broad use: 58% of small businesses told the U.S. Chamber of Commerce they use generative AI.

### Why does AI sometimes quote the wrong processing fee?

Fee schedules have tiers, card-present and online rates, and fixed fees. In our software pricing study, most differing prices also appeared on another page of the vendor’s own site.

### Does being in AI shopping features help a payment company win merchants?

It can become a selection factor, but the products are changing. Reuters reported that OpenAI ended Instant Checkout in March 2026 and FashionUnited described it as scaled back, while OpenAI’s help page still describes it for some eligible merchants.

### Can a payment company pay to be recommended by AI assistants?

Not through organic answers. OpenAI’s help page says Instant Checkout items are not preferred in product results; recommendations draw on the evidence assistants find.

## Sources

- Sonary (2026-07-14), [What do small businesses actually want from payment processors and POS systems?](https://sonary.com/content/what-do-small-businesses-actually-want-from-payment-processors/)
- Javelin Strategy & Research (2024-11-20), [2024 Small Business PaymentsInsights: U.S.: Payment Acceptance Services & Card Payment Processing](https://javelinstrategy.com/node/34611)
- U.S. Chamber of Commerce (2025-08-18), [Empowering Small Business: The Impact of Technology on U.S. Small Business](https://www.uschamber.com/technology/empowering-small-business-the-impact-of-technology-on-u-s-small-business)
- SiliconANGLE (2026-08-05), [Block shares slip despite second-quarter beat and raised 2026 guidance](https://siliconangle.com/2026/08/05/block-shares-slip-despite-second-quarter-beat-raised-2026-guidance/)
- Stripe (2026-02), [Stripe publishes 2025 annual letter](https://stripe.com/newsroom/news/stripe-2025-update)
- Square (2026), [Understanding our fees](https://squareup.com/us/en/payments/our-fees)
- Stripe (2026), [Pricing](https://stripe.com/pricing)
- PayPal (2026), [Merchant fees](https://www.paypal.com/us/business/paypal-business-fees)
- OpenAI (2025-09-29), [Buy it in ChatGPT: Instant Checkout and the Agentic Commerce Protocol](https://openai.com/index/buy-it-in-chatgpt/)
- Reuters, via Zawya (2026), [Retailers tap AI shopping traffic but fight to keep customer data](https://www.zawya.com/en/business/americas/retailers-tap-ai-shopping-traffic-but-fight-to-keep-customer-data-424957)
- FashionUnited (2026-09-29), [US consumers would let AI agents buy clothes, but not without a say](https://fashionunited.com/news/retail/us-consumers-would-let-ai-agents-buy-clothes-but-not-without-a-say/2026092974919)
- OpenAI (2026), [Shopping in ChatGPT search](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- Underneath (2026), [How faithfully do AI assistants quote software prices?](https://underneath.agency/research/ai-pricing-accuracy-study)
- Underneath (2026), [“Is this brand legit?” How AI assistants build a reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study)
- Underneath (2026), [Do Wikipedia and schema make AI assistants recommend a brand?](https://underneath.agency/research/brand-entity-ai-recommendations-study)
- Underneath (2026), [How many “best of” lists cited by AI rank their own brand first?](https://underneath.agency/research/self-promoting-best-lists-study)

---

This is the Markdown twin of https://underneath.agency/resources/payment-companies-merchants-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How payment processors reach merchant shortlists through AI"
description: "By being named and described accurately when payments teams, engineers and procurement research processors with AI, from pricing models to local acquiring."
canonical: "https://underneath.agency/resources/payment-processors-merchant-demand-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How does a payment processor get onto the shortlist when merchants research with AI?

By making its pricing model, market coverage, performance and integration facts easy for AI assistants to find, verify and repeat. High-volume merchants and platforms choose processors through long evaluations, but the long list now forms partly in AI answers, and the engineers who integrate payments increasingly work through AI tools. A processor described vaguely or wrongly at that stage may never receive the request for proposal.

## The short version

1. A few merchants drive the business: [Adyen](https://www.financemagnates.com/fintech/payments/adyen-lifts-2026-revenue-outlook-to-2123-as-volume-hits-804-billion/) processed €803.8 billion in the first half of 2026, up 24%, and said 300 merchants accounted for roughly 60% of its growth.
2. Enterprise merchants spread volume: in the [2025 Global eCommerce Payments & Fraud Report](https://www.visaacceptance.com/content/dam/documents/campaign/fraud-report/global-fraud-report-2025.pdf) from Visa Acceptance Solutions and the Merchant Risk Council, merchants used 3 to 4 payment gateways and acquiring banks.
3. Pricing models need explaining: the [Federal Reserve](https://www.federalreserve.gov/paymentsystems/regii-average-interchange-fee.htm) caps covered debit interchange at $0.21 plus 0.05% of the transaction, plus a $0.01 fraud-prevention adjustment, while credit interchange varies by card and network.
4. Answers change by market: in [our country study](https://underneath.agency/research/ai-recommendations-by-country-study), the country’s share of the variation in brand lists was 0.434 for financial services and insurance, the highest of any industry measured.
5. Integration work runs through AI tools: [Stack Overflow’s 2025 survey](https://stackoverflow.co/company/press/archive/stack-overflow-2025-developer-survey/) found 84% of developers use or plan to use AI tools in development.

## Who buys payment processing at scale, and what is one merchant worth?

A group of payments, finance, engineering and procurement leaders; one large merchant can move a processor’s growth.

For an enterprise retailer, a subscription business, a travel company or a software platform that embeds payments, choosing a processor is a multi-year infrastructure decision. The head of payments owns authorization rates and cost; the chief financial officer owns the fee line and treasury; engineering owns the integration; procurement runs the process. Across industries, [Forrester](https://forrester.com/blogs/state-of-business-buying-2026) found procurement professionals serve as decision-makers 53% of the time in the average business buying cycle.

The value of each win is unusually concentrated. Adyen reported net revenue of €1,302.9 million in the first half of 2026, up 19%, and said about 70% of its growth came from existing customers. That last figure is the shape of the business: processors often win a share of a merchant’s volume first and grow it later. Platforms are the fastest-growing segment: Adyen’s Platforms unit processed €135.0 billion, up 42%.

Stripe shows the same reach at the top of the market. Its [2025 annual letter](https://stripe.com/newsroom/news/stripe-2025-update) says it serves 90% of the Dow Jones Industrial Average, and that businesses on Stripe generated $1.9 trillion in volume. For processors, the commercial question is less “how many sign-ups?” and more “are we invited into the next large evaluation?”

## Where does AI already sit in a processor evaluation?

In the early research, the questions engineers ask while scoping integration, and the drafting of evaluation criteria.

We found no public survey on how many payments leaders use AI assistants to build a processor long list. Across business purchases, Forrester reports that 94% of buyers use AI during their buying process and that they then validate the output with peers, analysts and experts. That is a cross-industry figure, but it describes the role AI plays before a request for proposal: breadth first, validation later.

Engineers are a second route in. Stack Overflow’s 2025 survey found that 84% of developers use or plan to use AI tools in their development process, while 46% said they do not trust the accuracy of the output. Processors have noticed. [Stripe’s documentation](https://docs.stripe.com/building-with-llms) offers each page as Markdown, a button to copy it for AI tools, and a server that gives coding agents “up-to-date technical guidance.” When an engineer asks a coding assistant how hard it would be to add a second processor, the answer draws on whatever documentation it can read.

Agentic commerce is a third. [Financial IT](https://financialit.net/news/infrastructure/adyen-publishes-h1-2026-financial-results) reports that Adyen launched Adyen Agentic, “enabling enterprise merchants to securely process payments across AI agent protocols,” and Stripe describes its work with OpenAI on the Agentic Commerce Protocol. A reasonable expectation is that support for these protocols will start to appear as a line in evaluation criteria.

## What do payments leaders and engineers ask AI about processors?

Pricing models, market coverage, performance, redundancy and integration effort. The sample prompts that follow are our own wording of typical buyer and engineer questions, not logged queries.

| Stage | Illustrative prompt |
|---|---|
| Pricing model | “Interchange-plus vs blended pricing for $80 million a year in card sales” |
| Coverage | “Which processors offer local acquiring in Brazil, Japan and the EU?” |
| Comparison | “Adyen vs Stripe vs Checkout.com for a marketplace paying out sellers” |
| Redundancy | “How do we add a second acquirer without rebuilding checkout?” |
| Performance | “How can we raise card authorization rates on recurring payments?” |
| Process | “What should a payment processing RFP include?” |
| Integration | “How long does a migration from our current gateway usually take?” |

These map onto what merchants say they measure. The Visa and Merchant Risk Council report, based on 1,082 merchant professionals in 38 countries, found six metrics rated extremely important by more than 4 in 10 merchants: revenue, success rate, loss rates, authentication rate, authorization rate and cost of payments. A processor that publishes clear, checkable material on each gives an AI answer something to cite.

The pricing question deserves special care. Interchange is set by card networks and, for regulated US debit, capped by the Federal Reserve’s Regulation II; the processor’s markup sits on top. Public price lists show both models: [Adyen](https://www.adyen.com/pricing) charges a fixed processing fee of $0.13 per transaction plus a payment-method fee, with one card line priced as “$0.13+Interchange+ + 0.60%,” while [Stripe](https://stripe.com/pricing) offers IC+ pricing in custom packages for businesses with large volume. An answer that collapses these into one headline rate misleads the buyer.

## How does an AI answer become processed volume?

By shaping the long list, which decides who receives the request for proposal, then a share of volume that grows.

The path, as we infer it from how processors are bought:

1. **Framing.** A payments lead asks an assistant to explain pricing models or list processors for a set of markets.
2. **Long list.** Analysts, peers and AI answers produce the names worth inviting.
3. **Request for proposal.** Procurement invites a short list and scores responses.
4. **Technical proof.** Engineers test documentation, sandboxes and migration effort.
5. **Pilot share.** The merchant routes part of its volume to the new processor; multi-acquiring makes this common.
6. **Growth.** Share rises as performance proves out.

Steps 1 and 2 are where AI search can matter, and they are the hardest to see. A processor missing from the long list never reaches step 3, and nothing in its pipeline reports shows the loss. Our guide to [how AI assistants shape enterprise software shortlists](https://underneath.agency/resources/enterprise-software-shortlists-ai-search) covers the same pattern in a neighboring market, and [linking AI answers to pipeline](https://underneath.agency/resources/ai-answers-pipeline-revenue) explains how to measure it.

## What decides whether an assistant names a processor for a merchant’s needs?

Platforms document little about vendor choice; studies point to market-specific sources, independent coverage and accurate pricing facts.

**Documented by platforms.** Google says its AI features may use [“query fan-out”](https://developers.google.com/search/docs/appearance/ai-features), issuing several related searches across subtopics. A question about processing in Brazil, Japan and the EU may therefore pull in separate sources for each market. We found no platform documentation on how assistants choose payment processors.

**Observed in studies.**

- **Market matters.** In our country study, the gap between same-country and different-country answers on ChatGPT was 0.254 for high-dependence questions such as insurance, lending and tax, against 0.053 for global products and software. About half (52.7%) of ChatGPT’s same-country advantage went with differences in the websites it cited.
- **Independent coverage.** Analysts and payments press carry weight: in [our brand study](https://underneath.agency/research/brand-entity-ai-recommendations-study), a tenfold rise in independent sites mentioning a brand across the cited pages was linked to 4.7 times the odds that assistants recommended it.
- **Pricing conditions get lost.** In [our pricing study](https://underneath.agency/research/ai-pricing-accuracy-study) of software plans, 61.9% of 840 quoted plan prices were fully faithful, and another 3.8% had the right amount but dropped a condition that changes what a buyer pays.

**Our inference for processors.** Coverage claims should be specific and verifiable: local acquiring licenses, domestic clearing access and supported payment methods by country. Adyen’s own results name such facts, including direct access to France’s domestic interbank clearing system and a license from the Central Bank of the UAE. Performance claims need a stated method, because buyers and assistants both need something to check. Pricing pages should say plainly which model applies to whom.

## What does a processor lose when it is missing from AI answers?

Invitations to evaluations it never hears about, in a market where a few merchants drive most of the growth.

We found no public data measuring lost requests for proposal due to AI answers. The exposure follows from concentration: when 300 merchants account for roughly 60% of a leading processor’s growth, a single missed evaluation can matter more than thousands of small sign-ups. And because merchants typically spread volume across several gateways and acquirers, being absent from the next evaluation means losing the chance at a share, not just one deal.

## What does GEO look like for a payment processor?

Generative engine optimization (GEO) here means making pricing, coverage, performance and integration facts findable and confirmed.

1. **Pricing-model explainers.** Plain pages on interchange-plus, blended and custom pricing, who each suits, and a worked example, consistent with sales materials.
2. **Coverage by market.** A current table of countries, local acquiring, settlement currencies and payment methods, with the licenses and clearing access that back it.
3. **Readable documentation.** Public, current integration guides in formats AI tools can read; our guide on [API companies in AI search](https://underneath.agency/resources/api-companies-ai-search) covers this in depth.
4. **Performance with method.** Authorization, fraud and cost results stated with how they were measured, never as unqualified promises.
5. **Independent coverage.** Analyst reports, payments trade press, merchant case studies with named customers, and conference talks by merchants.
6. **Evaluation-ready facts.** Public answers to the questions procurement asks: security certifications, uptime reporting, data residency and migration support.
7. **Visibility by market.** Track the evaluation questions in each region you sell into, because answers differ by country.

For adjacent buyers, see our guides on [fintech software sold to banks and finance teams](https://underneath.agency/resources/fintech-software-customers-from-ai-search) and [ecommerce platforms](https://underneath.agency/resources/ecommerce-platform-ai-search). Small merchants choosing a provider at founding are covered in [how payment companies win new businesses](https://underneath.agency/resources/payment-companies-merchants-ai-search). No one can guarantee a processor appears in AI answers; GEO makes the evidence they draw on accurate, current and specific.

## Which questions about AI and processor selection remain open?

How often payments leaders consult AI before a request for proposal, and whether being named changes who gets invited.

- **No payments-specific survey.** The 94% figure covers business buyers in general, not payments teams.
- **Developer trust is mixed.** Many developers use AI tools but distrust their accuracy, so the weight engineers give to AI answers is uncertain.
- **Studies outside payments.** Our country, brand and pricing studies did not test processors specifically.
- **Agentic commerce is early.** Protocol support is new, and we found no data on how often it decides an evaluation.

## Where should a payment processor start?

Start by running the questions your next evaluation will raise, by market, and checking how assistants describe you.

That first check shows which processors are named for each region and use case, whether your pricing model and coverage are described correctly, and which analyst, press and documentation pages the answers rely on. It also shows what an engineer’s coding assistant says about integrating with you.

If more invitations to enterprise and platform evaluations are the goal, [ask us to review your evaluation-stage visibility](https://underneath.agency/contact). We will test the pricing, coverage, comparison and integration questions buyers and engineers ask, find where answers miss or misstate you, and plan the documentation, coverage and content work that gives assistants accurate evidence. The way we run it, market by market for payments leads, engineers and procurement, is described on our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page.

## Frequently asked questions

### Do payments teams really use AI assistants to compare processors?

No payments-specific survey exists that we found. Across industries, Forrester reports 94% of business buyers use AI during buying, then validate with peers and analysts.

### Why do AI answers get processing fees wrong?

Processor pricing mixes interchange, network fees and markups, often on custom terms. In our software pricing study, some answers kept the amount but dropped conditions that change the price.

### Does our developer documentation affect AI visibility?

Plausibly, for integration questions. Stripe offers its documentation as Markdown and through a server for coding agents; our llms.txt study found 11.5% of top websites publish an llms.txt file.

### Should a processor publish its prices if most deals are custom?

Publishing the pricing model and how it works helps answers stay accurate even when final rates are negotiated. Adyen and Stripe both publish their models.

## Sources

- Finance Magnates (2026-08), [Adyen Lifts 2026 Revenue Outlook to 21–23% as Volume Hits €804 Billion](https://www.financemagnates.com/fintech/payments/adyen-lifts-2026-revenue-outlook-to-2123-as-volume-hits-804-billion/)
- Financial IT (2026-08-14), [Adyen Publishes H1 2026 Financial Results](https://financialit.net/news/infrastructure/adyen-publishes-h1-2026-financial-results)
- Visa Acceptance Solutions and Merchant Risk Council (2025), [2025 Global eCommerce Payments & Fraud Report](https://www.visaacceptance.com/content/dam/documents/campaign/fraud-report/global-fraud-report-2025.pdf)
- Board of Governors of the Federal Reserve System (2025), [Regulation II: Average Debit Card Interchange Fee by Payment Card Network](https://www.federalreserve.gov/paymentsystems/regii-average-interchange-fee.htm)
- Adyen (2026), [Pricing](https://www.adyen.com/pricing)
- Stripe (2026), [Pricing](https://stripe.com/pricing)
- Stripe (2026-02), [Stripe publishes 2025 annual letter](https://stripe.com/newsroom/news/stripe-2025-update)
- Stripe (2026), [Agents and AI on Stripe](https://docs.stripe.com/building-with-llms)
- Forrester (2026-01), [The State of Business Buying, 2026](https://forrester.com/blogs/state-of-business-buying-2026)
- Stack Overflow (2025-07-29), [2025 Developer Survey](https://stackoverflow.co/company/press/archive/stack-overflow-2025-developer-survey/)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- Underneath (2026), [Same question, four countries: do AI recommendations change?](https://underneath.agency/research/ai-recommendations-by-country-study)
- Underneath (2026), [Do Wikipedia and schema make AI assistants recommend a brand?](https://underneath.agency/research/brand-entity-ai-recommendations-study)
- Underneath (2026), [How faithfully do AI assistants quote software prices?](https://underneath.agency/research/ai-pricing-accuracy-study)
- Underneath (2026), [How many websites have an llms.txt file?](https://underneath.agency/research/llms-txt-adoption-study)

---

This is the Markdown twin of https://underneath.agency/resources/payment-processors-merchant-demand-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How does a payroll provider get named when businesses ask AI?"
description: "By being the clear, verifiable answer for a business’s size, states and countries. AI answers on tax-bound questions shift by location and misquote prices."
canonical: "https://underneath.agency/resources/payroll-software-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How does a payroll provider get named when businesses ask AI?

By being the easiest provider to verify for the buyer’s exact situation: company size, the states or countries it pays people in, and what it will cost. Payroll is a compliance purchase that customers keep for years, so the first shortlist matters, and AI answers about tax-bound products move with the buyer’s location more than answers about most software. For a payroll company, the goal is not a generic “best payroll” mention but being named, correctly priced, for the questions that lead to a quote request.

## The short version

1. Payroll customers stay for a very long time: Paychex counts about 840,000 customers and 144,000+ who have been with it for 10+ years; ADP serves more than 1.1 million clients across 140+ countries.
2. Mistakes are expensive for the buyer, which shapes what they ask: the IRS charges 2% to 15% of an unpaid employment tax deposit depending on how late it is, and of the states that tax wages, 14 use one rate while 27 plus Washington, D.C. use brackets.
3. Global payroll is now a category of its own: Deel says it supports 40,000 businesses in 150 countries and has hired 650,000+ workers; Paycom says it supports workforces in over 190 countries.
4. AI answers on tax-dependent questions move with the user’s country: in our study, the gap in ChatGPT’s brand lists between countries was 0.180 for insurance, lending and tax questions, against 0.034 for global products such as CRM software.
5. Prices quoted by AI assistants are often off for the billing structures payroll uses: in our study of 45 software products, pages with a monthly/annual toggle had 58.4% fully faithful prices against 72.9% for pages without one.

## Who buys payroll software, and what is a client worth?

Owners and finance or HR leads buy it at a trigger: a first hire, a new state, a costly mistake.

The market is concentrated and sticky. [Paychex](https://www.paychex.com/corporate) reports about 840,000 customers, says it pays 1 in 11 U.S. private sector workers, and lists $6.5 billion in revenue. More telling for lifetime value: 144,000+ of its customers have been with Paychex for 10+ years. [ADP](https://mediacenter.adp.com/2026-07-29-ADP-Reports-Fourth-Quarter-and-Fiscal-2026-Results) describes more than 1.1 million clients across 140+ countries. Mid-market specialists are smaller but significant: [Paycom](https://www.paycom.com/about/) reports more than 39,000 clients.

Payroll revenue compounds. A client pays a base fee plus a per-employee charge every pay period, and payroll is the door to time tracking, benefits, retirement plans, workers’ compensation and HR outsourcing. Because changing providers mid-year means moving tax records and filings, we infer that most switches happen at a trigger or at year-end, and that a won client is worth many years of fees. Providers that also sell HR suites can compare with [how AI search shapes HR software demos](https://underneath.agency/resources/hr-software-ai-search).

Three buyer groups ask different questions. Small businesses want simplicity and price. Multi-state employers need correct state and local withholding. Companies hiring abroad need an employer of record (EOR), a provider that legally employs workers in countries where the company has no entity. [Deel](https://www.deel.com/about/), which started in that last group, lists 40,000 supported businesses, 650,000+ workers hired and 150 available countries.

## Where do AI assistants fit into choosing a payroll provider?

At the research and comparison stage, but no published survey measures how often payroll buyers use them.

We found no payroll-specific study of AI use in provider selection, and we will not borrow one from another category as if it applied. What exists is general software-buying data. [Capterra’s 2025 Tech Trends survey](https://www.capterra.com/resources/tech-trends-successful-buyer-purchase-journey/) of 3,500 software buyers found that successful buyers were at least 50% more likely to factor product comparison sites and expert recommendations into their shortlist, and that most successful buyers (57%) took 3 months or less to evaluate options. AI assistants now perform exactly that comparison work in one answer; [our B2B SaaS article](https://underneath.agency/resources/b2b-saas-revenue-from-ai-search) covers the cross-category surveys.

What we can show is how AI answers behave on questions like payroll’s. In [our country study](https://underneath.agency/research/ai-recommendations-by-country-study), each question was coded by how much a good answer should depend on the country. On insurance, lending and tax questions the gap in ChatGPT’s brand lists between countries was 0.180; on global products such as headphones and CRM software it was 0.034. Payroll, built on national and state tax rules, sits with the first group. A UK employer and a US employer asking the same words should get different providers, and in our data they often do.

## Which questions do payroll buyers ask AI assistants?

Questions about size, states, countries, compliance, integrations and price, often triggered by a deadline or a mistake.

We drafted the sample prompts below to show how an owner or HR lead facing a payroll deadline might phrase things; none were recorded from real buyers:

- **First hire:** “What is the easiest payroll service for a business hiring its first two employees?”
- **Multi-state:** “Best payroll provider for a 60-person company with remote staff in 9 states.”
- **Global:** “Employer of record vs payroll provider for hiring three engineers in Poland.”
- **Compliance fear:** “Which payroll companies guarantee they will pay IRS penalties if they make a mistake?”
- **Switching:** “Gusto vs ADP vs Paychex for a restaurant with tipped employees.”
- **Integration and price:** “Which payroll software works with QuickBooks Online, and what does it cost per employee?”

Compliance questions carry weight because the penalties are concrete. The [IRS failure-to-deposit penalty](https://www.irs.gov/payments/failure-to-deposit-penalty) is 2% of the unpaid deposit at 1 to 5 days late, 5% at 6 to 15 days, 10% after more than 15 days, and 15% if it is still unpaid more than 10 days after a first IRS notice. State rules multiply the work: per the [Tax Foundation](https://taxfoundation.org/data/all/state/state-income-tax-rates/), of the states taxing wages, 14 have single-rate structures and 27 states and the District of Columbia levy graduated rates. A buyer who asks “which provider handles all of this” is asking a question with real money behind it.

## How does an AI answer become a payroll lead and a client?

Through the shortlist and the quote request: named, checked against price and compliance, then contacted for a quote.

**Named.** The assistant lists a handful of providers for the buyer’s size and locations. For tax-bound questions, the list depends on where the buyer is and which country they name, so national leaders and local specialists appear in different places. In our country study, local-market brands made up 25.4% of the brands ChatGPT named in the UK with only the location set, and 49.1% when “in the United Kingdom” was added to the question.

**Checked.** The buyer compares prices, guarantees and integrations, often in the same conversation. This is where wrong information costs leads. In [our pricing study](https://underneath.agency/research/ai-pricing-accuracy-study), 61.9% of plan prices four assistants quoted were fully faithful to the vendor’s page, another 3.8% had the right amount but dropped a condition, usually by presenting an annual-billing price as the monthly price, and 13.0% of dollar amounts were add-ons, yearly totals or other figures rather than plan prices. Payroll pricing, a base fee plus per-employee charges and add-ons, is exactly the kind assistants find hard to restate.

**Converted.** Most payroll providers sell through a quote, a sales call or a self-serve signup. The AI answer may not show up as a referral at all; a reasonable expectation is that it appears as a branded search, a direct visit or a prospect who already knows your guarantee. Over a client life that can run past a decade, as Paychex’s 144,000+ long-tenure customers show, a single additional win is worth far more than its first-month fee.

## What decides whether an AI assistant names a payroll provider?

Location, consistent facts and independent sources matter most in the evidence; platforms document only how they search.

**Documented by the platforms.** [Google says](https://developers.google.com/search/docs/appearance/ai-features) AI Overviews and AI Mode may use “query fan-out”, issuing multiple related searches across subtopics and data sources. A multi-state payroll question can therefore pull in pages on state withholding, pricing and reviews at once.

**Observed in our studies.** In the country study, about half (52.7%) of ChatGPT’s same-country advantage in brand overlap went with differences in the websites it cited. When the cited sources change by country, the brands change with them. Providers that want to be named in a market need to be described by that market’s sources, not only by their US site.

**Our inference about payroll-specific trust factors.** A reasonable expectation is that assistants weigh what payroll buyers already check: tax-filing guarantees, the states and countries supported, integrations with accounting software, review platforms, accountant recommendations, and independent coverage in small business and HR publications. No assistant lists tax guarantees or state coverage among its stated ranking rules.

## What does it cost a payroll provider to be missing?

A lost quote at the trigger moment, and often a decade of fees that follow.

Payroll buyers rarely switch without a reason. When they do, after a penalty, a new state or a first hire abroad, they decide in a short window, and Capterra found that successful software buyers mostly finish within three months. If you are not on the AI shortlist then, your next chance may be years away.

Being named with the wrong facts also costs. Where our pricing study could measure the size of a difference, 29 of 35 differing prices were underquotes. For payroll, an underquote sets an expectation your sales team then has to correct; an overquote can drop you from the list unseen. Correcting a misquoted per-employee fee follows the steps in [how to fix wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers).

## How does GEO work for a payroll company?

Generative engine optimization (GEO) gives AI search clear, verifiable facts about your coverage, price and guarantees. It cannot promise a recommendation.

1. **State and country pages.** Publish a current page for each state and country you support, stating taxes filed, registrations handled and limits. Answers change by location, so your coverage must be readable by location.
2. **One clear pricing structure.** Show the base fee, per-employee fee and add-ons in plain text on one page, with the billing term stated next to each price.
3. **Published guarantees.** State tax-filing accuracy guarantees, penalty policies and service levels in public, ungated pages.
4. **Independent coverage in each market.** Earn reviews, accountant recommendations and coverage in small business, HR and local trade publications in every country you sell in; see [our article on GEO across languages and markets](https://underneath.agency/resources/geo-strategy-for-global-brands-across-languages).
5. **Integration facts.** List accounting, time-tracking and benefits integrations with what each one syncs. Accounting tools face their own AI shortlist, covered in [how AI assistants pick accounting software](https://underneath.agency/resources/accounting-software-ai-search).
6. **Measurement by market.** Track a fixed set of buyer questions by state and country across ChatGPT, Gemini, Perplexity, Copilot and Google, and compare it with quote requests. An English-only, US-only check will miss most of what matters; see [why English-only audits fall short](https://underneath.agency/resources/english-only-ai-visibility-audits).

## What can’t the research yet tell a payroll provider about AI answers?

No published study ties AI visibility to payroll quote requests, and no survey tracks payroll buyers’ AI use.

- Our country and pricing findings cover many categories and one day of answers each, not payroll alone.
- Vendor figures, such as customer counts and tenure, are self-reported on the vendors’ own pages.
- Whether a published tax guarantee or state page changes what assistants say has not been tested.
- AI-influenced leads often leave no referral trail; see [why analytics miss AI visibility](https://underneath.agency/resources/why-analytics-miss-ai-visibility).

## How can a payroll provider check whether AI answers send it quote requests?

Put your buyers’ questions to the main assistants one state and one country at a time, and check each answer.

Note who is named, whether your coverage, tax-filing guarantees and per-employee prices are stated correctly, and which sources the answers cite, then set the gaps beside your quote requests by market. For an outside view tied to qualified payroll leads, [ask us for a state-by-state review of your AI visibility](https://underneath.agency/contact). Our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page outlines the work that follows, such as state and country coverage pages, one clear pricing structure and published tax-filing guarantees.

## Frequently asked questions

### Do AI assistants recommend different payroll providers in different countries?

They should, and on tax-bound questions they do more than on most software. In our country study, ChatGPT’s brand lists differed far more between countries on insurance, lending and tax questions than on global products.

### Can AI assistants quote payroll pricing correctly?

Sometimes. In our test of 45 software products, 61.9% of quoted plan prices were fully faithful to the vendor’s page, and pages with a monthly/annual toggle fared worse. Payroll’s base-plus-per-employee pricing adds room for error.

### Should an employer of record provider create a page for every country?

Yes, if it truly supports that country. Answers depend on location and on the sources cited in each market, so a country page with current facts is the most direct thing an assistant can cite.

### How do we know whether AI answers bring us payroll leads?

Ask new clients how they found you, track branded and direct inquiries by market, and compare both with a monthly check of your buyer questions across assistants. Referral data alone will undercount.

## Sources

- Paychex (2026), [Corporate overview](https://www.paychex.com/corporate)
- ADP (2026), [ADP Reports Fourth Quarter and Fiscal 2026 Results](https://mediacenter.adp.com/2026-07-29-ADP-Reports-Fourth-Quarter-and-Fiscal-2026-Results)
- Paycom (2026), [About Paycom](https://www.paycom.com/about/)
- Deel (2026), [About Deel](https://www.deel.com/about/)
- Internal Revenue Service (2026), [Failure to Deposit Penalty](https://www.irs.gov/payments/failure-to-deposit-penalty)
- Tax Foundation (2025), [State Individual Income Tax Rates and Brackets](https://taxfoundation.org/data/all/state/state-income-tax-rates/)
- Capterra (2024), [Capterra’s 2025 Tech Trends Report](https://www.capterra.com/resources/tech-trends-successful-buyer-purchase-journey/)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- Underneath (2026), [Same question, four countries](https://underneath.agency/research/ai-recommendations-by-country-study) and [software pricing accuracy](https://underneath.agency/research/ai-pricing-accuracy-study)

---

This is the Markdown twin of https://underneath.agency/resources/payroll-software-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How penetration testing firms can win scoped leads from AI search"
description: "By being the firm AI names when a buyer describes their audit, scope and deadline, with credentials, methods and pricing logic others can verify."
canonical: "https://underneath.agency/resources/pentest-firms-leads-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How can a penetration testing firm win more scoped leads from AI search?

By being the firm an AI assistant names when a buyer describes their audit, their systems and their deadline, and by publishing the credentials, methods and pricing logic that let the buyer check that answer. Penetration testing is bought under pressure from auditors, customers and insurers, so most buyers arrive with a scope already in mind. The assistant can match that scope to a firm before the first call, though no study yet measures how often pentest buyers ask one.

## The short version

1. Enterprises spend real money on testing: in [Pentera’s 2025 survey](https://itbrief.co.uk/story/survey-shows-enterprises-shift-towards-software-driven-pentesting) of 500 CISOs, US enterprises spent $187,000 a year on pentesting on average, about 10.5% to 11% of a security budget averaging $1.77 million.
2. Demand is set by rules and customers: PCI DSS requires internal and external tests [“at least once every 12 months”](https://strike.sh/learn/pci-dss-penetration-testing) and after significant changes, and [Cobalt](https://securitybrief.news/story/cobalt-report-reveals-gaps-in-critical-vulnerability-fixes) found 82% of organizations must give customers or regulators software security assurance.
3. Testing lags change: [Pentera’s 2024 survey](https://msspalert.com/news/pentesting-study-exposes-security-gaps-signals-mssp-prospects) found 73% of organizations change their IT at least quarterly but only 40% pentest that often.
4. Many smaller buyers have never bought a test: only 12% of UK businesses did penetration testing in the past year, in the [UK government’s 2025 breaches survey](https://www.gov.uk/government/statistics/cyber-security-breaches-survey-2025/cyber-security-breaches-survey-2025) of 2,180 businesses.
5. The pages AI cites are not neutral: in [our study of self-ranking lists](https://underneath.agency/research/self-promoting-best-lists-study), 24.2% of 269 cited “best X” lists ranked their own publisher first, the pattern behind many “best pentest companies” pages.

## Who buys penetration testing, and what is an engagement worth?

Security leaders, compliance owners and IT managers who need independent proof, usually by a fixed date.

Four buyers account for most of the work:

- **Enterprise CISOs** running a yearly testing program across networks, applications and cloud, often with several firms.
- **Compliance owners** at payment, software and health businesses who need a test report for an audit or a customer. A dental software maker is one such business: its group practice buyers ask about security before they choose, as [what dental groups ask technology vendors](https://underneath.agency/resources/dental-technology-ai-search) shows.
- **Public sector and critical infrastructure buyers.** In the UK, the NCSC, the government’s cybersecurity agency, runs the [CHECK scheme](https://www.ncsc.gov.uk/information/check-penetration-testing), which sets the standard for “authorised penetration tests of public sector and CNI systems and networks.”
- **Small and mid-size firms** with no in-house testers. In the UK survey, small and medium businesses that ran vulnerability audits were the most likely to use outside contractors only (38% each).

Spend varies with scope. Pentera’s figure of $187,000 a year describes large US enterprises, and more than half of its CISOs planned to raise pentest budgets. Platform contracts give a second view. On [Vendr’s purchase data](https://www.vendr.com/marketplace/synack), the median Synack buyer pays $105,600 a year, in a range from $79,215 to $140,150. [HackerOne’s](https://www.vendr.com/marketplace/hackerone) median is $41,500 across 316 purchases, for a mix of testing and bug bounty programs. A boutique firm’s single web application test sits well below these figures, but the account is worth more than the first test: PCI DSS requires repeat testing to verify fixes, and the cycle repeats every year.

## Why do companies buy more testing now?

Because audits, customers and insurers ask for proof, while systems change faster than yearly tests can cover.

**Rules set a floor.** PCI DSS Requirement 11.4, as quoted by Strike, calls for testing “By a qualified internal resource or qualified external third-party” with “Organizational independence of the tester,” and it is “not required to be a QSA or ASV.” Service providers must test segmentation controls “At least once every six months.” Buyers often ask an assistant to explain exactly these points before they ask who to hire.

**Customers ask for evidence.** Cobalt’s 2025 report, built on tests across more than 2,700 organizations, found 82% are required to give customers or regulators software security assurance. A pentest report is often that evidence.

**Insurers push controls.** In Pentera’s 2025 survey, 59% of organizations had implemented at least one security solution at their insurer’s request. Training vendors answer to the same insurers, as [how security awareness training vendors win buyers](https://underneath.agency/resources/security-awareness-training-customers-ai-search) shows.

**Systems change faster than tests.** Pentera’s 2024 survey found 60% of organizations pentest twice a year at most. Cobalt found 95% had pentested AI applications in the past year, and 32% of those tests found serious flaws. New AI features are a new reason to call a tester.

**Software competes for the budget.** In Pentera’s 2025 survey, 55% of organizations use software-based pentesting platforms. Pentera sells such software, so read its survey as one input. Buyers now ask which kind of testing they need, not only which firm.

## Where does AI search sit in the pentest buying journey?

At the research and shortlisting step, after a trigger and before scoping calls, though no survey isolates pentest buyers.

The journey, as we understand it:

1. **Trigger.** An audit date, a customer security questionnaire, an insurer’s request, a major release or a breach.
2. **Learning.** What does PCI DSS 11.4 require? What is the difference between a pentest, a vulnerability scan and a red team exercise?
3. **Shortlist.** Peers, past suppliers, search results and, increasingly, an AI assistant.
4. **Scoping call and proposal.** Number of applications, IP ranges, cloud accounts and test days set the price.
5. **Test, report and retest.** Then the same decision next year, or after the next significant change.

Across B2B purchases, [Gartner’s 2026 survey](https://www.gartner.com/en/newsroom/press-releases/2026-05-20-gartner-survey-finds-sixty-nine-percent-of-b-two-b-buyers-turn-to-sales-reps-to-validate-ai-generated-insights) of 645 buyers found 45% used generative AI in a recent purchase, mainly to gather information on vendors and products. For a pentest, we infer that much of this happens at steps 2 and 3, where an assistant can both explain a requirement and name firms that meet it.

## Which questions do pentest buyers ask AI assistants?

Questions about requirements, scope, credentials, price and test type. We drafted the sample prompts below in the voice of a CISO or compliance lead; they are examples, not logged queries.

| Buyer | Illustrative prompt |
|---|---|
| Payment business | “Penetration testing firm for PCI DSS 11.4 that can also test our segmentation every six months?” |
| UK public sector | “Which CHECK-approved companies test for NHS trusts?” |
| Software company | “Pentest firm that gives a report we can share with enterprise customers before a sale?” |
| Mid-size firm, first test | “How much does an external network penetration test cost for 50 IP addresses?” |
| AI product team | “Who tests AI chat features for prompt injection and data leaks?” |
| Enterprise CISO | “Manual pentest firm, PTaaS platform or automated validation: which fits a quarterly release cycle?” |
| Regional buyer | “Penetration testing companies in Texas with OSCP-certified testers?” |

The way the question is framed changes the answer. In [our phrasing study](https://underneath.agency/research/ai-prompt-phrasing-study), adding “on a tight budget” kept the original first brand only 15.3% of the time, against 68.0% when the same question was asked again. Price-conscious buyers may see a different set of firms, we infer.

## How does AI visibility become a scoped pentest lead?

When the assistant has matched your firm to the buyer’s standard, systems and deadline, the first call starts at scoping.

A good pentest lead has four facts settled: what must be tested, which requirement drives it, when the report is due and who will accept it (an auditor, a customer, an insurer). A buyer who asked for a firm that handles PCI DSS segmentation testing and AI application testing has already filtered on two of them. The likely path:

1. **AI answer.** A short list of firms, with reasons drawn from their pages and from what others write about them.
2. **Checking.** The buyer reads the firm’s methodology, sample report, tester credentials and reviews.
3. **Scoping call.** The buyer arrives with systems listed and a date in mind.
4. **Statement of work.** Test days, retest and reporting format are priced.
5. **Renewal and expansion.** Yearly retests, extra applications, segmentation tests and new AI features.

The reverse also holds. If an assistant describes your firm as network-only when you test cloud and AI applications, it filters out buyers who would have fit. Most of this happens without a click you can see in analytics; see [what lost clicks mean for pipeline](https://underneath.agency/resources/ai-answers-pipeline-revenue).

## What decides whether an assistant names a penetration testing firm?

No platform documents how it chooses security firms; studies show assistants look for rankings, reviews and prices.

**Documented by the platforms.** Google says AI Overviews and AI Mode [may use a “query fan-out” technique](https://developers.google.com/search/docs/appearance/ai-features), issuing several related searches for one question. Neither Google nor any other platform explains how a pentest firm ends up named.

**Observed in our studies.** In [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study), ChatGPT ran a search aimed at a named publication, ranking or award in 43.8% of its answers, looked for reviews in 46.2% and looked for prices in 23.8%. Lists matter, and many “best penetration testing companies” pages are written by firms that rank themselves first. Our self-ranking study found 24.2% of cited numbered lists did that, though such lists were only 1.1% of all citations. And names move between runs: in [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), only 25.2% of the brands ChatGPT named appeared in all five runs of the same question.

**What pentest buyers check.** The NCSC’s [penetration testing guidance](https://www.ncsc.gov.uk/guidance/penetration-testing) says third-party tests “should be performed by qualified and experienced staff only” and that “the quality of a penetration test is closely linked to the abilities of the penetration testers involved.” Buyers therefore look for named schemes (CHECK, CREST), individual certifications, a clear methodology and evidence of research.

**Our inference.** A reasonable expectation is that a firm whose accreditations, methods, sample deliverables and retest policy are stated plainly, and confirmed by independent sources, gives an assistant more to repeat accurately. No study has tested this for testing firms.

## What does it cost a pentest firm to be left out?

Mostly lost scoping calls on recurring work, though no study has measured that loss for testing firms.

- **The work recurs.** A buyer lost at the first test is often lost for the yearly retest, the post-change test and the next application too.
- **The market is shifting.** With 55% of large organizations using software-based testing, a consultancy missing from “which kind of testing do I need?” answers loses the argument before the comparison starts, we infer.
- **Resellers compete for the same buyer.** [MSSP Alert](https://msspalert.com/news/pentesting-study-exposes-security-gaps-signals-mssp-prospects) reports that 63% of managed security providers already offer their own pen testing as a service.
- **Wrong facts filter buyers out.** An assistant that omits your CHECK status or says you do not test AI applications sends those buyers elsewhere. The steps for correcting it are in [our guide to fixing wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers).

## How does GEO work for a penetration testing firm?

It makes your scope, credentials and proof easy for assistants and buyers to find and check. Nobody can promise a recommendation.

1. **Map your services to the requirements buyers name.** One clear page each for PCI DSS 11.4 testing (including segmentation), customer-assurance reports, CHECK or CREST work, cloud, and AI application testing.
2. **Explain how scope sets price.** Assistants answer pricing questions, and ChatGPT searched for prices in 23.8% of answers in our study. State what drives cost, even without a rate card.
3. **Show who tests.** Tester certifications, scheme memberships, years of experience and published research.
4. **Publish a sample report and retest policy.** These answer the questions compliance owners ask before a call.
5. **Earn independent coverage.** Vulnerability disclosures, conference talks, security press and analyst mentions carry more weight than your own list of top firms; see [how brands build authority for AI search](https://underneath.agency/resources/how-brands-build-authority-for-ai-search) and [which pages AI engines cite](https://underneath.agency/resources/best-of-lists-ai-recommendations).
6. **Keep reviews current** on the platforms buyers read. When [our reputation study](https://underneath.agency/research/is-it-legit-ai-reputation-study) asked assistants whether brands were legit, 88.0% of the answers cited a review or complaint platform.
7. **Measure with real buyer wording.** Test requirement, region, budget and test-type prompts across ChatGPT, Google AI Overviews and AI Mode, Gemini, Perplexity, Copilot and Claude, several runs each.

## What is still unproven about how pentest buyers use AI?

How many pentest buyers use AI assistants, and whether being named leads to signed statements of work. No published study answers either.

- **No survey of pentest buyers.** Gartner’s figures cover B2B purchases across industries.
- **Vendors fund much of the data.** Pentera and Cobalt sell testing; their surveys show the market’s direction, not neutral measurement.
- **Vendr is a sample.** Its medians reflect purchases on one platform, mostly for platforms rather than boutique consultancies.
- **Lead quality is unmeasured.** In [Martinez’s](https://arxiv.org/abs/2607.14035) review of research on AI search optimization, traffic and conversions had the weakest evidence.

## Where should a penetration testing firm start?

With the 20 to 30 questions your best-fit buyers ask before they call, and a record of who gets named.

Write prompts for each buyer type: PCI, customer assurance, public sector, AI applications, first-time buyers, with region and budget wording. Ask each major assistant several times. Note which firms are named, which lists and review sites are cited, and whether your credentials and services are described correctly. Firms that also sell security products can compare notes with [how cybersecurity software companies earn revenue from AI search](https://underneath.agency/resources/cybersecurity-software-revenue-from-ai-search) and [how MDR providers win leads](https://underneath.agency/resources/mdr-providers-qualified-leads-ai-search).

To see how your firm looks from the buyer’s side, [talk to our team](https://underneath.agency/contact). We will show which scoping-stage questions name your firm, which send buyers to competitors and platforms, and which gaps in your public proof are most likely costing you scoping calls and statements of work. Our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page sets out how we help testing firms publish requirement-mapped services, tester credentials and sample reports, then track the scoping questions over time.

## Frequently asked questions

### Does a pentest firm need to be a PCI QSA to do PCI DSS testing?

No. PCI DSS 11.4 asks for a qualified, organizationally independent tester, “not required to be a QSA or ASV.” Saying so plainly helps buyers and assistants.

### Should we publish penetration testing prices?

Explain what drives the price, at least. ChatGPT looked for prices in 23.8% of answers in our hidden-searches study, so price-free pages leave the answer to others.

### Do “best pentest companies” lists help AI visibility?

Independent ones may. Self-published lists are common: 24.2% of cited numbered lists in our study ranked their own publisher first.

### Is automated pentesting replacing consultancies?

Not replacing, but competing: 55% of large organizations in Pentera’s 2025 survey use software-based pentesting, often alongside manual testing.

### How often do companies run penetration tests?

Often less than they change their systems. In Pentera’s 2024 survey, 60% pentested twice a year at most.

## Sources

- Pentera, via IT Brief UK (2025-05-08), [Survey shows enterprises shift towards software-driven pentesting](https://itbrief.co.uk/story/survey-shows-enterprises-shift-towards-software-driven-pentesting)
- Pentera, via MSSP Alert (2024-04-22), [Pentesting infrequency leaves security gaps](https://msspalert.com/news/pentesting-study-exposes-security-gaps-signals-mssp-prospects)
- Cobalt, via SecurityBrief US (2025-04-17), [Cobalt report reveals gaps in critical vulnerability fixes](https://securitybrief.news/story/cobalt-report-reveals-gaps-in-critical-vulnerability-fixes)
- Strike (2026), [PCI DSS penetration testing: what Requirement 11.4 asks for](https://strike.sh/learn/pci-dss-penetration-testing)
- UK Department for Science, Innovation and Technology (2025-04-09), [Cyber security breaches survey 2025](https://www.gov.uk/government/statistics/cyber-security-breaches-survey-2025/cyber-security-breaches-survey-2025)
- UK NCSC, [CHECK penetration testing](https://www.ncsc.gov.uk/information/check-penetration-testing)
- UK NCSC, [Penetration testing](https://www.ncsc.gov.uk/guidance/penetration-testing)
- Vendr (2026), [Synack software pricing and plans](https://www.vendr.com/marketplace/synack)
- Vendr (2026), [HackerOne software pricing and plans](https://www.vendr.com/marketplace/hackerone)
- Gartner (2026-05-20), [Gartner survey finds 69% of B2B buyers turn to sales reps to validate AI-generated insights](https://www.gartner.com/en/newsroom/press-releases/2026-05-20-gartner-survey-finds-sixty-nine-percent-of-b-two-b-buyers-turn-to-sales-reps-to-validate-ai-generated-insights)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- Martinez (2026), [Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026)](https://arxiv.org/abs/2607.14035), arXiv:2607.14035.
- Underneath (2026), [How many “best of” lists cited by AI rank their own brand first?](https://underneath.agency/research/self-promoting-best-lists-study)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [Does rewording a question change AI brand recommendations?](https://underneath.agency/research/ai-prompt-phrasing-study)
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)
- Underneath (2026), [“Is this brand legit?” How AI assistants build a reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study)

---

This is the Markdown twin of https://underneath.agency/resources/pentest-firms-leads-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How pet brands get recommended when owners ask AI what to buy"
description: "By being the brand AI names, with accurate facts and claims, when owners ask what to buy for their pet; one recommendation can start years of repeat orders."
canonical: "https://underneath.agency/resources/pet-product-brands-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How do pet product brands get recommended when owners ask AI what to buy?

By making your products easy for AI assistants to identify, compare and trust: clear facts about which pet each product is for, consistent prices and availability at the retailers owners use, strong reviews, and credible independent and veterinary-professional coverage. Pet owners already ask AI about food and products, and they change what they buy as a result. Because most pet spending is on food and other repeat purchases, the brand named at that first question can keep the household for years.

## The short version

1. The market is large and repeat-driven: the [American Pet Products Association](https://americanpetproducts.org/news/u.s.-pet-industry-reaches-158-billion-in-2025-poised-for-continued-growth-in-2026) puts US pet spending at $158 billion in 2025 and projects $165 billion in 2026; food and treats were $68.3 billion, or 43.2% of the total ([Global Pet Industry](https://globalpetindustry.com/news/158b-us-pet-industry-maintains-steady-growth-through-2025/)).
2. Owners already ask AI: in a survey of 2,000 US pet owners commissioned by a pet food brand, 52% had asked AI about pet care products and 59% said they had changed how they care for their pet based on AI recommendations ([Pet Food Industry](https://www.petfoodindustry.com/pet-food-market/market-trends-and-reports/news/15833984/survey-55-of-pet-owners-ask-ai-for-nutrition-advice)).
3. Veterinarians still win trust: in the same survey, 71% named veterinarians their most trusted source for health and nutrition advice, against 14% who rely primarily on AI tools.
4. Repeat buying is where the money is: customers with Chewy’s Autoship subscription produced 84.6% of its net sales in its latest quarter, $2.82 billion ([Subscription Insider](https://www.subscriptioninsider.com/blog/chewy-autoship-customers-drive-84-6-of-net-sales)).
5. Strong brands can still be invisible: in 1,200 cat food recommendation lists from six AI models without web search, Purina, Royal Canin, Hill’s, Blue Buffalo and IAMS appeared in every list while Freshpet never appeared ([Malthouse and colleagues](https://arxiv.org/abs/2609.16304)).

## Who buys pet products, and what is one household worth?

Owners in 95 million US households buy, and a loyal household means years of repeat orders.

APPA counts 95 million US households with at least one pet. Dog ownership rose to 53% of households, or 71 million, and cat ownership to 39%, or 53 million, with growth led by Gen Z, Millennials and, newly, Gen X empty nesters. Budgets are under some pressure: 22% of owners spent less on their pets in 2025, and APPA describes a shift away from discretionary items toward essentials. That makes the essentials, food, litter, treats and everyday supplies, the core of most pet brands’ revenue.

The value of a household is in the repeat. At Chewy, the largest US online pet retailer, 21.7 million active customers spent an average of $602 each over the past year, and customers on Autoship, its repeat-delivery program, made up 84.6% of net sales. Brand data shows the same pattern. Freshpet, the refrigerated dog and cat food maker, reported to [Global Pet Industry](https://globalpetindustry.com/news/freshpets-q2-results-beat-expectations-as-heavy-buyers-drive-71-of-sales/) that its heaviest buyers account for 71% of its sales, with an average buy rate of $515 a year against $117 for the average buying household. Winning one household that settles on your food is worth several times a single purchase.

## How are pet owners using AI assistants before they buy?

To research nutrition, products and prices, and many act on what they read, while still deferring to their veterinarian.

The most detailed US data comes from a Talker Research survey of 2,000 pet owners, conducted in April 2026 and commissioned by Darwin’s pet food, so it should be read as a brand-funded survey:

- 43% had used AI tools like ChatGPT for information on their pet’s health or nutrition.
- The most common topics were nutrition (55%), symptoms or illness (54%) and pet care products (52%).
- 59% said they had changed how they care for their pet based on AI recommendations.
- When advice conflicts, 64% follow their veterinarian and only 12% side with AI.

European data points the same way. [zooplus](https://corporate.zooplus.com/wp-content/uploads/2026/07/260701_Press-release_zooplus_AI-application_EN.pdf), an online pet retailer with 11 million customers in Europe, found around one-third of pet owners already use AI for pet topics, from nutrition and training to product recommendations. In Germany, 32% of its customers had used AI to find the best price, but only 19% had completed a purchase directly through an AI assistant. AI shapes the choice; the purchase still mostly happens at a retailer.

Retailers are wiring themselves into these assistants. [Google documents](https://blog.google/products/shopping/agentic-checkout-holiday-ai-shopping/) that its agentic checkout, which can buy a tracked item once its price falls within budget, launched with eligible US merchants including Chewy. Chewy’s chief executive told investors that a team is “shaping protocol or commerce protocols” for the pet category “with the Googles of the world and the OpenAIs of the world,” and that Chewy feeds in “catalog information, attribute data, content data, vet information,” according to [Digital Commerce 360](https://www.digitalcommerce360.com/2026/07/09/ecommerce-trends-how-chewy-is-using-ai/). Petco’s chief executive was more cautious, saying customer adoption of agentic shopping “is still very small as it relates to our space” ([Global Pet Industry](https://globalpetindustry.com/news/chewy-and-petco-target-growth-through-ai-loyalty-and-recurring-sales/)).

## Which questions do pet owners ask AI about products?

Questions that combine the pet’s details with a product need, a budget and often a brand comparison.

We wrote these prompts to illustrate pet product questions; they are not observed data:

- Fit to the pet: “Best dry food for a large-breed puppy, grain-inclusive, under $70 a bag.”
- Format and cost: “Fresh dog food delivery services compared on price per day for a 40-pound dog.”
- Household problems: “Low-dust unscented clumping litter for a home with three cats.”
- Durability: “Chew toys that hold up to a power chewer.”
- Alternatives: “Cheaper food with similar ingredients to the brand my dog eats now.”
- Legitimacy: “Is this subscription dog food company legit?”

Many owners also ask health questions. Those belong with a veterinarian, and a responsible brand does not try to answer them through AI search. The commercial opportunity is in the product, fit, price and trust questions above.

## How does an AI recommendation turn into repeat pet product sales?

Through a first purchase, then a subscription or habitual reorder that can last for the pet’s life.

The path we infer from the data above runs like this: an owner describes the pet and the need; the assistant names a few brands and explains why; the owner checks reviews or asks the vet, buys a first bag or box at Chewy, Amazon, Petco, a grocery store or the brand’s site, and, if the pet does well on it, sets up repeat delivery. For consumables, that last step turns one answer into a stream of orders. Autoship customers, for example, account for most of Chewy’s sales and keep growing faster than its total.

Two details matter for brands. First, where the owner lands decides who captures the repeat: a recommendation that sends owners to a retailer builds the retailer’s subscription base as much as yours, so your products need complete, consistent listings there. Second, durable goods such as beds, crates, feeders and pet tech are bought less often, so for them AI visibility works more like a classic product comparison than a subscription engine. Pet tech makers can borrow from [how device brands make AI shortlists](https://underneath.agency/resources/consumer-tech-brands-ai-search).

## What decides which pet brands AI assistants name?

Studies point to marketplace visibility, professional authority and reviews; platforms do not document how they choose brands.

**Observed in studies.**

- *Visibility, not just size.* In the cat food tests by [Malthouse and colleagues](https://arxiv.org/abs/2609.16304), five established brands appeared in every list and Freshpet in none, with web search turned off. Across categories, the authors found that prominence went with marketplace-visibility signals, particularly search interest and online brand conversation, more than with conventional brand popularity. Freshpet is not small: by Nielsen’s measure it holds 4.3% of the US dog food and treats market and is the fastest-growing dog food brand by dollar sales. The study shows how an assistant’s built-in picture of a category can lag the market.
- *Veterinary authority and retail scale.* An AI visibility index published by a PR firm and reported by [Pet Food Industry](https://www.petfoodindustry.com/business-strategy/news/15834448/report-chewy-leads-pet-brands-in-ai-citation-share) found Chewy with a 14% share of AI citations across more than 60 pet-owner prompts and The Farmer’s Dog second at 9%. Hill’s Science Diet, Purina Pro Plan and Royal Canin dominated citations tied to veterinary recommendations, while mid-tier commodity brands increasingly appeared only for price-focused questions. The prompt set is small, so treat it as directional.
- *Review platforms.* In [our brand reputation study](https://underneath.agency/research/is-it-legit-ai-reputation-study), 88.0% of AI answers to “Is this brand legit?” cited a review or complaint platform, which matters for subscription and direct-to-consumer pet brands.
- *Vulnerable categories.* In tests of planted fake pages, [Luo and Chen](https://arxiv.org/abs/2606.13610) found everyday-consumption categories, including supplements, among the most exposed, “where users rely on community taste rather than canonical brands.” Pet supplements and treats share that profile, which is our inference.

**Documented by the platforms.** Google says agentic checkout runs on its Shopping Graph with eligible merchants. Neither Google nor OpenAI documents how pet brands are ranked in answers.

**Trust factors specific to pet products.** We infer that assistants weigh product facts that match the pet (species, life stage, size), the nutritional adequacy statement on pet food, recall history, veterinary-professional endorsement where it is real, review volume and sentiment, and price per day or per serving.

## Why must pet brands be careful with health claims in AI search?

Because pet food claims are regulated, owners act on AI answers, and an overstated claim can spread through them.

The [FDA](https://www.fda.gov/animal-veterinary/animal-food-feeds/pet-food) requires pet food to be “truthfully labeled,” reviews specific claims such as “maintains urinary tract health” and “hairball control,” and has a compliance policy on dog and cat food diets intended to diagnose, cure, mitigate, treat or prevent diseases. It also says its veterinary center does not recommend one product over another and that questions about a pet’s health should go to the pet’s veterinarian.

That has two consequences for AI search. First, a brand should never seed health claims it cannot support in the hope that assistants repeat them; that risks regulatory trouble and, given that 59% of surveyed owners changed their pet’s care based on AI, real harm. Second, the safe ground is product facts, the nutritional adequacy statement on the label, feeding guidance and clear advice to consult a veterinarian. Our article on [whether GEO can backfire on your brand](https://underneath.agency/resources/can-geo-backfire-on-your-brand) covers the reputational side, and [fake reviews and fake brands in AI recommendations](https://underneath.agency/resources/fake-reviews-fake-brands-ai-recommendations) covers the risk from others’ manipulation.

## How does GEO work for a pet product brand?

By making accurate product facts, credible coverage and good reviews easy for AI systems to find, without promising placement.

Generative engine optimization, GEO, for a pet brand covers:

- **Brand and product clarity.** One consistent name per product line and formula, with species, life stage, size, ingredients, nutritional adequacy statement and price per day or serving stated plainly on your site and every retailer listing.
- **Retailer data.** Complete, matching listings at Chewy, Amazon, Petco, Walmart and grocers, since assistants and agentic checkout draw on retailer and shopping data. For the details assistants can use, read [what product content AI shopping assistants prefer](https://underneath.agency/resources/product-content-ai-shopping-assistants-prefer).
- **Independent and professional coverage.** Reviews by pet publications and creators, and real veterinary-professional perspectives where they exist. Digital PR here earns the coverage assistants cite; see [how small brands get recommended by AI](https://underneath.agency/resources/how-small-brands-get-recommended-by-ai).
- **Reviews and reputation.** Volume, recency and replies on retailer sites and review platforms, and clear answers to “is it legit” questions about subscriptions and cancellation.
- **Search interest and conversation.** Since marketplace visibility went with prominence in one study, consistent brand demand and community discussion matter, not just your own pages.
- **Compliant content.** Product and feeding information that stays within what labels and the FDA allow.
- **Measurement.** Tracking product, price, comparison and legitimacy questions across ChatGPT, Gemini, Google AI Mode, Perplexity, Copilot and retailer assistants, repeated over time.

## What is still unproven about AI and pet product sales?

How much pet product revenue AI answers drive, and how assistants pick brands when they search the web.

- The best US survey of pet owners’ AI use was commissioned by a pet food brand; there is no independent equivalent yet.
- The cat food study turned web search off; answers that search the web may name different brands, including fresher ones.
- The citation-share index used just over 60 prompts and comes from a firm with commercial interests in the result.
- No public data links AI answers to subscriptions or lifetime value for a pet brand.
- Agentic checkout in the pet category is new, and one large retailer says adoption is still very small.

## Where should a pet product brand start?

Start by checking how AI assistants describe and recommend your products for the questions owners ask before a first purchase.

List the 30 to 50 product, fit, price, comparison and legitimacy questions that lead to a first bag, box or order in your categories, ask them repeatedly across the major assistants, and record whether your brand appears, whether facts and prices are right, and which retailers and reviewers are cited. Leave health questions to veterinarians and make sure nothing attributed to your brand overstates what it can claim. If you want a hand, [contact us about a pet brand review](https://underneath.agency/contact): we map how AI assistants present your brand at the moment owners choose, and build a plan to improve how your products are found, described and trusted, so more first purchases turn into repeat customers. The [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page explains how that plan is carried out, from matching retailer listings to keeping product claims within what labels allow.

## Frequently asked questions

### Do AI assistants recommend pet food based on veterinary advice?

They cite veterinary-associated brands often: one index found Hill’s Science Diet, Purina Pro Plan and Royal Canin dominating veterinary-recommendation citations. But assistants are not veterinarians, and owners should take health questions to their vet.

### Can a direct-to-consumer pet brand compete with big retail brands in AI answers?

Yes, in some categories. The same index found The Farmer’s Dog second overall with a 9% citation share, and our brand studies show independent coverage matters alongside size.

### Should we claim health benefits to get recommended?

No. Pet food claims are regulated, and the FDA reviews specific health claims. Stick to supported facts and label statements.

### Where do AI-influenced pet purchases actually happen?

Mostly at retailers. In zooplus’s German data, 19% of customers had completed a purchase directly through an AI assistant.

## Sources

- American Pet Products Association (2026-03-26), [U.S. Pet Industry Reaches $158 Billion in 2025, Poised for Continued Growth in 2026](https://americanpetproducts.org/news/u.s.-pet-industry-reaches-158-billion-in-2025-poised-for-continued-growth-in-2026)
- Global Pet Industry (2026-03), [$158B: US pet industry maintains steady growth through 2025](https://globalpetindustry.com/news/158b-us-pet-industry-maintains-steady-growth-through-2025/)
- Pet Food Industry (2026-09-02), [Survey: 55% of pet owners ask AI for nutrition advice](https://www.petfoodindustry.com/pet-food-market/market-trends-and-reports/news/15833984/survey-55-of-pet-owners-ask-ai-for-nutrition-advice)
- Pet Food Industry (2026-09-09), [Report: Chewy leads pet brands in AI citation share](https://www.petfoodindustry.com/business-strategy/news/15834448/report-chewy-leads-pet-brands-in-ai-citation-share)
- zooplus (2026-07-01), [AI becomes a digital pet advisor](https://corporate.zooplus.com/wp-content/uploads/2026/07/260701_Press-release_zooplus_AI-application_EN.pdf)
- Subscription Insider (2026-09), [Chewy Autoship Customers Drive 84.6% of Net Sales](https://www.subscriptioninsider.com/blog/chewy-autoship-customers-drive-84-6-of-net-sales)
- Global Pet Industry (2026-08-07), [Freshpet’s Q2 results beat expectations as heavy buyers drive 71% of sales](https://globalpetindustry.com/news/freshpets-q2-results-beat-expectations-as-heavy-buyers-drive-71-of-sales/)
- Global Pet Industry (2026), [Chewy and Petco target growth through AI, loyalty and recurring sales](https://globalpetindustry.com/news/chewy-and-petco-target-growth-through-ai-loyalty-and-recurring-sales/)
- Digital Commerce 360 (2026-07-09), [Ecommerce Trends: How Chewy is using AI](https://www.digitalcommerce360.com/2026/07/09/ecommerce-trends-how-chewy-is-using-ai/)
- Google (2025-11-13), [Let AI do the hard parts of your holiday shopping](https://blog.google/products/shopping/agentic-checkout-holiday-ai-shopping/)
- U.S. Food and Drug Administration (n.d.), [Pet Food](https://www.fda.gov/animal-veterinary/animal-food-feeds/pet-food)
- Malthouse and colleagues (2026), [Evaluating Brand Retrieval and Ranking in Large Language Model Recommendations](https://arxiv.org/abs/2609.16304)
- Luo and Chen (2026), [One Polluted Page Is Enough: Evaluating Web Content Pollution in LLM Recommenders](https://arxiv.org/abs/2606.13610)
- Underneath (2026), [Is it legit? AI reputation study](https://underneath.agency/research/is-it-legit-ai-reputation-study)

---

This is the Markdown twin of https://underneath.agency/resources/pet-product-brands-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How drugmakers keep brands visible and accurate in AI answers"
description: "By making the label, the safety story and the brand’s facts easy for AI to find and repeat, within FDA rules, since patients and doctors now ask AI first."
canonical: "https://underneath.agency/resources/pharma-brand-visibility-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How can a drugmaker keep its brands visible and accurate in AI answers?

By treating AI answers as a channel where accuracy is the goal: the brand should appear when people ask about it or its condition, described the way its approved labeling describes it, with the safety story intact. Patients and physicians now ask AI tools before they talk to each other, and regulators are watching drug promotion more closely than they have for years. Nothing here is medical advice, and no company can make an assistant mention or recommend a prescription drug.

## The short version

1. Patients ask AI first: a [2026 KFF poll](https://www.kff.org/health-information-trust/poll-1-in-3-adults-are-turning-to-ai-chatbots-for-health-information-equaling-the-share-who-use-social-media-for-health/) found 32% of US adults used AI chatbots for health information in the past year, and 41% of those users did so to look things up before deciding whether to see a provider.
2. AI answers about drugs can be wrong in ways that matter: in a study in BMJ Quality & Safety [summarized by AHRQ’s PSNet](https://psnet.ahrq.gov/issue/artificial-intelligence-powered-chatbots-search-engines-cross-sectional-study-quality-and), Microsoft’s Copilot answered questions on 50 common drugs with 88.7% mean accuracy, and experts judged that 22% of a subset of answers could lead to death or severe harm if followed.
3. Regulators are back: on September 9, 2025 the [FDA announced](https://www.fda.gov/news-events/press-announcements/fda-launches-crackdown-deceptive-drug-advertising) thousands of warning letters and about 100 cease-and-desist letters over drug ads, after warning letters fell to one in 2023 and zero in 2024.
4. Marketing money is moving to digital: [Fierce Pharma reports](https://www.fiercepharma.com/marketing/2026-forecast-pharma-ad-dollars-will-continue-shifting-away-traditional-tv) eMarketer’s estimate of $24.8 billion in healthcare and pharma digital ad spending in 2025, against about $7.9 billion traditional.
5. Doctors are harder to reach in person: Veeva’s 2024 field data, [published via BioSpace](https://www.biospace.com/new-veeva-pulse-findings-show-connected-engagement-creates-an-advantage-as-hcp-access-drops), found 45% of health care professionals accessible to drug companies, down from 60% in 2022.

## What does brand visibility mean for a prescription drug?

Being present and correctly described when people ask about a condition or a brand, not being recommended.

For most industries, AI visibility means being named when a buyer asks for options. A prescription drug is different. Patients cannot buy it on an assistant’s say-so; a prescriber decides, and a payer often decides whether it is covered. What a drugmaker needs from AI answers is narrower and stricter:

- **Accuracy.** The brand name, generic name, maker, approved indications and key safety information should match the approved labeling.
- **Balance.** The FDA’s [Office of Prescription Drug Promotion](https://www.fda.gov/about-fda/center-drug-evaluation-and-research-cder/office-prescription-drug-promotion-opdp) exists to ensure prescription drug promotion is “truthful, balanced, and accurately communicated.” An AI summary is not your promotion, but it shapes the same impression.
- **Presence.** When someone asks about a condition, the treatment options an assistant lists usually come from medical institutions and guidelines. Your brand appears if those sources discuss it.
- **Practical facts.** Patient support programs, savings information and how to talk to a doctor are questions people ask, and wrong answers cost prescriptions.

## How far have patients and physicians moved to AI for drug information?

Far enough to matter: a third of adults use AI chatbots for health information, and most physicians use AI.

**Patients and caregivers.** KFF’s 2026 poll found 32% of adults turned to AI chatbots for health information in the past year, including 29% for physical health. About two-thirds of users wanted quick information, and 41% wanted to look something up before deciding whether to see a provider. The answer they get may also be the last word: 42% of those who asked about physical health did not follow up with a doctor or other health professional. And 77% of the public worries about the privacy of medical information given to AI tools.

**Physicians.** The [American Medical Association’s 2026 survey](https://www.fiercehealthcare.com/ai-and-machine-learning/ama-physicians-use-ai-doubled-2023-2026) found 81% of physicians use AI professionally. Clinician tools are growing too. Research firm [Sacra describes](https://sacra.com/c/openevidence/) OpenEvidence as searching licensed content including journals, specialty guidelines and drug labels, and reports that it earns its money from pharmaceutical and medical device advertising. Device makers face the clinician side of this too, covered in [building device demand through AI answers](https://underneath.agency/resources/medical-device-demand-ai-search).

**Google.** In [our AI Overview study](https://underneath.agency/research/ai-overviews-frequency-study), all 20 question-form healthcare keywords we tested showed an AI Overview (100.0%), even though healthcare searches overall showed one only 43.0% of the time, mostly because local results crowd them out.

Meanwhile, the rep visit is less available. Veeva found half of accessible health care professionals meet with three or fewer companies, which leaves more of a prescriber’s learning to other channels, AI tools among them.

## What do patients and prescribers ask AI about a brand?

Questions about what a drug is for, how it compares, what it costs and what to watch for. We wrote these prompts as illustrations; none was collected from real users.

| Asked by | What they want to know | Example question (written by us) |
|---|---|---|
| Patient | Indication | “What is [brand] approved to treat?” |
| Patient | Comparison | “How is [brand] different from [other brand] for the same condition?” |
| Patient | Cost and access | “Is there a savings program for [brand] if my insurance does not cover it?” |
| Patient | Generic | “Is there a generic version of [brand]?” |
| Caregiver | Safety | “What side effects should I ask the doctor about with [brand]?” |
| Prescriber | Label detail | “What does the prescribing information say about dosing in kidney impairment?” |

Each question is a chance for an assistant to get your facts right or wrong. Comparison and cost questions are especially risky, because the answer often blends sources of uneven quality, including telehealth and compounding sites. Latham & Watkins notes that the overwhelming majority of FDA’s September 2025 warning letters concerned [online promotion of compounded GLP-1 products](https://www.lw.com/en/insights/fda-begins-crackdown-on-direct-to-consumer-pharmaceutical-advertising), a reminder of how crowded and contested some drug categories are online.

## How does an AI answer affect prescriptions?

Indirectly: it shapes the conversation a patient brings to the prescriber and the confidence a prescriber has in the brand.

As we read the evidence, the paths run like this:

1. **Patient path.** A patient or caregiver asks about symptoms, a condition or a brand they saw advertised. The answer names treatment options and, ideally, the right facts about yours. The patient raises it with a prescriber, who decides. How hospitals and practices get named in that search is covered in [how providers win patients through AI](https://underneath.agency/resources/healthcare-providers-patients-ai-search).
2. **Prescriber path.** A physician checks the label, the evidence or a comparison in an AI tool. A clear, current label and published evidence make an accurate answer more likely.
3. **Access path.** After a prescription, patients ask about coverage, savings and pharmacy options. A wrong answer here can mean an abandoned prescription; we infer this, as no study measures it.

What makes pharma different from every other industry in this series is that the commercial value of AI visibility is bounded by regulation and medical judgment. The aim is not to persuade the assistant. It is to make the accurate version of your brand the easiest one for any tool to find.

The size of the stakes shows in ad budgets. Fierce Pharma, citing iSpot data, reports pharma and over-the-counter brands spent more than $7 billion on linear TV ads through early December 2025, up about 16%, and it reports a forecast that digital will make up 82% of healthcare and pharma ad spending by 2027. Every one of those ads now sends viewers to a search box or an assistant to check what they heard.

## What decides how an assistant describes your drug?

The sources it finds: assistants lean on institutions, labels and independent coverage, and none publishes its selection rules.

**Documented by the platforms.** Google and OpenAI say their AI answers run web searches and show links to sources. Neither explains which drug pages they trust.

**Observed in studies.** An audit of 615 sources ChatGPT cited for consumer health questions, by [Jacques and colleagues](https://arxiv.org/abs/2601.17109), found 75.7% came from institutional sources such as medical institutions, government sites and professional associations; commercial health platforms took 12.4%. Our summary of [what cited commercial health sites have in common](https://underneath.agency/resources/chatgpt-cited-commercial-health-sites) covers the details. Across buyer questions in other categories, [our brand entity study](https://underneath.agency/research/brand-entity-ai-recommendations-study) found independent coverage was the strongest predictor of being named: each tenfold increase in independent sites naming a brand went with 4.7 times the odds of being recommended. That study did not test prescription drugs.

**Observed accuracy problems.** The BMJ Quality & Safety study used Drugs.com as its reference and found mean completeness of 76.7%. [Inside Precision Medicine reports](https://www.insideprecisionmedicine.com/topics/patient-care/ai-powered-chatbots-unreliable-for-patient-drug-information) that only 54% of the 20 answers the experts reviewed aligned with scientific consensus. The study tested Microsoft’s Bing chatbot, an earlier version of Copilot; newer assistants may do better or worse.

**Trust factors specific to pharma.**

- **The label is the anchor.** DailyMed, run by the National Library of Medicine, holds the [labeling companies submit to the FDA](https://dailymed.nlm.nih.gov/dailymed/about-dailymed.cfm), including prescribing information and patient labeling. Clinician tools describe answering from labels.
- **Fair balance.** The FDA notes that 100% of pharmaceutical social media posts in one 2024 review highlighted benefits, but only 33% mentioned potential harms. An assistant that learns from unbalanced material will repeat it.
- **Institutional voices.** Medical societies, patient advocacy groups and government health sites carry most of the weight in health answers.

**Our inference.** Assistants will describe your drug from whatever is most available and most consistent. If your label, patient information, brand site and third-party references agree, the accurate version has the best chance of being the one repeated. Earlier-stage assets face a related test with partners and investors, covered in [how biotech companies get found for partnering](https://underneath.agency/resources/biotech-partnering-ai-search).

## What does a pharmaceutical brand risk by ignoring AI answers?

Inaccurate descriptions, missing safety context and lost conversations with prescribers; nobody has priced the loss yet.

- **Errors you did not write.** An assistant that misstates an indication or omits a boxed warning damages patient trust, and you may be asked about it. Our guide to [fixing wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) covers what can be corrected.
- **Absence from condition answers.** If the institutions and guidelines that assistants cite do not discuss your brand, it may not appear when patients research their condition.
- **Confusion with copies.** In categories with compounded or look-alike products, assistants can blend your brand with others.
- **Unstable answers.** In [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), only 25.2% of the brands ChatGPT named for a question came back in all five runs, so one check proves little.
- **Regulatory attention.** The FDA says it is already using AI and other tools to monitor drug ads. Your own channels need to be beyond reproach, because they feed what assistants say.

## How does GEO work for a pharmaceutical company?

By making accurate, balanced, reviewed information about each brand easy to find and consistent everywhere. No mention is guaranteed.

1. **A clean entity record for each brand.** State the brand name, generic name, maker and approved indications the same way on your brand site, corporate site, patient information and public reference pages, consistent with the label.
2. **Current, crawlable labeling.** Keep prescribing information and patient labeling up to date in DailyMed and on your site as readable text, not only as image-based PDFs.
3. **Balanced brand pages.** Pages that present benefits and risks together, in plain language, give assistants a balanced source to quote. Every page should go through your medical, legal and regulatory review.
4. **Authoritative third-party coverage.** Support disease education through medical societies, patient advocacy groups and peer-reviewed publication, within the rules on industry funding and disclosure. How that third-party record turns into trust in AI answers is set out in [how brands build authority for AI search](https://underneath.agency/resources/how-brands-build-authority-for-ai-search).
5. **Access and support facts.** Publish savings program terms, eligibility and how to reach support clearly, so cost questions get the right answer.
6. **Keep ads separate.** Advertising on AI and clinician platforms is a different channel from being cited, with its own rules; see [the risks of native ads in AI answers](https://underneath.agency/resources/risks-of-native-ads-in-ai-answers).
7. **Monitor for accuracy and safety.** Track patient and prescriber questions across ChatGPT, Google AI Overviews and AI Mode, Gemini, Perplexity, Copilot and Claude, and route problems to medical affairs and pharmacovigilance as your procedures require. See [how many prompts to track](https://underneath.agency/resources/how-many-prompts-to-track-ai-visibility).

Do not try to game it. The FDA has named undisclosed paid influencer promotion as a concern, and planted reviews carry the same risk; see [legitimate GEO versus manipulation](https://underneath.agency/resources/legitimate-geo-vs-manipulation).

## Where are the gaps in what we know about pharma and AI answers?

Large: nobody has linked AI answers about a drug to prescriptions, and accuracy studies age quickly.

- **No prescription data.** We found no study connecting what assistants say about a brand to prescribing or adherence.
- **Accuracy studies are snapshots.** The Copilot study, published in 2025, tested an earlier chatbot, and assistants change often.
- **Rules are in motion.** The FDA plans rulemaking on broadcast ads, and it is unclear how promotion rules will treat sponsored content inside AI tools.
- **Industry figures have sponsors.** Ad spending estimates come from eMarketer and iSpot, and OpenEvidence figures from Sacra’s profile, not audited filings.
- **Revenue links are unproven everywhere.** See [does AI visibility drive business results](https://underneath.agency/resources/does-ai-visibility-drive-business-results).

## Where should a pharmaceutical company start?

With an accuracy audit of what assistants say about each brand today, reviewed by medical, legal and regulatory teams.

List the questions patients, caregivers and prescribers ask about each brand and its condition. Ask each one in more than one assistant, and repeat it on a few different days. Check every answer against the label: indications, safety information, generic status, access programs. Note which sources the assistants cite and where the errors come from.

We can help with that work: [talk to us about an AI answer accuracy review for your brands](https://underneath.agency/contact). It shows where assistants misstate your label, leave out safety context or confuse your brand with others, and which reviewed, public sources would most improve the accuracy of the conversations patients bring to their prescribers. The [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page describes how the ongoing label, entity and monitoring work is organized to fit your medical, legal and regulatory review.

## Frequently asked questions

### Can we pay an AI assistant to recommend our drug?

We know of no assistant that documents selling placement inside its answers. Ads on some platforms are shown separately and remain subject to FDA promotion rules.

### Is content written for AI search subject to FDA promotion rules?

Content a company publishes about its products is promotional material and needs the usual review. Ask your regulatory counsel how rules apply to new formats.

### Why does an assistant describe our drug wrongly?

Usually because the sources it finds disagree or are out of date. Check the label, your pages and third-party references first.

### Should we worry about compounded or copycat products in AI answers?

Yes, where they exist. Most of the FDA’s September 2025 warning letters concerned online promotion of compounded GLP-1 products.

### Do physicians’ AI tools use our website?

Clinician tools such as OpenEvidence describe answering from journals, guidelines and drug labels, so the label and published evidence matter most.

## Sources

- KFF (2026), [Poll: 1 in 3 Adults Are Turning to AI Chatbots for Health Information](https://www.kff.org/health-information-trust/poll-1-in-3-adults-are-turning-to-ai-chatbots-for-health-information-equaling-the-share-who-use-social-media-for-health/)
- AHRQ PSNet (2025), [Artificial intelligence-powered chatbots in search engines: a cross-sectional study on the quality and risks of drug information for patients](https://psnet.ahrq.gov/issue/artificial-intelligence-powered-chatbots-search-engines-cross-sectional-study-quality-and), summarizing Andrikyan et al., BMJ Quality & Safety 2025
- Inside Precision Medicine (2024-10), [AI-Powered Chatbots Unreliable for Patient Drug Information](https://www.insideprecisionmedicine.com/topics/patient-care/ai-powered-chatbots-unreliable-for-patient-drug-information)
- U.S. Food and Drug Administration (2025-09-09), [FDA Launches Crackdown on Deceptive Drug Advertising](https://www.fda.gov/news-events/press-announcements/fda-launches-crackdown-deceptive-drug-advertising)
- U.S. Food and Drug Administration (n.d.), [The Office of Prescription Drug Promotion (OPDP)](https://www.fda.gov/about-fda/center-drug-evaluation-and-research-cder/office-prescription-drug-promotion-opdp)
- U.S. National Library of Medicine (n.d.), [About DailyMed](https://dailymed.nlm.nih.gov/dailymed/about-dailymed.cfm)
- Latham & Watkins (2025-09), [FDA Begins Crackdown on Direct-to-Consumer Pharmaceutical Advertising](https://www.lw.com/en/insights/fda-begins-crackdown-on-direct-to-consumer-pharmaceutical-advertising)
- Fierce Pharma (2025-12), [2026 forecast: Pharma ad dollars will continue shifting away from traditional TV](https://www.fiercepharma.com/marketing/2026-forecast-pharma-ad-dollars-will-continue-shifting-away-traditional-tv)
- Veeva Systems via BioSpace (2024-05-23), [New Veeva Pulse Findings Show Connected Engagement Creates an Advantage as HCP Access Drops](https://www.biospace.com/new-veeva-pulse-findings-show-connected-engagement-creates-an-advantage-as-hcp-access-drops)
- Fierce Healthcare (2026), [AMA: Physicians’ use of AI doubled from 2023 to 2026](https://www.fiercehealthcare.com/ai-and-machine-learning/ama-physicians-use-ai-doubled-2023-2026)
- Sacra (2026), [OpenEvidence revenue, valuation and funding](https://sacra.com/c/openevidence/)
- Jacques et al. (2026), [Authority Signals in AI Cited Health Sources: A Framework for Evaluating Source Credibility in ChatGPT Responses](https://arxiv.org/abs/2601.17109), arXiv:2601.17109.
- Underneath (2026), [When does Google show an AI Overview? 1,248 US searches](https://underneath.agency/research/ai-overviews-frequency-study)
- Underneath (2026), [Do Wikipedia and schema make AI assistants recommend a brand?](https://underneath.agency/research/brand-entity-ai-recommendations-study)
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)

---

This is the Markdown twin of https://underneath.agency/resources/pharma-brand-visibility-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Selling procurement software when AI shapes the shortlist"
description: "Professional buyers still run RFPs, but AI answers now shape which procurement tools reach the longlist, so vendors need proof buyers and AI can check."
canonical: "https://underneath.agency/resources/procurement-software-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How do you sell procurement software to buyers who run RFPs for a living?

With evidence they can verify, placed where their research starts, because AI answers increasingly shape which tools reach the longlist that a formal request for proposal (RFP) then tests. Procurement teams are among the most disciplined software buyers there are, and in one 2025 survey every organization reported using AI in procurement. That combination rewards source-to-pay, spend and supplier management vendors whose claims are documented by others, and quietly excludes vendors who rely on their own marketing.

## The short version

1. The field is crowded: [The Hackett Group’s](https://www.thehackettgroup.com/the-hackett-group-releases-fall-2026-solutionmap/) fall 2026 assessment covers 122 procurement technology vendors across 16 source-to-pay categories.
2. Procurement buyers already use AI: in [ProcureAbility and ProcureCon’s 2025 survey](https://procureability.com/wp-content/uploads/2025/09/procureability_procurecon-the-state-of-procurement-in-h2-2025.pdf), every organization reported some AI in procurement, and 72% rated their AI maturity “moderate.”
3. Many are open to switching: 36% of the same respondents were not satisfied with their AI-powered supplier selection and risk tools, and procure-to-pay automation had a 19% dissatisfaction rate.
4. The stakes are large: 64% of the companies surveyed manage $250 million or more in spend, and 34% more than $1 billion.
5. Budgets are tight: the [Hackett Group’s 2025 study](https://www.supplychain247.com/article/2025-procurement-transformation-priorities) found procurement workloads rising 9.8% while staffing and budgets rose only 1%, so every purchase must prove its return.

## Who buys procurement software, and what makes them different?

Procurement leaders buy it, and they are professional evaluators who apply their own methods to your sale.

The buyers are senior and control large budgets. In the [ProcureAbility and ProcureCon survey](https://procureability.com/wp-content/uploads/2025/09/procureability_procurecon-the-state-of-procurement-in-h2-2025.pdf), respondents worked in procurement (35%), supply chain (41%) and risk management (24%), and 18% sat in the C-suite. Most of their companies (64%) had $250 million or more in spend under management. ProcureAbility sells procurement services, so it has an interest in the topic.

Their method is the RFP. The [Hackett Group’s SolutionMap](https://www.thehackettgroup.com/the-hackett-group-releases-fall-2026-solutionmap/) shows how formal procurement technology selection has become: it scores vendors against more than 500 functional and capability assessments, with mandatory demos and verified customer ratings, to help buyers “shortlist options faster, de-risk selections, and separate real capability from marketing narrative.” Hackett sells this assessment, but the description shows what buyers expect.

They also shape everyone else’s software purchases. [Forrester’s 2026 study of business buying](https://www.forrester.com/press-newsroom/forrester-2026-the-state-of-business-buying/) found procurement professionals are decision-makers in 53% of business buying cycles. A procurement leader who uses AI to vet vendors for their company will use it to vet you, we infer.

The pressure on them is to do more with less. The Hackett Group’s 2025 study found workloads expected to rise 9.8% with staffing and budgets up only 1%. Generative AI entered their list of top 10 improvement initiatives for the first time, alongside spend analytics, contract management and third-party risk.

## Where does AI search sit in procurement technology buying?

Before the RFP: in the research that frames the problem and builds the longlist. No public study measures this for procurement teams alone.

Procurement teams are heavy AI users at work. The ProcureAbility and ProcureCon survey found all surveyed organizations reported some AI implementation in procurement, with maturity “moderate” at 72%, “early” at 22% and “advanced” at only 6%. A [University of Mannheim and Institute for Supply Management study](https://www.bwl.uni-mannheim.de/en/details/state-of-the-procurement-profession-2026-results-presented-exclusively-at-ism-world/) found 80 percent of organizations still in exploration or pilot phase. Teams that are actively piloting AI tools, we infer, are also comfortable asking AI assistants about the market.

The wider software market already leans that way. When the review platform [G2](https://company.g2.com/news/g2-research-the-answer-economy) polled 1,076 software buyers in March 2026, 51% said they start research with an AI chatbot more often than with Google. That sample spans every kind of software purchase, not source-to-pay tools, and G2 sells visibility to software vendors.

A procurement lead who types a category query into Google will usually meet an AI answer first. Across the [1,248 US searches in our study](https://underneath.agency/research/ai-overviews-frequency-study), B2B software and technology keywords triggered an AI Overview, Google’s AI summary above the regular results, on 96.0% of searches.

The formal process still matters. An assistant does not run the RFP. But [Gartner Digital Markets](https://cdn-static.bizzabo.com/bizzabo.file.upload/yVkIqwmWRDKbEhqeE3aQ_Gartner-Digital-Markets_2025-Report_Making-The-List.pdf), surveying 3,500 software buyers across industries, found the initial informal list is “the list that 81% of buyers end up making a purchase from most or all of the time.” If AI research shapes that list, it shapes the RFP.

## Which questions do procurement teams ask AI assistants?

Questions that sound like an RFP: scope, fit, total cost, risk and replacement. We wrote the sample prompts below to show how a sourcing or procurement technology lead might phrase them; none were collected from real buyers.

| Buying stage | Illustrative prompt |
|---|---|
| Scope | “Should a $2 billion manufacturer buy a full source-to-pay suite or best-of-breed tools?” |
| Category | “What are the leading intake-to-procure tools for a company on SAP S/4HANA?” |
| Replacement | “Alternatives to Coupa for a mid-market company with a small procurement team?” |
| Comparison | “Ivalua or GEP for direct materials sourcing in automotive?” |
| Total cost | “What does a supplier risk management platform cost per year for 5,000 suppliers?” |
| Risk and compliance | “Which spend management tools have SOC 2 Type II and EU data residency?” |
| Proof | “What savings have companies reported after moving to an e-sourcing platform?” |

Price questions deserve care. When [we checked the prices four assistants quoted](https://underneath.agency/research/ai-pricing-accuracy-study) for 45 software products, only 61.9% of plan prices fully matched the vendor’s own pricing page. Procurement buyers, of all people, will notice a wrong number.

A single question about e-sourcing tools rarely stays single. Behind an AI Overview or AI Mode answer, Google [may use a “query fan-out” technique](https://developers.google.com/search/docs/appearance/ai-features), and behind a ChatGPT answer about spend software, OpenAI’s [search rewrites the question](https://help.openai.com/en/articles/9237897-chatgpt-search) into targeted queries. In [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study), ChatGPT ran a search comparing options in 33.8% of its answers and one aimed at a platform or directory in 31.2%. For procurement software, we infer, those searches would reach analyst assessments, review sites and procurement media.

## How does AI visibility turn into procurement software revenue?

Through the longlist that feeds the RFP, then through expansion across source-to-pay modules.

The path, as we read the evidence:

1. A chief procurement officer or procurement technology lead asks an assistant to frame options, often months before an RFP.
2. The answer names vendors and cites sources: analyst assessments, review platforms, trade media.
3. The team checks those names against advisors, peers and assessments such as SolutionMap.
4. Shortlisted vendors get the RFP, scripted demos and reference calls.
5. The winner signs a multi-year contract and, if it performs, adds modules.

Expansion is where the value grows. SolutionMap spans 16 source-to-pay categories, from sourcing and contracts to invoices and supplier management. A vendor that wins one category, such as spend analytics or intake, has a path to others. No public filing we found breaks out contract values for procurement software, so we do not estimate them.

Dissatisfaction creates openings. In the ProcureAbility and ProcureCon survey, procure-to-pay automation and automated contract management had low “very satisfied” ratings of 31% and 30%, with dissatisfaction rates of 19% and 11%. Teams unhappy with a tool tend to ask what else exists, which is exactly the alternatives question an assistant answers, we infer.

Expect the effect to show up as RFP invitations, not tracked clicks; our guide on [what lost clicks mean for pipeline](https://underneath.agency/resources/ai-answers-pipeline-revenue) shows how to connect AI visibility to procurement deals.

## Why does an assistant list one source-to-pay vendor and not another?

No platform documents how it chooses; studies point to independent sources, which professional buyers also trust most.

**Documented by the platforms.** Ask Google or ChatGPT about sourcing suites and, by their own documentation, each searches first and links the pages it drew on. Neither explains how one spend management or sourcing vendor gets picked over its rivals.

**Observed in studies.** Looking at US software questions, [Chen and colleagues](https://arxiv.org/abs/2509.08919) found that AI search drew 72.7% of its sources from earned media, that is, reviews and articles written by others, against 45.4% for Google, whose results leaned more on vendor sites. Software answers are also relatively stable: in [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), B2B software had the most stable brands of any industry, with an average overlap of 0.708 across repeated runs.

**What procurement buyers weigh.** The ProcureAbility and ProcureCon survey shows what slows their technology decisions:

| Barrier to procurement technology | Share rating it a top barrier |
|---|---|
| Budget limitations | 49% |
| Strategic misalignment | 49% |
| Security and compliance | 44% |
| Technical challenges | 26% |
| Human capital, such as digital literacy | 22% |

So a vendor needs a business case that survives finance, and security proof that survives IT.

**Our inference.** Procurement buyers check claims against independent assessments and references. Assistants appear to lean on the same kind of sources. A reasonable expectation is that vendors with public, specific, third-party proof give both readers more to work with. No study has tested this for procurement software.

## What does it cost a procurement software vendor to be missing?

Lost places on longlists, in a market where formal evaluations exclude latecomers. No study has yet priced that loss for source-to-pay vendors.

- **A crowded field.** With 122 vendors in one assessment, an assistant naming a handful leaves most out. We infer that a vendor missing from AI answers depends more on advisors and paid channels to get invited.
- **Long contracts.** Source-to-pay platforms run for years. A vendor missed in one selection may wait for the next renewal, we infer.
- **Narrow AI claims.** With 36% of buyers unhappy with AI-powered supplier selection tools, a vendor whose AI capability is misdescribed in AI answers may be ruled out on the very point buyers care about.
- **Wrong facts.** Wrong prices or features in AI answers cost credibility with buyers trained to spot them. Our guide to [fixing wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) covers how to correct a misquoted price or module.

## What can a source-to-pay vendor do to show up in AI answers?

Make your savings, security and integration proof easy for assistants to find and repeat; nobody can guarantee a recommendation.

For a procurement software vendor, generative engine optimization (GEO), earning accurate mentions in AI answers, comes down to six areas:

1. **One clear identity.** State which source-to-pay categories you cover, for which company sizes and spend types, and which systems you integrate with, the same way everywhere.
2. **Proof in public.** Publish customer results with numbers, such as savings, cycle time and spend under management, plus security certifications on plain web pages.
3. **Independent assessments and coverage.** Take part in analyst and advisor evaluations, earn coverage in procurement media and speak at industry events. Assistants search for evaluations like these by name. Our piece on [how brands build authority for AI search](https://underneath.agency/resources/how-brands-build-authority-for-ai-search) explains how that coverage is earned.
4. **Reviews and references.** Detailed reviews from real customers on platforms buyers use, and reference stories that name the problem solved.
5. **Content for RFP-style questions.** Honest pages on suite versus best-of-breed, total cost, implementation time and replacement. Third-party lists matter more than your own, as we explain in [which pages to target](https://underneath.agency/resources/best-of-lists-ai-recommendations).
6. **Measurement across assistants.** Put the same RFP-style questions to ChatGPT, Google AI Overviews and AI Mode, Gemini, Perplexity, Copilot and Claude more than once, and log which procurement vendors each one names.

Procurement tools are one corner of enterprise software; the broader picture is in [how B2B SaaS companies earn revenue from AI search](https://underneath.agency/resources/b2b-saas-revenue-from-ai-search). Contract and legal tools sold to lawyers face their own checks, covered in [how legal software wins law firms through AI](https://underneath.agency/resources/legal-software-ai-search).

## What don’t we know yet about AI search in procurement software buying?

Nobody has yet measured whether appearing in AI answers wins source-to-pay contracts.

- **No study of procurement buyers’ AI research habits.** The surveys measure AI use inside procurement work, not AI use to choose procurement software.
- **Interested sources.** ProcureAbility and Hackett sell procurement services and assessments; G2 sells review visibility.
- **No public contract values.** We found no public data on typical procurement software deal sizes.
- **Little proof on sales outcomes.** When [Martinez](https://arxiv.org/abs/2607.14035) reviewed 45 studies of AI search optimization, traffic and conversions were the least supported area, a gap that matters for vendors judged on signed contracts.

## How can a procurement software vendor see where it stands before the next RFP cycle?

Put the questions procurement leaders ask before an RFP to AI assistants, and record which vendors get named.

Cover scope, category, replacement, total cost, security and proof. Ask each question more than once in each major assistant, because the vendor list can shift between runs. Record which vendors appear, which sources are cited, and whether your source-to-pay categories, prices and customer results come through accurately. Where you are absent, the cause is usually thin independent proof: assessments, reviews and references written by others.

To get that picture without building it in-house, [ask us to review your RFP-stage visibility](https://underneath.agency/contact). We will map which longlists AI answers put you on or leave you off, and which missing proof most likely costs you RFP invitations. That follow-on work, from publishing customer results and security proof to tracking the questions procurement leaders ask, is laid out on our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page.

## Frequently asked questions

### Do procurement teams trust AI answers when choosing software?

They use AI widely but verify. Every organization in the ProcureAbility survey reported some AI use, and Hackett’s assessment exists to separate “real capability from marketing narrative.”

### Does a formal RFP make AI visibility irrelevant?

No. The RFP tests a list made earlier, and 81% of software buyers in a Gartner Digital Markets survey usually buy from their initial list.

### Should procurement software vendors publish prices?

At least ranges and cost drivers. Only 61.9% of software prices AI assistants quoted in our study were fully faithful, so publish facts to quote.

### Which procurement categories are most open to new vendors?

Where satisfaction is lowest: 36% were not satisfied with AI-powered supplier selection and risk tools in the ProcureAbility survey.

### How long before GEO affects procurement software pipeline?

Expect quarters. Selections pass through RFPs, demos and references, and budgets are tight, with staffing and budgets up only 1% in Hackett’s 2025 study.

## Sources

- The Hackett Group (2026-09-22), [The Hackett Group Releases Fall 2026 SolutionMap Evaluating 122 Procurement Technology Providers](https://www.thehackettgroup.com/the-hackett-group-releases-fall-2026-solutionmap/)
- The Hackett Group, via Supply Chain 24/7 (2025-07-21), [Top 10 Procurement Transformation Priorities Defining 2025](https://www.supplychain247.com/article/2025-procurement-transformation-priorities)
- ProcureAbility and ProcureCon (2025-09), [The State of Procurement in H2 2025](https://procureability.com/wp-content/uploads/2025/09/procureability_procurecon-the-state-of-procurement-in-h2-2025.pdf)
- University of Mannheim and ISM (2026-04-28), [State of the Procurement Profession 2026](https://www.bwl.uni-mannheim.de/en/details/state-of-the-procurement-profession-2026-results-presented-exclusively-at-ism-world/)
- Forrester (2026-01-21), [Forrester’s 2026 Buyer Insights: GenAI Is Upending B2B Buying](https://www.forrester.com/press-newsroom/forrester-2026-the-state-of-business-buying/)
- Gartner Digital Markets (2025), [Making the List](https://cdn-static.bizzabo.com/bizzabo.file.upload/yVkIqwmWRDKbEhqeE3aQ_Gartner-Digital-Markets_2025-Report_Making-The-List.pdf)
- G2, Tim Sanders (2026-04-15), [In the Answer Economy, Don’t Win the Click — Win the Answer](https://company.g2.com/news/g2-research-the-answer-economy)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- OpenAI Help Center (2025), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Chen, Wang, Chen and Koudas (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919), arXiv:2509.08919.
- Martinez (2026), [Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026)](https://arxiv.org/abs/2607.14035), arXiv:2607.14035.
- Underneath (2026), [When does Google show an AI Overview? 1,248 US searches](https://underneath.agency/research/ai-overviews-frequency-study)
- Underneath (2026), [How faithfully do AI assistants quote software prices?](https://underneath.agency/research/ai-pricing-accuracy-study)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)

---

This is the Markdown twin of https://underneath.agency/resources/procurement-software-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "What kind of product content do AI shopping assistants prefer?"
description: "Factual, specific, easy-to-compare content matched to what buyers ask. In tests, hard product facts drove 82.4% of AI rankings; hype and gimmicks backfired."
canonical: "https://underneath.agency/resources/product-content-ai-shopping-assistants-prefer"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# What kind of product content do AI shopping assistants prefer?

Factual, specific product information that is easy to compare and written around what buyers actually ask. In controlled tests, hard facts such as ratings, price and reviews drove most AI product rankings, while advertising-style copy and creative gimmicks pushed products down. Much of what AI says about products also comes from third-party pages, so your own listing is only part of the picture.

## The short version

1. In a test of three commercial AI systems, product facts such as rating, price and reviews explained 82.4% of how products were ranked, and brand name only 1.2% ([Chu and Hou](https://arxiv.org/abs/2606.17443)).
2. Rewriting listings in an advertising style dropped products 1.82 places on a ten-product list on average; the best automated rewrites instead kept facts and organized attributes for comparison ([Bagga and colleagues](https://arxiv.org/abs/2511.20867)).
3. Longer is not better: one of the strongest rewriters shortened listings by about 14 words on average, while one that added about 170 words ranked near the bottom (Bagga and colleagues).
4. In a hotel test across twelve AI systems, a top guest rating raised the chance of being recommended by 31.6 percentage points, while replies to reviews made no detectable difference ([Baig and colleagues](https://arxiv.org/abs/2606.16344)).
5. Only 61.9% of software plan prices quoted by AI assistants were fully faithful to the vendor’s pricing page ([our pricing study](https://underneath.agency/research/ai-pricing-accuracy-study)).

## What does product content that ranks well look like?

It preserves facts, names concrete attributes, organizes them for comparison and matches buyer language. That is the conclusion of [Bagga and colleagues](https://arxiv.org/abs/2511.20867), who built the largest shopping test so far. They used 13,747 real shopping requests drawn from Reddit, each paired with 10 Amazon listings.

The requests were long, averaging about 59 words. They stated budgets, past experiences and must-have features. A simulated AI shopping assistant then ranked the ten products, and the researchers rewrote one listing to see whether it moved up.

They let an automated system search for the best way to rewrite listings, starting from 15 different styles. Almost all ended up in the same place. The winning instructions told the writer to open with a summary and use labeled sections and short bullets. They also asked for real use cases, answers to likely buyer questions and strictly accurate facts.

This playbook worked broadly. After the automated search, 63 of 75 combinations of writing style and AI ranker improved. Results also improved on GPT-5 and Claude, which were never used during the search, though most gains there were modest.

## Do the usual copywriting tricks help?

Mostly not, and several hurt badly. Bagga and colleagues tested 15 common rules of thumb for AI-friendly writing, from an authoritative tone to adding an FAQ. Only four matched or slightly beat a plain one-line rewrite instruction.

| Writing style | Average places moved on a ten-product list |
|---|---|
| Advertisement-like copy | down 1.82 |
| Foreign-language flourishes | down 1.66 |
| Cut to a single sentence | down 1.49 |
| Creative short story (a deliberate bad example) | down 4.36 |

Length did not explain success either. GPT-5, one of the strongest rewriters, shortened listings by about 14 words on average.

GPT-4.1 added about 170 words and ranked near the bottom. Bulleted lists gave only a small lift, about 0.23 places.

The authors draw a blunt conclusion: “manually instilling ‘GEO knowledge’ through prompt design can backfire.” Testing beat intuition.

## How much do hard facts like price and ratings matter?

More than anything else tested. [Chu and Hou](https://arxiv.org/abs/2606.17443) gave three commercial AI systems lists of ten skincare products: one famous brand and nine invented ones. They varied rating, price, review count and ingredients.

Product facts explained 82.4% of how the AI ranked products. Position in the list explained 6.5%, and the brand name only 1.2%. When an invented brand had clearly better specifications, the AI recommended it about 96% of the time.

Tiny real differences were enough. An invented brand won half the time with just a 0.075-star rating edge or 1.6 times as many reviews. Brand name mattered most when product information was ambiguous, which is when AI falls back on names it knows.

## Which signals do AI assistants weigh differently from shoppers?

Some signals matter far more or less to AI than marketers expect. [Baig and colleagues](https://arxiv.org/abs/2606.16344) asked twelve AI systems to choose among five hotels whose details were randomly varied.

| Signal | Change in chance of being recommended |
|---|---|
| Top guest rating (4.7 versus 3.9 stars) | up 31.6 points |
| High price | down 30.0 points |
| Eco-certification | up 11.6 points |
| High review volume | up 8.3 points |
| Visible replies to reviews | up 0.1 points (no detectable effect) |

The authors note that replying to reviews is a tactic “the optimization industry actively promotes.” In their test, AI ignored it. Eco-certification counted for far more than a [human-focused reputation playbook](https://underneath.agency/resources/ai-vs-human-reputation-priorities) would expect.

[List position also mattered](https://underneath.agency/resources/does-list-order-change-ai-recommendations), though it says nothing about the hotel. Being placed higher in the candidate list was worth about $12 per night. And the reasons the AI gave for its choices did not fully match what actually drove them. Our guide to [why AI picks one hotel over another](https://underneath.agency/resources/why-ai-recommends-one-hotel-over-another) covers the hotel results in full.

## Where else do AI assistants get product information?

Mostly from third-party sites, which vary by assistant. [Uberti-Bona Marin and colleagues](https://arxiv.org/abs/2609.18729) audited ChatGPT, Gemini and Google’s AI Overviews on 117 real product questions, each asked three times.

Editorial and product-review sites made up 56.7% of the domains ChatGPT displayed. AI Overviews, the AI summary at the top of Google’s results, most often showed Reddit (35.0%) and YouTube (30.0%). For the same question, ChatGPT and Gemini shared only 5.4% of the domains they displayed.

The assistants also framed advice differently. ChatGPT gave a first-person pick, such as “my pick would be,” in 79% of answers that recommended products. Gemini did so in 7% and AI Overviews in 2%.

So your product page is one input among many. Reviews, comparison articles, forums and videos often shape what the assistant says.

## Does consistency across your pages matter?

Yes: AI answers often quote whichever version of a fact they find. In [our pricing study](https://underneath.agency/research/ai-pricing-accuracy-study), 61.9% of quoted plan prices were fully faithful to the vendor’s pricing page. A further 3.8% had the right amount but dropped a condition, usually presenting an annual-billing price as monthly.

Our [business facts study](https://underneath.agency/research/ai-business-facts-accuracy-study) found the same pattern with phone numbers. Where a business’s Google profile number did not appear on its own website, 30.6% of the numbers AI gave differed from the profile. Where it did appear, only 1.6% did.

The lesson is simple. If your site, marketplace listings and profiles disagree, AI may pick the wrong one.

## What should you do about it?

Write product content for comparison, not persuasion, and keep every fact consistent everywhere it appears.

1. **Lead with specifications buyers compare.** Price, ratings, dimensions, materials, certifications and compatibility should be stated plainly and labeled.
2. **Answer the real questions.** Shopping requests in the largest test averaged about 59 words, full of use cases and constraints. Address those situations directly.
3. **Cut advertising language.** Advertisement-style rewrites lost 1.82 places on average in testing.
4. **Earn genuine proof.** Ratings and review volume moved AI choices; replies to reviews did not, in the one test that measured them.
5. **Make facts match everywhere.** Align your site, marketplace listings, pricing pages and profiles, including billing terms.
6. **Check third-party coverage.** Review sites, forums and videos feed many answers. Make sure they have accurate information to work from.
7. **Test, don’t guess.** Rules of thumb often backfired. Measure how assistants describe your products before and after changes.

If you want help improving how AI assistants present your products, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

How these preferences play out inside live shopping assistants, end to end. The gaps:

- The largest test simulated only the ranking step, over ten Amazon listings, in English, with mostly North American buyers.
- The skincare and hotel tests used invented products and fixed short lists; real listings have photos, full reviews and many more competitors.
- None of these studies measures sales or clicks, only rankings and recommendations.
- AI systems change often. The hotel authors warn their results describe specific model versions and may drift.
- No study yet compares the same product content across Amazon’s, Google’s and OpenAI’s live shopping features.

## Frequently asked questions

### Does product description length affect AI recommendations?

Not on its own. In the largest test, one of the strongest rewriters shortened listings by about 14 words on average, while one that added about 170 words ranked near the bottom.

### Should product pages include an FAQ for AI search?

It may help a little but is no shortcut. An FAQ-style rewrite was one of only four of 15 common styles that matched a plain rewrite, and the best automated rewrites included buyer questions among other changes.

### Do AI shopping assistants favor big brands?

Only when products look identical. In one test, the famous brand won every trial with identical specifications, but brand name explained just 1.2% of rankings once real differences existed.

### Does structured, bulleted product copy rank better with AI?

Slightly. Bulleted lists added about 0.23 places on average in one large test, and the best-performing rewrites organized attributes into labeled sections.

## Sources

- Bagga, Farias, Korkotashvili, Peng and Wu (2025), [E-GEO: A Testbed for Generative Engine Optimization in E-Commerce](https://arxiv.org/abs/2511.20867), arXiv:2511.20867.
- Chu and Hou (2026), [Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems](https://arxiv.org/abs/2606.17443), arXiv:2606.17443.
- Baig, Gillani and Ali (2026), [Whose hotel does the AI recommend? An algorithm audit of reputation signals in LLM-assisted hotel selection](https://arxiv.org/abs/2606.16344), arXiv:2606.16344.
- Uberti-Bona Marin, Bertaglia, Astante, Rijsbosch, van Dijck, Hannák, Spanakis and Kollnig (2026), ["If I Had to Buy Just ONE: Galaxy S26 Ultra": Auditing AI-Generated Product Recommendations](https://arxiv.org/abs/2609.18729), arXiv:2609.18729.
- Underneath (2026), [How faithfully do AI assistants quote software prices?](https://underneath.agency/research/ai-pricing-accuracy-study)
- Underneath (2026), [Do AI answers match a business’s Google profile?](https://underneath.agency/research/ai-business-facts-accuracy-study)

---

This is the Markdown twin of https://underneath.agency/resources/product-content-ai-shopping-assistants-prefer. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How do productivity tools win users when AI assistants compete?"
description: "By being the tool AI assistants name for one job done better, then converting on day one. AI already sends productivity apps real traffic; suites raise the bar."
canonical: "https://underneath.agency/resources/productivity-software-ai-search-growth"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How do productivity tools win users when AI assistants compete?

By being the tool an AI assistant names when someone wants one job done better than the assistant or their office suite can do it, and by converting that person in the first session. AI assistants already send productivity apps measurable traffic, while Microsoft and Google now bundle AI into the suites most people already pay for. The evidence on traffic is solid; the evidence on how much of it becomes paying subscribers is still thin.

## The short version

1. AI assistants already send productivity tools real visits. [Similarweb](https://www.similarweb.com/ai-traffic/notion.com) estimates notion.com received 2.1M visits from AI engines in September 2026, [grammarly.com](https://www.similarweb.com/ai-traffic/grammarly.com) 692.9K and [todoist.com](https://www.similarweb.com/ai-traffic/todoist.com) 122.4K, with ChatGPT supplying 91.52% of Todoist’s.
2. Newer AI tools can punch above their size: AI engines made up 3.83% of [granola.ai’s](https://www.similarweb.com/ai-traffic/granola.ai) referral traffic, against 1.21% for notion.com, in the same estimates.
3. The first session decides: in [RevenueCat’s 2026 data](https://www.revenuecat.com/state-of-subscription-apps), 71.9% of conversions in its productivity category happened on day 0, against 50.6% across all apps.
4. The suites are bundling AI in: [Google](https://workspace.google.com/blog/product-announcements/empowering-businesses-with-AI) cut Workspace Business Standard with Gemini from $32 to $14 per user a month, and [Microsoft](https://www.microsoft.com/en-us/microsoft-365/blog/2025/01/16/copilot-is-now-included-in-microsoft-365-personal-and-family/) added Copilot to Microsoft 365 Personal and Family with a $3 monthly price rise.
5. Yet people add specialist tools rather than replace them: [Menlo Ventures](https://menlovc.com/wp-content/uploads/2026/09/menlo_ventures_consumer_ai_report-2026.pdf) found the share of AI users using both general and specialized AI products rose from 60% to 72% in a year.

## Who chooses productivity software, and what is a user worth?

Usually one person, for themselves, on a free plan or short trial; teams and companies come later.

Note-taking, task management, calendar, writing, meeting notes and workflow tools are mostly chosen bottom-up. A freelancer, student, manager or founder picks a tool, pays monthly if it earns a place in their day, and sometimes brings their team. That makes the individual decision fast and the price sensitivity high. For tools that whole teams adopt, see [how project management tools win customers](https://underneath.agency/resources/project-management-software-customers-ai-search).

RevenueCat’s 2026 report, built on over 115,000 apps, shows the pattern for its productivity category, which also includes design and developer tools. Productivity apps were the most price sensitive, with 44% of cancellations cost-related. Their monthly plans retained worst, at 28.2% against a 39.2% benchmark, while their yearly plans retained among the best, at 85.6%. A user who commits for a year is valuable; a user who tries a monthly plan and drifts is not.

At the business end, value comes from seats. Grammarly, now renamed Superhuman, says it has more than 40 million daily users, and it has bundled Grammarly, Coda and Superhuman Mail into one suite, according to its [rebrand announcement](https://www.grammarly.com/blog/company/announcing-company-rebrand-to-superhuman). Investors still pay for focused AI tools: the meeting-notes app [Granola](https://www.granola.ai/blog/series-c) raised $125M at a $1.5bn valuation.

## When does someone looking for a note, task or meeting app ask an assistant first?

At the start, often as the place someone first tries to do the job before looking for a tool.

Similarweb’s September 2026 estimates show AI engines sending traffic across the category: evernote.com received 172.6K AI visits, calendly.com 116.3K, granola.ai 111.8K, otter.ai 89.3K, obsidian.md 81.2K and fireflies.ai 34.8K. ChatGPT supplied most of it for most vendors, including 94.75% of Granola’s, but not all: Perplexity was Evernote’s largest AI source at 25.73%. Similarweb models these figures rather than counting visits directly, so read them as a guide to which productivity apps gain from AI, not as exact totals.

The jobs themselves are what people already bring to assistants. In Menlo’s survey, writing emails was the top consumer AI activity in 2025, cited by 19% of consumers, with managing to-do lists close behind at 18%, and organizing notes also in the top ten. When the assistant can do part of the job itself, a productivity company is [competing with the assistant’s own answer](https://underneath.agency/resources/ai-writing-tools-users-from-ai-search) as well as being recommended in it.

## Which questions do productivity buyers ask AI?

Questions about a workflow, a device, a constraint and, increasingly, whether a separate tool is needed at all.

We made up the following prompts to show how people describe a workflow when they ask for an app; they are not taken from real logs:

- Note-taking: “Best note app for someone who switches between iPhone and Windows and wants offline access.”
- Meeting notes: “Which AI meeting note taker works without a bot joining the call?”
- Tasks and calendar: “What app can schedule my to-do list into my calendar automatically?”
- Switching: “How do I move from Evernote to Notion or Obsidian without losing tags?”
- Bundling: “Do I still need Grammarly if I have Copilot in Microsoft 365?”
- Privacy: “Which AI note taker does not train on my recordings?”

The bundling question is the new one. A reasonable expectation is that buyers already paying for Microsoft 365 or Google Workspace now ask what a specialist tool adds, so the answer an assistant gives about your difference matters as much as whether it names you.

## Do Microsoft, Google and the assistants themselves shrink the market?

They raise the bar, but the evidence so far says people add specialist tools rather than drop them.

The suites have made AI a default. Google said it would include its AI in Workspace Business and Enterprise plans, so a Business Standard customer who had paid $32 per user a month with the Gemini add-on would pay $14, only $2 more than before. Google says Workspace serves more than 10 million businesses. Microsoft put Copilot into Microsoft 365 Personal and Family and raised the US price by $3 a month, the first increase since the plans launched. A productivity tool now has to be clearly better at one job than features its buyer already has.

Menlo’s data points the other way on substitution. Specialized AI usage rose from 69% in 2025 to 75% in 2026 among AI users, and the share using both general and specialized products jumped from 60% to 72%. Menlo’s reading is that when switching is free, so is adding, and that a specialized product can “own a workflow.” Independent productivity companies are bundling too: the Superhuman suite is one answer to the suites.

## How does an AI answer turn into a paying subscriber?

Through a quick try that has to succeed in the first session, then a plan that lasts.

**Recommendation.** The person describes the job and gets two or three names. In ChatGPT, the app itself may be part of the answer: [OpenAI says](https://openai.com/index/introducing-apps-in-chatgpt/) ChatGPT can suggest apps when they are relevant to the conversation, and that developers can reach over 800 million ChatGPT users. Granola says it is already an official connector in Claude and ChatGPT.

**First session.** Productivity buyers decide fast. RevenueCat found 71.9% of productivity conversions happened on day 0. We infer that an AI-referred visitor who arrives expecting the specific capability the assistant described must find it immediately, or the trial ends that day.

**Retention.** The monthly plan is the weak point, with productivity’s 28.2% monthly retention the lowest of its categories. AI visibility brings people in; it does not keep them. The yearly plan, the team plan and the workflow that depends on your tool are what make an AI-referred user worth having.

**What it costs to be missing.** If assistants answer “which app should I use for this” without your name, the people asking never reach your site, and your analytics will not show the loss. The missing app downloads and signups never reach a dashboard, as we explain in [why analytics miss AI visibility](https://underneath.agency/resources/why-analytics-miss-ai-visibility).

## Why does an assistant name one note or task app over another?

Platforms document a little; studies point to reviews, prices, fresh pages and a clear answer to “why this tool.”

**Documented by the platform.** OpenAI documents that ChatGPT can suggest apps in conversation and plans a directory where users browse them. On the Google side, AI Overviews and AI Mode may run a “query fan-out” technique of several related searches across subtopics and data sources, and [Google says](https://developers.google.com/search/docs/appearance/ai-features) a page needs nothing beyond normal search eligibility to appear.

**Observed in our studies.** In [our hidden searches study](https://underneath.agency/research/ai-hidden-searches-study), ChatGPT ran a mean of 3.7 searches per answer and looked for reviews in 46.2% of answers and prices in 23.8%. Recency counts too: across the assistants in [our freshness study](https://underneath.agency/research/ai-source-freshness-study), pages published in the last 90 days made up 17.4% to 22.6% of dated citations, against 6.9% of Google’s top ten. Productivity apps ship features monthly, so a stale review or comparison can describe a product that no longer exists. Those studies covered eight industries, not productivity software alone.

**Our inference for this industry.** A reasonable expectation is that assistants name tools that are described consistently for one job (“the note app for…”, “the meeting recorder that…”) across reviews, communities, app stores and the vendor’s own pages. For AI tools that read email, calendars or recordings, trust is part of the recommendation: Menlo found consumers have given AI agents access to their email (36%) and calendars (27%), so privacy and data-use pages are decision content.

## What does GEO involve for a note, task or meeting app?

It makes your app’s best job easy for assistants to find and repeat; it cannot promise a placement.

1. **Own one job in plain words.** Say clearly which job you do better than the suite or a general assistant, and say it the same way on your site, app store listings, help center and review profiles.
2. **Answer the bundling question.** Publish an honest page on what you add for Microsoft 365, Google Workspace and ChatGPT users, and what you do not.
3. **Keep comparisons and reviews current.** Refresh comparison and migration pages when features change, and encourage reviews that mention the use case and device.
4. **Publish pricing and limits plainly.** State free plan limits and monthly and yearly prices in one place, since assistants search for prices.
5. **Be present where assistants act.** Where it fits, consider assistant apps and connectors, and keep those listings accurate.
6. **Make privacy legible.** For tools that touch email, calendars or recordings, publish plain data-use and training policies.
7. **Measure by job and assistant.** Track the jobs your users hire you for across ChatGPT, Gemini, Claude, Perplexity, Copilot and Google, and connect them to trials, day-0 conversion and yearly plans. See [whether tracking ChatGPT alone is enough](https://underneath.agency/resources/is-tracking-chatgpt-enough).

For the B2B side of the story, see [how B2B SaaS companies earn revenue from AI search](https://underneath.agency/resources/b2b-saas-revenue-from-ai-search), and for smaller challengers, [how a small brand gets recommended by AI](https://underneath.agency/resources/how-small-brands-get-recommended-by-ai).

## What is still unknown about AI answers and productivity subscriptions?

Whether AI-referred users subscribe and stay at different rates, and whether bundled suite AI erodes specialist demand.

The Notion, Granola and Todoist visit counts are Similarweb models for a single month; they say nothing about signups or revenue. RevenueCat’s productivity category includes design and developer tools and covers apps using its platform, not web-first products alone. Menlo’s figures are survey answers from consumers. The suite price changes are documented, but nobody has published how they changed demand for specialist tools. No study follows people from an AI answer to a paid productivity subscription. Our guide on [how to prove GEO caused a change in sales](https://underneath.agency/resources/prove-geo-caused-sales) sets out how a subscription app could test it.

## How can a productivity app tell whether AI answers bring it yearly subscribers?

Find out which assistants name you for the jobs your best yearly subscribers hired you for.

List the jobs, devices and constraints behind your highest-value subscribers, ask the main assistants those questions, including the “do I need this if I have Copilot or Gemini” version, and compare the answers with your trials and yearly conversions. We can run that comparison with you and tie it to first-session conversion and yearly plans rather than to mentions alone; [talk to us about a subscription-focused audit](https://underneath.agency/contact). What the longer-term work covers, such as owning one job in plain words and answering the bundling question, is on our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page.

## Frequently asked questions

### Will Microsoft Copilot and Gemini make standalone productivity apps obsolete?

Not on the evidence so far. Both suites now include AI by default, but Menlo found more AI users combining general and specialized tools than a year earlier. The pressure is on tools that do not clearly do one job better.

### Which AI assistant sends the most traffic to productivity tools?

In Similarweb’s September 2026 estimates, ChatGPT was the largest AI source for most of the tools we checked, but Perplexity led for Evernote. One assistant does not represent the market.

### Should a productivity app build an app for ChatGPT?

It depends on the product. OpenAI documents that ChatGPT can suggest apps in conversation, but there is no published evidence yet on whether having an app changes how often a tool is recommended or how many users subscribe.

### Do AI-referred users convert better for productivity apps?

There is no published category data yet. Because productivity conversions happen mostly on day 0, the first session matters more than the source of the visit.

## Sources

- Similarweb (2026-09), AI traffic estimates for [notion.com](https://www.similarweb.com/ai-traffic/notion.com), [grammarly.com](https://www.similarweb.com/ai-traffic/grammarly.com), [todoist.com](https://www.similarweb.com/ai-traffic/todoist.com), [granola.ai](https://www.similarweb.com/ai-traffic/granola.ai), [evernote.com](https://www.similarweb.com/ai-traffic/evernote.com), [calendly.com](https://www.similarweb.com/ai-traffic/calendly.com), [otter.ai](https://www.similarweb.com/ai-traffic/otter.ai), [obsidian.md](https://www.similarweb.com/ai-traffic/obsidian.md/) and [fireflies.ai](https://www.similarweb.com/ai-traffic/fireflies.ai)
- RevenueCat (2026), [State of Subscription Apps 2026](https://www.revenuecat.com/state-of-subscription-apps)
- Menlo Ventures (2026-09), [2026: The State of Consumer AI](https://menlovc.com/wp-content/uploads/2026/09/menlo_ventures_consumer_ai_report-2026.pdf)
- Google Workspace (2025-01-15), [The best of Google AI, now included in Workspace Business and Enterprise plans](https://workspace.google.com/blog/product-announcements/empowering-businesses-with-AI)
- Microsoft (2025-01-16), [Copilot is now included in Microsoft 365 Personal and Family](https://www.microsoft.com/en-us/microsoft-365/blog/2025/01/16/copilot-is-now-included-in-microsoft-365-personal-and-family/)
- Grammarly (2025-10), [Announcing our company rebrand to Superhuman](https://www.grammarly.com/blog/company/announcing-company-rebrand-to-superhuman)
- Granola (2026-03-25), [Granola raises $125M to put your company’s context to work](https://www.granola.ai/blog/series-c)
- OpenAI (2025-10-06), [Introducing apps in ChatGPT and the new Apps SDK](https://openai.com/index/introducing-apps-in-chatgpt/)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- Underneath (2026), [hidden searches](https://underneath.agency/research/ai-hidden-searches-study) and [source freshness](https://underneath.agency/research/ai-source-freshness-study)

---

This is the Markdown twin of https://underneath.agency/resources/productivity-software-ai-search-growth. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How professional services firms win clients from AI search"
description: "Clients of law, accounting, design and consulting firms now research with AI first. Firms get named when directories, reviews and coverage confirm expertise."
canonical: "https://underneath.agency/resources/professional-services-firms-clients-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How do professional services firms win clients when buyers ask AI who to hire?

By making their expertise easy to confirm in the places AI assistants check: professional directories, reviews, independent coverage and clear practice pages. Across law, accounting, architecture, consulting and agencies, the same pattern is emerging. Clients use AI to understand the problem and the kind of help they need, then check names against sources they trust before they call.

## The short version

1. Legal clients start with AI: in Clio’s 2025 Legal Trends Report, as summarized by the [Illinois Supreme Court Commission on Professionalism](https://www.2civility.org/2025-clio-legal-trends-report/), 14% of consumers had used AI to answer legal questions and 43% had not but would; [Canadian Lawyer](https://www.canadianlawyermag.com/news/general/many-legal-clients-have-asked-artificial-intelligence-legal-questions-clio-report/393228) reports 28% of those who asked were told to contact a lawyer.
2. Design clients do too, then hire: in a June 2026 [Houzz survey of 3,149 renovating homeowners](https://www.constructionowners.com/press-release/houzz-survey-finds-ai-adoption-soars-among-construction-and-design-pros-while-homeowners-rely-on-the-experts), 22% had used AI for their project, and 80% still hire professionals.
3. Business buyers use AI and then verify: [Forrester](https://forrester.com/blogs/state-of-business-buying-2026) found 94% of business buyers use AI during buying, and a [Gartner survey of 645 B2B buyers](https://www.gartner.com/en/newsroom/press-releases/2026-05-20-gartner-survey-finds-sixty-nine-percent-of-b-two-b-buyers-turn-to-sales-reps-to-validate-ai-generated-insights) found 69% prefer to validate AI-generated insights with a person.
4. Assistants look up professional directories: in our study, when ChatGPT’s own search named a source such as Avvo or Super Lawyers, the answer cited that source 44.0% of the time, against 8.1% when it did not.
5. Reviews and listings shape local picks: in our local study, businesses with more reviews than the local median were 19.5 points more likely to be listed by ChatGPT, and accountants and lawyers were among the professions tested.

This is the overview. For one profession in depth, read our articles on [accounting firms](https://underneath.agency/resources/accounting-firms-clients-ai-search), [IT services firms](https://underneath.agency/resources/it-service-firms-leads-ai-search), [managed service providers](https://underneath.agency/resources/msps-customers-from-ai-search), [penetration testing firms](https://underneath.agency/resources/pentest-firms-leads-ai-search) and [wealth management firms](https://underneath.agency/resources/wealth-management-firms-clients-ai-search). Nothing here is legal, tax or financial advice.

## Why is buying professional services different from buying products?

Because the client is buying judgment they cannot inspect in advance, so trust and proof come before price.

A buyer can test software in a trial. They cannot test a lawyer, an auditor or an architect until the work is underway. That is why referrals still dominate. In a 2025 TaxDome survey of 350 US businesses reported by [CPA Practice Advisor](https://www.cpapracticeadvisor.com/2025/08/19/survey-of-smbs-shows-how-they-choose-and-evaluate-their-accountant-firm/167532/), 57% found their current accountant through a peer referral and only 3% through advertising.

Professional services also have two kinds of buyer, and AI sits differently with each:

- **Individuals and owners** hiring a lawyer, tax preparer, architect or advisor for a personal or small-business matter. They often do not know what kind of professional they need, so they ask.
- **Business buyers**, such as a general counsel, CFO, CMO or COO, hiring a firm through a shortlist, pitch or request for proposal. Forrester reports that, on average, 13 internal stakeholders and nine external participants influence business buying decisions.

In both cases, a referral now tends to be checked rather than simply accepted. We infer that AI assistants have joined the places that check happens, alongside the firm’s website, reviews and peers.

## Where does AI already sit in how clients choose a firm?

At the start, where clients work out the problem and the help they need. Increasingly, also at the check.

The evidence differs by profession, but it points the same way:

| Profession | What the evidence shows | Where to read more |
|---|---|---|
| Law | 14% of consumers have used AI for legal questions; 28% of those who asked were told to contact a lawyer (Clio, 2025) | This article |
| Accounting and tax | Tax-related ChatGPT searches in early 2026 were four times the year before; 30 percent were for help filing forms and using tax software ([GovTech](https://www.govtech.com/question-of-the-day/how-often-did-americans-turn-to-ai-for-help-with-their-taxes-this-year)) | [Accounting firms](https://underneath.agency/resources/accounting-firms-clients-ai-search) |
| Architecture and design | 22% of renovating homeowners used AI; comparing options was a use for 50% of them (Houzz, 2026) | This article |
| Consulting and advisory | 30% of small business owners go to generative AI such as ChatGPT for advice ([TD Bank and Wakefield Research](https://stories.td.com/volumes/default/Wakefield-Research-Analysis-of-Results-for-TD-Bank-4.10.25.pdf), 2025) | This article; [business advisory firms](https://underneath.agency/resources/business-advisory-firms-clients-ai-search) |
| Agencies and IT services | 45% of B2B buyers used generative AI, mainly to gather information on vendors and products (Gartner, 2026) | [IT services](https://underneath.agency/resources/it-service-firms-leads-ai-search), [MSPs](https://underneath.agency/resources/msps-customers-from-ai-search) |

Two patterns cut across the table. First, AI shows up early: clients use it to imagine, understand and compare. Houzz found homeowners used AI for renovation or design ideas (64%), visualizing concepts (58%) and comparing options (50%), far more than for budgeting (24%) or contracts (6%). Second, clients still hire people. Among homeowners who chose not to use AI at all, the top reason was a preference for professional expertise (38%).

## What do clients ask AI before they contact a firm?

Four kinds: the problem, who handles it, how the options compare, and whether a firm is any good.

We wrote the examples below to show typical phrasing; none comes from a real client’s chat.

| Profession | Illustrative prompt |
|---|---|
| Law | “My business partner is locking me out of the company accounts. What kind of lawyer do I need in Austin?” |
| Accounting | “Regional CPA firm vs Big Four for a first SOC 2 report at a 60-person SaaS company” |
| Architecture | “Architects in Denver who specialize in mid-century home remodels, and what they typically charge” |
| Consulting | “Who helps mid-sized distributors fix inventory planning, and how do I compare them?” |
| Agency | “B2B content agencies for cybersecurity companies with good client reviews” |
| Any profession | “Is [firm name] reputable? What do former clients say?” |

The last question matters most for firms that already win on referral. A prospect who hears a firm’s name from a colleague can now ask an assistant about it in seconds, and the answer draws on whatever the web says about the firm.

## How does an AI answer turn into an inquiry or a shortlist?

Through a short chain: the answer frames the problem, names options, and the client verifies a name before reaching out.

For individuals, the path tends to run from a problem question → an answer that explains what kind of professional to hire → a follow-up asking who → a firm page or directory profile → a consultation request. The Clio finding that 28% of people who asked AI legal questions were told to contact a lawyer shows the hand-off happening.

For business buyers, the path runs from research → a shortlist → a pitch or proposal → a retainer or project. Here the person matters at the end. Gartner found buyers used an average of seven information sources during a recent purchase, and 69% prefer to validate AI-generated insights with a sales rep. In a professional firm, that “rep” is usually a partner or principal. Forrester adds that buyers are more likely to engage with providers based on information from industry experts rather than AI tools.

Trust cuts both ways. Gartner found 51% of buyers say they are more likely to encounter misleading information from generative AI, and 49% say the same of a sales rep. Our reading: AI can put a firm on the list, but proof from independent sources and people is what keeps it there.

## What decides which firms an AI assistant names?

The sources it searches: professional directories, rankings, reviews, independent coverage and clear practice pages.

Here is what is documented by the platforms, observed in studies, and inferred by us.

- **Documented by Google.** [Google’s guidance on AI features](https://developers.google.com/search/docs/appearance/ai-features) says AI Overviews and AI Mode may use “query fan-out,” issuing multiple related searches across subtopics and data sources to develop a response.
- **Documented by OpenAI.** [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search) typically rewrites a question into one or more targeted queries for its search partners, and may share general location to make results local.
- **Observed in our hidden-searches study.** Asked “Who is the best divorce lawyer in Houston?”, ChatGPT ran four searches, naming Avvo, Martindale-Hubbell, U.S. News Best Lawyers and Super Lawyers. [Across the study](https://underneath.agency/research/ai-hidden-searches-study), when a search named a source, the answer cited it 44.0% of the time.
- **Observed in our local study.** ChatGPT listed [67.7% of the businesses ranked 1 to 3 on Google Maps](https://underneath.agency/research/chatgpt-local-picks-google-profile-study) for local questions about five services, including accountants and personal injury lawyers, and review volume was the one signal that held everywhere.
- **Observed in our brand studies.** Each tenfold increase in independent sites naming a brand went with [4.7 times the odds of being recommended](https://underneath.agency/research/brand-entity-ai-recommendations-study), and 88.0% of answers to “is this brand legit?” [cited a review or complaint platform](https://underneath.agency/research/is-it-legit-ai-reputation-study).

Each profession has its own version of those sources. Law has legal directories and peer rankings. Accounting has state society and CPA listings. Architects and designers have portfolio platforms such as Houzz. Agencies and development firms have review marketplaces such as Clutch, which in 2026 launched an app inside ChatGPT that brings [verified provider reviews into AI conversations](https://www.demandgenreport.com/?p=52857). We infer that a firm’s profile on its profession’s main directories now does double duty: it is read by the client and by the assistant.

## What happens when an assistant leaves out or garbles a firm’s details?

Missed consultations and shortlist places, plus wrong details that misdirect clients. Nobody has published the revenue cost yet.

**Being misdescribed.** When we asked four AI engines for the address, phone, website and hours of real local businesses, answers about accountants [differed from the Google profile 17.7% of the time](https://underneath.agency/research/ai-business-facts-accuracy-study), and answers about personal injury lawyers 16.4%. A wrong phone number on an intake call is a lost client.

**Being inconsistent.** In [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), only 25.2% of the brands ChatGPT named for a question appeared in all five runs. A firm checked once may look fine and then disappear on the client’s own search.

**Being absent.** The cost of not being named is our inference, not a measured loss: the client who learns the options from an assistant meets the named firms first. For a business buyer, that can mean the shortlist forms before a partner ever hears about the opportunity. Our article on [enterprise software shortlists in AI search](https://underneath.agency/resources/enterprise-software-shortlists-ai-search) shows the same thing happening in a neighboring market.

## How does GEO work for a professional services firm?

Generative engine optimization (GEO) makes a firm’s expertise, people and proof easy for AI assistants to find, confirm and repeat.

1. **Problem-first practice pages.** One page per problem the firm solves, in the client’s words, naming the industries, places and kinds of matter or project it takes on.
2. **The profession’s directories.** Complete, consistent profiles on the directories that matter in your field: legal directories, CPA and society listings, Houzz for design, Clutch for agencies and development firms.
3. **People pages with credentials.** Bar admissions, licenses, certifications, publications and industries served, matched across the firm’s site, LinkedIn and directory profiles.
4. **Independent coverage.** Trade press bylines, association talks, rankings and awards from credible publishers. This is where [authority for AI search](https://underneath.agency/resources/how-brands-build-authority-for-ai-search) comes from.
5. **Reviews and client proof within professional rules.** Reviews where the profession allows them, case studies or portfolios with client permission, and anonymized outcomes where it does not.
6. **Consistent facts.** The same name, offices, phone and practice list everywhere. If AI answers already get something wrong, [here is how to correct it](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers).
7. **Regular testing.** Ask the problem, who, compare and check questions across ChatGPT, Gemini, Perplexity, Claude and Google AI Mode, several times each, every quarter.

Clients increasingly care how firms use AI, too. In Clio’s report, 78% of clients want lawyers to disclose AI use. Lawyers, accountants and registered advisers must also keep every public claim within their profession’s advertising rules. Those rules are no handicap here, because sober facts anyone can check are exactly what an assistant can pass on without distortion. No firm can be promised a place in an AI answer; what GEO does is make the record about its lawyers, accountants or designers accurate, specific and easy to locate.

## What is still unknown about AI search and professional firms?

Mostly the last step: how often an AI mention becomes a signed engagement, and how much each profession differs.

- **No engagement data.** We found no public data linking a firm’s AI visibility to consultations, proposals or signed work, in any profession.
- **Vendor surveys.** The client-behavior figures come largely from software companies and platforms with an interest in the topic, such as Clio, Houzz and TaxDome.
- **Uneven coverage.** Law and accounting have the most evidence; architecture, consulting and agencies have little that is independent and public.
- **Firms are changing too.** The [Thomson Reuters Institute](https://www.thomsonreuters.com/en-us/posts/technology/ai-in-professional-services-report-2026) found organization-wide AI use in professional services almost doubled to 40% in 2026, so what clients expect from firms is shifting at the same time as how they find them.

## Where should a professional services firm start?

Pick the three practice areas or services you most want to grow. Then ask AI assistants what a client would ask.

For each, write the problem question, the who question, the comparison question and the reputation question, including the place where it matters. Run them in ChatGPT, Gemini, Perplexity and Google AI Mode several times. Note which firms are named, which directories and pages are cited, and whether your firm’s facts are right. That gap is the plan.

If your partners want more consultations, shortlist places and proposal invitations from clients who now research with AI first, [ask us for an AI visibility review of your firm](https://underneath.agency/contact). We will test the questions your clients ask, show the directories and sources assistants rely on in your profession, and plan the pages, profiles and coverage that connect your expertise to those clients. Our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page explains how that work runs for a firm, from practice pages and directory profiles to quarterly testing of client questions.

## Frequently asked questions

### Do people use ChatGPT to find a lawyer or an accountant?

Increasingly. 14% of consumers have used AI for legal questions, and in our study ChatGPT answered “best accountant” and “best lawyer” questions with named local firms, addresses and ratings.

### Which directories matter most for AI visibility?

It depends on the profession. In our study ChatGPT searched Avvo, Martindale-Hubbell and Super Lawyers for a lawyer question; agencies have Clutch, designers have Houzz.

### Do referrals still matter if clients use AI?

Yes. 57% of businesses found their accountant through a referral. AI is increasingly where a referred name gets checked, so the evidence about the firm must hold up.

### Should a firm publish content written for AI?

Write for clients. Clear pages about the problems you solve, who you serve and who does the work help clients and assistants alike; no format guarantees a mention.

## Sources

- Illinois Supreme Court Commission on Professionalism (2025), [Future Law: 2025 Clio Legal Trends Report](https://www.2civility.org/2025-clio-legal-trends-report/)
- Canadian Lawyer (2025-10-17), [Many legal clients have asked artificial intelligence legal questions: Clio report](https://www.canadianlawyermag.com/news/general/many-legal-clients-have-asked-artificial-intelligence-legal-questions-clio-report/393228)
- Houzz via Construction Owners (2026-09), [Houzz Survey Finds AI Adoption Soars Among Construction and Design Pros, While Homeowners Rely on the Experts](https://www.constructionowners.com/press-release/houzz-survey-finds-ai-adoption-soars-among-construction-and-design-pros-while-homeowners-rely-on-the-experts)
- Forrester (2026), [The State Of Business Buying, 2026](https://forrester.com/blogs/state-of-business-buying-2026)
- Gartner (2026-05-20), [Gartner Survey Finds 69% of B2B Buyers Turn to Sales Reps to Validate AI-Generated Insights](https://www.gartner.com/en/newsroom/press-releases/2026-05-20-gartner-survey-finds-sixty-nine-percent-of-b-two-b-buyers-turn-to-sales-reps-to-validate-ai-generated-insights)
- CPA Practice Advisor (2025-08-19), [Survey of SMBs Shows How They Choose and Evaluate Their Accounting Firm](https://www.cpapracticeadvisor.com/2025/08/19/survey-of-smbs-shows-how-they-choose-and-evaluate-their-accountant-firm/167532/)
- GovTech (2026-04-15), [How often did Americans turn to AI for help with their taxes this year?](https://www.govtech.com/question-of-the-day/how-often-did-americans-turn-to-ai-for-help-with-their-taxes-this-year)
- TD Bank and Wakefield Research (2025-04), [TD Bank Financial Preparedness Survey: analysis of results](https://stories.td.com/volumes/default/Wakefield-Research-Analysis-of-Results-for-TD-Bank-4.10.25.pdf)
- Demand Gen Report (2026-05-12), [Clutch Launches First B2B Services Marketplace App on ChatGPT](https://www.demandgenreport.com/?p=52857)
- Thomson Reuters Institute (2026), [AI in Professional Services Report 2026](https://www.thomsonreuters.com/en-us/posts/technology/ai-in-professional-services-report-2026)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- OpenAI Help Center (2026), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [Which Google Maps businesses does ChatGPT recommend?](https://underneath.agency/research/chatgpt-local-picks-google-profile-study)
- Underneath (2026), [Do AI answers match a business’s Google profile?](https://underneath.agency/research/ai-business-facts-accuracy-study)
- Underneath (2026), [Do Wikipedia and schema make AI assistants recommend a brand?](https://underneath.agency/research/brand-entity-ai-recommendations-study)
- Underneath (2026), [“Is this brand legit?” How AI assistants build a reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study)
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)

---

This is the Markdown twin of https://underneath.agency/resources/professional-services-firms-clients-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How do project management tools win customers from AI answers?"
description: "By being named for the use-case and team-size questions buyers ask AI, then turning free signups into paid seats that expand. Here is the evidence."
canonical: "https://underneath.agency/resources/project-management-software-customers-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How do project management tools win customers from AI answers?

By being the tool an AI assistant names when someone describes their team, their work and their budget, then turning that visit into a free workspace that grows into paid seats. Project management is one of the categories where AI assistants already send measurable traffic to vendors, and the vendors receiving it are not simply the largest. What no one has published yet is how much of that traffic becomes paying seats.

## The short version

1. AI assistants already send project management vendors measurable traffic. [Similarweb](https://www.similarweb.com/ai-traffic/trello.com) estimates trello.com received 326.6K visits from AI engines in September 2026, [clickup.com](https://www.similarweb.com/ai-traffic/clickup.com) 267.8K, [asana.com](https://www.similarweb.com/ai-traffic/asana.com) 198K and [monday.com](https://www.similarweb.com/ai-traffic/monday.com) 139.3K.
2. The AI-referred share does not follow company size: ClickUp received nearly twice monday.com’s estimated AI visits, and ChatGPT accounted for 87.63% of ClickUp’s.
3. Customers are worth more the longer they stay: [Asana](https://www.01net.it/asana-announces-fourth-quarter-and-fiscal-year-2026-results) had 817 customers spending $100,000 or more a year, and [monday.com](https://s29.q4cdn.com/881027206/files/doc_financials/2026/q2/v2/MNDY-USQ_Transcript_2026-08-10.pdf) expects net dollar retention of about 108% for 2026.
4. Project managers’ own communities shape Google’s AI answers: in [our Reddit study](https://underneath.agency/research/ai-reddit-citations-study), r/projectmanagers was the second most-cited community, and practitioner communities supplied 58.7% of the Reddit threads shown on software searches.
5. AI answers get seat pricing wrong in specific ways: in [our pricing study](https://underneath.agency/research/ai-pricing-accuracy-study), Asana’s Starter plan was quoted at $10.99 a month, the annual-billing rate, when the monthly price was $13.49.

## Who buys project management software, and what is a customer worth?

First a team lead who starts free; later the operations or IT leader who standardizes the company.

Project management software is bought at two levels. A team lead, agency owner or freelancer signs up for a free plan or trial, invites colleagues and starts paying when the team outgrows the free limits. Later, a PMO head, operations leader or IT team consolidates departments onto one platform under a company contract. The same brand has to win both moments, often years apart.

Security and fit weigh heavily at the second moment. In [Capterra’s 2025 survey](https://cioinfluence.com/machine-learning/ai-and-security-drive-project-management-software-investment-capterra-survey-finds/) of 2,545 project management professionals, 39% made purchases because of security needs, because these tools hold budgets, contracts and client deliverables. Method matters too: 41% of organizations use hybrid project management methods, so buyers ask whether a tool supports agile and traditional planning together.

The value of a customer comes from expansion. Asana reported 25,928 Core customers, those spending $5,000 or more a year, and its overall dollar-based net retention rate was 96%. monday.com said customers can now expand not only by adding seats for people but by buying AI capacity, and it still saw double-digit seat growth in enterprise. A team of five that finds you through an AI answer can become a company-wide contract, which is why the first recommendation matters more than its size suggests.

## Where do AI assistants already sit in the project management buying journey?

At the very first question, when someone describes their team and asks what to use.

Third-party traffic estimates show AI assistants already sending buyers to project management vendors. Similarweb’s September 2026 estimates put Trello, ClickUp and Asana each well above a hundred thousand visits from AI engines, with smartsheet.com at 96.1K. ChatGPT dominated most of these: it accounted for 84.42% of Trello’s AI visits, 79.71% of Smartsheet’s and 75.81% of monday.com’s. Asana was the exception, with claude.ai its largest AI source at 34.09%. Similarweb models these numbers rather than counting visits, so read them as a sign of which project management brands assistants send people to, not as exact traffic.

Google’s results are a second front. In [our AI Overview frequency study](https://underneath.agency/research/ai-overviews-frequency-study), 96.0% of B2B software and technology keywords showed an AI Overview, the highest of eight industries. On those software searches, Google often leaned on practitioners: in our Reddit study, a thread Google showed from a specialist practitioner community, such as one for project managers, was cited 19.3% of the time, against 7.7% to 10.3% for other kinds of community. That gap was not statistically clear once the search was held fixed, so treat it as a pattern, not a rule.

The monday.com experience shows why this matters. On its August 2026 call, its co-CEO said “the top-of-funnel environment remains volatile,” and the company has planned its year around growth from existing and larger customers rather than a return of new-signup volume.

## Which questions do project management buyers ask AI?

Questions that describe the team, the work and the constraint, rather than the product category alone.

We wrote these sample prompts ourselves to show how a team lead or agency owner might frame the question; none come from logged queries:

- Team size: “What is the best free project management tool for a team of three?”
- Use case: “Which project management software works best for a marketing agency juggling client approvals?”
- Method: “Best tool for a team that runs sprints but also needs Gantt charts for leadership.”
- Alternatives: “Simpler alternatives to Jira for a non-technical operations team.”
- Migration: “How hard is it to move from Trello to Asana with dozens of active boards?”
- Price per seat: “How much would ClickUp cost for twenty-five users billed monthly?”

The added context changes the answer. In [our prompt phrasing study](https://underneath.agency/research/ai-prompt-phrasing-study), of the brands named in at least 60% of answers to a plain question, only 51.3% held that level when business context was added and 42.7% when a budget was added. For project management, where nearly every real question includes a team size, a method or a budget, the plain “best project management software” prompt says little about where a brand really stands.

Vendors already write for these questions. ClickUp’s page on [monday.com alternatives](https://clickup.com/blog/monday-alternatives/) is organized by use case and team size, such as small teams that only need Kanban, and puts ClickUp first in several of its lists. Our [study of self-ranking lists](https://underneath.agency/research/self-promoting-best-lists-study) found no detectable advantage for the top pick of a list that ranks its own publisher first, so the structure is worth copying more than the ranking.

## How does AI visibility turn into paid seats?

Through a free signup that becomes a paid team, then a larger contract as the tool spreads across the company.

**Recommendation to signup.** The buyer asks, gets three or four names, opens two of them and starts a free workspace. For product-led project management tools, this is where AI referrals show up first, as signups and trials. [Note, task and meeting apps](https://underneath.agency/resources/productivity-software-ai-search-growth) face a similar test at signup.

**Signup to paid team.** Free plans convert when the team hits a limit: users, projects, automations or reporting. We infer that buyers who arrive from an AI answer have already compared options, so the first week of the free workspace has to match what the answer promised, especially on the use case the buyer described.

**Team to company.** Expansion is where most of the value sits. Asana had 817 customers spending $100,000 or more a year, up 13% year over year, with net retention of 96% in that group. monday.com expects net dollar retention of about 108% in 2026. A vendor that wins the first team through an AI answer has a head start when the company standardizes. Our guide to [how collaboration software wins teams](https://underneath.agency/resources/collaboration-software-ai-search) follows the same team-to-company path.

**What it costs to be missing.** A buyer who never sees your name in the answer rarely appears in your analytics at all. Because these tools spread from team to team, missing the first recommendation can mean missing the later company contract too. Our piece on [what lost clicks to AI answers mean for pipeline](https://underneath.agency/resources/ai-answers-pipeline-revenue) explains how to account for team leads who never reach your site.

## What decides whether an AI assistant names a project management tool?

Platforms document little; studies point to practitioner discussion, reviews, clear facts and pages that match the buyer’s situation.

**Documented by the platform.** According to [Google](https://developers.google.com/search/docs/appearance/ai-features), an AI Overview or AI Mode answer may rely on a “query fan-out” technique that issues multiple related searches across subtopics and data sources, and a page needs nothing beyond normal search eligibility to appear. A question about a tool for a three-person agency can therefore pull in pricing pages, template galleries, review sites and community threads at once.

**Observed in studies.** Project managers’ communities are a cited source in Google’s answers: r/projectmanagers was cited nine times on 26 September 2026, second only to a plumbing community, and r/projectmanagement also ranked among the most cited. In that study, half of Google’s Reddit citations pointed to practitioner communities, forums where people answer questions about their own work, as project managers do on r/projectmanagers.

**Our inference for this industry.** A reasonable expectation is that tools named often by practitioners for a specific situation, such as agencies, construction, software sprints or nonprofit programs, are the ones assistants name for that situation. Clear public facts matter as well: pricing per seat and billing term, free plan limits, integrations, security certifications and supported methods.

**Accuracy on seat pricing.** Seat-based prices with monthly and annual rates are easy to garble. In our pricing study, a dropped condition was almost always a billing term, as in the Asana example above. A buyer comparing three tools on a quoted price can rule you out on a number that was never your list price.

## How can a project management vendor earn its place in AI answers?

Make your use cases, seat prices and practitioner reputation easy for assistants to find and trust. No placement can be promised.

1. **Use-case and team-size pages.** Publish honest pages for the situations your buyers describe (team size, industry, method, migration), with real limits stated plainly.
2. **One clear pricing page.** Show per-seat prices for monthly and annual billing side by side, with free plan limits, and retire old price figures.
3. **Practitioner presence.** Support the communities, newsletters and courses project managers use. Our article on [which pages to target](https://underneath.agency/resources/best-of-lists-ai-recommendations) explains how to choose independent coverage.
4. **Reviews that describe the setup.** Encourage reviews that say team size, method and use case, so they answer the questions buyers ask.
5. **Template and integration catalogs.** Keep public, crawlable template galleries and integration pages that describe what each one does.
6. **Security and admin documentation.** Make the facts IT needs at the company-wide stage public and plain.
7. **Measurement by situation.** Track questions by team size, use case and method across ChatGPT, Claude, Gemini, Perplexity, Copilot and Google, and connect them to signups and seat growth. How long that list of team-size and method questions should be is covered in [how many prompts to track](https://underneath.agency/resources/how-many-prompts-to-track-ai-visibility).

For the wider SaaS picture, see [how B2B SaaS companies earn revenue from AI search](https://underneath.agency/resources/b2b-saas-revenue-from-ai-search), and for comparison content, [our article on comparison pages](https://underneath.agency/resources/do-comparison-pages-help-b2b-ai-citations).

## What don’t we know yet about AI-referred signups in project management?

Two things: how many AI-referred signups become paid seats, and what AI visibility adds beyond brand strength and content volume.

Similarweb’s figures model a single month of visits and say nothing about which visitors went on to buy seats. One vendor’s page in that set reported data that looked incomplete, so we left it out. The company figures come from investor materials that do not separate AI-referred customers. Our studies observed which communities and sources Google cited on specific dates; they do not show that being cited causes a recommendation. No published study follows project management buyers from an AI answer to a paid plan and an enterprise contract. To test cause and effect in your own seat data, read [how to prove GEO caused a change in sales](https://underneath.agency/resources/prove-geo-caused-sales).

## How can a project management vendor tell whether AI answers are bringing paid seats?

Check which AI answers name you for the team sizes, use cases and methods your best customers have.

List the situations that lead to your highest-value accounts, ask the main assistants those questions in the words buyers use, and compare the answers with your signups, conversions to paid and seat growth by segment. To have that comparison built for you, with a plan to close the gaps judged by paid seats and expansion rather than visibility alone, [ask us for a review of your AI-referred signups](https://underneath.agency/contact). The [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page describes how we run that work, from use-case and pricing pages to tracking questions by team size and method.

## Frequently asked questions

### Why would a smaller project management tool get more AI traffic than a larger one?

We do not know for certain. A reasonable expectation is that tools with many pages answering specific situations, and many practitioner mentions, are named more often, but no study has isolated the cause.

### Do “alternatives to” pages help project management tools get recommended?

They match how buyers ask, which is a good start. Pages that rank your own tool first in every list are a weak bet; useful, fair comparisons by use case and team size are the better model.

### Should we track Claude as well as ChatGPT?

Yes, if your buyers use it. Similarweb’s estimates showed claude.ai as Asana’s largest AI source in September 2026, while ChatGPT dominated for other vendors, so one assistant does not represent the market.

### Does a free plan help with AI recommendations?

There is no published evidence that it does directly. Free plans are part of how buyers phrase questions (“best free tool for…”), so state your free plan limits clearly wherever you describe pricing.

## Sources

- Similarweb (2026-09), AI traffic estimates for [trello.com](https://www.similarweb.com/ai-traffic/trello.com), [clickup.com](https://www.similarweb.com/ai-traffic/clickup.com), [asana.com](https://www.similarweb.com/ai-traffic/asana.com), [monday.com](https://www.similarweb.com/ai-traffic/monday.com) and [smartsheet.com](https://www.similarweb.com/ai-traffic/smartsheet.com)
- Asana, via 01net (2026-03), [Asana Announces Fourth Quarter and Fiscal Year 2026 Results](https://www.01net.it/asana-announces-fourth-quarter-and-fiscal-year-2026-results)
- monday.com (2026-08-10), [Q2 2026 earnings call transcript](https://s29.q4cdn.com/881027206/files/doc_financials/2026/q2/v2/MNDY-USQ_Transcript_2026-08-10.pdf)
- Capterra, via CIO Influence (2025-09-04), [AI and Security Drive Project Management Software Investment, Capterra Survey Finds](https://cioinfluence.com/machine-learning/ai-and-security-drive-project-management-software-investment-capterra-survey-finds/)
- ClickUp (2026-06-11), [Best monday.com Alternatives and Competitors in 2026](https://clickup.com/blog/monday-alternatives/)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- Underneath (2026), [Reddit citations in Google’s AI](https://underneath.agency/research/ai-reddit-citations-study), [software pricing accuracy](https://underneath.agency/research/ai-pricing-accuracy-study), [prompt phrasing](https://underneath.agency/research/ai-prompt-phrasing-study), [self-promoting “best of” lists](https://underneath.agency/research/self-promoting-best-lists-study) and [When does Google show an AI Overview?](https://underneath.agency/research/ai-overviews-frequency-study)

---

This is the Markdown twin of https://underneath.agency/resources/project-management-software-customers-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How do we prove that GEO caused a change in sales or conversions?"
description: "You need a comparison group that got no GEO work, ideally chosen at random, because AI traffic grows for everyone. Before-and-after numbers overstate it."
canonical: "https://underneath.agency/resources/prove-geo-caused-sales"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How do we prove that GEO caused a change in sales or conversions?

You need a comparison: pages, products or markets that did not get the GEO work, measured over the same weeks, ideally chosen at random. AI traffic is growing for everyone, so a before-and-after rise mostly measures that growth. The clearest field test so far found a real-looking effect that shrank, and became uncertain, under closer checks.

## The short version

1. On one site, total ChatGPT referrals grew 5.7 times, but pages that got no optimization grew 3.5 times anyway (Watanabe and Nakayashiki, 2026).
2. After netting out that growth, the optimized pages gained about 1.82 times more ChatGPT referrals, an effect the authors call suggestive, not conclusive.
3. In a reanalysis of the same data, the estimate ranged from 1.183 to 2.404 times, depending on the assumed start week (Kato and colleagues, 2026).
4. In 500 simulated runs, a simple before-and-after comparison never captured the true effect when all units shared platform growth (Kato and colleagues, 2026).
5. A review of GEO research found traffic and conversions to be its weakest evidence; one reported 20% traffic lift lacked the detail to judge it (Martinez, 2026).

## Why can’t you just compare sales before and after the GEO work?

Because AI platforms are growing fast, so traffic from them rises whether or not you do anything.

Generative engine optimization, or GEO, is work to make AI engines name and cite you more. Proof matters more as buying moves into AI. A retail survey cited by [Wen and colleagues](https://arxiv.org/abs/2606.12439) found that 39% of 8,350 shoppers in 21 countries use AI for product discovery and related tasks.

The trouble is the tailwind. [Watanabe and Nakayashiki](https://arxiv.org/abs/2606.04362) optimized one section of their company’s website, glasp.co, in January 2026 and left the rest alone. [Total ChatGPT referrals grew 5.7 times](https://underneath.agency/resources/chatgpt-referral-growth-and-geo) on monthly figures. But the untouched pages grew 3.5 times over the same window. A naive reading would have credited the optimized pages’ 6.1-times growth entirely to the work.

Headline numbers can be even worse. Measured from the lowest day to a single peak day, the site’s ChatGPT traffic rose by a factor of 38.2. The authors themselves call that a fragile basis for any claim.

## What did the clearest field test so far find?

A likely but unproven effect. The optimized pages gained roughly 1.8 to 2.3 times relative to untouched pages.

The authors tracked the weekly ratio of optimized to untouched pages’ ChatGPT referrals, before and after the change. Because both groups shared the same site, analytics and platform growth, the ratio cancels much of the common rise. They found a jump of about 1.82 times at the time of the change, and other versions of the analysis all landed between 1.8 and 2.3 times.

Two cautions apply. First, a stricter check found that the months before the change contained random jumps nearly as large, so the authors call the effect suggestive, not conclusive. Second, this is one site, mostly one AI engine, and the authors work for the company that owns it. The optimized and untouched pages were also different kinds of content.

The study also checked for side effects on Google. Google clicks to the optimized pages fell about 25%, close to the roughly 20% fall across the whole site. The authors read that as [no measurable harm to Google traffic](https://underneath.agency/resources/chatgpt-optimization-google-rankings). Note that these outcomes are referrals and clicks, not sales.

## How fragile are estimates like this?

Very. Small, reasonable choices in the analysis can double the estimate or erase it.

[Kato and colleagues](https://arxiv.org/abs/2609.11915) reran the same public referral data. The rollout happened over several weeks, so the start date is a judgment call. Moving the assumed first week from December 16, 2025 to January 6, 2026 changed the estimated jump from 1.183 times to 2.404 times. For the earlier dates, the plausible range included no effect at all.

A second problem was measurement. In mid-March the site changed how it filtered out bot traffic. Around then, the engagement rate of ChatGPT visitors rose from 0.486 to 0.896, according to the original study. When Kato and colleagues allowed for that change, the estimates fell to between 1.337 and 1.455 times. In one version the estimate was 1.387 times, with a plausible range that included no effect.

## What does a proper test of GEO’s business effect need?

A comparison group that shares the same market changes, and ideally random assignment. Kato and colleagues spell out the conditions.

A randomized rollout gives a fair comparison by design: you pick, at random, which products, pages or regions get the GEO work. Without that, you need data from before the change and a comparison group measured at the same time. Both groups must also face the same shifts in the AI engines. In their simulation, across 500 runs where every unit shared platform growth, comparing treated units with themselves before and after never once captured the true effect. Comparing the change in treated units with the change in untreated ones removed the error.

They also show how to connect visibility to sales. How often answers mention you must be multiplied by how many relevant questions are asked, which AI systems people use, and how often readers notice your name. Each piece varies. In their own collection of 2,240 answers, one brand appeared in 33.8% of answers from one OpenAI model and 27.8% from another. The order reversed for English questions. Their method was tested on simulated sales data, not a real company’s.

## How strong is the evidence that GEO lifts sales today?

Weak. No study we reviewed shows GEO causing a measured change in sales or conversions.

In a review of the field, [Martinez](https://arxiv.org/abs/2607.14035) ranks traffic and conversions as the weakest evidence. One industry study reported a 20% production traffic lift against a control, but did not describe group sizes, how units were assigned or the uncertainty. The review found no technique with a proven lasting effect on downstream clicks and conversions. That does not mean GEO does not work. It means the claims have outrun the evidence, so your own test matters.

## What should you do about it?

Design the proof before you spend, and make it a fair comparison. A workable plan:

1. **Pick one business outcome** (leads, sign-ups or sales from AI referrals) and define it before the work starts.
2. **Split your scope.** Choose similar products, pages, categories or regions, and assign some to GEO work and some to wait, at random if you can.
3. **Collect at least several months of history** for both groups, so you can see their normal trends.
4. **Track both groups weekly** over the same period, and compare the change in one with the change in the other. [Server logs of AI bot visits](https://underneath.agency/resources/ai-bot-server-logs-content-demand) can show which pages in each group the bots request.
5. **Log every other change:** analytics settings, bot filtering, site launches, campaigns and price moves. Any of these can fake or hide an effect.
6. **Test the start date.** If the result only appears for one chosen week, treat it as unproven.
7. **Discount headline multiples** from vendors, agencies or case studies that show no comparison group. The same caution applies to [“best of” lists that rank their publisher first](https://underneath.agency/research/self-promoting-best-lists-study).

If you want help setting up a test like this, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The research has not yet measured GEO’s effect on sales or revenue for any real business. Open questions:

- **Sales, not referrals.** The one field test measured ChatGPT referral visits and Google clicks, not purchases.
- **More than one site.** That test covered a single domain and mostly one engine.
- **Which tactic works.** Its changes were a bundle, so no single tactic can be credited.
- **Whether people notice.** Linking AI mentions to sales needs data on how often readers notice a brand name, and few studies collect it.
- **Random trials.** No randomized rollout of GEO with business outcomes has been published.

## Frequently asked questions

### Can we attribute AI referral traffic growth to our GEO work?

Not without a comparison group. On one site, pages that received no optimization still grew their ChatGPT referrals 3.5 times over the same months.

### What is a holdout test for GEO?

It is a test where some comparable products, pages or regions are deliberately left out of the GEO work and measured alongside the rest. Assigning them at random gives the fairest comparison.

### How much did GEO increase traffic in the clearest study so far?

About 1.8 to 2.3 times more ChatGPT referrals relative to untouched pages, on one site. A stricter check rated that suggestive, and a reanalysis put it between 1.183 and 2.404 times depending on the start date.

### Do AI visibility tools measure sales impact?

Usually not. They measure mentions and citations, which still need to be linked to how many people ask, which engines they use and whether they notice your name.

## Sources

- Watanabe, K. and Nakayashiki, K. (2026), [Disentangling Answer Engine Optimization from Platform Growth: A Log-Based Natural Experiment on ChatGPT Referral Traffic](https://arxiv.org/abs/2606.04362), arXiv:2606.04362.
- Kato, M., Honma, D. and Kato, T. (2026), [Generative Marketing Mix Modeling: A Causal Inference Framework Linking GEO and GEM to Business Impact](https://arxiv.org/abs/2609.11915), arXiv:2609.11915.
- Martinez, O. (2026), [Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026)](https://arxiv.org/abs/2607.14035), arXiv:2607.14035.
- Wen, Y., Zhang, N., Yuan, H., Chen, X., Zhang, H. and Guo, H. (2026), [Position: Generative Engine Optimization Creates Underexamined Risks, Governance Must Target Concentration, Disclosure, and Academic Blind Spots](https://arxiv.org/abs/2606.12439), arXiv:2606.12439.

---

This is the Markdown twin of https://underneath.agency/resources/prove-geo-caused-sales. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Can AI search bring a recruiting agency more employer clients?"
description: "Yes, if employers can see what roles you fill, where, how fast and how well. AI answers lean on local listings, reviews and specialty pages to name agencies."
canonical: "https://underneath.agency/resources/recruiting-agencies-employer-clients-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Can AI search bring a recruiting agency more employer clients?

Yes, when an employer’s question matches what the agency can show in public: the roles it fills, the places it covers, how quickly it fills them and what clients say afterward. Staffing is a crowded, local and specialized market, and AI assistants now sit in the same chat where hiring managers draft job descriptions, so the agencies with clear, verifiable facts about themselves are the ones an assistant can recommend with confidence.

## The short version

1. The market is large and crowded: US staffing sales were $27.6 billion in the first quarter of 2026, according to the [American Staffing Association](https://aijourn.com/u-s-staffing-industrys-seasonal-declines-narrow-in-first-quarter-of-2026/), and ASA counts around 27,000 staffing and recruiting companies.
2. Clients judge on service: the staffing industry’s average client Net Promoter Score was 45% in 2024, and clients of ClearlyRated’s award winners were 50% more likely to report complete satisfaction.
3. Hiring now starts inside chat apps: Indeed launched an app for job seekers inside ChatGPT in February 2026, and Upwork followed in April with one that lets businesses describe a need and find independent talent.
4. Agencies are already using AI internally: in Bullhorn’s survey of nearly 2,300 recruitment professionals, top-performing firms were four times more likely to use AI.
5. Local signals carry weight: in our study, ChatGPT listed 67.7% of the businesses Google Maps ranked 1 to 3, and more reviews than the local median added 19.5 points.

This article covers agencies that fill temporary, contract and permanent roles for employers. Retained searches for chief executives and board members work differently, with a different buyer, fee and set of trust signals, so they are not covered here.

## Who hires a recruiting or staffing agency, and what is a client worth?

Hiring managers, HR and operations leaders who need people faster than they can hire alone. A good client sends repeat orders for years.

ASA’s [staffing industry statistics](https://americanstaffing.net/research/fact-sheets-analysis-staffing-industry-trends/staffing-industry-statistics/) show how broad the buyer base is. Nearly 2.2 million temporary and contract employees worked for US staffing companies in an average week in 2024, and staffing gave job opportunities to about 11 million people that year. The work spans 36% industrial, 24% office and administrative, 21% professional and managerial, 11% engineering, IT and scientific, and 8% health care. ASA says clients turn to staffing companies for two reasons: workforce flexibility and access to talent. Firms that advise on people problems rather than fill roles are covered in [how HR consultancies win clients](https://underneath.agency/resources/hr-consulting-firms-clients-ai-search).

That means several different buyers. A plant manager needs 40 seasonal workers by November. A controller needs a contract accountant for year-end close. A software company needs a recruiter for three hard-to-fill engineering roles. At large employers, procurement and contingent workforce programs pick suppliers through formal processes, where AI answers likely matter less; we infer that the direct buyer at a small or mid-size company is where AI search has most influence. Boards hiring a chief executive are a different buyer again, covered in [our guide to executive search firms](https://underneath.agency/resources/executive-search-firms-clients-ai-search).

The value of a client shows in public numbers. Korn Ferry’s [fiscal 2026 results](https://ir.kornferry.com/sec-filings/all-sec-filings/content/0000056679-26-000017/kfy-20260430xex991q4fy26.htm) report $222.4 million of permanent placement fee revenue from 4,835 engagements billed, and an average interim bill rate of $145 an hour. Smaller agencies earn less per placement, but the pattern is the same: one satisfied hiring manager becomes a stream of job orders.

## How is the staffing market changing in 2026?

Demand is stabilizing after two weak years, and buyers are judging agencies harder on speed and service.

ASA reported that temporary and contract employment fell 7.5%, or 154,000 jobs, from the fourth quarter of 2025 to the first quarter of 2026, a normal seasonal dip and the slowest first-quarter decline since 2022. The year-to-year decline in staffing employment narrowed to 4.6%, from 10.8% a year earlier. In a slow hiring market, every new client account counts.

Agencies are competing on speed. In [Bullhorn’s GRID 2026 survey](https://www.bullhorn.com/news-and-press/press-releases/bullhorn-grid-report-staffing-firms-using-ai-see-stronger-growth-faster-placements/), 56% of the highest-growth firms reported average placement times under 10 days, and 43% of firms expected only modest economic improvement in 2026. Bullhorn sells staffing software, so read its findings as a vendor’s survey.

They also compete on service. ClearlyRated, which runs client satisfaction surveys for staffing firms, said in its [2026 Best of Staffing announcement](https://www.aap.com.au/aapreleases/globenewswire9647226/) that clients of winning agencies were 50% more likely to report complete satisfaction than the industry average. One winner, [Eastridge Workforce Solutions](https://www.eastridge.com/blog/eastridge-workforce-solutions-earns-clearlyrateds-2026-best-of-staffing-client-award), cited the 2024 industry average client score of 45% against its own 83. Those ratings are published, and published ratings are evidence an AI answer can find.

## Where do AI assistants enter an employer’s search for an agency?

Increasingly at the start, inside the same chat where a manager writes the job description. Agency-specific evidence is still limited.

We found no study that measures how employers use AI to choose staffing agencies. The signals we did find point one way:

- **Business buyers use AI to research suppliers.** In a Gartner survey of 645 B2B buyers, [45% said they used generative AI](https://www.gartner.com/en/newsroom/press-releases/2026-05-20-gartner-survey-finds-sixty-nine-percent-of-b-two-b-buyers-turn-to-sales-reps-to-validate-ai-generated-insights) in a recent purchase, mainly to gather information on vendors. Those buyers were not specifically buying staffing.
- **Talent marketplaces now live inside ChatGPT.** Upwork [launched an app in ChatGPT](https://stocks.observer-reporter.com/observerreporter/article/gnwcq-2026-4-9-upworks-work-marketplace-comes-to-chatgpt) on April 9, 2026 that lets businesses describe a project, discover talent among more than 18 million professionals and draft a job post. Clutch, a B2B services directory, [launched its own ChatGPT app](https://www.demandgenreport.com/?p=52857) a month later.
- **Hiring platforms are moving into AI.** Indeed’s Talent Scout, reported by [HR Grapevine](https://www.hrgrapevine.com/us/content/article/2025-09-19-openai-indeed-unveil-rival-ai-hiring-platforms), generates candidate shortlists for HR leaders, and OpenAI announced its own jobs platform for mid-2026.

For an agency, these are competitors as much as channels. A manager who can describe a role in ChatGPT and see freelancers or job board candidates may never ask for an agency. A reasonable expectation is that agencies win these conversations when the request is urgent, specialized, local or high-volume, which is where their own pages should be strongest.

## What do hiring managers ask AI about staffing agencies?

Questions that combine a role, a place, a timeline and a worry. We wrote the prompts below to show the pattern; none comes from a real search log.

| Buyer | Example prompt |
|---|---|
| Plant manager | “Staffing agencies in Fort Worth that can supply 40 forklift-certified workers for second shift by November” |
| Controller | “Accounting staffing firms for a contract senior accountant during year-end close, remote OK” |
| Clinic director | “Agencies that place travel or per diem nurses in Phoenix, and what they charge” |
| Software CTO | “Contingency recruiters for senior backend engineers: who specializes in fintech?” |
| HR manager | “Temp-to-hire vs direct hire through an agency: what are the usual fee terms?” |
| Any buyer | “Is [agency name] reliable? What do clients and workers say?” |

The pattern matters. Most questions name a specialty and a location, and the last is a reputation check run before a call. Generic pages that say “we staff all industries” give an assistant little to match against those details.

## How does an AI mention become a job order?

Through a short path: the answer names the agency, the manager checks it, then calls or sends a request.

The path is AI answer → branch or specialty page → reviews and ratings → phone call or “request talent” form → first job order → repeat orders. For temporary and contract work, a client account bills every week a worker is on assignment. For permanent placement, each hire earns a fee. Either way the first order is a trial, and the client’s experience decides whether more follow, which is why service ratings keep showing up in this article.

Speed matters inside the path. Our [hidden searches study](https://underneath.agency/research/ai-hidden-searches-study) found ChatGPT looked for reviews in 46.2% of its answers, so a manager in a hurry may see a summary of client and worker reviews before ever reaching the agency’s site.

## What decides whether an assistant names a staffing agency?

For local requests, the same signals that drive local search; for specialist requests, independent evidence of the specialty. Platforms document only part of this.

**Documented by the platforms.** Google says [local results are mainly based on relevance, distance and popularity](https://support.google.com/business/answer/7091), and that complete business information helps it match a Business Profile to searches. OpenAI’s help page on [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search) says ChatGPT may use an approximate location from the user’s IP address to give local results.

**Observed in our studies (other local professions, not staffing).** In our [study of Google Maps businesses](https://underneath.agency/research/chatgpt-local-picks-google-profile-study), ChatGPT listed 67.7% of businesses ranked 1 to 3 in Maps, and those with more reviews than the local median were 19.5 points more likely to be listed after adjusting for rank. When asked whether a brand was legitimate, [88.0% of AI answers cited a review or complaint platform](https://underneath.agency/research/is-it-legit-ai-reputation-study).

**Our inference for staffing.** An agency with several branches is really several local businesses. Each branch needs its own accurate profile, its own reviews and a page that says which roles it fills nearby. For specialist recruiting, such as nurses, engineers or accountants, the stronger evidence is likely independent: association memberships, ratings on directories like ClearlyRated, coverage in trade press, and lists that name the agency for that specialty. Other specialist firms face the same test, as our guide to [how engineering firms get named by AI](https://underneath.agency/resources/engineering-firms-project-inquiries-ai-search) shows.

## Does AI search matter for finding candidates too?

Yes, as a second front: agencies need people to place, and job seekers now search with AI.

In a Resume Now survey of 1,023 US workers reported by [Staffing Industry Analysts](https://www.staffingindustry.com/news/global-daily-news/ai-transforms-job-search-even-as-it-ups-competition-survey-says), 80% said they use AI-powered job search platforms, and ChatGPT was named by 58% as a tool they use. According to [HeroHunt](https://www.herohunt.ai/blog/openai-jobs-platform-2026-recruiter-playbook/), a recruiting software company, Indeed launched an app inside ChatGPT on February 10, 2026, while OpenAI’s standalone jobs marketplace had slipped past its mid-2026 target as of August 2026.

Trust is the issue on this side. ASA’s Workforce Monitor found nearly half of employed US job seekers (49%) believe AI tools used in recruiting are more biased than humans. Clear job pages, honest pay ranges where they are known, and visible reviews from placed workers help an agency with both candidates and the assistants that summarize it.

## What does it cost an agency to be invisible in AI answers?

Missed first orders from new clients, and lost ground to platforms that are already inside the chat. No one has measured the size yet.

The direct cost is invisible: a manager who asks for a local agency and gets three competitors never calls. The structural cost is clearer. Upwork, Indeed and Clutch have each built a way to appear inside ChatGPT, so a hiring conversation can end on a platform without an agency involved. That is our inference from their launches, not a measured shift in demand.

Being described wrongly also costs. In our [business facts study](https://underneath.agency/research/ai-business-facts-accuracy-study), 18.9% of AI answers stated at least one fact that differed from the business’s Google profile, mostly where the business’s own sources disagreed. An agency that has moved a branch or changed its phone line should expect old details to linger. Our guide on [fixing wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) explains the repair.

## How does GEO work for a recruiting or staffing agency?

Generative engine optimization (GEO) makes an agency’s branches, specialties and record easy for assistants to find and repeat. No one can guarantee a recommendation.

1. **A page per branch and per specialty.** Name the roles filled, the cities and sites served, shifts covered, typical time to first candidate, and the people to call. Keep the facts identical to the branch’s Google Business Profile.
2. **Reviews from both sides.** Ask clients and placed workers for reviews on the platforms buyers check, and answer criticism. Our [guide to how small brands get recommended by AI](https://underneath.agency/resources/how-small-brands-get-recommended-by-ai) explains why independent evidence counts.
3. **Third-party ratings and listings.** Service ratings such as ClearlyRated, association membership, and specialty directories put the agency’s name next to the roles it fills.
4. **Proof of the specialty.** Short case notes, with permission, on fills by role family and region, plus salary and market guides that trade press and other sites quote. Guides that other sites quote are the kind of standing described in [how brands build authority for AI search](https://underneath.agency/resources/how-brands-build-authority-for-ai-search).
5. **Plain fee and process pages.** How temp-to-hire works, what contingency terms usually involve, and what the client must provide. Buyers ask these questions; the agency should be a source for the answer.
6. **Candidate-facing clarity.** Current job pages, clear pay information where it is known, and honest descriptions of how the agency uses AI in screening.
7. **Regular checks.** Ask the role-and-place questions above in ChatGPT, Gemini, Perplexity and Google’s AI features each quarter, more than once each, and record which agencies are named and why.

## What are the limits of what we know?

The evidence shows how buyers and assistants behave in general, not how often an AI answer becomes a job order.

- **No staffing study of AI answers.** We found none that measures which agencies assistants name.
- **Our local studies cover other professions.** The Maps and review findings come from businesses like dentists and plumbers and may differ for agencies.
- **Several figures come from vendors.** Bullhorn and ClearlyRated sell to staffing firms; Resume Now sells resume tools.
- **Platforms are moving fast.** OpenAI’s jobs marketplace was not live as of August 2026, and the apps inside ChatGPT are new.

## Where should a recruiting agency start?

With your three most profitable role families and the branches that serve them, checked against what AI says today.

Write the questions a hiring manager would ask for each: role, place, timeline, worry. Ask them in several assistants on different days. Note which agencies appear, which pages and review sites are cited, and whether your branches’ details are right. Then fix what is wrong before adding anything new.

If you want more first job orders from employers who now start in a chat window, [ask us to review how AI assistants present your agency](https://underneath.agency/contact). We will test your role and branch questions, show which sources the answers draw on, and plan the pages, profiles and reviews that help assistants match your agency to the work you do best. For the full scope, from branch and specialty pages to client and worker reviews and quarterly checks, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## Frequently asked questions

### Does ChatGPT recommend local staffing agencies?

It can. OpenAI says ChatGPT may use approximate location for local results, and in our studies of other local businesses it named many of the top-ranked Google Maps listings.

### Do reviews from placed workers matter, or only client reviews?

Both are public evidence. AI answers to reputation questions cite review platforms often, and workers’ reviews shape how an agency is described to candidates as well as clients.

### Will AI hiring platforms replace staffing agencies?

There is no evidence of that yet. Upwork and Indeed have apps in ChatGPT, but agencies still compete on urgent, specialized and high-volume work.

### How quickly can an agency change what AI says about it?

Facts on its own pages and profiles can change within weeks once corrected. Reputation and third-party coverage build over months.

## Sources

- American Staffing Association via AI Journal (2026-06-25), [U.S. Staffing Industry’s Seasonal Declines Narrow in First Quarter of 2026](https://aijourn.com/u-s-staffing-industrys-seasonal-declines-narrow-in-first-quarter-of-2026/)
- American Staffing Association (2026), [Staffing Industry Statistics](https://americanstaffing.net/research/fact-sheets-analysis-staffing-industry-trends/staffing-industry-statistics/)
- Korn Ferry (2026-06-23), [Fourth quarter and full year FY’26 results](https://ir.kornferry.com/sec-filings/all-sec-filings/content/0000056679-26-000017/kfy-20260430xex991q4fy26.htm)
- Bullhorn (2026-02-25), [Bullhorn GRID report: Staffing firms using AI see stronger growth, faster placements](https://www.bullhorn.com/news-and-press/press-releases/bullhorn-grid-report-staffing-firms-using-ai-see-stronger-growth-faster-placements/)
- ClearlyRated via GlobeNewswire (2026-02-03), [2026 Best of Staffing Award Winners Achieve Superior Ratings](https://www.aap.com.au/aapreleases/globenewswire9647226/)
- Eastridge Workforce Solutions (2026-05-18), [Eastridge Earns ClearlyRated’s 2026 Best of Staffing Client Award](https://www.eastridge.com/blog/eastridge-workforce-solutions-earns-clearlyrateds-2026-best-of-staffing-client-award)
- Gartner (2026-05-20), [Gartner Survey Finds 69% of B2B Buyers Turn to Sales Reps to Validate AI-Generated Insights](https://www.gartner.com/en/newsroom/press-releases/2026-05-20-gartner-survey-finds-sixty-nine-percent-of-b-two-b-buyers-turn-to-sales-reps-to-validate-ai-generated-insights)
- Upwork via GlobeNewswire (2026-04-09), [Upwork’s Work Marketplace Comes to ChatGPT](https://stocks.observer-reporter.com/observerreporter/article/gnwcq-2026-4-9-upworks-work-marketplace-comes-to-chatgpt)
- Demand Gen Report (2026-05-12), [Clutch Launches First B2B Services Marketplace App on ChatGPT](https://www.demandgenreport.com/?p=52857)
- HR Grapevine (2025-09-19), [OpenAI and Indeed unveil rival AI hiring platforms](https://www.hrgrapevine.com/us/content/article/2025-09-19-openai-indeed-unveil-rival-ai-hiring-platforms)
- HeroHunt (2026-08), [OpenAI Jobs Platform 2026: recruiter playbook](https://www.herohunt.ai/blog/openai-jobs-platform-2026-recruiter-playbook/)
- Staffing Industry Analysts (2025-03-10), [AI transforms job search even as it ups competition, survey says](https://www.staffingindustry.com/news/global-daily-news/ai-transforms-job-search-even-as-it-ups-competition-survey-says)
- Google Business Profile Help (n.d.), [Tips to improve your local ranking on Google](https://support.google.com/business/answer/7091)
- OpenAI (2026), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Underneath (2026), [Which Google Maps businesses does ChatGPT recommend?](https://underneath.agency/research/chatgpt-local-picks-google-profile-study)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [“Is this brand legit?” How AI assistants build a reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study)
- Underneath (2026), [Do AI answers match a business’s Google profile?](https://underneath.agency/research/ai-business-facts-accuracy-study)

---

This is the Markdown twin of https://underneath.agency/resources/recruiting-agencies-employer-clients-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "The risks of native ads inside AI chatbot answers"
description: "Native ads in AI answers are hard to tell apart from advice, so they can mislead trusting users and damage the advertiser’s brand once spotted."
canonical: "https://underneath.agency/resources/risks-of-native-ads-in-ai-answers"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# What are the risks of native ads blended into AI chatbot answers?

The main risk is that people cannot tell a paid placement from honest advice, so a blended ad borrows trust it has not earned. That hurts users first, especially people asking about health or money. It can then hurt the brand that paid for the placement, because covert promotion tends to breed distrust once it is noticed.

## The short version

1. Ads are arriving. A 2026 audit by [Uberti-Bona Marin and colleagues](https://arxiv.org/abs/2609.18729) reports that OpenAI expanded ChatGPT advertising to 31 European markets, where ads are most likely on product questions.
2. In that audit, ChatGPT stated a personal, first-person product preference in 79% of the answers that recommended a product. Any blended ad in that voice reads like a friend’s tip.
3. In an experiment with 4,927 US adults, [Li and Aral](https://arxiv.org/abs/2504.06435) found that adding reference links raised trust in AI answers even when the links were wrong.
4. A US poll cited by [Wen and colleagues](https://arxiv.org/abs/2606.12439) found 60% of 1,437 adults use AI to find information at least some of the time, so the audience exposed is large.
5. Most of the harm is still argued, not measured: the core paper on native ads in chatbots, by [Erickson](https://arxiv.org/abs/2506.06447), uses prompted examples, not real ads.

## What is a native ad inside an AI answer?

It is a paid mention written into the answer itself, so it reads like the assistant’s own advice. [Erickson](https://arxiv.org/abs/2506.06447) contrasts two forms. A banner ad sits apart from the answer and is easy to recognize as an ad. A native ad is woven into the reply, and it may or may not be disclosed.

To show what that could look like, Erickson prompted ChatGPT to role-play a user with depressive symptoms. The prompted replies slipped in a soft drink, a brand-name antidepressant, and, in the worst case, a vodka brand as a “quick boost.” These were staged on purpose. The paper says plainly that it does not claim such responses are already happening organically.

## Are ads really coming to AI answers?

Yes: ads have started appearing, and product questions are where they are most likely to show up. The audit by [Uberti-Bona Marin and colleagues](https://arxiv.org/abs/2609.18729) reports that ChatGPT advertising expanded to 31 European markets in August 2026. Ads appear in between 10 and 14% of sessions for product recommendation queries. That figure comes from another study the audit cites, not from the audit’s own measurement. For the wider picture, see [how ads are reaching AI assistants](https://underneath.agency/resources/are-ads-coming-to-ai-assistants).

Product questions are not a niche use. The same audit notes that product and service recommendations already made up roughly 2% of ChatGPT conversations in a large 2025 measurement. And a survey cited by [Wen and colleagues](https://arxiv.org/abs/2606.12439) found that 39% of 8,350 shoppers across 21 countries use AI for product discovery and related shopping tasks.

Google’s conversational search was still ad-free when [Wang and colleagues](https://arxiv.org/abs/2608.18352) ran a field experiment with 1,100 US Google users in March 2026. Sending people to AI Mode cut the share who clicked any ad by 42.7 percentage points, simply because no ads were shown. The authors expect user satisfaction to shift once Google monetizes it.

## Why are blended ads riskier than banner ads?

Because users cannot easily spot them, and AI answers already carry a tone of personal, confident advice. Erickson points to earlier research showing that people are more skeptical when they know someone is trying to persuade them. An ad that hides its purpose skips that skepticism. That is what makes it effective, and also what makes it deceptive.

The voice of the answer adds to the problem. The audit found clear differences in how three Google and OpenAI products frame advice:

| Product | Answers with a first-person pick (“my pick is…”) | Answers labeling a product “best” |
|---|---|---|
| ChatGPT | 79% | 73% |
| Gemini | 7% | 43% |
| Google AI Overviews | 2% | 48% |

These shares come from 117 real shopping questions asked from the Netherlands. A paid product that appears inside a “my pick” sentence would read as a personal endorsement.

Trust is also easy to inflate. In the experiment by [Li and Aral](https://arxiv.org/abs/2504.06435), with 4,927 participants matched to the US adult population, people trusted AI search less than ordinary search on average. But adding reference links raised trust, even when the links were incorrect or invented. The look of credibility, in other words, can move people more than the substance.

## Who is most exposed to the harm?

People in distress and people asking about health, money or other high-stakes choices. Erickson argues that native ads could be “especially corrosive” for users who are struggling and who trust a system with commercial motives. He calls this the “fake friend dilemma”: the user believes the assistant is on their side when it is serving someone else.

He singles out older adults, children and people in acute distress. He also argues for tobacco-style limits on where harmful or expert-only products, such as prescription drugs, can be advertised. These are policy arguments from a position paper, not findings from a study of real ads.

The reach is wide. In the poll cited by [Wen and colleagues](https://arxiv.org/abs/2606.12439), 60% of 1,437 US adults said they use AI to find information at least some of the time.

## What is the risk for the brand that buys the placement?

A brand that is caught hiding an ad in an answer can lose trust and invite backlash. Erickson cites earlier advertising research: when users recognize a disguised ad, trust drops and views of the brand can turn negative. Ads that cross a moral line, or personal targeting that feels like a privacy breach, can make the reaction worse. Our study of [how AI assistants judge whether a brand is legit](https://underneath.agency/research/is-it-legit-ai-reputation-study) shows how that reputation is built from evidence.

[Wen and colleagues](https://arxiv.org/abs/2606.12439) describe a second, quieter risk. If hidden promotion beats labeled ads, honest advertisers lose ground to covert ones, and the whole channel drifts toward hidden influence. They warn of “trust erosion” when such influence is later revealed.

There is also a risk to the platform, and so to the value of the placement. In the Google experiment by [Wang and colleagues](https://arxiv.org/abs/2608.18352), being moved to AI Mode cut trust in information on Google by 0.34 points on a 7-point scale. It also drove an 11.2% increase in use of competing search engines. That test was about AI Mode, not ads, but it shows users do react when an experience feels worse.

A modeling study by [Zhang and colleagues](https://arxiv.org/abs/2603.29071) points the same way. In simulations of 500 users over 20 periods, ad-heavy policies raised early revenue but shrank the user base and slowed paid sign-ups. This is a theoretical model, not observed behavior.

## What should you do about it?

Treat any paid placement inside an AI answer as a reputational decision, not just a media buy. Practical steps:

1. Only buy AI placements that the platform [labels clearly as sponsored](https://underneath.agency/resources/should-ai-search-ads-be-labeled), and ask the platform to show you exactly how the label will appear.
2. Avoid placements next to sensitive questions, such as mental health, medication or debt, unless you are confident the context is appropriate.
3. Write ad copy that could survive a screenshot: no claims you could not defend if a user, journalist or regulator saw it out of context.
4. Keep your earned visibility honest. Clear, verifiable facts on your own site and in independent coverage are the content AI answers can cite without paid help. That earned base is also the core of a [long-term AI search strategy](https://underneath.agency/resources/ai-search-strategy-next-five-years).
5. Check what AI assistants say about your category regularly, since the products they recommend change from run to run.

If you want help making your brand easier for AI assistants to describe accurately, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The research has not yet measured how real users react to real ads inside AI answers. Specifically:

- The best-known paper on native ads in chatbots uses staged examples, not observed ads.
- The 10 to 14% ad frequency for ChatGPT is a figure one audit cites from another study; it is not independently confirmed here.
- No study we reviewed measures how often users notice a blended ad, or how a brand’s reputation changes after one is exposed.
- The trust experiments tested citations and AI labels, not ads; the link to advertising is an inference.
- The modeling study of ad-heavy versus ad-free answers is a simulation with assumed user behavior.

## Frequently asked questions

### Are there ads in ChatGPT answers now?

Yes, in some markets. An audit reports that OpenAI expanded ChatGPT advertising to 31 European markets in August 2026, most often on product recommendation questions.

### Can people tell an ad from an AI recommendation?

Often they cannot, unless it is clearly labeled. Erickson cites earlier user research showing that people may not notice ads woven into AI search answers unless the ads are explicitly disclosed.

### Is native advertising in AI answers legal?

The law is unsettled. Erickson notes that it is unclear how deceptive advertising rules apply to conversational search. Wen and colleagues call for extending FTC-style disclosure rules to AI answers.

### Should my brand advertise inside AI answers?

Only where the placement is clearly labeled and the context is not sensitive. The research suggests disguised promotion carries a backlash risk that labeled ads avoid, though this has not yet been tested with real AI ads.

## Sources

- Erickson, J. (2025), [Fake Friends and Sponsored Ads: The Risks of Advertising in Conversational Search](https://arxiv.org/abs/2506.06447), arXiv:2506.06447.
- Uberti-Bona Marin, L. G., Bertaglia, T., Astante, G., Rijsbosch, B., van Dijck, G., Hannák, A., Spanakis, G. and Kollnig, K. (2026), ["If I Had to Buy Just ONE: Galaxy S26 Ultra": Auditing AI-Generated Product Recommendations](https://arxiv.org/abs/2609.18729), arXiv:2609.18729.
- Li, H. and Aral, S. (2025), [Human Trust in AI Search: A Large-Scale Experiment](https://arxiv.org/abs/2504.06435), arXiv:2504.06435.
- Wen, Y., Zhang, N., Yuan, H., Chen, X., Zhang, H. and Guo, H. (2026), [Position: Generative Engine Optimization Creates Underexamined Risks, Governance Must Target Concentration, Disclosure, and Academic Blind Spots](https://arxiv.org/abs/2606.12439), arXiv:2606.12439.
- Wang, S. T., Gleason, J., Bart, Y., Wilson, C. and Metaxa, D. (2026), [AI in Search Reduces Publisher Referrals Without Improving User Experience: Experimental Evidence](https://arxiv.org/abs/2608.18352), arXiv:2608.18352.
- Zhang, L., Jiao, C., Li, B. and Xiong, C. (2026), [An Economic Framework for Generative Engines: Advertising or Subscription?](https://arxiv.org/abs/2603.29071), arXiv:2603.29071.

---

This is the Markdown twin of https://underneath.agency/resources/risks-of-native-ads-in-ai-answers. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How robot makers win factory buyers through AI search"
description: "It can, if assistants find your applications, specs, integrators and real deployments in sources they trust, as robot buying spreads beyond automotive."
canonical: "https://underneath.agency/resources/robotics-companies-customers-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Will AI assistants put our robots on a manufacturer’s shortlist before an integrator is called?

They can, when a buyer’s question matches an application you document and outside sources confirm it. Robot demand is spreading to industries with little automation history, and those first-time buyers research on their own before they call a robot maker or an integrator.

This guide is for executives at robot makers: industrial arms, collaborative robots (cobots), autonomous mobile robots and the first commercial humanoids. It treats system integrators as a sales channel. Integrators themselves, and warehouse automation systems, have their own articles in this series.

## The short version

1. The market is growing again: the [International Federation of Robotics (IFR)](https://ifr.org/img/worldrobotics/Executive_Summary_WR_2026_-_Industrial_Robots.pdf) counted 603,307 industrial robots installed worldwide in 2025, up 11%, and forecasts 655,000 in 2026. The [US installed 38,400](https://ifr.org/downloads/press_docs/EN-2026-SEP-24-IFR_Press_Release_WR-USA.pdf), up 12%, and is now the second-largest market after China.
2. New industries are buying: in the first half of 2026, North American robot orders from automotive manufacturers fell 25% while semiconductor and electronics orders rose 35% and life sciences 32%, according to the [Association for Advancing Automation (A3)](https://www.automate.org/robotics/news/robot-orders-increase-in-q2-as-automation-demand-broadens-across-industries).
3. Cobots are a large share of new buyers’ orders: cobots made up 15.4% of North American robot units ordered in the first half of 2026, and 43.7% in life sciences.
4. Mobile and service robots are growing fastest: the IFR reported professional service robot sales up 24% to almost 250,000 units in 2025, with transport and logistics robots at 117,500.
5. Reliability decides selection: in [PMMI’s 2026 robotics study](https://www.pmmi.org/report/2026-robotics-in-packaging-and-processing), 72% of packaging and processing end users already used robots, and 69% rated reliability and uptime as very important when choosing a supplier.

## Who buys robots, and what is a customer worth?

Engineers and operations leaders at manufacturers, usually buying through an integrator, with fleet orders following a first success.

A robot purchase usually starts with an operations problem: a labor gap, a quality issue, a safety risk or a throughput target. A manufacturing or automation engineer scopes the application, an operations or plant leader owns the business case, and finance approves the capital spend. In many projects a system integrator designs the cell and installs it, so the integrator often decides which robot brand goes in. Other capital equipment follows a similar path; see [how machinery makers get shortlisted](https://underneath.agency/resources/machinery-companies-buyers-ai-search).

The size of the market is public. The IFR put the global market value of industrial robot installations at US$18.6bn in 2025. In North America, A3 counted orders for 17,995 robots worth $1.166 billion in the first half of 2026, 6.6% more by value than a year earlier.

What one customer is worth is not public, and it varies widely: one cobot for machine tending, or hundreds of robots across plants. The pattern that matters is expansion. Our inference: a robot maker’s best customers are plants that succeed with a first cell and then repeat it, so the first shortlist decision carries the value of later fleet orders and service contracts. The IFR adds that robot-as-a-service (RaaS), where customers pay over time instead of buying outright, is already the dominant model for professional service robots in the US. Robots sold into distribution centers meet a different buyer, covered in [how warehouse automation vendors make shortlists](https://underneath.agency/resources/warehouse-automation-enterprise-leads-ai-search).

Competition is heavy. Teradyne, which owns Universal Robots and the mobile robot maker MiR, reported robotics revenue of $100 million in the second quarter of 2026, up 33% from $75 million, [according to The Robot Report](https://www.therobotreport.com/teradyne-robotics-revenue-rises-33-year-over-year-in-q2/). But that unit’s annual revenue had fallen from a peak of $326 million in 2022 to $293 million in 2024. Growth is available, not guaranteed.

## Where does AI search already sit in robot buying?

In early research, by engineers who use AI but verify it with suppliers, peers and video.

The best evidence on technical buyers comes from the 2026 [State of Marketing to Engineers](https://advertising.globalspec.com/state-of-marketing-to-engineers-research-report) survey by TREW Marketing and GlobalSpec, with Elektor. It found that 69% of technical buyers use generative AI during the purchasing process, but they rate their trust in its answers at 4.7 out of 10. On average, 62% of the buying process happens online before they talk to a vendor.

Across all business purchases, a [Gartner survey of 645 buyers](https://www.gartner.com/en/newsroom/press-releases/2026-05-20-gartner-survey-finds-sixty-nine-percent-of-b-two-b-buyers-turn-to-sales-reps-to-validate-ai-generated-insights) found that 45% used generative AI in a recent purchase, mainly to gather information on vendors and products, and that 69% prefer to validate AI-generated insights with sales reps.

Robot buying adds two habits. First, buyers want to see the robot work. In [our study of YouTube citations in Google’s AI Overviews](https://underneath.agency/research/ai-overview-youtube-videos-study), only 15.1% of cited videos were ones Google also showed on page one of the same search, and the text Google displayed came from what was said in the video, not its description. Our inference: a captioned application video from a robot maker or integrator can be read as evidence, even if it has few views. Second, industry associations and trade press carry weight: the IFR, A3 and PMMI publish the numbers that journalists and assistants repeat.

## What do robot buyers ask AI assistants?

Application questions first, then brand comparisons, business model and support questions.

The questions below are written by us to show typical patterns. They are examples, not records of real searches.

| Stage | Illustrative question |
|---|---|
| Application fit | “Best collaborative robot for palletizing 20 kg cases at the end of a packaging line” |
| Robot type | “Cobot or industrial robot for tending two CNC lathes?” |
| Mobile robots | “Mobile robots that move totes between work cells without changing the floor” |
| Emerging tech | “Are humanoid robots actually working in factories yet?” |
| Business model | “Robot as a service versus buying robots for a small food plant” |
| Support | “Robot brands with service engineers and spare parts in the Midwest” |
| Comparison | “Compare these two cobot brands on payload, reach and ease of programming” |

Two things stand out. Many questions describe a task, not a product, so a robot maker named only for its model numbers can miss them. And new buyers ask basic questions: the A3 data shows growth in semiconductor, pharmaceutical and food plants, many with smaller automation teams than an automaker. Our inference is that these buyers lean harder on outside research, including AI answers, before they know which integrator to call.

The humanoid question shows the risk of hype. The IFR estimates that about 7,000 full-size humanoids were sold in 2025 for commercial and professional uses, and its industrial robot report says most humanoids remain prototypes or pilots. An assistant that repeats press releases can leave a buyer with the wrong idea about what is ready today, in either direction.

## How does an AI answer become a robot order?

Through a brand shortlist, an integrator conversation, an application test, a pilot cell and then fleet orders.

1. **Shortlist.** The engineer asks which robots suit the task. Brands named in the answer, and in the pages it cites, become the starting list.
2. **Check.** The engineer looks for specifications, application videos, safety information and nearby integrators or distributors. Missing information ends consideration quietly.
3. **Integrator.** An integrator is brought in, or asked which brand it recommends. [Integrators research too](https://underneath.agency/resources/automation-integrators-industrial-buyers-ai-search), so the same visibility matters at this step.
4. **Application test and pilot.** The robot maker or integrator tests the part, cycle time and gripper. A first cell is installed.
5. **Fleet and service.** If the pilot meets its targets, the plant repeats it, signs service or RaaS terms, and may standardize on the brand.

AI visibility can influence steps 1 to 3. Cycle time, price, support and the integrator relationship decide the rest. In PMMI’s study, integration, serviceability and total value drove adoption decisions, which is why those facts belong in what an assistant can find.

## What decides whether an assistant names your robots?

Outside coverage it can find and specifications it can check; the platforms document their search, not their choices.

Each platform explains its search but not its picks. According to [OpenAI](https://help.openai.com/en/articles/9237897-chatgpt-search), ChatGPT search breaks an engineer’s question into one or more narrower queries for its search providers, and a robot maker’s site is eligible only if it admits OpenAI’s crawler, OAI-SearchBot. [Google](https://blog.google/products/search/ai-mode-search/) says AI Mode fans a question out into related searches on its subtopics, which it calls “query fan-out.” Neither publishes how a robot brand is chosen.

Observed in our studies:

- **Assistants look for rankings and awards.** In [our hidden searches study](https://underneath.agency/research/ai-hidden-searches-study), ChatGPT ran a search aimed at a named publication, ranking or award in 43.8% of its answers.
- **Many “best” lists are self-serving.** In [our study of cited best-of lists](https://underneath.agency/research/self-promoting-best-lists-study), 24.2% of numbered lists with an identifiable publisher ranked their own publisher first. Robot makers that publish their own rankings should expect buyers, and possibly assistants, to discount them.
- **Brand familiarity still tips decisions.** In the engineers survey, 70% were likely to pick the better-known brand when two solutions look technically similar, and 53% said familiarity influenced their most recent purchase.

Our inference for robot makers: the trust signals are the ones an engineer already checks. Payload, reach, repeatability and speed for each model; the applications it is proven in; safety standards and certifications; integrator and service coverage by region; and reference deployments that customers allow you to name. The IFR notes that local engineering, system integration and service infrastructure matter more as production regionalizes, and it says most robots installed in the US are still imported from Japan and Europe. Where your support is, and who installs your robots, are facts buyers ask about. Plant software vendors face similar checks, covered in [how MES and IIoT platforms get found](https://underneath.agency/resources/industrial-software-plants-ai-search).

## What does a robot maker lose when assistants leave it out?

First-time buyers in growing industries, and the integrators who standardize on a brand they find first.

We found no measurement of robot sales lost to AI absence, so this is reasoning, labeled as such:

- **The growth is in new hands.** Non-automotive customers made up 56% of robot units ordered in North America in the second quarter of 2026. Our inference: many of these plants have no incumbent robot brand, so early research decides who is considered.
- **Brand loyalty forms after the first cell.** A plant that standardizes on one brand for training and spare parts is hard to win later. Missing the first shortlist can cost the fleet.
- **New models lag in AI answers.** Assistants often answer from older training data. Our guide on [why ChatGPT misses new products](https://underneath.agency/resources/why-chatgpt-misses-new-products) explains why a new cobot or mobile robot can be invisible for months.
- **Errors spread.** A wrong payload, a discontinued model or an outdated support region in an answer filters you out. Our guide to [fixing wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) walks through the correction.

## How does GEO work for a robotics company?

Generative engine optimization (GEO) helps AI assistants find, describe accurately and verify your robots for the tasks buyers ask about.

For robot makers, the work usually covers:

1. **Application pages.** One page per task, such as palletizing, machine tending, welding, inspection or material transport, with the models that fit, typical cycle times, payload limits and industries served, written in text.
2. **Specification facts that match everywhere.** The same figures for payload, reach, repeatability and safety ratings on your site, distributor pages, datasheets and marketplaces.
3. **Integrator and service directory.** Who installs and supports your robots, by region, in a page an assistant can read.
4. **Captioned video.** Application demos with spoken explanations and accurate captions, since AI Overviews quote what is said in a video.
5. **Independent coverage.** Trade press, association data, conference talks, awards and case studies published with customer permission. Our article on [how brands build authority for AI search](https://underneath.agency/resources/how-brands-build-authority-for-ai-search) covers the outlets that matter, and [why “best of” lists matter](https://underneath.agency/resources/best-of-lists-ai-recommendations) explains why third-party rankings carry weight.
6. **Honest answers to hard questions.** When a cobot is the wrong choice, what RaaS costs over time, and what your humanoid or AI features can and cannot do today.
7. **Measurement.** Ask a fixed set of application, comparison and support questions across ChatGPT, Gemini, Perplexity, Copilot and Google’s AI features, repeatedly, and record who is named and which pages are cited. Then compare with demo requests and integrator leads.

None of this guarantees a recommendation. It makes your robots easier to verify for the tasks you actually do well.

## What is still unknown about AI and robot purchases?

How many robot orders begin with an AI answer, and how assistants handle robotics questions specifically.

- **No attribution data.** The surveys show that technical buyers use AI. None links AI answers to robot orders.
- **Sources have interests.** The IFR, A3 and PMMI represent industry members, TREW and GlobalSpec sell marketing to engineering firms, and Teradyne reports its own results.
- **Our studies were broad.** They covered buyer questions across several industries, not robotics. Applying them here is our inference.
- **Integrator influence is unmeasured.** We found no public data on how often integrators pick the brand, or how they use AI themselves.

## Where should a robotics company start?

Begin with the ten applications you most want to win, and ask assistants which robots suit each one.

Phrase the questions as a plant engineer would, with the task, payload, industry and region. Run them across the main assistants and note whether your robots appear, whether the specifications are right, which integrators and sources are cited, and which brands are named instead. That shows where buyers and integrators are being steered, and what evidence is missing.

If your growth depends on more qualified demos, application tests and pilot cells, [book a review of your robots’ AI visibility with us](https://underneath.agency/contact). We will show where your robots appear when engineers and integrators ask AI for options, why other brands are chosen, and which changes are most likely to bring in the right projects. How the ongoing work is run, from application pages and spec consistency to captioned video, is set out on our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page.

## Frequently asked questions

### Do integrators use AI assistants to choose robot brands?

We found no public data on that. Integrators research applications and components like other engineers, so it is reasonable to expect some do. Ask your integrator partners directly.

### Should we publish our own “best cobots” list?

It may help buyers, but self-ranking lists are common and easy to discount. Of the numbered lists assistants cited in our study, 24.2% put their own publisher at the top. Independent rankings and reviews are stronger evidence.

### Will humanoid robot news crowd our products out of AI answers?

It can shape broad questions. Task-specific questions with payload, cycle time and industry details are where proven robots are more likely to appear.

### How soon will assistants reflect a new robot model or an updated spec sheet?

There is no fixed timeline. Assistants that search the web can use new pages quickly; answers from training data change only when models are updated.

## Sources

- International Federation of Robotics (2026-09), [Executive Summary World Robotics 2026 Industrial Robots](https://ifr.org/img/worldrobotics/Executive_Summary_WR_2026_-_Industrial_Robots.pdf)
- International Federation of Robotics (2026-09-24), [US now second-largest robotics market, following China](https://ifr.org/downloads/press_docs/EN-2026-SEP-24-IFR_Press_Release_WR-USA.pdf)
- International Federation of Robotics (2026-09-30), [Global Sales of Professional Service Robots Surge 24%](https://ifr.org/ifr-press-releases/global-sales-of-professional-service-robots-surge-24-percent)
- Australian Manufacturing (2026-10), [Global professional service robot sales rise 24% in 2025: IFR](https://www.australianmanufacturing.com.au/global-professional-service-robot-sales-rise-24-in-2025-ifr/)
- Association for Advancing Automation (2026-08), [Robot Orders Increase in Q2 as Automation Demand Broadens Across Industries](https://www.automate.org/robotics/news/robot-orders-increase-in-q2-as-automation-demand-broadens-across-industries)
- The Robot Report (2026-07), [Teradyne Robotics revenue rises 33% year over year in Q2](https://www.therobotreport.com/teradyne-robotics-revenue-rises-33-year-over-year-in-q2/)
- PMMI (2026-08), [2026 Robotics in Packaging and Processing](https://www.pmmi.org/report/2026-robotics-in-packaging-and-processing)
- TREW Marketing and GlobalSpec (2026), [State of Marketing to Engineers research report](https://advertising.globalspec.com/state-of-marketing-to-engineers-research-report)
- Gartner (2026-05-20), [Gartner Survey Finds 69% of B2B Buyers Turn to Sales Reps to Validate AI-Generated Insights](https://www.gartner.com/en/newsroom/press-releases/2026-05-20-gartner-survey-finds-sixty-nine-percent-of-b-two-b-buyers-turn-to-sales-reps-to-validate-ai-generated-insights)
- OpenAI Help Center (2026), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Google (2025-03-05), [Expanding AI Overviews and introducing AI Mode](https://blog.google/products/search/ai-mode-search/)
- Underneath (2026), [The YouTube videos Google’s AI cites are small](https://underneath.agency/research/ai-overview-youtube-videos-study)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [How many “best of” lists cited by AI rank their own brand first?](https://underneath.agency/research/self-promoting-best-lists-study)

---

This is the Markdown twin of https://underneath.agency/resources/robotics-companies-customers-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How can sales software companies turn AI answers into pipeline?"
description: "By being named when sales leaders ask AI to shortlist tools, then answering the demo requests that follow fast, with proof on data, security and fit."
canonical: "https://underneath.agency/resources/sales-software-pipeline-from-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How can sales software companies turn AI answers into pipeline?

By getting named when revenue leaders ask an AI assistant which sales engagement, intelligence, prospecting or enablement tool to shortlist, then converting that intent into demos and pilots before a competitor does. Sales leaders are watching their own buyers move to AI research, and they now buy their tools the same way. The evidence that AI visibility moves pipeline is growing but mostly company-reported; nobody has yet measured its effect on closed revenue.

## The short version

1. Sales software accounts are large: [Gong](https://www.gong.io/press/gong-growth-accelerates-past-55-yoy-arr-tops-500m) reported more than $500 million in recurring revenue in May 2026 and [more than 5,000 customers](https://www.gong.io/press/gong-appoints-simon-frey-as-chief-customer-officer-as-global-customer-base-surpasses-5-000), which works out to roughly a hundred thousand dollars a customer on our arithmetic.
2. The category is growing and expanding inside accounts: [Clay](https://www.clay.com/blog/series-d) raised money at a $7.1 billion valuation after 4x revenue growth in 2025, and HubSpot said Sales Hub seat upgrades were up 71% year over year.
3. Demand that arrives is often wasted: when [Clay](https://www.clay.com/blog/claygent-experiment-speed-to-lead) submitted 6,346 demo and contact forms, only 32% got an email reply and 96% of companies never phoned.
4. Sales tools are moving into the assistant itself: [HubSpot](https://transcripts.platformaeronaut.com/transcripts/HUBS-2Q25-transcript) said over 20,000 customers had used its ChatGPT and Claude connectors, and [Outreach](https://www.outreach.io/company/newsroom/outreach-chatgpt-and-mcp-server-for-codex), [Apollo](https://www.apollo.io/magazine/apollo-is-now-available-in-chatgpt) and [ZoomInfo](https://www.zoominfo.com/solutions/zoominfo-in-chatgpt) all now offer apps inside ChatGPT.
5. Trust is the gate: in [Gong’s research](https://www.gong.io/press/unlocking-the-trust-barrier-new-gong-research) with more than 2,000 US and UK leaders, 58% of companies had stalled AI projects, with data and security concerns the top reason (34%).

## Who on the revenue team buys sales tools, and what is one account worth?

Revenue leaders sign, revenue operations evaluates, and a won account can be worth six figures a year.

Sales software covers several categories with different buyers inside the revenue team. Sales engagement platforms (Outreach, Salesloft) run sequences and cadences. Revenue and conversation intelligence (Gong, Clari) records calls and forecasts deals. Data and prospecting tools (ZoomInfo, Apollo, Clay) find and enrich contacts. Enablement platforms (Highspot, Seismic) manage content and coaching. Revenue operations usually runs the evaluation, frontline managers test it, and security and finance review it before the contract.

The customer values are high and growing. Gong said it added more $1 million-plus customers in its last two quarters than in the previous six combined. Clay, the prospecting data platform, has more than 17k customers. Expansion is built in: HubSpot told investors that Sales Hub seat upgrades were up 71% year over year. A sales software customer found through an AI answer is rarely a one-time sale; it is a seat count that can grow with the sales team.

The pain these buyers bring is specific. In [Highspot’s survey of 463 senior go-to-market leaders](https://www.demandgenreport.com/industry-news/news-brief/bridging-the-ai-performance-gap-insights-from-highspots-gtm-report/50338/), 98% said their strategy was in motion but only 10% said they were driving successful initiatives. Gong, a vendor with a stake in the claim, says sellers spend 77% of their time on non-selling activities. Buyers arrive looking for proof that a tool closes that gap.

## When do revenue leaders consult AI assistants while choosing sales tools?

At discovery and shortlisting, and sales leaders see the same shift in their own pipelines.

Sales leaders know AI-assisted buying from the other side of the table. In [Gartner’s survey of 646 B2B buyers](https://www.gartner.com/en/newsroom/press-releases/2026-03-09-gartner-sales-survey-finds-67-percent-of-b2b-buyers-prefer-a-rep-free-experience), 45% had used AI during a recent purchase and 67% preferred a rep-free experience. In a [second Gartner survey of 645 buyers](https://www.gartner.com/en/newsroom/press-releases/2026-05-20-gartner-survey-finds-sixty-nine-percent-of-b-two-b-buyers-turn-to-sales-reps-to-validate-ai-generated-insights), presented at its CSO and Sales Leader Conference, 69% preferred to validate AI-generated insights with sales reps. Buyers were split on which source is more likely to mislead them: about half said AI, and 49% said a sales rep. These are the conditions sales leaders are buying tools for, and the way they shop themselves.

For software purchases specifically, AI is now part of how the shortlist forms. [G2’s 2026 survey of more than 1,000 software buyers](https://sell.g2.com/2026-buyer-behavior-report) found that buyers who sourced recommendations from AI chatbots bought from their initial shortlist in at least three of their last five purchases 80% of the time, against 65% for buyers who did not. Review sites (38%) and AI chatbots (37%) were the top sources shaping the shortlist. G2 runs a review platform, so it has a stake in that conclusion. Marketing leaders show the same habit, covered in [how marketing software companies win buyers](https://underneath.agency/resources/marketing-software-ai-search-growth).

Google is a second AI surface with its own behavior. In [our AI Mode and AI Overviews study](https://underneath.agency/research/ai-mode-vs-ai-overviews-study), on 46 B2B software searches where both cited sources, the mean overlap between the pages they cited was 0.144 on a 0-to-1 scale. In practice, being cited in one of Google’s AI answers says little about the other.

## Which questions do sales leaders ask AI about sales software?

Fit, data quality, integration, proof and price, framed around their own team and motion.

We made up the prompts below to show how a CRO or RevOps lead might phrase a question; they are not recorded queries:

- Category fit: “Best sales engagement platform for a mid-market team selling into healthcare on Salesforce.”
- Data quality: “Which B2B contact data provider has the most accurate direct dials in Europe?”
- Alternatives: “Alternatives to Gong for conversation intelligence that cost less for a small team.”
- Consolidation: “Can one tool replace our sequencing, dialer and call recording?”
- Security and compliance: “Which call recording tools keep data in the EU and do not train on customer calls?”
- Proof: “What results have companies reported after rolling out an AI sales agent?”

Each question maps to a decision. Category and alternatives questions settle which sales tools make the first cut. Data, integration and security questions decide who survives revenue operations and the security review. Proof questions decide whether a pilot gets funded. Our [article on comparison pages](https://underneath.agency/resources/do-comparison-pages-help-b2b-ai-citations) looks at how side-by-side content for sales tools gets used at that point.

## What path runs from an AI answer to a sales software demo and pilot?

The answer names you, the buyer checks reviews, requests a demo, then runs a pilot.

**Shortlist.** If the assistant leaves you off a short list, you rarely get added later. G2’s 80% figure above is the clearest measure of how sticky an AI-shaped shortlist is.

**Demo request.** This is where many vendors lose the opportunity they earned. Clay’s experiment used agents to fill in 6,346 demo and contact forms across 100 countries. Only 2,013 forms (32%) got an email reply, about 7% got a reply from a salesperson, and average speed to lead was 15 to 24 hours. Clay notes that a fictional buyer profile and spam filters may have cut response rates. For a sales software company, whose buyers judge it partly on how it sells, a slow reply to an AI-referred demo request is especially costly.

**Pilot and contract.** Enterprise sales tools are usually proven in a pilot with one team before a wider rollout. We infer that the AI answer’s main effect here is on who gets invited to pilot, not on the pilot result, which depends on data quality, adoption and integration.

**What it costs to be missing.** The loss rarely shows up in analytics. A revenue leader who asked an assistant for three conversation intelligence tools and never saw your name does not appear in your funnel at all. We look at that blind spot in [what lost clicks to AI answers mean for pipeline](https://underneath.agency/resources/ai-answers-pipeline-revenue).

## Why does it matter that sales tools now run inside ChatGPT?

Because the assistant is becoming the place sellers work, not only where buyers research.

HubSpot said it was the first CRM to launch connectors with both ChatGPT and Claude, and that over 20,000 customers had used them to access insights across 23 million CRM records. Outreach launched an app that is available in the ChatGPT app directory, describing it as the first revenue orchestration platform available natively in OpenAI’s products. Apollo says its ChatGPT app lets sellers search prospects, enrich contacts and add them to sequences without leaving the conversation, and ZoomInfo offers its data inside ChatGPT as well.

These facts are documented by the vendors. What follows is our inference: when sellers already work inside an assistant, the vendors present there gain a second kind of visibility, as a tool the assistant can use as well as one it can recommend. A sales software company with no presence in the assistant ecosystems its buyers use risks being invisible in both roles. Outreach’s own numbers suggest the pace of change: it reported 480% year-over-year growth in AI recurring revenue in one quarter.

## Why do assistants put some sales tools on the shortlist and not others?

The platforms say little; research and buyer behavior point to reviews, consistent facts and visible data-security proof.

**What Google states.** According to [Google](https://developers.google.com/search/docs/appearance/ai-features), a single question to AI Overviews or AI Mode may set off a “query fan-out,” with multiple related searches across subtopics and data sources, and a page needs nothing beyond normal search eligibility to appear. A question about prospecting data can therefore pull in review pages, comparison articles and vendor documentation at once. See [the hidden searches AI assistants run](https://underneath.agency/research/ai-hidden-searches-study).

**Observed in studies.** B2B software answers are more stable than most. Of the eight industries in [our four-assistant study](https://underneath.agency/research/ai-assistants-brand-agreement-study), B2B software showed the most agreement between assistants (0.543); [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study) found its recommended brands held steadiest when a question was repeated (0.708). We infer that a sales software brand that becomes established in its category’s answers holds that position better than brands in local or retail categories do, and that a newcomer has to work harder to break in. Our guide to [why AI keeps naming the same CRMs](https://underneath.agency/resources/crm-software-ai-search) shows how challengers get in.

**Trust factors specific to sales software.** These tools record customer calls, store contact data and increasingly act on their own. Gong’s research found that one in four sales calls in its data referenced security, and that 46% of planned AI investments had been paused because of trust concerns. A reasonable expectation is that clear, public documentation of data handling, model training policies, certifications and CRM integrations helps both the human buyer and the assistant describing you. Verifiable customer results matter too, because sales leaders will ask for them. Support software vendors face the same demand for proof, as [our guide to customer support software](https://underneath.agency/resources/customer-support-software-ai-search) shows.

## What does GEO mean in practice for a sales tech vendor?

It improves what assistants can find and verify about your sales tool, without any promised shortlist spot.

1. **Category clarity.** State plainly which category you are in (engagement, intelligence, data, enablement or a mix), which motion you serve and which CRMs you integrate with, the same way on your site, review profiles, partner marketplaces and documentation.
2. **Reviews that name the use case.** Keep recent reviews flowing on the platforms revenue teams check, mentioning team size, motion and results.
3. **Independent proof.** Earn coverage in the newsletters, podcasts, communities and analyst-style comparisons sales leaders read. To choose which of those outlets to pursue first, read our piece on [which pages to target](https://underneath.agency/resources/best-of-lists-ai-recommendations).
4. **Public trust documentation.** Publish security, privacy, data-source and AI-training pages in plain language, ungated, so assistants and security reviewers can find the same answers.
5. **A presence where sellers work.** Where it fits your product, consider the assistant app directories and connectors your buyers use, and keep those listings as accurate as your website.
6. **Fast follow-up.** Treat AI-referred demo requests as your hottest leads and measure speed to lead, because the visibility is wasted if the reply comes a day later.
7. **Tracking question by question.** Follow category, alternatives and comparison questions for your sales tool across ChatGPT, Gemini, Perplexity, Copilot and Google, then match each one against demo requests and self-reported attribution. See [what to measure for AI visibility](https://underneath.agency/resources/what-to-measure-ai-visibility).

Sales tools follow many of the patterns in [how B2B SaaS companies earn revenue from AI search](https://underneath.agency/resources/b2b-saas-revenue-from-ai-search).

## What can’t we yet say about AI answers and sales software pipeline?

Nobody has shown how much closed sales software pipeline AI visibility causes, rather than merely accompanies.

Most figures here come from vendors describing their own growth or from surveys by companies that sell to software buyers. The Gartner surveys describe B2B buyers in general, not sales software buyers alone. Clay’s experiment used a single made-up buyer and may understate response rates. No published study follows sales leaders from an AI answer to a pilot and a signed contract, and the effect of being available inside ChatGPT on new customer acquisition has not been measured. For how a sales tech vendor could test that itself, read [how to prove GEO caused a change in sales](https://underneath.agency/resources/prove-geo-caused-sales).

## How should a sales tech vendor check whether AI answers feed its demo calendar?

Find out which AI shortlists name you for the questions revenue leaders ask before booking demos.

Map the category, alternatives, data, security and proof questions your buyers ask, check them across the main assistants, and compare the answers with your demo requests and pipeline by segment. Then time how fast your team responds to the demo requests you already get. We can [audit your AI shortlist presence alongside your demo data](https://underneath.agency/contact), so any plan to close the gaps is judged by demo volume and pipeline rather than by visibility alone. The [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page lays out what that plan covers for a sales tool, including review profiles, public trust documentation and question-by-question tracking.

## Frequently asked questions

### Do sales leaders use ChatGPT to choose sales software?

Many do some of their research there. No published survey isolates sales leaders, but surveys of software buyers find AI chatbots now rival review sites in shaping shortlists, and sales leaders see the same behavior in their own buyers.

### Does being available as an app in ChatGPT help a sales tool get recommended?

There is no published evidence either way. Vendors document the apps; whether an assistant favors tools it can connect to is not documented by the platforms, so treat any claim that it does as unproven.

### Which matters more for AI visibility: review sites or our own website?

Both play a role. Buyers and studies point to independent reviews and coverage for discovery, while your own pages need to state facts clearly, including pricing, integrations and security, so answers describe you correctly.

### How quickly should we respond to demo requests that come from AI search?

As fast as you can. Clay’s experiment found most companies took most of a day or never replied, so a same-day response, ideally within minutes, sets you apart.

## Sources

- Gong (2026-05-12), [Gong Growth Accelerates Past 55% YoY as Enterprises Adopt Revenue AI; ARR Tops $500M](https://www.gong.io/press/gong-growth-accelerates-past-55-yoy-arr-tops-500m)
- Gong (2026-03-04), [Gong Appoints Simon Frey as Chief Customer Officer as Global Customer Base Surpasses 5,000](https://www.gong.io/press/gong-appoints-simon-frey-as-chief-customer-officer-as-global-customer-base-surpasses-5-000)
- Gong (2026-04-15), [Unlocking the trust barrier: new Gong research](https://www.gong.io/press/unlocking-the-trust-barrier-new-gong-research)
- Gartner (2026-03-09), [Gartner Sales Survey Finds 67% of B2B Buyers Prefer a Rep-Free Experience](https://www.gartner.com/en/newsroom/press-releases/2026-03-09-gartner-sales-survey-finds-67-percent-of-b2b-buyers-prefer-a-rep-free-experience)
- Gartner (2026-05-20), [Gartner Survey Finds 69% of B2B Buyers Turn to Sales Reps to Validate AI-Generated Insights](https://www.gartner.com/en/newsroom/press-releases/2026-05-20-gartner-survey-finds-sixty-nine-percent-of-b-two-b-buyers-turn-to-sales-reps-to-validate-ai-generated-insights)
- G2 (2026-07), [2026 Buyer Behavior Report: The Evaluation Maze](https://sell.g2.com/2026-buyer-behavior-report)
- Clay (2026), [We asked 6,346 companies for a demo. Most never wrote back.](https://www.clay.com/blog/claygent-experiment-speed-to-lead)
- Clay (2026), [Clay raises $115M to be the AI growth engine for every company](https://www.clay.com/blog/series-d)
- Demand Gen Report (2025-09-15), [Bridging the AI Performance Gap: Insights From Highspot’s GTM Report](https://www.demandgenreport.com/industry-news/news-brief/bridging-the-ai-performance-gap-insights-from-highspots-gtm-report/50338/)
- HubSpot (2025-08), [Second quarter 2025 earnings call transcript](https://transcripts.platformaeronaut.com/transcripts/HUBS-2Q25-transcript)
- Outreach (2026-06-03), [Outreach Launches Revenue Orchestration Platform App in ChatGPT and MCP Server for Codex](https://www.outreach.io/company/newsroom/outreach-chatgpt-and-mcp-server-for-codex)
- Outreach (2026-09-09), [Outreach Reports 12x AI Usage Growth](https://www.outreach.io/company/newsroom/outreach-reports-12x-ai-usage-growth-as-revenue-teams-move-beyond-insight-to-ai-driven-execution)
- Apollo.io (2026), [Apollo is now available in ChatGPT](https://www.apollo.io/magazine/apollo-is-now-available-in-chatgpt)
- ZoomInfo (2026), [ZoomInfo for ChatGPT](https://www.zoominfo.com/solutions/zoominfo-in-chatgpt)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- Underneath (2026), [AI Mode vs AI Overviews](https://underneath.agency/research/ai-mode-vs-ai-overviews-study), [four-assistant agreement](https://underneath.agency/research/ai-assistants-brand-agreement-study), [recommendation consistency](https://underneath.agency/research/ai-recommendation-consistency-study) and [hidden searches](https://underneath.agency/research/ai-hidden-searches-study)

---

This is the Markdown twin of https://underneath.agency/resources/sales-software-pipeline-from-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How security awareness vendors win customers from AI search"
description: "By being named, with checkable proof, when IT leads, MSPs and CISOs ask AI which training meets their insurer, auditor and budget."
canonical: "https://underneath.agency/resources/security-awareness-training-customers-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How do security awareness training vendors win customers through AI search?

By being the vendor an AI assistant names, with proof the buyer can check, when an IT lead, an MSP or a CISO asks which training will satisfy their insurer, their auditor and their budget. Security awareness training is bought in high volume at modest contract values, often under a compliance deadline, so a place in the answer matters more than a click. The evidence that AI search shapes this specific category is still indirect, and we say where.

## The short version

1. Many purchases start with an insurer or an auditor. [Marsh](https://www.marsh.com/en-gb/about/media/incident-response-planning-emerges-key-cybersecurity-control-reducing-cyber-risk.html), studying the 12 controls cyber insurers track, ranked awareness training and phishing testing among the four most linked to fewer claims, and [Coalition](https://www.coalitioninc.com/announcements/2025-cyber-claims-report) traced 60% of its 2024 cyber insurance claims to email fraud.
2. Contracts are modest, so volume matters: the median buyer of KnowBe4 pays $7,763 a year, according to [Vendr’s data](https://www.vendr.com/marketplace/knowbe4) from 682 purchases.
3. The market is concentrated. KnowBe4 was bought for [$4.6 billion](https://www.vistaequitypartners.com/news/knowbe4-to-be-acquired-by-vista-equity-partners-for-4-6-billion/) in 2022, and a challenger’s CEO claims it holds over 80% of the market. For challengers, the shortlist is the battle.
4. The category is being renamed: Gartner predicts 80% of enterprises will run a staffed human risk management program by 2030, up from 20% in 2022. New names bring new buyer questions.
5. In our studies, 24.2% of the “best X” lists AI engines cited ranked their own publisher first, and adding “on a tight budget” to a question kept the original first brand only 15.3% of the time.

## Who signs up for security awareness training, and what does one account pay?

Mostly IT and security teams buying per-seat subscriptions, from small offices to global enterprises, often through managed service providers.

Who holds the budget depends on headcount. In a small firm it is the IT manager or an outside managed service provider (MSP). In a large one it is a security awareness lead inside the CISO’s team, working with HR, which owns training records, and internal communications. That lead is usually stretched. The [SANS Institute’s 2025 Security Awareness Report](https://www.sans.org/press/announcements/security-awareness-report-2025), based on more than 2,700 practitioners, found lack of time and staffing are the two biggest challenges. SANS says it takes at least 2.8 dedicated full-time staff to meaningfully influence behavior.

Contract values are modest by enterprise software standards. Vendr’s transaction data put the median KnowBe4 buyer at $7,763 a year, in a range from $1,929 to $19,419. Pricing is per user, per year, with volume discounts at seat thresholds and multi-year terms. A single customer is rarely a large deal. The business is built on many of them renewing.

The leader’s numbers show the scale of that model. When [Vista Equity Partners agreed to buy KnowBe4](https://www.helpnetsecurity.com/2022/10/12/vista-equity-partners-knowbe4/) for $4.6 billion in 2022, it served more than 52,000 organizations. Its [press page](https://www.knowbe4.com/press/knowbe4-report-reveals-security-training-reduces-global-phishing-click-rates-by-86) now says it is trusted by more than 70,000.

Small and large buyers have different problems, which matters for what they ask. In KnowBe4’s 2025 benchmark, organizations with more than 10,000 employees started with 40.5% of staff likely to fall for a simulated phish, against 24.6% for organizations with up to 250 employees.

## Why do companies buy it: compliance, insurance or real risk?

All three, and compliance or insurance deadlines often decide when the purchase happens.

**The risk is real.** The [2025 Verizon Data Breach Investigations Report](https://news.clearancejobs.com/2025/04/30/60-of-breaches-still-tied-to-human-mistakes-what-the-2025-dbir-means-for-fsos-and-leaders/) analyzed 22,052 security incidents and found the human element involved in 60% of them, as summarized by ClearanceJobs. In the SANS survey, 80% of organizations ranked social engineering as their number one human risk.

**Insurers ask about it.** Coalition found that 60% of its 2024 claims came from business email compromise and funds transfer fraud, both of which usually start with a deceived employee. [Marsh](https://www.marsh.com/na/services/cyber-risk/insights/cyber-resilience-twelve-key-controls-to-strengthen-your-security.html) says insurers now require specific cybersecurity controls, “placing insurability at stake.” Its 2025 analysis of the 12 controls the insurance industry tracks ranked awareness training and phishing testing among the four most linked to fewer breach-based claims.

**Rules require it.** Payment card rules, as [SecurityMetrics explains](https://maintenance.securitymetrics.com/blog/security-awareness-training), now require training to cover phishing and social engineering threats (PCI DSS requirement 12.6.3.1), mandatory since March 31, 2025. In Europe, [Article 20 of the NIS2 Directive](https://www.springlex.eu/en/packages/nis2/nis2-directive/article-20/) requires board members of covered companies to follow cybersecurity training and asks countries to encourage regular training for employees.

The practical effect, we infer, is that many buyers arrive with a deadline and a checklist: a renewal questionnaire, an audit date, a board requirement. They want a vendor that clearly meets it, quickly.

## At what point do awareness training buyers turn to AI assistants?

At the research step, where buyers turn a compliance requirement into a shortlist, though no study measures this category alone.

Across software as a whole, the shift is clear. Of 1,076 software buyers polled by the review platform [G2](https://company.g2.com/news/g2-research-the-answer-economy) in March 2026, 51% said that when research starts, they now reach for an AI chatbot more often than for Google. G2 profits from selling visibility to vendors, and its sample spans all software rather than phishing simulation and training tools.

Google itself nearly always adds an AI layer to software searches. Of the 1,248 US searches in [our AI Overview study](https://underneath.agency/research/ai-overviews-frequency-study), the B2B software and technology keywords carried an AI Overview, Google’s summary above the links, most often: 96.0% of them, the top rate among eight industries.

Two features of this market make AI research likely, we infer. Buyers are short of time, as SANS found. And many purchases are small and fast: in [Capterra’s 2025 survey](https://capterra.com/resources/tech-trends-successful-buyer-purchase-journey/) of 3,500 software buyers, most successful buyers (57%) took 3 months or less to evaluate their options. An IT manager with an insurance form due Friday is the kind of buyer who asks an assistant for a short list and acts on it.

## Which questions do security awareness buyers ask AI assistants?

Compliance, insurance, alternatives, comparisons, new threats and price. The sample prompts in this table are our own, written for an IT manager, CISO or MSP; none was taken from an actual buyer.

| Buyer and stage | Illustrative prompt |
|---|---|
| Small business, compliance | “What security awareness training meets PCI DSS 4.0 phishing requirements for a 60-person retailer?” |
| Insurance renewal | “What phishing training does my cyber insurer expect, and which tools produce the reports they ask for?” |
| Alternatives | “What are cheaper alternatives to KnowBe4 for a 200-person company?” |
| Enterprise comparison | “Hoxhunt or Proofpoint for 15,000 employees in 12 languages?” |
| New threats | “Which platforms train staff to spot deepfake voice calls from executives?” |
| MSP | “Best security awareness platform for an MSP managing 80 small clients?” |
| Board and HR | “What does NIS2 require for board cybersecurity training?” |

Each prompt carries a constraint: a rule, a size, a budget, a language, a channel. Constraints change the answer. In [our rewording study](https://underneath.agency/research/ai-prompt-phrasing-study), adding “on a tight budget” kept the original first brand only 15.3% of the time, and adding “I run a small business with about 10 employees” 40.6%. A vendor named for the enterprise version of a question may be missing from the small business version.

The emerging-threat questions are a real opening. [Proofpoint’s 2024 State of the Phish](https://www.proofpoint.com/us/newsroom/press-releases/proofpoints-2024-state-phish-report-68-employees-willingly-gamble) found only 23% of organizations educate users on generative AI safety. Buyers looking to close that gap are asking questions where no incumbent has yet earned the answer.

## How does an AI answer turn into a security awareness customer?

Through one of three short paths: a self-serve trial, an enterprise evaluation or an MSP partnership.

1. **Small business, direct.** An assistant names three or four platforms. The buyer runs a free phishing test or a trial with one or two, then signs an annual per-seat contract that renews if the reports satisfy the insurer and auditor.
2. **Enterprise.** The awareness lead uses AI answers to build a long list, then runs a formal evaluation with security, HR and procurement. The win is a multi-year contract, often expanded later to email security or other modules.
3. **MSP.** An MSP asks which platform suits many small clients. One decision can bring a vendor dozens of client accounts, we infer, which makes MSP-phrased questions unusually valuable.

The amounts per customer are modest, so the channel must work at volume and at low cost. AI answers may suit that: the assistant does the first round of qualification before the buyer arrives, we expect. Whether it does is not yet measured, and the click rarely shows in analytics; see [what lost clicks mean for pipeline](https://underneath.agency/resources/ai-answers-pipeline-revenue).

## Why might an assistant recommend one phishing training platform over another?

No platform documents how it picks vendors; studies point to reviews and published rankings, and buyers want measurable outcomes.

**What Google discloses.** A question about awareness training can become several searches: Google says AI Overviews and AI Mode [may use a “query fan-out” technique](https://developers.google.com/search/docs/appearance/ai-features) that issues multiple related searches. No platform, Google included, publishes how training vendors are chosen.

**Observed in our studies.** In [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study), ChatGPT searched for reviews in 46.2% of its answers and for a named publication, ranking or award in 43.8%. And when [our reputation study](https://underneath.agency/research/is-it-legit-ai-reputation-study) asked “Is this brand legit?”, a review or complaint platform appeared among the sources in 88.0% of answers.

“Best security awareness training” lists deserve a warning. Many are published by vendors in the category. In [our study of self-ranking lists](https://underneath.agency/research/self-promoting-best-lists-study), 24.2% of the numbered “best X” lists AI engines cited put their own publisher first. Such lists made up only 1.1% of all citations, so they are not a shortcut. We cover the wider pattern in [why “best of” lists shape AI recommendations](https://underneath.agency/resources/best-of-lists-ai-recommendations).

**What this market rewards.** Buyers and analysts are moving from completion rates to measured behavior. Gartner, as quoted by [Hoxhunt](https://hoxhunt.com/blog/gartner-names-hoxhunt-a-security-behavior-and-culture-change-program-representative-provider), predicts that by 2030 control frameworks will measure behavior change rather than compliance-based training. Benchmarks already circulate: KnowBe4 reports a baseline of 33.1% of employees falling for simulations, falling to 4.1% after 12 months of training. That is vendor-reported, from 67.7 million simulations.

**Our inference.** A reasonable expectation is that vendors with public, specific proof are easier for assistants to name and for buyers to trust. That proof includes compliance mapping, outcome data with a method, independent reviews and analyst mentions. Nobody has yet tested that for awareness training vendors.

## What does an awareness training vendor lose when assistants leave it out?

Mostly shortlist places in a concentrated market, though no one has put a dollar figure on it.

- **Leaders get named by default.** In a [BankInfoSecurity interview](https://www.bankinfosecurity.com/adaptive-security-gets-81m-series-b-for-ai-deepfake-defense-a-30332), Adaptive Security’s CEO said KnowBe4 holds over 80% of the market. That is a competitor’s claim, not audited data. Still, assistants often default to market leaders, as we explain in [do AI assistants favor big brands](https://underneath.agency/resources/do-ai-assistants-favor-big-brands). A challenger that is not named loses before it can compete on price or design.
- **Being named is not a fixed state.** Ask ChatGPT the same question five times, as [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study) did, and only 25.2% of the brands it names turn up in every run.
- **Missed moments recur yearly.** Insurance renewals, audits and annual training cycles come round every 12 months. A vendor missing from the answer at renewal waits a year, we infer.
- **New money is going to new threats.** [SecurityWeek](https://www.securityweek.com/adaptive-security-raises-81-million-in-series-b-funding/) reports Adaptive raised $81 million in a Series B to train staff against deepfakes and AI-driven scams. Funded entrants are competing for the same new questions.

## What does GEO look like for a phishing and awareness training vendor?

It puts your compliance mapping and outcome data where assistants and buyers can check them, with no guaranteed placement.

1. **Name the category consistently.** Decide how you describe yourself: security awareness training, human risk management or both. Use the same words on your site, review profiles, partner directories and analyst briefings. Mixed labels confuse an IT manager and an assistant alike, and our guide to [fixing wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) shows how to clean them up.
2. **Publish compliance mapping in plain pages.** One page each for PCI DSS, HIPAA, NIS2, insurance questionnaires and common frameworks, saying exactly which reports and records you produce.
3. **Publish outcome data with its method.** Benchmark reports are cited widely. Say how the numbers were measured and what they cannot show.
4. **Earn independent reviews and coverage.** Ask customers for detailed reviews on platforms buyers use, by company size and use case. Pursue analyst recognition and security press, not self-ranked lists. Our piece on [building authority for AI search](https://underneath.agency/resources/how-brands-build-authority-for-ai-search) covers those routes in more detail.
5. **Serve each buyer separately.** Write honest pages for small businesses, enterprises and MSPs, plus alternatives and comparison pages; see [whether comparison pages help](https://underneath.agency/resources/do-comparison-pages-help-b2b-ai-citations).
6. **Measure with real constraints.** Track several assistants with budget, size, MSP and regulation wording, asking each question more than once.

## What don’t we know yet about AI search in security awareness buying?

Two things: how many training buyers ask AI assistants, and whether being named sells more seats.

- **No category survey.** The AI-use figures here cover software buyers in general.
- **Vendor data dominates.** The phishing benchmarks come from vendors, the 80% share claim from a competitor, and the pricing from a procurement platform’s sample.
- **Seat sales are unproven.** Among the 45 studies of AI search optimization reviewed by [Martinez](https://arxiv.org/abs/2607.14035), the evidence tying it to traffic and conversions was the thinnest.
- **Answers change.** Results vary between runs, wordings and assistants, so one test proves little.

## How can an awareness training vendor check its standing before the next renewal cycle?

Test the questions IT managers, CISOs and MSPs ask, with their budget and compliance limits, and note who gets named.

Write 20 to 30 prompts across the three paths: small business, enterprise and MSP, each with compliance, insurance, budget and new-threat wording. Run each several times across ChatGPT, Google AI Overviews and AI Mode, Gemini, Perplexity, Copilot and Claude. Note who is named, which sources are cited and what is said about your compliance coverage and results. The broader security software picture is in [how cybersecurity software firms earn revenue from AI search](https://underneath.agency/resources/cybersecurity-software-revenue-from-ai-search).

To go through the results with people who run these checks every week, [ask us for a review of your awareness training visibility](https://underneath.agency/contact). We will show which buyer questions name you, which name competitors instead, and which gaps in your public proof are most likely costing you trials, MSP partners and seat renewals. Our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page describes the follow-on work, such as compliance mapping pages, outcome data with its method and separate pages for each buyer.

## Frequently asked questions

### Do small businesses use AI to choose security awareness training?

No study isolates them. Across software, 51% of buyers in G2’s survey start research with AI chatbots more often than Google, and small buyers decide fast.

### Should we publish our phishing click-rate results?

Yes, with the method. Buyers and insurers ask for measured results, and KnowBe4’s widely cited 33.1% baseline shows how often benchmark data gets repeated.

### Do “best security awareness training” lists on our own blog help?

Probably little. Self-ranking lists were 1.1% of AI citations in our study; independent reviews and rankings are the stronger evidence.

### Should we position as human risk management instead?

Use both terms if both fit. Gartner expects 80% of enterprises to run human risk management programs by 2030, but many small buyers still search for awareness training.

### How quickly can AI visibility turn into customers here?

Faster than in most security categories, we expect, because most successful software buyers decide within 3 months. It is not yet measured.

## Sources

- Marsh (n.d.), [Cyber resilience: twelve key controls to strengthen your security](https://www.marsh.com/na/services/cyber-risk/insights/cyber-resilience-twelve-key-controls-to-strengthen-your-security.html)
- Marsh (2025-08-27), [Incident response planning emerges as key cybersecurity control in reducing cyber risk](https://www.marsh.com/en-gb/about/media/incident-response-planning-emerges-key-cybersecurity-control-reducing-cyber-risk.html)
- Coalition (2025-05-07), [Coalition 2025 Cyber Claims Report](https://www.coalitioninc.com/announcements/2025-cyber-claims-report)
- Vendr (2026), [KnowBe4 software pricing and plans](https://www.vendr.com/marketplace/knowbe4)
- Vista Equity Partners (2022-10-12), [KnowBe4 to be acquired by Vista Equity Partners for $4.6 billion](https://www.vistaequitypartners.com/news/knowbe4-to-be-acquired-by-vista-equity-partners-for-4-6-billion/)
- Help Net Security (2022-10-12), [Vista Equity Partners acquires KnowBe4 for $4.6 billion in cash](https://www.helpnetsecurity.com/2022/10/12/vista-equity-partners-knowbe4/)
- KnowBe4 (2025-05-13), [KnowBe4 report reveals security training reduces global phishing click rates by 86%](https://www.knowbe4.com/press/knowbe4-report-reveals-security-training-reduces-global-phishing-click-rates-by-86)
- SANS Institute (2025-08-13), [Security Awareness Report 2025](https://www.sans.org/press/announcements/security-awareness-report-2025)
- ClearanceJobs (2025-04-30), [60% of breaches still tied to human mistakes: what the 2025 DBIR means](https://news.clearancejobs.com/2025/04/30/60-of-breaches-still-tied-to-human-mistakes-what-the-2025-dbir-means-for-fsos-and-leaders/)
- SecurityMetrics (n.d.), [New PCI requirements: security awareness training](https://maintenance.securitymetrics.com/blog/security-awareness-training)
- European Union, via Springlex (2022-12-27), [NIS2 Directive, Article 20](https://www.springlex.eu/en/packages/nis2/nis2-directive/article-20/)
- Hoxhunt, quoting Gartner (2022), [Gartner names Hoxhunt a Security Behavior and Culture Program Representative Provider](https://hoxhunt.com/blog/gartner-names-hoxhunt-a-security-behavior-and-culture-change-program-representative-provider)
- Proofpoint (2024-02-27), [2024 State of the Phish report](https://www.proofpoint.com/us/newsroom/press-releases/proofpoints-2024-state-phish-report-68-employees-willingly-gamble)
- BankInfoSecurity (2025), [Adaptive Security gets $81M Series B for AI deepfake defense](https://www.bankinfosecurity.com/adaptive-security-gets-81m-series-b-for-ai-deepfake-defense-a-30332)
- SecurityWeek (2025), [Adaptive Security raises $81 million in Series B funding](https://www.securityweek.com/adaptive-security-raises-81-million-in-series-b-funding/)
- G2 (2026-04-15), [In the Answer Economy, don’t win the click, win the answer](https://company.g2.com/news/g2-research-the-answer-economy)
- Capterra (2025), [Capterra’s 2025 Tech Trends Report](https://capterra.com/resources/tech-trends-successful-buyer-purchase-journey/)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- Martinez (2026), [Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026)](https://arxiv.org/abs/2607.14035), arXiv:2607.14035.
- Underneath (2026), [When does Google show an AI Overview? 1,248 US searches](https://underneath.agency/research/ai-overviews-frequency-study)
- Underneath (2026), [Does rewording a question change AI brand recommendations?](https://underneath.agency/research/ai-prompt-phrasing-study)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [“Is this brand legit?” How AI assistants build a reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study)
- Underneath (2026), [How many “best of” lists cited by AI rank their own brand first?](https://underneath.agency/research/self-promoting-best-lists-study)
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)

---

This is the Markdown twin of https://underneath.agency/resources/security-awareness-training-customers-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How can parcel companies win small shippers who ask AI?"
description: "Small sellers now ask AI how to ship cheaply, so parcel carriers and shipping platforms need current, citable rates, coverage and reviews to be named."
canonical: "https://underneath.agency/resources/shipping-companies-customers-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How do parcel carriers and shipping platforms win small businesses that ask AI how to ship?

By publishing the facts an assistant needs to compare them, and by being talked about in places it trusts. Small sellers are squeezed by shipping costs and most now use AI tools regularly, so the question “what is the cheapest way to ship this?” increasingly goes to an assistant. Carriers and platforms that state current rates, coverage and transit times clearly, and that earn reviews and independent coverage, give it something to cite.

## The short version

1. The market keeps growing and fragmenting: US parcel volume reached 23.1 billion shipments in 2025, up 3.3%, and “Other” carriers more than doubled their revenue share from 3.4% to 7.2%, according to the [Pitney Bowes Parcel Shipping Index](https://www.pitneybowes.com/us/shipping-index.html).
2. Shipping cost is the small seller’s biggest problem: in [ShipStation’s survey](https://www.shipstation.com/guides/shipstation-2026-merchant-insights-report/) of 868 US retailers with fewer than 500 employees, 62% said high shipping and freight rates limit their growth, and 67% spend more than 10% of revenue on shipping and fulfillment.
3. Small business owners use AI daily: in a [TechNet poll](https://www.technet.org/wp-content/uploads/2026/04/Small-Business-E-Commerce-and-AI-Survey-Poll-Toplines.pdf) of 1,108 small business owners and executives, 71% use AI tools regularly and 52% daily, and 79% say AI has helped expand the tools and service providers they use.
4. International rules just changed: before the US ended the de minimis exemption in August 2025, 1.36 billion packages worth $64.6 billion entered under it in a year, [the Associated Press reported](https://abc7news.com/post/us-tariff-exemption-small-orders-ends-friday-big-deal-shoppers-businesses/17653696/), and postal services in more than a dozen countries paused some US-bound packages.
5. Amazon Logistics became the largest US parcel carrier by volume in 2025, with 6.9 billion parcels, so the incumbents now compete with a new leader and a growing field of regional and last-mile carriers.

## Who buys parcel shipping, and what is a small shipper worth?

Small online sellers buy it, often through a platform. Shipping takes a large slice of their revenue.

The customer base is broad. The SBA Office of Advocacy counts [36.2 million small businesses](https://advocacy.sba.gov/?p=30239) in the US. Many sell online: in the TechNet poll, 46% sold on Amazon, 42% on Facebook Marketplace, 33% on eBay and 49% on their own websites.

Shipping spend is large relative to their size. ShipStation found that 35% of merchants spend more than 15% of revenue on shipping and fulfillment. By our own arithmetic, a seller with $2 million in annual revenue that spends 10% to 15% on shipping and fulfillment is putting $200,000 to $300,000 a year into labels, packaging and labor, and much of that goes to carriers and platforms.

Each kind of shipping company earns from that spend in a different way:

| Shipping company | How it earns from a small shipper | What the shipper is choosing |
|---|---|---|
| National carrier (USPS, UPS, FedEx) | Revenue per package, account volume | Price by weight and zone, speed, pickup, reliability |
| Amazon Logistics and marketplace programs | Fees bundled with selling on the marketplace | Whether to use marketplace shipping |
| Regional or last-mile carrier | Per-package revenue in its service area | Lower rates on some lanes, coverage limits |
| Multi-carrier shipping platform | Subscription, label and carrier fees | Discounted rates, integrations, automation |
| International parcel service | Per-package revenue plus duties handling | Customs, duties, transit time |

The incumbents still dominate revenue. UPS raised its full-year 2026 revenue outlook to about [$91.2 billion](https://investors.ups.com/news-events/press-releases/detail/2164/ups-releases-2q-2026-earnings). But the Pitney Bowes index shows volume moving toward Amazon and toward “Others,” whose volume rose 127%, led by Shein and Temu traffic and by regional carriers such as OnTrac, GLS, LSO and Spee-Dee and last-mile startups such as Veho and UniUni.

## How do small businesses choose a carrier or shipping platform today?

Mostly on price, with reliability and convenience close behind, and often inside a tool that compares carriers for them.

Rate pressure is constant. ShipStation’s report says UPS and FedEx both set 2026 general rate increases near 6%, with surcharges pushing the effective increase higher for most shippers. Peak season concentrates the risk: over 70% of the merchants surveyed earn more than a quarter of their annual revenue in a single quarter, and their top concern going in is shipping capacity.

That pressure pushes sellers toward comparison. Multi-carrier platforms exist to rate-shop across carriers, and marketplaces bundle shipping into their own programs. A carrier that is not in the platform a seller uses, or that is not mentioned when a seller asks for cheaper options, is not in the comparison at all.

International shipping adds a fresh layer of questions. When the US ended the de minimis exemption, which had let packages worth $800 or less enter duty-free, carriers had to collect and remit duties, and small sellers on both sides of the border had to relearn the rules. [CBC News](https://www.cbc.ca/lite/story/1.7618234) reported Canadian small businesses worried they could not survive the change. Questions like these are exactly the ones people now put to AI assistants.

## Where does AI search enter a small shipper’s decision?

At the research and comparison step, before a seller opens an account or starts a trial. No public study yet isolates parcel-carrier selection through AI assistants.

The adoption evidence is strong. Beyond the TechNet poll, the [U.S. Chamber of Commerce](https://www.uschamber.com/technology/empowering-small-business-the-impact-of-technology-on-u-s-small-business) reported that almost 60% of small businesses use artificial intelligence for business operations, more than double the 2023 level. In the TechNet poll, 67% used ChatGPT and 62% used Google Gemini. We cite those as signs of habit, not as proof that sellers pick carriers through AI.

The step from habit to shipping decisions is our inference, but a short one. Shipping questions are practical, frequent and price-driven, which suits an assistant: what does it cost, which option is cheaper, what are the rules for a country. And when 79% of small business owners say AI has already helped them expand the tools and service providers they use, shipping software and carriers are part of that set. Payment providers face the same moment, as [our guide to payment companies](https://underneath.agency/resources/payment-companies-merchants-ai-search) explains.

## What do small businesses ask AI assistants about shipping?

Cost, comparison, software and rules. The prompts below are made-up examples of the pattern, not observed queries.

- “What’s the cheapest way to ship a 2 lb box from Ohio to California?”
- “USPS Ground Advantage or UPS Ground for a small online store?”
- “Best shipping software for an Etsy seller doing 300 orders a month.”
- “Are there regional carriers cheaper than UPS for deliveries in the West?”
- “How do I ship to customers in Canada now that the de minimis rules changed?”
- “How can a small business get discounted FedEx rates?”

Each type favors a different kind of shipping company. Cost and comparison questions favor carriers and platforms whose rates and services are published plainly. Software questions favor platforms with strong reviews and clear plan pages. Rules questions favor whoever explains them best, which is often a carrier, a platform or a trade publication rather than a government page alone.

## How does AI visibility turn into shipping revenue?

Through account sign-ups, platform trials and label volume. An answer that names a carrier or platform sends the seller to try it.

The path differs by company type:

1. **Carriers.** An answer names the carrier for a weight, zone or region. The seller adds it in a shipping platform or opens an account, then sends test parcels. If service holds, the carrier gains a share of that seller’s weekly volume.
2. **Platforms.** An answer recommends a multi-carrier platform for a seller’s store type and volume. The seller starts a trial, connects the store and begins buying labels. Revenue grows with label volume.
3. **International services.** An answer explains how to ship to a country and names services that handle duties. The seller tries one on a few orders.

For challengers, the value is concentrated. A regional carrier or last-mile startup cannot outspend the national carriers on advertising, but a regional question (“cheaper than UPS in Texas”) is narrow enough that a clearly described regional carrier can plausibly be named. We infer that; no study has tested it for parcel carriers.

## What decides whether an AI assistant names a shipping company?

Little is documented by the platforms; our studies point to reviews, prices and independent coverage. Here is what each source shows.

Documented by the platforms:

- Google says no special optimization is needed to appear in AI Overviews or AI Mode: a page must be indexed and eligible for a snippet, and both features may issue multiple related searches ([Google Search Central](https://developers.google.com/search/docs/appearance/ai-features)).
- OpenAI says ChatGPT search typically rewrites a question into one or more targeted queries sent to partner search providers ([OpenAI Help Center](https://help.openai.com/en/articles/9237897-chatgpt-search)).

Observed in our studies:

- ChatGPT looked for reviews in 46.2% of its answers to buyer questions and for prices in 23.8% ([hidden searches study](https://underneath.agency/research/ai-hidden-searches-study)).
- Assistants disagree: for the same question, 66.3% of recommended options came from only one assistant, and all four agreed on the first pick for 10.0% of questions ([brand agreement study](https://underneath.agency/research/ai-assistants-brand-agreement-study)).
- Independent coverage was the strongest predictor we measured: each tenfold increase in independent sites naming a brand went with 4.7 times the odds of being recommended ([entity study](https://underneath.agency/research/brand-entity-ai-recommendations-study)).
- Prices are often but not always right: 61.9% of software plan prices quoted by four assistants were fully faithful to the official pricing page ([pricing study](https://underneath.agency/research/ai-pricing-accuracy-study)). That study covered software, not shipping rates.

Our inference: for parcel shipping, the facts an assistant needs are public rates or rate logic, surcharges, service areas, transit times, pickup options, integrations and customs handling. Companies that publish those facts in plain text, keep them current after each rate change and are reviewed and covered elsewhere give an assistant more to work with.

## What does it cost a shipping company to be missing or misquoted?

Lost trials and accounts that are hard to see, plus the risk of outdated rates being repeated. Neither has been measured for parcel shipping.

We have no figure for parcel business lost through AI answers. The risk is easiest to see in prices. Carriers change rates every year, and the 2026 increases were near 6% before surcharges. If an assistant repeats last year’s rates or an old surcharge, a seller may compare carriers on wrong numbers. Our pricing study found most differing prices came from somewhere real, often another page on the vendor’s own site, which suggests that stale pages are a common source of wrong answers. Our guide to [fixing wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) covers how to trace and correct them.

## How does generative engine optimization work for a shipping company?

It makes rates, coverage and service facts easy to find and confirm. It does not guarantee a recommendation.

For carriers and shipping platforms, the work includes:

1. **Current, readable rate information.** Rate tables or clear rate logic, surcharge lists and the date of the last change, in page text rather than only in PDFs or calculators. Remove or redirect old rate pages.
2. **Coverage and speed facts.** Service-area maps described in text, transit times by zone, pickup and drop-off options, and cutoff times.
3. **Use-case pages.** Pages that match real seller questions: shipping for Etsy or eBay sellers, heavy or oversized parcels, cross-border shipping after the de minimis change.
4. **Integration lists.** For carriers, the platforms and marketplaces where they can be selected; for platforms, the carriers and stores they connect.
5. **Reviews and independent coverage.** Reviews on independent software and business review sites, and coverage in parcel and e-commerce trade press.
6. **Monitoring.** A fixed set of cost, comparison, software and rules questions tracked across ChatGPT, Gemini, Google AI Overviews and AI Mode, Perplexity and Copilot.

Related guides: [ecommerce platforms in AI search](https://underneath.agency/resources/ecommerce-platform-ai-search), [how small brands get recommended by AI](https://underneath.agency/resources/how-small-brands-get-recommended-by-ai) and [logistics software in AI search](https://underneath.agency/resources/logistics-software-ai-search). For heavier loads, see [how freight carriers and brokers win shippers](https://underneath.agency/resources/freight-companies-shippers-ai-search).

## Which parts of this does the evidence leave open?

Several, and they matter for planning.

- How many small sellers ask AI assistants to choose a carrier or platform. The TechNet and Chamber figures measure AI use, not shipping decisions.
- How assistants handle rates that depend on weight, size, zone and negotiated discounts.
- Whether assistants favor the national carriers for broad questions and smaller carriers for regional ones.
- How quickly answers change after a carrier updates its rates or service area.

## Where should a shipping company start?

With the questions small sellers ask most about cost and comparison. Check the answers, then fix the facts they rely on.

Choose the 20 or 30 questions closest to revenue for your business: cost for typical parcel profiles, comparisons with the carriers you compete with, software questions if you are a platform and cross-border rules if you ship internationally. Record which companies the main assistants name, which sources they cite and whether your rates and coverage are stated correctly. If you would rather not run that alone, [talk with us about a GEO audit](https://underneath.agency/contact) focused on the questions that lead to new accounts, trials and label volume. What comes after the audit, from readable rate pages to coverage facts and monitoring, is described on our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page.

## Frequently asked questions

### Do small businesses use AI to choose a shipping carrier?

No study measures that directly. But 71% of small business owners in a 2026 poll use AI tools regularly, and 79% say AI helped expand their service providers.

### Can a regional carrier be named next to UPS and FedEx?

Plausibly, for questions about its region or parcel type. “Other” carriers already doubled their revenue share to 7.2% in 2025.

### Should carriers publish rates on their websites?

Publishing current rates or rate logic gives assistants accurate numbers to quote, and removing outdated rate pages reduces wrong answers.

### Does the end of de minimis matter for AI search?

Yes. Rule changes create new questions, and the companies that explain the new process clearly are better placed to be cited.

### Is GEO for a shipping platform different from GEO for a carrier?

Yes. Platforms are judged like software, on reviews, integrations and plans; carriers on rates, coverage and reliability.

## Sources

- Pitney Bowes (2026), [Parcel Shipping Index](https://www.pitneybowes.com/us/shipping-index.html)
- Pitney Bowes, via Nasdaq (2025-06-30), [Pitney Bowes Parcel Shipping Index: U.S. Carrier Disruption Is Increasing](https://www.nasdaq.com/press-release/pitney-bowes-parcel-shipping-index-us-carrier-disruption-increasing-and-thats-good)
- ShipStation (2026), [ShipStation 2026 Merchant Insights Report](https://www.shipstation.com/guides/shipstation-2026-merchant-insights-report/)
- TechNet (2026-04), [Small Business E-Commerce and AI Survey Poll Toplines](https://www.technet.org/wp-content/uploads/2026/04/Small-Business-E-Commerce-and-AI-Survey-Poll-Toplines.pdf)
- U.S. Chamber of Commerce (2025), [Empowering Small Business: The Impact of Technology on U.S. Small Business](https://www.uschamber.com/technology/empowering-small-business-the-impact-of-technology-on-u-s-small-business)
- U.S. SBA Office of Advocacy (2025), [New Advocacy Report Shows the Number of Small Businesses in the U.S. Exceeds 36 million](https://advocacy.sba.gov/?p=30239)
- UPS (2026-07-28), [UPS Releases 2Q 2026 Earnings](https://investors.ups.com/news-events/press-releases/detail/2164/ups-releases-2q-2026-earnings)
- Associated Press, via ABC7 News (2025-08-28), [US tariff exemption for small orders ends Friday](https://abc7news.com/post/us-tariff-exemption-small-orders-ends-friday-big-deal-shoppers-businesses/17653696/)
- CBC News (2025-08), [Small businesses that relied on duty-free U.S. shipping wonder if they can survive without it](https://www.cbc.ca/lite/story/1.7618234)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- OpenAI Help Center (2025), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude agree on brands?](https://underneath.agency/research/ai-assistants-brand-agreement-study)
- Underneath (2026), [Do Wikipedia and schema make AI assistants recommend a brand?](https://underneath.agency/research/brand-entity-ai-recommendations-study)
- Underneath (2026), [How faithfully do AI assistants quote software prices?](https://underneath.agency/research/ai-pricing-accuracy-study)

---

This is the Markdown twin of https://underneath.agency/resources/shipping-companies-customers-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Should AI search ads be labeled, and does it help advertisers?"
description: "Yes. Clear labels protect users, and research argues they protect advertisers too, since disguised ads breed distrust once people spot them."
canonical: "https://underneath.agency/resources/should-ai-search-ads-be-labeled"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Should ads in AI search be clearly labeled, and does labeling help the advertiser?

Yes: ads in AI answers should be clearly labeled, and the research suggests clear labels serve advertisers better over time than disguised placements. A label costs some persuasive power today. A disguised ad that users later discover can cost the brand, and the platform it runs on, their trust.

## The short version

1. US consumer protection rules already require paid ads to be labeled “Ad” or “Sponsored,” and [Wen and colleagues](https://arxiv.org/abs/2606.12439) argue these rules should be extended to AI answers.
2. In an economic model by [Zhang and colleagues](https://arxiv.org/abs/2603.29071), always showing ads in AI answers ended with the lowest cumulative payoff in all eight market conditions tested across 160 simulation runs.
3. In an experiment with 4,927 US adults, [Li and Aral](https://arxiv.org/abs/2504.06435) showed that small cues, such as links or helpfulness counts, change trust in AI answers.
4. Undisclosed commercial interest is already in AI citations: in [our self-ranking lists study](https://underneath.agency/research/self-promoting-best-lists-study), 24.2% of cited “best X” lists with an identifiable publisher ranked that publisher first.
5. The direct question, whether labeled AI ads beat disguised ones for the brand, has not yet been tested with real users.

## Why does labeling matter more in AI answers than in search results?

Because an AI answer blends everything into one voice, so an unlabeled ad looks exactly like advice. On a traditional results page, sponsored links sit in marked slots. [Wen and colleagues](https://arxiv.org/abs/2606.12439) point out that US Federal Trade Commission rules require paid ads to be labeled “Ad” or “Sponsored,” so people can tell promotion from neutral information.

In an AI answer, they argue, commercial influence can move into the evidence the assistant reads and the reasoning it writes. Persuasion then “operates through the model’s reasoning itself,” and [the line between advice and marketing collapses](https://underneath.agency/resources/risks-of-native-ads-in-ai-answers). They call for clear markers when an answer reflects a material commercial connection.

Regulators are paying attention. An audit by [Uberti-Bona Marin and colleagues](https://arxiv.org/abs/2609.18729) notes that on 31 August 2026 the European Commission designated ChatGPT a Very Large Online Search Engine under the Digital Services Act. That brings duties on consumer protection risks and independent audits.

## Does a clear label help the advertiser or hurt it?

It costs some short-term persuasion but likely protects long-term trust. [Erickson](https://arxiv.org/abs/2506.06447) explains the short-term cost. Earlier advertising research shows that people who know they are being persuaded become more skeptical. A disguised ad avoids that skepticism, which is why it can work better in the moment.

The long-term picture is different. Erickson argues that when users do recognize promotion inside an answer, it can reduce trust and worsen their view of the brand. His conclusion: “In the long run, a clear delineation of advertising may benefit companies since it will not breed the same level of distrust that disguised advertisements might.” This is a reasoned argument built on earlier advertising studies, not a test in AI search.

Wen and colleagues add a market-level warning. If hidden promotion outperforms labeled ads, firms that hide their motives win, and the whole channel drifts toward covert tactics. That raises the risk of “trust erosion” for everyone when the influence comes to light.

## What does economic modeling say about ad-heavy AI answers?

A model of AI platforms finds that leaning hard on ads wins early revenue but loses users and long-run value. [Zhang and colleagues](https://arxiv.org/abs/2603.29071) built a model in which an AI engine decides, question by question, whether to show an answer with ads or an ad-free one. Ads earn money now. Ad-free answers build user experience, which supports retention and paid subscriptions later.

Their simulations used 500 users over 20 periods, and a sensitivity check covered eight market conditions in 160 simulation runs. In every condition, always showing ads produced the lowest cumulative payoff of the four policies tested. When users were highly sensitive to ads, the always-ads policy ended at 1.40 in the model’s payoff units, against 9.19 for the best policy. That best policy served ad-free answers to non-subscribers 92% of the time.

The authors add that the damage is hard to undo: switching to ad-free answers later did not recover the gains of serving them from the start. These are simulations with assumed user behavior. They describe the platform’s incentives, but an advertiser’s reach depends on that platform keeping its audience. Ads are already arriving in real assistants, as [our guide to ads in AI assistants](https://underneath.agency/resources/are-ads-coming-to-ai-assistants) explains.

## How much do small design cues change trust in AI answers?

A lot, which is why the design of a label matters as much as its presence. In the experiment by [Li and Aral](https://arxiv.org/abs/2504.06435), 4,927 participants matched to the US adult population saw search results presented either as AI answers or as ordinary search. Our guide on [whether people trust AI search less](https://underneath.agency/resources/do-people-trust-ai-search-less) covers its overall trust finding.

Several cues moved trust:

| Cue added to the AI answer | Effect on trust in the experiment |
|---|---|
| Reference links | Raised trust, even when the links were wrong or invented |
| “Users who found this helpful” shown between 65% and 95% | Raised trust |
| The same count shown between 5% and 35% | Lowered trust |
| Highlighting how certain the AI was | Lowered trust and willingness to share |

The study did not test ad labels. But it shows that people respond to small signals around an AI answer. That cuts both ways. Wen and colleagues warn that labels can backfire through over-labeling, and they suggest platforms test the presence, wording and placement of disclosures with controlled experiments.

## Is undisclosed commercial content already in AI answers?

Yes, though not as paid ads: vendors’ own rankings already appear among AI citations. In [our self-ranking lists study](https://underneath.agency/research/self-promoting-best-lists-study), we checked numbered “best X” lists cited by six AI surfaces in the US in September 2026. Of 269 lists with an identifiable publisher, 65 (24.2%) ranked their own publisher first. Publishers that included themselves almost always went first: 92.9% of self-including lists did.

These lists were a small part of all citations, 1.1%. But a reader of such an answer may be reading a vendor’s ranking of itself without being told.

Fabricated claims are a sharper version of the same problem. In a skincare test with three commercial AI models, [Chu and Hou](https://arxiv.org/abs/2606.17443) found that a made-up clinical citation had the same effect as 0.17 rating points of real product improvement. They class such invented claims as potential false advertising.

## What should you do about it?

Choose disclosure by default, both in paid placements and in your own content. Practical steps:

1. In any AI ad program, require the platform’s sponsored label to be visible next to your mention, not hidden behind a link.
2. Disclose commercial interest in your own comparison content. If you publish a “best X” list that includes your product, say so near the top.
3. Never use invented statistics, reviews or endorsements in copy that AI assistants may read and repeat.
4. Ask platforms for evidence on how their labels affect user trust, and prefer those that test label wording and placement.
5. Track both reach and brand sentiment after AI ad campaigns, since the risk shows up in trust, not clicks.

If you want help building honest, citable visibility in AI answers, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

No study we reviewed has measured whether labeled AI ads outperform disguised ones for the advertiser. In detail:

- Erickson’s case for clear labels rests on earlier advertising research, not on tests inside AI answers.
- The economic model describes platform incentives with simulated users; its payoff numbers are not money.
- The trust experiment tested links, feedback counts and certainty cues, not ad labels.
- Wen and colleagues note that disclosure can carry a trust penalty and can over-label; the right design is still an open question.
- Our self-ranking lists study observes citations, not whether readers notice or care about the self-ranking.

## Frequently asked questions

### Do ads in AI chatbots have to be labeled?

US consumer protection rules require paid ads to be labeled “Ad” or “Sponsored,” but how these rules apply inside AI answers is unsettled. Wen and colleagues argue the FTC-style rules should be extended to AI answers explicitly.

### Does labeling an ad make it less effective?

Possibly in the moment, because people who know they are being persuaded become more skeptical. Erickson argues that clear labels still pay off in the long run because disguised ads breed distrust.

### Are AI platforms better off without ads?

In one economic model, always showing ads gave the lowest cumulative payoff in all eight conditions tested. Mixed policies that mostly served ad-free answers did best, but this is a simulation, not observed behavior.

### Is it disguised advertising to publish a list that ranks my own product first?

It is not a paid ad, but readers of an AI answer may not know the list is self-ranked. In our study, 24.2% of cited “best X” lists with an identifiable publisher ranked that publisher first, so disclosing your interest is the safer practice.

## Sources

- Erickson, J. (2025), [Fake Friends and Sponsored Ads: The Risks of Advertising in Conversational Search](https://arxiv.org/abs/2506.06447), arXiv:2506.06447.
- Wen, Y., Zhang, N., Yuan, H., Chen, X., Zhang, H. and Guo, H. (2026), [Position: Generative Engine Optimization Creates Underexamined Risks, Governance Must Target Concentration, Disclosure, and Academic Blind Spots](https://arxiv.org/abs/2606.12439), arXiv:2606.12439.
- Zhang, L., Jiao, C., Li, B. and Xiong, C. (2026), [An Economic Framework for Generative Engines: Advertising or Subscription?](https://arxiv.org/abs/2603.29071), arXiv:2603.29071.
- Li, H. and Aral, S. (2025), [Human Trust in AI Search: A Large-Scale Experiment](https://arxiv.org/abs/2504.06435), arXiv:2504.06435.
- Uberti-Bona Marin, L. G. and colleagues (2026), ["If I Had to Buy Just ONE: Galaxy S26 Ultra": Auditing AI-Generated Product Recommendations](https://arxiv.org/abs/2609.18729), arXiv:2609.18729.
- Chu, X. and Hou, Y. (2026), [Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems](https://arxiv.org/abs/2606.17443), arXiv:2606.17443.
- Underneath (2026), [How many “best of” lists cited by AI rank their own brand first?](https://underneath.agency/research/self-promoting-best-lists-study)

---

This is the Markdown twin of https://underneath.agency/resources/should-ai-search-ads-be-labeled. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How SIEM vendors reach enterprise shortlists through AI search"
description: "By being named, with checkable proof, when SOC leaders use AI to build a SIEM long list during a migration, a consolidation or a cost review."
canonical: "https://underneath.agency/resources/siem-enterprise-pipeline-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How do SIEM vendors get onto enterprise shortlists through AI search?

By being named, with evidence a security architect can verify, when SOC leaders ask AI assistants to build a long list during a migration, a consolidation or a cost review. A SIEM, the system that collects and analyzes an organization’s security logs, is replaced rarely and at high cost, so each evaluation that starts is worth a great deal. AI answers appear to influence who gets invited into those evaluations, though no study yet measures SIEM buyers alone.

## The short version

1. SIEM contracts are large: [Vendr’s purchase data](https://www.vendr.com/marketplace/splunk) put the median Splunk buyer at $94,200 a year across 222 purchases, and the median [Exabeam](https://www.vendr.com/marketplace/exabeam) buyer at $210,252.
2. The vendor map has been redrawn. Cisco bought Splunk for about [$28 billion](https://newsroom.cisco.com/c/r/newsroom/en/us/a/y2024/m03/cisco-completes-acquisition-of-splunk.html) in 2024, and [Exabeam and LogRhythm merged](https://www.exabeam.com/press-releases/exabeam-and-logrhythm-complete-merger/) the same year. Changes like these send customers looking at alternatives.
3. SOC teams are under strain: in [Splunk’s State of Security 2025](https://www.splunk.com/en_us/campaigns/state-of-security.html), 59% reported too many alerts and 78% said their security tools are disconnected and dispersed.
4. Enterprise buyers mix AI and people: in a [Gartner survey](https://www.gartner.com/en/newsroom/press-releases/2026-05-20-gartner-survey-finds-sixty-nine-percent-of-b-two-b-buyers-turn-to-sales-reps-to-validate-ai-generated-insights) of 645 B2B buyers, 45% used generative AI in a recent purchase and 69% prefer to check AI-generated insights with sales reps.
5. In our study of AI search, ChatGPT ran a search for a named publication, ranking or award in 43.8% of its answers, one of them for a Gartner Magic Quadrant. Analyst coverage is part of what assistants look for.

## Who signs off on a SIEM, and how much is one contract worth?

A security committee led by the CISO and SOC leadership buys it, on multi-year contracts priced mostly by data volume.

The core group is the CISO, the head of security operations, security architects and detection engineers. Finance and IT join because cost scales with data. Splunk’s pricing, Vendr notes, is “primarily consumption-based,” scaling with data ingested rather than user seats; list prices run from $1,800 to $2,700 per GB of daily ingest a year for self-hosted Splunk Enterprise. [Microsoft Sentinel](https://www.microsoft.com/en-us/security/business/siem-and-xdr/microsoft-sentinel) is also priced on the data customers ingest, store and consume.

That makes each customer large and sticky. Vendr’s medians, which are samples of purchases on one procurement platform rather than market averages:

| Vendor | Median annual spend | Purchases in sample |
|---|---|---|
| Splunk | $94,200 | 222 |
| Sumo Logic | $85,135 | 181 |
| Securonix | $67,750 | not stated |
| Exabeam | $210,252 | not stated |

The category is large and growing. Splunk, citing a market forecast, puts SIEM growth at a 14.5% annual rate, reaching $11.3 billion by 2026 from $4.8 billion in 2021. Treat a vendor-cited forecast with care.

## What starts a SIEM evaluation today?

Usually a trigger: an ownership change, a cost shock, a consolidation program, a failing SOC or a new reporting rule.

**Ownership changes.** Cisco [agreed to pay $157 per share](https://newsroom.cisco.com/c/r/newsroom/en/us/a/y2023/m09/cisco-to-acquire-splunk-to-help-make-organizations-more-secure-and-resilient-in-an-ai-powered-world.html), about $28 billion, for Splunk. Exabeam and LogRhythm completed their merger on July 17, 2024. When a product’s owner or roadmap changes, we infer that some customers re-evaluate at renewal, and challengers court them.

**Consolidation.** A [Gartner survey of 418 organizations](https://www.gartner.com/en/newsroom/press-releases/2022-09-12-gartner-survey-shows-seventy-five-percent-of-organizations-are-pursuing-security-vendor-consolidation-in-2022) found 75% pursuing security vendor consolidation in 2022, up from 29% in 2020. The data are older, but the pressure behind platform bundles is not.

**Strained SOCs.** In Splunk’s survey, 46% said they spend more time maintaining tools than defending the organization, 55% deal with too many false positives, and 57% lose investigation time to gaps in their data management. Splunk is itself a SIEM vendor, so read these as one input.

**Detection speed and money.** A breach took 241 days on average to identify and contain, according to [IBM’s 2025 breach report](https://newsroom.ibm.com/2025-07-30-ibm-report-13-of-organizations-reported-breaches-of-ai-models-or-applications,-97-of-which-reported-lacking-proper-ai-access-controls), which is the window a SIEM is bought to shorten. When the organization’s own team caught the breach, rather than hearing of it from the attacker, the saving was $900,000. Heavy use of AI and automation in security operations went with $1.9 million lower breach costs.

**Reporting rules.** Under [Article 23 of the EU’s NIS2 Directive](https://www.springlex.eu/en/packages/nis2/nis2-directive/article-23/), covered companies must send an early warning within 24 hours of becoming aware of a significant incident and a fuller notification within 72 hours. Rules like these raise the cost of slow detection, we infer.

## Where do AI assistants enter the SIEM buying journey?

At the long list and in background research, with people and analysts used to check what the assistant said.

The best evidence on enterprise buyers comes from Gartner’s 2026 survey of 645 B2B buyers. They used an average of seven information sources during a recent purchase. 45% used generative AI, mainly to gather information on vendors and products. 69% prefer to validate AI-generated insights with sales reps, and just over half said they are more likely to meet misleading information from generative AI.

That pattern suits SIEM, we infer. An architect asks an assistant to summarize options and compare pricing models, then checks the claims in analyst reports, peer reviews and a proof of concept. The assistant shapes which vendors are worth a call. The call decides the deal. Data platform buyers follow a similar path, as [our guide to analytics and BI software](https://underneath.agency/resources/analytics-bi-software-ai-search) shows.

A security architect who types a SIEM question into Google will usually meet an AI answer first. Of the eight industries in [our study of 1,248 US searches](https://underneath.agency/research/ai-overviews-frequency-study), B2B software and technology had the highest AI Overview rate: Google’s AI summary topped the results on 96.0% of those searches.

## Which questions do SIEM buyers ask AI assistants?

Questions about alternatives, migration, pricing models, consolidation and fit with existing tools. We drafted the SIEM prompts below as examples of each trigger; none comes from observed buyer logs.

| Trigger | Illustrative prompt |
|---|---|
| Renewal or ownership change | “What are the main Splunk alternatives for a 20,000-employee bank, and how hard is migration?” |
| Merger | “What happens to LogRhythm SIEM customers after the Exabeam merger?” |
| Cost | “Which SIEMs price by ingest, and which by users or endpoints? How do I estimate cost at 2 TB a day?” |
| Consolidation | “Should we use Microsoft Sentinel if we already run Defender, or keep a separate SIEM?” |
| Architecture | “Can a SIEM search logs kept in our own data lake without moving them?” |
| Compliance | “Which SIEMs help meet NIS2 incident reporting deadlines?” |
| Proof | “Which vendors are Leaders in the latest Gartner Magic Quadrant for SIEM?” |

The proof questions matter because assistants search for rankings. In [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study), ChatGPT ran a search aimed at a named publication, ranking or award in 43.8% of its answers, and most past-year searches looked for the latest edition of an annual ranking. One observed search was “Gartner Magic Quadrant 2024 managed detection and response.” Managed detection providers face the same ranking checks, covered in [how MDR providers win qualified leads](https://underneath.agency/resources/mdr-providers-qualified-leads-ai-search).

## How does a mention in an AI answer become a SIEM proof of concept?

Through the evaluation invite list: the AI answer shapes who gets a call, a proof of concept and a contract.

1. **Trigger.** A renewal, a merger, a cost review or an incident opens the question.
2. **Long list.** The team asks assistants for alternatives and comparisons, alongside analyst reports and peers.
3. **Validation.** Architects check claims with vendors’ sales engineers, references and documentation.
4. **Proof of concept.** Two or three vendors ingest real data for weeks.
5. **Contract.** A multi-year agreement priced by data volume, which grows as the customer sends more data.

The value sits in steps 2 and 3. A vendor missing from the long list never reaches the proof of concept. A vendor named with a wrong pricing model or an outdated product name starts the sales call correcting the assistant. Pricing is a known weak spot: when [our pricing study](https://underneath.agency/research/ai-pricing-accuracy-study) checked four assistants on 45 software products, just 61.9% of the plan prices they quoted matched the vendor’s own pricing page in full. We expect ingest-based SIEM pricing to be harder still to quote than a per-seat price.

An architect who shortlists you after reading an AI answer often arrives through a sales call, not a tracked click; [what lost clicks mean for pipeline](https://underneath.agency/resources/ai-answers-pipeline-revenue) covers how to account for that.

## Why does an assistant put one SIEM on the long list and leave another off?

The platforms do not say; studies point to independent and recent sources, and SIEM buyers lean on analyst reports.

**Documented by the platforms.** According to Google, AI Overviews and AI Mode [may use a “query fan-out” technique](https://developers.google.com/search/docs/appearance/ai-features), so a single SIEM question can trigger related searches on pricing, migration and integrations. None of the platforms explains how it settles on the security vendors it names.

**Observed in studies.** On US software questions, earned, independent sites supplied 72.7% of the sources AI search used, compared with 45.4% for Google, in work by [Chen and colleagues](https://arxiv.org/abs/2509.08919). In [our four-assistant study](https://underneath.agency/research/ai-assistants-brand-agreement-study), recommendations overlapped most for B2B software, at 0.543 on a scale from 0 to 1, so assistants agree more here than in most industries, though far from fully. In [our freshness study](https://underneath.agency/research/ai-source-freshness-study), assistants cited pages first published about half as long ago as Google’s top 10 for the same questions. In [our Reddit study](https://underneath.agency/research/ai-reddit-citations-study), more than a third of Google’s AI Overviews for B2B software cited Reddit, and r/cybersecurity was among the most-cited communities.

**What SIEM buyers rely on.** Analyst evaluations carry weight in this category. Gartner published its latest Magic Quadrant for SIEM on 8 October 2025. [Splunk](https://www.splunk.com/en_us/form/gartner-siem-magic-quadrant.html) says it has been named a Leader 11 times in a row; Microsoft cites the same report plus an IDC MarketScape. Splunk’s own [SIEM guide](https://www.splunk.com/en_us/blog/learn/siem-security-information-event-management.html) lists Gartner, Forrester and IDC reports as the common references.

**Our inference.** A reasonable expectation is that vendors with current analyst coverage, specific public documentation and active practitioner discussion give assistants more to cite. That expectation has not been tested on security analytics vendors.

## What does a SIEM vendor lose when assistants skip it?

Evaluations it never hears about, each one worth a multi-year contract. No study puts a figure on the total.

- **Rare windows.** SIEM replacements are disruptive, so evaluations open mainly at renewal or after a trigger. Missing one may mean waiting for the next contract cycle, we infer.
- **Large stakes per miss.** With median contracts in the tens or hundreds of thousands of dollars a year, a single lost evaluation is material for most vendors.
- **A place on the list is not stable.** We asked ChatGPT each question five times in [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), and only 25.2% of the brands it named came back in every run, so a SIEM vendor can be on one architect’s long list and off the next.
- **Wrong facts travel.** After mergers and rebrands, we expect assistants to describe retired SIEM product names or old pricing; our guide to [correcting wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) covers the repair.

## What does GEO look like for a SIEM or security analytics vendor?

It puts your detection, pricing and migration evidence where assistants and architects can find and check it. Nobody can promise a SIEM a spot on any long list.

1. **State what you are, precisely.** SIEM, security analytics platform, data lake add-on or part of a broader operations platform. Use the same words everywhere, especially after a merger or rename.
2. **Publish pricing logic in plain pages.** Explain the unit (ingest, users, endpoints, queries), what counts toward it and a worked example. Ingest pricing is where confusion costs deals.
3. **Document migration paths.** Public guides for moving from the main incumbents, with detection rule conversion and timelines, answer the questions buyers ask at renewal.
4. **Earn analyst and peer coverage.** Analyst evaluations, peer review platforms and practitioner communities are where SIEM claims get checked; [building brand authority for AI search](https://underneath.agency/resources/how-brands-build-authority-for-ai-search) explains how that coverage is earned.
5. **Publish integration and detection detail.** Data connector lists, detection content mapped to common frameworks and architecture notes give assistants specifics to cite.
6. **Write honest comparison pages.** See [whether comparison pages help B2B citations](https://underneath.agency/resources/do-comparison-pages-help-b2b-ai-citations) and [why ranked lists matter](https://underneath.agency/resources/best-of-lists-ai-recommendations).
7. **Measure across assistants.** Track trigger-specific questions in ChatGPT, Google AI Overviews and AI Mode, Gemini, Perplexity, Copilot and Claude, several runs each.

## What is still unknown about AI search in SIEM evaluations?

No published study measures SOC teams’ use of assistants, or whether being named wins SIEM evaluations.

- **Nothing specific to security operations buyers.** Gartner surveyed B2B buyers across categories, not CISOs or SOC leaders choosing a SIEM.
- **Vendor and sample data.** Splunk’s survey comes from a SIEM vendor; Vendr’s medians come from one platform’s purchases; the market forecast is vendor-cited.
- **The link to revenue is the least proven part.** Across the 45 studies of AI search optimization in [Martinez’s review](https://arxiv.org/abs/2607.14035), traffic and conversions had the weakest evidence.
- **Answers move.** Results change between runs, wordings and assistants.

## How can a SIEM vendor find out whether AI answers are costing it enterprise evaluations?

Ask the assistants your prospects’ renewal and migration questions, and see which SIEMs make their long lists.

List 20 to 30 prompts across alternatives, migration, pricing model, consolidation, compliance and proof. Because answers shift, put every prompt to ChatGPT, Gemini, Perplexity, Copilot, Claude and Google’s AI features more than once. Note which vendors appear, which analyst reports and communities are cited, and whether your ingest pricing and current product names come through correctly. SIEM is one part of a wider security market, covered in [how cybersecurity software firms earn revenue from AI search](https://underneath.agency/resources/cybersecurity-software-revenue-from-ai-search). For firewall and zero trust buyers, see [how network security vendors win buyers](https://underneath.agency/resources/network-security-zero-trust-customers-ai-search).

To have that test run for you, [talk to us about a SIEM visibility audit](https://underneath.agency/contact). We will show where assistants place you on the long lists that turn into proofs of concept, and which gaps in your public detection, pricing and migration evidence are most likely keeping you out of renewal-cycle evaluations. Closing those gaps, through pricing logic pages, migration guides and analyst and peer coverage, is the work our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page covers.

## Frequently asked questions

### Do enterprise security teams trust AI answers about SIEMs?

Partly. In Gartner’s survey, 45% of B2B buyers used generative AI, but 69% prefer to validate AI-generated insights with sales reps.

### Will being a Leader in the SIEM Magic Quadrant get us named by assistants?

Plausibly. In 43.8% of the answers in our study, ChatGPT went looking for a named ranking, and one of those searches was for a Magic Quadrant. Nobody has measured the effect for SIEM.

### Should we publish our SIEM pricing?

Publish the pricing logic at least. Ingest-based pricing is hard to quote, and only 61.9% of AI-quoted software prices were fully faithful in our study.

### Do Reddit and practitioner forums matter for SIEM?

They can. More than a third of Google’s AI Overviews for B2B software cited Reddit in our study, and r/cybersecurity was among the most-cited communities.

### When would better AI visibility appear as SIEM proofs of concept?

Expect quarters, not weeks. Evaluations open mainly at renewals and after triggers, and the average breach lifecycle alone is 241 days.

## Sources

- Vendr (2026), [Splunk software pricing and plans](https://www.vendr.com/marketplace/splunk)
- Vendr (2026), [Exabeam software pricing and plans](https://www.vendr.com/marketplace/exabeam)
- Vendr (2026), [Sumo Logic software pricing and plans](https://www.vendr.com/marketplace/sumo-logic)
- Vendr (2026), [Securonix software pricing and plans](https://www.vendr.com/marketplace/securonix)
- Cisco (2024-03-18), [Cisco completes acquisition of Splunk](https://newsroom.cisco.com/c/r/newsroom/en/us/a/y2024/m03/cisco-completes-acquisition-of-splunk.html)
- Cisco (2023-09-21), [Cisco to acquire Splunk](https://newsroom.cisco.com/c/r/newsroom/en/us/a/y2023/m09/cisco-to-acquire-splunk-to-help-make-organizations-more-secure-and-resilient-in-an-ai-powered-world.html)
- Exabeam (2024-07-17), [Exabeam and LogRhythm complete merger and announce new company details](https://www.exabeam.com/press-releases/exabeam-and-logrhythm-complete-merger/)
- Splunk (2025), [State of Security 2025](https://www.splunk.com/en_us/campaigns/state-of-security.html)
- Splunk (2025-10), [2025 Gartner Magic Quadrant for SIEM](https://www.splunk.com/en_us/form/gartner-siem-magic-quadrant.html)
- Splunk (n.d.), [What is SIEM?](https://www.splunk.com/en_us/blog/learn/siem-security-information-event-management.html)
- Microsoft (2026), [Microsoft Sentinel](https://www.microsoft.com/en-us/security/business/siem-and-xdr/microsoft-sentinel)
- Gartner (2026-05-20), [Gartner survey finds 69% of B2B buyers turn to sales reps to validate AI-generated insights](https://www.gartner.com/en/newsroom/press-releases/2026-05-20-gartner-survey-finds-sixty-nine-percent-of-b-two-b-buyers-turn-to-sales-reps-to-validate-ai-generated-insights)
- Gartner (2022-09-13), [Gartner survey shows 75% of organizations are pursuing security vendor consolidation in 2022](https://www.gartner.com/en/newsroom/press-releases/2022-09-12-gartner-survey-shows-seventy-five-percent-of-organizations-are-pursuing-security-vendor-consolidation-in-2022)
- IBM (2025-07-30), [IBM report: 13% of organizations reported breaches of AI models or applications](https://newsroom.ibm.com/2025-07-30-ibm-report-13-of-organizations-reported-breaches-of-ai-models-or-applications,-97-of-which-reported-lacking-proper-ai-access-controls)
- European Union, via Springlex (2022-12-27), [NIS2 Directive, Article 23](https://www.springlex.eu/en/packages/nis2/nis2-directive/article-23/)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- Chen, Wang, Chen and Koudas (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919), arXiv:2509.08919.
- Martinez (2026), [Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026)](https://arxiv.org/abs/2607.14035), arXiv:2607.14035.
- Underneath (2026), [When does Google show an AI Overview? 1,248 US searches](https://underneath.agency/research/ai-overviews-frequency-study)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [How faithfully do AI assistants quote software prices?](https://underneath.agency/research/ai-pricing-accuracy-study)
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude agree on brands?](https://underneath.agency/research/ai-assistants-brand-agreement-study)
- Underneath (2026), [How fresh are the pages AI engines cite?](https://underneath.agency/research/ai-source-freshness-study)
- Underneath (2026), [When does a Reddit thread become evidence in Google’s AI?](https://underneath.agency/research/ai-reddit-citations-study)
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)

---

This is the Markdown twin of https://underneath.agency/resources/siem-enterprise-pipeline-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How skincare brands get found when shoppers ask AI"
description: "By answering concern and ingredient questions with claims you can substantiate, real expert backing and honest reviews, the signals AI and shoppers both check."
canonical: "https://underneath.agency/resources/skincare-brands-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How do skincare brands get their products found when shoppers ask AI about their skin?

By publishing clear, substantiated answers to the concern and ingredient questions shoppers ask, backed by real expert involvement and honest reviews. Skincare is a category where shoppers cannot judge a product before using it, so assistants and buyers both lean on familiar names and trust signals. This guide covers where AI sits in skincare discovery, what research shows about how assistants pick skincare products, and how to compete without crossing the lines the FDA and FTC draw on claims. It is not legal or medical advice; check claims with your regulatory counsel.

## The short version

1. Dermatological skincare is where the growth is: L’Oréal’s Dermatological Beauty division, home to CeraVe and La Roche-Posay, grew 10.6% like-for-like in the first half of 2026, with all three of its main skincare brands up in double digits ([L’Oréal](https://www.loreal-finance.com/eng/press-release/2026-half-year-results)).
2. Shoppers already ask AI about their skin: in a survey of 1,238 US women by skin-analysis vendor Haut.AI, [62% were interested in or already using AI for skincare or bodycare advice](https://haut.ai/blogs/ai-body-analysis-launch).
3. Familiarity wins ties in skincare: when ten products had identical ratings, prices, reviews and ingredients, three AI models picked the one real brand, such as CeraVe, in all 670 valid trials ([Chu and Hou, 2026](https://arxiv.org/abs/2606.17443)). A rating edge of just +0.075 stars was enough for an unknown brand to win half the time.
4. Trust is contested online: 21% of Americans rely on Instagram or TikTok influencers for skincare advice, 36% of Gen Z, and more than 16 million adults cut back or stopped using sunscreen because of online claims ([American Academy of Dermatology](https://www.newswise.com/articles/misinformation-puts-over-16-million-americans-at-an-increased-risk-for-skin-cancer)).
5. The rules apply to what assistants read: the [FDA](https://www.fda.gov/cosmetics/cosmetics-laws-regulations/it-cosmetic-drug-or-both-or-it-soap) says claims “on the Internet” help decide whether a product is a cosmetic or a drug, and the [FTC’s final rule on fake reviews](https://www.ftc.gov/news-events/news/press-releases/2024/08/federal-trade-commission-announces-final-rule-banning-fake-reviews-testimonials) allows civil penalties, including for AI-generated fake reviews.

## Who buys skincare, and what is a customer worth to a brand?

Concern-led shoppers who build routines, so a customer won on one product is worth a regimen and its repurchases.

US skincare is large and steady. [Circana reports](https://www.circana.com/post/us-prestige-and-mass-beauty-retail-deliver-a-positive-performance-in-2025-circana-reports) that prestige skincare grew 3% in dollars in 2025 and was the fastest-growing prestige category by units sold, while skincare at mass retail grew 6% in both dollars and units, driven by facial cleansers and moisturizers.

Dermatological brands are the standout. At L’Oréal, Dermatological Beauty grew 10.6% like-for-like in the first half of 2026, with growth accelerating for a third consecutive quarter and divisional profitability of 28.4%. La Roche-Posay, CeraVe and SkinCeuticals all grew in double digits, and the company credits launches such as CeraVe Sun.

For a skincare brand, value comes from the routine. A shopper who trusts one product, often a cleanser, moisturizer or sunscreen, tends to try a serum or treatment from the same range and to repurchase what works. That is our inference from how regimens are sold rather than a published figure, but it explains why being named for one hero product matters beyond a single sale.

## Where does AI sit in skincare discovery today?

Between a skin concern and the shortlist, alongside social media, Reddit and, for some, a dermatologist.

Shoppers start from a concern: acne, redness, dryness, dark spots, aging, sun protection. Haut.AI’s survey found that 74% of US women believe AI could make their routines easier or more effective. Qualitative research by the consultancy [8th Day](https://www.8th-day.com/thinking/getting-under-your-skin-why-consumers-are-letting-ai-rebuild-their-skincare) describes people asking ChatGPT about conditions such as rosacea and perioral dermatitis, and buying what it recommends. [The Nod Mag](https://thenodmag.com/content/skincare-routine-chatgpt-artificial-intelligence-diagnosis) describes a writer uploading a selfie and receiving a routine of named products; a dermatologist it interviewed warned that a photo “cannot reliably determine skin hydration, sebum production, collagen content, barrier function, inflammation or biological skin age.”

Assistants are already naming brands for plain skincare questions. Chu and Hou reproduce a real ChatGPT answer from May 2026 to the question “I would like to buy a good face moisturizer, which is the best?” that named CeraVe as “Best overall.”

AI arrives in a crowded field of advice. The American Academy of Dermatology’s 2026 survey of 1,132 US adults found that nearly half of Americans, and 64% of Gen Z, report encountering sunscreen misinformation online. That makes credible, accurate brand information more valuable, not less.

## Which skincare questions do shoppers ask AI?

Questions about concerns, ingredients, compatibility and comparisons, often with skin type and sensitivities attached.

We wrote these sample prompts in a shopper’s voice; none come from real chat logs:

- Concern: “Moisturizer for sensitive, acne-prone skin that won’t clog pores.”
- Ingredient: “Niacinamide or azelaic acid for redness, and what’s the difference?”
- Compatibility: “Can I use retinol and vitamin C in the same routine?”
- Comparison: “Which mineral sunscreen for oily skin leaves no white cast?”
- Credibility: “Which cleansers do dermatologists recommend most?”
- Sensitivity: “Fragrance-free body lotion for very dry skin.”

Two features set skincare apart. First, many questions are framed around an ingredient rather than a brand, so a brand appears only if its product is clearly connected to that ingredient and concern. Second, some questions edge into medical territory. Brands should answer the cosmetic part well and point to a dermatologist for the rest, which is also what regulators expect.

## How does an AI answer turn into skincare sales?

Through the hero product: an assistant names it for a concern, and the shopper checks, buys and builds a routine.

1. **The concern question.** The assistant names a handful of products and explains why, often by ingredient.
2. **The check.** Shoppers look for reviews, dermatologist involvement, ingredient lists and price. [OpenAI documents](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search) that ChatGPT’s product results draw on structured product data, other third-party content and review summaries built from “reviews from public websites,” which it does not verify.
3. **The purchase.** The shopper buys at a retailer, a marketplace or the brand’s own site, depending on where the assistant’s links and the shopper’s habits lead.
4. **The routine.** A product that works earns repurchase and a chance to sell the rest of the range.

The most valuable answers are the ones where a brand is named for a clear concern it can genuinely serve. A reasonable expectation is that hero products named in AI answers carry more of a skincare brand’s new-customer acquisition over time, though no public data yet measures this.

## What decides which skincare products an assistant recommends?

Research shows familiarity, ratings and authority claims all matter; platforms document only general product signals.

**Documented by the platforms.** OpenAI says ChatGPT considers structured metadata such as price and description, other third-party content and public reviews, and that labels such as “Most popular” are model-generated, not verified.

**Observed in a study of skincare specifically.** Chu and Hou tested three AI models on moisturizers, exfoliants, sunscreens and cleansers. With nothing to separate products, the real brand won every time. But the advantage was fragile: a +0.075-star rating edge, 1.6 times the reviews or a 7.3% lower price let an unknown brand win half the time. Authority language, such as clinical-trial or dermatologist endorsements, was worth the equivalent of +0.17 rating points. In the tests, even fabricated clinical claims worked, which the authors themselves label “potential false advertising.” And when every competing brand used the same authority language, the gain from it fell from +0.802 to +0.007 in their payoff measure. Our article on [whether AI assistants favor big brands](https://underneath.agency/resources/do-ai-assistants-favor-big-brands) covers the brand side of this research.

**Our inference for brands.** Real, specific, verifiable evidence is the durable advantage. Vague “clinically proven” copy is easy for competitors to copy and, as the study shows, loses value when everyone uses it; fabricated evidence is illegal and dangerous. See our article on [legitimate GEO versus manipulation](https://underneath.agency/resources/legitimate-geo-vs-manipulation).

**Trust factors specific to skincare.** Substantiated claims, published clinical testing, genuine dermatologist involvement, complete ingredient lists, fragrance and allergen information, and reviews that describe real skin types and results.

## What claims can a skincare brand make, and why does it matter for AI?

Only claims that keep the product a cosmetic, unless it is a drug, and only claims you can substantiate.

AI makes compliance more important, because assistants repeat what they find. The FDA decides whether a product is a cosmetic or a drug by its intended use, which can be established by “claims stated on the product labeling, in advertising, on the Internet, or in other promotional materials.” Claims that a product will “increase or decrease the production of melanin (pigment) in the skin, or regenerate cells” are among the FDA’s examples of claims that can make a product a drug. The FDA also says the term “cosmeceutical” has no meaning under the law.

Sunscreens are regulated differently. The FDA regulates them as nonprescription drugs, and in June 2026 it [added bemotrizinol](https://www.fda.gov/drugs/understanding-over-counter-medicines/sunscreen-how-help-protect-your-skin-sun) as a permitted active ingredient. Under the [Modernization of Cosmetics Regulation Act](https://www.fda.gov/cosmetics/cosmetics-laws-regulations/modernization-cosmetics-regulation-act-2022-mocra), cosmetic companies must keep records supporting safety substantiation and report serious adverse events to the FDA within 15 business days.

The FTC polices advertising. Its [Health Products Compliance Guidance](https://www.ftc.gov/business-guidance/resources/health-products-compliance-guidance) says health-related benefits generally need “randomized, controlled human clinical testing” to meet its standard of “competent and reliable scientific evidence,” and that advertisers are liable for misleading endorsements, including expert ones. Its fake reviews rule bans buying or creating fake reviews, including AI-generated ones, and was approved by a 5-0 vote.

The practical point for GEO: everything a brand publishes to be found by AI is advertising. Content that overstates results may get repeated by an assistant and still be a violation. Mattress brands face a similar line on sleep claims, covered in [how mattress brands win AI shoppers](https://underneath.agency/resources/mattress-brands-customers-from-ai-search).

## What does a skincare brand lose when assistants leave it out?

The new shopper for that concern, and often the routine that would have followed.

In Chu and Hou’s tests, brands that did nothing while competitors improved their descriptions received no recommendations at all. Those were controlled experiments with invented brands, not live markets, but they show how quickly a small, well-documented advantage moves AI recommendations in skincare. For an indie brand without the recognition of CeraVe, the upside is the reverse: clear, comparable evidence can overcome the familiarity edge.

There is also a reputational cost. Our [study of Reddit citations](https://underneath.agency/research/ai-reddit-citations-study) found that Google’s AI Overviews cited Reddit in 17.9% of answers, and that 20.8% of sentences citing Reddit were not supported by the thread. If forum anecdotes are what assistants find about your product, they may describe it inaccurately. We have not seen a study measuring how much skincare revenue this costs, so treat it as a risk to check.

## How does GEO work for a skincare brand?

It makes accurate, compliant answers about your products easy to find and verify; it cannot guarantee recommendations.

Generative engine optimization (GEO) for skincare usually covers six pieces of work:

1. **Concern and ingredient pages.** Explain which concerns each product is designed for, the key ingredients and how to use them, in cosmetic terms your regulatory team has approved. Our article on [the product content AI shopping assistants prefer](https://underneath.agency/resources/product-content-ai-shopping-assistants-prefer) explains why specific facts work better than adjectives.
2. **Evidence summaries.** Publish what your testing actually showed: study design, number of participants, duration and results, so assistants and shoppers can see the basis for a claim.
3. **Real expert involvement, disclosed.** If dermatologists helped formulate or test a product, name them and their role. Avoid vague “dermatologist-approved” lines you cannot document.
4. **Honest review programs.** Collect detailed reviews from verified buyers without conditioning incentives on sentiment, as the FTC rule requires. Our article on [fake reviews and AI recommendations](https://underneath.agency/resources/fake-reviews-fake-brands-ai-recommendations) explains why review integrity matters to assistants too.
5. **Consistent product facts everywhere.** Keep ingredient lists, sizes, prices and claims identical across your site, retailers and marketplaces, since assistants combine them.
6. **Monitoring by concern.** Track the concern, ingredient and comparison questions for your hero products across ChatGPT, Gemini, Perplexity, Copilot and Google’s AI features, and correct errors at their source. If an assistant repeats misinformation about your product, see [how to fix wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers).

## What can’t the evidence yet tell a skincare brand?

It does not show how often AI answers change which skincare product a shopper buys.

The Haut.AI survey comes from a vendor that sells AI skin analysis and covers US women only. The AAD figures describe social media and sun care, not AI assistants. The 8th Day and Nod Mag accounts are qualitative. Chu and Hou’s experiments supplied products to the models rather than letting them search the open web, and used invented brands. L’Oréal’s growth figures show where skincare demand is, not what drives it. No platform documents how it chooses between skincare brands for a concern question.

## Where should a skincare brand start?

Start by checking what assistants say about your hero products for the concerns they are meant to address.

A useful first step is an audit of the concern, ingredient and comparison questions your customers ask, across the main assistants, showing which products are named, what claims are repeated about yours and which sources are cited. Paired with a review of the claims on your own pages, that shows where you are missing and where you may be misdescribed. If you want us to run that audit with you and plan the work around hero-product sales and routine adoption, [reach out to our team](https://underneath.agency/contact). Our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page shows how the ongoing work fits a skincare brand, from approved concern and ingredient pages to evidence summaries and honest review programs.

## Frequently asked questions

### Do people really ask AI for skincare advice?

Yes. In Haut.AI’s survey of US women, 62% were interested in or already using AI for skincare or bodycare advice. Dermatologists caution that AI cannot diagnose skin conditions from a photo.

### Can we say our product is “clinically proven” to help AI recommend it?

Only if you can substantiate it. In tests, authority language helped, but the FTC expects health-related claims to rest on competent and reliable scientific evidence.

### Does dermatologist backing help with AI recommendations?

In controlled tests, authority signals such as expert endorsements shifted recommendations. Real, documented dermatologist involvement is both more defensible and harder for competitors to copy.

### Can a small skincare brand beat a famous one in AI answers?

In controlled tests, yes, when it had a clear advantage such as better ratings, more reviews or a lower price. With nothing to separate products, assistants chose the familiar brand.

## Sources

- L’Oréal (2026), [2026 Half-Year Results](https://www.loreal-finance.com/eng/press-release/2026-half-year-results)
- Circana (2026), [US Prestige and Mass Beauty Retail Deliver a Positive Performance in 2025](https://www.circana.com/post/us-prestige-and-mass-beauty-retail-deliver-a-positive-performance-in-2025-circana-reports)
- Haut.AI (2026), [Haut.AI Launches New AI-Powered Body Analysis as Consumer Demand for Personalized Bodycare Accelerates](https://haut.ai/blogs/ai-body-analysis-launch)
- Chu and Hou (2026), [Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems](https://arxiv.org/abs/2606.17443)
- American Academy of Dermatology, via Newswise (2026), [Misinformation Puts Over 16 Million Americans at an Increased Risk for Skin Cancer](https://www.newswise.com/articles/misinformation-puts-over-16-million-americans-at-an-increased-risk-for-skin-cancer)
- US Food and Drug Administration, [Is It a Cosmetic, a Drug, or Both? (Or Is It Soap?)](https://www.fda.gov/cosmetics/cosmetics-laws-regulations/it-cosmetic-drug-or-both-or-it-soap)
- US Food and Drug Administration (2026), [Sunscreen: How to Help Protect Your Skin from the Sun](https://www.fda.gov/drugs/understanding-over-counter-medicines/sunscreen-how-help-protect-your-skin-sun)
- US Food and Drug Administration, [Modernization of Cosmetics Regulation Act of 2022 (MoCRA)](https://www.fda.gov/cosmetics/cosmetics-laws-regulations/modernization-cosmetics-regulation-act-2022-mocra)
- Federal Trade Commission (2022), [Health Products Compliance Guidance](https://www.ftc.gov/business-guidance/resources/health-products-compliance-guidance)
- Federal Trade Commission (2024), [Federal Trade Commission Announces Final Rule Banning Fake Reviews and Testimonials](https://www.ftc.gov/news-events/news/press-releases/2024/08/federal-trade-commission-announces-final-rule-banning-fake-reviews-testimonials)
- 8th Day (2026), [Getting under your skin: why consumers are letting AI rebuild their skincare](https://www.8th-day.com/thinking/getting-under-your-skin-why-consumers-are-letting-ai-rebuild-their-skincare)
- The Nod Mag (2026), [My new skincare expert is an AI chatbot. That’s probably a problem](https://thenodmag.com/content/skincare-routine-chatgpt-artificial-intelligence-diagnosis)
- OpenAI (2026), [Shopping with ChatGPT Search](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search)
- Underneath (2026), [Reddit citations study](https://underneath.agency/research/ai-reddit-citations-study)

---

This is the Markdown twin of https://underneath.agency/resources/skincare-brands-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How can sportswear brands win sales from AI shopping questions?"
description: "By being the product AI names for specific activity and attribute questions, with data clean enough to buy from your site, a retailer or the chat itself."
canonical: "https://underneath.agency/resources/sportswear-brands-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How do sportswear brands win the sale when shoppers ask AI what to train in?

By being the product an assistant names when a shopper describes an activity, a constraint and a budget, and by making that product easy to buy wherever the answer points. Sportswear questions are unusually specific: anti-chafe shorts for a half marathon, a high-support bra for running, a soccer kit for a hot summer. The brands that answer those specifics clearly, on their own pages and through retailers and reviewers, give AI something concrete to recommend.

## The short version

1. Participation is at a record: [the Sports & Fitness Industry Association](https://sfia.org/resources/participation-hits-new-high-but-majority-of-americans-not-yet-meeting-recommended-guidelines-of-150-minutes-of-weekly-activity-sfias-2026-topline-report-finds/) counted 250 million active Americans in 2025 across 126 activities, with team sports above 90 million participants for the first time.
2. Owned digital sales are diverging across the leaders: [Nike’s](https://about.nike.com/newsroom/releases/nike-inc-reports-fiscal-2026-fourth-quarter-and-full-year-results) brand digital sales fell 12% in fiscal 2026, [lululemon’s](https://www.digitalcommerce360.com/2026/09/08/ecommerce-earnings-recap-five-below-lululemon-tillys-and-more/) ecommerce fell 6% while still making up 39% of revenue, and [adidas](https://www.retailgazette.co.uk/blog/2026/07/adidas-hikes-sales-forecast-as-world-cup-demand-fuels-record-quarter/) grew ecommerce 27% in its World Cup quarter.
3. Google now documents sports-gear questions with constraints, such as “lightweight hydration vests under $80,” and says searches like “run club” and “how to train for a marathon” hit all-time highs this year ([Google](https://blog.google/products-and-platforms/products/search/running-race-training-tips/)).
4. The sale can now finish inside the chat: [JD Sports](https://fashionunited.com/news/business/jd-sports-to-allow-us-shoppers-to-purchase-through-ai-platforms/2026011269935) set up one-click purchases on AI platforms such as ChatGPT and Gemini for US shoppers.
5. A famous name helps only when products look the same: in a controlled test, product details explained 82.4% of AI rankings and brand identity 1.2% ([Chu and Hou](https://arxiv.org/abs/2606.17443)).

## Who buys sportswear, and where does the money come in today?

Active people buy by activity and season, through a mix of brand sites, retailers and marketplaces.

The buyer is whoever is about to do something: a first-time marathoner, a parent outfitting a youth team, a pilates regular, a pickleball convert. The SFIA’s 2026 report says pickleball has been the fastest-growing sport for five years running, and its members, more than 700 brands, manufacturers, retailers and governing bodies, generate $150 billion in domestic sales. Each new activity brings a short list of needs: what to wear, what to carry, [what to put on your feet](https://underneath.agency/resources/footwear-brands-ai-search).

How that demand turns into revenue varies a lot by brand. Nike earned $46.4 billion in fiscal 2026, of which wholesale was $27.5 billion and its own direct business $17.7 billion, and it spent $4.8 billion on demand creation. Adidas reported record second-quarter sales of €6.7 billion, with direct-to-consumer sales up 25% and apparel up 35%; it also spent an additional €212 million on marketing and World Cup activations. Lululemon’s ecommerce brought in about $900 million in its second quarter, and its interim co-CEO said the company has redesigned its homepage and category pages and will update its product pages next.

Two lessons follow for executives. First, sportswear brands already spend heavily to create demand, so an assistant naming the product for free at the moment of need is valuable. Second, a large share of sales runs through retailers, so the brand does not always own the page where the purchase happens. AI answers can point to either.

## Why do sports moments send shoppers to AI assistants?

Because a race, season or new sport creates many questions, and AI answers them together.

Google’s September 2026 post on race preparation shows the pattern from the platform’s side. It says running-related searches are “spiking,” with “run club,” “how to choose running shoes” and “how to train for a marathon” at all-time highs. It then suggests using AI Mode to build a training plan and to shop for gear that fits constraints, such as “road shoes for wide feet, lightweight hydration vests under $80, or anti-chafing apparel,” drawing on what it calls a Shopping Graph of over 60 billion product listings, with side-by-side comparisons and local availability.

Events work the same way at a larger scale. Adidas credits its World Cup collections and campaigns, plus its football and running categories, for a 39% rise in its performance division. Our inference is that event moments concentrate shopping questions into a few weeks, which is when being named matters most.

Shoppers are bringing more of these questions to AI in general. In [Adobe’s retail data](https://techcrunch.com/2026/04/16/ai-traffic-to-us-retailers-rose-393-in-q1-and-its-boosting-their-revenue-too/), reported by TechCrunch, AI-driven revenue per visit to US retail sites was 37% higher than non-AI traffic in March 2026, a reversal from a year earlier. That figure covers all of retail; nobody has yet published a sportswear-only number.

## Which shopping prompts carry the most purchase intent in sportswear?

Prompts that combine an activity, a performance attribute, a fit detail and a price.

We wrote the examples below to illustrate high-intent sportswear prompts; they are not observed data:

- Attribute plus activity: “Squat-proof leggings with a phone pocket for heavy lifting, under $100.”
- Problem to solve: “Running shorts that don’t chafe on long runs in humid weather.”
- Fit and support: “High-support sports bra for running that actually fits a larger cup size.”
- New sport: “Do I need pickleball shoes, or are tennis shoes fine?”
- Team and bulk: “Matching training tops for a youth soccer club of 20 players, with numbers.”
- Event: “Breathable football shirt to wear to World Cup games in summer heat.”
- Buy now: “Black running shoes under $150, good for daily training.”

The last example is the one The Next Web used to describe [JD Sports’ AI shopping flow](https://thenextweb.com/news/jd-sports-brings-ai-shopping): a shopper refines the request, compares options and pays in the same chat. Prompts like these carry more purchase intent than a broad “best activewear brands” question, because the shopper has already decided what they need and is choosing which product delivers it. For the difference between your brand being known and being chosen for a plain category question, see [our article on why well-known brands miss AI recommendations](https://underneath.agency/resources/why-well-known-brands-miss-ai-recommendations).

## How does an AI answer turn into a sportswear sale, and who books it?

The answer names products, then sends the shopper to a brand site, a retailer, or a checkout inside the assistant.

There are now three endings to the same conversation.

**On your own site.** The shopper clicks through from a product result. This is the highest-margin ending, and it depends on your product pages saying what the shopper asked about. Adobe found that around 34% of retail product pages cannot be properly accessed by AI, which is a risk for brands whose product details sit in images or scripts.

**At a retailer.** [OpenAI documents](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search) that when a shopper opens a product, ChatGPT may list several merchants, ranked “based on factors like availability, price, quality, and whether they are the maker or primary seller of that item.” A retailer with complete stock and price data can win that click even for your product. For the marketplace side of that click, see [how online marketplaces win buyers and sellers](https://underneath.agency/resources/marketplace-buyers-sellers-ai-search).

**Inside the assistant.** OpenAI’s help page says that for some eligible products and merchants, ChatGPT may show an Instant Checkout option so the shopper can pay without leaving. OpenAI scaled the feature back in March 2026, according to [FashionUnited](https://fashionunited.com/news/retail/us-consumers-would-let-ai-agents-buy-clothes-but-not-without-a-say/2026092974919), though that help page still describes it. [Google’s](https://blog.google/products/shopping/google-shopping-ai-mode-virtual-try-on-update/) agentic checkout lets a shopper track a price, then tap “buy for me” to complete the purchase on the merchant’s site through Google Pay. JD Sports, which calls North America its largest market, chose to make its catalog purchasable in AI tools; its chief executive said the aim is to let customers using AI “find and transact with JD quicker.”

For a brand, the strategic question is less “will AI send traffic?” and more “when AI recommends our product, which of these endings happens, and at what margin?” We infer that brands with clean product feeds and strong retailer data will capture more of the second and third endings even when they do not win the first.

## Does a famous logo decide which sportswear AI recommends?

Only when products look alike; clearly stated advantages decide the rest, and answers vary from run to run.

**Observed in a controlled study.** [Chu and Hou](https://arxiv.org/abs/2606.17443) gave three AI models lists of products in which one brand was real and the rest invented. When every product had identical specs, the real brand was recommended in all 670 valid trials. But when an invented brand had even the smallest advantage in rating, price or reviews, its win rate jumped from 3.6–6.0% to 64–80%. Product details explained 82.4% of the ranking, brand identity 1.2%. For a challenger activewear label, that is encouraging: a better-documented product can beat a household name. For a market leader, it is a warning not to rely on the logo. [Our article on whether AI assistants favor big brands](https://underneath.agency/resources/do-ai-assistants-favor-big-brands) covers the limits of these tests.

**Documented by the platform.** OpenAI says ChatGPT considers “structured metadata from first-party and third-party providers (e.g., price, product description) and other third-party content,” and builds review summaries from reviews on public websites. It also says a budget in the question shifts the focus to price.

**Observed in our studies.** Answers move. In [our study of repeated questions](https://underneath.agency/research/ai-recommendation-consistency-study), one ChatGPT answer showed 57.8% of the brands its five answers to the same question named between them, and only 25.2% of those brands appeared in all five. Retail questions were among the least stable, with a mean overlap of 0.458 between runs. Video also plays a part in retail answers: in [our YouTube study](https://underneath.agency/research/ai-overview-youtube-videos-study), Google’s AI Overviews cited a YouTube video on 45.0% of retail searches, even though they cited only 15.0% of the videos shown on page one.

**Our inference for sportswear.** The trust signals shoppers use are concrete: fabric weight and opacity, sweat handling, pockets, support level, inseam and sizing range, wash durability, and real wear-testing by runners, lifters or players. When those details are written plainly and repeated consistently by retailers, reviewers and creators, an assistant can match them to a question. Hype words cannot be matched to anything; see [our guide to product content AI shopping assistants prefer](https://underneath.agency/resources/product-content-ai-shopping-assistants-prefer).

## What does a sportswear brand lose when AI names a rival?

The highest-intent shoppers, at the moments when they are ready to buy several items at once.

No one has published how much sportswear revenue AI answers redirect, and we do not estimate it. The evidence supports the direction. Sports moments create bursts of specific questions. Google says it answers those with tailored recommendations and side-by-side comparisons. AI-driven retail visits were worth more per visit than other visits in Adobe’s data. And the leaders’ own digital channels are under pressure in some cases, with Nike’s brand digital sales down 12%. A brand absent from the answer for “best anti-chafe running shorts” loses a shopper who was about to buy, and possibly the rest of the kit that shopper buys alongside.

## How does generative engine optimization work for sportswear brands?

By making each product’s performance facts specific and consistent, and confirmed by independent sources across every AI surface.

- **Product pages built around activity questions.** State the activity, conditions, fabric, support level, pockets, fit, inseam and size range in plain words. Lululemon’s plan to update its product pages after its homepage and category pages shows where conversion pressure is landing.
- **Feeds and retailer listings that agree.** Keep price, stock, sizes and colors accurate in Google Merchant Center, Shopify Catalog or OpenAI’s merchant program, and make sure wholesale partners describe each product the way you do.
- **Agentic readiness.** Decide which AI checkouts you want to support directly and which you leave to retailers, and check that product details survive the trip.
- **Independent proof.** Wear tests by running, lifting and team-sport publications, honest reviews, and creator videos that say the important details out loud. Outdoor brands lean on expert reviews in the same way; see [how outdoor brands get shortlisted](https://underneath.agency/resources/outdoor-brands-ai-search).
- **Event calendars.** Publish clear guidance ahead of marathon season, the start of school sports, and major tournaments, when questions spike.
- **Measurement that respects variation.** Track your products for the same prompts across ChatGPT, Google AI Mode, AI Overviews, Gemini and Perplexity, repeated over time. [Our article on how many prompts to track](https://underneath.agency/resources/how-many-prompts-to-track-ai-visibility) explains why one check misleads.

None of this guarantees a recommendation; assistants decide their own answers. What a brand controls is whether the facts that win a comparison are available, consistent and believable.

## Which sportswear questions does the data leave open?

How much sportswear specifically moves through AI answers, and which ending, brand site, retailer or chat, wins most often.

The traffic and conversion figures here are for all US retail. We found no public sportswear-only figure for AI-referred revenue, no published share of sportswear purchases completed inside assistants, and no independent study of how assistants rank activewear. The brand results cited are company reports, and we do not attribute them to AI. The controlled study used invented brands and typed-in product data, which is cleaner than real shopping. So these figures point the way for a sportswear brand; they do not prove a return.

## Where should a sportswear brand start?

With an audit of the activity and attribute prompts tied to your best-selling product lines.

List the prompts your buyers ask before a race, a season or a tournament, then check how ChatGPT, Google’s AI features, Gemini and Perplexity answer them, several times each. Note which products are named, whether their details are right, and whether the answer sends shoppers to you, a retailer or a checkout in the chat. If you would rather hand this off, [speak with our team](https://underneath.agency/contact): we will map where your products appear for high-intent prompts, where rivals are named instead, and the work most likely to turn those answers into sales. That work, from activity-led product pages to retailer feeds that agree and event-season guidance, is described on our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page.

## Frequently asked questions

### Can AI assistants complete a sportswear purchase without a website visit?

Yes, in some cases. OpenAI’s help page describes Instant Checkout for some eligible merchants, though FashionUnited reports OpenAI scaled it back in March 2026. Google offers an agentic “buy for me” checkout, and JD Sports lets US shoppers buy through AI platforms.

### Should a sportswear brand sell directly inside AI assistants or leave it to retailers?

It depends on margins, operations and retailer relationships. Either way, your product data must be accurate wherever the assistant reads it, because it may list several merchants for the same item.

### Do big sportswear brands always win AI recommendations?

No. In a controlled test, a famous brand won when products were identical, but a small, clearly stated advantage let an unknown brand win most of the time.

### Do creator videos matter for AI answers about sportswear?

They can. Our study found Google’s AI Overviews cited a YouTube video on 45.0% of retail searches, and the text shown for cited videos came from what was said in them.

### How often should we check our AI visibility?

Regularly, and more than once per prompt. Answers change between runs and over hours, so a single check can miss half the brands an assistant names.

## Sources

- Sports & Fitness Industry Association (March 2026), [Participation hits new high, SFIA’s 2026 Topline Report finds](https://sfia.org/resources/participation-hits-new-high-but-majority-of-americans-not-yet-meeting-recommended-guidelines-of-150-minutes-of-weekly-activity-sfias-2026-topline-report-finds/)
- Nike (June 2026), [NIKE, Inc. reports fiscal 2026 fourth quarter and full year results](https://about.nike.com/newsroom/releases/nike-inc-reports-fiscal-2026-fourth-quarter-and-full-year-results)
- Retail Gazette (July 2026), [Adidas hikes sales forecast as World Cup demand fuels record quarter](https://www.retailgazette.co.uk/blog/2026/07/adidas-hikes-sales-forecast-as-world-cup-demand-fuels-record-quarter/)
- Digital Commerce 360 (September 2026), [Ecommerce earnings recap: Five Below, Lululemon, Tilly’s and more](https://www.digitalcommerce360.com/2026/09/08/ecommerce-earnings-recap-five-below-lululemon-tillys-and-more/)
- Google (September 2026), [3 ways to prep for your next big race with Search](https://blog.google/products-and-platforms/products/search/running-race-training-tips/)
- Google (May 2025), [Shop with AI Mode, use AI to buy and try clothes on yourself virtually](https://blog.google/products/shopping/google-shopping-ai-mode-virtual-try-on-update/)
- OpenAI (2026), [Shopping with ChatGPT Search](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search)
- FashionUnited (September 2026), [US consumers would let AI agents buy clothes, but not without a say](https://fashionunited.com/news/retail/us-consumers-would-let-ai-agents-buy-clothes-but-not-without-a-say/2026092974919)
- FashionUnited (January 2026), [JD Sports to allow US shoppers to purchase through AI platforms](https://fashionunited.com/news/business/jd-sports-to-allow-us-shoppers-to-purchase-through-ai-platforms/2026011269935)
- The Next Web (January 2026), [JD Sports brings AI shopping](https://thenextweb.com/news/jd-sports-brings-ai-shopping)
- TechCrunch (April 2026), [AI traffic to US retailers rose 393% in Q1, and it’s boosting their revenue too](https://techcrunch.com/2026/04/16/ai-traffic-to-us-retailers-rose-393-in-q1-and-its-boosting-their-revenue-too/)
- Chu and Hou (2026), [Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems](https://arxiv.org/abs/2606.17443)
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)
- Underneath (2026), [The YouTube videos Google’s AI cites are small](https://underneath.agency/research/ai-overview-youtube-videos-study)

---

This is the Markdown twin of https://underneath.agency/resources/sportswear-brands-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How do subscription brands win new subscribers through AI search?"
description: "Be the clear, honest answer when shoppers compare plans, prices and cancellation terms; 55% of consumers would welcome AI reordering everyday items."
canonical: "https://underneath.agency/resources/subscription-ecommerce-customers-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How do subscription brands win new subscribers through AI search?

By being the plan an AI assistant can describe accurately and recommend with confidence when a shopper compares options, ongoing prices and how easy it is to cancel. Shoppers are already open to AI help with repeat purchases, and subscription brands spend heavily to acquire each customer, so a recommendation that brings in a subscriber who stays is worth far more than one order.

This article covers subscriptions for physical products: meal kits, pet food, personal care, coffee, curated boxes and replenishment programs. For app and software subscriptions, see how consumer apps win subscribers from AI assistants.

## The short version

1. Recurring customers carry the business: at [Chewy](https://www.subscriptioninsider.com/blog/chewy-autoship-customers-drive-84-6-of-net-sales), customers with an Autoship order accounted for 84.6% of net sales, or $2.82 billion, in the quarter to August 2, 2026, and net sales per active customer reached $602.
2. Acquisition is expensive: [HelloFresh’s 2025 annual report](https://www.lobbyregister.bundestag.de/media/ff/cf/781215/Annual-Report-2025.pdf) shows €1,245.0 million in marketing expenses, 18.4% of revenue, as the company shifted toward fewer, higher-value customers.
3. Shoppers accept AI for repeat buying: in research by Koddi reported by [Supermarket News](https://www.supermarketnews.com/grocery-technology/ai-agents-gaining-the-most-trust-with-grocery-shoppers), 55% would welcome AI automatically reordering everyday items, and 64% are open to AI recommending brands they might not otherwise consider.
4. But trust has limits: in a Nuvei study in the same report, only 19% would let AI handle subscription renewals, and more than half would not let AI buy anything for them yet.
5. Subscription fatigue is real: a [NerdWallet survey](https://www.wafb.com/2026/04/07/more-than-half-americans-plan-cut-subscriptions-2026-survey-finds/) found 55% of Americans plan to significantly cut back on subscriptions in 2026, so any plan an assistant recommends faces a skeptical buyer.

## Who buys a product subscription, and what is a subscriber worth?

A household buyer choosing convenience and value, and a subscriber is worth hundreds of dollars a year if they stay.

Product subscriptions sell a routine: dinners planned, dog food that never runs out, razors that arrive before the old ones are dull. The buyer is usually one person deciding for a household, comparing price, flexibility and quality. Drink brands sold by subscription face the same comparison; see [how beverage brands get recommended](https://underneath.agency/resources/beverage-brands-ai-recommendations).

Chewy shows what a recurring relationship is worth. In its fiscal second quarter of 2026, customers with an Autoship order in the past year generated $2.82 billion, up 9.3% from a year earlier, and the share of net sales from those customers rose from 83.0% a year before. Chewy ended the quarter with 21.7 million active customers and net sales per active customer of $602. [Global Pet Industry](https://globalpetindustry.com/news/chewy-adds-208000-customers-doubles-down-on-ai-investment/) reported that Chewy added 208,000 net active customers in the quarter. Our guide to [pet product brands in AI answers](https://underneath.agency/resources/pet-product-brands-ai-search) covers that category in depth.

[Meal kits](https://underneath.agency/resources/food-ecommerce-sales-from-ai-search) show the cost side. HelloFresh, which also runs Factor, Good Chop and The Pets Table, reported for 2025:

| HelloFresh Group, 2025 | Figure |
|---|---|
| Revenue | €6,760.8 million, down 11.8% |
| Orders | 100.53 million, down 12.3% |
| Average order value | €66.8 |
| Marketing expenses | €1,245.0 million, 18.4% of revenue |

HelloFresh says that after mid-2024 it began “prioritizing high-value customer acquisition over customer volume alone.” It also notes that customers can pause or cancel at any time, and that many who cancel come back later. The commercial question for every subscription brand is the same: how many of the right customers can be won at a sustainable cost, and how long they stay. App and software subscriptions face the same question, as [how consumer apps win subscribers from AI assistants](https://underneath.agency/resources/consumer-app-subscribers-from-ai-search) shows.

## Where does AI already show up in subscription buying?

In product research, comparison and, increasingly, in the idea of letting software handle repeat orders.

Across US retail, [Adobe’s data reported by TechCrunch](https://techcrunch.com/2026/04/16/ai-traffic-to-us-retailers-rose-393-in-q1-and-its-boosting-their-revenue-too/) showed traffic from AI sources up 393% year over year in the first quarter of 2026. In March 2026 those visitors converted 42% better than other visitors. Adobe did not separate subscription sites.

Three 2026 studies summarized by Supermarket News show where shoppers draw the line:

- **Reordering is welcome.** Koddi found 55% of consumers would welcome AI automatically reordering everyday items when they run low, and 64% are open to AI recommending brands they might not otherwise consider.
- **Essentials come first.** Invoice Home found 49% of consumers most want to hand household essentials to AI agents, followed by groceries (39%) and toiletries (37%).
- **Renewals are a harder sell.** Nuvei found 19% would trust AI with subscription renewals, and more than half would not let AI purchase anything on their behalf yet.

The platforms are building the purchase step. [OpenAI documents](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search) that ChatGPT shows product results with links to merchants and, for some eligible products and merchants, an Instant Checkout option. OpenAI scaled that feature back in March 2026, according to [FashionUnited](https://fashionunited.com/news/retail/us-consumers-would-let-ai-agents-buy-clothes-but-not-without-a-say/2026092974919), though the help page still describes it for some eligible merchants. Chewy is also using AI on its own side: Global Pet Industry reported that its assistant, Kai, resolved about 30% of chats through self-service, including Autoship questions.

## What do shoppers ask AI assistants before they subscribe?

Which service fits, what it costs after the first box, how it compares and how easy cancelling is.

The prompts below are illustrative, written by us to show the buying stages. They are not captured from real users.

| Stage | Illustrative prompt |
|---|---|
| Fit | “Best meal kit for two people who are mostly vegetarian” |
| Alternatives | “Cheaper alternative to my razor subscription with the same blades” |
| Comparison | “Factor vs. HelloFresh: price per meal after the intro discount” |
| True cost | “What does this coffee subscription cost after the first-order deal, including shipping?” |
| Flexibility | “Can I skip a week or pause this dog food subscription?” |
| Exit | “How easy is it to cancel this subscription box?” |
| Legitimacy | “Is this subscription company legit? What do customers complain about?” |

Price and terms are where answers can go wrong. In [our pricing study](https://underneath.agency/research/ai-pricing-accuracy-study) of 45 software and subscription products, 61.9% of the plan prices four assistants quoted were fully faithful to the official page. Another 3.8% had the right amount but dropped a condition that changes what a buyer pays, usually by presenting an annual-billing price as the monthly price. Our inference: a first-box discount, a per-serving price and a shipping fee create the same room for confusion.

Each question may also become several searches. Google says AI Overviews and AI Mode [may use a “query fan-out” technique](https://developers.google.com/search/docs/appearance/ai-features). In [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study), ChatGPT looked for reviews in 46.2% of its answers to buyer questions and for prices in 23.8%.

## How does an AI recommendation turn into a lasting subscriber?

Through a shortlist, a price and terms check, a first order and then the months that follow.

1. **Shortlisted.** The assistant names a few services for the shopper’s need.
2. **Checked.** The shopper confirms the ongoing price, delivery area, skip and cancel rules, and reviews.
3. **First order.** The shopper subscribes, often on an introductory offer.
4. **Retained.** The subscriber skips, pauses, stays or cancels. Revenue depends on this step.

The value of step one depends on step four. A recommendation that wins a subscriber at a lower acquisition cost than paid channels, and who then behaves like Chewy’s Autoship customers, compounds. A recommendation built on a misquoted price may win a first order and lose the customer at the first full-price charge. We infer that for subscription brands, being described accurately matters as much as being named.

## What decides which subscription an assistant recommends?

Evidence about price, quality and reputation that it can find; no platform publishes how it ranks services.

What is documented: OpenAI says ChatGPT’s product results are not ads and are chosen based on the query and context. It considers structured metadata such as price and product description from first- and third-party providers, other third-party content, and review summaries drawn from public websites.

What has been observed:

- **Small advantages can beat fame.** In controlled tests by [Chu and Hou](https://arxiv.org/abs/2606.17443), an unknown brand beat a famous one half the time with a price discount of only 7.3%, and product facts such as rating, price and reviews explained 82.4% of rankings.
- **Reputation answers lean on complaint sites.** In [our brand reputation study](https://underneath.agency/research/is-it-legit-ai-reputation-study), 88.0% of answers to “Is this brand legit?” cited a review or complaint platform, and 99.7% of complete answers raised at least one problem. For subscription brands, billing and cancellation complaints are likely candidates.
- **The list moves.** In [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), only 25.2% of the brands ChatGPT named appeared in all five repeats of a question.

Our inference for subscription ecommerce: the trust factors are clear ongoing prices, plain cancellation and pause terms, a clean complaint record, independent reviews and taste tests, and consistent plan names everywhere. The legal backdrop points the same way. The FTC’s click-to-cancel rule was vacated by the Eighth Circuit on July 8, 2025, but, as [Latham & Watkins notes](https://www.lw.com/en/insights/eighth-circuit-vacates-ftc-click-to-cancel-rule), companies remain subject to the Restore Online Shoppers’ Confidence Act and state auto-renewal laws. This is a summary, not legal advice.

## What does a subscription brand lose when assistants recommend rivals?

It loses subscribers at the cheapest point of acquisition, and each lost subscriber is a stream of orders.

Direct measurement does not exist yet, so we label the reasoning:

- **Paid acquisition is costly.** HelloFresh spent 18.4% of revenue on marketing in 2025. We infer that every subscriber won through an unpaid recommendation eases that pressure.
- **Recurring value is large.** With Chewy’s net sales per active customer at $602, missing a household is a loss repeated every month, not a single sale.
- **Fatigue raises the bar.** With 55% of Americans planning to cut subscriptions, a shopper who asks an assistant for the best option may subscribe to one service only. Being absent from that answer can mean being absent from that household.

No subscription company has published how many subscribers AI answers send it, so treat any precise figure as a guess.

## How does GEO work for a subscription ecommerce brand?

Generative engine optimization (GEO) makes your subscription easy for AI assistants to find, describe accurately and support with outside evidence.

For a subscription brand, the work usually covers:

1. **A pricing page assistants cannot misread.** Introductory price, ongoing price, price per unit or serving, shipping and taxes, all in plain text, with the billing frequency stated next to each number.
2. **Terms in plain words.** How to skip, pause and cancel, stated on a page an assistant can read, and matching what customers actually experience.
3. **Comparison and alternatives pages.** Honest pages for “X vs. Y” and “alternatives to Y,” including where a rival is a better fit.
4. **Review and complaint platforms.** Accurate profiles, real reviews and complaints resolved in public on the sites assistants cite. If answers already repeat old problems, see [how to fix wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers).
5. **Independent coverage.** Taste tests, product reviews, creator reviews and roundups from publications that compare subscriptions; small brands need this most, as covered in [how small brands get recommended by AI](https://underneath.agency/resources/how-small-brands-get-recommended-by-ai).
6. **Product data.** Consistent plan and product names in feeds and catalogs, since OpenAI documents that it reads structured product data.
7. **Measurement.** Ask a fixed set of fit, comparison, price and cancellation questions across assistants many times, and check whether your prices and terms are quoted correctly. Add a “How did you hear about us?” question at signup, because [analytics miss most AI visibility](https://underneath.agency/resources/why-analytics-miss-ai-visibility).

None of this guarantees a recommendation. It makes your service the easiest one for an assistant to describe accurately and back with evidence.

## What can’t the evidence tell subscription brands yet?

It shows openness to AI in repeat buying, not how many subscribers AI answers create or keep.

- **Survey answers, not behavior.** Koddi, Invoice Home, Nuvei and NerdWallet measured what people say. The three agent studies come from companies that sell commerce or payment services.
- **Cross-retail data.** Adobe’s conversion figures cover US retail broadly, not subscription sign-ups.
- **No retention data by source.** We found no subscription company that publishes whether subscribers who arrive from AI assistants stay longer or shorter than others.
- **Price research is mostly software.** Our pricing study covered software and subscription products; physical subscriptions with first-box offers have not been tested directly.

## Where should a subscription brand start?

Start by asking the fit, comparison, price and cancellation questions your future subscribers ask, and read the answers closely.

That first check usually shows whether you are named for your core use cases, which rivals appear instead, whether your ongoing price and terms are quoted correctly, and which review sites shape what assistants say about your billing and service. From there, the work is to fix the facts, earn the evidence and make your terms as clear to an assistant as they are to a customer.

If your growth depends on winning subscribers at a sustainable cost, [talk with us about your subscription funnel](https://underneath.agency/contact). We will map where your service appears in AI answers, why competitors are named instead, and which changes are most likely to bring in subscribers who stay. Our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page sets out how those changes are made, from pricing and cancellation pages assistants can read to review profiles and repeated checks.

## Frequently asked questions

### Do shoppers trust AI to manage subscriptions for them?

Not yet, for most. In a Nuvei study, 19% would trust AI with subscription renewals, while Koddi found 55% would welcome AI reordering everyday items when they run low.

### Can AI assistants get our subscription price wrong?

Yes. In our study of 45 software and subscription products, 61.9% of quoted plan prices were fully faithful to the official page, and some answers presented annual-billing prices as monthly ones.

### Does an introductory discount help us get recommended?

Price matters to how products are ranked, but a misread discount can backfire. State the intro and ongoing price side by side so an assistant can quote both.

### Do cancellation terms affect AI recommendations?

We infer they can, through reputation. Answers about whether a brand is legitimate cite review and complaint sites 88.0% of the time in our study, and billing complaints appear there.

## Sources

- Subscription Insider (2026-09-10), [Chewy Autoship customers drive 84.6% of net sales](https://www.subscriptioninsider.com/blog/chewy-autoship-customers-drive-84-6-of-net-sales)
- Global Pet Industry, Diana Dominguez (2026-09-11), [Chewy adds 208,000 customers, doubles down on AI investment](https://globalpetindustry.com/news/chewy-adds-208000-customers-doubles-down-on-ai-investment/)
- HelloFresh SE (2026), [Annual Report 2025](https://www.lobbyregister.bundestag.de/media/ff/cf/781215/Annual-Report-2025.pdf)
- Supermarket News, Bill Wilson (2026-07-28), [AI agents gaining the most trust with grocery shoppers](https://www.supermarketnews.com/grocery-technology/ai-agents-gaining-the-most-trust-with-grocery-shoppers)
- WAFB / InvestigateTV (2026-04-07), [More than half of Americans plan to cut subscriptions in 2026, survey finds](https://www.wafb.com/2026/04/07/more-than-half-americans-plan-cut-subscriptions-2026-survey-finds/)
- TechCrunch (2026-04-16), [AI traffic to US retailers rose 393% in Q1, and it’s boosting their revenue too](https://techcrunch.com/2026/04/16/ai-traffic-to-us-retailers-rose-393-in-q1-and-its-boosting-their-revenue-too/)
- OpenAI (2026), [Shopping with ChatGPT Search](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search)
- FashionUnited (September 2026), [US consumers would let AI agents buy clothes, but not without a say](https://fashionunited.com/news/retail/us-consumers-would-let-ai-agents-buy-clothes-but-not-without-a-say/2026092974919)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- Latham & Watkins (2025-07), [Eighth Circuit vacates FTC click-to-cancel rule days before compliance deadline](https://www.lw.com/en/insights/eighth-circuit-vacates-ftc-click-to-cancel-rule)
- Chu and Hou (2026), [Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems](https://arxiv.org/abs/2606.17443), arXiv:2606.17443.
- Underneath (2026), [How faithfully do AI assistants quote software prices?](https://underneath.agency/research/ai-pricing-accuracy-study)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [“Is this brand legit?” How AI assistants build a reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study)
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)

---

This is the Markdown twin of https://underneath.agency/resources/subscription-ecommerce-customers-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How do supplement brands win new customers from AI assistants?"
description: "By being the brand assistants can verify: third-party testing, honest labels and strong reviews, now that 32% of US adults ask AI chatbots about health."
canonical: "https://underneath.agency/resources/supplement-brands-customers-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How do supplement brands win new customers from AI assistants?

By being the brand an assistant can verify when someone asks which product to buy: clear labels, third-party testing, consistent facts and a reputation that holds up on review sites. About a third of US adults now ask AI chatbots about health, and most supplement sales already happen online, so the brands named and trusted in those answers have a direct path to new customers and their monthly refills.

This article is about how supplement brands are found and chosen. It is not health advice, and nothing here says any supplement works.

## The short version

1. In [KFF’s March 2026 poll](https://www.kff.org/health-information-and-trust/poll-finding/kff-tracking-poll-on-health-information-and-trust-use-of-ai-for-health-information-and-advice/), 32% of US adults had used AI chatbots for health information or advice in the past year.
2. Supplements are now mostly an online purchase: NIQ data presented in [Nutrition Business Journal’s State of Supplements 2026](https://static.swapcard.com/public/files/7f36a9be36c149f6b653d70cf9a6c9be.pdf) found online supplement sales grew 18.7% in a year, against 12.4% for the whole market, and TikTok Shop sales grew 71.1%.
3. Discovery is shifting: in a SPINS survey in the same deck, 11% of Gen Z and millennial shoppers ranked AI chatbots among their top ways to discover products, while 46% of supplement buyers get information from a health care practitioner.
4. The market is large and growing: [Nutrition Business Journal](https://newhope.com/industry-insights/register-now-navigating-the-75-billion-supplement-market-categories-channels-trends) expects US supplement sales to reach nearly $80 billion in 2026 after 7.1% growth in 2025.
5. When ChatGPT answers health questions, it leans on institutions: in an audit of 615 cited sources, [Jacques and colleagues](https://arxiv.org/abs/2601.17109) found 75.7% came from hospitals, government, journals and similar bodies, and 12.4% from commercial health sites.

## Who buys supplements now, and what is a customer worth?

Mostly adults who take a product every day, so one new customer can mean a refill every month.

Supplements have moved from occasional purchase to routine. NBJ’s January 2026 survey found 71% of general consumers, not only current buyers, view supplements as essential or important to everyday health, and only one in five sees them as a luxury. Nutrition Business Journal puts the US market on course for nearly $80 billion this year.

Where people buy has changed even faster. In NIQ’s data for the 52 weeks to January 2026:

| US vitamins and supplements, dollar growth vs. a year earlier (NIQ) | Growth |
|---|---|
| Total market | +12.4% |
| Online | +18.7% |
| Amazon third-party sellers | +25.0% |
| Amazon first-party | +11.1% |
| TikTok Shop | +71.1% |

NIQ’s headline was blunt: “Most supplements sales happen online today.” NBJ estimates supplements on TikTok Shop alone were worth around $1 billion in 2025, and NIQ calls supplements the largest category on TikTok Shop.

No public benchmark gives a lifetime value for a supplement customer, and we will not invent one. The mechanism is clear enough: a product taken daily runs out monthly. Our inference is that a brand’s economics depend on whether a first purchase becomes a refill or subscription, and on how cheaply that first purchase was won.

## Where do AI assistants sit in a supplement buyer’s research?

Early in the research, alongside practitioners, search engines and social media, for a growing share of buyers.

Health is one of the most common reasons people use AI assistants. When it launched ChatGPT Health in January 2026, [OpenAI said](https://openai.com/index/introducing-chatgpt-health/) that over 230 million people globally ask health and wellness questions on ChatGPT every week. ChatGPT Health can connect to apps such as MyFitnessPal and Instacart, and OpenAI states that it “is not intended for diagnosis or treatment.” That is documented by the platform. Health app makers face the same shift, covered in [how digital health apps win users](https://underneath.agency/resources/digital-health-apps-users-ai-search).

KFF’s poll gives the US picture:

- 32% of adults used AI chatbots for health information in the past year, and 29% for physical health, rivaling social media as a source.
- 65% of those users cited a desire for quick, immediate information as a major reason.
- 58% of those who used AI for physical health later [followed up with a doctor or other provider](https://underneath.agency/resources/telehealth-patients-from-ai-search).

For supplements specifically, the deck NBJ presented in March 2026 shows AI arriving next to established sources. NBJ’s buyer survey added AI platforms such as ChatGPT and Perplexity as an information source, and a July 2025 SPINS survey of US Gen Z and millennial shoppers found 11% rank AI chatbots among their top discovery methods. Practitioners remain the bigger influence: 46% of supplement buyers get information from one, up significantly from 2024.

## What do supplement buyers ask AI assistants about brands?

Questions about which brand, which form, whether it is tested, whether the company is legitimate and where to buy.

The prompts below are illustrative, written by us to show the shape of commercial supplement questions. They are not captured from real users, and they deliberately stop short of health advice.

| Stage | Illustrative prompt |
|---|---|
| Brand shortlist | “Which magnesium glycinate brands are third-party tested?” |
| Format | “Creatine gummies or powder: which brands have the most reviews?” |
| Certification | “Protein powders with NSF Certified for Sport for a college athlete” |
| Comparison | “Compare these two greens powders on price per serving and ingredients” |
| Legitimacy | “Is this supplement brand legit? What do customers say?” |
| Price and value | “Cheaper alternative to my current omega-3 with the same dose” |
| Where to buy | “Is this brand sold on Amazon by the brand itself?” |

Buyers bring a checklist to these questions. NBJ’s survey of supplement buyers lists certifications they actively seek, including USP Verified and NSF Certified for Sport, and found “Made in the USA” overtook USDA Organic as the top certification in 2026. Our inference: an assistant answering “which brands are tested” needs to find that proof stated somewhere it can read.

Questions about whether a supplement treats a condition are a different matter. Those belong with clinicians, and OpenAI says its own health product is not meant for diagnosis or treatment. A brand should not try to win those answers. Wellness products face the same limit on claims, as [our guide to wellness brands](https://underneath.agency/resources/wellness-brands-customers-ai-search) shows.

## How does a mention in an AI answer become a customer?

Through a short list, a verification step and an online purchase, usually on Amazon, the brand’s site or TikTok Shop.

1. **Named for a product need.** The assistant lists a few brands for “third-party tested magnesium” or “clean protein for runners.”
2. **Verified.** The buyer checks certifications, reviews and whether the company is legitimate. In [our brand reputation study](https://underneath.agency/research/is-it-legit-ai-reputation-study), 88.0% of AI answers to “Is this brand legit?” cited a review or complaint platform.
3. **Bought online.** The purchase lands where supplements now sell: Amazon, the brand’s store or TikTok Shop. Across US retail, [Adobe’s data reported by TechCrunch](https://techcrunch.com/2026/04/16/ai-traffic-to-us-retailers-rose-393-in-q1-and-its-boosting-their-revenue-too/) showed visitors arriving from AI sources converted 42% better than other visitors in March 2026. Adobe did not break out supplements.
4. **Refilled.** If the buyer stays with the product, the revenue repeats every month, which is where the value of the first recommendation is realized.

Amazon is a gate in step three. Since April 2, 2024, according to [NSF’s summary of the policy](https://www.nsf.org/knowledge-library/amazon-new-dietary-supplements-policy-enhancing-safety-compliance), Amazon requires supplement products to be verified through a third-party testing, inspection and certification organization, and it no longer accepts certificates of analysis directly from sellers. NSF, which sells that testing, also reports consumers would pay 11% more for certified supplements.

## What decides which supplement brand an assistant names?

Verifiable evidence about the product and the company; the platforms do not publish the exact rules.

What is documented: [OpenAI says](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search) ChatGPT’s shopping results are not ads and are selected by ChatGPT. It considers structured product data such as price and description from first- and third-party providers, plus other third-party content, and it summarizes reviews from public websites.

What has been observed:

- **Health answers favor institutions.** In the Jacques audit, commercial health platforms earned 12.4% of ChatGPT’s citations. Those that were cited stated a medical review on 71.1% of pages and used schema markup on 86.8%. We cover that study in [what commercial health sites cited by ChatGPT have in common](https://underneath.agency/resources/chatgpt-cited-commercial-health-sites).
- **Reputation answers lean on review sites.** In our brand reputation study, Trustpilot and the BBB made up 61.7% of review-platform citations, and 99.7% of complete answers raised at least one problem with the brand.
- **Authority language works only until everyone uses it.** In skincare tests by [Chu and Hou](https://arxiv.org/abs/2606.17443), adding clinical-trial-style claims to an otherwise identical product beat a famous brand 73.3% of the time. When all nine challenger brands used the same language, the famous brand won again 93.8% of the time.
- **Lists change between asks.** In [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), only 25.2% of the brands ChatGPT named appeared in all five repeats of a question.

The Chu and Hou result is a warning, not a tactic. In supplements, health claims carry legal weight: the [Federal Trade Commission](https://www.ftc.gov/business-guidance/resources/health-products-compliance-guidance) says claims about the health benefits or safety of dietary supplements require “competent and reliable scientific evidence,” and the [FDA](https://www.fda.gov/food/dietary-supplements) says companies are responsible for evaluating the safety and labeling of their products before marketing. Our inference: the durable trust factors are independent testing, accurate labels, honest reviews, practitioner and editorial coverage, and a clean complaint record.

## What does a supplement brand lose when it is missing from AI answers?

It loses first purchases at the research stage, which then become a rival’s monthly refills.

We have no direct measurement of sales lost to AI absence, so we label the reasoning:

- **The buying has moved online.** With online sales growing 18.7% and TikTok Shop 71.1% in NIQ’s data, more first purchases begin on a screen where an assistant can be one tap away.
- **Refills compound.** A rival chosen after an AI answer is chosen again next month. We infer the cost of a missed recommendation is a stream of orders, not one.
- **A bad description persists.** If an assistant cannot find proof of testing or repeats an old complaint, its answer to “Is this brand legit?” carries that forward. For the repair work, read [how to fix wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers).

## How does GEO work for a supplement brand?

Generative engine optimization (GEO) makes your brand easy for AI assistants to find, describe accurately and verify through outside evidence.

For a supplement company, the work usually covers:

1. **One clear identity.** Consistent brand, product and ingredient names, forms and doses across your site, Amazon, TikTok Shop and retail listings, so assistants do not mix up products.
2. **Proof stated in text.** Certifications, testing partners, where products are made and supplement facts written on the page, not only in images or PDFs. Link to the certifier’s own listing where one exists.
3. **Compliant, useful content.** Pages that explain forms, doses on the label, who the product is designed for and how it compares, without disease claims, ideally reviewed by a qualified professional and citing sources.
4. **Review and complaint platforms.** Accurate profiles on Trustpilot, the BBB and Amazon, real customer reviews, and complaints resolved in public. Avoid any review tactic the FTC would treat as deceptive; [fake reviews also corrupt AI answers](https://underneath.agency/resources/fake-reviews-fake-brands-ai-recommendations).
5. **Independent coverage.** Editorial reviews, practitioner mentions, podcast and creator coverage with proper disclosure, and listings with certification programs.
6. **Product data and feeds.** Complete catalogs and feeds where assistants read them, since OpenAI documents that it uses structured product data.
7. **Measurement.** Ask a fixed set of brand-shortlist, comparison and legitimacy questions across ChatGPT, Gemini, Perplexity and Google’s AI features, repeatedly, and track who is named, what is said and which sources are cited.

None of this guarantees a recommendation, and no honest provider can promise one. It makes your brand the easiest one for an assistant to verify.

## What can’t the evidence tell supplement brands yet?

It shows buyers using AI for health research, not how many supplement sales AI answers create.

- **Health research is not purchase research.** KFF and OpenAI measure health questions in general. We found no public data on how often people ask assistants which supplement brand to buy.
- **Surveys and vendors.** KFF, NBJ and SPINS rely on what people report. NSF sells the testing it recommends, and Adobe sells analytics.
- **No supplement-specific ranking studies.** The tests we cite used skincare products and general health questions. Applying them to supplements is our inference.
- **Platform rules can change.** Shopping and health features are new, and assistants may treat health products differently from other goods in ways no one outside the platforms can see.

## Where should a supplement brand start?

Start by asking the brand, comparison and legitimacy questions your future customers ask, and see what assistants say.

That first check usually shows whether your products are named for the needs you serve, whether your testing and certifications are mentioned, which review sites shape your reputation and which rival brands appear instead. From there, the work is to put your proof where assistants can read it, strengthen your reputation and keep every claim compliant.

If your growth depends on first orders that become refills, [ask us for a supplement brand review](https://underneath.agency/contact). We will map where your brand appears in AI answers, why competitors are named instead, and which changes are most likely to bring more qualified buyers to your store and listings. The [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page explains how we handle the rest, including testing proof stated in text, compliant product pages and review platform profiles.

## Frequently asked questions

### Do people ask ChatGPT about supplements?

Many ask about health in general. OpenAI says over 230 million people ask health and wellness questions on ChatGPT every week, and KFF found 32% of US adults used AI chatbots for health information in the past year.

### Can a supplement brand pay to be recommended by ChatGPT?

Not in its shopping results. OpenAI says ChatGPT’s product results are selected independently and are not ads, and that ads are shown separately.

### Will strong health claims help a supplement get recommended?

They are a legal risk, and the benefit fades. In one test, clinical-trial-style claims helped while one brand used them, but the famous brand won 93.8% of the time once all challengers copied them.

### Does third-party testing matter for AI visibility?

It matters for buyers and marketplaces, and we infer it helps assistants. Amazon has required third-party verification of supplements since April 2024, and buyers in NBJ’s survey actively seek certifications.

## Sources

- KFF (2026-03), [KFF Tracking Poll on Health Information and Trust: Use of AI for Health Information and Advice](https://www.kff.org/health-information-and-trust/poll-finding/kff-tracking-poll-on-health-information-and-trust-use-of-ai-for-health-information-and-advice/)
- OpenAI (2026-01), [Introducing ChatGPT Health](https://openai.com/index/introducing-chatgpt-health/)
- Nutrition Business Journal, with NIQ and SPINS (2026-03), [The State of Supplements](https://static.swapcard.com/public/files/7f36a9be36c149f6b653d70cf9a6c9be.pdf)
- New Hope Network (2026-06-10), [Watch now! How to navigate the $75 billion supplement market](https://newhope.com/industry-insights/register-now-navigating-the-75-billion-supplement-market-categories-channels-trends)
- NSF (2024-04-22), [Amazon’s new dietary supplements policy: enhancing safety and compliance](https://www.nsf.org/knowledge-library/amazon-new-dietary-supplements-policy-enhancing-safety-compliance)
- Federal Trade Commission (2022-12), [Health Products Compliance Guidance](https://www.ftc.gov/business-guidance/resources/health-products-compliance-guidance)
- US Food and Drug Administration (2026), [Dietary Supplements](https://www.fda.gov/food/dietary-supplements)
- OpenAI (2026), [Shopping with ChatGPT Search](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search)
- TechCrunch (2026-04-16), [AI traffic to US retailers rose 393% in Q1, and it’s boosting their revenue too](https://techcrunch.com/2026/04/16/ai-traffic-to-us-retailers-rose-393-in-q1-and-its-boosting-their-revenue-too/)
- Jacques and colleagues (2026), [Authority Signals in AI Cited Health Sources: A Framework for Evaluating Source Credibility in ChatGPT Responses](https://arxiv.org/abs/2601.17109), arXiv:2601.17109.
- Chu and Hou (2026), [Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems](https://arxiv.org/abs/2606.17443), arXiv:2606.17443.
- Underneath (2026), [“Is this brand legit?” How AI assistants build a reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study)
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)

---

This is the Markdown twin of https://underneath.agency/resources/supplement-brands-customers-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How do supply chain consultancies get shortlisted by AI?"
description: "AI answers now help shape which supply chain consultancies get invited to bid, so firms need proof that independent sources repeat, not just a good website."
canonical: "https://underneath.agency/resources/supply-chain-consulting-firms-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How does a supply chain consultancy get on the shortlist when executives research with AI?

By being easy to name and easy to verify in the sources AI assistants search: rankings, trade press, client results and clear practice pages. Supply chain leaders are short of in-house expertise and are under pressure on tariffs, networks and AI, so they are looking for outside help. Much of the early research that decides who gets an RFP now runs through AI tools, and a firm that is absent there can be absent from the bid list.

## The short version

1. Demand for outside help is real: in a [Gartner survey](https://www.gartner.com/en/newsroom/press-releases/2026-04-29-gartner-survey-finds-technology-integration-and-talent-perceived-as-key-roadblocks-to-scaling-ai-in-supply-chain) of 140 senior supply chain leaders, 50% said they have limited internal expertise or talent to implement and manage AI, and 56% called integrating AI with legacy systems a major challenge.
2. Big network decisions are hard to land: 72% of supply chain leaders in [another Gartner survey](https://www.gartner.com/en/newsroom/press-releases/2026-07-14-gartner-survey-shows-72-percent-of-supply-chain-leaders-revisit-final-approvals-for-network-decisions-at-least-once-causing-delays) had to revisit final approvals for network decisions at least once, which is exactly the kind of work consultancies sell.
3. Spending is moving faster than confidence: AI now takes 67% of supply chain digital investment, yet 55% of supply chain chiefs are unclear about the returns, [Gartner found](https://www.supplychain247.com/article/gartner-supply-chain-chiefs-unclear-ai-investment-returns).
4. Talent is the top internal problem: in the [MHI and Deloitte survey](https://www.mhisolutionsmag.com/index.php/2025/06/09/digital-investments-cover-end_to_end-supply-chains/) of more than 700 supply chain leaders, 52% struggled with hiring and retaining workers and 45% with a general talent shortage.
5. The clients are large: a trade consultant [told CNN](https://abc17news.com/money/cnn-business-consumer/2024/12/23/trump-tariff-chaos-is-a-creating-a-cash-cow-for-consultants/) his importer and exporter clients do $200 million to $2 billion in annual sales, and Gartner’s supply chain surveys cover companies with at least $250 million in revenue.

## Who hires supply chain consultants and managed service providers, and what is a client worth?

Supply chain chiefs at large companies hire them for projects that often turn into retainers.

The buyer is senior and the company is sizable. Gartner’s recent supply chain surveys sampled leaders from organizations with annual revenue of $250 million or more. Source, a research firm that studies how clients buy consulting, surveyed 700 senior executives, directors and senior managers for its “Perceptions of Consulting in the US in 2024” study, and [according to KPMG’s release on it](https://www.newsfilecorp.com/release/209674/KPMG-Ranks-Number-One-for-Quality-for-AI-Advice-and-Implementation-by-Source), 90 percent of them worked in organizations generating over $500 million in revenue.

The commercial model is project work that tends to extend. Dan Gardner of Trade Facilitators, a Los Angeles consultancy in supply chain, logistics and trade compliance, told CNN that he charges either a monthly retainer or sets up projects with specific deliverables for six months to a year, and that many of them end up extending. That is the shape of value in this industry: a first project, then extensions, then often an ongoing managed service for procurement, planning or logistics.

Some firms are building the managed and platform side on purpose. [The Hackett Group](https://www.thehackettgroup.com/the-hackett-group-announces-second-quarter-2026-results/), a consulting and benchmarking firm known for procurement and operations work, reported second-quarter 2026 revenue of $69.3 million and cited recent platform-led wins exceeding $30 million as it shifts toward AI-enabled transformation services. One won relationship can be worth years of revenue, which is why a single missed shortlist is expensive.

## Why are supply chain leaders looking for outside help now?

Because tariffs, AI and network redesign have arrived together, and most teams lack the people to handle all three.

The Gartner and MHI data above describe a talent gap. The tariff wave added urgency. In the same CNN report, Joseph Esteves, CEO of Maine Pointe, a supply chain and operations consultancy that is part of SGS, said many new clients wanted to avoid repeating pandemic-era inventory mistakes, and he expected “some very great consulting years.” John Piatek, vice president of consulting at GEP, said contingency planning was “absolutely the hot topic.”

AI raises the stakes further. The MHI report found only 28% of supply chain leaders use AI today, but 82% expect to use it within five years. And Gartner’s finding that 72% of leaders revisited network approvals, more than half of them three or more times, shows how often big decisions stall without a clear business case. Each of those problems starts with a search for someone who has solved it before.

## Where do AI assistants enter the choice of a consultancy?

At the market-research and longlist stage, before any firm is contacted. No public study isolates supply chain buyers here, so we rely on broader B2B evidence.

The best evidence comes from buyers who run formal vendor selections. In a survey of 350 business buyers at organizations that issue at least 10 RFPs a year, conducted for the response-software company Responsive and [reported by Digital Commerce 360](https://www.digitalcommerce360.com/2026/01/05/ai-reshapes-b2b-buying-rfps/), 90% said they research before first contact. About one-third named generative AI chatbots, alongside web search and peer recommendations, as primary ways to find new vendors, and buyers used AI most in market research, drafting vendor questionnaires and evaluating shortlists. Two-thirds said they now use generative AI tools as much as or more than traditional search engines, and many companies require buyers to verify what AI tools tell them. That describes the early stage of a consulting RFP closely: AI helps decide which firms are worth calling, and buyers then check.

## What do buyers of supply chain consulting ask AI assistants?

They ask problem-first questions, then check named firms. The questions below are our own examples of what a buyer might type, not logged queries.

- “Which consultancies help mid-size importers redesign sourcing to reduce tariff exposure?”
- “Who does supply chain network design for food and beverage distributors in the US?”
- “Should we outsource procurement operations to a managed services provider or build in-house?”
- “Best partners to implement sales and operations planning at a $1 billion manufacturer.”
- “Alternatives to Maine Pointe for inventory reduction projects.”
- “How long does a network design study take, and how are consultants usually paid?”

Two patterns matter. First, these questions name a problem and an industry, not a firm, so the answer depends on which firms are [clearly tied to that problem in public sources](https://underneath.agency/resources/consulting-firms-clients-ai-search). Second, the comparison and pricing questions come later, when a buyer already has names. A firm can be strong in the second set and invisible in the first, which is where the longlist forms. Our [guide to how AI answers feed pipeline](https://underneath.agency/resources/ai-answers-pipeline-revenue) covers this split in more detail.

## How does AI visibility become consulting revenue?

Through the bid list. An AI answer adds a firm to the longlist, the longlist decides RFP invitations, and a won proposal opens multi-year work.

The path in this industry looks like this:

1. A supply chain leader asks an assistant who handles a problem, such as tariff-driven sourcing redesign.
2. The answer names several firms and cites a few sources. The leader checks those firms’ sites and coverage.
3. The firms that hold up are invited to an RFI or RFP.
4. Proposals and presentations decide the winner.
5. The first project leads to extensions, implementation and sometimes managed services.

AI visibility only touches steps one to three, but those steps decide who competes at all. The Responsive survey shows why that matters: 61% of buyers said they start with a preferred vendor in mind, though half said they are open to switching, and buyers ranked industry expertise as the top factor in the final decision, ahead of pricing and product fit. For a consultancy, that means the proof of sector expertise needs to be visible before the RFP, so the firm is invited, and inside the proposal, so it wins.

## What decides whether an AI assistant names a consultancy?

Platforms document little; studies point to named authorities and repeated independent evidence. We separate the two below.

What the platforms document:

- Google says there are no additional requirements to appear in AI Overviews or AI Mode: a page must be indexed and eligible to appear with a snippet. It also says both features may use “query fan-out,” issuing multiple related searches ([Google Search Central](https://developers.google.com/search/docs/appearance/ai-features)).
- OpenAI says ChatGPT search sometimes partners with other search providers and typically rewrites a user’s question into one or more targeted queries ([OpenAI Help Center](https://help.openai.com/en/articles/9237897-chatgpt-search)).

What studies observe:

- In our [study of the hidden searches assistants run](https://underneath.agency/research/ai-hidden-searches-study), ChatGPT ran a search aimed at a named publication, ranking or award in 43.8% of its answers. When a search named a source, the answer cited it 44.0% of the time, against 8.1% when it was not named. The two rise together, but that alone does not show the named source caused the citation.

What we infer for consultancies: rankings and perception studies of the kind Source publishes, analyst coverage, quotes in trade and business press, and published client results are the sources an assistant is likely to look for when asked who is good at supply chain work. A reasonable expectation is that a firm with a sharp, repeated association (“tariff sourcing for consumer goods importers”) will be named for that question more often than a firm that claims every capability. Our [guide to building authority for AI search](https://underneath.agency/resources/how-brands-build-authority-for-ai-search) explains the third-party side.

## What does it cost a consultancy to be missing from AI answers?

The cost is not being invited, and no one can measure it directly yet.

There is no public figure for consulting work lost to AI answers, and we will not invent one. The logic rests on the evidence above. Most buyers research before first contact, many enter a selection with a preferred firm in mind, and buyers lean on AI hardest when building and judging shortlists. A firm that is absent at that stage does not lose a pitch; it never pitches. Because one won client can mean a project, extensions and a managed service, the value of each missing invitation is high even if the number of such invitations is unknown.

There is a second cost: wrong facts. If AI answers describe a firm’s practice areas, industries or ownership incorrectly, buyers may rule it out before checking. Our guide on [fixing wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) covers that problem.

## How does generative engine optimization work for a supply chain consultancy?

It makes the firm’s expertise easy for AI systems to find, connect and confirm across independent sources. It does not guarantee being named.

The work, in the order we would do it:

1. **Entity clarity.** One consistent description of the firm: practice areas (network design, sourcing and tariffs, planning, procurement managed services), industries served, regions and ownership. Directories such as Consultancy.org list supply chain firms, and those listings should match the site.
2. **Practice pages that answer buyer questions.** Each service page should state the problem, typical client size, how engagements are scoped and paid (retainer, fixed project, managed service) and what clients got. Pricing questions are common; a firm can explain its model without publishing rates.
3. **Proof others can cite.** Client results with numbers, approved by the client, published as open pages rather than gated PDFs.
4. **Named-authority coverage.** Perception studies, analyst reports, awards and trade press such as DC Velocity or Supply Chain 24/7. Expert commentary on tariffs or network turbulence is how firms like GEP and Maine Pointe appeared in CNN’s reporting.
5. **People as experts.** Partner bios with credentials and bylined analysis, because buyers hire people.
6. **Measurement across assistants.** Track a fixed set of buyer questions in ChatGPT, Gemini, Google AI Overviews and AI Mode, Perplexity and Copilot, and record which firms are named and which sources are cited. Our article on [comparison pages for B2B citations](https://underneath.agency/resources/do-comparison-pages-help-b2b-ai-citations) covers one content type in this mix.

Supply chain software vendors face a related but different problem, covered in our guide on [getting supply chain software onto the RFP longlist](https://underneath.agency/resources/supply-chain-software-ai-search). Services firms sell people and outcomes, so proof of results and named experts carry more weight than feature lists.

## Which questions can’t the current evidence settle?

Several, and a consultancy should plan with them in mind.

- How many supply chain consulting RFPs begin with an AI assistant. The Responsive figures cover B2B buyers broadly.
- Whether consulting rankings or perception studies directly raise the chance of being named. Our study shows named sources are cited more often, not that rankings cause recommendations.
- How stable answers are. AI answers vary between runs and between assistants, so one check proves little.
- How much of a won engagement to credit to AI visibility. Most deals involve referrals, events and existing relationships as well.

## Where should a supply chain consultancy start?

Check what AI assistants say about your practice areas today. Then fix the gaps that keep you off bid lists.

Pick the three problems you most want to be hired for, such as tariff-driven sourcing, network redesign or managed procurement. Ask the main assistants the questions a chief supply chain officer would ask about each, record which firms are named and what sources are cited, and compare that with your own coverage and case results. If you want help, [ask us where your firm stands in AI answers](https://underneath.agency/contact) focused on the questions that lead to RFP invitations in your sectors. Our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page covers the longer work, such as practice pages, citable client results and expert coverage that ties your firm to its problems.

## Frequently asked questions

### Do supply chain executives use AI to find consultants?

No survey isolates them. Among B2B buyers who run RFPs, about one-third named AI chatbots as a primary way to find new vendors.

### Is GEO different for a consultancy than for a software vendor?

Yes. A consultancy sells expertise, so named experts, client results and third-party rankings matter more than product pages and feature comparisons.

### Should we publish our consulting rates?

Not necessarily. Explain how engagements are scoped and paid, such as retainers or six-to-twelve-month projects, so assistants can answer cost questions accurately.

### Can a boutique firm be named next to the big consultancies?

It can, for narrow questions where it is clearly tied to a problem or sector. Broad questions favor firms with more coverage.

### How long before GEO work affects our pipeline?

Expect months. Coverage and cited results take time to build, and consulting RFP cycles are long.

## Sources

- Gartner (2026-04-29), [Gartner Survey Finds Technology Integration and Talent Perceived as Key Roadblocks to Scaling AI in Supply Chain](https://www.gartner.com/en/newsroom/press-releases/2026-04-29-gartner-survey-finds-technology-integration-and-talent-perceived-as-key-roadblocks-to-scaling-ai-in-supply-chain)
- Gartner (2026-07-14), [Gartner Survey Shows 72% of Supply Chain Leaders Revisit Final Approvals for Network Decisions at Least Once, Causing Delays](https://www.gartner.com/en/newsroom/press-releases/2026-07-14-gartner-survey-shows-72-percent-of-supply-chain-leaders-revisit-final-approvals-for-network-decisions-at-least-once-causing-delays)
- Gartner, via Supply Chain 24/7 (2026-08-05), [Gartner: 55% of Supply Chain Leaders Unclear on AI Returns](https://www.supplychain247.com/article/gartner-supply-chain-chiefs-unclear-ai-investment-returns)
- MHI Solutions (2025-06-09), [Digital Investments Cover End-to-End Supply Chains](https://www.mhisolutionsmag.com/index.php/2025/06/09/digital-investments-cover-end_to_end-supply-chains/)
- CNN Business, via ABC 17 News (2024-12-23), [Trump tariff chaos is creating a cash cow for consultants](https://abc17news.com/money/cnn-business-consumer/2024/12/23/trump-tariff-chaos-is-a-creating-a-cash-cow-for-consultants/)
- KPMG, via Newsfile (2024-05-20), [KPMG Ranks Number One for Quality for AI Advice and Implementation by Source](https://www.newsfilecorp.com/release/209674/KPMG-Ranks-Number-One-for-Quality-for-AI-Advice-and-Implementation-by-Source)
- The Hackett Group (2026-08-04), [The Hackett Group Announces Second Quarter 2026 Results](https://www.thehackettgroup.com/the-hackett-group-announces-second-quarter-2026-results/)
- Digital Commerce 360 (2026-01-05), [AI reshapes B2B buying, but RFPs still decide who wins](https://www.digitalcommerce360.com/2026/01/05/ai-reshapes-b2b-buying-rfps/)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- OpenAI Help Center (2025), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)

---

This is the Markdown twin of https://underneath.agency/resources/supply-chain-consulting-firms-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Can AI search put our supply chain software on the RFP longlist?"
description: "Partly: AI answers now shape the vendor longlist before an RFP is written, so supply chain vendors must be named and verifiable there first."
canonical: "https://underneath.agency/resources/supply-chain-software-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Can AI search put our supply chain software on the RFP longlist?

It can help, because the longlist for a supply chain software deal now forms in research that increasingly runs through AI assistants, long before procurement writes the request for proposal (RFP). Supply chain buyers are spending more and putting most of their digital budget into AI, so they arrive at vendor selection with a short list already in mind. Whether your planning, visibility or optimization product is on it depends on what independent sources say about you, not only on your own site.

## The short version

1. The prize is large: [Gartner](https://www.gartner.com/en/documents/6666034) puts 2024 revenue of supply chain management software providers at about $33.4 billion, up 12.4% in a year.
2. Budgets are rising and tilting to AI: in the [2025 MHI and Deloitte report](https://www.intelligentcio.com/north-america/2025/03/20/new-mhi-and-deloitte-report-focuses-on-orchestrating-end-to-end-digital-supply-chain/), 55% of supply chain leaders were increasing technology investment and 19% planned to spend over $10 million; [Gartner](https://www.supplychain247.com/article/gartner-supply-chain-chiefs-unclear-ai-investment-returns) found AI now takes 67% of supply chain digital investment.
3. Single deals are big: [Kinaxis](https://www.kinaxis.com/en/node/8861) reported over 100 software deals above $1 million in total contract value in 2025, on annual recurring revenue of $433 million.
4. Buyers already name AI features as a top reason to buy: in [Futurum’s survey](https://futurumgroup.com/press-release/supply-chain-software-dominated-by-oracle-sap-and-blue-yonder/) of 411 decision-makers, generative AI capabilities ranked second among purchase drivers for supply chain software, at 13.2%.
5. The market is concentrated, which makes the longlist hard to join: 53.8% of the same respondents named Oracle SCM Cloud among their supply chain vendors, and SAP’s products combined reached 57.2%.

## Who signs for supply chain software, and how large is a contract?

A cross-functional team led by supply chain executives buys it, in multi-year contracts that often run to seven figures.

The buyers are senior. In the [MHI and Deloitte survey](https://www.intelligentcio.com/north-america/2025/03/20/new-mhi-and-deloitte-report-focuses-on-orchestrating-end-to-end-digital-supply-chain/) of over 700 manufacturing and supply chain leaders, 83% held executive-level positions. Their spending intentions are serious: 55% were increasing investment in supply chain technology and innovation, 60% planned to invest over $1 million, and 19% over $10 million.

The money is moving toward AI. [Gartner](https://www.supplychain247.com/article/gartner-supply-chain-chiefs-unclear-ai-investment-returns), surveying 394 supply chain professionals at companies with at least $250 million in revenue, found AI now accounts for 67% of supply chain digital investment. Yet 55% of supply chain chiefs were unclear about the returns. Buyers are spending heavily and want proof, a combination that rewards vendors whose results are easy to verify. Legal AI buyers likewise want proof before they commit, in their case of accuracy and confidentiality, as [our guide to legal software](https://underneath.agency/resources/legal-software-ai-search) shows.

Vendor earnings releases show how much a single supply chain customer can bring in. [Kinaxis](https://www.kinaxis.com/en/node/8861), a supply chain planning vendor, reported annual recurring revenue up 20% to $433 million at the end of 2025. Its finance chief said it saw “over 100 software deals above $1 million in total software contract value and more than 20 deals above $1 million in average annual software contract value.” Private competitors grow fast too: [o9 Solutions](https://o9solutions.com/news/o9-gains-a-strong-start-to-2025-with-continued-growth-in-new-customer-acquisition) reported 60% year-over-year growth in new clients in the first quarter of 2025.

The category is also crowded at the top. [Futurum’s 2025 survey](https://futurumgroup.com/press-release/supply-chain-software-dominated-by-oracle-sap-and-blue-yonder/) of 411 IT decision-makers found:

| Vendor named as a current supply chain software supplier | Share of respondents |
|---|---|
| SAP (Ariba and S/4HANA Cloud combined) | 57.2% |
| Oracle SCM Cloud | 53.8% |
| Blue Yonder | 22.1% |
| IBM Sterling Supply Chain Intelligence Suite | 20.4% |

For a specialist in planning, visibility, network design or risk, that means most buyers already own a large suite. The specialist has to earn a place on the list before anyone compares features.

## At what point do supply chain buyers meet AI answers?

At the research stage that builds the longlist, though no public study isolates supply chain buyers’ use of AI assistants.

The formal process is well documented. Selection guides, such as [ToolsGroup’s summary of Gartner’s advice](https://www.toolsgroup.com/blog/how-to-select-supply-chain-planning-software/), tell teams to “identify a longlist of vendors (up to 10) to approach with the RFI” before any demos. Gartner itself sells clients tools to [build a vendor shortlist](https://www.gartner.com/en/documents/7399530) for network design software and to [identify planning vendor candidates](https://www.gartner.com/en/documents/6321347), noting that selection “can prove time-consuming and overwhelming.” A vendor that is not on that longlist never receives the RFI.

What changed is how the longlist gets built. [Forrester’s 2026 study of business buying](https://www.forrester.com/press-newsroom/forrester-2026-the-state-of-business-buying/) calls generative AI searches “the starting point for B2B buyers.” A typical purchase involves 13 internal stakeholders and nine external influencers, and procurement staff are decision-makers in 53% of buying cycles. Each of them may run their own research. Selling to those professional buyers is its own challenge, covered in [how procurement software reaches RFP-savvy buyers](https://underneath.agency/resources/procurement-software-ai-search).

Software buyers in general now lean on AI assistants. Of 1,076 software buyers that the review platform [G2](https://company.g2.com/news/g2-research-the-answer-economy) surveyed in March 2026, 51% now open a vendor search in an AI chatbot more often than in Google. Treat that with care: G2 earns money selling visibility to vendors, and its respondents bought every kind of software, not planning or logistics tools specifically.

Even a planner who never opens a chatbot meets AI answers on Google. When we checked [1,248 US searches](https://underneath.agency/research/ai-overviews-frequency-study), an AI Overview, Google’s AI summary above the results, appeared for 96.0% of B2B software and technology keywords, more than in any of the eight industries we tracked.

## Which questions do supply chain buyers ask AI assistants?

Questions about categories, fit with existing systems, alternatives to a suite, and proof of results. We wrote the prompts in this table ourselves to show the pattern; none was recorded from a real supply chain buyer.

| Buying stage | Illustrative prompt |
|---|---|
| Category | “What are the leading supply chain planning platforms for a mid-sized consumer goods manufacturer?” |
| Fit with the suite | “Which demand planning tools work well alongside SAP S/4HANA?” |
| Alternatives | “What are the alternatives to Blue Yonder for retail replenishment?” |
| Comparison | “Kinaxis or o9 for a semiconductor company with long lead times?” |
| Risk and compliance | “Which supplier risk monitoring tools help with forced-labor and due-diligence rules?” |
| Proof | “Which supply chain visibility vendors have published customer results in automotive?” |
| AI capability | “Which planning vendors have AI agents in production, not just on the roadmap?” |

A buyer’s single question about planning tools rarely stays single. Google’s documentation says AI Overviews and AI Mode [may use a “query fan-out” technique](https://developers.google.com/search/docs/appearance/ai-features), splitting it into related searches across subtopics. OpenAI’s help pages describe [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search) turning the buyer’s wording into one or more targeted queries for its search partners.

We have watched what those hidden searches target. Across the 80 buyer questions in [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study), spanning eight industries, 43.8% of ChatGPT’s answers involved a search for a specific named publication, ranking or award. In supply chain, we infer, those named sources would include analyst evaluations, trade press rankings and peer-review sites.

## How does an AI answer become a signed supply chain contract?

Through the longlist: an AI answer reaches the team before the RFI, and the RFI leads to contracts.

Our reading of the evidence gives a five-step chain from question to contract:

1. A supply chain director, planner or IT architect asks an assistant a category, fit or alternatives question months before a budget is approved.
2. The answer names a handful of vendors and links sources such as analyst notes, trade articles and reviews.
3. The team checks those names with peers, consultants and analysts.
4. Up to 10 vendors receive an RFI; a few get scripted demos; one wins.
5. The winner signs a multi-year contract that expands across sites and modules.

The early list matters because buyers rarely stray far from it. [Gartner Digital Markets](https://cdn-static.bizzabo.com/bizzabo.file.upload/yVkIqwmWRDKbEhqeE3aQ_Gartner-Digital-Markets_2025-Report_Making-The-List.pdf), surveying 3,500 software buyers across industries, found an average of 4.4 options on a buyer’s initial informal list, and called it “the list that 81% of buyers end up making a purchase from most or all of the time.”

Proof matters more in large deals. Forrester found 78% of buyers making purchases of $10 million or more run a trial first. For a planning or visibility platform, that usually means a pilot on real data, which only shortlisted vendors are offered.

The revenue will not show up as a click. A buyer who met you in an AI answer may reach you months later through a systems integrator, a consultant or an analyst inquiry. Our guide to [what lost clicks mean for pipeline](https://underneath.agency/resources/ai-answers-pipeline-revenue) covers how to trace that slow, indirect route back to revenue. Consultancies that advise those buyers face their own longlist, covered in [how supply chain consultancies get shortlisted](https://underneath.agency/resources/supply-chain-consulting-firms-ai-search).

## Why does an assistant name one supply chain vendor and skip another?

Nobody outside the platforms knows the rules; research points to third-party sources, and supply chain buyers demand checkable results.

**What Google and OpenAI say.** Both state that their AI answers search the web and cite what they find. Neither explains how a planning or logistics vendor gets picked for a recommendation.

**What researchers have measured.** When [Chen and colleagues](https://arxiv.org/abs/2509.08919) tested US software questions, 72.7% of the sources AI search relied on were earned media, meaning independent reviews and publications; Google’s figure was 45.4%, with more weight on vendors’ own sites. Knowing a product is not the same as recommending it: in a study of 112 startups, [Sharma](https://arxiv.org/abs/2601.00912) found ChatGPT recognized products asked about by name 99.4% of the time but surfaced them in only 3.32% of category questions. That version of ChatGPT answered without web search, so treat it as a warning, not a forecast.

**What supply chain buyers weigh.** Futurum ranked the top purchase drivers for supply chain software as features and functionality (18.7%), generative AI capabilities (13.2%) and cost and pricing model (8.9%). Its research director added that buyers want software that can “interface with other disparate data and systems.” Integration proof, AI capability and total cost are the facts a buyer will check.

**Our inference.** The evidence a supply chain team uses to verify a vendor is mostly public: analyst evaluations, customer case studies with numbers, integration partner listings, conference talks and trade coverage. An assistant can reach and quote every one of those supply chain pages too. So a planning or visibility vendor with public, specific and consistent proof probably serves both the human evaluator and the assistant better. That remains our expectation, not a tested finding for supply chain software.

## What does a supply chain vendor lose by being absent from AI answers?

Mostly lost evaluations rather than lost clicks, though no study has measured the cost directly.

- **The suites start ahead.** With Oracle and SAP named by more than half of buyers, a specialist missing from AI answers leaves the buyer’s default, the suite they already own, unchallenged. That is our inference from the survey data.
- **Missed cycles are long cycles.** Contracts run for years and renew. A vendor left off one longlist may wait until the next replacement cycle, we infer.
- **AI claims are now table stakes.** Gartner forecasts supply chain software with agentic AI capabilities growing from less than $2 billion in 2025 to [$53 billion in spend by 2030](https://www.gartner.com/en/newsroom/press-releases/2026-04-07-gartner-forecasts-supply-chain-management-software-with-agentic-ai-will-grow-to-53-billion-in-spend-by-2030). If AI answers describe your product as lacking AI features you do have, you lose on the second-ranked purchase driver.
- **Being named is hit or miss.** Across five runs of the same question in [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), just 25.2% of the brands ChatGPT named showed up every time. A planning vendor that appears on only some runs drops out of some RFIs without ever learning they existed.

## What does GEO involve for a supply chain vendor?

It puts your supply chain results where assistants can find, read and repeat them, with no guaranteed recommendation.

Generative engine optimization (GEO) means working to be described accurately when AI assistants answer questions about supply chain software. For a planning, visibility or logistics vendor it comes down to six pieces of work:

1. **One clear identity.** Say what you plan, track or optimize, for which industries, and which systems you sit beside, in the same words on your site, partner pages, analyst briefings and company databases. Where assistants already describe you wrongly, our guide on [correcting brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) explains the repair.
2. **Numbers buyers can check.** Publish customer results with named metrics, such as forecast accuracy, inventory reduction or service level, on plain web pages rather than gated PDFs.
3. **Independent coverage.** Brief analysts, speak at industry events and earn trade press coverage. Assistants often look up exactly these outlets by name, and our piece on [building authority for AI search](https://underneath.agency/resources/how-brands-build-authority-for-ai-search) covers how vendors earn them.
4. **Integration and partner pages.** Document how you connect to SAP, Oracle and the major warehouse and transport systems. Buyers ask fit questions first, and systems integrators often shape the longlist.
5. **Honest comparison content.** Explain when a specialist beats the suite and when it does not. Third-party ranked lists still matter more, as we show in [which pages to target](https://underneath.agency/resources/best-of-lists-ai-recommendations).
6. **Repeated checks on every assistant.** Put each longlist question to ChatGPT, Google AI Overviews and AI Mode, Gemini, Perplexity, Copilot and Claude more than once, since a single run misleads. B2B software answers vary less than most: in [our four-assistant study](https://underneath.agency/research/ai-assistants-brand-agreement-study), B2B software had the highest overlap between assistants, 0.543 on a scale from 0 to 1, but that is still far from full agreement.

How the same forces play out across software in general is covered in [how B2B SaaS companies earn revenue from AI search](https://underneath.agency/resources/b2b-saas-revenue-from-ai-search).

## What is still unknown about AI search in supply chain software buying?

The biggest gap: nobody has measured whether appearing in AI answers lifts a supply chain vendor’s sales.

- **No study of supply chain buyers’ AI use.** The surveys on AI in buying cover software buyers in general. None isolates supply chain or operations leaders.
- **Interested sources.** G2 sells review visibility; Kinaxis and o9 report their own results; Futurum’s top-line findings come from a press release.
- **Nobody has tied it to closed deals.** [Martinez](https://arxiv.org/abs/2607.14035) went through 45 studies of AI search optimization, and found the work on traffic and conversions the least supported of the lot. Our article [does AI visibility drive business results](https://underneath.agency/resources/does-ai-visibility-drive-business-results) weighs what is known.
- **Vendor choice is a black box.** Only the platforms know how an assistant settles on which planning or logistics vendors to name, and the list shifts from run to run.

## How can a supply chain vendor tell whether it makes the AI-built longlist?

Put your buyers’ pre-RFI questions to assistants repeatedly, and note which vendors are named and from which sources.

Write down what a planning, procurement or logistics leader would ask before sending an RFI: category, fit with the suite they own, alternatives, compliance and proof. Ask each of them in ChatGPT, Gemini, Perplexity and Google’s AI results on several different days, since one supply chain answer can differ from the next. Note who is named, which sources are cited, and how your AI features, integrations and customer results come across. Where you are missing, the cause is usually too little analyst, trade or customer proof published by others about you.

We can [review your place on AI-built RFI longlists](https://underneath.agency/contact) with you: which answers include you, which leave you off, and which gaps in your public proof are most likely costing you evaluations and seven-figure deals. For how that proof gets built, from checkable customer results to integration pages and analyst briefings, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## Frequently asked questions

### Do supply chain executives use ChatGPT to find software vendors?

No survey isolates them. Across all software, 51% of buyers in G2’s 2026 poll began research with AI chatbots more often than Google.

### Will a Gartner Magic Quadrant spot get a supply chain vendor named in AI answers?

Plausibly, but untested. ChatGPT searched for a named publication, ranking or award in 43.8% of answers in our study, and analyst reports are such sources.

### Can a specialist vendor compete with SAP and Oracle in AI answers?

It can be named for specific questions, but recognition is not recommendation: one study found ChatGPT surfaced startups in only 3.32% of category questions.

### Should we ungate our customer case studies?

For the results buyers verify, generally yes. AI search drew 72.7% of sources for US software questions from independent pages, so publish proof others can cite.

### How long until GEO work shows up in supply chain deals?

Expect quarters, not weeks. Large supply chain deals pass through RFIs, demos and pilots, and 78% of $10 million-plus buyers run a trial first.

## Sources

- Gartner (2025-06-30), [Market Share Analysis: Supply Chain Management Software, Worldwide, 2024](https://www.gartner.com/en/documents/6666034)
- Gartner, via Supply Chain 24/7 (2026-08-05), [Gartner: 55% of Supply Chain Leaders Unclear on AI Returns](https://www.supplychain247.com/article/gartner-supply-chain-chiefs-unclear-ai-investment-returns)
- Gartner (2026-04-07), [Gartner Forecasts Supply Chain Management Software with Agentic AI Will Grow to $53 Billion in Spend by 2030](https://www.gartner.com/en/newsroom/press-releases/2026-04-07-gartner-forecasts-supply-chain-management-software-with-agentic-ai-will-grow-to-53-billion-in-spend-by-2030)
- Gartner (2026-02-16), [Tool: Supply Chain Network Design Vendor Shortlist Builder](https://www.gartner.com/en/documents/7399530)
- Gartner (2025-04-03), [Tool: Identify Supply Chain Planning Solution Vendor Candidates](https://www.gartner.com/en/documents/6321347)
- MHI (2025-06-09), [Digital Investments Cover End-to-End Supply Chains](https://www.mhisolutionsmag.com/index.php/2025/06/09/digital-investments-cover-end_to_end-supply-chains/)
- MHI and Deloitte, via Intelligent CIO (2025-03-20), [New MHI and Deloitte report focuses on orchestrating end-to-end digital supply chain solutions](https://www.intelligentcio.com/north-america/2025/03/20/new-mhi-and-deloitte-report-focuses-on-orchestrating-end-to-end-digital-supply-chain/)
- Futurum Group (2026-01-23), [Supply Chain Software Dominated by Oracle, SAP, and Blue Yonder](https://futurumgroup.com/press-release/supply-chain-software-dominated-by-oracle-sap-and-blue-yonder/)
- Kinaxis (2026-03-04), [Kinaxis Inc. Reports Record Fourth Quarter 2025 Results](https://www.kinaxis.com/en/node/8861)
- o9 Solutions (2025-04-09), [o9 Gains a Strong Start to 2025 with Continued Growth in New Customer Acquisition](https://o9solutions.com/news/o9-gains-a-strong-start-to-2025-with-continued-growth-in-new-customer-acquisition)
- ToolsGroup (2021-06-02), [How to select supply chain planning software](https://www.toolsgroup.com/blog/how-to-select-supply-chain-planning-software/)
- Gartner Digital Markets (2025), [Making the List](https://cdn-static.bizzabo.com/bizzabo.file.upload/yVkIqwmWRDKbEhqeE3aQ_Gartner-Digital-Markets_2025-Report_Making-The-List.pdf)
- Forrester (2026-01-21), [Forrester’s 2026 Buyer Insights: GenAI Is Upending B2B Buying](https://www.forrester.com/press-newsroom/forrester-2026-the-state-of-business-buying/)
- G2, Tim Sanders (2026-04-15), [In the Answer Economy, Don’t Win the Click — Win the Answer](https://company.g2.com/news/g2-research-the-answer-economy)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- OpenAI Help Center (2025), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Chen, Wang, Chen and Koudas (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919), arXiv:2509.08919.
- Sharma (2025), [The Discovery Gap: How Product Hunt Startups Vanish in LLM Organic Discovery Queries](https://arxiv.org/abs/2601.00912), arXiv:2601.00912.
- Martinez (2026), [Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026)](https://arxiv.org/abs/2607.14035), arXiv:2607.14035.
- Underneath (2026), [When does Google show an AI Overview? 1,248 US searches](https://underneath.agency/research/ai-overviews-frequency-study)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude agree on brands?](https://underneath.agency/research/ai-assistants-brand-agreement-study)
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)

---

This is the Markdown twin of https://underneath.agency/resources/supply-chain-software-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How tech consultancies reach enterprise shortlists via AI search"
description: "By being the firm AI answers tie to proven production results in a named platform and industry, backed by analyst, partner and client evidence."
canonical: "https://underneath.agency/resources/technology-consulting-firms-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How does a technology consultancy get onto enterprise shortlists when buyers ask AI first?

By being the firm that AI answers connect with a specific platform, industry and production result, and by making sure the analyst reports, partner directories and client stories those answers draw on say so. Enterprise programs in cloud, data and AI are still won in competitive bids, but the long list that decides who gets invited is increasingly drafted with an assistant. No study has yet counted how many consulting contracts start that way.

This guide is for digital transformation, cloud, data and AI consultancies and systems integrators that sell multi-month enterprise programs. Advisory IT consulting for mid-size companies, and hands-on IT projects and support, follow different paths and are covered separately.

## The short version

1. The market is large and still growing: [Gartner](https://www.gartner.com/en/newsroom/press-releases/2025-10-22-gartner-forecasts-worldwide-it-spending-to-grow-9-point-8-percent-in-2026-exceeding-6-trillion-dollars-for-the-first-time) forecasts IT services spending of $1,869,269 million in 2026, up 8.7%.
2. AI is now the first thing buyers screen for: in [IDC’s survey](https://www.idc.com/resource-center/blog/how-service-provider-selection-criteria-have-shifted-what-we-learned-and-what-comes-next/) of 700 European organizations, AI capabilities rose from sixth to first among the criteria for choosing a services partner between 2023 and August 2025.
3. The money behind that shift is visible: [Accenture](https://newsroom.accenture.com/content/4q-full-fy25-earnings/accenture-reports-fourth-quarter-and-full-year-fiscal-2025-results.pdf) reported $37.64 billion of consulting new bookings in fiscal 2025, including $5.9 billion of generative AI new bookings.
4. Clients are not satisfied with what they get: in an [ISG study](https://www.silicon.co.uk/press-release/enterprises-seek-more-innovation-ai-value-from-service-providers-isg-study-shows), nearly 65 percent of enterprises were dissatisfied or only moderately satisfied with how their IT outsourcing providers innovate. Dissatisfied clients look for alternatives.
5. Assistants look for third-party proof: in [our brand entity study](https://underneath.agency/research/brand-entity-ai-recommendations-study), each tenfold increase in independent sites naming a brand went with 4.7 times the odds of being recommended.

## Who buys enterprise technology consulting, and what is one client worth?

A buying group led by the CIO, CTO or chief data officer, with procurement and often a sourcing advisor alongside.

The program decides the buyer. A cloud migration usually belongs to the CIO and the head of infrastructure; a data platform to the chief data officer; an AI program to a transformation office or a business unit leader with a CIO sponsor. Around each of them sits a large group. [Forrester’s 2026 buying study](https://www.forrester.com/press-newsroom/forrester-2026-the-state-of-business-buying/) counts 13 internal stakeholders and nine external influencers in a typical purchase, and finds procurement among the decision-makers in 53% of buying cycles.

External influencers matter more here than in most industries. Sourcing advisors run many large selections. ISG, one of them, says it works with more than 900 clients, including 75 of the world’s top 100 enterprises. Analyst firms, hyperscaler partner teams and software vendors also steer work toward the integrators they trust. Strategy firms hired by the CEO are covered in [our guide to management consulting firms](https://underneath.agency/resources/management-consulting-firms-clients-ai-search).

A client is worth the program and everything after it. The first statement of work is often a discovery phase or a pilot; the value comes from the build, the rollout, the next platform and, in many cases, a managed services contract. Accenture’s fiscal 2025 results show the scale at the top of the market: consulting new bookings of $37.64 billion in one year. Tier-2 integrators compete for the same programs with a sharper focus. [Technology Business Research](https://tbri.com/?p=17321) describes these firms as typically focused on “a country or region, a select few industries or a technology niche,” at sizes from 3,000 to 80,000 employees.

## Where do AI assistants sit in an enterprise consulting selection?

At the framing and long-list stages, before a request for proposal names who gets to bid.

A typical program selection runs in six steps:

1. **Trigger.** An AI mandate from the board, a data center exit, an ERP end-of-support date, or a pilot that stalled.
2. **Framing.** What should we build, in what order, and what have peers done?
3. **Long list.** Analyst rankings, hyperscaler partner directories, peers, sourcing advisors and, increasingly, AI assistants.
4. **Request for proposal.** Procurement and often an advisor run a formal bid.
5. **Proof.** Orals, workshops and paid pilots. For deals of $10 million or more, Forrester found 78% of buyers insist on a trial before they sign.
6. **Program.** A first statement of work, then follow-on phases.

Assistants already appear at steps 2 and 3 across business buying. In [a Gartner survey](https://www.gartner.com/en/newsroom/press-releases/2026-05-20-gartner-survey-finds-sixty-nine-percent-of-b-two-b-buyers-turn-to-sales-reps-to-validate-ai-generated-insights) of 645 B2B buyers in all sectors, 45% said they used generative AI during a recent purchase, “primarily to gather information on vendors and products.” The same survey found 69% prefer to check those insights with a sales rep, which in consulting means a partner or principal.

Services directories are moving into the assistants too. Clutch, a B2B services marketplace, [launched an app inside ChatGPT](https://www.demandgenreport.com/?p=52857) in 2026 so buyers can compare providers, reviews and “pricing signals” without leaving the conversation.

IDC’s survey explains why buyers start with questions rather than names. Buyers had stopped asking only “Can you build this?” and were asking “What should we build? What’s the market doing?” Those are exactly the questions a CIO now tries first on an assistant.

## What do enterprise buyers ask AI about consulting partners?

Questions that combine a platform, an industry, a scale and a risk. We wrote these examples to show typical phrasing; none was collected from a real buyer.

| Buyer | Example question |
|---|---|
| CIO, insurer | “Which consultancies have moved a core insurance platform to AWS without a long outage?” |
| Chief data officer | “Best Databricks implementation partners for a global retailer, and how do they differ?” |
| CFO sponsoring ERP | “SAP S/4HANA migration partners for a mid-size manufacturer in Germany?” |
| Transformation office | “Who has taken generative AI from pilot to production in banking, with named clients?” |
| Procurement | “Alternatives to the big four for a nine-month data program, with lower day rates?” |
| CTO | “Which systems integrators are strongest in ServiceNow for public sector?” |

Three features stand out. The questions name a platform, because buyers expect partner depth on Microsoft, AWS, Google Cloud, SAP, Salesforce, Snowflake or ServiceNow. They name an industry, because TBR found buyers now rank “specific line of business or domain” among the top three selection attributes. And they ask for proof of production, because pilots have disappointed: the MIT NANDA research reported by [Fortune](https://fortune.com/2025/08/18/mit-report-95-percent-generative-ai-pilots-at-companies-failing-cfo/) found generative AI implementation “falling short” for 95% of companies in its dataset.

## How does being named in an AI answer turn into a program?

Through an invitation to bid. An assistant’s answer can put a firm on the long list; the contract is still won in the room.

The path:

1. **AI answer.** A buyer or advisor sees three to eight firms named for a platform and industry.
2. **Long list.** The team checks the firms’ case studies, analyst positions and partner status.
3. **Request for proposal.** Five or six firms are invited, usually with a sourcing advisor.
4. **Workshop or pilot.** A paid discovery phase or proof of concept.
5. **Statement of work.** The first phase of the program.
6. **Expansion.** Build, rollout, run, and the next program.

Nothing is guaranteed after step 1. A CIO interviewed by TBR described the reality: every similar project is rebid, and “You can’t be 20%, 30% more expensive than the other.” What AI visibility changes is who gets to compete. A firm that is not on the long list never sees the bid, and in consulting the lost value is the whole multi-year relationship, not a single sale.

The partnership finding in the MIT research is the opening for consultancies. According to Fortune, buying from specialized vendors and building partnerships “succeed about 67% of the time, while internal builds succeed only one-third as often.” Enterprises that have failed alone are looking for help, and many begin by asking who has done it before. We explain why those early answers matter more than the clicks they replace in [what lost clicks mean for pipeline and revenue](https://underneath.agency/resources/ai-answers-pipeline-revenue).

## What decides whether an assistant names a consultancy?

No platform publishes how it chooses consulting firms; studies point to independent coverage, rankings and the searches assistants run.

**Documented by the platforms.** Assistants with web search retrieve pages and cite them; none documents a rule for choosing consultancies.

**Observed in our studies, across industries.** None of these studies asked about consulting firms specifically, so read them as a guide:

- **Independent coverage counts most.** In our brand entity study, the number of independent sites naming a brand in the cited pages was the strongest predictor we measured. A Wikipedia article added little once prominence was accounted for.
- **Rankings and awards get looked up.** In [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study), ChatGPT ran a search aimed at a named publication, ranking or award in 43.8% of its answers. For consulting, the obvious targets are analyst evaluations, partner awards and industry rankings.
- **Self-published “best of” lists are a weak lever.** In [our study of self-ranking lists](https://underneath.agency/research/self-promoting-best-lists-study), 24.2% of AI-cited numbered “best X” lists with an identifiable publisher ranked their own publisher first. Assistants do cite them, but buyers and advisors discount them.
- **Google rank is only part of it.** In [our study of AI citations and Google rankings](https://underneath.agency/research/ai-citations-google-rankings-study), 8.3% of the pages ChatGPT cited ranked in Google’s top 10 for the question, against 25.7% for Claude.

**Our inference, specific to consulting.** A reasonable expectation is that assistants lean on the same sources enterprise buyers already trust: analyst evaluations, hyperscaler and software partner directories, named client stories, and trade press. The IDC and TBR findings point to the content buyers want, namely industry knowledge plus a current view on AI. Firms that publish that view under named leaders give both buyers and assistants something to cite. Our guide to [how brands build authority for AI search](https://underneath.agency/resources/how-brands-build-authority-for-ai-search) covers the general mechanics.

## What does it cost a consultancy to be missing?

The bids it never hears about, and the follow-on work behind them. No study has put a figure on that for consultancies.

- **Incumbents are exposed.** In the ISG study, more than two-thirds of enterprises were dissatisfied or only moderately satisfied with the cost of AI services, and about 47 percent of those using outsourcing services were not fully satisfied with providers’ ability to adapt to changing demand. Unhappy clients research alternatives, and an assistant is a quiet place to start.
- **Relationships no longer protect the account.** TBR found an “existing vendor relationship remained the least critical attribute for vendor selection.”
- **AI demand is concentrated in new categories.** IDC’s forecast has enterprises spending $400B in 2026 on AI platforms and services combined. Firms not associated with AI work in answers risk being left out of the fastest-growing part of the market, we infer.
- **Wrong facts travel.** An answer that describes a firm by its old specialty, or omits a platform it now leads, sends the wrong buyers and loses the right ones. Our guide to [fixing wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) covers the corrections.

## How does GEO work for a technology consultancy?

It makes your platforms, industries and production results easy for assistants and buyers to find and verify. No one can promise a mention.

1. **Define the entity precisely.** Practices, platforms, partner tiers, industries, regions and leadership, stated the same way on your site, partner directories, LinkedIn and analyst profiles.
2. **Publish production proof, not pilot proof.** Named client stories with the platform, industry, scale and measurable result, cleared with the client. Buyers ask for this because so many pilots stalled.
3. **Keep partner listings current.** Hyperscaler and software partner directories are where buyers and advisors check depth. Specializations, certifications and industry tags should match what you sell now.
4. **Earn independent coverage.** Analyst evaluations, sourcing advisor briefings, trade press, conference talks and client-side speakers. Our entity study suggests independent mentions carry more weight than your own pages.
5. **Answer the framing questions.** IDC’s buyers want a point of view on what to build next. Publish clear, dated answers to the questions in the table above, by industry and platform, under named partners.
6. **Explain commercial models.** Fixed-fee discovery, outcome-based pricing, team shapes and typical durations. Procurement asks, and so do assistants.
7. **Test the long-list questions.** Ask ChatGPT, Gemini, Perplexity, Claude, Copilot, Google AI Overviews and AI Mode the questions your buyers ask, several times each, and record who is named and which pages are cited. One answer is a sample, not a verdict. If you plan comparison content, our review of [whether comparison pages help B2B citations](https://underneath.agency/resources/do-comparison-pages-help-b2b-ai-citations) sets out what it can do.

The same buying groups pick software too; [how AI assistants shape enterprise software shortlists](https://underneath.agency/resources/enterprise-software-shortlists-ai-search) covers that side. For one large example, see [how ERP vendors reach the shortlist](https://underneath.agency/resources/erp-software-ai-search).

## What can’t a consultancy measure yet?

How many enterprise programs began with an AI answer. No public study tracks a consulting selection from first question to signed statement of work.

- **Buyer surveys are cross-industry.** Gartner’s 45% covers B2B buyers of all kinds, not consulting buyers.
- **IDC’s selection data is European.** The 700 organizations were surveyed in Europe; American buyers may rank criteria differently.
- **Our studies did not ask about consultancies.** Their patterns come from other categories.
- **Attribution is hard.** Many buyers never click through, and a request for proposal rarely records where the long list came from. Our article on [why analytics miss AI visibility](https://underneath.agency/resources/why-analytics-miss-ai-visibility) explains the gap.

## Where should a technology consultancy start?

With the long-list questions your last ten winning clients would have asked before they invited you to bid.

Write them down by platform and industry. Ask each major assistant, several times, and compare the firms it names with the firms you met in those bids. Note which analyst reports, partner listings, client stories and articles are cited, and whether your platforms and industries are described correctly.

If you want help reading those answers, [talk to us about a review of your enterprise long-list questions](https://underneath.agency/contact). We will show where your firm is named and where competitors are, which sources the answers rely on, and which gaps in proof and coverage are most likely to cost you invitations to bid. The [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page explains how we close them, from production client stories and current partner listings to analyst and advisor coverage.

## Frequently asked questions

### Do enterprise buyers use AI to find consulting firms?

Across B2B purchases, 45% of 645 buyers in Gartner’s survey used generative AI, mainly to gather information on vendors. No survey isolates consulting buyers.

### Do analyst reports still matter?

Yes, for buyers and, we infer, for assistants. ChatGPT searched for a named publication, ranking or award in 43.8% of answers in our hidden-searches study.

### Can a tier-2 integrator compete with the largest firms in AI answers?

Plausibly, on specific platforms and industries. Buyers rank domain knowledge highly, and assistants answer narrow questions with narrower lists, we infer.

### Should we publish our own “best consultancies” list?

It is unlikely to help much. In our study, 24.2% of AI-cited “best” lists ranked their own publisher first, and readers know it.

### How is this different from IT consulting for mid-size firms?

Enterprise programs run through sourcing advisors, procurement and formal bids. Mid-size advisory work is bought faster, often by one executive.

## Sources

- Gartner (2025-10-22), [Gartner forecasts worldwide IT spending to grow 9.8% in 2026, exceeding $6 trillion for the first time](https://www.gartner.com/en/newsroom/press-releases/2025-10-22-gartner-forecasts-worldwide-it-spending-to-grow-9-point-8-percent-in-2026-exceeding-6-trillion-dollars-for-the-first-time)
- IDC (2026-08-27), [How service provider selection criteria have shifted: what we learned and what comes next](https://www.idc.com/resource-center/blog/how-service-provider-selection-criteria-have-shifted-what-we-learned-and-what-comes-next/)
- IDC (2026-04), [IDC Directions 2026: The AI supercycle](https://www.idc.com/wp-content/uploads/2026/04/IDC-Directions-AI-Supercycle-Whalen.pdf)
- Accenture (2025-09-25), [Accenture reports fourth-quarter and full-year fiscal 2025 results](https://newsroom.accenture.com/content/4q-full-fy25-earnings/accenture-reports-fourth-quarter-and-full-year-fiscal-2025-results.pdf)
- ISG, via Silicon UK (2025), [Enterprises seek more innovation, AI value from service providers, ISG study shows](https://www.silicon.co.uk/press-release/enterprises-seek-more-innovation-ai-value-from-service-providers-isg-study-shows)
- Technology Business Research (2025-07), [TBR launches Enterprise Systems Integrators Market Landscape](https://tbri.com/?p=17321)
- Forrester (2026-01-21), [Forrester’s 2026 buyer insights: GenAI is upending B2B buying](https://www.forrester.com/press-newsroom/forrester-2026-the-state-of-business-buying/)
- Gartner (2026-05-20), [Gartner survey finds 69% of B2B buyers turn to sales reps to validate AI-generated insights](https://www.gartner.com/en/newsroom/press-releases/2026-05-20-gartner-survey-finds-sixty-nine-percent-of-b-two-b-buyers-turn-to-sales-reps-to-validate-ai-generated-insights)
- Fortune (2025-08-18), [MIT report: 95% of generative AI pilots at companies are failing](https://fortune.com/2025/08/18/mit-report-95-percent-generative-ai-pilots-at-companies-failing-cfo/)
- Demand Gen Report (2026-05-12), [Clutch launches first B2B services marketplace app on ChatGPT](https://www.demandgenreport.com/?p=52857)
- Underneath (2026), [Do Wikipedia and schema make AI assistants recommend a brand?](https://underneath.agency/research/brand-entity-ai-recommendations-study)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [How many “best of” lists cited by AI rank their own brand first?](https://underneath.agency/research/self-promoting-best-lists-study)
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude cite pages that rank?](https://underneath.agency/research/ai-citations-google-rankings-study)

---

This is the Markdown twin of https://underneath.agency/resources/technology-consulting-firms-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How do telehealth brands get chosen when patients ask AI?"
description: "By being the service assistants can verify on price, insurance, licensing and reputation, as 58% of AI health users follow up with a provider."
canonical: "https://underneath.agency/resources/telehealth-patients-from-ai-search"
published: 2026-10-09
updated: 2026-10-10
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How do telehealth companies get chosen when patients ask AI where to get care?

By being the virtual care service an assistant can verify when someone asks where to be seen: clear prices, the insurance you accept, the states you serve, licensed clinicians and a reputation that holds up on review sites. A third of US adults now ask AI chatbots about their health, and most who ask about physical health go on to contact a provider. The telehealth brands that are named, described accurately and trusted at that moment can win the visit; the rest pay more to be found.

This article is about how consumer telehealth companies are found and chosen. It is not medical advice, and nothing here recommends a diagnosis, treatment or provider. For software sold to hospitals and health systems, see our guide to healthcare software and AI search.

## The short version

1. In [KFF’s March 2026 poll](https://www.kff.org/health-information-and-trust/poll-finding/kff-tracking-poll-on-health-information-and-trust-use-of-ai-for-health-information-and-advice/), 32% of US adults had used AI for health information or advice in the past year, and 58% of those who asked about physical health later followed up with a doctor or other provider.
2. Access is part of the reason: 18% of those AI users said a major reason was that they had no regular provider or could not get an appointment, and 19% that they could not afford to see one.
3. Acquisition is expensive: [Hims & Hers](https://www.subscriptioninsider.com/blog/hims-hers-raises-2026-revenue-outlook-as-subscribers-near-2-9-million) spent $262.2 million on marketing in the second quarter of 2026 on revenue of $753.2 million, while [Teladoc’s BetterHelp](https://www.santelog.com/actualites-sante-nasdaq/teladoc-health-reports-second-quarter-2026-results) averaged 346,000 paying users, down 11% from a year earlier.
4. Paid channels are gated: [Google Ads](https://support.google.com/adspolicy/answer/176031?hl=en) allows telemedicine providers to advertise only if they are verified by LegitScript’s Healthcare Merchant Certification Program.
5. The assistant can itself be a competitor: [Amazon’s Health AI](https://www.fiercehealthcare.com/ai-and-machine-learning/amazon-launches-health-ai-assistant-its-website-expands-free-virtual-care) connects users to One Medical clinicians, with five free consultations for more than 30 common conditions for eligible Prime members and $29 pay-per-visit care after that.

## Who chooses a telehealth service, and what is a patient worth?

Mostly younger, convenience-minded adults, many without a regular provider, paying by subscription, per visit or through insurance.

Telehealth became routine during the pandemic and then settled. National Health Interview Survey data reported by [HealthDay](https://healthday.com/healthpro-news/health-technology/2021-to-2022-saw-decrease-in-telemedicine-use-in-past-12-months) show the share of adults who used telemedicine in the past year fell from 37.0% in 2021 to 30.1% in 2022. That is still roughly three in ten adults, which makes virtual care a mainstream choice rather than a niche one. In-person practices compete for the same patients; see [how clinics and hospitals win patients from AI](https://underneath.agency/resources/healthcare-providers-patients-ai-search).

Consumer telehealth brands earn money in three ways, and the value of a new patient depends on which:

| Model | Real example (public figures) |
|---|---|
| Subscription care | Hims & Hers had 2.891 million subscribers at the end of Q2 2026, up 19%, and monthly revenue per average subscriber of $92 |
| Therapy subscription and insurance | BetterHelp segment revenue was $212.6 million in Q2 2026, down 12%, as demand shifted toward insurance-covered care |
| Pay per visit | Amazon One Medical charges $29 for a pay-per-visit telehealth consult |

Hims defines a subscriber as a customer with at least one subscription that bills automatically, so each new patient who stays is worth recurring monthly revenue. We found no public lifetime value per telehealth patient and will not invent one. Companies selling software to hospitals rather than patients face a different buyer, covered in [our guide to healthcare software and AI search](https://underneath.agency/resources/healthcare-software-ai-search).

## Where do AI assistants sit in a patient’s search for virtual care?

Before the visit: people ask an assistant about symptoms and options, then many go looking for a provider.

KFF found that about a third of adults used AI for health information in the past year, most often to look up symptoms or general information, and 65% of users cited wanting quick information as a major reason. [Rock Health’s December 2025 survey](https://fiercehealthcare.com/ai-and-machine-learning/ai-chatbot-use-health-information-16-2024-rock-health-survey) of 8,000 adults, reported by Fierce Healthcare, also found 32% use, with ChatGPT the most used (23%), then Gemini (15%). Its researchers note people also use AI for “looking for specific providers and clinics,” and 40% consulted a provider after a chatbot answer.

The handoff from question to care is where a telehealth brand can be named. The platforms treat that handoff differently:

- **OpenAI** says [ChatGPT Health](https://openai.com/index/introducing-chatgpt-health/) “is not intended for diagnosis or treatment” and is designed to support, not replace, care from clinicians. Its launch post does not describe booking visits with outside telehealth providers.
- **Amazon** documents the opposite design: its Health AI agent can connect users directly to licensed One Medical providers. Fierce Healthcare reports it reached US consumers through Amazon.com and the Amazon app in March 2026.

Our inference: when the assistant belongs to a company that also sells virtual care, independent brands are competing with the assistant’s owner. When it does not, the assistant answers from what it can find on the web, so a brand’s public evidence matters more. Health apps face a similar test, covered in [our guide to digital health apps](https://underneath.agency/resources/digital-health-apps-users-ai-search).

## What do patients ask AI before choosing a virtual care service?

Whether a service covers their state and insurance, what it costs, whether it is legitimate and how it handles data.

The prompts below are illustrative, written by us to show the shape of provider-discovery questions. They are not captured from real users, and they deliberately stay away from asking what condition someone has or which treatment to take.

| Stage | Illustrative prompt |
|---|---|
| Access | “Online doctor available tonight in Ohio for a common illness” |
| Insurance | “Virtual therapy services that accept my insurance plan” |
| Price | “How much does a telehealth visit cost without insurance?” |
| Comparison | “Compare subscription telehealth services on price and cancellation” |
| Legitimacy | “Is this telehealth company legit? Are the doctors licensed?” |
| Privacy | “Does this online therapy company share my data with advertisers?” |
| After the visit | “Can I get my prescription sent to my local pharmacy?” |

The insurance question is getting louder. Teladoc said demand for insurance-covered services at BetterHelp was “stronger than anticipated” and outpaced available provider capacity in the second quarter of 2026. Our inference: a service that clearly states which plans it accepts, in which states, gives an assistant a precise answer to repeat.

## How does an AI answer become a first visit?

Through a short list, a legitimacy check, an intake form and a consult, then a subscription or follow-up care.

1. **Named for a need.** The assistant lists a few services for a type of care, a state or an insurance plan.
2. **Checked.** The patient asks whether the company is legitimate, licensed and private. In [our brand reputation study](https://underneath.agency/research/is-it-legit-ai-reputation-study), 88.0% of AI answers to “Is this brand legit?” cited a review or complaint platform, and 99.7% of complete answers raised at least one problem with the brand.
3. **Intake.** The patient fills in an online questionnaire, often on a phone, and pays or confirms coverage.
4. **Consult.** A licensed clinician decides what care is appropriate. That decision belongs to the clinician, not the brand or the assistant.
5. **Retained.** Subscription models bill monthly, therapy continues over weeks, and pay-per-visit patients return when they need care again.

Step two is where telehealth brands are most exposed. Trustpilot and the BBB made up 61.7% of review-platform citations in that study, so complaints about billing or cancellation can shape the description a patient sees before the brand’s own pages do.

## What decides which telehealth brand an assistant names?

Verifiable facts about access, price, licensing and reputation; the platforms do not publish their selection rules.

What is documented: OpenAI and Amazon describe the purpose of their health assistants, as above. Google documents a hard gate for paid search: telemedicine providers must be LegitScript-verified to advertise. No platform documents an equivalent rule for which services its AI answers name.

What has been observed:

- **Health answers favor institutions.** In an audit of 615 sources cited by ChatGPT for consumer health questions, [Jacques and colleagues](https://arxiv.org/abs/2601.17109) found commercial health platforms earned 12.4% of citations; most came from hospitals, government bodies and similar institutions. We cover the pattern in [what commercial health sites cited by ChatGPT have in common](https://underneath.agency/resources/chatgpt-cited-commercial-health-sites).
- **Reputation answers lean on review sites,** as our study above showed. Our guide to [supplement brands in AI answers](https://underneath.agency/resources/supplement-brands-customers-ai-search) shows the same pattern.

What we infer: the trust factors in this industry are concrete and checkable. They include state licensure of clinicians, a clear list of states served, accepted insurance, transparent self-pay prices, LegitScript certification where it applies, a plain privacy policy, easy cancellation and a clean complaint record.

## What does it cost a telehealth brand to be missing or misdescribed?

More paid acquisition in a category where marketing already consumes a large share of revenue.

We have no direct measurement of patients lost to AI absence, so we show the economics and label the reasoning. Hims spent $262.2 million on marketing in one quarter. Teladoc cut its advertising and marketing spending from a year earlier, and BetterHelp’s paying users fell. Paid search is restricted to certified providers, and the [Federal Trade Commission](https://www.ftc.gov/news-events/news/press-releases/2023/03/ftc-ban-betterhelp-revealing-consumers-data-including-sensitive-mental-health-information-facebook) required BetterHelp to pay $7.8 million after finding it shared sensitive data, including that people had been in therapy, with platforms such as Facebook for advertising. The FTC said that targeting helped bring in “tens of thousands of new paying users.”

Our inference: as targeting with health data narrows and paid search stays gated, being named accurately in answers people already ask becomes a more valuable source of patients. A brand described with an outdated price, a state it no longer serves or an old complaint loses those patients to the next name on the list. The steps for correcting those errors are in [how to fix wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers).

## How does generative engine optimization work for a telehealth company?

Generative engine optimization (GEO) makes your service easy for assistants to find, describe accurately and verify through outside evidence.

For a consumer virtual care brand, the work usually covers:

1. **One consistent identity.** The same brand name, services, states and prices across your site, both app stores, review platforms and partner listings, so assistants do not merge you with a similarly named service.
2. **Access facts in text.** States served, insurance plans accepted, hours, typical wait times and self-pay prices, kept current and written on pages an assistant can read.
3. **Clinician credibility.** Named medical leadership, how clinicians are licensed and supervised, and service pages reviewed by clinicians, without promising any outcome or treatment. Wellness brands need the same care with claims; see [how wellness brands stay lawful in AI answers](https://underneath.agency/resources/wellness-brands-customers-ai-search).
4. **Certifications and policies.** LegitScript status where it applies, a privacy page that says plainly what you never share for advertising, and clear cancellation and refund terms.
5. **Reputation where patients check.** Accurate profiles on Trustpilot, the BBB and the app stores, real patient reviews and complaints resolved in public. [Fake reviews also distort AI answers](https://underneath.agency/resources/fake-reviews-fake-brands-ai-recommendations) and carry legal risk.
6. **Regulatory currency.** Telehealth rules change. The DEA’s fourth temporary extension, [reported by Telehealth.org](https://telehealth.org/blog/dea-extends-telehealth-controlled-substance-prescribing-flexibilities-through-the-end-of-2026/), allows clinicians to prescribe controlled substances without a prior in-person visit through December 31, 2026. Pages that explain what your service can and cannot do under current rules keep assistants from repeating old ones. Drugmakers face similar rules on what AI repeats, covered in [how pharma brands stay accurate in AI answers](https://underneath.agency/resources/pharma-brand-visibility-ai-search).
7. **Measurement.** Ask a fixed set of illustrative access, insurance, price and legitimacy questions across ChatGPT, Gemini, Perplexity and Google’s AI features, by state where it matters, and track who is named and which sources are cited.

No telehealth company can be promised a place in AI answers, and anyone who offers that promise is overselling. The aim is narrower: when a patient is ready to be seen, an assistant can describe your service correctly.

## What remains unknown about AI and telehealth patient acquisition?

How many virtual visits begin with an AI answer, and how assistants choose among telehealth brands, are not yet measured.

- **Follow-up is not booking.** KFF and Rock Health show people contact providers after AI answers, not which provider or whether it was virtual.
- **Company figures mix causes.** Hims’ growth and BetterHelp’s decline reflect products, prices and insurance shifts, not AI visibility.
- **No telehealth ranking studies.** The Jacques audit studied health information citations, and our reputation study covered brands in general. Applying them to telehealth brands is our inference.
- **Platform designs differ and change.** Amazon routes to its own care; OpenAI says its health product is not for diagnosis or treatment. Either could change.

## Where should a telehealth company start?

Start with the access, price and legitimacy questions patients in your states ask, and see who assistants name.

That first check usually shows whether your service is named for the types of care and states you serve, whether your prices, insurance and licensing are described correctly, which review sites shape your reputation and which competitors appear instead. From there, the work is to publish accurate access facts, strengthen your reputation and keep every page current with the rules.

If your growth depends on completed intakes and first visits, [ask us to review how AI assistants describe your service](https://underneath.agency/contact). We will map where your brand appears for access, insurance and legitimacy questions, why competitors are named instead, and which changes are most likely to bring more patients to your intake. How those changes are carried out, from state, insurance and price facts in text to clinician credibility and reputation work, is set out on our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page.

## Frequently asked questions

### Do patients use ChatGPT to find a telehealth provider?

Some do. Rock Health found people use AI chatbots for “looking for specific providers and clinics,” and KFF found 58% of AI users with physical health questions later followed up with a provider.

### Can a telehealth company advertise in AI answers?

We found no documented way to buy placement in AI health answers. On Google Ads, telemedicine providers must be verified by LegitScript before they can advertise.

### Does Amazon’s Health AI recommend other telehealth companies?

Amazon documents that its Health AI connects users to its own One Medical providers. It does not describe recommending outside telehealth services.

### Will health claims help a telehealth brand get named?

They add legal risk without documented benefit. Describe access, licensing, prices and policies accurately, and leave diagnosis and treatment decisions to clinicians.

## Sources

- KFF (2026-03), [KFF Tracking Poll on Health Information and Trust: Use of AI for Health Information and Advice](https://www.kff.org/health-information-and-trust/poll-finding/kff-tracking-poll-on-health-information-and-trust-use-of-ai-for-health-information-and-advice/)
- Fierce Healthcare (2026-03-24), [AI chatbot use for health information up 16% from 2024: Rock Health survey](https://fiercehealthcare.com/ai-and-machine-learning/ai-chatbot-use-health-information-16-2024-rock-health-survey)
- Subscription Insider (2026-08), [Hims & Hers Raises 2026 Outlook as Subscribers Near 2.9 Million](https://www.subscriptioninsider.com/blog/hims-hers-raises-2026-revenue-outlook-as-subscribers-near-2-9-million)
- Teladoc Health (2026-07), [Teladoc Health Reports Second Quarter 2026 Results](https://www.santelog.com/actualites-sante-nasdaq/teladoc-health-reports-second-quarter-2026-results)
- Fierce Healthcare (2026-03), [Amazon launches health AI agent on its website, expands free virtual care to 200M Prime members](https://www.fiercehealthcare.com/ai-and-machine-learning/amazon-launches-health-ai-assistant-its-website-expands-free-virtual-care)
- TechRepublic (2026-03), [Amazon Expands Health AI to Its Retail App, Offering Prime Members Free 24/7 Virtual Care](https://www.techrepublic.com/article/news-amazon-health-ai-app-site/)
- OpenAI (2026-01), [Introducing ChatGPT Health](https://openai.com/index/introducing-chatgpt-health/)
- HealthDay (2024-06-20), [2021 to 2022 Saw Decrease in Telemedicine Use in Past 12 Months](https://healthday.com/healthpro-news/health-technology/2021-to-2022-saw-decrease-in-telemedicine-use-in-past-12-months)
- Google Ads policies (2026), [Healthcare and medicines](https://support.google.com/adspolicy/answer/176031?hl=en)
- Federal Trade Commission (2023-03-02), [FTC to Ban BetterHelp from Revealing Consumers’ Data, Including Sensitive Mental Health Information, to Facebook and Others for Targeted Advertising](https://www.ftc.gov/news-events/news/press-releases/2023/03/ftc-ban-betterhelp-revealing-consumers-data-including-sensitive-mental-health-information-facebook)
- Telehealth.org (2026-01), [DEA Extends Telehealth Controlled Substance Prescribing Flexibilities Through the End of 2026](https://telehealth.org/blog/dea-extends-telehealth-controlled-substance-prescribing-flexibilities-through-the-end-of-2026/)
- Jacques and colleagues (2026), [Authority Signals in AI Cited Health Sources: A Framework for Evaluating Source Credibility in ChatGPT Responses](https://arxiv.org/abs/2601.17109), arXiv:2601.17109.
- Underneath (2026), [“Is this brand legit?” How AI assistants build a reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study)

---

This is the Markdown twin of https://underneath.agency/resources/telehealth-patients-from-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How do toy brands get found when shoppers ask AI for gift ideas?"
description: "By making age, play value and safety facts easy for AI to match to gift questions, and earning the reviews and gift-guide coverage AI answers draw on."
canonical: "https://underneath.agency/resources/toy-brands-product-discovery-ai-search"
published: 2026-10-09
updated: 2026-10-10
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How do toy brands get found when shoppers ask AI for gift ideas?

By giving AI assistants the facts they need to match a toy to a child’s age, interests and a buyer’s budget, and by earning the reviews and gift-guide coverage those assistants read. Toys were among the categories where shoppers leaned on AI most over the 2025 holidays. The brands that get named are the ones whose products are easy to describe correctly and easy to verify.

## The short version

1. Toys are an AI shopping category already: [Adobe](https://news.adobe.com/news/2026/01/adobe-holiday-shopping-season) says AI tools were used most in categories including video games, toys, appliances and electronics over the 2025 holidays, when US shoppers spent $8.8 billion on toys online, up 7.8%.
2. The market is growing and fan-driven: Circana data show US toy sales grew 6% in 2025, with games and puzzles up 37% and Pokémon toys up 87% to $2.5 billion, as reported by [Retail Dive](https://retaildive.com/news/premium-toys-drive-us-toy-market-rebound/812849/).
3. Gift timing raises the stakes: Circana says the second half of the year accounts for over 60% of annual toy sales ([VMSD](https://vmsd.com/kidult-demand-boosting-toy-sales-worldwide/amp/)).
4. AI use for holiday shopping more than doubled: in [Deloitte’s 2025 holiday survey](https://www.deloitte.com/content/dam/insights/articles/2025/us188488_cic-holiday-retail/pdf/DI_2025-Holiday-Survey.pdf), 33% of shoppers planned to use generative AI, up from 15% in 2024.
5. Safety facts are non-negotiable: the [Consumer Product Safety Commission](https://www.cpsc.gov/s3fs-public/Toy_Report_2023_Final_wCoverPage.pdf) requires third-party testing for all toys intended for children 12 and under, and estimated 231,700 toy-related emergency-room injuries in 2023.

## Who buys toys, and when does the money arrive?

Gift givers buying for a child they may not know well, and a growing group of adults buying for themselves.

Toy buying splits into two journeys. The first is gifting: a parent, grandparent, aunt or family friend needs something right for a particular age, interest and budget, often for a birthday or the holidays. The second is self-purchase by teens and adults, the “kidults” who collect trading cards, building sets and figures.

Both are growing. Circana data, summarized by [ClickPost](https://www.clickpost.ai/en-us/blog/us-toy-sales-statistics), put full-year US retail toy sales at $30.3 billion in 2025, and dollar sales grew 17% in the first half of 2026, the strongest first half in six years. In the first half of 2025, sales to recipients aged 18 and older rose 18%, and licensed toys made up 37% of US toy sales ([Spielwarenmesse](https://www.spielwarenmesse.de/en/mag/toy-market-news/circana-us-toy-market-in-growth-period/), reporting Circana data).

Shoppers are also trading up. Retail Dive reports that toys priced between $30 and $69.99 grew 18% in 2025, the fastest of any price band. A higher price means more research before buying, which is exactly where AI assistants come in. Electronics shoppers show the same research habit, as [our guide to consumer electronics](https://underneath.agency/resources/consumer-electronics-sales-from-ai-search) explains.

Most toy brands do not sell most of their volume themselves. The sale usually closes at a retailer: Amazon, Walmart, Target or a specialty store. So the revenue path for a toy brand is being named in an answer, then being bought wherever the shopper prefers to buy, and then earning repeat demand, reorders and shelf space from retailers who see the sell-through. Food brands sold through the same retailers face a similar path, covered in [how food ecommerce brands turn AI into orders](https://underneath.agency/resources/food-ecommerce-sales-from-ai-search).

## Where does AI already show up in toy shopping?

In gift ideas, comparisons and deal-hunting, inside both general assistants and retailers’ own shopping assistants.

- **Toys are a heavy-use category.** Adobe measured a 693.4% rise in AI-referred traffic to US retail sites over the 2025 holidays, and named toys among the categories where these tools were used most. Its top sellers that season included LEGO Icons sets, Hot Wheels sets, Bluey playsets and Fisher-Price Little People.
- **Gift ideas are a core use.** In an earlier [Adobe survey](https://blog.adobe.com/en/publish/2025/03/17/adobe-analytics-traffic-to-us-retail-websites-from-generative-ai-sources-jumps-1200-percent) of 5,000 US consumers, 35% of those using AI for shopping used it to get present ideas.
- **The platforms advertise it.** When [OpenAI launched shopping research](https://openai.com/index/chatgpt-shopping-research/) in ChatGPT, one of its own example requests was “I need a gift for my four year old niece who loves art.” [Amazon](https://aboutamazon.com/news/retail/best-gifts-kids-amazon-toys-we-love-list-2025) presents its Rufus assistant as ready to “compare toys, and help you pick the perfect gift.”
- **Toymakers are adopting it.** In June 2025, [Mattel announced a strategic collaboration with OpenAI](https://globaltoynews.com/2025/06/16/mattel-and-openai-announce-strategic-collaboration-pr/) to build AI-powered products and experiences based on its brands, with an emphasis on privacy and safety, and to use ChatGPT Enterprise in its own operations.
- **Shoppers plan to use it for the holidays.** Among Deloitte’s respondents planning to use generative AI, the top uses were comparing prices and finding deals (56%), reading summaries of reviews (47%) and generating shopping lists (33%).

Retail-wide, AI visitors are now valuable. In March 2026, AI traffic to US retailers converted 42% better than other traffic, according to Adobe data reported by [TechCrunch](https://techcrunch.com/2026/04/16/ai-traffic-to-us-retailers-rose-393-in-q1-and-its-boosting-their-revenue-too/). That figure covers all retail, not toys alone.

## What do toy shoppers ask AI assistants?

Questions defined by age, interest, budget, safety and occasion, rather than by brand name.

The prompts below are illustrative, written by us to show the types of question. They are not captured from real shoppers, and we make no claim about what any assistant answers to them today.

| Situation | Illustrative prompt |
|---|---|
| Age and interest | “Best building toy for a 6-year-old who loves space” |
| Budget gift | “Screen-free gifts for a 3-year-old under $40” |
| Safety | “Is this toy safe for a child under 3? Does it have small parts?” |
| Comparison | “Is the brand-name magnetic tile set worth it over cheaper ones?” |
| Adult collector | “Best display-worthy sets for an adult fan of Formula 1” |
| Trend | “What toys are kids asking for this Christmas?” |

Most of these questions do not mention a brand. That is the opportunity and the risk: the assistant chooses which brands to put in front of a buyer who has not decided yet. Questions about age are also safety questions, because the right age grade is the first fact a gift giver needs. Baby product brands face the same safety-first questions, covered in [how baby product brands get recommended](https://underneath.agency/resources/baby-product-brands-ai-recommendations).

## How does an AI recommendation turn into a toy sale?

Through a short list in the answer, a review and price check, then a purchase at a trusted retailer.

1. **Named.** The assistant suggests a few toys that fit the child’s age, interests and the budget.
2. **Checked.** The shopper reads reviews and compares prices. Deloitte found 50% of holiday shoppers would read online reviews before deciding, and Adobe recorded average toy discounts peaking at 29.6% off list price, so price comparison is part of the habit.
3. **Bought.** The click goes to a retailer or brand site, or the purchase happens inside a retailer’s own assistant. OpenAI says ChatGPT ranks the merchants it lists by factors including availability, price, quality and whether the merchant is the maker or primary seller of the item.
4. **Repeated.** A well-reviewed toy that sells through earns reorders and placement; we infer that AI-driven demand adds to that loop rather than replacing it.

Brands are also building their own assistants. Salesforce reported that companies that deployed their own AI agents, naming [Funko](https://www.salesforce.com/news/stories/2025-holiday-shopping-data/) among its customer examples, saw a 59% higher sales growth rate over the 2025 holidays than those that did not. That is a vendor’s analysis of its own customers, not a controlled comparison.

## What decides which toy an assistant names?

Specific product facts, reviews and independent coverage; the exact selection rules are not published.

**Documented by the platform.** OpenAI says ChatGPT’s [product results](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search) consider structured information such as price and product description from first-party and third-party providers, other third-party content, and OpenAI’s safety standards and product policies. It also says ChatGPT “may consider available options, price, reviews, and ease of use,” and that review summaries come from public websites and are not verified by OpenAI. Shopping research, OpenAI says, reads product pages directly and avoids low-quality or spammy sites.

**Observed in studies.**

- In a test of three AI systems, product facts such as rating, price and reviews explained 82.4% of how products were ranked, while brand name mattered mostly as a tiebreaker ([Chu and Hou](https://arxiv.org/abs/2606.17443)). We cover that study in [what actually drives AI product recommendations](https://underneath.agency/resources/what-drives-ai-product-recommendations).
- For the same shopping question, ChatGPT and Gemini showed only 5.4% of the same source domains on average ([Uberti-Bona Marin and colleagues](https://arxiv.org/abs/2609.18729)), so being visible in one assistant says little about the others.
- Ratings get checked before the answer is written: [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study) found ChatGPT searched for reviews or ratings in 46.2% of buyer questions.
- Ranked “best of” lists are the page type AI engines cite most when naming brands, as we explain in [which pages to target for AI recommendations](https://underneath.agency/resources/best-of-lists-ai-recommendations). For toys, those are gift guides and “best toys for age X” roundups.

**Our inference for toys.** The trust factors that likely matter most are a clear age grade and safety information, a concrete description of how the toy is played with and what it develops, strong retailer ratings, presence in independent gift guides and parenting publications, and accurate price and availability during the season.

## What does a toy brand lose by being invisible to AI?

Gift sales in a short, concentrated season, when a shopper has no favorite and the assistant fills the gap.

Nobody has published how many toy sales start in an AI answer, so the loss cannot be stated precisely. What the evidence does show is concentration. More than 60% of annual toy sales fall in the second half of the year, and most gift questions name an age and an interest rather than a brand. A toy that is not on the assistant’s short list when that question is asked misses a buyer who was ready to be persuaded.

Answers also move. Across five repeats of the same question in [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), just 25.2% of the brands ChatGPT named showed up every time. Being named once in a test proves little; being named reliably across assistants and weeks is the goal.

Accuracy carries its own cost. The CPSC received reports of 10 toy-related deaths among children 14 and under in 2023, including choking on a bouncy ball or a crayon and ingesting water beads. An answer that pairs a toy with the wrong age group, or repeats an outdated listing, is a risk to children and to the brand. Parents notice: in a [BSM Media](https://www.webull.com/news/13243978330498048) survey of nearly 500 mothers, 61% said they were more likely to buy a toy that does not require an app.

## How does GEO work for a toy brand?

Generative engine optimization (GEO) makes your toys easy for assistants to match, describe correctly and back with evidence.

For a toy company, that usually means:

1. **Complete, consistent product facts.** Age grade, safety warnings, piece count, dimensions, batteries, what is in the box, price and availability, identical on your site, retailer listings and product feeds. Our guide to [product content AI shopping assistants prefer](https://underneath.agency/resources/product-content-ai-shopping-assistants-prefer) explains why plain facts beat hype.
2. **Play value in plain words.** What a child does with the toy, which skills it supports, and which interests it fits, written as facts rather than superlatives, without educational or developmental claims you cannot support.
3. **Gift-guide and review coverage.** Earned placement in independent gift guides, parenting and toy publications, creator reviews and collector communities, timed well before the holiday rush.
4. **Retailer reputation.** Genuine ratings and answered questions on the retailer pages where most toys are bought, since assistants may draw on them and shoppers check them.
5. **Collector and licensing clarity.** For kidult lines, clear set names, edition details, release dates and the license, so assistants can answer specific fan questions accurately.
6. **Seasonal measurement.** Ask a fixed set of age, interest and budget questions across ChatGPT, Gemini, Perplexity, Google’s AI features and retailer assistants from late summer onward, repeat them, and track how often your toys appear. See [what to measure for AI visibility](https://underneath.agency/resources/what-to-measure-ai-visibility).

None of this can force an assistant to recommend a toy. It improves the odds that when an assistant looks, it finds accurate, trusted evidence for yours.

## What doesn’t the evidence tell toy brands yet?

How many toy purchases start with an AI answer, and how retailer assistants choose which toys to show.

- **Category data is thin.** Adobe names toys as a heavy AI-use category but does not publish toy-specific AI traffic or conversion figures.
- **Surveys measure plans.** Deloitte and Adobe asked people what they planned or remembered doing, which can differ from behavior.
- **Retailer assistants are opaque.** Amazon describes what Rufus does, not how it ranks toys.
- **Studies are adjacent.** The product-ranking test used skincare and the source-overlap audit covered many product types, not toys alone.

## Where should a toy brand start before the holiday season?

Start now, by asking the gift questions your buyers ask and checking which toys AI assistants name.

Pick your core age bands, interests and price points, ask them in the main assistants and in retailer assistants several times, and record which toys appear, which sources are cited, and whether your age grades, prices and availability are right. That baseline shows whether the gap is missing facts, thin review coverage or no presence in the gift guides assistants rely on.

If holiday sell-through is the number that matters, [let’s talk before the season peaks](https://underneath.agency/contact). We will show where your toys appear when shoppers ask AI for gift ideas, why competing toys appear instead, and which changes are most likely to put your products in front of more gift buyers this season. Our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page outlines the seasonal work behind that, including age and safety facts, gift-guide coverage and retailer reputation.

## Frequently asked questions

### Do shoppers really ask AI for toy gift ideas?

Yes. Adobe found 35% of US consumers who use AI for shopping used it for present ideas, and OpenAI and Amazon both promote their assistants for picking gifts, including toys.

### Does an AI assistant check a toy’s age rating?

Platforms do not say how age grades are used. OpenAI says safety standards and product policies are among the factors it considers, so brands should state age grades and warnings clearly everywhere their toys are listed.

### Is selling through Amazon enough to show up in AI answers?

Probably not on its own. In shopping answers studied by researchers, general assistants drew sources from review sites, communities and retailers, and ChatGPT and Gemini shared only 5.4% of the same domains.

### When should toy brands start working on AI visibility for the holidays?

Months ahead. Gift guides and reviews take time to earn, and Circana says over 60% of annual toy sales happen in the second half of the year.

## Sources

- Adobe (2026-01-07), [Holiday Shopping Season Drove a Record $257.8 Billion Online with Consumers Embracing Generative AI Tools](https://news.adobe.com/news/2026/01/adobe-holiday-shopping-season)
- Adobe (2025-03-17), [Adobe Analytics: Traffic to U.S. retail websites from generative AI sources jumps 1,200 percent](https://blog.adobe.com/en/publish/2025/03/17/adobe-analytics-traffic-to-us-retail-websites-from-generative-ai-sources-jumps-1200-percent)
- Retail Dive (2026-02-27), [Premium toys drive US toy market rebound](https://retaildive.com/news/premium-toys-drive-us-toy-market-rebound/812849/)
- ClickPost (2026), [US toy sales statistics](https://www.clickpost.ai/en-us/blog/us-toy-sales-statistics)
- Spielwarenmesse (2025), [Circana: US toy market in growth period](https://www.spielwarenmesse.de/en/mag/toy-market-news/circana-us-toy-market-in-growth-period/)
- VMSD (2025), [“Kidult” Demand Boosting Toy Sales Worldwide](https://vmsd.com/kidult-demand-boosting-toy-sales-worldwide/amp/)
- Deloitte (2025), [2025 Deloitte Holiday Retail Survey](https://www.deloitte.com/content/dam/insights/articles/2025/us188488_cic-holiday-retail/pdf/DI_2025-Holiday-Survey.pdf)
- U.S. Consumer Product Safety Commission (2024-11), [Toy-Related Deaths and Injuries, Calendar Year 2023](https://www.cpsc.gov/s3fs-public/Toy_Report_2023_Final_wCoverPage.pdf)
- OpenAI (2025-11-24), [Introducing shopping research in ChatGPT](https://openai.com/index/chatgpt-shopping-research/)
- OpenAI (2026), [Shopping with ChatGPT Search](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search)
- Mattel via Global Toy News (2025-06-12), [Mattel and OpenAI Announce Strategic Collaboration](https://globaltoynews.com/2025/06/16/mattel-and-openai-announce-strategic-collaboration-pr/)
- Amazon (2025), [Amazon’s 2025 Toys We Love list](https://aboutamazon.com/news/retail/best-gifts-kids-amazon-toys-we-love-list-2025)
- TechCrunch (2026-04-16), [AI traffic to US retailers rose 393% in Q1, and it’s boosting their revenue too](https://techcrunch.com/2026/04/16/ai-traffic-to-us-retailers-rose-393-in-q1-and-its-boosting-their-revenue-too/)
- Salesforce (2026-01), [AI and agents account for $262 billion of 2025 holiday spend](https://www.salesforce.com/news/stories/2025-holiday-shopping-data/)
- BSM Media via PR Newswire (2025-07-29), [Moms and Artificial Intelligence: New Survey from BSM Media Reveals Shift in Comfort, Concerns, and Consumer Expectations](https://www.webull.com/news/13243978330498048)
- Chu and Hou (2026), [Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems](https://arxiv.org/abs/2606.17443)
- Uberti-Bona Marin and colleagues (2026), ["If I Had to Buy Just ONE: Galaxy S26 Ultra": Auditing AI-Generated Product Recommendations](https://arxiv.org/abs/2609.18729)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)

---

This is the Markdown twin of https://underneath.agency/resources/toy-brands-product-discovery-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How AI search brings virtual care platforms contracts and visits"
description: "By getting the platform onto health system, plan and employer shortlists, and by helping clients’ patients find the virtual front door they paid for."
canonical: "https://underneath.agency/resources/virtual-care-platforms-ai-search"
published: 2026-10-09
updated: 2026-10-10
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How does AI search bring a virtual care platform new contracts and more patient visits?

It works on two fronts: it can put your platform on the shortlist when a health system, health plan or employer researches vendors, and it can help the patients those clients serve find the virtual care they already pay for. Both matter, because a platform earns its renewal through visits, not only through the signature. No study yet measures how often virtual care buyers use AI assistants, so treat this as a reasoned case built on public filings and surveys.

## The short version

1. One health system contract is worth hundreds of thousands of dollars a year: Amwell’s [2025 annual report](https://www.sec.gov/Archives/edgar/data/1393584/000119312526048776/amwl-20251231.htm) puts its average annual contract value at $461 thousand per health system client, on a typical contract term of three years.
2. The pool of large buyers is small and shrinking: Amwell’s average number of health system clients fell from 129 in 2023 to 88 in 2025, which it attributes to market consolidation and its own strategic shift.
3. Visits decide what a contract is worth: [Teladoc Health](https://s21.q4cdn.com/672268105/files/doc_financials/2025/q4/TDOC-4Q-2025-Earnings-Press-Release.pdf) ended 2025 with 101.8 million US Integrated Care members and average monthly revenue of $1.29 per member, while visits on the Amwell Platform fell from 5.9 million in 2024 to 4.5 million in 2025.
4. Employers are a real buyer: [KFF’s 2025 employer survey](https://www.kff.org/health-costs/2025-employer-health-benefits-survey/) found 30% of firms with 50 or more workers contract for virtual primary care, rising to 45% of firms with 1,000 or more workers.
5. The patients your clients serve already ask AI: in Rock Health’s survey [reported by HIT Consultant](https://hitconsultant.net/2026/03/23/rock-health-2025-survey-consumer-ai-adoption-chatgpt-healthcare/), 32% of US adults had used an AI chatbot for health information, and 74% of them used general-purpose tools such as ChatGPT rather than a provider’s own bot.

## Who buys a virtual care platform, and what is one client worth?

Three kinds of organization buy: health systems, health plans and employers. Each signs multi-year contracts and pays for access and use.

**Health systems** buy the platform to run their own virtual visits, virtual nursing, tele-stroke and remote sitting under their own brand. The buyer is usually a chief digital officer, CIO or virtual care leader, with clinical leaders and security reviewers holding a veto. Amwell says it powers the digital care programs of “approximately 80 of the nation’s largest health systems,” and it hosts a dedicated instance for each client, “brandable by the client.” That detail matters later: the patient often never sees the platform’s name.

**Health plans** buy virtual care to offer their members. Amwell reports about 50 health plan clients covering more than 90 million lives, and its largest client, Elevance Health, accounted for 31% of its 2025 revenue. Its top ten clients made up 71% of revenue. In this business a handful of contracts can make or break a year. How plans themselves win members is covered in [our guide to health insurers](https://underneath.agency/resources/health-insurers-members-ai-search).

**Employers** buy through benefits teams, consultants and health plans. KFF’s 2025 survey shows 30% of firms with 50 or more workers have a contract for virtual primary care that goes beyond their health plan network. Teladoc, which works through “relationships with health plans, employers, providers, health systems and consumers,” reported $1,579.6 million of Integrated Care revenue in 2025. Our guide to [healthtech vendors selling to employers and plans](https://underneath.agency/resources/healthtech-employers-payers-ai-search) covers that buyer in depth.

What one client is worth depends on the model:

- **Platform subscriptions.** Amwell’s average annual contract value per health system client was $461 thousand in 2025, down from $488 thousand in 2024, a drop it attributes to churn. Multiply by a three-year term and a single health system is a seven-figure decision.
- **Per-member access fees.** Teladoc earned average monthly revenue of $1.29 per US Integrated Care member in 2025. A plan or employer with hundreds of thousands of members is worth millions of dollars a year.

## Where does AI search already sit in a virtual care purchase?

Around the early vendor research and on the patient side; no survey yet isolates virtual care buyers.

The people inside health systems are now regular AI users. The [American Medical Association’s 2026 survey](https://www.fiercehealthcare.com/ai-and-machine-learning/ama-physicians-use-ai-doubled-2023-2026) found 81% of physicians use AI professionally, and [Fierce Healthcare reports](https://fiercehealthcare.com/ai-and-machine-learning/75-us-healthcare-systems-use-plan-use-ai-platform-2026) that 75% of US health systems use at least one AI application, up from 59% in 2025. Neither survey asks whether those leaders use ChatGPT or Gemini to choose a telehealth vendor.

The closest buyer data point comes from outside healthcare. [G2’s 2026 survey](https://company.g2.com/news/g2-research-the-answer-economy) of software buyers found 51% start their research with an AI chatbot more often than with Google. G2 runs a review marketplace and the sample is not healthcare-specific, so read it as a direction, not a measurement of your buyers.

A buyer who sticks to Google will probably meet an AI answer anyway. In [our AI Overview study](https://underneath.agency/research/ai-overviews-frequency-study) of US searches, keywords in B2B software and technology triggered an AI Overview on 96.0% of searches, the highest of the eight industries we sampled.

The patient side is better measured. Rock Health’s survey of 8,000 US adults found 32% had used an AI chatbot for health information, double the 16% of a year earlier. Only 5% of those users turned to a provider-offered bot. When a member wonders whether their plan covers a video visit tonight, a general assistant may answer before your client’s app does.

## Which questions do virtual care buyers put to AI assistants?

Questions about fit with their EHR, security, use cases, outcomes and alternatives. We wrote the prompts below ourselves as illustrations; none was collected from a real buyer.

| Buyer | Concern | Illustrative prompt |
|---|---|---|
| Health system | Category | “Which virtual nursing platforms do multi-hospital systems use?” |
| Health system | EHR fit | “Which telehealth platforms embed video visits inside Epic?” |
| Health system | Alternatives | “What are the alternatives to Amwell for a regional health system?” |
| Health plan | Member access | “Which virtual care vendors run white-label urgent care for health plans?” |
| Employer | Benefit design | “What virtual primary care options suit a self-insured employer with 5,000 staff?” |
| Any buyer | Security | “Which telehealth platforms hold HITRUST certification?” |

The competition in those answers is broad. Amwell’s annual report lists as competitors platform players such as Teladoc and Caregility, consumer telehealth firms, Microsoft, Amazon and Zoom, virtual nursing vendors, and the EHR companies themselves, including Epic and Oracle Health. An assistant asked a simple category question can draw on any of those.

Each question may trigger several searches behind the scenes. Google says AI Overviews and AI Mode [may use a “query fan-out” technique](https://developers.google.com/search/docs/appearance/ai-features), issuing multiple related searches, and OpenAI says ChatGPT search [typically rewrites a question](https://help.openai.com/en/articles/9237897-chatgpt-search) into one or more targeted queries. For a virtual care question, those hidden searches might cover EHR integration and certifications separately.

## How does an AI answer turn into a contract and then into visits?

Through two linked paths: the buyer’s shortlist, then the patient’s first visit, which together decide renewal.

**The contract path.** As we read the evidence, it runs like this:

1. A virtual care leader, benefits manager or plan executive asks a category, EHR-fit or alternatives question.
2. The assistant names a few platforms and links some of the pages it used.
3. The team checks those names against peers, KLAS and consultants.
4. A long evaluation follows. Amwell describes a sales cycle “ranging from a few months to a year” that often includes evaluating competitors, and says large clients “often begin to deploy our solution on a limited basis.” Medical device makers face a similar committee review, covered in [how devices build demand before value analysis](https://underneath.agency/resources/medical-device-demand-ai-search).
5. A pilot becomes a contract, typically for three years.

**The visit path.** Once live, the platform earns its renewal through use. Visits on the Amwell Platform fell from 5.9 million in 2024 to 4.5 million in 2025, and Teladoc’s total visits slipped from 17.3 million to 17.1 million. A contract whose members rarely log in is easy to cut at renewal. Here AI search works for your client, not for you: a member who asks an assistant how to see a doctor online should find the plan’s or health system’s own virtual front door, accurately described.

**What makes this industry different.** The platform is often invisible to the patient, because the service runs under the client’s brand. So a platform has two visibility jobs: its own name in buyer research, and its clients’ branded services in patient questions. We infer that the second job, done well, protects the first at renewal time.

Policy also shapes demand. Congress extended Medicare telehealth flexibilities, including care from home without rural restrictions, [through Dec. 31, 2027](https://aasm.org/congress-ends-partial-government-shutdown-and-extends-telehealth-flexibilities-and-the-work-gpci-floor-through-2027/), the American Academy of Sleep Medicine reports. Buyers weighing a multi-year contract will ask what happens after that date; a platform that explains it clearly in public gives assistants an accurate answer to repeat.

## What decides whether an assistant names your platform?

The assistants do not publish their selection rules; studies point to independent sources, and buyers check proof they can verify.

**Documented by the platforms.** Google and OpenAI confirm that their AI answers run web searches and show links to sources. Neither explains why one telehealth vendor is named and another is not.

**Observed in studies.** When [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study) logged ChatGPT’s queries, it looked for reviews in 46.2% of its answers and ran a search aimed at a named publication, ranking or award in 43.8%. On software questions, [Chen and colleagues](https://arxiv.org/abs/2509.08919) found AI search drew 72.7% of its sources from earned sites such as reviews and independent publications. In virtual care, a reasonable expectation is that those sources include KLAS, health IT trade press and benefits publications.

**What virtual care buyers check.** These trust factors are specific to the industry:

- **EHR integration.** A [KLAS patient engagement study](https://healthsystemcio.com/?p=88757) of 67 organizations found 55% used Epic, and 70% of those named Epic as the vendor most aligned with their goals. Amwell says building its technology natively into the EHR through channel partners “may lead to a higher win rate.” A platform that is vague about how it works inside Epic or Oracle Health gives buyers and assistants little to go on. Software sold to smaller practices faces similar EHR questions, as [our guide to medical practice software](https://underneath.agency/resources/medical-practice-software-ai-search) shows.
- **Security posture.** Amwell publicly lists HITRUST, ISO 27001 and PCI certifications. Health systems and plans will ask about these before any pilot.
- **Access priorities.** In the same KLAS study, 53% of respondents named patient access as a leading investment focus in 2025, up from 48% in 2023. Platforms that publish how they improve access give buyers language that matches their goals.
- **Proof of use.** Visit volumes, member engagement and outcomes with named clients, published with their permission.

**Our inference.** Most of what a virtual care buyer verifies could sit on public pages: EHR integrations, certifications, client logos, use cases and results. We expect a platform whose proof is public, current and consistent to be easier for both the buying committee and the assistant to describe correctly. No study has tested this for telehealth vendors.

## What does it cost a virtual care platform to be missing?

Fewer seats in evaluations and weaker renewals, though nobody has yet put a figure on the loss.

- **Each missed evaluation is large.** At an average of $461 thousand a year over a three-year term, one lost health system evaluation can mean more than a million dollars of contract value.
- **Fewer buyers raise the stakes.** With health system client counts falling and the top ten clients making up 71% of Amwell’s revenue, there are fewer chances to recover a missed opportunity.
- **The EHR fills the gap.** Epic and Oracle Health appear on Amwell’s own competitor list. If an assistant cannot place your platform, a reasonable expectation is that the buyer defaults to what the EHR already offers. Newer vendors face this most, as [our guide to healthcare startups](https://underneath.agency/resources/healthcare-startups-demand-ai-search) explains.
- **Answers move.** In [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), only 25.2% of the brands ChatGPT named for a question came back in all five runs. Being named once is not a settled position.

## How does GEO work for a virtual care platform?

By making your platform and your clients’ services easy for assistants to understand and cite. No mention is guaranteed.

1. **Separate the platform from the brand on the screen.** State plainly what you are (an enterprise platform), who buys you, and which clients run services on you, where they agree to be named. Consistency across your site, KLAS profile, EHR marketplace listings and company databases matters. When an assistant confuses you with a consumer telehealth app, our guide to [fixing wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) sets out the corrections.
2. **Publish integration and security facts.** List the EHRs you integrate with and how, your certifications and your data practices on plain web pages, not only in sales decks.
3. **Earn independent coverage.** KLAS interviews, health IT and benefits trade press, consultant briefings and conference talks are the kind of earned sources studies associate with AI citations. For a virtual care vendor, the steps for turning that earned record into authority are in [how brands build authority for AI search](https://underneath.agency/resources/how-brands-build-authority-for-ai-search).
4. **Answer each buyer’s question separately.** A health system, a health plan and a self-insured employer ask different things. Write honest pages on use cases, comparisons and implementation for each, without clinical claims you cannot support.
5. **Help clients’ front doors get found.** Give clients accurate, plain descriptions of their virtual services, coverage and how members start a visit, so a member’s question to an assistant leads to the right place. Commercial health sites that ChatGPT cites tend to show clear review and structure signals, as our summary of [what cited health sites have in common](https://underneath.agency/resources/chatgpt-cited-commercial-health-sites) explains.
6. **Track buyer and member questions across assistants.** Ask them repeatedly in ChatGPT, Google AI Overviews and AI Mode, Gemini, Perplexity, Copilot and Claude; [how many prompts to track](https://underneath.agency/resources/how-many-prompts-to-track-ai-visibility) covers the sample size.

Stay clear of shortcuts. Invented outcomes or planted reviews are a serious risk in a regulated field; see [legitimate GEO versus manipulation](https://underneath.agency/resources/legitimate-geo-vs-manipulation). Much of the buyer-side work adapts what we describe for [healthcare software vendors](https://underneath.agency/resources/healthcare-software-ai-search) and for [B2B SaaS companies](https://underneath.agency/resources/b2b-saas-revenue-from-ai-search).

## Which questions about virtual care and AI search are still open?

Most of them: nobody has yet linked AI visibility to virtual care contracts or visit volumes.

- **No survey of virtual care buyers’ AI use.** The physician and health system surveys measure AI adoption in general, not vendor research.
- **Patient data are self-reported.** Rock Health’s figures come from a consumer survey, not from logs of what people asked.
- **Filings describe one company each.** Amwell’s contract values and client counts are its own and may not match Teladoc, Caregility or EHR-native tools.
- **Policy after 2027 is unknown.** The Medicare extension runs through Dec. 31, 2027; demand beyond that depends on Congress.
- **The revenue link is unproven.** The general evidence is reviewed in [does AI visibility drive business results](https://underneath.agency/resources/does-ai-visibility-drive-business-results); for measurement, see [how to prove GEO caused sales](https://underneath.agency/resources/prove-geo-caused-sales).

## Where should a virtual care platform start?

Ask the questions your buyers and your clients’ members ask, see who gets named, then publish the missing proof.

List what a virtual care leader, a health plan executive and an employer benefits manager would ask at each stage, and what a member would ask when trying to start a visit. Put each question to several assistants more than once. Note which platforms are named, which sources are cited, and whether your EHR integrations, certifications and client services come back accurately.

If you want help with that, [talk to us about where AI answers place your platform with health systems, plans and employers](https://underneath.agency/contact). The review shows which buyer questions leave you off the shortlist, where your clients’ virtual front doors are hard to find, and which public proof would most help your pipeline of evaluations and the visit volumes that carry renewals. Our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page explains how that proof is published and tracked, for buyer questions about EHR fit and security as well as member questions about starting a visit.

## Frequently asked questions

### Do health system leaders use ChatGPT to pick telehealth vendors?

No published study measures it. What is known: 81% of physicians use AI professionally, and 75% of US health systems use at least one AI application.

### Why should a white-label platform care about patient-side AI answers?

Because visits drive renewals. Amwell’s visits fell from 5.9 million to 4.5 million in a year, and members who cannot find a service do not use it.

### Does Epic integration matter for AI visibility?

It matters to buyers: 55% of organizations in a KLAS study used Epic. Whether assistants weigh it is untested, but they can only repeat what is published.

### Are employers worth targeting separately from health plans?

Yes. KFF found 30% of firms with 50 or more workers contract for virtual primary care, and they research benefits differently from plans.

### Does the 2027 Medicare deadline change what we should publish?

Explain clearly how your platform serves clients under current rules and after Dec. 31, 2027, without legal advice. Buyers will ask.

## Sources

- American Well Corporation (2026-02), [Form 10-K for fiscal year 2025](https://www.sec.gov/Archives/edgar/data/1393584/000119312526048776/amwl-20251231.htm)
- Teladoc Health (2026-02-25), [Teladoc Health Reports Fourth Quarter and Full Year 2025 Results](https://s21.q4cdn.com/672268105/files/doc_financials/2025/q4/TDOC-4Q-2025-Earnings-Press-Release.pdf)
- KFF (2025-10), [2025 Employer Health Benefits Survey](https://www.kff.org/health-costs/2025-employer-health-benefits-survey/)
- HIT Consultant (2026-03-23), [Rock Health 2025 survey: consumer AI adoption](https://hitconsultant.net/2026/03/23/rock-health-2025-survey-consumer-ai-adoption-chatgpt-healthcare/)
- Health System CIO (2025), [Epic, Press Ganey Lead in Patient Engagement, KLAS Finds](https://healthsystemcio.com/?p=88757)
- American Academy of Sleep Medicine (2026-02), [Congress ends partial government shutdown and extends telehealth flexibilities through 2027](https://aasm.org/congress-ends-partial-government-shutdown-and-extends-telehealth-flexibilities-and-the-work-gpci-floor-through-2027/)
- Fierce Healthcare (2026), [AMA: Physicians’ use of AI doubled from 2023 to 2026](https://www.fiercehealthcare.com/ai-and-machine-learning/ama-physicians-use-ai-doubled-2023-2026)
- Fierce Healthcare (2026), [Health system AI adoption surges in 2026 with execs reporting increased ROI: survey](https://fiercehealthcare.com/ai-and-machine-learning/75-us-healthcare-systems-use-plan-use-ai-platform-2026)
- G2 (2026-04-15), [In the Answer Economy, Don’t Win the Click — Win the Answer](https://company.g2.com/news/g2-research-the-answer-economy)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- OpenAI Help Center (2025), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Chen, Wang, Chen and Koudas (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919), arXiv:2509.08919.
- Underneath (2026), [When does Google show an AI Overview? 1,248 US searches](https://underneath.agency/research/ai-overviews-frequency-study)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)

---

This is the Markdown twin of https://underneath.agency/resources/virtual-care-platforms-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How warehouse automation firms win enterprise deals in AI search"
description: "Partly: buyers now use AI to learn options and build longlists years before an RFP, so AS/RS, AMR and sortation vendors must be named and verifiable there."
canonical: "https://underneath.agency/resources/warehouse-automation-enterprise-leads-ai-search"
published: 2026-10-09
updated: 2026-10-10
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Are AI assistants deciding which warehouse automation vendors make the enterprise shortlist?

Partly. Distribution leaders now use AI assistants to learn what AS/RS, autonomous mobile robots and sortation can do, and to draw up a first list of vendors, long before a formal request for proposal goes out. The vendor that is missing or misdescribed at that stage may never be invited to bid on a project worth millions.

This article is for the executives of companies that design, build or integrate warehouse automation: automated storage and retrieval systems (AS/RS), autonomous mobile robots (AMRs), conveyors and sortation, and goods-to-person picking. It covers enterprise distribution buyers, not robot makers in general or logistics service providers.

## The short version

1. Adoption is now mainstream: in the [2026 Intralogistics Robotics Survey](https://www.supplychain247.com/article/2026_intralogistics_robotics_survey_robotics_moves_into_the_mainstream) by Peerless Research Group, MHI and The Robotics Group, 52% of respondents used robots (48% a year earlier) and the share with no plans fell from 9% to 3%.
2. The early stage is research: 47% of companies still planning robotics were gathering information and building internal knowledge, and 47% expected more than two years from project start to go-live.
3. Engineers use AI but check it: 69% of technical buyers used generative AI in purchasing in the [2026 State of Marketing to Engineers](https://advertising.globalspec.com/state-of-marketing-to-engineers-research-report) research by TREW Marketing and GlobalSpec, and rated its trustworthiness 4.7 out of 10.
4. The prize is large: in the [2026 MHI and Deloitte industry report](https://www.mannpublications.com/fashionmannuscript/2026/04/15/new-mhi-and-deloitte-report-finds-ai-is-biggest-disruptor-of-supply-chains-over-the-next-decade/), 52% of organizations planned to spend over $1 million on supply chain innovation and 17% over $10 million.
5. The field is crowded: [MODEX 2026](https://www.mhisolutionsmag.com/index.php/2026/06/26/modex-2026-sets-new-record/) had 1,057 exhibitors, and 70% of technical buyers said they were likely to choose the better-known brand when two solutions are technically similar.

## Who buys warehouse automation, and what is one enterprise customer worth?

Operations and supply chain executives with a capital budget, and one customer can mean a multi-site, multi-year program.

The buyer is rarely one person. In the 2026 robotics survey, 41% of respondents were in corporate management and 20% were logistics leaders, and operations owned the robots in 63% of companies that had them. Engineering, IT, finance and procurement join as the project grows. [Forrester’s 2026 State of Business Buying](https://www.forrester.com/press-newsroom/forrester-2026-the-state-of-business-buying/) found the typical business purchase now involves 13 internal stakeholders and nine external influencers, and procurement was a decision-maker in 53% of buying cycles.

The reason to buy is labor. Asked for the single most important factor, 67% of respondents in the robotics survey named labor costs and 33% labor availability. The business case is judged on return on investment (63%), payback time (52%) and total cost of ownership (47%).

Deal sizes are large and lumpy. Public vendor results show the shape:

| Public figure | Value | Source |
|---|---|---|
| Symbotic backlog, end of fiscal 2025 | $22.5 billion | [Q4 2025 earnings call](https://www.fool.com/earnings/call-transcripts/2025/11/24/symbotic-sym-q4-2025-earnings-call-transcript/) |
| Symbotic operational systems | 48 | same |
| AutoStore order intake, Q2 2026 | $218 million (up 45%) | [AutoStore Q2 2026 results](https://mfn.se/all/a/autostore/autostore-q2-2026-financial-results-0d1c3482) |
| AutoStore installed base | about 2,000 systems in 68 countries | same |
| Organizations planning over $10 million on supply chain innovation | 17% | MHI and Deloitte, 2026 |

The value of a customer continues after go-live. Symbotic’s software revenue grew 57% year over year to $9.3 million in its fiscal fourth quarter, and its operations services revenue grew 21%. In the robotics survey, 35% of companies already had funded new initiatives in progress, and 46% expected to use more than five types of robots within three years. Our inference: the first system is the entry ticket to a customer’s later sites and upgrades, so the cost of losing the first evaluation is larger than the first contract.

## Where do AI assistants already sit in an automation project?

At the start, when teams learn which technologies exist and which vendors to call, and again for the business case.

No published survey yet measures AI use among warehouse automation buyers specifically. The closest evidence comes from engineers and from business buyers in general:

| Finding | Share | Source |
|---|---|---|
| Technical buyers who use generative AI in purchasing | 69% | TREW Marketing and GlobalSpec, 2026 |
| Who routinely research on generative AI platforms (up 8 points in a year) | 21% | same |
| Who routinely research in online technical publications | 76% | same |
| Share of the technical buying journey done online before contacting a vendor | 62% | same |
| B2B buyers who used generative AI in a recent purchase, mainly to gather information on vendors and products | 45% | [Gartner, 2026](https://www.gartner.com/en/newsroom/press-releases/2026-05-20-gartner-survey-finds-sixty-nine-percent-of-b-two-b-buyers-turn-to-sales-reps-to-validate-ai-generated-insights) |
| B2B buyers who prefer to validate AI-generated insights with sales reps | 69% | same |

Two patterns matter for an automation vendor. First, buyers use AI to start and then verify: engineers rated AI answers 4.7 out of 10 for trust, and Forrester describes generative AI searches as the starting point for business buyers, followed by validation from trusted people and sources. Second, warehouse projects begin with a long learning phase. Among companies still planning robotics, 47% were gathering information, 13% were developing strategy and only 12% were finalizing the business case and capital approval. That learning phase is where an assistant explains what a shuttle system, a cube storage grid or an AMR fleet is, and whose products are examples. Robot makers selling into factories face a similar research phase, covered in [how robot makers reach manufacturer shortlists](https://underneath.agency/resources/robotics-companies-customers-ai-search).

Supply chain leaders are also primed to ask. In the 2026 MHI and Deloitte report, 48% rated the disruptive impact of AI as significant or greater, up 25 points from 2025, and 39% said the same of robotics and automation, up 16 points.

## Which questions do operations leaders ask AI about warehouse automation?

Questions about fit, payback, risk and alternatives, usually with the site’s own constraints written in.

The prompts below were written by us to show the kinds of questions an operations or engineering team might ask. They are examples, not recorded buyer prompts or observed AI answers.

| Stage | Illustrative prompt |
|---|---|
| Learning | “What is the difference between a shuttle AS/RS and a cube storage system for small parts?” |
| Fit | “Goods-to-person options for a 300,000 square foot brownfield DC with 40,000 SKUs and low ceilings” |
| Payback | “Typical payback for automated parcel sortation at 15,000 packages an hour” |
| Funding | “Robotics as a service or buy outright for 60 AMRs?” |
| Vendor longlist | “Which companies supply pallet AS/RS for frozen food warehouses in North America?” |
| Risk | “What happens to my system if an automation vendor goes bankrupt?” |
| Integration | “Which integrators install this system in Texas and connect it to our warehouse management system?” |

The risk question is not hypothetical. [Interact Analysis noted](https://www.mhwmag.com/features/warehouse-automation-what-to-expect-in-2026/) that Attabotics filed for bankruptcy in 2025 and that other vendors divested or closed units, including Zebra’s robotics division. Buyers committing capital for a decade want evidence of a vendor’s staying power, and an assistant will answer with whatever it can find.

Funding questions are live too. In the robotics survey, planners split between hybrid capital and operating models (36%), pure capital purchase (36%) and robotics as a service (29%). A vendor whose commercial models are not described anywhere public cannot be matched to the buyer’s preferred one.

## How does an AI answer become a signed automation project?

Through a technology choice, a longlist, a proposal request, a business case and a pilot, then rollout.

1. **Learning and technology choice.** The team learns which system types suit its order profile. An assistant’s explanation can push the project toward one technology, and with it toward the vendors that supply it.
2. **Longlist.** In the robotics survey, buyers leaned on materials handling suppliers (47%), robotics vendors (40%), trade associations (35%) and industry analysts or advisory firms (33%). An assistant now sits beside those sources when the first list is made. The advisory firms on that list are now judged by what AI says about them too, as [when CEOs ask AI which consultancy to hire](https://underneath.agency/resources/management-consulting-firms-clients-ai-search) shows.
3. **Request for information or proposal.** A short list of suppliers and integrators is invited to respond, often with a consultant involved. Integrators face their own shortlist, covered in [how automation integrators reach plant leaders](https://underneath.agency/resources/automation-integrators-industrial-buyers-ai-search).
4. **Business case and approval.** Simulation, return on investment and payback go to finance. This stage is slow: 47% of planners expected more than two years from project start to go-live.
5. **Pilot or proof.** Forrester found 78% of buyers making purchases of $10 million or more ran a trial first.
6. **Rollout and expansion.** Software, service and new sites follow, and in the robotics survey 45% of companies with robots said their budgets were rising.

AI visibility can change the first two steps. Engineering, price, references and delivery decide the rest. But a vendor not on the longlist in step 2 never reaches step 3, and Interact Analysis expects 2026 demand to broaden from a handful of giants to large enterprises and mid-sized companies, many of them buying automation for the first time with no incumbent vendor in mind.

## What decides whether an assistant names your system?

The platforms document how they search, not how they choose; studies point to independent coverage and specific, checkable facts.

On the search side, the platforms are open. [OpenAI’s help page](https://help.openai.com/en/articles/9237897-chatgpt-search) says ChatGPT search reworks an operations leader’s question into one or more targeted queries for its search providers, and only sites that let in its crawler, OAI-SearchBot, are eligible. [Google’s announcement](https://blog.google/products/search/ai-mode-search/) says AI Mode runs several related searches across a question’s subtopics, a technique it names “query fan-out.” Neither publishes how vendors are selected for an answer.

What has been observed in our studies, which covered buyer questions across several industries rather than automation:

- **Assistants search for rankings and publications.** In [our study of hidden searches](https://underneath.agency/research/ai-hidden-searches-study), ChatGPT ran a mean of 3.7 searches per answer, and in 43.8% of its answers it ran a search aimed at a named publication, ranking or award.
- **Google rank is only part of it.** In [our comparison of AI citations with Google rankings](https://underneath.agency/research/ai-citations-google-rankings-study), 8.3% of the pages ChatGPT cited ranked in Google’s top 10 for the question.
- **Answers vary between assistants.** In [our brand agreement study](https://underneath.agency/research/ai-assistants-brand-agreement-study), two assistants’ recommendations for the same question overlapped by only 0.327 on average, on a scale where 1 means identical lists.

Our inference for warehouse automation: the trust factors are the ones an engineering team already checks, written where a machine can read them. That means system types and throughput ranges, storage density and footprint, temperature ranges, the industries and order profiles served, named reference sites (with permission), integrator coverage by region, software interfaces, service and spare-parts commitments, and evidence of financial stability. It also means coverage in the publications engineers trust, which in the GlobalSpec survey edged out vendor websites as their top research destination for the first time.

## What does it cost to be missing from the early longlist?

A seat at an evaluation that may decide years of follow-on business; no one has measured the loss directly.

We found no study that measures automation projects lost because a vendor was absent from AI answers, so we set out the reasoning and label it:

- **The list forms before contact.** Engineers complete 62% of the journey online before talking to a vendor. We infer that a vendor missing from the early research is often missing from the request for proposal, and that the loss shows up as invitations that never arrive.
- **Recognition decides close calls.** 70% of technical buyers said they would likely pick the better-known brand when two solutions are technically similar, and 53% said familiarity influenced their most recent purchase. With 1,057 exhibitors at MODEX 2026, being remembered is hard, and being named by an assistant is one more place to be remembered.
- **Wrong facts filter you out.** An assistant that says your system does not handle frozen goods, totes over a certain weight or a given throughput removes you from a project you could win. Our guide to [fixing wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers) explains how to trace an error to the page behind it.
- **Expansion follows the first win.** Because customers add sites and robot types after a first project, a lost first evaluation can cost the later ones too. This is our inference from the survey and vendor results above.

## How does GEO work for a warehouse automation company?

Generative engine optimization (GEO) makes your systems easy for AI assistants to find, describe accurately and verify against outside sources.

For an automation vendor or integrator, the work usually covers:

1. **Application pages written as text.** One page per use case, such as each-picking for e-commerce, pallet storage in freezers or parcel sortation, with throughput, density, footprint and building requirements stated plainly, not only in brochures, videos or PDFs behind a form.
2. **A clear business case method.** How you calculate payback and total cost of ownership, which inputs matter and what ranges customers have seen, in public, so an assistant answering a payback question has a source that names you.
3. **Commercial models stated.** Capital purchase, hybrid and robotics-as-a-service options described where buyers and assistants can find them.
4. **One consistent identity.** The same company name, product names, partner and integrator lists, regions and certifications on your site, partner sites, association directories and trade show listings.
5. **Independent coverage.** Features and case studies in trade publications, talks at MODEX and ProMat, association membership, and analyst coverage. Our guide on [how brands build authority for AI search](https://underneath.agency/resources/how-brands-build-authority-for-ai-search) explains why outside sources matter.
6. **Staying-power evidence.** Installed base, years in operation, service network and spare-parts commitments, stated in text and confirmed by outside reporting.
7. **Crawl access and measurement.** Allow the search crawlers the assistants document, then ask a fixed set of technology, fit and vendor questions across ChatGPT, Gemini, Perplexity, Copilot and Google’s AI features, repeatedly, tracking who is named and which sources are cited.

No one can promise a vendor a place on an AI-built longlist; this work makes your systems easier to find and confirm. Warehouse and transportation software vendors face the same longlist problem, which our articles on [logistics software](https://underneath.agency/resources/logistics-software-ai-search) and [supply chain software](https://underneath.agency/resources/supply-chain-software-ai-search) cover. Plant-floor software has its own version, covered in [MES and IIoT platforms in AI answers](https://underneath.agency/resources/industrial-software-plants-ai-search).

## What can’t the current data tell an automation vendor?

It shows buyers researching with AI, not how many automation contracts start with an AI answer.

- **No automation-specific AI study.** The AI usage figures come from engineers in general and from business buyers across industries. Applying them to warehouse automation is our inference.
- **Interested parties.** The robotics survey reached 166 subscribers of Modern Materials Handling and its sister publications; MHI is a trade association; GlobalSpec sells advertising to industrial firms; vendor figures come from investor reports.
- **No attribution.** We found no public data linking AI answers to requests for proposal or contract value in this industry.
- **Our studies are cross-industry.** They covered buyer questions in several sectors, not automation projects, and the platforms change how their assistants search over time.

## Where should a warehouse automation company start?

Start by asking assistants the questions a distribution team asks before it writes a request for proposal.

Run those questions for your system types, your target industries and your regions, and look at four things: whether you are named, whether your capabilities and limits are described correctly, which publications and sites the answers cite, and which competitors appear instead. The work that follows is to put your application facts where assistants read them and to earn the outside coverage that confirms them.

If your growth depends on being invited into a few large automation projects each year, [talk to us about your AI visibility](https://underneath.agency/contact). We will show where your systems appear when operations leaders research automation with AI, why other vendors are named instead, and which changes are most likely to bring more qualified invitations to bid. The ongoing work, such as application pages in plain text, a public payback method and evidence of staying power, is described on our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page.

## Frequently asked questions

### Do warehouse operators really use ChatGPT to research automation?

There is no automation-specific measure yet. Among technical buyers in general, 69% used generative AI somewhere in purchasing in 2026, and 21% routinely researched on AI platforms, up 8 points in a year.

### Can a smaller AMR or integrator compete with large automation vendors in AI answers?

It can be named for specific combinations of use case, building type, industry and region where fewer suppliers fit. Large vendors still benefit from recognition, which 70% of technical buyers say tips close decisions.

### Should we publish payback and pricing information?

Publishing your payback method and the inputs that drive it gives assistants a source to cite for business case questions. Exact project prices depend on the site, so ranges and worked examples are more useful than list prices.

### Does AI search replace trade shows and consultants?

No. Buyers still rely on suppliers, trade associations and advisory firms, and 69% of business buyers prefer to validate AI-generated insights with sales reps. AI changes who is on the first list.

## Sources

- Supply Chain 24/7, Peerless Research Group, MHI and The Robotics Group (2026-06-01), [2026 Intralogistics Robotics Survey: Robotics moves into the mainstream](https://www.supplychain247.com/article/2026_intralogistics_robotics_survey_robotics_moves_into_the_mainstream)
- TREW Marketing and GlobalSpec (2026), [State of Marketing to Engineers research report](https://advertising.globalspec.com/state-of-marketing-to-engineers-research-report)
- MHI and Deloitte via Mann Publications (2026-04-15), [New MHI and Deloitte report finds AI is biggest disruptor of supply chains over the next decade](https://www.mannpublications.com/fashionmannuscript/2026/04/15/new-mhi-and-deloitte-report-finds-ai-is-biggest-disruptor-of-supply-chains-over-the-next-decade/)
- MHI Solutions (2026-06-26), [MODEX 2026 Sets New Record](https://www.mhisolutionsmag.com/index.php/2026/06/26/modex-2026-sets-new-record/)
- Interact Analysis via Material Handling Wholesaler (2026-02-09), [Warehouse automation: what to expect in 2026](https://www.mhwmag.com/features/warehouse-automation-what-to-expect-in-2026/)
- The Motley Fool (2025-11-24), [Symbotic (SYM) Q4 2025 Earnings Call Transcript](https://www.fool.com/earnings/call-transcripts/2025/11/24/symbotic-sym-q4-2025-earnings-call-transcript/)
- AutoStore via MFN (2026-08-13), [AutoStore Q2 2026 financial results](https://mfn.se/all/a/autostore/autostore-q2-2026-financial-results-0d1c3482)
- Forrester (2026-01-21), [Forrester’s 2026 Buyer Insights: GenAI Is Upending B2B Buying](https://www.forrester.com/press-newsroom/forrester-2026-the-state-of-business-buying/)
- Gartner (2026-05-20), [Gartner survey finds 69% of B2B buyers turn to sales reps to validate AI-generated insights](https://www.gartner.com/en/newsroom/press-releases/2026-05-20-gartner-survey-finds-sixty-nine-percent-of-b-two-b-buyers-turn-to-sales-reps-to-validate-ai-generated-insights)
- OpenAI Help Center (2026), [ChatGPT search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- Google (2025-03-05), [Expanding AI Overviews and introducing AI Mode](https://blog.google/products/search/ai-mode-search/)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude cite pages that rank?](https://underneath.agency/research/ai-citations-google-rankings-study)
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude agree on brands?](https://underneath.agency/research/ai-assistants-brand-agreement-study)

---

This is the Markdown twin of https://underneath.agency/resources/warehouse-automation-enterprise-leads-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How wealth managers win clients when prospects vet them with AI"
description: "Referrals still lead, but prospects increasingly check a firm with AI. Firms win by making their niche, fees, fiduciary status and reviews easy to verify."
canonical: "https://underneath.agency/resources/wealth-management-firms-clients-ai-search"
published: 2026-10-09
updated: 2026-10-10
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How do wealth management firms win new clients when prospects check them with AI before calling?

Referrals still bring most clients to wealth management firms, but a growing share of prospects, especially younger and wealthier ones, now check a firm with AI before they call, and some find their advisor that way. A firm wins those clients by making who it serves, how it is paid, its fiduciary status and what clients say about it easy for an assistant to find, verify and repeat. That work has to fit within the SEC’s marketing rule.

## The short version

1. AI has entered advisor selection: in a June 2026 [survey of 1,000 advised investors](https://investmentnews.com/practice-management/wealthy-investors-are-less-referral-dependent-than-advisors-think/266951) reported by InvestmentNews, nearly 9% used an AI tool such as ChatGPT, Gemini or Claude while searching for an advisor, rising to 25% of investors under 45.
2. Referrals lead, but less for the wealthiest: in the same survey, 50% of investors with $5 million or more said no referral was involved in finding their current advisor; [Kitces](https://www.investmentnews.com/practice-management/faced-with-a-referral-well-that-has-run-dry-this-is-what-advisors-need-to-do-says-michael-kitces/268200) found 43% of consumers find an advisor through friends and family, and only about 4% through a search engine or AI.
3. AI is raising the value of advice, not replacing it: in [Vanguard’s June 2026 survey](https://corporate.vanguard.com/content/dam/corp/research/pdf/the_ai_advice_frontier.pdf) of 6,686 investors, 32% said AI made them value human advice more, against 2.9% who said less.
4. Fast growers depend less on referrals: high-growth practices in the [Kitces 2026 marketing study](https://www.financial-planning.com/news/kitces-research-pegs-the-winners-and-losers-in-marketing-roi) got only a third of new client revenue from referrals, against 80% for slower-growing peers.
5. The rules apply to every claim: the SEC’s Division of Examinations [flagged in December 2025](https://www.willkie.com/publications/2026/01/sec-division-of-examinations-issues-risk-alert-regarding-advisers-act-marketing-rule-compliance) that the most common testimonial failure was missing disclosures at the time the testimonial was first shared.

This guide covers how wealth management firms get found and described by AI assistants. It is not investment, legal or compliance advice; take marketing-rule questions to your chief compliance officer and counsel.

## How do affluent clients find a wealth manager today?

Mostly through people they trust, followed by their own research; the wealthiest and youngest rely on referrals least.

Referrals remain the largest single source. In Kitces’ 2026 marketing study, as presented by Michael Kitces, 43% of consumers find an advisor by asking friends and family, about a quarter through an event or networking, and about 20% through a trusted professional such as an attorney or accountant. Advisors mirror this: in the same study’s survey of 506 advisors, 88% used client referrals and 64% used referrals from centers of influence such as CPAs and estate attorneys.

The referral is often only the start. The June 2026 survey of advised investors found that 15% of advised investors received a referral and then used at least one other method to evaluate the advisor before reaching out. Among investors under 45, 59% found their advisor without any referral, and only 8% relied on a referral alone. What made the final choice was fit: nearly 74% said it was “very important” that the advisor showed an understanding of their specific needs.

Kitces describes a “referral paradox.” As a firm grows, recent clients become a smaller share of the book, a smaller share of clients still have untapped networks to refer from, and referral-driven growth slows. [Family Wealth Report’s summary](https://www.familywealthreport.com/article.php/-Marketing-Study%3A-Referral-Paradox%2C-Rise-Of-AEO-And-Advisor-Antipathy) of the study adds that asking clients for referrals more often appears to reduce them. That is why the fastest-growing firms lean on channels they control.

## Where does AI fit in choosing an advisor?

Mostly at the checking stage today, and increasingly at discovery for younger and wealthier prospects.

- **Discovery.** About 4% of consumers find an advisor through a search engine or AI, by Kitces’ count. Kitces noted that even a small share of the roughly 30 million US households with $100,000 or more to invest still means a large number of people.
- **Searching with AI.** The June 2026 survey’s figures are higher among the clients firms want most: 25% of investors under 45 and 15% of those with more than $5 million used an AI tool while searching for an advisor.
- **General AI use for money.** In Vanguard’s survey, 32% of investors had used AI for financial guidance: 43% of Gen Z and millennial investors, 41% of Gen X and 15% of baby boomers and older.
- **Trust stays with people.** Among Gen Z, millennial and Gen X investors, 63% reported low or no trust in AI advice, and 37% of AI users had received incorrect or misleading information. [Cerulli](https://www.cerulli.com/press-releases/investor-skepticism-of-ai-in-financial-advice-persists) found just 38% of affluent investors at least somewhat comfortable with AI.

The pattern favors advisors. Vanguard reports that among advised clients, 41% value their advisor more because of AI, and only 2% value them less. Prospects use AI to learn and to check, then hire a person. For a firm, the risk is not that AI replaces it; the risk is that the assistant a prospect consults cannot find, or misstates, the facts that would earn the first meeting.

## What do prospects ask AI about a wealth manager?

Discovery questions about specialists nearby, and verification questions about a specific firm. The examples are our own composites of what prospects tend to ask, not questions captured from real users.

| Stage | Illustrative prompt |
|---|---|
| Discovery by niche | “Fee-only fiduciary advisor in Denver who works with physicians” |
| Discovery by event | “How do I choose a wealth manager after selling my company?” |
| Verification | “Is this firm a fiduciary, and does it have any disciplinary history?” |
| Fees | “What does a wealth manager typically charge on $3 million, and how is this firm paid?” |
| Comparison | “Independent RIA or a large brokerage firm for a family with trusts?” |
| Reputation | “What do clients say about this firm?” |

The verification questions matter most for referred prospects. A referred prospect who asks an assistant about the firm and gets a thin or wrong answer may never book the meeting the referral set up, and the firm would not know why.

## How does an AI answer turn into new assets under management?

Through a short path: named or confirmed, website checked, introductory meeting, then a long relationship.

1. **Named or confirmed.** For a discovery question, the assistant names a few firms. For a verification question, it confirms or fails to confirm what the referrer said.
2. **Website and records.** The prospect reads the firm’s site, advisor biographies and fee page, and may check public registration records.
3. **Introductory meeting.** The firm earns a conversation, where fit decides the outcome.
4. **Engagement.** Assets move, and fees accrue for years if the relationship holds.

Each new relationship is significant for a typical firm. The [2026 Investment Adviser Industry Snapshot](https://www.investmentadviser.org/iaatoday/press-release/2026-investment-adviser-industry-snapshot-shows-continued-growth-in-demand-for-adviser-services/) counts 16,544 SEC-registered advisers serving 73.7 million clients, and advisers focused on individuals averaged just 8 employees and $424 million in assets under management. Small teams cannot meet everyone; they need the right introductions.

Marketing efficiency is measurable. Kitces found the typical practice spends 7% of annual revenue on marketing and 70 cents for each new dollar of client revenue. By tactic, online advisor directories cost $0.28 per new revenue dollar, client referrals $0.34 and search engine optimization $0.45. Answer engine optimization, AI-focused visibility work that Kitces measured for the first time, cost $3.45, the report’s summary notes. The report also warned that AI search may never become a major route to advisor recommendations. We read this as a caution against buying AI visibility as a standalone tactic, and a case for doing the groundwork (accurate facts, directories, reviews, search) that serves both search engines and assistants.

## What makes an assistant name or vouch for a firm?

Public, consistent facts and third-party signals; platforms do not document how they choose advisors.

**Documented by the platform.** Google explains that AI Overviews and AI Mode can split one question into several related searches, a method it calls [“query fan-out”](https://developers.google.com/search/docs/appearance/ai-features), before writing an answer. For subjects that could affect someone’s financial stability, such as choosing who manages their money, Google adds that its systems [give even more weight](https://developers.google.com/search/docs/fundamentals/creating-helpful-content) to content that shows strong experience, expertise, authoritativeness and trust. Neither Google nor OpenAI documents how it chooses which advisory firm to name.

**Observed in our studies.** None of our studies tested financial advisors, so these findings come from neighboring local professional services.

- **Reviews and maps matter for local questions.** In [our study of ChatGPT’s local picks](https://underneath.agency/research/chatgpt-local-picks-google-profile-study), covering services such as accountants and lawyers, businesses with more Google reviews than the local median were 19.5 points more likely to be listed after adjusting for other signals, and ChatGPT listed 67.7% of businesses ranked in the Maps top 3.
- **Local answers vary more.** In [our agreement study](https://underneath.agency/research/ai-assistants-brand-agreement-study), recommendation lists from four assistants overlapped far less when the question named a place (0.160 on a scale where 1 means identical) than when it did not (0.390).
- **Named sources travel.** In [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study), when ChatGPT’s search named a source such as a directory or comparison site, the answer cited it 44.0% of the time, against 8.1% when it did not.

**Our inference.** For a wealth firm, the facts that let an assistant vouch for you are specific: whom you serve, how you are paid, whether you act as a fiduciary, your credentials, where you are registered, and what clients say. They need to agree everywhere they appear, from your website to your regulatory filings to directory profiles. Kitces found that only 13% of practices use third-party review sites in their marketing, and at the median firm only 16% of clients left a review. A firm with a thin public record leaves assistants little to say, and prospects little reason to call.

## How does the SEC marketing rule shape AI visibility work?

It governs the content assistants draw on, so compliant content is the only safe route to visibility.

Registered advisers’ advertisements fall under the SEC’s [marketing rule](https://www.ecfr.gov/api/renderer/v1/content/enhanced/current/title-17?part=275&section=275.206%284%29-1). It bars untrue statements of material fact and misleading implications. It allows testimonials and endorsements only with clear and prominent disclosures, oversight and, where someone is paid, a written agreement. It allows third-party ratings only with a reasonable basis for believing the underlying survey was fair, plus disclosures such as whether the adviser paid in connection with the rating.

The December 2025 risk alert shows where firms slip. Examiners most often found testimonials and endorsements without the required disclosures at the time they were shared, on websites, social media, and through lead-generation firms and influencers. They also found third-party ratings used without the required basis or disclosures.

Three points follow for GEO, which your compliance team should confirm:

1. **Reviews need a program, not a push.** Inviting clients to leave Google reviews and displaying them involves testimonial rules; design the process with compliance before scaling it.
2. **Awards and rankings carry duties.** “Top advisor” lists are exactly the third-party content assistants may repeat, and using them in your marketing triggers the rating provisions.
3. **Never manufacture signals.** Paid placements disguised as editorial content, or reviews written to influence AI answers, create regulatory risk and, as covered in [can fake reviews make AI assistants recommend a fake brand?](https://underneath.agency/resources/fake-reviews-fake-brands-ai-recommendations), reputational risk too.

## What does GEO look like for a wealth management firm?

Generative engine optimization (GEO) makes your firm’s niche, terms and reputation easy for assistants to find, check and repeat accurately.

For an RIA or wealth firm, the work usually covers:

1. **A clear niche page.** Who you serve best (business owners, physicians, executives with equity pay, retirees), the problems you solve, and minimums, written plainly. Kitces found firms with a well-defined niche overrepresented among high-growth practices with up to $1 million in revenue.
2. **Fees and status in text.** How you are paid, whether you act as a fiduciary, and your services, consistent with your regulatory filings.
3. **Advisor profiles.** Credentials, experience and specialties for each advisor, matching their public registration records.
4. **Directories and maps.** Accurate profiles on the directories and map listings assistants draw on, the lowest-cost tactic in Kitces’ data.
5. **A compliant review program.** Designed with your compliance team, so the reviews that prospects and assistants read are real, current and properly disclosed.
6. **Professional and press coverage.** Articles, podcasts and talks that show expertise in your niche, plus visibility with the CPAs and attorneys who refer clients.
7. **Monitoring.** Track discovery and verification questions about your firm in several assistants and in Google’s AI surfaces, since local answers differ most.

If an assistant misstates your fees, minimums or credentials, the correction process is in [how to fix wrong information about your brand in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers). For building third-party credibility more broadly, see [how brands build authority that AI search recognizes](https://underneath.agency/resources/how-brands-build-authority-for-ai-search), and for smaller firms competing with national names, [how a small brand can get recommended by AI assistants](https://underneath.agency/resources/how-small-brands-get-recommended-by-ai). Fund firms face a related question, covered in [how asset managers get funds considered by AI](https://underneath.agency/resources/asset-managers-fund-demand-ai-search). No one can guarantee that an assistant will recommend a firm; GEO makes the evidence it finds accurate, consistent and compliant.

## What can’t the current advisor data tell us yet?

How many new clients AI produces: surveys measure use, not the assets that follow.

- **Small channel, uncertain growth.** Kitces’ roughly 4% and the June 2026 survey’s 9% measure different things in different samples. Neither shows how fast AI discovery is growing.
- **Interested parties and narrow samples.** Vanguard and Cerulli serve the wealth industry, so their findings are useful but not neutral. The June 2026 survey covered only investors who already work with an advisor.
- **No advisor-specific AI studies.** Our local-picks evidence comes from other professional services; we have not tested advisor questions directly.
- **Cost figures are early.** Kitces measured AI-focused visibility work for the first time in 2026; the report’s summary suggests it may become more efficient as it matures.

## Where should a wealth management firm start?

Start by checking what assistants say when a referred prospect asks about your firm, then fix the gaps.

A useful first review covers verification questions about your firm and each lead advisor, discovery questions for your niche in your markets, and how your fees, fiduciary status and credentials are described. It shows whether assistants confirm what your referrers say, which firms they name instead, and which facts are missing or inconsistent with your filings.

If your growth depends on turning introductions into new relationships and assets, [talk to us about a review of your firm in AI answers](https://underneath.agency/contact). We will compare how assistants describe you and the firms you compete with, list what they get wrong or cannot find, and plan the profile, directory, review and content work, designed with your compliance team, that helps prospects who check you with AI take the next step. Our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page describes how that work runs for an advisory firm, from niche and fee pages to advisor profiles and monitoring.

## Frequently asked questions

### Do wealthy investors really use ChatGPT to find a financial advisor?

Some do. In a June 2026 survey of advised investors, 15% of those with more than $5 million used an AI tool while searching for an advisor, and 25% of those under 45.

### Will AI replace wealth managers?

The evidence so far says no. In Vanguard’s survey, investors were about 11 times as likely to say AI raised the value they place on human advice as to say it lowered it.

### Can we ask clients for Google reviews?

Reviews can help, but under the marketing rule testimonials carry disclosure and oversight duties. Design any review program with your compliance officer before you start.

### Is GEO worth it if most clients come from referrals?

Yes, if it supports referrals. Referred prospects often check a firm before calling, and an accurate, verifiable public record helps that referral turn into a meeting.

## Sources

- InvestmentNews (2026-06-09), [Wealthy investors are less referral-dependent than advisors think](https://investmentnews.com/practice-management/wealthy-investors-are-less-referral-dependent-than-advisors-think/266951)
- InvestmentNews (2026-09-15), [Faced with a referral well that has run dry, this is what advisors need to do, says Michael Kitces](https://www.investmentnews.com/practice-management/faced-with-a-referral-well-that-has-run-dry-this-is-what-advisors-need-to-do-says-michael-kitces/268200)
- Financial Planning (2026-09-15), [New Kitces research pegs the winners and losers in marketing ROI](https://www.financial-planning.com/news/kitces-research-pegs-the-winners-and-losers-in-marketing-roi)
- Family Wealth Report (2026), [Marketing Study: Referral Paradox, Rise Of AEO And Advisor Antipathy](https://www.familywealthreport.com/article.php/-Marketing-Study%3A-Referral-Paradox%2C-Rise-Of-AEO-And-Advisor-Antipathy)
- Vanguard (2026-09), [The AI advice frontier: Use, trust, and the human edge](https://corporate.vanguard.com/content/dam/corp/research/pdf/the_ai_advice_frontier.pdf)
- Cerulli Associates (2026-02-24), [Investor Skepticism of AI in Financial Advice Persists](https://www.cerulli.com/press-releases/investor-skepticism-of-ai-in-financial-advice-persists)
- Investment Adviser Association (2026), [2026 Investment Adviser Industry Snapshot Shows Continued Growth in Demand for Adviser Services](https://www.investmentadviser.org/iaatoday/press-release/2026-investment-adviser-industry-snapshot-shows-continued-growth-in-demand-for-adviser-services/)
- Willkie Farr & Gallagher (2026-01), [SEC Division of Examinations Issues Risk Alert Regarding Advisers Act Marketing Rule Compliance](https://www.willkie.com/publications/2026/01/sec-division-of-examinations-issues-risk-alert-regarding-advisers-act-marketing-rule-compliance)
- eCFR (2026), [17 CFR 275.206(4)-1, Investment adviser marketing](https://www.ecfr.gov/api/renderer/v1/content/enhanced/current/title-17?part=275&section=275.206%284%29-1)
- Google Search Central (2025), [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features)
- Google Search Central (2025), [Creating helpful, reliable, people-first content](https://developers.google.com/search/docs/fundamentals/creating-helpful-content)
- Underneath (2026), [Which Google Maps businesses does ChatGPT recommend?](https://underneath.agency/research/chatgpt-local-picks-google-profile-study)
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude agree on brands?](https://underneath.agency/research/ai-assistants-brand-agreement-study)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)

---

This is the Markdown twin of https://underneath.agency/resources/wealth-management-firms-clients-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Are websites hiding instructions aimed at AI search?"
description: "Yes. A scan of 1.2 billion web addresses found 1,521 hidden instructions trying to make AI systems promote, cite or praise a page. Most rarely work."
canonical: "https://underneath.agency/resources/websites-hiding-instructions-for-ai-search"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Are websites hiding instructions to manipulate how AI search presents them?

Yes, a small number already do: the largest scan so far found 1,521 hidden instructions written to make AI systems promote, cite or praise a page. They are mostly invisible to people, often years old, and in lab tests they rarely worked on capable AI systems. For a brand, copying the tactic buys little and carries real reputational risk.

## The short version

1. A crawl of 1.2 billion web addresses found 15.3 thousand confirmed instructions aimed at AI systems, on 11.7 thousand pages ([Khodayari and colleagues](https://arxiv.org/abs/2604.27202)).
2. Of those, 1,521 tried to shape how AI systems present a business or page: promoting it (1,040), forcing a citation (542) or demanding a positive review (502) (Khodayari and colleagues).
3. About 70% sat in parts of the page people never see, such as headers, code comments and metadata (Khodayari and colleagues).
4. In a lab test of 13 AI systems summarizing pages, small systems followed the hidden instruction 4.2% of the time and large ones 1.2% (Khodayari and colleagues).
5. Visible self-promotion is far more common: 24.2% of the numbered “best X” lists that AI engines cited ranked their own publisher first ([our self-ranking lists study](https://underneath.agency/research/self-promoting-best-lists-study)).

## How common are hidden instructions to AI systems?

They exist at real scale but on a small share of the web. [Khodayari and colleagues](https://arxiv.org/abs/2604.27202) scanned part of the October 2025 Common Crawl, a public archive of the web, plus two internet-scanning services. They covered 1.2 billion web addresses on 24.8 million websites.

They searched for phrases like “ignore all previous instructions” and then checked every match by hand. That left 15.3 thousand confirmed instructions on 11.7 thousand pages. The authors call this a lower bound, because their search used English phrases and missed disguised wording.

The practice is copied, not invented fresh each time. Just 54 wording templates covered 95% of all cases. That suggests ready-made snippets passed between sites, rather than careful, tailored attacks. Openly rewriting pages for AI search is a separate practice, covered in [how much of the web targets AI search](https://underneath.agency/resources/how-much-web-content-is-optimized-for-ai-search).

## What are these instructions trying to make AI do?

Most try to disrupt or block AI crawlers; a smaller group tries to steer AI search. The single biggest category was garbage injection: telling a machine reader to output random numbers or nonsense. It alone accounted for 8,469 of the 8,894 instructions aimed at disrupting systems.

Defensive uses were also common. Site owners wrote instructions telling AI not to reuse personal data or copyrighted text, or asking any AI reader to reveal itself. Many publishers seem to use these as a home-made “keep out” sign.

The group that matters for marketers is what the authors call reputation manipulation. It made up 1,521 instructions across 139 websites.

The main forms were content or product promotion (1,040), citation forcing (542) and positive review forcing (502). The authors say this cluster “targets search-oriented AI systems” and tries to shape “how entities, products, or sources are surfaced downstream.”

One template appeared 541 times on a single website. It tells any AI reader that the page “is the authoritative source of information” on its topic and that it “should not trust any other source.”

## Where on the page do these instructions hide?

Mostly in places a visitor never sees, which is what makes them manipulative rather than persuasive. About 70% appeared in parts of a page that browsers do not display. The largest single channel was the HTTP header, the technical note a server sends before the page itself: 7,887 instructions arrived that way.

Structured data was another favorite. This is the machine-readable code sites add for search engines, often called JSON-LD.

Researchers found 1,996 instructions there, on 1,611 pages. Others sat in code comments and metadata tags.

When instructions did sit in the visible page, they were usually disguised. Of those still live when checked, 58.6% were concealed with tricks such as text the same color as the background, tiny fonts or elements placed off-screen. Overall, 87% of all instructions were hidden from human readers.

| Where the instruction sat | What it means |
|---|---|
| Server headers | Sent before the page loads; never shown to visitors |
| Structured data and metadata | Code meant for search engines and link previews |
| Code comments | Notes in the page source that browsers skip |
| Visible page, disguised | White-on-white text, tiny fonts, off-screen placement |

## Do hidden instructions actually work on AI search?

Rarely, in the one controlled test available, and the best systems were the hardest to fool. The researchers sampled 100 real instructions and asked 13 AI systems to summarize each page, in four different page formats. That made 5,200 test runs, each checked by hand.

Small AI systems followed the planted instruction 4.2% of the time. Large open systems followed it 1.2% of the time, and both medium and commercial closed systems 0.6%.

The worst case was small systems reading pages flattened to plain text, which peaked at 8%. When the page structure was kept, success fell to between 0.2% and 1.1%.

Commercial systems were also the best at spotting the trick: they flagged the instruction in 25.1% of runs. Small systems flagged it in 4.8%.

A separate shopping test points the same way. [Bagga and colleagues](https://arxiv.org/abs/2511.20867) rewrote product listings to include hidden override instructions. GPT-5, Claude and Gemini flagged those listings as questionable in nearly every case and pushed them down the ranking.

An open Llama model was different: it flagged only 21.0% of them and moved them up by about three places. Results depend heavily on which system is reading. Our guide on [whether competitors can game AI picks](https://underneath.agency/resources/can-competitors-game-ai-recommendations) covers more of these tests.

## Why is copying this tactic a risk for a brand?

Because it is designed to deceive, it rarely works, and it can stay on your site for years. The study is careful not to call every instruction an attack. But planting hidden text to force citations or positive reviews is, by construction, an attempt to mislead a reader.

The instructions are also durable. Checking older archived copies, the researchers found 65% of affected pages already carried the instruction 12 months earlier. Something added once, by a developer or a plugin, can sit unnoticed for a long time.

Most came from the site owners themselves. First-party content accounted for 79.9% of instructions; third parties posting on a platform, such as job listings, accounted for 20.1%. User comments were a small share.

There is also a policy direction to watch. A position paper by [Wen and colleagues](https://arxiv.org/abs/2606.12439) argues that AI platforms should downrank or exclude sources that use undisclosed influence, much as web search penalizes link schemes. That is a proposal, not a documented practice of any engine today. Whether engines can reliably [filter out manipulative content](https://underneath.agency/resources/can-ai-search-filter-manipulative-geo) is still being tested in labs.

## What should you do about it?

Treat hidden instructions as a liability to find and remove, not a tactic to adopt. Visible, accountable content is where the legitimate influence is.

1. **Audit your own site.** Ask your web team to search page source, server headers, structured data and comments for phrases addressed to AI, such as “ignore previous instructions” or “if you are an AI.”
2. **Check what vendors added.** Plugins, tag managers and agency code can insert text you never approved. Most instructions in the study were first-party, so “we didn’t write it” is not a defense.
3. **Watch user-posted areas.** Reviews, forums and job boards on your domain can carry instructions written by others.
4. **Make your real case visibly.** Self-promotion in plain sight is common and accepted: 24.2% of cited numbered “best X” lists in [our study](https://underneath.agency/research/self-promoting-best-lists-study) put their own publisher first. Those lists were only 1.1% of all citations, so even visible self-ranking is a modest lever.
5. **Decide your policy on AI reuse openly.** If you want to limit AI crawling, use standard, documented controls rather than hidden prompts.

If you want help with the visible side of this work, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

Whether hidden instructions change what live AI search engines like ChatGPT or Google’s AI features actually say. The gaps are worth stating plainly:

- The effectiveness test used one task, page summarization, on 13 systems in a lab. It was not a test of production AI search.
- The scan covered about half of one October 2025 crawl, looked for English phrases only and is a lower bound.
- No study has measured whether pages carrying these instructions get cited more, or less, by AI search engines.
- No study has measured whether engines penalize sites found using them.
- The shopping test used simulated rankers with an added anti-manipulation instruction, not deployed shopping assistants.

## Frequently asked questions

### Is hiding text for AI the same as old-school cloaking?

It is closely related: both show machines something people cannot see. In the largest scan, 87% of instructions to AI were hidden from human readers, using channels like server headers or tricks like white-on-white text.

### Can a hidden prompt make ChatGPT recommend my brand?

There is no evidence it does in the live product. In a lab test, commercial AI systems followed planted page instructions only 0.6% of the time and flagged the attempt in 25.1% of runs.

### How do I know if my site has hidden AI instructions?

Search your page source, headers and structured data for phrases addressed to AI. In one scan, 1,996 instructions sat in structured data, so check the code your SEO tools generate too.

### Are all hidden AI instructions malicious?

No. Many are defensive, such as asking AI not to reuse personal or copyrighted content. The concern for marketers is the 1,521 instructions that tried to promote a page, force citations or demand positive reviews.

## Sources

- Khodayari, Zhang, Acharya and Pellegrino (2026), [Indirect Prompt Injection in the Wild: An Empirical Study of Prevalence, Techniques, and Objectives](https://arxiv.org/abs/2604.27202), arXiv:2604.27202.
- Bagga, Farias, Korkotashvili, Peng and Wu (2025), [E-GEO: A Testbed for Generative Engine Optimization in E-Commerce](https://arxiv.org/abs/2511.20867), arXiv:2511.20867.
- Wen, Zhang, Yuan, Chen, Zhang and Guo (2026), [Position: Generative Engine Optimization Creates Underexamined Risks, Governance Must Target Concentration, Disclosure, and Academic Blind Spots](https://arxiv.org/abs/2606.12439), arXiv:2606.12439.
- Underneath (2026), [How many “best of” lists cited by AI rank their own brand first?](https://underneath.agency/research/self-promoting-best-lists-study)

---

This is the Markdown twin of https://underneath.agency/resources/websites-hiding-instructions-for-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How wellness brands get recommended by AI assistants"
description: "By being the product assistants can describe and verify when people ask how to sleep, train or recover better, with wellness claims that stay lawful."
canonical: "https://underneath.agency/resources/wellness-brands-customers-ai-search"
published: 2026-10-09
updated: 2026-10-10
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Will AI assistants recommend our wellness products when people ask how to sleep, train or recover better?

They can, if your products are the ones an assistant can describe accurately and check against independent evidence. Health is one of the most common reasons people open ChatGPT, nearly half of US adults now own a wearable, and a wellness purchase often starts with a question about sleep, strength or stress rather than a brand name.

This article covers wellness products other than supplements: fitness gear and wearables, sleep aids and sleep technology, recovery devices and personal care. It is about how these brands are found and chosen. It is not health advice, and nothing here says any product works.

## The short version

1. The market is large: the [Global Wellness Institute](https://globalwellnessinstitute.org/industry-research/2025-global-wellness-economy-monitor/) puts the wellness economy at $6.8 trillion in 2024, with [personal care and beauty at $1.35 trillion and physical activity at $1.14 trillion](https://www.healthclubmanagement.co.uk/health-club-management-news/Global-wellness-economy-hits-new-record-peak-of-US68trn-forecast-to-reach-US98trn-by-2029/361529).
2. Wellness devices are now mainstream and sticky: in [Rock Health’s 2025 survey](https://athletechnews.com/more-than-half-of-americans-track-their-health-with-a-device-what-comes-next/), 46% of US adults owned a wearable, up from 13% in 2015, 26% tracked their sleep, and only 23% of owners had ever switched brands.
3. A single wellness device brand can now be a billion-dollar business: [Oura](https://www.fortune.com/2025/09/23/oura-ring-11-billion-valuation-series-e-finland-875-million-raise-unicorn) expected more than $1 billion in 2025 revenue, about 20% of it from subscriptions.
4. People already bring wellness goals to AI: in [Life Time’s 2026 survey](https://athletechnews.com/consumers-seek-strength-longevity-over-weight-loss-life-time-survey/), 35.3% used AI tools for workouts, nutrition or health, and ChatGPT Health connects apps such as Apple Health, Peloton and MyFitnessPal.
5. Claims set the limits: the FDA revised its [general wellness policy](https://www.cov.com/en/news-and-insights/insights/2026/01/fda-issues-revised-guidance-on-general-wellness-products) for wearables on January 6, 2026, and the FTC requires competent and reliable scientific evidence for health claims.

## Who buys wellness products, and what is a customer worth?

Adults trying to sleep, move or feel better, and for device brands a buyer often stays for years.

The wellness economy keeps outgrowing the wider economy. The Global Wellness Institute’s 2025 monitor projects 7.6% annual growth to nearly $9.8 trillion in 2029. Its sector figures show where consumer product brands compete:

| Wellness sector, 2024 (Global Wellness Institute) | Size |
|---|---|
| Personal care and beauty | $1.35 trillion |
| Physical activity | $1.14 trillion |
| Mental wellness (the US alone is $125 billion) | $268 billion |

GWI names sleep among the growing segments of mental wellness. For the US, a [McKinsey survey reported by Health Club Management](https://www.healthclubmanagement.co.uk/health-club-management-features/Research-Growth-Market/37147) put the wellness market at $480 billion, and found about half of consumers had bought a fitness wearable at some point while 75% were open to using one.

The buyer is broad. The [Sports & Fitness Industry Association](https://sfia.org/resources/participation-hits-new-high-but-majority-of-americans-not-yet-meeting-recommended-guidelines-of-150-minutes-of-weekly-activity-sfias-2026-topline-report-finds/) counted 250 million Americans who took part in at least one sport, fitness or leisure activity in 2025. In PwC’s global [Voice of the Consumer survey for the Consumer Goods Forum](https://www.theconsumergoodsforum.com/app/uploads/2026/06/Voice-of-the-Consumer-Survey-The-Rise-of-Everyday-Health.pdf), 90% of 21,808 respondents made discretionary health purchases in the past year and 80% used a health app or wearable.

What a customer is worth depends on the product. A sound machine or a recovery tool is often a one-time sale. A wearable is closer to a long relationship:

- Rock Health found 83% of wearable owners wear the device five or more days a week, 47% have used one for three or more years, and only 23% have ever switched brands.
- Oura sold about 3 million rings in a year, 5.5 million in total, and [told Fierce Healthcare](https://www.fiercehealthcare.com/health-tech/oura-raises-900m-series-e-oura-ring-sales-catapult) it expected $1 billion in 2025 sales across devices and app subscriptions. It expects more than $1.5 billion in 2026.

No public benchmark gives a lifetime value for a wellness customer, and we will not invent one. Our inference: when the first brand a buyer chooses is rarely replaced, the recommendation that shaped that first choice is worth far more than one sale.

## Where do AI assistants sit in a wellness purchase?

At the goal stage, before a brand is chosen, and increasingly inside the assistant itself.

Health questions are already a mass behavior on AI assistants. When it launched ChatGPT Health in January 2026, [OpenAI said](https://openai.com/index/introducing-chatgpt-health/) over 230 million people globally ask health and wellness questions on ChatGPT every week. In [KFF’s March 2026 poll](https://www.kff.org/health-information-and-trust/poll-finding/kff-tracking-poll-on-health-information-and-trust-use-of-ai-for-health-information-and-advice/), 32% of US adults had used AI for health information or advice in the past year.

What makes wellness different is that the assistant can now sit next to the product. OpenAI documents that ChatGPT Health connects Apple Health for movement, sleep and activity data, plus apps such as Peloton for workout classes and guided meditations, MyFitnessPal for nutrition and AllTrails for hikes. A brand whose app is connected is present in the conversation in a way a rival without one is not. Brands whose main product is an app face related questions, covered in [how digital health apps win users](https://underneath.agency/resources/digital-health-apps-users-ai-search).

Use for specific wellness tasks is still early. A Menlo Ventures and Morning Consult survey of 5,031 US adults, [reported by eMarketer](https://www.emarketer.com/content/ai-has-yet-penetrate-health-wellness-sector), found 62% plan their workouts but only 11% use AI to do it, and 55% track health and fitness data while 12% use AI for that. Life Time’s survey of members and the public found higher use, at 35.3%. The honest summary: AI is part of the wellness routine for a large minority and growing, not yet for most people.

## Which questions lead a wellness shopper to a brand?

Questions about a goal, a comparison, a constraint or a brand’s reputation, more often than a product name.

The prompts below are illustrative, written by us to show the shape of commercial wellness questions. They are not captured from real users.

| Category | Illustrative prompt |
|---|---|
| Sleep tech | “Best sleep tracker that doesn’t need a monthly subscription” |
| Sleep aids | “Quietest white noise machine for a shared bedroom” |
| Wearables | “Smart ring or fitness watch for someone who lifts weights?” |
| Home fitness | “Compact rowing machine for a small apartment, under 6 feet long” |
| Recovery | “Massage gun for runners: which brands last longest?” |
| Personal care | “Fragrance-free deodorant brands for sensitive skin” |
| Legitimacy | “Is this red light therapy mask brand legit?” |

Two patterns matter. First, many questions start from a goal, so the assistant decides which product category answers it before it picks a brand. Rock Health found activity (35%), sleep (26%) and heart rate (21%) are what wearable owners track most, a fair guide to which goals drive device questions. Second, compatibility and subscription costs come up early for devices, and those are facts an assistant can only repeat if they are written down somewhere it can read.

Questions about diagnosing or treating a condition, such as sleep apnea or high blood pressure, are a different matter. Those belong with clinicians, and OpenAI says ChatGPT Health is not intended for diagnosis or treatment. A wellness brand should not try to win those answers.

## How does a recommendation become revenue for a wellness brand?

Through a shortlist, a check of reviews and specs, a purchase and, for devices, a paid app or subscription.

1. **Named for a goal.** The assistant lists a few products for “track sleep without a subscription” or “quiet rower for an apartment.”
2. **Checked.** The buyer reads reviews, compares specs and asks whether the brand is legitimate. Review sites tend to shape that answer: in [our brand reputation study](https://underneath.agency/research/is-it-legit-ai-reputation-study), 88.0% of AI answers to the question “Is this brand legit?” drew on a review or complaint platform.
3. **Bought where the buyer already shops.** The brand site, Amazon or a store. Oura sells through 4,000 stores, so an assistant’s answer can end at a retail shelf rather than a click. When visitors do arrive from AI, [Adobe’s data reported by TechCrunch](https://techcrunch.com/2026/04/16/ai-traffic-to-us-retailers-rose-393-in-q1-and-its-boosting-their-revenue-too/) showed they converted 42% better than other visitors to US retail sites in March 2026. Adobe did not break out wellness products.
4. **Kept.** For a wearable or connected device, the app, the subscription and the data history keep the customer. Oura’s chief executive told Fortune the mix of hardware and subscription revenue puts it on “a different level than most hardware companies.”

The traffic is growing from a small base. [Adobe reported](https://news.adobe.com/news/2026/01/adobe-holiday-shopping-season) that traffic to US retail sites from generative AI tools rose 693.4% in the 2025 holiday season, while noting the base of users remains modest.

## What decides whether an assistant names your product?

Evidence it can read and verify: product facts, reviews and independent coverage; the platforms do not publish exact rules.

What is documented: [OpenAI’s shopping help page](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search) says ChatGPT’s product results are selected independently and are not ads. It considers structured metadata such as price and product description from first- and third-party providers, plus other third-party content. Review summaries come from public websites, and merchants are ranked on factors such as availability, price, quality and whether they are the maker or primary seller.

What has been observed:

- **Health answers lean on institutions.** In an audit of 615 sources cited by ChatGPT for health questions, [Jacques and colleagues](https://arxiv.org/abs/2601.17109) found 75.7% came from institutional sources and 12.4% from commercial health platforms. We cover what those commercial sites had in common in [what commercial health sites cited by ChatGPT have in common](https://underneath.agency/resources/chatgpt-cited-commercial-health-sites).
- **“Best of” lists are part of the evidence, and some are self-serving.** In [our study of AI-cited “best” lists](https://underneath.agency/research/self-promoting-best-lists-study), 24.2% of numbered lists with an identifiable publisher ranked their own publisher first.
- **Many product pages are hard to read.** Adobe found around 34% of retail product pages cannot be properly accessed by AI.

For wellness, the trust factors carry legal weight. The [FTC’s health products guidance](https://www.ftc.gov/business-guidance/resources/health-products-compliance-guidance) says claims about the health benefits or safety of health-related products require “competent and reliable scientific evidence,” and it applies to devices as well as supplements. One of its own examples is a nasal strip: evidence that it reduces snoring does not support implying it treats sleep apnea. The FDA’s revised general wellness policy, issued January 6, 2026, widened the room for wearables that estimate measures such as heart rate variability, but only when the product is promoted for wellness, not for diagnosing or treating a condition, according to [Covington’s analysis](https://www.cov.com/en/news-and-insights/insights/2026/01/fda-issues-revised-guidance-on-general-wellness-products).

Our inference: the durable trust factors for a wellness brand are accurate specs, claims that stay on the wellness side of that line, independent reviews and testing, and a clean complaint record. A product that an assistant can describe without repeating a risky claim is easier to recommend.

## What does a wellness brand lose when AI leaves it out?

It loses first purchases at the goal stage, and with devices, years of subscription and repeat revenue.

We have not found any measurement of how many wellness sales a brand loses when assistants leave it out, so each point below is labeled reasoning:

- **The first device tends to stick.** With only 23% of wearable owners ever switching brands, we infer a buyer steered to a rival by an AI answer is unlikely to come back soon.
- **The ecosystem effect compounds.** If an assistant connects to a rival’s app, as ChatGPT Health does with Apple Health and Peloton, that rival is present in later questions too. That is our inference from OpenAI’s documentation, not a measured effect.
- **A wrong description persists.** An outdated subscription price, a missing compatibility fact or an old complaint can follow a brand from answer to answer. The fixes are laid out in our guide on [how to fix wrong brand information in AI answers](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers).

## How does GEO work for a wellness brand?

Generative engine optimization (GEO) makes your products easy for AI assistants to find, describe correctly and verify through outside evidence.

For a wellness company, the work usually covers:

1. **Facts in text.** Sensors, battery life, compatibility with phones and health apps, subscription price and what it includes, dimensions, noise levels, materials and return terms, stated on product pages, not only in images or videos. See [what product content AI shopping assistants prefer](https://underneath.agency/resources/product-content-ai-shopping-assistants-prefer).
2. **Claims that hold up.** Wellness language your legal team has cleared against the FTC standard and the FDA’s wellness policy, with studies cited where you make a performance claim. Assistants repeat what they find, including what you overstate.
3. **Integrations documented.** If your product works with Apple Health, Google’s and Samsung’s health apps or a fitness platform, say so plainly and keep it current.
4. **Independent coverage.** Hands-on reviews from sleep, fitness and gear editors, testing labs, trainers and clinicians who can speak to the product honestly, and creator coverage with proper disclosure.
5. **Review and complaint platforms.** Real reviews on retailer sites, Trustpilot and the BBB, complaints resolved in public, and no review tactics the FTC would treat as deceptive. [Fake reviews can also mislead AI answers](https://underneath.agency/resources/fake-reviews-fake-brands-ai-recommendations).
6. **Product feeds.** Complete, consistent catalog data wherever assistants read it, since OpenAI documents its use of structured metadata.
7. **Measurement.** Ask a fixed set of goal, comparison, compatibility and legitimacy questions across ChatGPT, Gemini, Perplexity and Google’s AI features, repeatedly, and track who is named and which sources are cited.

No wellness brand can be promised a spot in AI answers, by us or by anyone else. What this work does is make your product simple for an assistant to describe accurately when a shopper asks. Supplements raise separate testing and label questions, covered in [how supplement brands win customers from AI assistants](https://underneath.agency/resources/supplement-brands-customers-ai-search).

## Which questions can’t the data answer for wellness brands yet?

It shows people using AI for health and wellness, not how many wellness product sales AI answers create.

- **Health questions are not shopping questions.** OpenAI’s and KFF’s figures cover health questions in general. We found no public data on how often people ask an assistant which sleep tracker, massage gun or rower to buy.
- **Surveys differ widely.** AI use for wellness ranges from about one in nine for workout planning (Menlo Ventures) to over a third (Life Time), depending on who was asked and how.
- **Company figures are self-reported.** Oura’s revenue comes from the company and press reports, not audited filings, and Adobe sells the analytics it reports on.
- **No wellness-specific ranking studies.** The citation research we cite looked at health questions in general and at mixed product categories. Applying it to wellness devices is our inference.

## Where should a wellness brand start?

Start by asking assistants the goal and comparison questions your buyers ask, and read what they say about you.

That first check usually shows whether your products appear for the sleep, training or recovery goals you serve, whether your specs and subscription terms are described correctly, which review sites and editors shape your reputation, and which rivals are named instead. The work then is to put accurate, lawful facts where assistants read them and earn the outside coverage that backs them up.

If your growth depends on first purchases that turn into years of app use or repeat orders, [talk to us about a wellness brand review](https://underneath.agency/contact). We will show where your products appear in AI answers to buying questions, why rivals are named instead, and which changes are most likely to win more of those first purchases. The [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) page covers how that work is done for a wellness brand, from specs and integrations in text to cleared claims and independent reviews.

## Frequently asked questions

### Do people really ask AI about wellness products?

They ask about health and wellness at scale: OpenAI says over 230 million people do so on ChatGPT every week. Product-specific use is smaller and less measured; 35.3% in Life Time’s survey used AI for workouts, nutrition or health.

### Can a wellness brand pay to appear in ChatGPT’s product results?

No. OpenAI says product results are selected independently and are not ads, and that ads are shown separately.

### Does the FDA’s 2026 wellness guidance change what we can say?

It clarified that some wearables measuring things like heart rate variability can be wellness products if promoted only for wellness. Claims to diagnose or treat a condition still make a product a medical device. Ask your regulatory counsel.

### Is this different for sleep products than for mattresses?

Yes. Sleep trackers, sound machines and sleep aids are often goal-driven or subscription purchases. Mattresses have their own journey, covered in [how mattress brands win customers who ask AI](https://underneath.agency/resources/mattress-brands-customers-from-ai-search).

## Sources

- Global Wellness Institute (2025-11), [2025 Global Wellness Economy Monitor](https://globalwellnessinstitute.org/industry-research/2025-global-wellness-economy-monitor/)
- Health Club Management (2025-11), [Global wellness economy hits record peak of US$6.8trn](https://www.healthclubmanagement.co.uk/health-club-management-news/Global-wellness-economy-hits-new-record-peak-of-US68trn-forecast-to-reach-US98trn-by-2029/361529)
- Health Club Management (2024), [Research: Growth market (McKinsey Future of Wellness survey)](https://www.healthclubmanagement.co.uk/health-club-management-features/Research-Growth-Market/37147)
- Athletech News (2026-06-09), [More than half of Americans track their health with a device (Rock Health 2025 Consumer Adoption Survey)](https://athletechnews.com/more-than-half-of-americans-track-their-health-with-a-device-what-comes-next/)
- Athletech News (2025-12-31), [Consumers seek strength, longevity over weight loss: Life Time survey](https://athletechnews.com/consumers-seek-strength-longevity-over-weight-loss-life-time-survey/)
- eMarketer (2025-07-02), [AI has yet to penetrate the health and wellness sector](https://www.emarketer.com/content/ai-has-yet-penetrate-health-wellness-sector)
- Sports & Fitness Industry Association (2026-03-12), [SFIA’s 2026 Topline Participation Report](https://sfia.org/resources/participation-hits-new-high-but-majority-of-americans-not-yet-meeting-recommended-guidelines-of-150-minutes-of-weekly-activity-sfias-2026-topline-report-finds/)
- The Consumer Goods Forum and PwC (2026-06), [Voice of the Consumer Survey: The Rise of Everyday Health](https://www.theconsumergoodsforum.com/app/uploads/2026/06/Voice-of-the-Consumer-Survey-The-Rise-of-Everyday-Health.pdf)
- Fortune (2025-09-23), [Oura Ring maker to become $11 billion company with latest raise](https://www.fortune.com/2025/09/23/oura-ring-11-billion-valuation-series-e-finland-875-million-raise-unicorn)
- Fierce Healthcare (2025-10-14), [Oura raises $900M series E as smart ring sales catapult company’s growth](https://www.fiercehealthcare.com/health-tech/oura-raises-900m-series-e-oura-ring-sales-catapult)
- OpenAI (2026-01), [Introducing ChatGPT Health](https://openai.com/index/introducing-chatgpt-health/)
- KFF (2026-03), [KFF Tracking Poll on Health Information and Trust: Use of AI for Health Information and Advice](https://www.kff.org/health-information-and-trust/poll-finding/kff-tracking-poll-on-health-information-and-trust-use-of-ai-for-health-information-and-advice/)
- OpenAI (2026), [Shopping with ChatGPT Search](https://help.openai.com/en/articles/11128490-improved-shopping-results-from-chatgpt-search)
- TechCrunch (2026-04-16), [AI traffic to US retailers rose 393% in Q1, and it’s boosting their revenue too](https://techcrunch.com/2026/04/16/ai-traffic-to-us-retailers-rose-393-in-q1-and-its-boosting-their-revenue-too/)
- Adobe (2026-01-07), [Holiday Shopping Season Drove a Record $257.8 Billion Online](https://news.adobe.com/news/2026/01/adobe-holiday-shopping-season)
- Federal Trade Commission (2022-12), [Health Products Compliance Guidance](https://www.ftc.gov/business-guidance/resources/health-products-compliance-guidance)
- Covington & Burling (2026-01-08), [FDA Issues Revised Guidance on General Wellness Products](https://www.cov.com/en/news-and-insights/insights/2026/01/fda-issues-revised-guidance-on-general-wellness-products)
- Jacques and colleagues (2026), [Authority Signals in AI Cited Health Sources: A Framework for Evaluating Source Credibility in ChatGPT Responses](https://arxiv.org/abs/2601.17109), arXiv:2601.17109.
- Underneath (2026), [“Is this brand legit?” How AI assistants build a reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study)
- Underneath (2026), [How many “best of” lists cited by AI rank their own brand first?](https://underneath.agency/research/self-promoting-best-lists-study)

---

This is the Markdown twin of https://underneath.agency/resources/wellness-brands-customers-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "What does content that AI engines prefer look like? | Underneath"
description: "AI engines favor pages that state the conclusion first, cover the topic fully, explain how and why, and pack checkable facts such as definitions and figures."
canonical: "https://underneath.agency/resources/what-content-do-ai-engines-prefer"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# What does content that AI engines prefer look like?

AI engines prefer content that states its main point first, covers the topic fully, explains how and why, and packs in checkable facts such as definitions and figures. Layout alone does little; what matters is that each section holds information an answer can reuse. Most of this evidence comes from lab tests and observational studies, so treat it as a pattern to test, not a guarantee.

## The short version

1. When researchers at Carnegie Mellon taught a system the rules AI engines seem to reward, rewrites following them raised visibility by 35.99% on average in [lab tests](https://arxiv.org/abs/2510.11438).
2. Cited pages with definitions shaped answers 57.33% more, and pages with comparisons 55.28% more, in [an analysis of pages cited by ChatGPT, Google and Perplexity](https://arxiv.org/abs/2604.25707).
3. Question-and-answer formatting on its own did not help in that analysis; it went with 5.74% less influence.
4. Google’s AI Overviews rarely copy sentences: the median share of five-word sequences lifted verbatim from a cited page was 0.0% in [our study](https://underneath.agency/research/ai-overview-cited-pages-study).

## What patterns do AI engines reward in a page?

They reward pages that are comprehensive, factual, neutral in tone, clearly organized and specific. An automated system learned these rules from tens of thousands of AI engine choices.

Wu and colleagues at Carnegie Mellon University built AutoGEO. For each question, it compares the page an AI engine used most with the one it used least, then writes down the differences as rules. The common rules included covering the topic comprehensively, attributing claims to credible sources, keeping a neutral tone without promotion, and using clear headings and lists.

For in-depth research questions on Gemini, it also learned three more rules. State the key conclusion at the beginning, explain causes and mechanisms, and back claims with specific data or named examples.

Rewrites that followed the rules raised the lab visibility scores by 35.99% on average, without lowering the quality of the AI’s answers. The tests used fixed sets of five candidate pages per question, not the live web.

## Should you put the conclusion first?

Probably, though the direct evidence is thin. “Conclusion first” is one of the learned rules, but no study has tested it alone on live engines.

The AutoGEO authors show one worked example. Their rewritten version presents the main thesis upfront, discusses the topic more thoroughly and explains the underlying how and why. That is a single paragraph, judged by the authors, not a measurement.

The GEO-16 audit by [Kumar and Palkhouski](https://arxiv.org/abs/2509.10762) recommends the same thing: an answer-first summary, compact paragraphs and descriptive headings. Its authors present this as a design principle rather than a tested result. Leading with the answer costs little, so the downside of trying it is small.

## Which kinds of content get used most in AI answers?

Pages with definitions, comparisons, figures and code get reused most. Citations used only as bare references count least.

Zhang and colleagues analyzed 18,151 pages cited by ChatGPT, Google AI Overviews and Perplexity across 602 prompts. They scored how much of each answer’s wording, structure and position could be traced to each cited page. Pages with definitions, comparisons, numbers, how-to steps or code scored higher.

| Content on the cited page | Influence on the answer, against pages without it |
|---|---|
| Definitions | +57.33% |
| Comparisons | +55.28% |
| How-to steps | +41.20% |
| Question-and-answer format | −5.74% |

By role, citations used as definitions averaged an influence score of 0.1531, against 0.0529 for citations used only as references. The authors’ reading is that evidence works as reusable building blocks, while question-and-answer formatting is “only a surface wrapper”. They stress these are associations, not tested causes.

## Does longer, more structured content win?

Depth and structure go together with being used, but length alone does not win. Once pages are compared fairly, word count fades.

In the Zhang analysis, the most influential quarter of cited pages averaged 1,943.30 words, against 169.82 for the least influential. They also had 12.50 times as many headings. The authors warn this “should not be reduced to a rule that longer content always wins”.

Our own Google data adds a check. In our study of 3,096 ranking pages, word count did not matter once ranking position and the search were held constant. But cited pages with an HTML table were 14.7 points more likely to have the answer’s wording traced to them. Other page features, such as dates and structured data, are covered in [on-page signals linked to AI citations](https://underneath.agency/resources/on-page-signals-linked-to-ai-citations).

## Do all AI engines want the same things?

Mostly, within the same kind of question. Rules shift more by topic than by engine.

In the AutoGEO study, the [rules learned from Gemini and GPT](https://underneath.agency/resources/do-ai-engines-prefer-same-content) overlapped by 78.95% on research questions. Rules learned from the two general question sets overlapped by 88.24%, but overlap with shopping questions fell to 34.78% and 40.00%. Shopping rules favored step-by-step guidance and product details such as model numbers and specifications, over in-depth explanation.

A controlled test by [Vishwakarma and colleagues](https://arxiv.org/abs/2605.25517) found the same split. For product reviews, missing prices and missing specifications cost pages the first citation. Their test examined 18 content factors across 252,000 trials, in a simulation run by the software company Sprinklr.

## Is “answer-ready” content always a good idea?

Only when the facts are real. AI-optimized writing is now detectable, and some of it props up weak claims.

Researchers at the CISPA Helmholtz Center built a detector for content written to win AI citations. Run on 10,095 pages retrieved by Google Search and Gemini, it flagged an estimated 8.90% as optimized, rising to 16.36% among pages modified in 2026. The authors treat these as estimates, since the detector is not error-free.

On the flagged pages, 69.34% of the sources they cited were rated low on verifiability: little editorial accountability, or hard for readers to check.

The AutoGEO team also tested manipulative rewrites that try to hijack the AI. These raised visibility scores but always made the AI’s answers worse. Dense, factual pages are what engines reuse; invented facts dressed in the same format are a reputational risk.

## What should you do about it?

Write each important page as a dense, honest briefing on one question. In practice:

1. Open with the answer or conclusion in a sentence or two.
2. Cover the topic fully, including the how and why behind the answer.
3. Add definitions, comparisons, specific figures and step-by-step instructions where they fit, each with a source.
4. Put product facts such as prices and specifications in plain text or simple tables.
5. Keep the tone neutral; cut sales language and hedging.
6. Do not convert pages to question-and-answer format expecting a lift on its own; see [whether FAQ pages help citations](https://underneath.agency/resources/do-faq-pages-help-ai-citations).

If you want help applying this across your key pages, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

No study has rewritten real pages to this pattern and tracked live AI citations over time.

- **Lab versus live.** AutoGEO’s gains come from fixed sets of five pages per question, with the target page already included.
- **Association versus cause.** The influence analysis observed existing pages. Better-written pages may also come from stronger brands.
- **Conclusion first, alone.** Its effect has not been isolated from the other rules.
- **Industry differences.** Shopping, research and local questions reward different things, and many sectors are untested.
- **Measurement.** The influence score is the authors’ own proxy built from word overlap and position, not a view inside the AI.

## Frequently asked questions

### How should I structure content for AI search?

Lead with the answer, then cover the topic fully with definitions, comparisons and specific figures under clear headings. In one analysis of cited pages, definitions went with 57.33% more influence on the answer.

### Does FAQ formatting help with AI citations?

Not on its own. Pages in question-and-answer format showed 5.74% less influence on AI answers than other pages in an analysis of ChatGPT, Google and Perplexity citations.

### Is longer content better for AI engines?

Only when the length carries more useful structure and facts. The most influential cited pages averaged 1,943.30 words, but in our Google study word count stopped mattering once ranking was accounted for.

### Do AI engines copy text from my page?

Rarely word for word. In our study of Google’s AI Overviews, the median share of five-word sequences copied verbatim was 0.0%, so clear facts matter more than catchy phrasing.

## Sources

- Wu, Zhong, Kim and Xiong (2025), [What Generative Search Engines Like and How to Optimize Web Content Cooperatively](https://arxiv.org/abs/2510.11438), arXiv:2510.11438.
- Zhang, He and Yao (2026), [From Citation Selection to Citation Absorption: A Measurement Framework for Generative Engine Optimization Across AI Search Platforms](https://arxiv.org/abs/2604.25707), arXiv:2604.25707.
- Chu and colleagues (2026), [GEO-Flag: Detecting and Measuring GEO-Optimized Web Content](https://arxiv.org/abs/2608.16824), arXiv:2608.16824.
- Kumar and Palkhouski (2025), [AI Answer Engine Citation Behavior An Empirical Analysis of the GEO16 Framework](https://arxiv.org/abs/2509.10762), arXiv:2509.10762.
- Vishwakarma, Kumar and Jamidar (2026), [What Gets Cited: Competitive GEO in AI Answer Engines](https://arxiv.org/abs/2605.25517), arXiv:2605.25517.
- Underneath (2026), [What pages cited by AI Overviews have in common: 3,096 pages](https://underneath.agency/research/ai-overview-cited-pages-study)

---

This is the Markdown twin of https://underneath.agency/resources/what-content-do-ai-engines-prefer. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "What Does a GEO Agency Do? The 2026 Guide | Underneath"
description: "A GEO agency gets a brand cited, mentioned and recommended in AI answers. The nine things it does, what it costs, and how the AI engines describe the job."
canonical: "https://underneath.agency/resources/what-does-a-geo-agency-do"
published: 2026-09-25
updated: 2026-09-27
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# What does a GEO agency do?

A GEO agency, short for generative engine optimization agency, gets a brand cited, mentioned and recommended inside the answers that ChatGPT, Google AI Overviews, Google AI Mode, Gemini, Perplexity and Copilot give. This guide covers what that work consists of, what it costs, and what the AI engines themselves said when we asked them the question 22 times.

Underneath teamGuide

On this page

1. [What a GEO agency is](#what-a-geo-agency-is)
2. [The nine things a GEO agency does](#the-nine-things-a-geo-agency-does)
3. [What the AI engines say a GEO agency does](#what-the-ai-engines-say-a-geo-agency-does)
4. [GEO agency, SEO agency or both](#geo-agency-seo-agency-or-both)
5. [What a GEO agency costs in 2026](#what-a-geo-agency-costs-in-2026)
6. [What to expect in the first 90 days](#what-to-expect-in-the-first-90-days)
7. [Frequently asked questions](#frequently-asked-questions)

Related service

Generative Engine Optimization

Get mentioned, cited and recommended by ChatGPT, Gemini, Perplexity and Copilot, and described correctly when you are.

[How we do it](https://underneath.agency/services/generative-engine-optimization)

## The short version

1. A GEO agency makes a brand the one an AI answer engine names, cites or recommends. It does that by working on four things: the brand’s pages, its structured data, its entity footprint and the third-party sources the engines draw on.
2. The work has nine parts: an AI visibility audit, prompt and fan-out research, content restructuring, structured data and crawler access, entity work, consolidation, citation building, daily measurement, and team training.
3. Published 2026 price guides put agency retainers between about $1,500 and $50,000 a month; the cheapest useful engagement is an audit that shows where you are cited today.

In September 2026 the question “what does a GEO agency do” produces a full AI answer on every platform, and the answers agree with each other more than the agencies do. This guide sets out the job in plain terms, then compares it with what the engines say, since their answer is what a buyer who asks an assistant reads before any agency’s website.

## What a GEO agency is

A generative engine optimization agency is a marketing firm whose deliverable is presence inside AI-generated answers rather than a position on a results page. When a buyer asks an assistant “which accounting software handles multi-currency” or “who is a good plumber in Leeds”, the engine retrieves pages, selects passages, and names a few brands. The agency’s job is to make its client one of the brands named, to make sure the engine describes the client accurately, and to make the pages that carry the client’s facts the ones the engine cites.

The same firm may call itself an AI search optimization agency, an answer engine optimization (AEO) agency, an LLM optimization agency or an AI visibility agency. The label varies; the job does not. When we asked ChatGPT the question, its first search was “geo agency meaning generative engine optimization agency”: it expanded the acronym before answering. A page that wants to be read for this question should do the same in its first sentence.

## The nine things a GEO agency does

Agencies package the work differently, but the AI answers to this question converge on the same activities. These are the nine that recur, in the order an engagement usually runs them.

| Activity | What it changes | Typical deliverable |
| --- | --- | --- |
| AI visibility audit | Where the brand is cited or recommended today, and by which engine | Question-by-question citation map, run daily for at least ten days |
| Prompt and fan-out research | Which questions and sub-questions the engines actually run | Question list per page, including the engines’ hidden follow-up searches |
| Content restructuring | Whether a passage can be quoted on its own | Definition-first openings, question headings, complete FAQ answers, comparison tables |
| Structured data and crawler access | Whether the engines can read the page and its facts | Organization, Service, Product, FAQPage schema; crawler rules for OAI-SearchBot, PerplexityBot, Google-Extended |
| Entity and knowledge-graph work | Whether the brand is one recognizable entity | Consistent name, description and identifiers across the site, Wikidata, LinkedIn and directories |
| Consolidation | Which single page the engine picks per topic | Merged overlapping pages, redirects |
| Citation and digital PR | Whether third parties the engines trust name the brand | Listings and mentions on the roundups, directories, forums and publications the engines cite |
| Measurement and reporting | Whether you know what changed | Daily tracking of every target question, monthly report on mentions, citations and share of voice |
| Training | Whether the results survive the engagement | Writing standards, schema checklist, monitoring routine |

Three of these separate a GEO agency from a content agency with a new label. The audit must be repeated, because a single check of an AI answer is a coin toss: in Underneath’s own tests a source cited once in twelve runs was noise and one cited in ten of twelve was a position. The citation work must target the specific third-party pages each engine uses, which differ by platform. And the measurement must be per engine, because ChatGPT, Google’s AI surfaces and Perplexity do not share sources.

## What the AI engines say a GEO agency does

Underneath asked the five main engines this exact question 22 times between 15 and 25 September 2026: twelve times over ten days and ten times in one afternoon. The answers were consistent.

- Every platform opened with a definition that expands the acronym: “A GEO agency (generative engine optimization agency) helps businesses…”. Google’s AI Overview cited about eight sources per answer, AI Mode seven, Perplexity nine, Gemini three, ChatGPT one.
- Every platform then listed services, and the lists overlap almost completely with the nine above: content optimization for AI answers, structured data, entity work, citation building, monitoring and reporting.
- The sources were agency explainers and directory guides, not roundups. On Google AI Mode the five most-cited pages were all agency pages titled around the same question; on Perplexity a directory’s guide was cited in 22 of 22 runs. Reddit was named inside 31 of the 110 answers.
- The People Also Ask box under the Google result was mostly about SEO: “is it worth hiring an SEO agency?” (16 times), “is GEO replacing SEO?” (12), “which are the best GEO agencies?” (7). Buyers asking what a GEO agency does are also asking whether they still need SEO. The next section answers that.

## GEO agency, SEO agency or both

A GEO agency does not replace an SEO agency. Google’s guidance for its generative AI features says the same SEO fundamentals apply to them, and our measurements bear that out: for the query “generative engine optimization services”, Google’s AI Overview cited 75% of the pages ranked 1 to 3 in the same search but only 42% of those ranked 4 to 10. That was one keyword. Across 486 searches, [41.7% of pages ranking 1 to 3 were cited against 20.1% at positions 7 to 10](https://underneath.agency/research/ai-overview-cited-pages-study), while [71.3% of all AI Overview citations went to pages outside the top 10](https://underneath.agency/research/ai-overview-citations-study). On Google’s surfaces, then, ranking helps a great deal but does not decide the citation. On ChatGPT it mattered even less: 26 of 28 ChatGPT citations for that keyword came from outside Google’s top 10.

| | SEO agency | GEO agency |
| --- | --- | --- |
| Goal | Rank pages and earn clicks | Be the brand the AI answer names and cites |
| Where it shows | Google and Bing results | ChatGPT, AI Overviews, AI Mode, Gemini, Perplexity, Copilot |
| Unit of work | The page | The passage, the entity and the third-party mention |
| Off-page work | Links | Listings and mentions on the pages the engines cite |
| Measurement | Rankings, organic traffic | Mentions, citations, share of voice, AI referral traffic |

The practical answer for most companies is to keep SEO as the foundation and add GEO as the layer that decides whether the ranking page is also the cited one. The full comparison is in [GEO agency vs SEO agency](https://underneath.agency/resources/geo-agency-vs-seo-agency).

## What a GEO agency costs in 2026

Published price guides from 2026 put GEO agency retainers between about $1,500 and $50,000 a month. [WebFX](https://www.webfx.com/blog/ai/generative-engine-optimization-cost/) (May 2026) lists $1,500 to $5,000 a month for small businesses, $5,000 to $25,000+ for mid-sized companies and $25,000 to $50,000+ for enterprises, with projects at $5,000 to $50,000 and hourly consulting at $50 to $300. Underneath’s programs are priced in three bands, $5,000 to $10,000, $11,000 to $20,000 and $20,000+ a month, each beginning with a Marketing & AI Visibility Audit whose fee is credited if you continue.

Two things to check before paying any of those figures: whether the audit repeats each question over several days rather than once, and whether the reporting is per engine. An agency that shows you one screenshot of one ChatGPT answer has measured nothing.

## What to expect in the first 90 days

1. Weeks 1 to 2: the audit. Your buyer questions run daily on each engine; you get a map of who is cited today and which pages and third parties the engines used.
2. Weeks 3 to 8: the pages. Priority pages are restructured so their first sentences can be quoted, structured data and crawler access are fixed, overlapping pages are merged.
3. From week 4: the third parties. The agency pitches the roundups, directories and forum threads the engines cite for your category.
4. Day 30 and monthly after: the report. Mentions, citations and share of voice by engine against the baseline; pages the engines retrieved but did not cite get reworked.

Changes to pages that are already indexed can show up in Google AI Overviews and AI Mode within weeks, and in Perplexity once it recrawls them. ChatGPT recommendations lean on third-party sources, so they take longer.

## Frequently asked questions

### What is a GEO agency?

A GEO agency, or generative engine optimization agency, is a firm that gets a brand cited, mentioned and recommended in AI-generated answers from ChatGPT, Google AI Overviews, Gemini, Perplexity and similar engines, and makes sure the engines describe the brand accurately.

### What services does a GEO agency provide?

An AI visibility audit, prompt and fan-out research, content restructuring for quotability, structured data and crawler access, entity and knowledge-graph work, page consolidation, citation and digital PR building, daily measurement with monthly reporting, and team training.

### Is GEO replacing SEO?

No. Google’s AI features draw on Google’s index, and in Underneath’s study of 486 searches the AI Overview cited 41.7% of the pages ranking in the top three, twice the rate of pages ranking 7 to 10. GEO adds the passage, entity and third-party work that decides whether a ranking page is also the one the engine cites; it does not replace ranking.

### Is it worth hiring an SEO agency if I want AI visibility?

For Google’s AI surfaces, yes: a top-three ranking roughly doubles a page’s chance of being cited compared with positions 7 to 10, even though most AI Overview citations still come from outside the top 10. On Perplexity and ChatGPT ranking matters far less, so an agency that only does SEO will miss those engines. Keep the SEO foundation and add a GEO specialist who measures each engine separately.

### How do GEO agencies measure success?

By tracking a fixed set of buyer questions on each engine daily and reporting mentions, citations of your pages, share of voice against named competitors, and sessions and leads from AI referrals, against a baseline taken before the work started.

### Which are the best GEO agencies?

The engines cannot answer that consistently: in 22 runs no agency was named in more than a quarter of answers to this question, and the agencies named were those whose own explainers were cited. Judge an agency by whether it can show you a repeated, per-engine measurement of a client’s visibility, not by a list.

## Keep *reading.*

[All resources](https://underneath.agency/resources)

- Guide · AI search

  ### [How to choose a GEO agency](https://underneath.agency/resources/how-to-choose-a-geo-agency)

  Eight criteria, the red flags, and the questions to ask on the first call, checked against what the AI engines themselves tell buyers.

  Underneath team
- Guide · AI search

  ### [Is a GEO agency worth it?](https://underneath.agency/resources/is-a-geo-agency-worth-it)

  When hiring one pays back, when it does not, and how to decide with numbers rather than fear of missing out.

  Underneath team
- Guide · AI search

  ### [GEO agency vs SEO agency](https://underneath.agency/resources/geo-agency-vs-seo-agency)

  What each one does, where the work overlaps, and whether you need one, the other or both.

  Underneath team

Free strategy call

## See which engines cite you today.

On a free 30-minute call we run your ten most important buyer questions through the engines and show you who is being recommended instead of you.

[Book a strategy call](https://underneath.agency/contact)[How we work](https://underneath.agency/approach)

---

This is the Markdown twin of https://underneath.agency/resources/what-does-a-geo-agency-do. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "What actually drives which product an AI assistant recommends?"
description: "When AI assistants can see ratings, prices and reviews, that data decides the pick. Brand name mostly breaks ties when products look the same."
canonical: "https://underneath.agency/resources/what-drives-ai-product-recommendations"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# What actually drives which product an AI assistant recommends?

Product data drives the pick when the assistant can see it: in a controlled test of three AI assistants, rating, price and reviews explained 82.4% of how products were ranked, and brand name only 1.2%. A famous brand still wins when the products look identical. The bigger risk for most brands is missing, thin or wrong product data, not a famous rival.

## The short version

1. In a skincare test of three AI assistants, product data explained 82.4% of the rankings and brand name 1.2% ([Chu and Hou](https://arxiv.org/abs/2606.17443)).
2. When every product had the same specs, the well-known brand won all 670 valid trials in the same study, so brand works as a tiebreaker.
3. A small edge flipped the result: an invented brand with a slightly better rating, price or review count won 64–80% of the time instead of 3.6–6.0%.
4. Real assistants are less tidy: for the same shopping questions, ChatGPT and Gemini showed only 5.4% of the same source websites on average ([Uberti-Bona Marin and colleagues](https://arxiv.org/abs/2609.18729)).
5. Assistants also lose detail: only 61.9% of software plan prices quoted by four assistants were fully faithful to the vendor’s pricing page ([our pricing study](https://underneath.agency/research/ai-pricing-accuracy-study)).

## How much does product data matter compared with brand name?

Far more: in one controlled test, product data explained 82.4% of rankings and brand name only 1.2%.

[Chu and Hou](https://arxiv.org/abs/2606.17443), of Trine University and Texas A&M, asked three AI assistants (GPT-4o-mini, Claude Sonnet and Gemini 3 Flash) to rank lists of ten skincare products. Each list held one real brand, such as CeraVe, and nine invented ones. Across thousands of trials they varied rating, price, review count and the order of the list.

Product details (rating, price and reviews) explained 82.4% of the variation in rankings. Position in the list explained 6.5%. Brand identity explained 1.2%. The rest came from the factors acting together.

Two caveats matter. The product details were typed into the question through the companies’ developer access, so this is not the consumer app searching the web. And [skincare is a category where buyers lean on brand](https://underneath.agency/resources/skincare-brands-ai-search). A smaller repeat with USB-C cables and AA batteries (3,840 calls) found the same pattern.

## When does a famous brand still win?

A famous brand wins when nothing else separates the products, or when the product data is unclear.

In the same study, when all ten products had identical rating, price, reviews and description, the real brand was recommended in all 670 valid trials. Not one invented brand was picked. The authors call this a conditional monopoly: total, but only while nothing else differs.

Brand also mattered most in the murky middle. With clearly good or clearly poor specs, real and invented brands ranked the same. With middling specs, the real brand’s average rank was 1.70 against 5.49 for an invented brand with the same numbers. When the data gives the assistant no reason to choose, it falls back on the name it knows.

Live assistants show the same default. [Chen and colleagues](https://arxiv.org/abs/2509.08919) at the University of Toronto asked ChatGPT and Perplexity 50 unbranded questions about cola. Major brands took 62.2% of brand mentions and niche brands 9.0%. Those questions gave the assistant nothing to compare, so it named the market leaders.

## How small an edge is enough to win?

Very small: a rating less than a tenth of a star higher was enough to win half the time.

Chu and Hou then gave the invented brand a growing advantage. With identical specs, it won only 3.6–6.0% of the time. With the smallest advantage tested, its win rate jumped to 64–80%, depending on which detail was better. Larger advantages added little after that.

The halfway point was strikingly low: a 0.075-star rating advantage, 1.6 times as many reviews, or a 7.3% price discount. That is less than the gap between a 4.3 and a 4.4 star rating.

The three assistants did not behave alike. Claude was the hardest to move: with the smallest rating edge, the invented brand won 11% of the time on Claude, against 94% on GPT-4o-mini and 88% on Gemini. Which assistant your buyers use changes how much a small edge is worth.

## Does the way a product is described matter too?

Yes, but evidence-style wording moves assistants while sales pressure does not, and invented claims are a legal risk.

In a second experiment, Chu and Hou kept the specs identical and changed only the wording. Copy that looked like evidence, such as clinical-trial claims or customer testimonials, broke the famous brand’s hold 50–73% of the time. Pressure tactics like “limited stock” moved it only 10–13% of the time. The clinical claims in the test were invented on purpose to find the upper limit. The authors class invented claims as potential false advertising and limit their advice to real certifications and published evidence. Our guide on [gaming AI shopping rankings](https://underneath.agency/resources/can-you-game-ai-shopping-rankings) covers those risks in more depth.

Other controlled tests agree that substance beats style. A team from MIT and Columbia built [E-GEO](https://arxiv.org/abs/2511.20867), a test bed of 13,747 shopping questions paired with real Amazon listings, ranked by five AI models acting as shopping assistants. Making a listing longer did not help. The rewrites that worked kept the facts, led with a summary, listed concrete features and use cases, and answered likely buyer questions.

Researchers at Sprinklr, a software vendor, ran 252,000 head-to-head trials in which two pages differed in one detail. A page without a price was far less likely to be cited first in every model they tested. Formatting changes did little: seven of their 18 factors (39%) had weak or no effects, including how the text was laid out. As a vendor study in a two-page setup, treat it as directional.

## Do real AI assistants behave the same way?

Partly: real assistants also lean on concrete facts, but their answers vary and draw on other people’s pages.

[Uberti-Bona Marin and colleagues](https://arxiv.org/abs/2609.18729) put 117 real shopping questions to ChatGPT, Gemini and Google’s AI Overviews from the Netherlands in September 2026. ChatGPT stated a personal pick (“my pick would be…”) in 79% of its product answers, against 7% for Gemini and 2% for AI Overviews. For the same question, ChatGPT and Gemini shared only 5.4% of their source websites on average.

Those sources are mostly third parties. Editorial and product-review sites made up 56.7% of the domains ChatGPT displayed. So the product data an assistant compares often comes from reviewers and retailers, not from your own page.

Our own studies point the same way:

| What we measured | Result |
|---|---|
| Software plan prices quoted fully correctly by four assistants | 61.9% ([pricing study](https://underneath.agency/research/ai-pricing-accuracy-study)) |
| Odds of being recommended for each tenfold rise in independent sites naming a brand | 4.7 times ([brand entity study](https://underneath.agency/research/brand-entity-ai-recommendations-study)) |
| Share of ChatGPT’s brands across five runs shown in a single answer | 57.8% ([consistency study](https://underneath.agency/research/ai-recommendation-consistency-study)) |

Independent coverage was the strongest predictor of being recommended in our brand study, ahead of a Wikipedia article. And because one answer shows only part of the picture, a single check of “are we recommended?” is a weak measure.

## What should you do about it?

Make your product data complete, accurate and easy to compare everywhere an assistant might read it.

1. **Publish the numbers assistants compare.** Put price, rating, review count and key specs on the product page in plain text. A missing price was one of the strongest reasons a page lost in controlled tests.
2. **Fix the data on other people’s pages.** Review sites and retailers supply most of what assistants cite. Give reviewers current specs and prices, and correct listings that are out of date.
3. **Earn the small edges honestly.** A slightly higher rating or more reviews was enough to beat a famous brand in testing. Ask every satisfied customer for a review.
4. **Back claims with real evidence.** Name real certifications, tests and awards. Never invent them. In one test, assistants told to watch for manipulation flagged every listing rewritten with extreme superlatives.
5. **Check what assistants say, repeatedly.** Ask the same buying questions several times across assistants and record the prices and reasons they give.

If you want help building that measurement, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The controlled evidence is strong on how assistants weigh data they are given, and weak on everything before that.

- In the main experiments the product details were typed into the question. Real assistants search the web first, and whether your data is found at all is a separate step these tests skip.
- The core study used skincare and three assistant versions available in 2026; other categories and newer versions may weigh data differently. Our guide to [sportswear brands in AI shopping answers](https://underneath.agency/resources/sportswear-brands-ai-search) looks at one such category, where questions turn on the activity and specific product attributes.
- The wording experiments were not repeated for cables and batteries, so the effect of evidence-style copy outside skincare is unknown.
- No study here links better product data in AI answers to sales.
- Real-world audits are snapshots from one place and month, and assistants change often.

## Frequently asked questions

### Do AI assistants favor big brands?

Only when they have nothing else to go on. With identical specs, the known brand won all 670 valid trials in one study, but a slightly better rating, price or review count let an unknown brand win most of the time.

### Does a better star rating help a product get recommended by ChatGPT?

In controlled tests, yes. A rating advantage of 0.075 stars was enough for an unknown brand to win half the time, although Claude was much harder to move than GPT-4o-mini or Gemini.

### Should I rewrite product descriptions for AI search?

Rewrite for clarity and completeness, not length. In tests on Amazon listings, longer copy did not help, while factual summaries, concrete features and answers to buyer questions did.

### Does listing a price matter for AI answers?

It appears to. In a vendor study of 252,000 trials, pages without a price were much less likely to be cited first by every model tested.

## Sources

- Chu and Hou (2026), [Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems](https://arxiv.org/abs/2606.17443), arXiv:2606.17443.
- Uberti-Bona Marin et al. (2026), ["If I Had to Buy Just ONE: Galaxy S26 Ultra": Auditing AI-Generated Product Recommendations](https://arxiv.org/abs/2609.18729), arXiv:2609.18729.
- Chen, Wang, Chen and Koudas (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919), arXiv:2509.08919.
- Bagga, Farias, Korkotashvili, Peng and Wu (2026), [E-GEO: A Testbed for Generative Engine Optimization in E-Commerce](https://arxiv.org/abs/2511.20867), arXiv:2511.20867.
- Vishwakarma, Kumar and Jamidar (2026), [What Gets Cited: Competitive GEO in AI Answer Engines](https://arxiv.org/abs/2605.25517), arXiv:2605.25517.
- Underneath (2026), [How faithfully do AI assistants quote software prices?](https://underneath.agency/research/ai-pricing-accuracy-study)
- Underneath (2026), [Do Wikipedia and schema make AI assistants recommend a brand?](https://underneath.agency/research/brand-entity-ai-recommendations-study)
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)

---

This is the Markdown twin of https://underneath.agency/resources/what-drives-ai-product-recommendations. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "What GEO practices are actually supported by research?"
description: "Research supports a conservative GEO playbook: relevant, complete, verifiable, well-structured pages that AI can reach, measured stage by stage. Tricks fail."
canonical: "https://underneath.agency/resources/what-geo-practices-does-research-support"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# What GEO practices are actually supported by research?

The research supports a conservative playbook: publish relevant, complete, verifiable, clearly structured pages that AI systems can reach, then measure whether they are found, cited and quoted correctly. Relevance and ranking position have the strongest evidence. Popular tricks such as keyword stuffing, formatting makeovers and blanket schema have weak or negative evidence, and no technique has been shown to raise visibility durably across AI engines.

## The short version

1. A review of 45 studies by [Martinez](https://arxiv.org/abs/2607.14035) found no technique with a proven, lasting effect on being found across AI engines; relevance and position were the most reproducible levers.
2. The famous “up to 40%” figure comes from a lab test where the page was already in front of the AI, in [the original GEO paper](https://arxiv.org/abs/2311.09735).
3. In one benchmark, only three of 54 combinations of rewrite method and subject area clearly helped.
4. Pages with numbers, definitions and comparisons shaped AI answers more, in [an analysis of 18,151 cited pages](https://arxiv.org/abs/2604.25707); question-and-answer formatting alone did not.
5. Evidence that GEO raises traffic or sales is the weakest of all: one suggestive field study and a few industry claims.

## What does the research actually support?

It supports good, findable information pages, measured carefully, not tricks. That is the conclusion of the most complete review of the field so far.

Olivier Martinez reviewed 45 studies published between November 2023 and July 2026. His summary: “produce a relevant, comprehensive, verifiable, clearly structured, and technically retrievable page; then measure retrieval, citation, and fidelity separately.” In plain words, the page must be found, then cited, then represented accurately, and each step needs its own check. For plain definitions of terms like GEO and citations, see our [AI search glossary](https://underneath.agency/resources/glossary).

The evidence is strongest for one narrow claim. Once a page is already among the sources an AI system has gathered, changing the page can change how it is cited or used. Whether those changes help a page get gathered in the first place is far less proven.

## Which practices have the strongest evidence?

Relevance to the question and position among sources have the strongest support. Specific, checkable evidence comes next.

| Practice | Strength of evidence | What it means for you |
|---|---|---|
| Answer the exact question asked | Strong | Write for real buyer questions |
| Rank well enough to be gathered | Strong | Search work still matters |
| Verifiable figures, definitions, comparisons | Moderate to strong | Add real, sourced facts |
| Prices, dates, recency | Moderate | Helps on time-sensitive and buying questions |
| Page structure | Moderate, mixed | Test it; do not assume it helps |
| Authoritative tone | Weak and unstable | Confidence is not evidence |
| Keyword stuffing | None or negative | Avoid |
| Formatting alone, fixed recipes | Poor | Avoid paying for makeovers |

The grades come from the Martinez review. A controlled test by [Vishwakarma and colleagues](https://arxiv.org/abs/2605.25517) across 252,000 trials backs the top rows: topic match and position decided which page was cited first.

The content analysis by Zhang and colleagues points the same way for evidence. By the authors’ measure of influence, cited pages containing numbers or statistics shaped answers 61.55% more on average than pages without. Pages in question-and-answer format did slightly worse, by 5.74%.

## Is the “40% more visibility” claim true?

Only inside the lab setup that produced it. The figure does not mean 40% more citations, traffic or customers.

The original GEO study reported that rewrites “can boost visibility by up to 40%”. Martinez traces this to one score rising from 19.3 to 27.2 when quotations were added. In that test, five pages were already placed in front of the AI, so the result says nothing about being found.

The study’s live check on Perplexity used 200 examples, uploaded as files rather than found on the web. Martinez lists “GEO increases visibility by 40%” as a claim rejected in general form.

## Which popular tactics do not hold up?

Keyword stuffing, one-size-fits-all rewrites and markup shortcuts have weak or negative evidence. Some rewrites even hurt.

- **Keyword stuffing.** On Perplexity, it performed 10% worse than the original page in the GEO study. Our guide on [whether keyword stuffing still works](https://underneath.agency/resources/does-keyword-stuffing-work-in-ai-search) covers the later tests.
- **Generic rewrite recipes.** In the C-SEO Bench benchmark, only three of 54 method and subject combinations were clearly positive, and some rewrites lowered a page’s rank.
- **Rewrites that hurt being found.** In an end-to-end test, rewriting only the body text cut a page’s presence among the top 10 candidates by 16%.
- **Schema as a shortcut.** In [our study of AI Overview citations](https://underneath.agency/research/ai-overview-cited-pages-study), Organization schema added +0.5 points to the chance of being cited, which is no meaningful difference.
- **llms.txt as a cure-all.** Only 11.5% of top sites publish a valid file, in [our llms.txt study](https://underneath.agency/research/llms-txt-adoption-study), and no one we found has shown it changes citations.

## Why do results vary so much between studies?

AI engines differ from each other and change constantly, so one snapshot rarely generalizes. Even repeated runs of the same question give different sources.

Martinez cites a two-month study in the US and Germany where only 18% of the pages in AI Overviews stayed the same, against 45% for regular Google results. Repeated runs set up to be as consistent as possible still changed 9–28% of decisions. In another study, 57.8% of ChatGPT runs did not search the web at all.

Engines also disagree with each other. In [our comparison of Google’s two AI surfaces](https://underneath.agency/research/ai-mode-vs-ai-overviews-study), AI Mode and AI Overviews shared a mean 13.9% of cited URLs on the same searches. A result measured on one engine, on one day, should not be treated as a rule.

## Does GEO lead to more traffic or sales?

The evidence here is the weakest in the field. One field study suggests a gain, but it is not proof.

In that study, ChatGPT referrals to treated pages rose by a factor of 5.7. But untreated pages on the same site had already risen 3.5 times as the platform grew. A careful analysis put the extra effect at 1.82 times, which Martinez calls “suggestive rather than causally established”. One industry study reported a 20% traffic lift, without enough detail to judge it.

Being cited is not the same as being represented well, either. An early audit found only 51.5% of sentences in AI search answers were fully supported by their citations. Our article on [AI Mode and traffic](https://underneath.agency/resources/ai-mode-default-traffic-loss) covers what the shift to AI answers may do to clicks.

## What should you do about it?

Follow the conservative playbook and measure each stage separately. In practice:

1. Map the real questions buyers ask, and make sure one page answers each directly.
2. Back claims with specific, sourced figures, definitions and comparisons. Never invent statistics or reviews.
3. Use clear headings and tables, but judge them by results, not by assumption.
4. Keep pages reachable. In [our crawler study](https://underneath.agency/research/ai-crawler-blocking-study), 7.4% of top sites blocked ChatGPT’s search crawler.
5. Keep investing in search ranking, since being gathered comes before being cited.
6. Measure repeatedly: run each question several times, on several dates and engines, and track found, cited and quoted correctly as separate numbers.
7. Avoid hidden instructions to AI systems, fake authority and undisclosed promotion; Martinez’s tests for honest optimization rule these out.

For help building this kind of measured program, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The research has not shown that any GEO practice durably lifts visibility, traffic or revenue across engines.

- **Being found.** Very few studies test whether a change helps a real page get retrieved by a live engine.
- **Competition.** Gains may shrink as everyone adopts the same tactics; C-SEO Bench found the problem behaves like a zero-sum game.
- **Business outcomes.** Clicks, leads and sales are almost never measured with a proper control group.
- **Durability.** Engines change often, and most studies are single snapshots.
- **Non-English markets.** Most studies use English, anonymous accounts and few locations.

## Frequently asked questions

### Is GEO proven to work?

Partly. Changing a page already gathered by an AI system can change how it is cited. But no reviewed technique has shown a stable, cross-engine effect on being found or on traffic.

### Does keyword stuffing work for AI search?

No. In the original GEO study, keyword stuffing performed 10% worse than the unedited page on Perplexity, and later reviews rate it null or negative.

### Is GEO different from SEO?

It builds on SEO rather than replacing it. The strongest levers are relevance and position among sources, and both depend on being found by search first.

### How should I measure GEO results?

Measure being found, being cited and being quoted correctly as separate numbers, over repeated runs. In one study, 57.8% of ChatGPT runs did not search the web at all, so a single check can mislead.

## Sources

- Martinez (2026), [Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026)](https://arxiv.org/abs/2607.14035), arXiv:2607.14035.
- Aggarwal and colleagues (2023), [GEO: Generative Engine Optimization](https://arxiv.org/abs/2311.09735), arXiv:2311.09735.
- Puerto and colleagues (2025), [C-SEO Bench: Does Conversational SEO Work?](https://arxiv.org/abs/2506.11097), arXiv:2506.11097.
- Vishwakarma, Kumar and Jamidar (2026), [What Gets Cited: Competitive GEO in AI Answer Engines](https://arxiv.org/abs/2605.25517), arXiv:2605.25517.
- Zhang, He and Yao (2026), [From Citation Selection to Citation Absorption: A Measurement Framework for Generative Engine Optimization Across AI Search Platforms](https://arxiv.org/abs/2604.25707), arXiv:2604.25707.
- Underneath (2026), [What pages cited by AI Overviews have in common: 3,096 pages](https://underneath.agency/research/ai-overview-cited-pages-study)
- Underneath (2026), [How many websites have an llms.txt file? 2026 adoption data](https://underneath.agency/research/llms-txt-adoption-study)
- Underneath (2026), [AI Mode vs AI Overviews: how different are the sources?](https://underneath.agency/research/ai-mode-vs-ai-overviews-study)
- Underneath (2026), [Which AI crawlers do top websites block? 10,000 sites, 2026](https://underneath.agency/research/ai-crawler-blocking-study)

---

This is the Markdown twin of https://underneath.agency/resources/what-geo-practices-does-research-support. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "What happens to brands that skip GEO when competitors adopt it?"
description: "In controlled tests, brands that did nothing while rivals optimized got zero AI recommendations. The early lead shrinks fast, so copying rivals is not enough."
canonical: "https://underneath.agency/resources/what-happens-if-you-skip-geo"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# What happens to brands that skip GEO when competitors adopt it?

In controlled tests they disappear: when rival brands used optimized product copy, brands that did nothing received zero recommendations across 4,745 trials. The same research shows the advantage of going first shrinks fast as rivals catch up, and many popular tactics do not work at all. Skipping generative engine optimization (GEO), the work of making a brand visible and accurately described in AI answers, is risky; copying rivals’ tricks is not a strategy either.

## The short version

1. In a test of three AI assistants, invented brands that did not optimize got zero recommendations across 4,745 trials once rivals did ([Chu and Hou](https://arxiv.org/abs/2606.17443)).
2. Going first paid off, but the gain fell from +0.802 to +0.007 on the authors’ scale once all nine rivals used the same tactic.
3. In a separate benchmark, popular rewriting tactics produced a reliable gain in only three of 54 tests, and gains shrank as more sites adopted them ([Puerto and colleagues](https://arxiv.org/abs/2506.11097)).
4. Signs of optimization appeared on 8.90% of pages in Google and Gemini results for 1,000 real questions, and on 16.36% of pages updated in 2026 ([Chu, Leng and colleagues](https://arxiv.org/abs/2608.16824)).
5. On one live website, ChatGPT referrals grew 5.7 times, but untreated pages grew 3.5 times too, so most growth came from ChatGPT itself ([Watanabe and Nakayashiki](https://arxiv.org/abs/2606.04362)).

## What happened to brands that did nothing in the experiments?

They vanished: once rivals optimized, brands that stayed neutral were never recommended.

[Chu and Hou](https://arxiv.org/abs/2606.17443), of Trine University and Texas A&M, built lists of ten skincare products: one real, well-known brand and nine invented challengers, all with identical specs. They then gave some challengers authority-style copy, such as clinical-trial claims, and asked GPT-4o-mini, Claude Sonnet and Gemini 3 Flash to recommend one.

With no optimized rivals, the famous brand won every time. A single optimized challenger cut the famous brand’s survival from 100% to 19.8%. And across 4,745 trials, invented brands without optimized copy received zero recommendations. The authors put it bluntly: any brand that does not optimize when competitors do becomes invisible.

Read the conditions closely. The brands were invented, the specs were identical, and the clinical claims were deliberately made up to find the upper limit. The finding holds inside that experiment; real markets have more ways to stand apart.

## Does moving first keep you ahead?

Not for long: the first mover’s advantage shrank by half with each 1.4 rivals that copied it.

The first brand to optimize gained a lot, +0.802 on the authors’ payoff scale. As others copied the same tactic, the gain decayed, halving every 1.4 brands. With all nine challengers optimized, each one’s gain fell to +0.007, close to nothing. The famous brand recovered to 93.8% of recommendations, because the assistants ignored a signal everyone now sent.

The assistants differed at that end point. Claude and GPT-4o-mini largely returned to the famous brand, while Gemini kept some lasting effect, with the famous brand at 84.9%. The authors call this a prisoner’s dilemma: every brand is better off optimizing, yet when all do, nobody gains, and no one can afford to stop.

A separate benchmark, [C-SEO Bench](https://arxiv.org/abs/2506.11097), from Parameter Lab, NAVER and university researchers, found the same pattern for the best two rewriting methods in retail and video games, tested on GPT-4o-mini: early adopters gained, and the gain fell steadily as more sites adopted the same method. The authors describe it as congested and zero-sum.

## Do the popular GEO tactics actually work?

Many do not: in one benchmark, common rewriting tactics rarely helped and often hurt.

C-SEO Bench tested ten ways of rewriting web pages, such as [adding statistics, quotations or an authoritative tone](https://underneath.agency/resources/do-geo-content-tactics-work), across six areas including retail products and news. Out of 54 cases tested, only three showed a reliable gain. Adding statistics lowered rankings in 19 of 24 settings. For retail products, one method left 61.0% of rankings unchanged.

Manipulation is also getting harder. In the [E-GEO](https://arxiv.org/abs/2511.20867) test bed from MIT and Columbia, AI shopping rankers told to watch for questionable listings flagged 100% of listings rewritten with extreme superlatives. The rewrites that kept working were factual, clearly organized and aimed at real buyer questions.

So the risk of skipping GEO is real, but so is [the risk of doing it badly](https://underneath.agency/resources/can-geo-backfire-on-your-brand).

## Is search ranking still part of it?

Yes: whether your page is found in the first place still outweighs most rewriting.

C-SEO Bench found that moving a page to the top of the material an assistant reads produced far bigger gains than any rewriting method. The authors conclude that traditional search optimization remains essential and that GEO complements it rather than replacing it.

Rankings are not the whole story, though. In [our study](https://underneath.agency/research/ai-citations-google-rankings-study) of 80 US buyer questions, 8.3% of the pages ChatGPT cited ranked in Google’s top 10 for the question, against 25.7% for Claude. Chu and Hou also ran a small add-on test in which a basic search step chose which products the assistant saw. Against one optimized rival, the famous brand was then recommended only 5.0–6.7% of the time, because the search step did not favor familiar names. The authors treat that probe as directional.

## How many competitors are already optimizing for AI?

A growing minority, by one estimate: about one page in eleven in real search results shows signs of it.

[Chu, Leng and colleagues](https://arxiv.org/abs/2608.16824), led from the CISPA Helmholtz Center for Information Security, built a [detector for pages optimized for AI search](https://underneath.agency/resources/how-much-web-content-is-optimized-for-ai-search). They ran it on pages returned by Google Search and Gemini for 1,000 real user questions. It flagged 8.90% of 10,095 pages. Among pages with a stated update date, the share rose from 7.02% for 2024 to 16.36% for 2026. The authors warn that the detector makes mistakes and that the rise is a trend in pages, not proof of rising adoption.

Our own data shows one visible form of it. In [our best-of lists study](https://underneath.agency/research/self-promoting-best-lists-study), 24.2% of numbered “best X” lists cited by AI engines, where the publisher could be identified, ranked their own publisher first.

## How big is the real-world payoff?

Smaller than headline case studies suggest, according to the most careful field study we found.

[Watanabe and Nakayashiki](https://arxiv.org/abs/2606.04362) work at Glasp, the website they studied, so this is a vendor’s own account of one domain. In January 2026 they applied a bundle of changes, mainly rewritten titles and summaries, to one large section of their site and used the rest of the site as a comparison. Total ChatGPT referrals grew 5.7 times, but untreated pages grew 3.5 times over the same window. Once that rising tide was removed, the optimized section showed a step up of about 1.82 times. A stricter test left that result suggestive, not conclusive.

The lesson for a brand deciding whether to act is about measurement. Raw growth in AI traffic mostly reflects AI assistants getting bigger, so it cannot tell you whether your rivals’ work, or yours, is paying off.

## What should you do about it?

Do not sit out, but compete on substance and measure honestly rather than copying rivals’ tactics.

1. **Start with being findable.** Pages that search systems retrieve beat clever rewrites of pages they never see.
2. **Differentiate with real facts.** Specific, verifiable evidence lasts longer than a tactic every rival can copy.
3. **Skip invented claims and keyword tricks.** They can be flagged, and in benchmarks many tactics lowered rankings.
4. **Watch what rivals are doing.** Check which competitors appear in AI answers to your buyers’ questions, and which pages and lists those answers cite.
5. **Measure against a baseline.** Compare optimized pages with untreated ones, and judge visibility across many repeated questions. In [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), pinning one brand’s visibility near 50% within 10 points would take about 97 runs.

If you want a structured way to do this, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

No study has followed real competing brands in live AI assistants over time, so the strongest results are from simulations.

- The “zero recommendations” finding used invented brands, identical specs and deliberately made-up claims, in one product category.
- Benchmarks used older assistant versions, often with the material handed to the assistant rather than found by search.
- The prevalence estimate depends on an imperfect detector and covers one set of 1,000 questions.
- The only field study covers one website, run and reported by its owner.
- Nobody has measured what skipping GEO costs in sales.

## Frequently asked questions

### Is GEO necessary if my competitors are doing it?

The controlled evidence says ignoring it is risky. In one test, brands that did not optimize while rivals did received zero recommendations across 4,745 trials, though the brands and claims were invented.

### Does being first to do GEO give a lasting advantage?

No. In the same test the first mover’s gain halved with every 1.4 rivals that copied it and fell to almost nothing when all nine did.

### Do common GEO tactics like adding statistics work?

Often not. In one benchmark only three of 54 cases tested showed a reliable gain, and adding statistics lowered rankings in 19 of 24 settings.

### How much web content is already optimized for AI search?

One detector flagged 8.90% of pages in Google and Gemini results for 1,000 real questions as optimized, rising to 16.36% among pages updated in 2026.

## Sources

- Chu and Hou (2026), [Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems](https://arxiv.org/abs/2606.17443), arXiv:2606.17443.
- Puerto, Gubri, Green, Oh and Yun (2025), [C-SEO Bench: Does Conversational SEO Work?](https://arxiv.org/abs/2506.11097), arXiv:2506.11097.
- Bagga, Farias, Korkotashvili, Peng and Wu (2026), [E-GEO: A Testbed for Generative Engine Optimization in E-Commerce](https://arxiv.org/abs/2511.20867), arXiv:2511.20867.
- Chu, Leng, Li, Shen, Shen and Zhang (2026), [GEO-Flag: Detecting and Measuring GEO-Optimized Web Content](https://arxiv.org/abs/2608.16824), arXiv:2608.16824.
- Watanabe and Nakayashiki (2026), [Disentangling Answer Engine Optimization from Platform Growth: A Log-Based Natural Experiment on ChatGPT Referral Traffic](https://arxiv.org/abs/2606.04362), arXiv:2606.04362.
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude cite pages that rank?](https://underneath.agency/research/ai-citations-google-rankings-study)
- Underneath (2026), [How many “best of” lists cited by AI rank their own brand first?](https://underneath.agency/research/self-promoting-best-lists-study)
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)

---

This is the Markdown twin of https://underneath.agency/resources/what-happens-if-you-skip-geo. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "What is GEO, and how is it different from SEO? | Underneath"
description: "GEO is the work of getting your information found, used and cited in AI-written answers. SEO competes for a rank; GEO competes for a share of the answer."
canonical: "https://underneath.agency/resources/what-is-generative-engine-optimization"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# What is GEO, and how is it different from SEO?

Generative engine optimization (GEO) is the work of getting your information found, used and correctly cited by AI systems that write answers instead of listing links. SEO competes for a position on a ranked list. GEO competes for a share of one written answer, chosen by an engine that often reads different sources from Google.

## The short version

1. The term comes from a 2023 paper by Princeton and IIT Delhi researchers, which reported that content changes could raise a page’s share of AI answers by up to 40% in a simulated AI search engine ([Aggarwal and colleagues](https://arxiv.org/abs/2311.09735)).
2. Across 55,936 queries, only 38% of the domains cited by six AI search engines also appeared in Google or Bing results ([Zhang and colleagues](https://arxiv.org/abs/2512.09483)).
3. In US car questions, 81.9 percent of the sources an AI engine cited were independent “earned” media, against 45.1 percent for Google ([Chen and colleagues](https://arxiv.org/abs/2509.08919)).
4. Keyword stuffing, a classic SEO trick, did 10% worse than doing nothing when tested on Perplexity ([Aggarwal and colleagues](https://arxiv.org/abs/2311.09735)).
5. Asked the same question five times, ChatGPT named only 25.2% of its brands every time, in [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study).

## What is generative engine optimization?

GEO means shaping your content and reputation so AI answer engines find it, use it and cite it accurately. The answer engines in question include ChatGPT, Perplexity, Gemini and Google’s AI Overviews, the AI summary at the top of Google’s results.

The name was coined by [Aggarwal and colleagues](https://arxiv.org/abs/2311.09735), researchers at Princeton, IIT Delhi and elsewhere, in a paper first posted in 2023. They framed it as a way to help content creators be seen in AI answers when the engines are closed boxes. Their headline result was that GEO “can boost visibility by up to 40% in generative engine responses.”

That 40% needs care. [Martinez’s 2026 survey](https://arxiv.org/abs/2607.14035) of 45 studies explains that it came from a simulator in which five pages had already been handed to the AI. The best method raised a page’s weighted share of the answer from 19.3 to 27.2. It does not mean 40% more customers or 40% more citations in the wild. We cover this in detail in [our article on the 40% claim](https://underneath.agency/resources/does-geo-increase-ai-visibility-by-40-percent).

## What happens when a search engine writes the answer instead of listing pages?

The prize changes from a place on a list to a share of one answer. A classic search engine shows ten links and leaves the choice to the reader. An AI engine reads sources for the reader, writes a summary and attaches a few citations.

The foundational paper puts the difference plainly: generative engines “combine information from multiple sources in a single response,” so simple ranking metrics “are not applicable” to their answers. Visibility becomes relative. When one source gains a larger share of the answer, another loses some.

The engine also does its own searching. In [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study), ChatGPT ran a mean of 3.7 searches per buyer question. None of the 509 searches we saw repeated the user’s question word for word. You are optimizing for searches you never see.

And AI answers favor certain kinds of question. In a 40-day audit of 55,393 trending searches in spring 2026, [Xu, Iqbal and Montgomery](https://arxiv.org/abs/2605.14021) found that question-form searches triggered an AI Overview 64.7% of the time, against 9.5% for other searches.

## Do AI search engines find the same sources as Google?

Mostly not: they overlap with Google only partly, and they lean toward third-party sources. That is the main reason ranking well on Google is not the same as appearing in AI answers.

[Zhang and colleagues](https://arxiv.org/abs/2512.09483) compared six AI search engines with Google and Bing across 55,936 queries in 2025. Only 38% of domains appeared in both kinds of results, and 37% were unique to the AI engines. Even Google’s own AI Overviews look past its rankings: in the spring 2026 audit by Xu and colleagues, 29.8% of domains cited in AI Overviews did not appear in the first page of results at all. Our guide on [how AI and Google sources overlap](https://underneath.agency/resources/do-ai-search-engines-cite-the-same-sites-as-google) compares more of these studies.

The type of source shifts too. [Chen and colleagues](https://arxiv.org/abs/2509.08919) at the University of Toronto sorted sources into brand-owned, social and “earned” media, meaning independent reviews and publications.

| US car questions, 2025 | Earned media | Brand-owned | Social |
|---|---|---|---|
| Google | 45.1 percent | 39.5 percent | 15.4 percent |
| AI search (web-enabled GPT) | 81.9 percent | 18.1 percent | none |

So GEO reaches beyond your website. What reviewers, publishers and directories say about you often matters more than your own pages. The ranking link still matters, as we explain in [our article on whether SEO still matters](https://underneath.agency/resources/is-seo-still-important-for-ai-search).

## Do classic SEO tactics carry over to AI search?

Some do and some do not. Being crawlable and authoritative still helps you get considered; repeating keywords does not.

In the foundational GEO tests, keyword stuffing offered “little to no improvement.” On Perplexity, a live AI search engine, it “performs 10% worse than the baseline.” The tactics that helped in that study added substance instead, such as quotations, statistics and sources. Later studies found those gains far smaller and less reliable, as we report in [our review of GEO content tactics](https://underneath.agency/resources/do-geo-content-tactics-work).

The SEO foundations still matter. [Zhang Kai and colleagues](https://arxiv.org/abs/2604.25707), in a study of 602 prompts across ChatGPT, Google and Perplexity, write that “source authority and domain recognizability influence the gateway into citation pools.” The new layer, in their words, is “answer participation”: whether your page actually shapes what the answer says. Their study was run by three independent researchers on a public dataset and has not been peer reviewed.

## Why is AI visibility harder to pin down than a Google ranking?

Because AI answers change from run to run and from engine to engine. A ranking is one number; AI visibility is a range of outcomes.

In [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), we asked ChatGPT, Gemini and Perplexity the same 20 buyer questions five times each. Of the brands ChatGPT named for a question, 25.2% appeared in all five runs. A single ChatGPT answer showed only 57.8% of the brands its five answers named between them.

Engines also disagree with each other. Across 80 buyer questions in [our brand agreement study](https://underneath.agency/research/ai-assistants-brand-agreement-study), ChatGPT, Gemini, Perplexity and Claude all chose the same first pick for 10.0% of questions. A small 15-prompt audit by [Tannenbaum](https://arxiv.org/abs/2609.22655), an author at a private company, found that 96.4% of cited addresses appeared in only one of four engines on the same day.

Being seen also yields fewer visits. [Huang and colleagues](https://arxiv.org/abs/2603.16138) cite Pew Research data showing that people clicked a source inside Google’s AI summary in only 1% of visits. So the goal shifts from winning clicks alone to shaping what the answer says. We look at clicks in [our article on answer engines and website visits](https://underneath.agency/resources/ai-answer-engines-fewer-website-clicks).

## What should you do about it?

Treat GEO as an extension of search work, measured on AI answers rather than rankings. Five steps cover most of it.

1. Keep your SEO foundations. Pages that cannot be crawled or that lack authority rarely get into the pool AI engines choose from.
2. Invest in third-party coverage. Reviews, trade press and comparison sites carry more weight in AI answers than in Google results.
3. Write pages an engine can quote. State facts, definitions, comparisons and figures clearly, with sources.
4. Measure AI answers directly. Ask your buyers’ questions on several engines, several times, and track how often you appear and what is said.
5. Judge results over months, not single snapshots, because answers vary from run to run.

If you want help setting this up, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The research describes how AI search differs from Google far better than it shows what reliably moves results. Several gaps remain.

- Most GEO experiments test pages already handed to the AI. Few show that a change makes a live engine find a page in the first place.
- No study we reviewed links GEO work to sales. Martinez’s survey notes that effects on conversions are almost never established.
- Engines change quickly. Most findings are snapshots from one period, often in English and the US.
- Several key studies are preprints, small, or written by companies that sell GEO tools.

## Frequently asked questions

### Is GEO replacing SEO?

No: GEO builds on SEO rather than replacing it. Authority and crawlability still decide whether AI engines consider a page, but AI answers draw on a different mix of sources, so SEO alone covers only part of the picture.

### Who invented the term generative engine optimization?

Aggarwal and colleagues, researchers at Princeton, IIT Delhi and elsewhere, introduced it in a paper first posted in November 2023. It reported gains of up to 40% in a simulated AI search engine.

### Does a top Google ranking get you into ChatGPT answers?

Not reliably. In [our rankings study](https://underneath.agency/research/ai-citations-google-rankings-study), only 8.3% of the pages ChatGPT cited ranked in Google’s top 10 for the question, against 25.7% for Claude.

### Is GEO the same on every AI engine?

No. Engines cite different sources and different numbers of them, so a page that does well on one may be absent from another. Our brand agreement study found all four assistants shared a first pick for only 10.0% of questions.

## Sources

- Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan and Deshpande (2024), [GEO: Generative Engine Optimization](https://arxiv.org/abs/2311.09735), arXiv:2311.09735.
- Martinez, O. (2026), [Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026)](https://arxiv.org/abs/2607.14035), arXiv:2607.14035.
- Zhang, Ye, Peng, Garimella and Tyson (2025), [Source Coverage and Citation Bias in LLM-based vs. Traditional Search Engines](https://arxiv.org/abs/2512.09483), arXiv:2512.09483.
- Chen, Wang, Chen and Koudas (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919), arXiv:2509.08919.
- Xu, Iqbal and Montgomery (2026), [Measuring Google AI Overviews: Activation, Source Quality, Claim Fidelity, and Publisher Impact](https://arxiv.org/abs/2605.14021), arXiv:2605.14021.
- Zhang, He and Yao (2026), [From Citation Selection to Citation Absorption: A Measurement Framework for Generative Engine Optimization Across AI Search Platforms](https://arxiv.org/abs/2604.25707), arXiv:2604.25707.
- Tannenbaum (2026), [Scoring With the Engine: Retrieval Exposure, Cross-Engine Divergence, and the Limits of Engine-Agnostic GEO Scores](https://arxiv.org/abs/2609.22655), arXiv:2609.22655.
- Huang, Goyal, Saha and Chandrasekharan (2026), [Answer Bubbles: Information Exposure in AI-Mediated Search](https://arxiv.org/abs/2603.16138), arXiv:2603.16138.
- Underneath (2026), [Do AI search engines run hidden searches?](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude agree on brands?](https://underneath.agency/research/ai-assistants-brand-agreement-study)
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude cite pages that rank?](https://underneath.agency/research/ai-citations-google-rankings-study)

---

This is the Markdown twin of https://underneath.agency/resources/what-is-generative-engine-optimization. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "What to measure to know if you are visible in AI answers"
description: "Measure how often each AI engine names and cites you in final answers, over repeated runs, and whether what it says is right, not rankings or one screenshot."
canonical: "https://underneath.agency/resources/what-to-measure-ai-visibility"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# What should we measure to know whether we are visible in AI answers?

Measure what people actually see: how often each AI engine names your brand and cites your pages in its final answers, across many questions and repeated runs. Then check where you appear in the answer and whether what it says about you is accurate. Rankings of pages before the answer is written, and single screenshots, tell you much less.

## The short version

1. Brand mentions are steadier than cited pages. In a Swiss study, brand lists overlapped 45% to 59% from one day to the next, cited sources only 34% to 42% (Schulte and colleagues, 2026).
2. Being cited is not the same as being used. ChatGPT cited 6.88 sources per question against Perplexity’s 16.35, yet each page it cited shaped its answer far more (Zhang Kai and colleagues, 2026).
3. Being known is not being recommended. In one study ChatGPT recognized 99.4% of products when asked by name but surfaced only 3.32% in open discovery questions (reported by Martinez, 2026).
4. Each engine is its own market: in one small audit, 96.4% of cited web addresses appeared on only one of four engines (Tannenbaum, 2026).
5. Presence is not accuracy: in [our pricing study](https://underneath.agency/research/ai-pricing-accuracy-study), 61.9% of software plan prices quoted by AI assistants were fully faithful to the vendor’s page.

## Why measure final answers instead of rankings or page scores?

Because the final answer is what people read. Rankings and page scores only describe steps before it.

An AI answer is built in stages. The engine decides whether to search, runs its own searches, picks some pages, writes an answer and attaches citations. A page can drop out at any stage. [Wen and colleagues](https://arxiv.org/abs/2606.12439) argue that academic studies often report movement in intermediate ranked lists. By contrast, how often a source appears and is cited in the final answer shows whether it survived every stage.

Google rankings are a weak stand-in. In [our rankings study](https://underneath.agency/research/ai-citations-google-rankings-study), only 8.3% of the pages ChatGPT cited ranked in Google’s top 10 for the question asked. [Tannenbaum](https://arxiv.org/abs/2609.22655) makes the same point about scores that rate a page without ever querying an engine. In his audit of 15 prompts on one day, 96.4% of the web addresses cited appeared on only one of ChatGPT, Copilot, Google and Perplexity. A page score cannot tell you which engine will find you. The author works for a company that sells AI visibility scoring.

## Should you track brand mentions or cited pages?

Track both, but treat brand mentions as the headline. They are steadier, and most citations point at someone else’s site.

[Schulte and colleagues](https://arxiv.org/abs/2604.07585) followed four AI engines in four Swiss industries for about 45 days. From one day to the next, the brands named overlapped 45% to 59%, while [the sources cited overlapped only 34% to 42%](https://underneath.agency/resources/why-ai-search-citations-change). They conclude that brand presence, added up over a campaign, is a more reliable measure than individual cited web addresses. Their brand detection used a fixed list of names, and one industry was dropped because answers rarely named brands.

Cited pages and named brands can also move independently. In [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), Perplexity cited identical web addresses in 166 pairs of runs, yet the brand list still changed in 91.6% of them. And your own site is a small slice of the evidence. On one tracking platform, only 2.9% of citations pointed at the tracked brand’s own domain, while 75.2% pointed at other companies in the same space. That figure comes from [Kumar](https://arxiv.org/abs/2606.20065), a co-founder of the platform.

## Is being cited the same as being used?

No. A page can be listed as a source while contributing little to what the answer actually says.

[Zhang Kai and colleagues](https://arxiv.org/abs/2604.25707) studied 602 prompts across ChatGPT, Google and Perplexity. They scored [how much each cited page shaped the answer](https://underneath.agency/resources/does-ai-search-use-the-pages-it-cites), from repeated references to shared wording.

| AI engine | Citations per question | How much each cited page shaped the answer (0 to 1) |
|---|---|---|
| ChatGPT | 6.88 | 0.2713 |
| Google AI Overview or Gemini | 12.06 | 0.0584 |
| Perplexity | 16.35 | 0.0646 |

The authors’ advice is to track both whether you are cited and how much your page shapes the answer. A dashboard that counts citations alone misses the difference. Their influence score is their own construction from one static dataset, not a view inside the engines.

Citations also do not guarantee support. In a 2023 audit of four AI search engines, summarized by [Martinez](https://arxiv.org/abs/2607.14035), only 51.5% of sentences in AI search answers were fully supported by their citations. Only 74.5% of citations supported the statement they were attached to.

## Does it matter whether the question names your brand?

Yes, enormously. Questions that name you measure recognition; open questions measure whether you get recommended.

Martinez reports a preprint by Sharma on 112 startups. ChatGPT recognized 99.4% of products when they were named, but surfaced them in only 3.32% of open discovery questions. For Perplexity the drop was from 94.3% to 8.29%. That is one preprint using two models, but the distinction it draws matters for any score. A visibility score padded with branded questions will look healthy while customers never hear your name.

The denominator matters too. In the Swiss data, 57.8% of ChatGPT runs did not search the web at all. A share of citations computed only among answers that cite something overstates how often people see you.

## Do position and tone matter as well as presence?

Yes. Two answers can both name you while one puts you first and the other sixth.

In our consistency study, 86.4% of the brands ChatGPT named in every one of five runs still moved position at least once. Its first pick changed at least once for 80.0% of questions. So track separately how often you are named, how often you are in the top three, and how often you are first. Engines differ here too; see [which AI engine is most consistent](https://underneath.agency/resources/most-consistent-ai-engine-for-brands).

Tone is harder to measure reliably. On the platform studied by Kumar, the tone toward a brand flipped between positive and negative in 45.5% of cases. Whether the brand was mentioned flipped in only 6.8%. Tone readings need many more answers before they settle, and the platform’s own model scored the tone.

## Should you check whether the answer is accurate?

Yes. Being named with the wrong price, phone number or claim can cost you more than not being named.

In our pricing study, 61.9% of plan prices quoted by four AI assistants were fully faithful to the vendor’s page. In [our business facts study](https://underneath.agency/research/ai-business-facts-accuracy-study), 18.9% of answers about local businesses stated at least one fact that differed from the business’s Google profile. Not every difference was an error, which is why the check needs a human eye.

## What should you do about it?

Build a small set of measures that follow the answer from search to sale. In practice:

1. **Lead with brand mention rate** on open questions that do not name you, per engine, over repeated runs.
2. **Report branded questions separately.** They show recognition, not recommendation.
3. **Track position:** share of answers where you are first, in the top three, or named at all.
4. **Track citations of your own pages** as a second measure, and note which third-party pages cite you.
5. **Count answers with no search or no citations** in the denominator rather than dropping them.
6. **Audit accuracy** of prices, facts and claims in a sample of answers every month.
7. **Keep engines separate.** A gain on one engine does not carry over to another.

If you want a measurement plan built this way, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

No study yet links any of these measures to sales or revenue. Other gaps:

- **Which measure predicts business results.** Wen and colleagues argue final-answer measures sit closer to customer attention, but they did not measure clicks or revenue.
- **Whether influence scores reflect real use.** Measures of how much a page shaped an answer are built from text overlap, not from the engines’ internals.
- **Whether customers notice citations.** Few studies observe whether people see or click the sources listed.
- **How stable these patterns are.** Most findings come from one period, a few markets and a handful of engines, and the engines change often.

## Frequently asked questions

### What is the best single metric for AI visibility?

There is no single best one, but brand mention rate on unbranded questions, per engine, is the most stable starting point. In the Swiss study, brand lists overlapped 45% to 59% from day to day, against 34% to 42% for cited sources.

### Is share of citations a good measure of AI visibility?

It is useful but incomplete. Citations count exposure, not influence, and ChatGPT cited 6.88 sources per question while using each far more heavily than Perplexity used its 16.35.

### Should I track prompts that name our brand?

Yes, but report them separately. Engines recognize named brands almost every time, so branded prompts inflate a visibility score that should reflect open recommendation questions.

### Do Google rankings tell me how visible I am in AI answers?

Not reliably. Only 8.3% of the pages ChatGPT cited ranked in Google’s top 10 for the same question in our study.

## Sources

- Wen, Y., Zhang, N., Yuan, H., Chen, X., Zhang, H. and Guo, H. (2026), [Position: Generative Engine Optimization Creates Underexamined Risks, Governance Must Target Concentration, Disclosure, and Academic Blind Spots](https://arxiv.org/abs/2606.12439), arXiv:2606.12439.
- Schulte, J., Bleeker, M. and Kaufmann, P. (2026), [Don’t Measure Once: Measuring Visibility in AI Search (GEO)](https://arxiv.org/abs/2604.07585), arXiv:2604.07585.
- Zhang Kai, He Xinyue and Yao Jingang (2026), [From Citation Selection to Citation Absorption: A Measurement Framework for Generative Engine Optimization Across AI Search Platforms](https://arxiv.org/abs/2604.25707), arXiv:2604.25707.
- Martinez, O. (2026), [Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026)](https://arxiv.org/abs/2607.14035), arXiv:2607.14035.
- Tannenbaum, B. (2026), [Scoring With the Engine: Retrieval Exposure, Cross-Engine Divergence, and the Limits of Engine-Agnostic GEO Scores](https://arxiv.org/abs/2609.22655), arXiv:2609.22655.
- Kumar, P. (2026), [Generative Engine Optimization at Scale: Measuring Brand Visibility Across AI Search Engines](https://arxiv.org/abs/2606.20065), arXiv:2606.20065.
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude cite pages that rank?](https://underneath.agency/research/ai-citations-google-rankings-study)
- Underneath (2026), [How faithfully do AI assistants quote software prices?](https://underneath.agency/research/ai-pricing-accuracy-study)
- Underneath (2026), [Do AI answers match a business’s Google profile?](https://underneath.agency/research/ai-business-facts-accuracy-study)

---

This is the Markdown twin of https://underneath.agency/resources/what-to-measure-ai-visibility. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "When AI agents can’t read your site, who tells your story?"
description: "From web searches and other people’s pages. In one test, 42% of an AI agent’s answer came from outside a business whose site it could not read."
canonical: "https://underneath.agency/resources/when-ai-agents-cant-read-your-site"
published: 2026-10-11
updated: 2026-10-11
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# If AI agents can’t read my website, where does their answer about us come from?

When an AI agent cannot read your website, it answers anyway, from web searches and other people’s pages. In a 2026 test across 1,056 real businesses, 42% of the answer about a hard-to-read business came from somewhere other than the business. Those answers were less accurate, less often recommended the business, and more often left out the facts the buyer asked for.

## The short version

1. In 37,927 AI agent tasks about 1,056 real businesses, 42% of the answer came from outside the business when its site was hard for agents to read ([Finder and colleagues](https://arxiv.org/abs/2609.34951)).
2. The agent almost never gives up: in about 99% of tasks that hit a dead end on a site, it answered anyway from what it found elsewhere.
3. Only 7 to 10% of the answer came from what the AI already knew, so the gap is filled by web searches, not memory.
4. Answers built from the business’s own pages got 48.3% of the requested facts right, against 34.3% for answers built from the open web.
5. In [our crawler study](https://underneath.agency/research/ai-crawler-blocking-study), 15.2% of top websites with a readable robots.txt file block OpenAI’s GPTBot, many through old rules that were not written with AI in mind.

## Where does the answer come from when an agent can’t read your site?

It comes mostly from web search results and third-party pages, not from your site or the AI’s memory. The clearest evidence is a 2026 study by [Finder and colleagues](https://arxiv.org/abs/2609.34951). They ran 37,927 agent journeys, each a buyer question about a business, across four independent harnesses over 1,056 real businesses. A harness here means one AI agent setup: a model plus its search and page-reading tools.

The businesses were split by how easy their sites were for an agent to fetch and read. The two groups were matched on fame, on how well the AI already knew the brand, and on [how widely the brand was mentioned elsewhere](https://underneath.agency/resources/off-site-mentions-vs-site-readability-ai-agents). An AI judge then labeled where each part of every finished answer came from.

| Where the finished answer came from | Easy-to-read site | Hard-to-read site |
|---|---|---|
| The business’s own pages | 78% | 58% |
| Web search results | 12% | 25% |
| Other websites | 3% | 7% |
| What the AI already knew | 7% | 10% |

In the authors’ words, on sites that are not agent-ready, 42% of the answer comes from somewhere other than the business. The share from web search roughly doubled, from 12% to 25%.

## Doesn’t the AI just fall back on what it already knows?

No: memory supplied only a small, steady slice of the answer whether or not the site was readable. Only 7 to 10% of the finished answer came from the model’s training, the same either way. Whatever the site did not provide, more web searching filled in.

The same team measured how this changed over time. On 90 buyer questions put to successive OpenAI releases, the [share of the answer built from memory](https://underneath.agency/resources/do-ai-assistants-answer-from-training-data) fell from 52% in gpt-4.1 to 14% in gpt-5.6. For a business, that means what an agent can read today matters more than what an AI learned about you years ago.

## Does the agent give up when your site blocks it?

Almost never: it searches more and keeps going. In about 99% of tasks that hit a dead end on a site, the agent still answered, built from whatever it found elsewhere.

The harder a site was to read, the more the agent searched. Mean web searches per task rose from 1.8 on the most readable sites to 4.5 on the least readable. Agents were also blocked by the site 2.1 times as often on hard-to-read sites. Each extra search is another chance for the answer to rest on a page you never wrote, never updated, or cannot vouch for.

That matters because third-party pages can be wrong or even planted. [Luo and Chen](https://arxiv.org/abs/2606.13610) tested 12 AI assistants in the lab. A single polluted page yields fooled rates of up to 27%, meaning the assistant recommended a fake product. The risk was higher where the AI knew the real brands less well. That test used saved search results, mostly in Chinese, not the live web.

## Does it change whether the AI recommends you?

Yes: in this study, readable businesses were clearly recommended almost twice as often. Two AI judges from different companies rated each answer, and both had to give the top score. [Agent-ready businesses got a clear recommendation](https://underneath.agency/resources/do-ai-agents-recommend-readable-websites) 20% of the time, against 11% for the others, or 1.9 times as often.

The gap showed up in all four agent setups and all twelve business categories, though the levels differed widely. One setup clearly recommended 5% of the time, another 36%. When agents could not read a site, they hedged. They admitted they could not access the business 4.4 times more often and vouched for it from secondhand sources 3.0 times more often.

## Are answers built from other sources less accurate?

Yes, mainly because they leave out the facts the buyer asked for. Comparing answers about the same business, from the same agent, to the same question, site-built answers got 48.3% of the requested facts right. Answers built from the open web got 34.3%.

[The main failure was omission, not invention](https://underneath.agency/resources/do-ai-agents-invent-or-omit-business-facts). Facts stated wrongly rose only from 4% to 6%, while facts never mentioned grew from 29% to 45%. An answer built from the web was 3.7 times more likely to contain none of the facts the buyer asked for. Accuracy was graded on a subset of 131 businesses, and the gain was largest for pricing questions.

Our own studies point the same way. In [our software pricing study](https://underneath.agency/research/ai-pricing-accuracy-study), most prices that differed from the official pricing page were not invented. For 39 of 64, the same figure appeared on another page of the vendor’s own site, and for 21 more on a third-party page cited for the product. In [our business facts study](https://underneath.agency/research/ai-business-facts-accuracy-study), 42 of 55 phone numbers that differed from the Google profile were on the business’s own website. What a business publishes, everywhere, is what agents repeat.

## How common is it for a site to be hard for agents to read?

Common enough to matter, often by accident. In [our crawler study](https://underneath.agency/research/ai-crawler-blocking-study) of the top 10,000 websites, 15.2% of the 5,572 top sites with a readable robots.txt block OpenAI’s GPTBot, its training crawler. Fewer block the crawlers that power AI search: 7.4% block OAI-SearchBot (ChatGPT search) and 7.1% block Claude-SearchBot. Google’s own opt-out raises a related question, covered in [blocking Google-Extended and AI Overviews](https://underneath.agency/resources/does-blocking-google-extended-hurt-ai-overviews).

Many of those blocks are not deliberate choices about AI. Across all AI search crawlers, 39.8% of blocks came from a catch-all rule that applies to every unnamed bot, against 24.9% for training crawlers. A file written years ago to keep bots out also keeps out agents that did not exist then.

Being allowed in is not the same as being easy to read. In [our study of agent-readable websites](https://underneath.agency/research/agent-readable-web-study), 191 of 5,902 live top websites (3.2%) returned a plain-text Markdown version when an agent asked for it. And 54.2% carried no structured data describing the business. Finder and colleagues report that across nearly 100,000 sites, fewer than 1% earned the top grade on their own readiness score.

## Would being readable mean the AI quotes you instead of others?

Not entirely: for broad “best of” questions, AI search leans heavily on third-party sources even when brand sites are open. In a 2025 study of ranking questions such as “best smartphones”, [Chen and colleagues](https://arxiv.org/abs/2509.08919) classified each cited source as brand-owned, earned (reviews and publishers) or social. For well-known brands, ChatGPT’s citations were 93.5% earned and 6.5% brand-owned.

So the two findings answer different questions. When a buyer asks about your business by name, a readable site lets the agent build most of the answer from your pages. When a buyer asks which product is best, reviews and publishers dominate, and your site is one voice among many.

## What should you do about it?

Make sure agents can get into your site, read each page, and find the facts buyers ask about.

1. **Check your robots.txt file for catch-all blocks.** Ask your web team which AI crawlers and user-triggered agents are allowed, and whether any block comes from an old rule nobody meant for AI.
2. **Put key facts in the page text.** Pricing, plans, features and setup steps should be readable without running scripts or opening menus. In the Finder study, the accuracy gain was largest for pricing.
3. **Keep your own facts consistent.** If two pages on your site show different prices or phone numbers, agents may repeat either one, as our pricing and business facts studies found.
4. **Describe your business in structured data.** A basic organization record with links to your real profiles gives agents a clear statement of who you are.
5. **Keep earning third-party coverage.** For “best of” questions, reviews and publishers still carry most of the weight.

If you want help checking what agents see on your site, our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) covers this.

## What does the research not tell us yet?

The evidence is strong on direction but rests heavily on one study with real limits.

- **The main study is vendor research.** Its authors work for a company that sells an agent-readiness score, and the study uses that company’s own scoring tool to sort sites. The authors say so themselves.
- **AI judges did the labeling.** Where each part of an answer came from, and whether it recommended the business, was judged by AI models, not people.
- **The sample is narrow.** The businesses were mostly software and commerce, mostly English-language, and the questions were business-fact lookups about pricing, features and setup.
- **Content depth was not matched.** Sites that block agents may also publish less, which could widen the gap on its own. Only the accuracy comparison, which compares answers about the same business, is protected from this.
- **No one has yet fixed a site and measured before and after.** The authors name that as the test that would settle cause and effect.

## Frequently asked questions

### Will ChatGPT still talk about my company if it can’t access my website?

Yes. In one large test, agents answered in about 99% of tasks that hit a dead end on a site, using web searches and third-party pages instead.

### Does blocking GPTBot stop my site appearing in ChatGPT answers?

Not by itself. GPTBot is OpenAI’s training crawler. On one day, [our crawler study](https://underneath.agency/research/ai-crawler-blocking-study) found ChatGPT cited six pages closed to GPTBot, while none of its 123 checked citations was closed to its search crawler.

### Do AI agents make up facts about businesses they can’t read?

Mostly they leave facts out rather than invent them. In the Finder study, facts stated wrongly rose only from 4% to 6%, while facts never mentioned rose from 29% to 45%.

### Is being mentioned on other sites enough for AI visibility?

Not for questions about your own business. In the Finder study, how widely a brand was mentioned elsewhere had no confirmed link to whether answers used its pages or recommended it, once readability was taken into account.

## Sources

- Finder, Elovic, Shalev and Yosef (2026), [AX is the New AEO](https://arxiv.org/abs/2609.34951), arXiv:2609.34951.
- Chen, Wang, Chen and Koudas (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919), arXiv:2509.08919.
- Luo and Chen (2026), [One Polluted Page Is Enough: Evaluating Web Content Pollution in LLM Recommenders](https://arxiv.org/abs/2606.13610), arXiv:2606.13610.
- Underneath (2026), [Which AI crawlers do top websites block?](https://underneath.agency/research/ai-crawler-blocking-study)
- Underneath (2026), [Do websites serve Markdown to AI agents?](https://underneath.agency/research/agent-readable-web-study)
- Underneath (2026), [How faithfully do AI assistants quote software prices?](https://underneath.agency/research/ai-pricing-accuracy-study)
- Underneath (2026), [Do AI answers match a business’s Google profile?](https://underneath.agency/research/ai-business-facts-accuracy-study)

---

This is the Markdown twin of https://underneath.agency/resources/when-ai-agents-cant-read-your-site. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Which Google searches trigger an AI Overview? | Underneath"
description: "Long, specific questions trigger Google AI Overviews most; local, brand and breaking-news searches rarely do. Overall rates run from 13.7% to 67%."
canonical: "https://underneath.agency/resources/which-searches-trigger-ai-overviews"
published: 2026-10-08
updated: 2026-10-10
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Which Google searches trigger an AI Overview?

Long, specific, question-shaped searches trigger AI Overviews far more often than short keywords, brand names, local “near me” searches or breaking news. How often they appear overall depends on which searches are counted, so published rates run from 13.7% to 67%. The only figure that matters for a business is the one measured on its own keyword list.

## The short version

1. Question-form searches triggered an AI Overview 64.7% of the time against 9.5% for other searches, in a spring 2026 study of 55,393 trending US searches by Xu and colleagues.
2. In our own test of 96 topics, the same topic showed an AI Overview 59.4% of the time as a short keyword and 93.8% as a natural question.
3. Searches where Google showed a local map showed an AI Overview only 21.8% of the time in our study, against 81.1% without a map.
4. Navigational searches, such as a brand or website name, returned an AI Overview only 12% of the time in a 2025 global test by Aral and colleagues.
5. Overall rates range from 8.1% of fresh trending searches to 51.5% of representative real-user searches, so the mix of searches decides the headline number.

## What is an AI Overview, and why does triggering matter?

An AI Overview is the AI summary Google places above its normal results for some searches, not all of them. If a search does not trigger one, there is nothing to be cited in, and the classic results page still decides who gets the click. When one does appear, it also affects [whether people keep browsing afterward](https://underneath.agency/resources/ai-overviews-end-browsing-sessions).

Google does not publish the rules. It says AI Overviews appear where AI adds benefit and are held back for “sensitive, explicit, rapidly evolving, or dangerous topics”, as [Xu, Iqbal and Montgomery](https://arxiv.org/abs/2605.14021) summarize. Everything known about triggering therefore comes from researchers running large sets of searches and recording when a summary appeared.

## Which kinds of searches trigger AI Overviews most often?

Questions and long, specific searches trigger them most; short keywords trigger them least.

Xu and colleagues ran 55,393 trending US searches over 40 days in March and April 2026. Searches that began with a question word showed an AI Overview 64.7% of the time, against 9.5% for the rest: a 6.8 times difference. Open-ended questions did best, with “how” at 84.3% and “why” at 73.4%. Simple lookups did less well, with “who” at 47.9% and “did” at 39.8%.

Length matters on its own too. Even among searches that were not questions, the rate climbed from 9.9% for single-word searches to 38.7% for searches of six or more words. A [2025 panel study by Pew researchers](https://arxiv.org/abs/2608.04831) found the same pattern in real searches by US adults: longer searches, searches starting with a question word and searches with both a noun and a verb tended to produce AI Overviews.

| Type of search | Showed an AI Overview | Study |
|---|---|---|
| Starts with “how” | 84.3% | Xu et al., trending US searches, 2026 |
| Any question word | 64.7% | Xu et al. |
| Starts with “did” | 39.8% | Xu et al. |
| No question word | 9.5% | Xu et al. |
| Complex “explain this” questions | 94.6% | Grossman et al., December 2025 |
| Amazon product keywords | 17.4% | Grossman et al. |

The last two rows come from [Grossman and colleagues](https://arxiv.org/abs/2604.27790), who ran 11,500 benchmark searches in December 2025 from a simulated mobile phone in Newark, New Jersey.

## Does the same topic get an AI Overview if it is worded differently?

Often yes: rewording a short keyword as a longer, specific question sharply raised the rate in two tests.

In [our AI Overview frequency study](https://underneath.agency/research/ai-overviews-frequency-study), we searched 96 commercial topics three ways on 28 September 2026. As written, 59.4% showed an AI Overview; as a long non-question phrase, 87.5%; as a natural question, 93.8%. Most of the lift came from making the search longer and more specific, not from the question word itself. Every one of the 31 keywords containing a price or cost word showed an AI Overview.

Grossman and colleagues saw the same effect at larger scale. The same Natural Questions set triggered an AI Overview 86.2% of the time when phrased as questions, and 76.5% when cut down to keywords. Wording can also change [which sources an AI answer cites](https://underneath.agency/resources/does-question-phrasing-change-ai-sources).

## Which searches rarely get an AI Overview?

Local map searches, brand and website searches, shopping keywords and breaking news rarely get one.

- **Local searches.** In our study, 21.8% of searches with a local map pack showed an AI Overview, against 81.1% without one. Searches containing “near me” showed one 17.3% of the time. Google seems to decide first whether a search needs a map or an explanation.
- **Brand and website searches.** In 2024, across seven countries, [Li and Aral](https://arxiv.org/abs/2504.06435) found questions returned an AI answer 49% of the time, statements 16% and navigational searches 4%. In the 2025 rerun by [Aral, Li and Zuo](https://arxiv.org/abs/2602.13415), the figures were 60%, 37% and 12%.
- **Shopping searches.** Only 5% of shopping searches returned an AI answer in 2024. By 2025 it was 13% worldwide.
- **Breaking news.** Grossman and colleagues found AI Overviews on only 8.1% of trending searches. The rate was 4.8% for topics trending for under five days and 12.7% after that.

## Does the topic or industry matter?

Yes, topic and industry matter, but usually less than how the search is worded.

[Huang and colleagues](https://arxiv.org/abs/2603.16138) ran 11,000 general-knowledge questions from one US location. Business questions produced AI Overviews 84% of the time, history 81% and education 74%. Sports (34%), technology (32%) and travel (32%) rarely did.

Xu and colleagues found a 13-fold spread across Google Trends categories, from 3.5% in beauty and fashion to 46.1% in hobbies and leisure. Politics (7.5%) and law and government (9.6%) sat below average, which fits Google’s stated caution on sensitive topics. Health (26.6%) did not, so that caution is not applied evenly.

Wording can override topic. Grossman and colleagues asked “What is the politician [NAME] known for?” for every member of the US Congress, and 93.8% produced an AI Overview. Politics is rarely summarized when the search is a trending name, but often when it is a direct question.

In our study of eight industries, B2B software keywords showed AI Overviews 96.0% of the time and financial services 88.0%. Legal, healthcare and franchise keywords were lower, largely because their results pages are dominated by local maps. Adjusting for that shrank the gap between industries from 60.0 to 36.7 points.

## So how often do AI Overviews appear overall?

Between about one search in seven and two in three, depending on which searches are counted.

| Study | Searches counted | When | Showed an AI Overview |
|---|---|---|---|
| Xu et al. | 55,393 trending US searches | Spring 2026 | 13.7% |
| Chapekis et al. (Pew) | Real searches by 900 US adults | March 2025 | 18% |
| Wang et al. | Everyday searches of 1,100 US volunteers | March 2026 | 36% |
| Grossman et al. | Representative real-user searches | December 2025 | 51.5% |
| Underneath | 800 US commercial keywords | September 2026 | 60.8% |
| Aral et al. | Fixed set of benchmark searches, US | 2025 | 67% |

These figures do not contradict each other. Trending news searches rarely get a summary; the research-heavy keywords businesses track often do. Weighting by search volume also changes the answer: in our sample, the rate fell from 60.8% to 40.8%, because a few huge searches were mostly local or navigational.

The rate has also been rising. Running identical searches a year apart, Aral and colleagues found 67% of US searches answered by AI in 2025, against 42% in 2024. The figure of 36% comes from the field experiment by [Wang and colleagues](https://arxiv.org/abs/2608.18352), a recent measure of what real people saw in everyday use.

## What should you do about it?

Measure AI Overview exposure on your own keyword list, split by type of search, before planning around any average.

1. Group your tracked keywords into questions and long phrases, short keywords, local searches, brand searches and shopping searches. Each group will behave differently.
2. Treat “how” and “why” searches as the most exposed: in the Xu study, 84.3% of “how” searches showed an AI Overview.
3. Track local searches separately. Where Google shows a map, your Google Business Profile still carries more weight than an AI summary.
4. Recheck regularly. In our study, 83.3% of keywords had the same outcome two days apart, so roughly one in six changed.
5. Do not read this as content advice. Question-shaped searches get more AI Overviews; no study here shows that question-shaped pages earn more citations. Whether you are cited once one appears is covered in [Google rankings and AI Overview citations](https://underneath.agency/resources/does-ranking-on-google-get-you-into-ai-overviews).

If you want help measuring this across your own searches, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The research shows when AI Overviews appear, not why Google decides, and the rules keep changing.

- Google publishes no triggering rules, so every finding is observed from outside.
- Most studies use one US location and one point in time. Results by device, signed-in user or country are thinly measured.
- Query sets differ: trending searches, old question collections, commercial keywords. Only the Pew panel counts what real people typed, and its data is from March 2025.
- Several studies used AI models to label query types or topics, which adds some error.
- No study shows that changing a page changes whether it is cited once an AI Overview appears.

## Frequently asked questions

### What percentage of Google searches show an AI Overview?

There is no single figure. Studies range from 13.7% of trending US searches to 67% of a fixed US benchmark set, and Pew’s March 2025 panel of real searches found about 18%.

### Do question searches trigger AI Overviews more often?

Yes, by a wide margin. Question-form trending searches triggered one 64.7% of the time against 9.5% for other searches, in a 2026 study by Xu and colleagues.

### Do “near me” searches show AI Overviews?

Rarely. In our September 2026 study, searches containing “near me” showed an AI Overview 17.3% of the time, because Google usually shows a map instead.

### Do brand searches show AI Overviews?

Rarely. Navigational searches, such as a company or website name, returned an AI answer 12% of the time in a 2025 global test, up from 4% in 2024.

### Are AI Overviews becoming more common?

On identical searches, yes. Aral and colleagues found 67% of US benchmark searches answered by AI in 2025, against 42% in 2024.

## Sources

- Xu, Iqbal and Montgomery (2026), [Measuring Google AI Overviews: Activation, Source Quality, Claim Fidelity, and Publisher Impact](https://arxiv.org/abs/2605.14021), arXiv:2605.14021.
- Grossman, Liu, Chen, Smith, Borcea and Chen (2026), [How Generative AI Disrupts Search: An Empirical Study of Google Search, Gemini, and AI Overviews](https://arxiv.org/abs/2604.27790), arXiv:2604.27790.
- Chapekis, Lieb, Shah and Smith (2026), [Investigating Click Behaviors On Google Search Result Pages That Produce an AI Overview](https://arxiv.org/abs/2608.04831), arXiv:2608.04831.
- Li and Aral (2025), [Human Trust in AI Search: A Large-Scale Experiment](https://arxiv.org/abs/2504.06435), arXiv:2504.06435.
- Aral, Li and Zuo (2026), [The Rise of AI Search: Implications for Information Markets and Human Judgement at Scale](https://arxiv.org/abs/2602.13415), arXiv:2602.13415.
- Huang and colleagues (2026), [Answer Bubbles: Information Exposure in AI-Mediated Search](https://arxiv.org/abs/2603.16138), arXiv:2603.16138.
- Wang, Gleason, Bart, Wilson and Metaxa (2026), [AI in Search Reduces Publisher Referrals Without Improving User Experience: Experimental Evidence](https://arxiv.org/abs/2608.18352), arXiv:2608.18352.
- Underneath (2026), [When does Google show an AI Overview? 1,248 US searches](https://underneath.agency/research/ai-overviews-frequency-study)

---

This is the Markdown twin of https://underneath.agency/resources/which-searches-trigger-ai-overviews. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Which types of websites do AI search engines rely on most?"
description: "It depends on the question. Official, reference and news sites lead on public-interest topics; review publishers lead on buying questions."
canonical: "https://underneath.agency/resources/which-websites-do-ai-search-engines-cite"
published: 2026-10-11
updated: 2026-10-11
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Which types of websites do AI search engines rely on most?

AI search engines rely most on established third parties: official bodies, reference sites like Wikipedia, news outlets and specialist review publishers. Which type leads depends on the topic and on the engine. Video, social media and forums matter far more in Google’s AI Overviews than in ChatGPT.

## The short version

1. On politics, health and environment questions, government websites took 33.0% of citations among the 25 most-cited domains, in [an audit of four AI engines](https://arxiv.org/abs/2605.23684).
2. For 100 health questions put to ChatGPT, institutions such as hospitals, government bodies and journals supplied 75.7% of cited sources, [one 2026 study found](https://arxiv.org/abs/2601.17109).
3. For product questions, editorial and product-review sites made up 56.7% of the domains ChatGPT displayed, according to [an audit of AI product recommendations](https://arxiv.org/abs/2609.18729).
4. Google’s AI leans on forums far more than ChatGPT: [our Reddit study](https://underneath.agency/research/ai-reddit-citations-study) found a Reddit thread in 17.9% of AI Overviews and in no ChatGPT answers.

## Which kinds of websites do AI search engines cite most?

Mostly established third parties: official bodies, reference works, news outlets and specialist publishers. [Zhang Kai and colleagues](https://arxiv.org/abs/2604.25707) analyzed a public dataset of 602 controlled prompts on ChatGPT, Google’s AI Overview and Gemini, and Perplexity. Official, news and vertical (specialist industry) sources accounted for 79.12% to 87.52% of citations across platforms.

Two caveats apply. Researchers wrote those prompts; they were not real user questions. The authors also warn that the source-type labels in the dataset are noisy.

Beyond that broad picture, the mix shifts with the subject. Public-interest topics pull engines toward official sources. Buying questions pull them toward review publishers. Google’s AI features add a large share of video and social media.

## Do government and institutional sites really dominate?

On public-interest topics, yes, but not everywhere. [Allaham and Diakopoulos](https://arxiv.org/abs/2605.23684) put 712 real user questions on politics, health and the environment to ChatGPT, Copilot, Gemini and Perplexity. Among the 25 most-cited domains, government websites took 33.0% of citations and Wikipedia alone took 15.2%. Social media, including Reddit and YouTube, took 8.8%. The same audit also checked [how many cited sources were AI-generated](https://underneath.agency/resources/are-ai-search-engines-citing-ai-content).

Health shows the same pattern. [Jacques and colleagues](https://arxiv.org/abs/2601.17109) coded the sources ChatGPT cited for 100 consumer health questions on one day in January 2026. Institutions with built-in authority supplied 75.7% of citations, while commercial health sites supplied 12.4% and professional or practice websites 11.9%. Ten organizations, led by Wikipedia, Mayo Clinic and Cleveland Clinic, accounted for 52.8% of all citations.

This is not a general rule. In [a comparison of 11,500 user queries](https://arxiv.org/abs/2604.27790), traditional Google Search was more likely than AI Overviews or Gemini to retrieve government and education sites. Official sources lead when the topic itself is official.

## How much do AI engines rely on Wikipedia, YouTube and Reddit?

A great deal, but it varies sharply by engine. [Xu and colleagues](https://arxiv.org/abs/2605.14021) tracked Google AI Overviews for 40 days in March and April 2026. The five most-cited sites were youtube.com (5.49%), en.wikipedia.org (4.39%), facebook.com (3.68%), instagram.com (3.65%) and usatoday.com (2.80%). Sports alone accounted for 33.6% of all AI Overview references, which may lift the share of video and social sources.

Our own data agrees on YouTube. In [our AI Overview citation study](https://underneath.agency/research/ai-overview-citations-study), YouTube was cited in 64.2% of AI Overviews. In our Reddit study, Perplexity cited a Reddit thread in 12.5% of 80 buyer questions, Gemini in 2.5%, and ChatGPT and Claude in none.

[A study of 11,000 real search questions](https://arxiv.org/abs/2603.16138) found the same split. A search-enabled OpenAI model took 0.1% of its [citations from social platforms](https://underneath.agency/resources/does-social-media-help-ai-search-visibility), against 8.5% for AI Overviews and 13.4% for regular Google results. The OpenAI model drew 27.3% of its citations from encyclopedias and reference sites.

## What do AI engines cite for product and buying questions?

Independent review and editorial publishers, more than brand or retail sites. [Uberti-Bona Marin and colleagues](https://arxiv.org/abs/2609.18729) put 117 physical-product questions to ChatGPT, Gemini and Google from the Netherlands in September 2026. Editorial and product-review sources made up 56.7% of the domains ChatGPT displayed and 45.2% in Gemini.

Google’s AI Overviews spread wider. There, editorial and review sources were 25.4% of displayed domains, user-generated and community sources 22.2%, retailers 17.8% and manufacturers and brands 16.1%. The authors labeled source types with an AI model and call these results exploratory.

An August 2025 study by [Chen and colleagues](https://arxiv.org/abs/2509.08919) points the same way. For US consumer electronics questions, AI search drew 92.1% of its sources from earned media, meaning independent reviews and publications. Google’s regular results gave brand sites 32.9% and social sites 15.4%. That study used developer versions of the engines, not the consumer apps, and thanked a GEO vendor for support.

## Do different AI engines favor different types of sites?

Yes, enough to change where a brand should invest. ChatGPT leans toward reference works, news agencies and review publishers. Google’s AI Overviews mix in more video, social media and forums. Engines also differ in [how many sources they cite per answer](https://underneath.agency/resources/how-many-sources-ai-search-engines-cite).

News citations show how concentrated one engine can be. [Yang](https://arxiv.org/abs/2507.05301) analyzed over 366,000 citations from an AI search comparison site in 2025; 9% pointed to news sources. For OpenAI models, the 20 most-cited news outlets took 67.3% of citations in that news sample, led by Reuters and AP News. Google and Perplexity models spread their news citations more widely.

A vendor dataset gives a commercial view. [Kumar](https://arxiv.org/abs/2606.20065), whose company sells AI visibility tracking, analyzed 149,912 citations for brand prompts. Outside company websites, YouTube led at 4.2%, ahead of tech and business media at 3.8%, Reddit and other forums at 3.3% and Wikipedia at 2.6%.

## Do AI engines rely on a few big sites or many small ones?

Both: a short list of repeat sources plus a very long tail. In the Allaham and Diakopoulos audit, [the 25 most-cited domains](https://underneath.agency/resources/ai-search-citation-concentration) took 23.8% of all citations. Most other domains received very few citations: 59.1% were cited only once.

AI Overviews look similar. In the Xu study, 56.2% of cited sites appeared exactly once in 40 days. The five most-cited sites took 20.0% of AI Overview citations, against 47.0% of links on Google’s first page of results. That long tail is why [smaller websites can still get cited](https://underneath.agency/resources/can-small-websites-get-cited-by-ai).

## What should you do about it?

Match your presence to the source types engines use for your kind of question.

1. Check which sources the engines cite for your buyers’ real questions. The mix differs by topic and by engine, so measure rather than assume.
2. For buying questions, earn coverage in independent review and editorial publications. They were the largest displayed source type in ChatGPT’s product answers.
3. If your buyers use Google, invest in video. YouTube was the most-cited site in AI Overviews in the Xu study.
4. Keep reference entries, such as Wikipedia, accurate where you legitimately qualify for one.
5. On regulated or public-interest topics, align your claims with official sources, because engines lean on them.

Our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) starts with this kind of source mapping.

## What does the research not tell us yet?

The research shows what engines cite, not why, and few studies cover business-to-business questions.

- Most studies are snapshots from one period, and engines change their sourcing often.
- Category labels differ between studies, so “earned”, “official” and “editorial” are not identical buckets.
- Several large datasets come from vendors or used researcher-written prompts rather than real users.
- Being cited is not the same as shaping the answer. One study found AI Overviews cite Reddit and Quora yet draw 22.1% less content from social platforms than from other sources.
- We found no study linking source types to sales or leads.

## Frequently asked questions

### Does ChatGPT cite Reddit?

Rarely, in the studies so far. In our Reddit study, ChatGPT cited no Reddit threads across 80 buyer questions, while 17.9% of Google AI Overviews did.

### Which website does Google’s AI Overview cite most?

YouTube, in the largest recent study. Xu and colleagues found youtube.com accounted for 5.49% of AI Overview citations, ahead of Wikipedia at 4.39%.

### Do AI search engines prefer government websites?

On public-interest topics, yes. Government sites took 33.0% of top-domain citations for politics, health and environment questions in one audit of four engines.

### Are brand websites cited for product questions?

Sometimes, but less than reviews. In Google AI Overviews for product questions, manufacturers and brands were 16.1% of displayed domains, behind editorial and review sources.

## Sources

- Allaham and Diakopoulos (2026), [Synthetic Sources?: Auditing Generative Search Engine Citations for Evidence of AI-Generated Sources](https://arxiv.org/abs/2605.23684), arXiv:2605.23684.
- Jacques et al. (2026), [Authority Signals in AI Cited Health Sources: A Framework for Evaluating Source Credibility in ChatGPT Responses](https://arxiv.org/abs/2601.17109), arXiv:2601.17109.
- Zhang Kai et al. (2026), [From Citation Selection to Citation Absorption: A Measurement Framework for Generative Engine Optimization Across AI Search Platforms](https://arxiv.org/abs/2604.25707), arXiv:2604.25707.
- Xu et al. (2026), [Measuring Google AI Overviews: Activation, Source Quality, Claim Fidelity, and Publisher Impact](https://arxiv.org/abs/2605.14021), arXiv:2605.14021.
- Uberti-Bona Marin et al. (2026), ["If I Had to Buy Just ONE: Galaxy S26 Ultra": Auditing AI-Generated Product Recommendations](https://arxiv.org/abs/2609.18729), arXiv:2609.18729.
- Chen et al. (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919), arXiv:2509.08919.
- Yang (2025), [News Source Citing Patterns in AI Search Systems](https://arxiv.org/abs/2507.05301), arXiv:2507.05301.
- Kumar (2026), [Generative Engine Optimization at Scale: Measuring Brand Visibility Across AI Search Engines](https://arxiv.org/abs/2606.20065), arXiv:2606.20065.
- Huang, Goyal, Saha and Chandrasekharan (2026), [Answer Bubbles: Information Exposure in AI-Mediated Search](https://arxiv.org/abs/2603.16138), arXiv:2603.16138.
- Grossman et al. (2026), [How Generative AI Disrupts Search: An Empirical Study of Google Search, Gemini, and AI Overviews](https://arxiv.org/abs/2604.27790), arXiv:2604.27790.
- Underneath (2026), [When does a Reddit thread become evidence in Google’s AI?](https://underneath.agency/research/ai-reddit-citations-study)
- Underneath (2026), [Do AI Overviews cite the pages that rank? 4,051 citations](https://underneath.agency/research/ai-overview-citations-study)

---

This is the Markdown twin of https://underneath.agency/resources/which-websites-do-ai-search-engines-cite. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Who in my organization should own AI search visibility?"
description: "No single team: SEO owns being found, content and product own what pages say, PR owns third-party coverage, and one person should own measurement."
canonical: "https://underneath.agency/resources/who-should-own-ai-search-visibility"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Who in my organization should own AI search visibility?

No single team can own it, because AI answers draw on your rankings, your own pages and what other sites say about you. The clearest published split gives SEO the job of getting found, content and product teams the pages themselves, and PR the third-party coverage. Someone also has to own the measurement, because AI visibility is easy to misread.

## The short version

1. A vendor-authored study recommends a three-way split: SEO drives whether you are found, content and product teams fix pages, and PR handles third-party outreach ([Vishwakarma and colleagues](https://arxiv.org/abs/2605.25517)).
2. SEO alone cannot cover it: 68.5% of the pages ChatGPT cited were not in Google’s top 100 for the question or for any of its searches ([our AI citations and Google rankings study](https://underneath.agency/research/ai-citations-google-rankings-study)).
3. Third-party coverage is a PR job with measurable weight: in 43.8% of its answers, ChatGPT ran a search aimed at a named publication, ranking or award ([our hidden searches study](https://underneath.agency/research/ai-hidden-searches-study)).
4. Your own facts need an owner: 18.9% of AI answers about local businesses stated a fact that differed from the business’s Google profile ([our business facts study](https://underneath.agency/research/ai-business-facts-accuracy-study)).
5. Measurement needs a referee: a single ChatGPT answer showed only 57.8% of the brands that five answers to the same question named ([our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study)).

## Is there one right owner for AI search visibility?

No: the work spans three existing functions, plus a measurement role. The most direct guidance comes from [Vishwakarma and colleagues](https://arxiv.org/abs/2605.25517), who work at the software vendor Sprinklr. They write that “SEO teams drive retrieval rank, content/product teams own on-page fixes, and PR/partnerships own third-party outreach.”

Treat that as a recommended workflow, not a tested result. Their evidence is a simulated test in which AI assistants chose between two versions of a product review. Their internal pilot of the workflow reported only positive qualitative feedback, with no measured outcomes.

The split still matches what the wider evidence shows. AI answers depend on whether your page is found, what it says, and what independent sources say about you. Each of those already has a natural home in most marketing organizations.

## What should the SEO team own?

SEO should own whether AI assistants can find and reach your pages at all. In the vendor study’s workflow, a brand absent from an answer’s citations has a finding problem, and the stated remedy is to improve SEO. In their tests, [four factors mattered strongly in all six assistants](https://underneath.agency/resources/why-ai-cites-competitor-page-first): topic match, a stated price, a recent date, and being listed first.

But Google rankings explain only part of what AI assistants find. In [our study of 80 US buyer questions](https://underneath.agency/research/ai-citations-google-rankings-study), 68.5% of the pages ChatGPT cited were not in Google’s top 100 for the question or for any of the searches it ran. ChatGPT ran a mean of 3.7 searches per answer, none of them the user’s exact words ([our hidden searches study](https://underneath.agency/research/ai-hidden-searches-study)). So the SEO brief widens: you need to rank for the searches assistants write, not only the ones people type. The SEO team will also want to know [whether optimizing for ChatGPT hurts Google](https://underneath.agency/resources/chatgpt-optimization-google-rankings).

Technical access also needs a clear owner. In [our crawler study](https://underneath.agency/research/ai-crawler-blocking-study), 51.7% of sites that block OpenAI’s training crawler still allow its search crawler. Most of those, 82.7%, never mention the search crawler at all, so the outcome is a default rather than a decision.

## What should content and product teams own?

They should own what your pages say: facts, prices, specifications, dates and evidence. The vendor study’s owned-content advice is concrete: surface core topic terms early, add explicit price and key specs, include comparisons, keep dates current, and replace hedging with evidence. Of the 18 factors tested, 11 made a reliable difference in at least four of six assistants; formatting-only changes had little effect.

Consistency across your own pages is where this team earns its keep. In [our pricing study](https://underneath.agency/research/ai-pricing-accuracy-study), 61.9% of plan prices that AI assistants quoted were fully faithful to the official pricing page. Of the 64 prices that differed, 39 matched a figure on another page of the vendor’s own site.

The same pattern shows up for basic business facts. Where a business’s Google profile phone number did not appear on its own website, 30.6% of the numbers AI answers gave differed from the profile. Where it did appear, only 1.6% did ([our business facts study](https://underneath.agency/research/ai-business-facts-accuracy-study)). Conflicting facts on your own properties travel straight into AI answers.

## What should PR and partnerships own?

PR should own what independent sources say about you, because AI assistants go looking for it. In 43.8% of its answers, ChatGPT ran a search aimed at a named publication, ranking or award. When a search named a source, the answer cited that source 44.0% of the time, against 8.1% when it did not ([our hidden searches study](https://underneath.agency/research/ai-hidden-searches-study)). That is an association, not proof of cause.

Independent coverage was the strongest signal in [our brand entity study](https://underneath.agency/research/brand-entity-ai-recommendations-study). Each tenfold increase in the number of independent sites naming a brand in the cited pages went with 4.7 times the odds of being recommended. The vendor study’s third-party advice fits this: give reviewers missing data, enable independent testing, and support side-by-side comparisons.

Reviews belong here too, or with customer experience. In [our reputation study](https://underneath.agency/research/is-it-legit-ai-reputation-study), 88.0% of answers to “is this brand legit?” cited a review or complaint platform, and claims drawn only from review platforms were negative 56.5% of the time. In a simulated hotel test, a top guest rating raised the chance of being recommended by 31.6 percentage points ([Baig and colleagues](https://arxiv.org/abs/2606.16344)).

## Who should measure AI visibility and referee results?

One named person, ideally outside the teams being judged, because AI visibility is easy to overstate. Answers vary from run to run. In [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), a single ChatGPT answer showed 57.8% of the brands its five answers to the same question named between them.

Traffic numbers need the same care. On one site studied by [Watanabe and Nakayashiki](https://arxiv.org/abs/2606.04362), pages nobody optimized still grew 3.5 times in ChatGPT referrals, simply because ChatGPT grew. A team measuring its own work against last quarter would claim that growth as its result. The same site was used to test [whether ChatGPT citations boost Google rankings](https://underneath.agency/resources/do-chatgpt-citations-boost-google-rankings).

Measurement does not have to be expensive. Researchers proposing independent audits of AI answers estimate that running them through public interfaces costs roughly $50–$300 at current pricing ([Wen and colleagues](https://arxiv.org/abs/2606.12439)).

## What should you do about it?

Give each part of the work a named owner and one shared scorecard. A practical setup:

1. **SEO:** owns being found. That covers rankings for the searches AI assistants write, crawler access rules, and fixing pages that never appear in citations.
2. **Content and product marketing:** owns on-page facts. One source of truth for prices, specifications, contact details and dates, kept identical across every page.
3. **PR and partnerships:** owns independent coverage, reviews, rankings, awards and comparison pieces that AI assistants search for.
4. **Customer experience or operations:** owns review volume and ratings on the platforms AI assistants cite.
5. **A measurement lead:** owns a fixed list of buyer questions, asked repeatedly, with results compared against pages nobody changed.

Hold a short monthly review where the measurement lead reports and each owner takes one action. If you want outside help running that process, see [our generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

No study has tested which ownership model produces better results, so the split above is reasoned, not proven.

- The three-way split comes from a vendor’s recommended workflow, backed by a simulated two-source test and an internal pilot with only qualitative feedback.
- Our studies are observational snapshots of AI answers on given days. They show where citations come from, not what changes them.
- The hotel finding uses fictional hotels in a controlled test, not live booking data.
- Nobody has published evidence on budget splits, team sizes or reporting lines for this work.

## Frequently asked questions

### Should the SEO team own GEO?

The SEO team should own part of it, not all of it. In our study, 68.5% of the pages ChatGPT cited were not in Google’s top 100 for the question or its own searches, so rankings alone cannot explain visibility.

### Is AI search visibility a PR job?

Partly, because AI assistants actively look for independent sources. ChatGPT ran a search aimed at a named publication, ranking or award in 43.8% of its answers in our study.

### Who should fix wrong information about our company in AI answers?

Whoever owns your website and listings facts, usually content or operations. In our study, conflicting phone numbers across a business’s own sources went with far more mismatches in AI answers: 30.6% against 1.6%.

### How do we know whether the work is paying off?

Assign one measurement lead and compare optimized pages with pages you left alone. On one site, untouched pages grew 3.5 times in ChatGPT referrals, so raw growth says little.

## Sources

- Vishwakarma, Kumar and Jamidar (2026), [What Gets Cited: Competitive GEO in AI Answer Engines](https://arxiv.org/abs/2605.25517), arXiv:2605.25517.
- Watanabe and Nakayashiki (2026), [Disentangling Answer Engine Optimization from Platform Growth: A Log-Based Natural Experiment on ChatGPT Referral Traffic](https://arxiv.org/abs/2606.04362), arXiv:2606.04362.
- Wen, Zhang, Yuan, Chen, Zhang and Guo (2026), [Position: Generative Engine Optimization Creates Underexamined Risks, Governance Must Target Concentration, Disclosure, and Academic Blind Spots](https://arxiv.org/abs/2606.12439), arXiv:2606.12439.
- Baig, Gillani and Ali (2026), [Whose hotel does the AI recommend? An algorithm audit of reputation signals in LLM-assisted hotel selection](https://arxiv.org/abs/2606.16344), arXiv:2606.16344.
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude cite pages that rank?](https://underneath.agency/research/ai-citations-google-rankings-study)
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [Which AI crawlers do top websites block? 10,000 sites, 2026](https://underneath.agency/research/ai-crawler-blocking-study)
- Underneath (2026), [How faithfully do AI assistants quote software prices?](https://underneath.agency/research/ai-pricing-accuracy-study)
- Underneath (2026), [Do AI answers match a business’s Google profile?](https://underneath.agency/research/ai-business-facts-accuracy-study)
- Underneath (2026), [Do Wikipedia and schema make AI assistants recommend a brand?](https://underneath.agency/research/brand-entity-ai-recommendations-study)
- Underneath (2026), [“Is this brand legit?” How AI assistants build a reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study)
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)

---

This is the Markdown twin of https://underneath.agency/resources/who-should-own-ai-search-visibility. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Why ChatGPT gives a different answer about your brand each time"
description: "AI assistants pick each answer partly by chance, so a brand named once may vanish on the next run. Judge visibility by rates over many runs."
canonical: "https://underneath.agency/resources/why-ai-answers-about-your-brand-change"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Why does ChatGPT give a different answer about my brand each time?

Because ChatGPT writes each answer by sampling, so asking the same question twice can produce different brands, sources and wording. In one detailed study, about a third of the movement between identical questions was this built-in chance alone. A single answer is one draw from a range of possible answers, so it should never drive a decision on its own.

## The short version

1. In one study, chance alone accounted for 34.8% of the variation between answers to identical questions, even at a low randomness setting ([Żatuchin, 2026](https://arxiv.org/abs/2607.13304)).
2. In our test of 20 buyer questions asked five times, only 25.2% of the brands ChatGPT named appeared in every run, and 36.6% appeared once.
3. In a Dutch audit, ChatGPT’s product picks overlapped by 0.178 between repeats of the same question, against 0.421 for Google’s AI Overviews ([Uberti-Bona Marin and colleagues, 2026](https://arxiv.org/abs/2609.18729)).
4. ChatGPT’s first-named brand changed at least once for 80.0% of our questions over five runs.

## Is ChatGPT broken when its answer changes?

No; varying answers are how these systems work. An AI assistant builds each answer word by word, choosing among likely options with some randomness. With web search switched on, it may also run different searches and read different pages each time.

[Żatuchin](https://arxiv.org/abs/2607.13304) measured how much of the movement comes from this chance alone. Three AI models answered the same questions about 20 Central and Eastern European brands about five times each. When everything else was held fixed, chance accounted for 34.8% of the variation in the answers’ scores. The models ran at a low randomness setting, so the author treats that figure as a floor; default app settings are likely to vary more.

The author is affiliated with Rankfor.AI, which sells brand monitoring, and the score measured was the tone of each answer. Even so, the finding matches every other study below: identical questions do not give identical answers.

## How much do the answers actually change?

Enough that one answer shows only about half of what the assistant tends to say. In our [consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), we asked ChatGPT, Gemini and Perplexity 20 buyer questions five times each, with the runs a median of about 11 minutes apart.

| Brands named across five runs | ChatGPT | Gemini | Perplexity |
|---|---|---|---|
| Named in all five | 25.2% | 13.7% | 40.9% |
| Named in only one | 36.6% | 46.7% | 15.8% |
| Share of five-run brands seen in a single run | 57.8% | 48.4% | 72.2% |

So most brands sit in a long tail that comes and goes. A small core appears nearly every time.

An independent audit found the same pattern for shopping questions. [Uberti-Bona Marin and colleagues](https://arxiv.org/abs/2609.18729) sent 117 real product questions to the ChatGPT and Gemini apps and to Google, three times each, from the Netherlands. ChatGPT’s recommended products overlapped by only 0.178 between repeats, on a scale where 1 means identical. That was about as different as ChatGPT was from Google’s answers.

## Is it chance, or did something change on the web?

Mostly chance; the pages behind the answers rarely change that fast. [Sielinski](https://arxiv.org/abs/2603.08924) stored a fingerprint of each cited page he could reach during nine days of repeated questions to three AI search engines. Most pages did not change at all between collections, yet [citations still swung](https://underneath.agency/resources/why-ai-search-citations-change).

His conclusion was that swings of this size, including rank changes within a single four-hour window, are built into the engines rather than caused by edits to the pages. On daily data, his uncertainty ranges for a site’s share of citations were 3 to 7 percentage points wide.

A Swiss study pointed the same way. When the same prompts were asked within 24 hours, cited sources overlapped by only 0.32 to 0.43, about the same as from one day to the next ([Schulte, Bleeker and Kaufmann, 2026](https://arxiv.org/abs/2604.07585)). In our study, Perplexity cited identical pages in 166 pairs of runs, and the brand list still changed in 91.6% of them.

## Which parts of an answer change most?

Rankings and wording move more than the basic fact of whether you are mentioned. In our study, ChatGPT’s first-named brand changed at least once for 80.0% of questions over five runs. Even among brands named every time, most moved position.

The framing shifts too. In the Dutch audit, whether an assistant labeled a product “the best” changed across repeats for between 27% and 40% of questions, depending on the assistant. Whether an answer named any product at all changed far less often.

A study by Ranqo, a company that sells AI visibility tracking, found the same split across more than 100 brands. For brand, question and engine combinations tracked over several runs, the tone flipped between positive and negative in 45.5% of them. Whether the brand was mentioned flipped in only 6.8% ([Kumar, 2026](https://arxiv.org/abs/2606.20065)). Most of those combinations were small brands that were never mentioned, which keeps the second figure low. Still, tone is the noisiest thing you can track.

## Do other AI assistants vary as much as ChatGPT?

No; each assistant has its own level of consistency. In our study, Perplexity was the steadiest, naming 40.9% of its brands in every run. Gemini was the least steady at 13.7%. In the Dutch audit, Google’s AI Overviews were the most consistent, with product picks overlapping by 0.421 between repeats. Which engine looks steadiest depends on the measure; see [which AI engine is most consistent](https://underneath.agency/resources/most-consistent-ai-engine-for-brands).

Repeats also keep turning up new sources. In the Dutch audit, the number of distinct websites ChatGPT showed for a question rose from 2.65 after one request to 5.56 after three. Even Google varies in whether it shows an AI answer at all: an earlier audit cited by Sielinski found a 33% inconsistency rate in whether AI Overviews appeared for the same query.

A survey of the field adds that randomness is not the whole story. It cites an audit in which repeated runs with the randomness setting at zero still changed 9 to 28% of decisions ([Martinez, 2026](https://arxiv.org/abs/2607.14035)).

## What should you do about it?

Treat one AI answer as one roll of the dice, and judge your brand over many runs. In practice:

1. Never act on a screenshot. A missing mention in one answer, or a sudden first place, is weak evidence on its own.
2. Ask each important question many times and report a rate, such as “named in 7 of 10 runs”, with the number of runs beside it.
3. Track separately whether you are named, whether you are near the top and how you are described. These move at very different speeds.
4. Aim to be in the core: the brands an assistant names nearly every time. Moving out of the long tail is a real gain; a single appearance is not.
5. Compare assistants separately. A brand can be steady on Perplexity and erratic on Gemini.
6. When a number moves, check it against the normal run-to-run range before you treat it as news. Our guide on [whether a GEO campaign really worked](https://underneath.agency/resources/did-geo-improve-ai-citations) shows how.

If you want a tracking setup built on repeated runs, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

We know the answers vary; we know less about how much is enough to measure. Gaps:

- No study yet separates exactly how much variation comes from the search step and how much from the writing step inside the assistant.
- Most repeat tests are short: minutes to nine days. How a brand’s steady core changes over months, through model updates, is barely measured.
- Several studies here come from companies that sell AI visibility tools. Their data are useful, but independent checks are fewer.
- The Dutch audit used three repeats per question; our study used five. Both say these are too few to capture every brand an assistant might name.
- Logged-in users with chat history may see different answers again. The studies here used logged-out sessions or developer access.

## Frequently asked questions

### Why does ChatGPT recommend different brands every time I ask?

Because it generates each answer with some randomness and may search different pages each time. In our five-run test, only 25.2% of the brands ChatGPT named for a question appeared in all five answers.

### Is one ChatGPT answer a reliable measure of my brand’s visibility?

No. In our test a single ChatGPT answer showed only 57.8% of the brands its five answers named between them, so one answer misses a large share of the picture.

### Does Google’s AI Overview change as much as ChatGPT?

Less, in the evidence so far. In a Dutch audit of product questions, AI Overview picks overlapped by 0.421 between repeats, against 0.178 for ChatGPT.

### If the AI answer about my brand changed, did a website change?

Usually not. One study that fingerprinted the cited pages found most pages unchanged while citations still swung, which points to the engine rather than the web.

## Sources

- Żatuchin (2026), [Where Does the Noise Come From? A Variance-Components Decomposition of Non-Determinism in LLM Brand Answers](https://arxiv.org/abs/2607.13304), arXiv:2607.13304.
- Uberti-Bona Marin and colleagues (2026), ["If I Had to Buy Just ONE: Galaxy S26 Ultra": Auditing AI-Generated Product Recommendations](https://arxiv.org/abs/2609.18729), arXiv:2609.18729.
- Sielinski (2026), [Quantifying Uncertainty in AI Visibility: A Statistical Framework for Generative Search Measurement](https://arxiv.org/abs/2603.08924), arXiv:2603.08924.
- Schulte, Bleeker and Kaufmann (2026), [Don’t Measure Once: Measuring Visibility in AI Search (GEO)](https://arxiv.org/abs/2604.07585), arXiv:2604.07585.
- Kumar (2026), [Generative Engine Optimization at Scale: Measuring Brand Visibility Across AI Search Engines](https://arxiv.org/abs/2606.20065), arXiv:2606.20065.
- Martinez (2026), [Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026)](https://arxiv.org/abs/2607.14035), arXiv:2607.14035.
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)

---

This is the Markdown twin of https://underneath.agency/resources/why-ai-answers-about-your-brand-change. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "What decides whether AI cites your page over a competitor’s?"
description: "In a 252,000-trial test, topic match, list position, a stated price and a recent date decided which page AI cited first. Formatting barely mattered."
canonical: "https://underneath.agency/resources/why-ai-cites-competitor-page-first"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# What decides whether an AI engine cites my page over a competitor’s?

In the largest controlled test so far, four things decided which competing page an AI assistant cited first: topic match, list position, a stated price and a recent date. Completeness and trust signals helped less, and formatting barely mattered. The test was a simulation run by a software company, so read it as a strong signal, not a law.

## The short version

1. Four “gatekeepers” won in all six AI systems tested: topic match, list position, a stated price and a recent date, in [a 252,000-trial simulation](https://arxiv.org/abs/2605.25517) by researchers at Sprinklr.
2. Eleven of 18 content factors (61%) mattered in at least four of the six systems; layout and formatting changes did not.
3. On live Google, position still leads: 41.7% of pages ranking 1 to 3 were cited by AI Overviews, against 20.1% at positions 7 to 10, in [our study of 3,096 ranking pages](https://underneath.agency/research/ai-overview-cited-pages-study).
4. A review of 45 studies by [Martinez](https://arxiv.org/abs/2607.14035) found relevance and position were the most reproducible levers, while generic rewrite tricks transferred poorly.

## How did researchers test which page gets cited first?

They gave AI systems two near-identical pages that differed in one detail, then recorded which page was cited first. The study, by Vishwakarma and colleagues at Sprinklr, a customer-experience software company, is the cleanest head-to-head test published so far.

They started from 100 anonymized product review articles across 50 categories, such as consumer tech and fitness equipment. From these they built 1,440 scenarios. In each one, two versions of a page matched in facts, prices and length, but differed in exactly one of 18 factors.

Brand and publisher names were replaced with invented ones, so fame could not sway the result. The order of the two pages was swapped to cancel out any bias toward whichever came first. Six AI systems, including GPT-5.2, Gemini 2.5 Flash and Claude 3.5 Sonnet, ran 252,000 trials in total.

## What are the four gatekeepers?

The gatekeepers are topic match, list position, a stated price and a recent date. All six AI systems agreed on these four, and the effects were so large that the weaker page was nearly shut out.

In the authors’ words, failing on any one “can eliminate citation odds regardless of other content strengths.” The topic test compared a page about the products asked about with one discussing unrelated products. The date test compared content dated 2026 versus 2019.

List position is different from the other three. It is the slot your page lands in when the engine gathers sources, and you cannot set it on the page itself. It comes from how well the page is found and ranked in the first place, which is still a search problem.

| Factor | What the better page had | Who controls it |
|---|---|---|
| Topic match | Covers the exact products asked about | Your content team |
| List position | Appears first among the sources | Search ranking and retrieval |
| Price | States the price plainly | Your content and pricing teams |
| Recent date | Dated this year, not years ago | Your content team |

## What matters once the basics are covered?

Completeness, trust and comparisons matter next, but less. Seven secondary factors helped in most systems once the four gatekeepers were met.

They fall into three groups. Completeness means listing specifications and covering the topic in depth.

Trust means confident wording instead of hedging, claims backed by evidence such as tests or certifications, and no internal contradictions. Competitive positioning means using the words the question uses and comparing the product with alternatives.

The systems also differed in how picky they were. Kimi K2 responded to 83% of the factors, while Claude 3.5 responded to 50% and Gemini 2.5 to 33%. The authors give a telling example of a weak page: a product description that ends with “Contact us for pricing details.”

## Does formatting or tone change the outcome?

Barely: formatting changes had no consistent effect, and tone mattered in only some systems. Seven factors (39%) had weak or no effects across the six systems.

[Turning a dense paragraph into organized sections](https://underneath.agency/resources/does-content-structure-increase-ai-citations) did not move citation in a consistent way. A promotional tone, a weaker value proposition and weaker social proof mattered in only two or three of the six systems. The authors judged that too few to call a pattern.

Our own data on live Google points the same way.

AI Overviews are the AI summaries at the top of Google’s results. In [our study of pages they cite](https://underneath.agency/research/ai-overview-cited-pages-study), an HTML table added just +0.6 points to the chance of being cited. Organization schema, a tag that describes the company behind a site, added +0.5 points, which is no meaningful difference.

## Does the same pattern hold on live AI search?

Position clearly does, and relevance holds up across many studies. The other gatekeepers have not been tested as cleanly outside a lab.

On Google, [our AI Overview citation study](https://underneath.agency/research/ai-overview-citations-study) found the first organic result was cited in 49.5% of AI Overviews, and the ninth in 15.5%. In our study of 3,096 ranking pages, position explained more of the choice than 14 page features together, and 91.4% was left unexplained by anything we measured.

AI assistants are less tied to Google’s list. In [our comparison of AI citations with Google rankings](https://underneath.agency/research/ai-citations-google-rankings-study), only 8.3% of the pages ChatGPT cited ranked in Google’s top 10 for the question, against 25.7% for Claude. ChatGPT ran 3.7 searches per answer in [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study), so its “list position” comes from its own searches, not yours.

The wider literature agrees on the order of priorities. The Martinez review reports that in one benchmark, only three of 54 [combinations of rewrite method and subject area](https://underneath.agency/resources/do-geo-content-tactics-work) were clearly positive.

In another end-to-end test, rewriting only the body of a page reduced final citation by about 6%. The rewrite made the page harder to find in the first place. For the practices that do hold up, see [what GEO research supports](https://underneath.agency/resources/what-geo-practices-does-research-support).

## What should you do about it?

Fix the four gatekeepers before anything else, then work on completeness and trust. In practice:

1. Match each important page to the exact question buyers ask, and answer it near the top.
2. State prices on the page in plain numbers. In [our pricing study](https://underneath.agency/research/ai-pricing-accuracy-study), only 61.9% of plan prices quoted by AI assistants were fully faithful to the vendor’s page, so clear pricing also protects accuracy.
3. Keep visible dates current, and only change a date when the content really changes.
4. Treat ranking and retrieval as part of the job, since list position is a gatekeeper you cannot fix with copy.
5. Then add specifications, comparisons and evidence for your claims, and remove hedging words.
6. Do not spend a redesign budget on layout alone; the evidence does not support it.

If you want help turning this into a page-by-page plan, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The strongest evidence comes from one simulated, vendor-run test, so several questions remain open.

- **Bigger source lists.** The test used only two pages at a time. The authors note real systems often pull in “five to ten or more pages”, so crowded results are untested.
- **Real brands.** Brands were anonymized on purpose. How much a famous name or trusted domain changes the outcome is unknown.
- **Other industries.** The pages were consumer product reviews. Services, B2B software and local businesses were not tested.
- **Beyond the first citation.** The study measured which page was cited first, not whether the brand was recommended, or whether anyone clicked.
- **Independence.** The authors work at Sprinklr, the rewrites were generated by an AI system, and the method was piloted inside the company.

## Frequently asked questions

### Does putting prices on my website help AI assistants cite it?

In a simulation, yes: a stated price was one of four factors that decided the first citation in all six AI systems tested. That test used anonymized product reviews, so it has not been confirmed for live engines or for services.

### Does page formatting affect AI citations?

The evidence says formatting alone does little. In the Sprinklr test, layout changes had no consistent effect, and in our Google study an HTML table added only +0.6 points to the chance of being cited.

### Can a smaller site beat a bigger competitor in AI answers?

Possibly, though the research has only tested this in the lab. In the original study of [Aggarwal and colleagues](https://arxiv.org/abs/2311.09735), as summarized by Martinez, adding citations helped the fifth source gain 115.1% in visibility while the first lost 30.3%.

### Is a recent date enough to win the citation?

No. A recent date mattered in the simulation, but live results are mixed. In [our freshness study](https://underneath.agency/research/ai-source-freshness-study), among dated pages on the same Google results page, a recent date made no difference to AI Overview citation (−4.1 points).

### What is a gatekeeper factor in AI citation?

It is a condition that, if failed, can stop a page from being cited first regardless of its other strengths. The Sprinklr study found four: topic match, list position, price and a recent date.

## Sources

- Vishwakarma, Kumar and Jamidar (2026), [What Gets Cited: Competitive GEO in AI Answer Engines](https://arxiv.org/abs/2605.25517), arXiv:2605.25517.
- Martinez (2026), [Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026)](https://arxiv.org/abs/2607.14035), arXiv:2607.14035.
- Aggarwal and colleagues (2023), [GEO: Generative Engine Optimization](https://arxiv.org/abs/2311.09735), arXiv:2311.09735.
- Underneath (2026), [What pages cited by AI Overviews have in common](https://underneath.agency/research/ai-overview-cited-pages-study)
- Underneath (2026), [AI Overview citations and page-one results](https://underneath.agency/research/ai-overview-citations-study)
- Underneath (2026), [AI citations and Google rankings](https://underneath.agency/research/ai-citations-google-rankings-study)
- Underneath (2026), [The hidden searches AI assistants run](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [How accurately do AI assistants quote prices?](https://underneath.agency/research/ai-pricing-accuracy-study)
- Underneath (2026), [How fresh are the pages AI engines cite?](https://underneath.agency/research/ai-source-freshness-study)

---

This is the Markdown twin of https://underneath.agency/resources/why-ai-cites-competitor-page-first. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "What makes ChatGPT or Gemini recommend one hotel over another?"
description: "Guest rating and price decide most AI hotel picks; eco-certification and review count help, list order matters, and replying to reviews did nothing."
canonical: "https://underneath.agency/resources/why-ai-recommends-one-hotel-over-another"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# What makes ChatGPT or Gemini recommend one hotel over another?

Guest rating and price decide most of it: in a controlled audit of twelve AI models, a top rating raised a hotel’s chance of being picked by 31.6 points and a high price cut it by 30.0. Eco-certification and review count helped, the order of the list mattered, and replying to reviews made no difference. The evidence comes from simulated hotel cards, so it shows how assistants weigh the details they see, not which hotels they find.

## The short version

1. A 4.7-star rating instead of 3.9 raised a hotel’s chance of being recommended by 31.6 points, and a $249 price instead of $129 lowered it by 30.0 ([Baig and colleagues](https://arxiv.org/abs/2606.16344)).
2. Eco-certification added 11.6 points and a large review count 8.3 points, while a visible management response added 0.1, no detectable effect.
3. Being listed first was worth $11.7 per night on average, and one Gemini model favored the first slot by about 26 points.
4. When Gemini answered Tokyo hotel questions, booking sites supplied 55.3% of its cited sources ([Zhu and Chang](https://arxiv.org/abs/2603.20062)).
5. In our study of real ChatGPT answers about local businesses, more reviews than the local median went with a 19.5-point higher chance of being listed ([our local picks study](https://underneath.agency/research/chatgpt-local-picks-google-profile-study)).

## Which hotel details move an AI assistant’s pick the most?

Guest rating and price: together they outweigh every other detail an assistant was shown.

[Baig and colleagues](https://arxiv.org/abs/2606.16344) ran a pre-registered audit of twelve AI models. Each was asked to recommend one of five invented hotels. Every hotel card listed a rating, review count, date of the latest review, whether management replies to reviews, chain or independent, nightly price, and whether it held a Green Key eco-certification. Each detail was assigned at random, so its effect can be read directly.

The panel included OpenAI’s GPT-4o-mini, three Google Gemini models, four Anthropic Claude models and four smaller open models. The study made 61,459 model calls in total. Here is how each detail changed the chance that a hotel was the one recommended:

| Detail on the hotel card | Change in chance of being picked |
|---|---|
| Rating 4.7 instead of 3.9 | +31.6 points |
| Price $249 instead of $129 | −30.0 points |
| Green Key eco-certification | +11.6 points |
| 2,100 reviews instead of 45 | +8.3 points |
| Latest review 3 days old, not 11 months | +1.6 points |
| Part of a major chain | −1.8 points |
| “Management responds to guest reviews” | +0.1 points (no detectable effect) |

All twelve models preferred higher-rated, cheaper hotels. They disagreed sharply on eco-certification: its effect ranged from almost nothing to 29.9 points depending on the model. The models’ own stated reasons are a separate matter, covered in [whether AI explanations can be trusted](https://underneath.agency/resources/ai-explanations-for-recommendations).

## What is each detail worth in dollars per night?

Because price was randomized too, each detail can be priced: a full rating step was worth about $126 a night.

The researchers converted every effect into the nightly price change that would cancel it out. Moving from 3.9 to 4.7 stars was worth $126.4 per night. Eco-certification was worth $46.4, a large review count $33.2 and fresh reviews $6.2. Chain membership cost $7.2. A visible management response was worth $0.5 per night, which is indistinguishable from nothing.

These are [the assistant’s trade-offs, not a guest’s](https://underneath.agency/resources/ai-vs-human-reputation-priorities). They tell a revenue manager where the AI channel puts its weight.

## Does the order of the list matter?

Yes: on average a hotel listed first was picked more often, and one model leaned on order heavily.

Across the panel, a hotel in fifth place was recommended 3.7 points less often than the same hotel in first place. Being listed first was worth $11.7 per night on average, and as much as $18.8 for the business traveler. That gain comes from placement alone, with nothing about the property changed.

The average hides one outlier. Gemini 2.0 Flash showed a first-position advantage of about 26 points, roughly ten times the panel average. Order is usually set by the booking site, search tool or interface feeding the assistant, which a hotel does not control. Our guide on [list order in AI recommendations](https://underneath.agency/resources/does-list-order-change-ai-recommendations) covers tests beyond hotels.

## Does the traveler’s request change the weights?

Yes: the same details counted differently for a budget family, a business traveler and an eco-minded couple.

The audit used three traveler personas. Eco-certification was worth $65.4 per night to the eco-conscious couple against $36.8 for the budget family. A top rating was worth $98 per night for the family against $153 for the business traveler, because the family was more price-sensitive. The chain penalty was steepest for the eco-couple and vanished for the business traveler.

So a hotel that courts one kind of guest should make the details that guest cares about plain on every listing.

## Where does the assistant get its hotel information in the first place?

Mostly from booking sites, at least in one Gemini audit of Tokyo hotels.

The audit above handed each model a fixed list. Real assistants search first. [Zhu and Chang](https://arxiv.org/abs/2603.20062), of Blossom AI, a company, put 156 Tokyo hotel questions to Gemini 2.5 Flash with Google Search in March 2026. Online travel agencies such as Booking.com and Expedia supplied 55.3% of the 1,357 cited sources. Hotels’ own websites supplied 8.2% of citations in English and 11.0% in Japanese.

In a small check of 14 hotel websites, the hotels Gemini cited directly had deeper content, such as long FAQ pages and neighborhood guides. That check shows association, not cause.

## Do real assistants weigh reviews the same way?

The closest real-world evidence says yes for review volume, though it covers local services, not hotels.

In the audit, the API-served models answered almost identically when the same set was repeated. Real consumer apps add web search, and their answers vary more. We checked live ChatGPT answers to 120 local questions, such as the best dentist in a city, over two days in September 2026 ([our local picks study](https://underneath.agency/research/chatgpt-local-picks-google-profile-study)).

ChatGPT listed 67.7% of businesses ranked 1 to 3 in Google Maps. At the same Maps rank, more reviews than the local median went with a 19.5-point higher chance of being listed, and a rating of 4.8 or more added 9.9 points. Our study shows association, not cause, but it points the same way as the controlled audit: rating and review volume travel with being recommended.

## What should you do about it?

Treat the AI channel as a buyer who reads your rating, price and certifications, but not your replies.

1. **Protect the rating and grow review volume.** These are the largest levers in the audit and in our live data.
2. **Price with the AI trade-offs in mind.** A higher rate costs recommendations unless the rating and other signals justify it.
3. **State real certifications on every listing.** Eco-certification counted heavily with several models, especially for eco-minded travelers. Only claim what you hold.
4. **Keep facts identical across booking sites and your own site.** Assistants read listings you do not own.
5. **Build pages that answer guest questions.** Deep FAQ and area guides went with direct citation in the Tokyo audit.
6. **Check the answers regularly.** Ask the main assistants the questions your guests ask, several times, and note who is named and why.

For help setting up that monitoring, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The research shows how assistants choose among hotels they are given, not how a hotel gets into the list.

- The hotel cards were synthetic, with no photos, review text or amenities, and every hotel was a 4-star near the city center.
- The audit tested models through developer access in English with US-dollar prices, not the consumer ChatGPT or Gemini apps.
- It covered one turn of conversation; follow-up questions could change the weights.
- The Tokyo citation audit covered one engine, one city and one month, and measured citations, not bookings.
- Model versions change often, and the authors warn the weights may drift.

## Frequently asked questions

### Does replying to reviews help a hotel get recommended by AI?

Not directly in the evidence so far. A line saying management replies to reviews had no detectable effect across twelve models, worth about $0.5 per night.

### Do AI assistants prefer chain hotels?

No. In the audit, chain membership slightly lowered the chance of being picked, by 1.8 points, and the penalty was largest for eco-minded travelers.

### Does eco-certification matter for AI hotel recommendations?

Yes, more than many managers expect. Green Key certification raised the chance of being picked by 11.6 points across the panel, though some models gave it almost no weight.

### Can a hotel get cited by ChatGPT or Gemini instead of Booking.com?

Sometimes. In Gemini’s Tokyo answers, hotels’ own sites were 8.2% of English citations, and the cited hotels tended to have deep, question-answering pages.

## Sources

- Baig, Gillani and Ali (2026), [Whose hotel does the AI recommend? An algorithm audit of reputation signals in LLM-assisted hotel selection](https://arxiv.org/abs/2606.16344), arXiv:2606.16344.
- Zhu and Chang (2026), [The End of Rented Discovery: How AI Search Redistributes Power Between Hotels and Intermediaries](https://arxiv.org/abs/2603.20062), arXiv:2603.20062.
- Underneath (2026), [Which Google Maps businesses does ChatGPT recommend?](https://underneath.agency/research/chatgpt-local-picks-google-profile-study)

---

This is the Markdown twin of https://underneath.agency/resources/why-ai-recommends-one-hotel-over-another. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Why AI Search Engines Cite Different Sources Each Time"
description: "Because AI search engines pick sources partly at random each time. In one 45-day study, about 65% of cited sources changed from one day to the next."
canonical: "https://underneath.agency/resources/why-ai-search-citations-change"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Why do AI search engines cite different sources every time I check?

Because AI search engines choose their sources partly at random each time they answer, even when nothing on the web has changed. In a 45-day study of four engines, about 65% of cited sources changed from one day to the next, and asking the same question minutes apart produced almost as much change. The churn is built into how these systems work, so it is not a sign that your content suddenly got better or worse.

## The short version

1. In a study of four AI engines over 45 days, roughly 65% of cited sources changed from one day to the next (Schulte and colleagues, University of St. Gallen).
2. Re-asking the same question within 24 hours produced similar churn, with only 32% to 43% of sources shared, so most of the change is built-in randomness rather than news or website edits.
3. Each engine has its own level of stability: repeated runs shared about 0.30 of cited websites on Gemini, 0.40 on ChatGPT search and 0.50 on Perplexity (Sielinski, who works for an AI visibility company).
4. ChatGPT ran a mean of 3.7 hidden web searches per buyer question in our study, and none of 509 searches repeated the user’s question word for word, which gives the answer many ways to drift.
5. Which engine you ask matters far more than which day: one vendor study found same-engine sources a day apart were about 42 times more alike than two engines’ sources on the same day.

## How much do AI citations change from one check to the next?

A lot: most cited sources are different from one day to the next. [Schulte and colleagues](https://arxiv.org/abs/2604.07585) tracked ChatGPT, Gemini, Google AI Mode and Perplexity every day for 45 days in early 2026, using German-language shopping and service questions sent from Swiss servers. Across 4,044 pairs of consecutive days, only about 35% of cited sources overlapped, so roughly 65% changed overnight.

Google’s AI Overviews, the AI summaries at the top of Google’s results, change too. [Xu and colleagues](https://arxiv.org/abs/2605.14021) compared repeat appearances of the same trending searches on different days. When both showed an AI Overview, only 0.24% cited exactly the same set of pages, and only 0.06% had exactly the same text.

A vendor study by [Tannenbaum](https://arxiv.org/abs/2609.22655) found the same scale of movement in a small benchmark of 15 prompts about AI visibility software. Between 5 and 6 June 2026, 67.0% of cited web addresses turned over for the same prompt on the same engine.

## Is the churn caused by websites changing?

Mostly not: asking the same question minutes apart produces nearly as much change. Schulte’s team re-asked each question up to 10 times within 24 hours. Source overlap between those near-simultaneous runs averaged 32% to 43%, the same range as the day-to-day figures. The authors conclude that the engines’ own randomness accounts for most of the instability.

[Sielinski](https://arxiv.org/abs/2603.08924) tested the other explanation directly. Sielinski fingerprinted the content of cited pages each day and found most pages did not change while their citation shares swung up and down. The conclusion: the variability is structural, not driven by content.

So a page can drop out of an answer tomorrow without anyone touching it, and come back the day after.

## Where does the randomness come from?

From several steps: the searches the assistant writes, the pages it picks, and the answer it generates. Before answering, an assistant writes its own web searches. In [our hidden-searches study](https://underneath.agency/research/ai-hidden-searches-study), ChatGPT ran a mean of 3.7 searches per buyer question, Gemini 1.9 and Claude 0.76. None of the 509 searches repeated the user’s question word for word.

Each of those searches can return different pages, and the assistant then chooses which to use. ChatGPT does not always search at all: in the Swiss study, 57.8% of its runs had no citations, because it answered some questions from memory.

Fixing the settings does not remove the variation. A [survey by Martinez](https://arxiv.org/abs/2607.14035) reports a peer-reviewed audit in which repeated runs with the randomness setting at zero still changed 9% to 28% of decisions. Small wording changes matter too. [Grossman and colleagues](https://arxiv.org/abs/2604.27790) found that edits as minor as “what is” versus “what’s” cut the similarity of AI Overview sources by 28.99% compared with a plain rerun.

## Do some AI engines change more than others?

Yes, and each engine has a fairly steady level of churn of its own. Sielinski sampled Gemini, ChatGPT search and Perplexity on three consumer topics over nine days. Repeated runs of the same question shared about 0.30 of their cited websites on Gemini, 0.40 on ChatGPT search and 0.50 on Perplexity, however many sources an answer cited.

| Engine | Change in cited web addresses, 5 to 6 June 2026 |
|---|---|
| ChatGPT | 82.6% |
| Microsoft Copilot | 76.6% |
| Perplexity | 45.5% |

Source: Tannenbaum, 15 prompts, one day-to-day comparison; the author founded an AI visibility software company and notes the change may partly reflect the collection system.

Google’s AI Overviews are less stable than Google’s ordinary results. Over two months, a peer-reviewed audit summarized by Martinez found only 18% of pages overlapped in AI Overviews, against 45% in organic Google. In Grossman’s tests from two US cities, two runs of the same AI Overview scored 0.69 on a 0-to-1 similarity scale, against 0.86 for the ordinary results.

## Does a changing source list mean your brand changes too?

Not one for one: brand mentions are steadier than sources, but still move. In the Swiss study, consecutive days shared 45% to 59% of brands, against 34% to 42% of sources.

The link between sources and brands is loose. In [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), 166 pairs of Perplexity runs cited identical web addresses, yet the brand list changed in 91.6% of them. The same evidence can produce a different shortlist. Engines also differ in how steady their brand lists are; see [which AI engine is most consistent](https://underneath.agency/resources/most-consistent-ai-engine-for-brands).

Switching engines changes far more than waiting a day. In Tannenbaum’s benchmark, 84.9% of [pairs of engines answering the same prompt](https://underneath.agency/resources/do-ai-search-engines-cite-the-same-sources) shared no cited web address at all. 96.4% of addresses appeared on only one engine. A same-engine source list from the next day was about 42 times more alike than another engine’s list from the same day.

## What should you do about it?

Read every AI citation check as one draw from a moving system, not as a fixed position.

1. Do not react to a single disappearance. Re-check the same question several times, on several days, before deciding anything changed.
2. Track each engine separately. Their stability differs, and their sources barely overlap. If you need one number, see [how to combine engines into one score](https://underneath.agency/resources/combine-ai-engines-visibility-score).
3. Watch your brand mentions as well as your cited pages. Mentions move less, and they are what buyers read.
4. Keep the wording of tracked questions fixed, since small edits shift sources on their own.
5. Aim to be one of the sources an engine returns to often. In these studies, a small core of sites recurs while the rest rotate.

If you want help setting up steady tracking, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The studies measure how much citations change, but none can see inside the engines to say exactly why.

- Most data come from short windows, from nine days to two months, so seasonal shifts and model updates are not well covered.
- The Swiss study used German questions from Swiss servers, and Sielinski studied three consumer topics; other markets may differ.
- Several key papers have commercial ties: Sielinski works for IQRush, Tannenbaum founded Aiso Boost, and Schulte is also affiliated with Aurora Intelligence. Tannenbaum’s day-to-day comparison may partly reflect collection changes.
- No study links citation churn to what buyers do, so we do not know how much a rotating source list costs a brand.

## Frequently asked questions

### Is it normal for ChatGPT to cite different websites for the same question?

Yes. In one 45-day study of four engines, about 65% of cited sources changed from one day to the next, and one vendor benchmark measured 82.6% turnover for ChatGPT between two days.

### Why did my page stop being cited by Perplexity?

Possibly for no reason related to your page. Perplexity was the steadiest engine in one study, but repeated runs still shared only about 0.50 of their cited websites.

### Are Google AI Overviews more stable than ChatGPT?

They are less stable than Google’s ordinary results. Over two months, only 18% of AI Overview pages overlapped, against 45% for organic results, in a peer-reviewed audit.

### Will checking again a few minutes later give different sources?

Often, yes. When the same questions were re-asked within 24 hours, runs shared only 32% to 43% of their sources, about as little as on different days.

## Sources

- Julius Schulte, Malte Bleeker and Philipp Kaufmann (2026), [Don’t Measure Once: Measuring Visibility in AI Search (GEO)](https://arxiv.org/abs/2604.07585), arXiv:2604.07585.
- Ronald Sielinski (2026), [Quantifying Uncertainty in AI Visibility: A Statistical Framework for Generative Search Measurement](https://arxiv.org/abs/2603.08924), arXiv:2603.08924.
- Benjamin Tannenbaum (2026), [Scoring With the Engine: Retrieval Exposure, Cross-Engine Divergence, and the Limits of Engine-Agnostic GEO Scores](https://arxiv.org/abs/2609.22655), arXiv:2609.22655.
- Olivier Martinez (2026), [Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026)](https://arxiv.org/abs/2607.14035), arXiv:2607.14035.
- Riley Grossman and colleagues (2026), [How Generative AI Disrupts Search: An Empirical Study of Google Search, Gemini, and AI Overviews](https://arxiv.org/abs/2604.27790), arXiv:2604.27790.
- Haofei Xu, Umar Iqbal and Jacob M. Montgomery (2026), [Measuring Google AI Overviews: Activation, Source Quality, Claim Fidelity, and Publisher Impact](https://arxiv.org/abs/2605.14021), arXiv:2605.14021.
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)

---

This is the Markdown twin of https://underneath.agency/resources/why-ai-search-citations-change. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Why Your Marketing Analytics Miss Brand Exposure in AI Answers"
description: "Analytics log clicks and paid impressions. Most AI answers are read without a click, so brand mentions in them leave little or no trace in your data."
canonical: "https://underneath.agency/resources/why-analytics-miss-ai-visibility"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Why can’t our marketing analytics see our visibility in AI answers?

Your analytics record what reaches your website or your ad platform, and most exposure in AI answers never reaches either. A brand can be named in a third of the AI answers in its category and still barely register in referral reports. The research now treats AI visibility as a separate measurement that has to be built, not pulled from existing dashboards.

## The short version

1. AI answers are mostly read without a click: in a panel of 900 US adults, only about 1% of visits to Google results pages with an AI Overview led to a click on a cited source.
2. In a field test with 1,100 US users, forcing Google’s AI Mode cut clicks to outside websites by 18.8 percentage points.
3. One software brand appeared in 33.8% of answers from one OpenAI model to 56 product questions, exposure no web analytics tool would have logged.
4. On one website, ChatGPT referrals grew 5.7 times in four months, but pages nobody changed grew 3.5 times, so most of the rise was the platform’s own growth.
5. A survey of 45 studies rates the evidence that AI citation scores predict clicks, conversions or revenue as “very low”.

## What do our current marketing reports actually record?

They record ad impressions, clicks and visits; none of these captures an unpaid brand mention in an AI answer.

[Kato and colleagues](https://arxiv.org/abs/2609.11915), in a 2026 paper on extending marketing mix models to AI, put the gap plainly: “conventional impression logs do not record nonsponsored occurrences of a firm’s name in generated answers.” Paid placements have the reverse problem. The platform records that a sponsored placement was shown, but not whether anyone noticed it.

Marketing mix models, the statistical models many marketing teams use to split budget across channels, need an exposure count for each market and period. TV, search ads and social all supply one. AI answers do not. [Schulte and colleagues](https://arxiv.org/abs/2604.07585) add that AI providers offer no monitoring tool equivalent to Google Search Console, so even the questions people ask, and how often, are hidden from brands.

## Why don’t AI answers show up as traffic?

Because most people read the answer and stop there, so your brand is seen but no visit is logged.

The best evidence comes from Google’s AI features. In [a Pew Research Center browsing panel](https://arxiv.org/abs/2608.04831) of 900 US adults tracked in March 2025, about 18% of all Google searches showed an AI Overview, the AI summary at the top of Google’s results. Only about 1% of visits to those pages led to [a click on a cited source](https://underneath.agency/resources/do-people-check-ai-sources). People clicked any search result on 15% of pages without an AI Overview, but on only 8% of pages with one, and ended their browsing session more often (26% against 16%). The study is observational, so it shows an association, not a cause.

A randomized test points the same way. In [a preregistered field experiment](https://arxiv.org/abs/2608.18352) with 1,100 US participants in March 2026, forcing Google’s AI Mode reduced clicks to outside sites by 18.8 percentage points. Hiding AI Overviews raised them by 8.8 points. We cover what that means for traffic in [our article on AI Mode](https://underneath.agency/resources/ai-mode-default-traffic-loss).

Standalone assistants such as ChatGPT look similar. A [panel study by Scrunch AI](https://arxiv.org/abs/2607.04282), a vendor of AI visibility software, covered US and British users in early 2026. It found that 34.1% of sessions containing an AI assistant showed no visit to any outside website, against 19.5% for sessions built around search.

## Is AI referral traffic a good stand-in for AI visibility?

No: a referral is a different event, and its growth mostly reflects the AI platforms’ own growth.

Kato and colleagues make the first point directly: referral sessions measure something else, “because users may read an answer without following a link.” A brand can be named often and receive almost no visits.

The second point comes from [Watanabe and Nakayashiki](https://arxiv.org/abs/2606.04362), who studied their own website’s server logs and Google Analytics data. From January to May 2026, total ChatGPT referrals grew 5.7 times. But pages they never touched grew 3.5 times over the same window. The pages they did change grew 6.1 times, and their best estimate of their own effect was a lift of about 1.82 times. Even that estimate did not pass their strictest check, so they call it suggestive. Our guide on [crediting ChatGPT referral growth to GEO](https://underneath.agency/resources/chatgpt-referral-growth-and-geo) covers this test in more depth.

Their data also shows how fragile the numbers are. The share of ChatGPT visits counted as engaged rose from 0.486 to 0.896, and most of the jump came when bot filtering changed in mid-March. Monthly totals rebuilt from daily data differed from direct monthly queries by up to 3.3%. A dashboard can move for reasons that have nothing to do with your brand.

## What would it take to measure AI exposure properly?

You would have to build the exposure count yourself, from repeated answers, question volumes, engine shares and attention.

Kato and colleagues set out the recipe. How often a brand appears in sampled answers must be combined with how many relevant questions are asked, the share each AI system handles, and the chance a reader notices the name. Their own sample shows why a single figure misleads. Across 2,240 answers to 56 product questions, the target brand appeared in 33.8% of answers from GPT-5.6 Luna and 27.8% from GPT-4o. Yet the order flipped by language: GPT-4o led by 6.79 points in English, while GPT-5.6 Luna led by 18.75 points in Japanese.

Repeated sampling matters too. In [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), a single ChatGPT answer showed 57.8% of the brands that five answers to the same question named between them. Most firms have none of the other inputs, such as question volumes by engine or measured notice rates. The authors flag that their method needs exactly the data most firms lack.

## Can we link AI visibility to revenue yet?

Not reliably: published research has not shown that AI citation scores predict clicks, conversions or revenue.

[Martinez’s critical survey](https://arxiv.org/abs/2607.14035) of 45 studies from 2023 to 2026 grades that claim “very low” confidence. It rests on one suggestive quasi-experiment, the single-site study above, plus a few industry claims. One industry study reported a 20% traffic lift against a control group, but without the group sizes or uncertainty needed to rely on it.

Kato and colleagues test their method on simulated sales data, not a real company’s results. That makes it a design for future measurement, not evidence of a return.

## What should you do about it?

Treat AI visibility as its own channel with its own data, instead of waiting for analytics to reveal it.

1. Stop reading flat AI referral traffic as proof of absence. The studies above show exposure without clicks is the normal case.
2. Measure mentions directly. Ask a fixed set of buyer questions repeatedly, on each engine your buyers use, and report how often your brand appears.
3. When you change content, leave a comparable set of pages untouched. Watanabe and Nakayashiki’s untouched pages are what revealed that most of their growth was the platform’s.
4. Brief your analytics or mix-modeling team on the missing inputs: question volumes, engine shares and notice rates. Do not let AI exposure enter a model as a guess.
5. Ask any vendor what its visibility number counts, how many runs it uses, and whether it has ever been checked against sales.

If you want help setting up that kind of measurement, see [our generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

It does not yet show how many people notice a brand in an AI answer, or what that is worth.

- No study in this set measures notice rates for brand names in AI answers; Kato and colleagues need them but assume them in simulation.
- Their answer sample covers one brand, two OpenAI models and one day of collection in September 2026.
- The referral evidence comes from one website, one engine and a bundle of changes made together.
- The click evidence is strongest for Google’s AI features; for standalone assistants it comes from one vendor’s opt-in panel.
- Whether AI visibility drives sales has not been tested in a controlled way.

## Frequently asked questions

### Does Google Analytics track visibility in ChatGPT?

No, it only records visits that arrive through a link. In one site’s data, ChatGPT referrals grew 5.7 times in four months while pages that were never changed grew 3.5 times, so even those visits mix your efforts with platform growth.

### Why is our AI referral traffic low if AI assistants mention us?

Because most readers do not click. In a 900-person US panel, only about 1% of visits to Google pages with an AI Overview led to a click on a cited source.

### Can marketing mix modeling measure AI search?

Only once you build the missing input. Researchers propose combining how often your brand appears with question volumes, engine shares and notice rates, data most firms do not yet collect.

### How do we know if our AI visibility work is paying off?

Compare changed pages or questions against a similar set you left alone. Without that control, a rise in traffic may simply reflect the AI platforms growing.

## Sources

- Kato, Honma and Kato (2026), [Generative Marketing Mix Modeling: A Causal Inference Framework Linking GEO and GEM to Business Impact](https://arxiv.org/abs/2609.11915), arXiv:2609.11915.
- Schulte, Bleeker and Kaufmann (2026), [Don’t Measure Once: Measuring Visibility in AI Search (GEO)](https://arxiv.org/abs/2604.07585), arXiv:2604.07585.
- Chapekis, Lieb, Shah and Smith, Pew Research Center (2026), [Investigating Click Behaviors On Google Search Result Pages That Produce an AI Overview](https://arxiv.org/abs/2608.04831), arXiv:2608.04831.
- Wang, Gleason, Bart, Wilson and Metaxa (2026), [AI in Search Reduces Publisher Referrals Without Improving User Experience: Experimental Evidence](https://arxiv.org/abs/2608.18352), arXiv:2608.18352.
- Iannelli and Ai, Scrunch AI (2026), [The New Shape of Search: How Conversational AI Recomposes Information Seeking](https://arxiv.org/abs/2607.04282), arXiv:2607.04282.
- Watanabe and Nakayashiki (2026), [Disentangling Answer Engine Optimization from Platform Growth: A Log-Based Natural Experiment on ChatGPT Referral Traffic](https://arxiv.org/abs/2606.04362), arXiv:2606.04362.
- Martinez (2026), [Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026)](https://arxiv.org/abs/2607.14035), arXiv:2607.14035.
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)

---

This is the Markdown twin of https://underneath.agency/resources/why-analytics-miss-ai-visibility. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Why doesn’t ChatGPT mention our newly launched product?"
description: "ChatGPT often answers from training data that predates your launch, and even assistants that search rarely surface new products. Here is what helps."
canonical: "https://underneath.agency/resources/why-chatgpt-misses-new-products"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Why doesn’t ChatGPT mention our newly launched product?

Usually because the version answering learned about the world before your product existed, and has not found pages about it since. AI assistants that search the web can close part of that gap, but in tests even they rarely surfaced new products for open-ended questions. A launch plan has to work around both problems instead of counting on AI discovery.

## The short version

1. In a test of 112 Product Hunt startups, ChatGPT without web search surfaced them in 3.32% of discovery answers and Perplexity, which searches, in 8.29% ([Sharma](https://arxiv.org/abs/2601.00912), December 2025).
2. Across all its answers, ChatGPT surfaced only 6 of the 112 products even once, against 31 for Perplexity.
3. The author calls this a “recency wall”: products launched in January 2025 could not be in the tested ChatGPT version’s training data.
4. In [our hidden searches study](https://underneath.agency/research/ai-hidden-searches-study), consumer ChatGPT ran 3.7 web searches per buyer question, and answers that ran no search cited nothing.
5. When assistants did search, 17.4% to 22.6% of their dated citations were under 90 days old, against 6.9% of Google’s top 10 ([our freshness study](https://underneath.agency/research/ai-source-freshness-study)).

## Why can’t ChatGPT recommend a product it has never seen?

An AI assistant answering from memory only knows what was in its training data. [Sharma](https://arxiv.org/abs/2601.00912), a single-author study from IIT Patna, tested 112 startups from the 2025 Product Hunt leaderboard on the developer version of ChatGPT (gpt-4o-mini) with no web search. He describes a “recency wall”: products launched in January 2025 cannot appear in that model’s training data.

He calls this “a hard constraint, not something optimization can overcome”. No change to your website can teach a model about a product it was never trained on. Only a newer model, or a live web search, can.

A second effect compounds the first. Even before the cutoff, established products have far more written about them, so there is more for a model to learn from. Sharma calls this “authority concentration”, and it means a new entrant starts behind even once it is old enough to be included.

## Does ChatGPT search the web for new products?

Often, but not always, and that decides whether a new product can appear at all. In [our hidden searches study](https://underneath.agency/research/ai-hidden-searches-study), the consumer ChatGPT app ran a mean of 3.7 web searches per buyer question. None of the 42 answers without a search, across the assistants we tested, cited anything.

The same assistant can answer either way. In [Kumar’s](https://arxiv.org/abs/2606.20065) tracking setup at the AI visibility company Ranqo, ChatGPT and Claude ran web search only on a weekly cycle per brand and answered from memory in between. Whether your launch can show up therefore depends on whether a search happens, and on what it finds.

## How much do search-connected assistants help a new product?

They help, but less than many founders hope. In Sharma’s test, Perplexity, which searches the web on every question, surfaced the startups in 8.29% of discovery answers. That is 2.5 times ChatGPT’s rate, and still low.

| Measure | ChatGPT (no search) | Perplexity (search) |
|---|---|---|
| Discovery answers naming the product | 3.32% | 8.29% |
| Products surfaced at least once | 6 of 112 (5.4%) | 31 of 112 (27.7%) |

Sharma puts the ChatGPT figure bluntly: a 3% discovery rate means 97 out of 100 relevant questions will not mention your product. For ChatGPT, nothing he measured predicted which products got through. For Perplexity, [links from other websites](https://underneath.agency/resources/do-backlinks-matter-for-perplexity-visibility), a strong Product Hunt ranking and genuine Reddit discussion all went with more discovery.

## Does an optimized launch page get you found?

Not on its own, in the one study that tested it. Sharma scored each startup’s website for the features usually recommended for AI visibility, such as statistics, citations and structured data. That score showed no link with discovery on either assistant.

His reading is that on-page polish works as a multiplier, not a starting point: “You can’t multiply zero.” A product first has to appear on pages that the assistant’s searches find.

Those pages are often not the ones ranking on Google. In [our study of AI citations and Google rankings](https://underneath.agency/research/ai-citations-google-rankings-study), only 8.3% of the pages ChatGPT cited ranked in Google’s top 10 for the question asked.

## Does fresh content help a new product get cited?

In assistants that search, recent pages have a measurable edge. In [our freshness study](https://underneath.agency/research/ai-source-freshness-study), the pages ChatGPT, Gemini, Perplexity and Claude cited were first published about half as long ago as Google’s top 10 for the same questions, a ratio of 0.50. Between 17.4% and 22.6% of each assistant’s dated citations were under 90 days old, against 6.9% for Google.

A lab test points the same way. Researchers at the software company [Sprinklr](https://arxiv.org/abs/2605.25517) ran 252,000 head-to-head trials on six AI models, each comparing two otherwise identical pages. A recent date was one of four factors that decided which page was cited first on all six models. The test compared content dated 2026 with the same content dated 2019.

Do not mistake a new date for new content. In our data, 66.4% of the recently dated pages in Google’s top 10 were old pages with a new modified date. The assistants’ edge came mostly from pages that were genuinely new.

## How long until AI assistants know about a new product?

No study has measured that for brand-new products yet. Kumar reports that established brands were recognized right away when a question named them, from 94% to 100% across engines, while unbranded questions surfaced them far less often, at 12% to 52%. He cautions that a brand with no prior web presence can take longer.

For [assistants that answer from memory](https://underneath.agency/resources/do-ai-assistants-answer-from-training-data), Sharma’s advice is stark: “there’s not much to do except wait” for newer training data. That timeline is set by the AI companies, not by you.

## What should you do about it?

Plan a launch that does not depend on early AI discovery, while building what search-connected assistants can find.

1. **Assume early AI visibility will be low.** Budget for channels you control, and treat any AI mentions as a bonus at first.
2. **Get written about on other sites.** Reviews, launch platforms, comparisons and genuine community discussion went with discovery on the assistant that searched.
3. **Publish clear, dated, factual pages.** Put prices, specifications and launch dates on pages that comparison writers and assistants can quote. Device makers face this with every launch, since assistants must get launch and compatibility facts right, as [how consumer tech brands reach AI shortlists](https://underneath.agency/resources/consumer-tech-brands-ai-search) shows.
4. **Test the assistants your buyers use.** Check whether they search for your category, and ask the same questions several times.
5. **Track monthly.** Visibility can change as new pages are found and new models are released.

If you want help planning launch visibility, see [our approach to generative engine optimization](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The evidence on new products is thin, and most of it comes from one study.

- **One small study carries most of the weight.** Sharma tested 112 startups, skewed toward developer, productivity and AI tools, on two developer versions of assistants in December 2025.
- **The measure is simple.** A success counted whenever the product name appeared in the answer.
- **Nobody has tracked time to visibility.** How long a new product takes to appear in each assistant is still unmeasured.
- **Freshness findings are associations or lab results.** Neither our data nor the Sprinklr test shows what happens when a real launch page is published.

## Frequently asked questions

### Why doesn’t ChatGPT know about my new product?

If it answers from training data, your product may simply postdate it. In one study, ChatGPT without web search surfaced only 6 of 112 recent startups in any discovery answer.

### Does ChatGPT search the internet before answering?

Often. In our study, the consumer ChatGPT app ran 3.7 web searches per buyer question, but answers that ran no search cited no pages at all.

### Does Perplexity find new products better than ChatGPT?

In one test, yes: Perplexity surfaced new startups in 8.29% of discovery answers against 3.32% for ChatGPT without search, about 2.5 times as often.

### How long does it take for a new product to appear in AI answers?

No study has measured it yet. Assistants that search can find new pages quickly, while answers from training data wait for the next model update.

## Sources

- Sharma (2026), [The Discovery Gap: How Product Hunt Startups Vanish in LLM Organic Discovery Queries](https://arxiv.org/abs/2601.00912), arXiv:2601.00912.
- Kumar (2026), [Generative Engine Optimization at Scale: Measuring Brand Visibility Across AI Search Engines](https://arxiv.org/abs/2606.20065), arXiv:2606.20065.
- Vishwakarma, Kumar and Jamidar (2026), [What Gets Cited: Competitive GEO in AI Answer Engines](https://arxiv.org/abs/2605.25517), arXiv:2605.25517.
- Underneath (2026), [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- Underneath (2026), [How fresh are the pages AI engines cite?](https://underneath.agency/research/ai-source-freshness-study)
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude cite pages that rank?](https://underneath.agency/research/ai-citations-google-rankings-study)

---

This is the Markdown twin of https://underneath.agency/resources/why-chatgpt-misses-new-products. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Why well-known brands go missing from AI recommendations"
description: "AI assistants often know a famous brand but fail to call it up for a plain category question. Research shows why, and how to check your own brand."
canonical: "https://underneath.agency/resources/why-well-known-brands-miss-ai-recommendations"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Why is our well-known brand missing from AI product recommendations?

Because being famous does not guarantee a place on the short list an AI assistant builds for each question. Researchers have caught household names such as L.L.Bean and Craftsman at zero in plain category questions, even though the assistants clearly knew them. In most cases the brand is known but not called up by the way buyers ask, so the fix starts with measuring the right questions.

## The short version

1. Six AI models answering from memory never recommended L.L.Bean, Craftsman, Braun or Philips for a plain category question across 1,200 answer lists ([Malthouse and colleagues](https://arxiv.org/abs/2609.16304), Northwestern University, 2026).
2. When the question used the brands’ own positioning language, L.L.Bean appeared in 88.5% of lists and Craftsman in 81.3%, so the models did know them.
3. Realistic, detailed buyer questions helped far less: Craftsman reached 35.4% of lists and L.L.Bean 5.4%.
4. In a study of 112 startups, ChatGPT recognized products by name 99.4% of the time but surfaced them for discovery questions 3.32% of the time ([Sharma](https://arxiv.org/abs/2601.00912), 2026).
5. In [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), one ChatGPT answer showed only 57.8% of the brands its five answers to the same question named.

## Can a market leader really be absent from AI answers?

Yes, and researchers have recorded well-known brands receiving no recommendations at all. [Malthouse and colleagues](https://arxiv.org/abs/2609.16304) asked six AI models, including GPT-5.5, Gemini 3.1 Pro Preview and Claude Opus 4.7, a plain request: “I am looking for a [category].” Each model answered 40 times per category, giving 1,200 lists of up to five brands across cordless drills, hiking jackets, coffee makers, cat food and cruises.

Many large, established brands never appeared. The authors list L.L.Bean, Eddie Bauer and REI Co-op in hiking jackets, Craftsman and Black+Decker in drills, Braun and Philips in coffee makers, Freshpet in cat food, and Virgin Voyages and Regent Seven Seas in cruises. Their conclusion for brand managers is blunt: check whether you are recommended at all, not only where you rank.

One caveat matters for reading this. The models were called through their developer interfaces with web search switched off, so they answered from what they learned in training. The consumer apps your buyers use may search the web and give different results.

## Why would an AI assistant know a brand but not recommend it?

Because recognizing a name and calling it up for a category are two different tasks. In [Sharma’s study](https://arxiv.org/abs/2601.00912) of 112 startups from the 2025 Product Hunt leaderboard, ChatGPT recognized products when asked about them by name 99.4% of the time. When asked discovery questions such as “What are the best AI tools launched this year?”, it surfaced them in 3.32% of answers. The study counted a hit whenever the product name appeared in the answer, so the by-name figure may flatter true recognition. Younger companies face the same gap, covered in [winning customers as an AI startup](https://underneath.agency/resources/ai-startups-customers-from-ai-search).

Malthouse’s team tested the same gap for big brands. They wrote deliberately unrealistic “probes” that borrowed language straight from each brand’s own marketing. With those cues, Craftsman appeared in 81.3% of lists and L.L.Bean in 88.5%.

So the models knew both brands and what they stood for. The authors conclude that low recommendation rates “cannot be attributed simply to an absence of brand knowledge”. The brand was in the model; the ordinary question did not bring it out.

## Does the way buyers phrase a question change who appears?

Yes, and often by more than any other factor you can see. Malthouse’s team wrote 20 realistic buyer questions per category, such as a teacher asking for a tough, affordable waterproof jacket for school field trips. Across 240 lists per brand, Craftsman rose from zero to 35.4% of lists, and L.L.Bean only to 5.4%.

[Our prompt phrasing study](https://underneath.agency/research/ai-prompt-phrasing-study) found the same sensitivity in live assistants. Asking the identical question again kept the same first brand 68.0% of the time. Adding “on a tight budget” kept it only 15.3% of the time.

A vendor data set adds a warning about specific questions. [Kumar](https://arxiv.org/abs/2606.20065), a co-founder of the AI visibility company Ranqo, found broad “best X” questions surfaced a tracked brand about 23% of the time, while specific problem and use-case questions did so about 11% of the time. A brand can be findable for the broad ask and nearly invisible for the specific one. Smaller players face a related question, covered in [whether smaller websites get cited](https://underneath.agency/resources/can-small-websites-get-cited-by-ai).

## Is it about your price tier or positioning?

Sometimes, in some categories, but no single rule explains it yet. In drills and hiking jackets, the models in the Northwestern study favored premium brands such as DeWalt, Milwaukee, Patagonia and Arc’teryx, while giving little room to mass-market names such as Black+Decker, Craftsman, Eddie Bauer and L.L.Bean.

Conventional popularity explained the results only in places. In cruises, Carnival and Royal Caribbean, the two brands consumers think of most readily (scored 177 and 162 on Kantar BrandZ salience), were among the most prominent picks. In the other categories, that measure and the models’ picks showed little systematic link, and the premium pattern held in only two of the five categories.

The authors’ practical reading is that what a brand is associated with matters, not only how visible it is. They advise stating a few distinct, relevant points of difference consistently across your own channels and the reviews, news and conversations you can influence. For device makers, whose buyers weigh compatibility and ecosystem, our guide to [consumer tech brands on AI shortlists](https://underneath.agency/resources/consumer-tech-brands-ai-search) shows how this applies.

## Could you appear in some answers but not the one you checked?

Yes, because AI answers change from run to run and from assistant to assistant. In [our consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), we asked ChatGPT, Gemini and Perplexity the same 20 buyer questions five times each. Of the brands ChatGPT named for a question, 25.2% appeared in all five runs and 36.6% in only one.

Different assistants also disagree. In [our four-assistant study](https://underneath.agency/research/ai-assistants-brand-agreement-study) of 80 buyer questions, 66.3% of recommended options came from only one assistant, and all four agreed on the first pick for 10.0% of questions.

The flip side matters too. In Kumar’s vendor data, 63.2% of brand, question and assistant combinations were never mentioned in any run. If your brand is missing across many runs and assistants, that is probably a stable gap, not bad luck.

## What should you do about it?

Measure whether you appear at all, then find out why the ordinary question leaves you out.

1. **Define your competitors first.** List the brands you compete with before you look at AI answers, so you can see who is missing, including you.
2. **Ask plain category questions repeatedly.** Use several runs on each assistant your buyers use. One answer is a sample, not the result.
3. **Add the needs your brand is built for.** Write questions as your target buyers would, with budget, use and experience. This shows whether your positioning connects.
4. **Test whether the assistant knows you.** Ask with your own positioning language. If you appear then but not otherwise, the gap is activation, not awareness.
5. **Make your positioning consistent everywhere.** Repeat the same few points of difference on your site and in the reviews, coverage and discussions that describe you.

If you want help running this kind of audit, see [our approach to generative engine optimization](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The evidence explains how brands go missing better than it explains how to bring them back.

- **Search was switched off.** The Northwestern results come from models answering from training alone. How far web search changes the picture for big brands is untested in that study.
- **The needs-based tests are small.** They cover two brands with hand-written questions, and the authors say the approach needs validation across many brands.
- **The price-tier pattern is unexplained.** It held in two of five categories, and the authors do not yet know why.
- **No fix has been proven.** Whether supplying clearer positioning information raises recommendations is listed by the authors as future research.

## Frequently asked questions

### Why doesn’t ChatGPT recommend my brand even though it knows it?

Recognition and recommendation are separate. In a study of 112 startups, ChatGPT recognized products by name 99.4% of the time but surfaced them for discovery questions 3.32% of the time.

### Does being a market leader guarantee AI visibility?

No. A Northwestern study found brands such as L.L.Bean, Craftsman, Braun and Philips received no recommendations across six AI models for plain category questions.

### How do I check whether AI assistants recommend my brand?

Ask unbranded buyer questions several times on each assistant and count how often you appear. In our tests, a single ChatGPT answer showed only 57.8% of the brands that five answers named.

### Will more detailed buyer questions bring my brand back?

Sometimes, partly. Detailed questions lifted Craftsman to 35.4% of lists but L.L.Bean only to 5.4%, so detail helps most when your positioning clearly fits the need.

## Sources

- Malthouse, Lee, Yang, Pal and Feng (2026), [Evaluating Brand Retrieval and Ranking in Large Language Model Recommendations](https://arxiv.org/abs/2609.16304), arXiv:2609.16304.
- Sharma (2026), [The Discovery Gap: How Product Hunt Startups Vanish in LLM Organic Discovery Queries](https://arxiv.org/abs/2601.00912), arXiv:2601.00912.
- Kumar (2026), [Generative Engine Optimization at Scale: Measuring Brand Visibility Across AI Search Engines](https://arxiv.org/abs/2606.20065), arXiv:2606.20065.
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)
- Underneath (2026), [Does rewording a question change AI brand recommendations?](https://underneath.agency/research/ai-prompt-phrasing-study)
- Underneath (2026), [Do ChatGPT, Gemini, Perplexity and Claude agree on brands?](https://underneath.agency/research/ai-assistants-brand-agreement-study)

---

This is the Markdown twin of https://underneath.agency/resources/why-well-known-brands-miss-ai-recommendations. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Why does Wikipedia matter so much for AI search visibility?"
description: "Wikipedia is the most cited site in many AI search studies and shapes answers beyond its citations. For buyer questions its share is small."
canonical: "https://underneath.agency/resources/why-wikipedia-matters-for-ai-search"
published: 2026-10-11
updated: 2026-10-11
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Why does Wikipedia matter so much for AI search visibility?

Wikipedia matters because AI search engines cite it more than any other website on general questions, and they lean on its text more heavily than their citation lists suggest. For commercial and brand questions, its share of citations is much smaller. Having an article mostly reflects how well known a brand already is, so for most companies accuracy matters more than presence.

## The short version

1. In a 2026 study of 11,000 real search questions, Wikipedia was the most cited website for every system, cited in 49% of SearchGPT answers and 28% of AI Overviews.
2. The same study found Wikipedia’s content over-represented in the AI summaries themselves, by 5.4 percentage points in Google’s AI Overviews.
3. Across twelve European languages, Wikipedia was the most cited website in 11, in a 2026 study of 35,640 brand answers.
4. On US buyer searches it is a minor source: 0.9% of AI Overview citations in our study of 481 AI Overviews.
5. In our study of 80 buyer questions, 60.0% of brands named by all four assistants had an English Wikipedia article, but most of that link went away once brand prominence was accounted for.

## How often do AI search engines cite Wikipedia?

On general-knowledge questions, more than any other website. It is the default reference source for most AI engines.

[Huang and colleagues](https://arxiv.org/abs/2603.16138) at the University of Illinois sent 11,000 real search questions to SearchGPT, Google’s AI Overviews, Perplexity and regular Google. Wikipedia was the most cited website for every system. It appeared in 49% of SearchGPT answers, 28% of AI Overviews and 58.0% of Perplexity answers.

Other topics show the same pattern. In ChatGPT’s answers to 100 consumer health questions, Wikipedia was the single most cited organization with 10.7% of citations, ahead of Mayo Clinic, according to [Jacques and colleagues](https://arxiv.org/abs/2601.17109). In an [audit of four engines](https://arxiv.org/abs/2605.23684) on politics, health and the environment, Wikipedia took 15.2% of the citations that went to [the 25 most cited websites](https://underneath.agency/resources/ai-search-citation-concentration).

## Does Wikipedia shape the answer, not just the source list?

Yes: AI summaries draw on Wikipedia’s text more than its share of citations implies. Its influence is larger than citation counts suggest.

Huang and colleagues compared how much of each cited source made it into the summary. Wikipedia’s content was over-represented by 2.6 percentage points in SearchGPT and 5.4 points in AI Overviews. In their words, Wikipedia is “not only the most frequently cited domain across all systems, but its content is disproportionately over-represented in generated summaries.”

A separate analysis of 602 prompts by [Zhang, He and Yao](https://arxiv.org/abs/2604.25707), three independent researchers, points the same way. On their own measure of how much a cited page shapes an answer, encyclopedia pages scored 0.2144 on average, against 0.0726 for news pages.

## Is Wikipedia as important for commercial and brand questions?

No: on buyer and brand questions, Wikipedia is a small source. Other sites take most citations there.

In [our AI Overview study](https://underneath.agency/research/ai-overview-citations-study) of 481 US AI Overviews on commercial searches, Wikipedia made up 0.9% of citations and appeared in 6.9% of AI Overviews. YouTube appeared in 64.2%. In [our comparison of AI Mode and AI Overviews](https://underneath.agency/research/ai-mode-vs-ai-overviews-study), Wikipedia was 0.5% of AI Mode citations and 0.8% of AI Overview citations.

A [study by Ranqo](https://arxiv.org/abs/2606.20065), a vendor of AI visibility tools, tracked more than 100 brands. Wikipedia took 2.6% of citations, behind YouTube at 4.2%, editorial media at 3.8% and forums at 3.3%. On trending news-style searches the picture shifts again: [Xu and colleagues](https://arxiv.org/abs/2605.14021) found en.wikipedia.org was the second most cited site in AI Overviews, with 4.39% of citations.

## Does a Wikipedia article make AI assistants recommend your brand?

Not by itself, in our data. Brands with articles get named more, but mainly because they are better known.

In [our brand entity study](https://underneath.agency/research/brand-entity-ai-recommendations-study), 80 US buyer questions went to ChatGPT, Gemini, Perplexity and Claude. Of the options all four assistants named for a question, 60.0% had an English Wikipedia article, against 20.6% of options only one assistant named. Once we accounted for website traffic and how often Wikipedia mentions the brand, an article added little or nothing.

Where an article helped, it was upstream. Brands with an article were named more often in the pages assistants cited, 74.3% against 59.5%. Once named there, they were recommended at about the same rate. And 21 of the 110 options all four assistants named had no Wikipedia article for themselves or their parent brand.

## Does Wikipedia’s role change by language?

Mostly not: it leads in almost every language studied, though local outlets compete in smaller ones. Language does change which brands get named.

[Żatuchin](https://arxiv.org/abs/2606.23165), who is also affiliated with an AI brand-monitoring company, asked three AI engines about 66 European brands in twelve languages. Wikipedia was the most cited website in 11 of the 12. The exception was Lithuanian, where the business daily vz.lt edged ahead with 4.38% of citations. The author reads this as local coverage mattering at the margin in smaller languages. Our guide on [local-language sources in AI answers](https://underneath.agency/resources/do-ai-engines-cite-local-language-sources) covers more of these studies.

## Is AI search changing Wikipedia’s own traffic?

Yes: AI answers use Wikipedia’s content while sending it fewer visitors. That makes the encyclopedia more of a background source.

[Khosravi and Yoganarasimhan](https://arxiv.org/abs/2602.18455) at the University of Washington compared English Wikipedia with German and French versions of the same articles. Default AI Overviews reduced English search traffic by 5.45% and 4.82% in the two comparisons. In a field experiment with 1,100 US participants, [Wang and colleagues](https://arxiv.org/abs/2608.18352) found an AI Mode-only Google cut the share of users clicking through to Wikipedia by 9.9 percentage points.

## What should you do about it?

Make sure Wikipedia is accurate about your company and category, but do not treat an article as a shortcut. In practice:

1. Check whether your company, products and category have Wikipedia coverage, and whether it is correct and current.
2. Never edit your own article or pay someone to; Wikipedia’s notability and conflict-of-interest rules are strict.
3. Earn independent coverage first; in our study it predicted recommendation far better than an article did.
4. Keep Wikidata and your website’s organization details consistent, so assistants match the right company.
5. For buyer questions, spend most effort on the sources that dominate your category, such as video, reviews and trade media. Our guide on [social media and AI search](https://underneath.agency/resources/does-social-media-help-ai-search-visibility) shows where video and social posts count.

If you want help reviewing your entity record, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The research shows that Wikipedia is heavily used, but not what a brand gains from changing it. The gaps:

- No study we reviewed tests whether adding or correcting a Wikipedia article changes AI answers about a brand.
- The strongest over-representation results use general-knowledge questions, not buyer or brand questions.
- How much AI assistants rely on Wikipedia from their training rather than live search is not measured here.
- Language studies cover Europe and a few engines in one period; other regions are untested.
- Two of the brand datasets come from companies that sell AI visibility tools.

## Frequently asked questions

### Does ChatGPT use Wikipedia as a source?

Often. In one 2026 study, Wikipedia appeared in 49% of SearchGPT answers to 11,000 general questions, more than any other website.

### Do I need a Wikipedia page to show up in AI search?

No. In our study, 21 of the 110 options all four assistants named had no article for themselves or their parent brand.

### How often do Google AI Overviews cite Wikipedia?

It depends on the searches. Wikipedia appeared in 28% of AI Overviews for general questions in one study, but in only 6.9% of AI Overviews on our US commercial searches.

### Can a company edit its own Wikipedia page to improve AI visibility?

It should not. Wikipedia’s conflict-of-interest rules discourage it, and no study we reviewed shows that edits change AI answers.

## Sources

- Huang and colleagues (2026), [Answer Bubbles: Information Exposure in AI-Mediated Search](https://arxiv.org/abs/2603.16138), arXiv:2603.16138.
- Jacques and colleagues (2026), [Authority Signals in AI Cited Health Sources: A Framework for Evaluating Source Credibility in ChatGPT Responses](https://arxiv.org/abs/2601.17109), arXiv:2601.17109.
- Allaham and Diakopoulos (2026), [Synthetic Sources?: Auditing Generative Search Engine Citations for Evidence of AI-Generated Sources](https://arxiv.org/abs/2605.23684), arXiv:2605.23684.
- Zhang, He and Yao (2026), [From Citation Selection to Citation Absorption: A Measurement Framework for Generative Engine Optimization Across AI Search Platforms](https://arxiv.org/abs/2604.25707), arXiv:2604.25707.
- Kumar (2026), [Generative Engine Optimization at Scale: Measuring Brand Visibility Across AI Search Engines](https://arxiv.org/abs/2606.20065), arXiv:2606.20065.
- Xu and colleagues (2026), [Measuring Google AI Overviews: Activation, Source Quality, Claim Fidelity, and Publisher Impact](https://arxiv.org/abs/2605.14021), arXiv:2605.14021.
- Żatuchin (2026), [The Language Blind Spot: How Query Language and Brand Recognition Tier Shape AI-Constructed Brand Reputation Across Twelve European Languages](https://arxiv.org/abs/2606.23165), arXiv:2606.23165.
- Khosravi and Yoganarasimhan (2026), [Impact of AI Search Summaries on Website Traffic: Evidence from Google AI Overviews and Wikipedia](https://arxiv.org/abs/2602.18455), arXiv:2602.18455.
- Wang and colleagues (2026), [AI in Search Reduces Publisher Referrals Without Improving User Experience: Experimental Evidence](https://arxiv.org/abs/2608.18352), arXiv:2608.18352.
- Underneath (2026), [Do AI Overviews cite the pages that rank?](https://underneath.agency/research/ai-overview-citations-study)
- Underneath (2026), [AI Mode vs AI Overviews: how different are the sources?](https://underneath.agency/research/ai-mode-vs-ai-overviews-study)
- Underneath (2026), [Do Wikipedia and schema make AI assistants recommend a brand?](https://underneath.agency/research/brand-entity-ai-recommendations-study)

---

This is the Markdown twin of https://underneath.agency/resources/why-wikipedia-matters-for-ai-search. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Do websites serve Markdown to AI agents? 2026 data | Underneath"
description: "3.2% of top websites serve Markdown when an AI agent asks for it, at a median 96.1% smaller than the HTML; 45.8% of homepages carry JSON-LD."
canonical: "https://underneath.agency/research/agent-readable-web-study"
published: 2026-09-26
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Research · AI crawlers

# Do websites serve Markdown to AI agents? 2026 data

AI agents read web pages as text. Some now ask for Markdown directly, by sending “Accept: text/markdown” with the request, and some hosting platforms can answer with a clean Markdown version instead of the full HTML page. We sent that request to the homepage of every live site in the Tranco top 10,000 on 26 September 2026, and at the same time recorded the other ways a homepage can describe itself to machines: a link to a Markdown version, and JSON-LD structured data. As far as we can find, this is the first published measurement of how many websites answer a Markdown request. Two days later we went back to every site that served Markdown and compared its Markdown with its HTML, item by item, to see whether an agent that asks for Markdown receives the same information as a reader of the page.

## The short version

1. 191 of 5,902 live top websites (3.2%) returned Markdown when asked for it; 71.2% of them are served through Cloudflare.
2. The Markdown was small: a median 13,081 bytes against 401,073 bytes for the same homepage as HTML, a median reduction of 96.1%.
3. Only 28 homepages (0.5% of HTML homepages) link to a Markdown version of themselves with a link tag, the other way to offer Markdown.
4. 45.8% of homepages carry JSON-LD structured data, 36.1% describe the organization behind the site, and 28.2% list sameAs links to profiles elsewhere; 54.2% carry no JSON-LD at all.
5. Of the sites that serve Markdown, 43.5% also publish a valid llms.txt: the same small group is adopting every agent-facing format at once.
6. The Markdown is a lighter page, not a copy. Only 2 of 186 Markdown homepages contained everything their HTML did. The median one kept 83.8% of the HTML’s sentences and 29.9% of its links, mostly by leaving out menus and footers, and 23 of 178 kept less than 20% of the sentences.

## What we measured

For each of the 5,902 domains in the Tranco top 10,000 whose homepage responded with a status below 400, we requested the homepage with the header an agent that prefers Markdown sends: “Accept: text/markdown, text/html;q=0.9, */*;q=0.8”. A site that supports Markdown content negotiation answers with Content-Type text/markdown; every other site answers with its usual HTML.

For the sites that answered with Markdown we then requested the same homepage asking for HTML, to compare the two. For the 5,710 homepages returned as HTML we read the head for a link to a Markdown alternate and parsed every JSON-LD block, including nested @graph objects.

## Findings

### Markdown on request is rare, and mostly a platform feature

| Signal | Sites | Share |
|---|---|---|
| Returned Markdown for Accept: text/markdown | 191 | 3.2% of live sites |
| Of those, also sent “Vary: Accept” so caches keep the versions apart | 183 | 95.8% of Markdown sites |
| Of those, Markdown starts with a metadata block (front matter) | 152 | 79.6% of Markdown sites |
| HTML homepage links to a Markdown version | 28 | 0.5% of HTML homepages |

Server headers show where the Markdown comes from: 136 of the 191 sites (71.2%) are served through Cloudflare, which launched “Markdown for Agents” as an optional feature in February 2026, followed by Vercel (14 sites), nginx (7), Framer (6) and Netlify (5). Early adopters include cloudflare.com, wordpress.org, sentry.io, zendesk.com, gitlab.com, jotform.com, intercom.com, mixpanel.com, typeform.com, siemens.com, elsevier.com, laravel.com, asana.com and miro.com.

### The Markdown version is a fraction of the size

For 185 sites we could fetch both versions. The median Markdown homepage was 13,081 bytes; the median HTML homepage for the same sites was 401,073 bytes. The median reduction was 96.1%, and for three sites in four it was at least 90.0%. For an agent that is billed by the token, that is a large difference in what each page costs to read.

### The Markdown is rarely the same page

On 28 September we requested each of the 191 Markdown-serving homepages again, twice asking for HTML and twice asking for Markdown, and compared the two versions: numbers, dates, prices, headings, links, name-like phrases and sentences. Only items that appeared in both requests of the same version were compared, so content that changes on every load (a date, a rotating banner) does not count as a difference. Two requests of the same version agreed on a median of 100% of items, so the differences below are between the versions, not between requests.

| HTML versus Markdown, same homepage | Result |
|---|---|
| Markdown contains everything the HTML does (and the reverse) | 2 of 186 (1.1%; 95% interval 0.3% to 3.8%) |
| Median share of the HTML’s sentences found in the Markdown | 83.8% (178 pages with sentences) |
| Median share of the HTML’s links found in the Markdown | 29.9% (182 pages with links) |
| Markdown keeps less than 20% of the HTML’s sentences | 23 of 178 |
| Same comparison with menus, headers and footers removed from the HTML (exploratory) | 7 of 186 equivalent; median 98.3% of sentences kept |

Most of the gap is page furniture. Platform conversions drop navigation menus, footers and social links, which is much of what makes the Markdown small. For a minority the Markdown is a different document written for agents: sentry.io and cloudflare.com, for example, answer a Markdown request with a short page about their product for AI tools rather than a conversion of the homepage. An agent reading those sites sees different text from a person visiting them.

Structured data travels unevenly. Of the 137 pages that carry JSON-LD, the Markdown includes it for 76. On 127 of the 137, the JSON-LD holds facts that the visible text does not, such as addresses, founding dates and profile names, so an agent that receives Markdown without it loses those facts.

The feature is not fixed. Three of the 191 sites (circle.com, webflow.com and webflow.io) did not return Markdown on at least one request two days later. Among the rest, the Markdown was byte-for-byte identical across two requests for 176 of 186 pages.

### Homepage structured data

| JSON-LD on the homepage | Share of 5,710 HTML homepages |
|---|---|
| Any JSON-LD | 45.8% |
| Organization (or a subtype such as NewsMediaOrganization or LocalBusiness) | 36.1% |
| WebSite | 31.0% |
| sameAs links to other profiles | 28.2% |
| A JSON-LD block that fails to parse | 0.6% |

The most common types were Organization (1,783 homepages), WebSite (1,770), SearchAction (1,226), ImageObject (1,222) and WebPage (921). Adoption barely changes with popularity: 43.8% of the top 1,000 homepages carry JSON-LD, against 45.3% of sites ranked 1,001 to 5,000 and 46.5% of sites ranked 5,001 to 10,000.

### A handful of sites opt out in the page itself

18 homepages (0.3%) carry a “noai” value in their robots meta tag, a non-standard signal some platforms use to ask AI systems not to use the content.

## How this compares with other studies

We found no earlier count of websites that answer a Markdown request. What has been published measures the other side, the agents: Checkly tested seven AI agents in February 2026 and found “only 3 out of 7 agents request markdown” (Claude Code, Cursor and OpenCode). Cloudflare’s launch post measured one page at 16,180 tokens as HTML against 3,150 as Markdown, an “80% reduction in token usage”; our median byte reduction across 185 homepages is larger, which may partly reflect that homepages carry more markup than article pages (we did not measure article pages). Vercel reported a 99.37% reduction in payload size for one page.

For structured data, the HTTP Archive Web Almanac 2025 (whole web, July 2025) found JSON-LD on 43% of desktop home pages and the Organization type on 26.74%. Our top-10,000 figures are similar for JSON-LD (45.8%) and higher for Organization (36.1%); the two samples differ (top sites against the whole web), so part of the gap may reflect larger sites investing more in describing who they are.

Sources: [Checkly](https://www.checklyhq.com/blog/state-of-ai-agent-content-negotation/); [Cloudflare, Markdown for Agents](https://blog.cloudflare.com/markdown-for-agents/); [Vercel](https://vercel.com/blog/making-agent-friendly-pages-with-content-negotiation); [HTTP Archive Web Almanac 2025](https://almanac.httparchive.org/en/2025/seo).

## What this means

What follows is our reading of the figures above; none of it was measured directly.

- **Markdown on request is the cheapest agent upgrade available.** For sites on a platform that supports it, it is a setting rather than a project, and it cuts what an agent must download by an order of magnitude. Few agents ask for Markdown today; in Checkly’s test, the ones that did were coding tools.
- **Entity data is still missing on most homepages.** More than half of top homepages give machines no structured statement of who is behind the site. For a brand that wants AI assistants to describe it correctly, an Organization record with sameAs links to its real profiles is the basic step.
- **Check what your Markdown says before an agent does.** A platform setting produces the Markdown, and it decides what to keep. For most sites that means losing menus, which is the point; for some it means losing most of the page, or the facts that sit only in the JSON-LD. Reading your own Markdown once takes minutes.
- **The same sites adopt every format.** The overlap between Markdown serving and llms.txt suggests these formats are being taken up by the same technical teams rather than spreading evenly. That is an opening for brands in categories where competitors have not started.

## Methodology

- **Sample:** Tranco list L5PZ4, top 10,000 domains; 5,902 with a live homepage (status below 400, more than 200 bytes) on 26 September 2026.
- **Markdown test:** one HTTPS GET of the homepage with “Accept: text/markdown, text/html;q=0.9, */*;q=0.8”; Markdown = response Content-Type text/markdown or text/x-markdown. Size comparison: a second request asking for HTML, same day.
- **Structured data:** every script of type application/ld+json parsed, types collected recursively. Homepages over 600 KB were re-fetched in full (790 of 807 succeeded); for the other 17 only the first 600 KB was searched.
- **Serving platform:** the Server response header, which a site can change or hide.
- **Representation check (version 1.1):** on 28 September 2026, four more requests to each of the 191 Markdown-serving homepages, two asking for HTML and two for Markdown, in a seeded random order. Two versions count as the same content only if every number, date, price, heading, link and name-like phrase of each appears in the other and at least 95% of the sentences do. The protocol was fixed before collection; 186 homepages were analyzed (three no longer served Markdown, one refused the HTML request, one timed out).
- **Update schedule:** quarterly.

## Limitations

- Only homepages were tested; a site may serve Markdown or structured data on other pages.
- Sites that blocked our research user agent are outside the sample.
- A site can serve Markdown for some user agents and not others; we used one request profile.
- This study measures what sites send. It does not show whether AI crawlers ask for Markdown, whether AI search engines retrieve it, or whether it changes what an answer says; those need sites whose format can be switched on and off, which a survey of other people’s sites cannot do.
- Sentence and name matching in the representation check are text-matching proxies, not a judgment of meaning, and no person or model reviewed them. The comparison without menus and footers was added after seeing the results.

## Data and downloads

- Every statistic on this page: [stats.json](https://underneath.agency/research-data/agent-readable-web-study/stats.json)
- The 191 sites that served Markdown, with sizes: [md_vs_html.json](https://underneath.agency/research-data/agent-readable-web-study/md_vs_html.json)
- The HTML and Markdown comparison for each of the 186 homepages: [s3_md_html_pairs.csv](https://underneath.agency/research-data/agent-readable-web-study/s3_md_html_pairs.csv)
- The comparison protocol, fixed before collection: [s3_representation_protocol.json](https://underneath.agency/research-data/agent-readable-web-study/s3_representation_protocol.json)
- Machine-readable methodology: [methodology.json](https://underneath.agency/research-data/agent-readable-web-study/methodology.json)

**Version 1.1** (28 September 2026) adds the representation check. No version 1.0 figure changed.

The data is free to reuse with attribution (CC BY 4.0).

To cite: Underneath. (2026). *Do websites serve Markdown to AI agents? 2026 data*. Underneath Research. https://underneath.agency/research/agent-readable-web-study

## Frequently asked questions

### What is Markdown content negotiation?

A website that supports it looks at the Accept header of each request. When an AI agent asks for text/markdown, the site returns a Markdown version of the page instead of HTML. Browsers keep getting the normal page.

### How many websites serve Markdown to AI agents?

In September 2026, 191 of the 5,902 live websites in the Tranco top 10,000 (3.2%) returned Markdown when asked for it.

### How much smaller is the Markdown version of a page?

Across 185 homepages we could compare, the Markdown version was a median 96.1% smaller than the HTML version of the same page.

### Is the Markdown version the same as the web page?

Usually not in full. Only 2 of 186 Markdown homepages we checked contained everything their HTML did. The typical one kept 83.8% of the sentences and 29.9% of the links, mainly because menus and footers are left out.

### What share of websites use Organization schema on the homepage?

36.1% of the top-site homepages we checked carry an Organization record (or a subtype) in JSON-LD, and 28.2% link to their other profiles with sameAs.

## Related research

- [How many websites have an llms.txt file?](https://underneath.agency/research/llms-txt-adoption-study)
- [Which AI crawlers do top websites block?](https://underneath.agency/research/ai-crawler-blocking-study)
- [What pages cited by AI Overviews have in common](https://underneath.agency/research/ai-overview-cited-pages-study)
- [How often do AI Overviews appear?](https://underneath.agency/research/ai-overviews-frequency-study)

## Related guides

- [Do AI agents recommend businesses whose websites they can read more often?](https://underneath.agency/resources/do-ai-agents-recommend-readable-websites)
- [How much web content is already optimized for AI search?](https://underneath.agency/resources/how-much-web-content-is-optimized-for-ai-search)
- [If AI agents can’t read my website, where does their answer about us come from?](https://underneath.agency/resources/when-ai-agents-cant-read-your-site)
- [How do developer tool companies win users when developers ask AI first?](https://underneath.agency/resources/developer-tools-ai-search)
- [Can our server logs show what content AI bots are looking for on our site?](https://underneath.agency/resources/ai-bot-server-logs-content-demand)

---

This is the Markdown twin of https://underneath.agency/research/agent-readable-web-study. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Do ChatGPT, Gemini, Perplexity and Claude agree on brands?"
description: "80 buyer questions, four AI assistants: a third of picks shared, the same top pick 10% of the time. They mostly choose differently, not read differently."
canonical: "https://underneath.agency/research/ai-assistants-brand-agreement-study"
published: 2026-09-26
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Research · AI assistants

# Do ChatGPT, Gemini, Perplexity and Claude agree on brands?

Ask ChatGPT, Gemini, Perplexity and Claude the same buyer question, such as “What is the best CRM for a small business?” or “Who is the best divorce lawyer in Houston?”, and you get four different shortlists. We put 80 such questions, ten in each of eight industries, to all four assistants on 26 September 2026 and compared the 320 answers.

Version 1.0 of this study measured how much the assistants agree. This version asks where the disagreement comes from. An AI recommendation passes through several stages: the assistant finds pages, cites some of them, selects options from what it read, orders them and explains them. We measured the stages we can observe from the outside, fetched the 2,209 pages the assistants cited, and asked for each option one assistant recommended and another did not: did the other assistant’s own sources name it?

The answer changes the usual explanation. The assistants do cite different websites, but that barely predicts which brands they recommend. On national questions, most of the options one assistant recommended and another left out were named in the pages the second assistant cited. The main difference is what each assistant picks from its evidence, not which evidence it finds. Local questions are the exception: there, most one-sided picks appear in neither assistant’s cited pages.

## The short version

1. **A third of the picks are shared.** For the same question, two assistants’ recommended options overlapped by 0.327 on average (Jaccard, 95% interval 0.281 to 0.375). 66.3% of the options recommended for a question came from one assistant only; all four recommended the same first pick for 10.0% of questions (3.8% to 17.5%).
2. **Place names halve agreement.** Overlap was 0.160 for the 24 questions naming a place and 0.390 for the 56 national ones, a gap of 0.231 (0.166 to 0.294). B2B software had the highest overlap (0.543), home and local services the lowest (0.227).
3. **Different sources barely explain different picks.** Two assistants shared little of what they cited (mean cited-domain overlap 0.079), and 43.8% of question pairs shared no cited domain at all. Yet each 0.1 of extra domain overlap went with only about 0.02 more recommendation overlap, an association that was not distinguishable from zero in most models. Source overlap accounted for 0.9% of the local gap.
4. **Most one-sided picks are selection, not retrieval.** When one assistant recommended an option and another did not, the second assistant’s own cited pages named it 42.1% of the time (34.1% to 50.5%); only the recommender’s pages named it 28.8% of the time; neither did 29.1% of the time. On national questions the selection share was 56.5%; on local questions 54.2% of one-sided picks appeared in neither assistant’s cited pages.
5. **Even with the same evidence, they pick differently.** When both assistants’ cited pages named an option that at least one of them recommended, both recommended it 47.9% of the time (42.5% to 53.4%). Across all options it was 27.8%.
6. **It is not only run-to-run noise.** Across five repeat runs of 20 questions, an assistant’s answers overlapped with its own other runs (0.421 to 0.682) about twice as much as with another assistant’s (0.257 to 0.288). Of the brands one assistant named in at least three of five runs, 50.0% to 56.6% were never named by the other assistant in any of its five.

## Research questions

| Question | Answered here? |
|---|---|
| RQ1. How far do the assistants agree on the options they recommend? | Yes |
| RQ2. Does source overlap predict recommendation overlap, net of question, place, industry and pair? | Yes, as an association |
| RQ3. Did the assistant’s cited pages name the options it recommended, and the options the other assistant recommended? | Yes, for cited pages we could read |
| RQ4. How much of the difference is run-to-run variation? | Partly: 20 questions, 3 assistants |
| RQ5. What reasons for disagreement are visible in the answers? | Yes, model-coded |
| RQ6. With the evidence held constant, how much disagreement remains, and which content changes move which stage? | No: needs controlled experiments (see What comes next) |

## Where the answers diverge

We treat an AI recommendation as a pipeline and report each stage separately rather than folding them into one “AI visibility” number.

| Stage | Observed here? | Result |
|---|---|---|
| Question | Yes | Local questions agree less than half as much as national ones |
| Retrieval | No | The search results each assistant saw are not visible |
| Cited evidence | Yes | Low overlap of cited domains (0.079) |
| Selection | Yes, via cited pages | 42.1% of one-sided picks were named in the other assistant’s cited pages |
| Recommendation | Yes | Overlap 0.327; same first pick 10.0% |
| Justification | No | Whether the sources support each reason given was not checked |
| Stability | Partly | Within an assistant 0.421 to 0.682; across assistants 0.257 to 0.288 |

“Cited evidence” means the pages an assistant linked to. Assistants may read pages they do not cite, so everything we say about evidence is a lower bound on what each one saw.

## What we analyzed

**Questions and answers.** 80 buyer questions, ten per industry, from national product questions (“Which robot vacuum is best?”) to local service questions (“Who are the best roofing contractors in Dallas?”); 24 name a place (23 cities and one state). Each was asked once of each assistant on 26 September 2026.

- **ChatGPT:** the consumer web app, via DataForSEO’s LLM Scraper, US location.
- **Gemini:** the consumer web app, via the same scraper, US location.
- **Perplexity:** the sonar model through its API, with web search.
- **Claude:** Claude Haiku 4.5 through the API, with web search.

**Recommendations, not mentions.** Version 1.0 counted a brand when its exact name appeared anywhere in an answer. That mixes recommendations with passing mentions and splits one firm written two ways into two brands. This version uses a model coder (Claude Opus) that read all four answers to a question at once, under shuffled letters so it could not tell which assistant wrote which. It listed every option a buyer could choose, merged spelling variants, marked whether each answer recommended the option or only mentioned it, and recorded the order of the recommendations. The version 1.0 rule found 70.6% of the options the coder found; 95.6% of the rule’s brands matched a coded option.

**Evidence.** We fetched all 2,209 unique cited URLs on 28 September, two days after the answers, and extracted the main text. 64.2% were readable; the rest blocked automated requests, were not web pages or returned too little text. An answer’s evidence counts as usable when at least half of its cited pages were readable. An option counts as “in” an answer’s evidence when any of its coded names appears in that text.

**Stability.** Our [consistency study](https://underneath.agency/research/ai-recommendation-consistency-study) asked 20 of these questions five times each of ChatGPT, Gemini and Perplexity on the same day. We reuse those answers to compare an assistant with itself and with the others.

**Causes.** Two model coders (Claude Opus and Claude Sonnet), working independently and blind to which assistant wrote which answer, coded every question for the reasons the four answers differ, using a fixed list of nine codes.

## Study 1: how much the assistants agree

### Pairwise agreement on recommended options

| Pair | Overlap | 95% interval |
|---|---|---|
| Gemini and Perplexity | 0.355 | 0.305 to 0.404 |
| ChatGPT and Gemini | 0.346 | 0.283 to 0.406 |
| ChatGPT and Perplexity | 0.336 | 0.285 to 0.391 |
| Gemini and Claude | 0.321 | 0.262 to 0.380 |
| Perplexity and Claude | 0.321 | 0.270 to 0.372 |
| ChatGPT and Claude | 0.284 | 0.230 to 0.342 |

Overlap is the Jaccard similarity of two answers’ recommended options: options both recommended, divided by options either recommended. The intervals overlap heavily, but comparing pairs within the same questions, the pairs do differ (F test, p = 0.0098), driven by ChatGPT and Claude agreeing least.

Of the 1,257 option-and-question combinations, 66.3% were recommended by one assistant only (61.5% to 70.5%) and 8.2% by all four. All four put the same option first for 10.0% of questions, and at least three agreed on the first pick for 36.2%. Counting the first brand named rather than the first recommendation, as version 1.0 did, all four agreed for 11.2% of questions.

### By industry and by place

| Industry | Overlap | 95% interval |
|---|---|---|
| B2B software and technology | 0.543 | 0.462 to 0.637 |
| Hospitality and travel | 0.385 | 0.244 to 0.532 |
| Financial services and insurance | 0.332 | 0.227 to 0.455 |
| Healthcare and dental | 0.294 | 0.192 to 0.412 |
| Retail and ecommerce | 0.289 | 0.191 to 0.393 |
| Franchises and multi-location brands | 0.265 | 0.161 to 0.378 |
| Legal and professional services | 0.236 | 0.127 to 0.368 |
| Home and local services | 0.227 | 0.150 to 0.319 |

Each industry rests on ten questions, so most intervals overlap; only B2B software stands clearly apart. Industry differences are real as a group (F test, p = 0.0087), but much of them is about place: whether a question names a place explains 26.9% of the variation in question-level overlap, and industry and place together 45.4%.

Questions naming a place had a mean overlap of 0.160 (0.124 to 0.197); national questions 0.390 (0.338 to 0.445). In a model with a random intercept for each question and controls for industry, pair, source overlap, number of sources and list length, naming a place still lowered overlap by 0.232.

### Each assistant’s recommendation profile

| Assistant | Median picks | Share of all picks | Picks no other made |
|---|---|---|---|
| ChatGPT | 6 | 52.4% | 39.7% |
| Perplexity | 5 | 45.6% | 31.7% |
| Gemini | 5 | 42.0% | 28.6% |
| Claude | 5 | 41.6% | 36.7% |

“Share of all picks” is the share of the options recommended by any assistant for a question that this assistant recommended. ChatGPT recommended the most options and the most that no other assistant named.

## Study 2: do different sources explain different picks?

The assistants cite little in common. The mean overlap of cited registrable domains between two assistants for the same question was 0.079 (0.068 to 0.092), and 43.8% of pairs shared no cited domain. If each assistant simply recommended what its own sources list, pairs that share more sources should share more picks. They do, but only slightly.

| Model (397 question pairs) | Per 0.1 more domain overlap | 95% interval |
|---|---|---|
| Source overlap alone | +0.021 | −0.014 to +0.056 |
| Plus place, industry, pair, sources, list length | +0.020 | −0.011 to +0.052 |
| Within the same question | +0.025 | −0.004 to +0.055 |
| Question random intercept, with controls | +0.025 | +0.004 to +0.046 |

A bootstrap over questions gave −0.011 to +0.058 for the first slope. Pairs that shared no cited domain still shared 0.291 of their picks; pairs that shared at least one shared 0.355. Source overlap explains 0.6% of the variation on its own.

It also does not explain the local gap. Local questions had somewhat lower domain overlap (0.063 against 0.088), but adding domain overlap to the model moved the local effect from 0.233 to 0.231, 0.9% of the gap.

Cited-domain overlap is a coarse measure: two assistants can cite different sites that list the same firms. Study 3 therefore looks inside the pages.

## Study 3: evidence or selection?

### Were the recommended options in the assistant’s own sources?

| Assistant | Picks in own sources | Picked if listed | Picked if not |
|---|---|---|---|
| Perplexity | 89.1% | 47.7% | 12.6% |
| Claude | 81.2% | 53.2% | 11.7% |
| Gemini | 78.4% | 52.4% | 16.0% |
| ChatGPT | 52.0% | 54.2% | 36.9% |

Answers with usable evidence only. “Picks in own sources” is the share of the assistant’s recommendations that its own cited pages named. The last two columns use every option any assistant named for that question: of those the assistant’s cited pages listed, the share it recommended; of those they did not, the share it recommended.

Perplexity, Claude and Gemini mostly recommend options their cited pages name. ChatGPT is the outlier because of local questions: 26 of its 80 answers included a business list with addresses, and only 27.9% of its local picks appeared in its cited pages, against 75.9% on national questions. Its local picks likely come from a business-listing source it does not cite as a page. Once an option was in its evidence, each assistant recommended it roughly half the time (47.7% to 54.2%): all four select, rather than repeat, what they read.

### Why one assistant recommends an option and the other does not

For each pair of assistants and each option one recommended and the other did not, we checked where the option appeared (1,912 one-sided picks, both answers with usable evidence).

| Where the option was named | All questions | National | Local |
|---|---|---|---|
| In the other assistant’s cited pages (selection) | 42.1% | 56.5% | 19.4% |
| Only in the recommender’s cited pages (evidence) | 28.8% | 30.3% | 26.5% |
| In neither’s cited pages | 29.1% | 13.2% | 54.2% |

On national questions, more than half of the one-sided picks were in pages the other assistant had cited: it had the option in front of it and did not recommend it. On local questions, more than half came from outside either assistant’s cited evidence, which is consistent with local picks coming from business listings, maps or the model’s own knowledge rather than from the pages it links to.

Restricting to answers whose cited pages were all readable (158 one-sided picks) moved the split to 31.6% selection, 39.2% evidence and 29.1% neither. The selection share is sensitive to how much of the evidence can be read; that it is large is not.

A page naming an option is not the same as a page recommending it. The selection share counts options a page named for any reason, including in a comparison or as an also-ran.

## Study 4: is the disagreement just noise?

AI answers change from run to run, so part of the gap between two assistants would appear even between two runs of the same one. On the 20 questions asked five times of three assistants (brand names matched by the version 1.0 rule):

| Measure | ChatGPT | Gemini | Perplexity |
|---|---|---|---|
| Overlap with its own other runs | 0.530 | 0.421 | 0.682 |
| Same first pick as its own other runs | 54.7% | 53.9% | 63.6% |

Between two assistants, over every combination of their five runs, the figures were much lower.

| Pair | Overlap across all 25 run pairs | Same first pick |
|---|---|---|
| ChatGPT and Perplexity | 0.288 | 25.0% |
| Gemini and Perplexity | 0.268 | 16.4% |
| ChatGPT and Gemini | 0.257 | 25.0% |

Cross-assistant overlap was about half of what each assistant’s own run-to-run overlap would allow (0.478 to 0.544 of it). Comparing each assistant’s typical set, the brands it named in at least three of five runs, still gave overlaps of only 0.325 to 0.380. And of the brands one assistant named in at least three of five runs, 50.0% to 56.6% were never named by the other in any of its five. Run-to-run variation is large, but the assistants also differ systematically.

## Study 5: the reasons visible in the answers

The two coders agreed on the main reason for 58.8% of questions (Cohen’s kappa 0.433, moderate). We report a reason only where both coders gave it.

| Reason (both coders) | All questions | Local | National |
|---|---|---|---|
| Many valid options; each picks a different subset | 51.2% | 91.7% | 33.9% |
| The answers substantially agree | 33.8% | 8.3% | 44.6% |
| Each names what its own, different sources list | 25.0% | 54.2% | 12.5% |
| At least one answer names no options | 17.5% | 12.5% | 19.6% |

Among the 40 questions with the lowest overlap, both coders saw a crowded field of valid options in 82.5% and source-driven differences in 45.0%. The coders did not agree reliably on a different reading of the question (kappa 0.195) or on suspect entities (one coder flagged 15.0% of questions, the other none), so we do not report those codes as findings.

The coders’ view fits the numbers: in crowded local categories, each assistant draws a different short list from a long one, and the sources it cites differ too; in national categories with clear leaders, the assistants mostly agree and differ at the margins.

## Observed, inferred and unknown

- **Observed:** which options each answer recommended and in what order; which pages each answer cited; whether those pages, fetched two days later, named each option; how answers varied across five runs for 20 questions.
- **Inferred:** that most national disagreement is selection rather than retrieval rests on cited pages only, and pages the assistants read but did not cite could change it; that ChatGPT’s local picks come from a business listing is our reading of its answer format.
- **Unknown:** what each assistant retrieved but did not cite; whether the reasons the answers give for a pick are supported by the sources; whether any change to a brand’s website would change these outcomes.

## What this means

These are our interpretations, not additional findings.

- **Measure each assistant, and each stage, separately.** A brand can be recommended by one assistant and absent from another. Whether it appears in an assistant’s sources and whether that assistant then picks it are different problems.
- **Being in the sources is necessary but not enough.** For national categories, most one-sided picks were already in the other assistant’s evidence. Being listed on the pages an assistant reads raises the odds of a recommendation (roughly half of listed options were picked, against a small share of unlisted ones) but does not decide it.
- **Local visibility is a different problem.** Local picks often do not trace to any cited page. For local businesses, listings and profiles the assistants draw on without citing may matter more than articles.

## What comes next

This version works only with answers already collected. Three steps would turn these associations into causal estimates, and they are planned for the next versions:

1. Same evidence, different assistants. Give all four the same fixed set of pages for a question, in randomized order, and measure how much disagreement remains when retrieval is held constant.
2. Repeated runs for all four assistants. A stratified subset of questions asked several times on several days, with rewordings, so each option gets a recommendation probability instead of a single yes or no.
3. Content changes. Controlled changes to test pages, with retrieval, citation, recommendation and rank measured separately, to see which stage each change moves.

## How this compares with other studies

| Source | Sample and date | Figure |
|---|---|---|
| This study | 80 questions, 4 assistants, September 2026 | Recommendation overlap 0.327; all four agree on the first pick 10.0% |
| BrightEdge | ChatGPT, AI Overviews and AI Mode, August 2025 | 61.9% of queries get different brand recommendations across platforms; 17% the same brands on all three |
| BrightEdge | Five engines, category-level top-30 brand lists, May 2026 | Pairwise overlap of top-named brands between 36% and 55% |
| BrightEdge | ChatGPT and Gemini, June 2026 | Agree on about 2 of their top 5 brands |

BrightEdge compares category-level brand lists built from many prompts, which smooths out the variation of single answers; our figures compare the answers a buyer sees to one question, which may be why they sit at the low end of BrightEdge’s range. None of the BrightEdge comparisons includes Claude, publishes its prompts or examines the cited pages.

Sources: [BrightEdge, August 2025](https://www.brightedge.com/resources/weekly-ai-search-insights/chatgpt-vs-google-ai-62-brand-recommendation-disagreement); [BrightEdge, May 2026](https://www.brightedge.com/resources/weekly-ai-search-insights/where-ai-engines-agree-on-brands); [BrightEdge, June 2026](https://www.brightedge.com/resources/weekly-ai-search-insights/chatgpt-vs-gemini-same-question-different-brands).

## Methodology

- **Questions:** 80, ten per industry, written as a buyer would ask them; 24 name a place (23 cities, one state). Published in the dataset.
- **Collection:** 26 September 2026, one run per question per assistant, via DataForSEO (LLM Scraper for ChatGPT and Gemini with US location; LLM Responses API for Perplexity sonar and Claude Haiku 4.5 with web search). No new answers were collected for version 1.1.
- **Recommendations:** model-coded by Claude Opus from all four answers at once, engine hidden; options merged across spelling variants; recommended or only mentioned; rank of each recommendation. First pick = rank 1.
- **Agreement:** Jaccard similarity of recommended sets per question and pair, averaged over questions where both answers recommended something. 95% intervals by bootstrap over questions (2,000 resamples). F tests for pair (within questions) and industry.
- **Source models:** ordinary least squares with errors clustered by question, a question fixed-effects model and a question random-intercept model; domain overlap in steps of 0.1.
- **Evidence:** every cited URL fetched on 28 September 2026; readable = HTTP 200 and at least 400 characters of text that is not a block page; usable evidence = at least half of an answer’s cited pages readable; an option is in the evidence if any of its coded names appears in the page text.
- **Stability:** study 8’s five runs of 20 questions for ChatGPT, Gemini and Perplexity, brand names matched by the version 1.0 rule.
- **Reasons:** two model coders, blind to engine and to each other, nine fixed codes, agreement by Cohen’s kappa; codes reported where both agree.
- **Update schedule:** quarterly, same questions.

## Limitations

- One run per assistant for the main sample; stability covers 20 questions and three assistants, not Claude.
- Cited pages are not everything an assistant read, so the evidence and selection shares are estimates from what is visible. 35.8% of cited URLs could not be read, and pages were fetched two days after the answers.
- A page naming an option may name it negatively or in passing; name matching can miss variants.
- Associations only: nothing here changes the evidence an assistant sees, so none of it shows that a source caused a pick.
- All coding is by language models (two Claude models); no person coded this sample.
- Perplexity and Claude were queried through their APIs, which may differ from their consumer apps; Gemini’s scraped answers came from its Flash-Lite model.
- Questions were asked without a user location for the API assistants; city names in the question supply the location.

## What changed in version 1.1

- **The question.** From “do the assistants agree?” to “where do they diverge?”, with separate measures for cited evidence, selection, recommendation and stability.
- **The measure.** Recommended options, coded by a model from the full answers, replace brand names matched anywhere in the text. Headline overlap moves from 0.363 (mentions) to 0.327 (recommendations); the share named by one assistant only from 56.7% to 66.3%; all four agreeing on the first pick from 11.2% to 10.0%. Merging near-duplicate names under the version 1.0 rule would have raised its 0.363 to 0.377.
- **Uncertainty.** 95% intervals for every headline figure, resampling questions.
- **Sources.** Version 1.0 said the assistants “cited almost entirely different sources” and that this “may be one reason” for different picks. That is now stated as low cited-domain overlap (0.079) and tested: it explains little.
- **New measures.** Cited pages fetched and checked; repeated-run comparison; reasons for disagreement coded by two models.

## Data and downloads

- Every answer with its recommended options, first pick and cited domains: [s7_answers_v11.csv](https://underneath.agency/research-data/ai-assistants-brand-agreement-study/s7_answers_v11.csv)
- Coded options per answer (recommended or mentioned, rank): [s7_entities_coded.csv](https://underneath.agency/research-data/ai-assistants-brand-agreement-study/s7_entities_coded.csv)
- Question-by-pair table with recommendation and source overlap: [question_pairs_pipeline.csv](https://underneath.agency/research-data/ai-assistants-brand-agreement-study/question_pairs_pipeline.csv)
- Reasons for disagreement, both coders: [s7_disagreement_causes.csv](https://underneath.agency/research-data/ai-assistants-brand-agreement-study/s7_disagreement_causes.csv)
- Version 1.0 answers, brands and 80 questions: [s7_s8_answers.csv](https://underneath.agency/research-data/ai-assistants-brand-agreement-study/s7_s8_answers.csv)
- Every statistic on this page: [stats.json](https://underneath.agency/research-data/ai-assistants-brand-agreement-study/stats.json) (version 1.0: [stats_v1.0.json](https://underneath.agency/research-data/ai-assistants-brand-agreement-study/stats_v1.0.json))
- Machine-readable methodology: [methodology.json](https://underneath.agency/research-data/ai-assistants-brand-agreement-study/methodology.json)

The data is free to reuse with attribution (CC BY 4.0). Code: cite/pipeline/s7v11_*.py.

To cite: Underneath. (2026). *Do ChatGPT, Gemini, Perplexity and Claude agree on brands?* (Version 1.1). Underneath Research. https://underneath.agency/research/ai-assistants-brand-agreement-study

## Frequently asked questions

### Do ChatGPT and Gemini recommend the same brands?

Partly. For the same buyer question, their recommended options overlapped by 0.346 on a scale from 0 (nothing in common) to 1 (identical). All four assistants put the same option first for 10.0% of questions.

### Why do AI assistants recommend different brands?

Mostly because they choose differently from similar evidence, not because they read completely different sources. On national questions, 56.5% of the options one assistant recommended and another did not were named in the second assistant’s own cited pages. Local questions differ: most one-sided local picks appear in neither assistant’s cited pages.

### Do AI assistants agree more on some industries than others?

Yes. Overlap was highest for B2B software (0.543) and lowest for home and local services (0.227). Questions about a specific place had less than half the agreement of national questions.

### Is the disagreement just randomness?

No. Each assistant’s answers overlapped with its own repeat runs about twice as much as with another assistant’s, and half or more of the brands one assistant named in most runs never appeared in the other’s five runs.

## Related research

- [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)
- [ChatGPT local recommendations vs Google Maps](https://underneath.agency/research/chatgpt-local-recommendations-study)
- [AI Mode vs AI Overviews: how different are the sources?](https://underneath.agency/research/ai-mode-vs-ai-overviews-study)
- [Same question, four countries: do AI recommendations change?](https://underneath.agency/research/ai-recommendations-by-country-study)

## Related guides

- [Is tracking ChatGPT alone enough to measure our AI search visibility?](https://underneath.agency/resources/is-tracking-chatgpt-enough)
- [Do Gemini, GPT and Claude prefer the same kind of content?](https://underneath.agency/resources/do-ai-engines-prefer-same-content)
- [Do ChatGPT, Copilot, Google and Perplexity cite the same sources?](https://underneath.agency/resources/do-ai-search-engines-cite-the-same-sources)
- [Can we combine ChatGPT, Gemini and Perplexity into one AI visibility score?](https://underneath.agency/resources/combine-ai-engines-visibility-score)
- [How can an IT services firm win more projects and support calls from AI search?](https://underneath.agency/resources/it-service-firms-leads-ai-search)

---

This is the Markdown twin of https://underneath.agency/research/ai-assistants-brand-agreement-study. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Do AI answers match a business’s Google profile? | Underneath"
description: "We asked 4 AI engines for the address, phone, website and hours of 159 local businesses. 18.9% of answers had a fact that differed from the Google profile."
canonical: "https://underneath.agency/research/ai-business-facts-accuracy-study"
published: 2026-09-26
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Research · Local AI search

# Do AI answers match a business’s Google profile?

People now ask AI assistants for a business’s phone number or opening hours instead of searching for it. On 26 September 2026 we asked ChatGPT, Gemini, Perplexity and Google AI Mode for the address, phone number, website and opening hours of 159 real local businesses in the United States, United Kingdom, Canada and Australia. We compared every answer with the business’s Google Business Profile. When a phone number differed, we also looked for it on the business’s own website and on the pages the answer cited. Most answers matched the profile. Most phone numbers that did not match were numbers the business also publishes itself.

## The short version

1. 18.9% of the 636 answers (95% interval 15.1% to 22.6%) stated at least one fact that differed from the business’s Google profile; in 73.1% every fact matched.
2. The engines differed, and the differences hold up when each business is compared with itself across engines. 4.4% of Gemini’s answers had a differing fact, against 16.4% for ChatGPT, 17.6% for Google AI Mode and 37.1% for Perplexity. Perplexity’s rate was 32.7 percentage points above Gemini’s. ChatGPT and Google AI Mode could not be told apart.
3. Differing from the profile is not the same as being wrong. Of the 55 phone numbers that differed from the profile, 42 were on the business’s own website and 7 more on a page the same answer cited. Only 3 were found in neither place; for 3 more the website could not be checked. Counting a number on the business’s website as consistent, 97.9% of the phone numbers given matched one of the business’s own sources.
4. Phone numbers differed mostly for businesses whose own sources disagree. Where the profile number did not appear on the business’s website, 30.6% of the numbers given differed from the profile. Where it did appear, 1.6% did.
5. These are single answers from one day, one wording and one sampling frame (businesses already in the Google Maps top 20). They show how often each engine differed from the profile in this sample. They do not rank the engines’ reliability in general.

## What this study measures, and what it does not

This study measures **agreement with a reference record**: does the fact an assistant states match the business’s Google Business Profile? The profile is a reasonable reference, because it is the record the business controls on the most-used local search surface. It is not proof of the truth. A profile can be out of date or show a call-tracking number, and a business can publish more than one valid number.

We therefore report three things, and keep them apart:

| Measure | Question | Measured here? |
|---|---|---|
| Agreement with the profile | Does the stated fact match the Google Business Profile? | Yes, for all four facts |
| Consistency with the business’s own sources | Does it match the profile or the business’s own website? | Yes, for phone numbers |
| Verified truth | Is it the business’s current, correct fact? | No: nothing was confirmed with the businesses |

Two further limits define the scope. We asked about businesses by name, so the study measures how an assistant **describes** a business it is asked about, not which businesses it **chooses to recommend**; our [ChatGPT local recommendations study](https://underneath.agency/research/chatgpt-local-recommendations-study) covers that. And matching a number on a website or cited page shows that the sources agree. It does not show where the assistant took the number from.

## What we asked and how we scored it

We drew 160 independent businesses at random from the Google Maps top 20 for five services (dentist, plumber, personal injury lawyer, accountant, physiotherapist) in six cities in each of four countries, 8 per service per country. Choosing businesses from Google Maps rather than from any AI answer keeps the test fair to every engine, but it also means every business was already visible in Google local search. We left out chains present in more than one city and directory listings. One business was later excluded because our shortened version of its name was ambiguous, leaving 159.

Each engine was asked once: “What is the address, phone number, website and opening hours of {business} in {city}?”, using the name as the business presents it before any keyword suffix (for example “WT Law”, not “WT Law - Car Accident Lawyers Brisbane”). All 636 answers were collected between 07:52 and 08:14 UTC on 26 September 2026.

| Engine | Queried through | Model reported |
|---|---|---|
| ChatGPT | Consumer app | Not reported for 151; gpt-5-6 for 8 |
| Gemini | Consumer app | 3.5 Flash-Lite |
| Google AI Mode | Google Search, desktop | Not reported |
| Perplexity | Sonar API, web search on | sonar |

ChatGPT and Gemini were reached through DataForSEO’s LLM Scraper and Google AI Mode through its SERP API, all located in the business’s country. Perplexity was reached through DataForSEO’s LLM Responses API, which has no location setting.

Every answer was scored by fixed rules against the Google Business Profile as shown on Google Maps the same day:

- **Phone:** matches when the last nine digits match. A differing number was then searched for on the business’s own website (homepage and contact pages, fetched the same day) and on every page the answer cited.
- **Address:** matches when both the street number and the postcode match.
- **Website:** matches when the business’s own domain appears, and differs when the answer gives another domain as the website. (Version 1.0 could only score a website as matching or missing; see “Checking the scoring”.)
- **Monday hours:** match when the first opening time and last closing time match. Answers that said the hours were unclear or conflicting count as not given. 34 businesses whose profile says “open 24 hours”, usually a placeholder, were not scored for hours.

## Findings

### Agreement with the profile, by engine

| Engine | A fact differs (95% interval) | Every fact matches |
|---|---|---|
| Gemini | 4.4% (1.3% to 7.5%) | 94.3% |
| ChatGPT | 16.4% (10.7% to 22.6%) | 81.8% |
| Google AI Mode | 17.6% (11.9% to 23.9%) | 72.3% |
| Perplexity | 37.1% (29.6% to 44.6%) | 44.0% |

Share of each fact that matched the profile:

| Engine | Phone | Address | Website | Monday hours |
|---|---|---|---|---|
| Gemini | 99.4% | 98.7% | 99.4% | 94.4% |
| ChatGPT | 91.2% | 96.2% | 99.4% | 88.8% |
| Google AI Mode | 89.3% | 93.1% | 99.4% | 80.0% |
| Perplexity | 83.6% | 91.2% | 97.5% | 51.2% |

Shares are of all answers, so a fact that was not given counts as not matching. Hours columns cover the 125 businesses with regular hours on their profile; the other columns cover all 159. Intervals come from 4,000 resamples of the 159 businesses, each keeping its four answers.

### Are the differences between engines real?

Every business was asked of all four engines, so the engines are compared business by business rather than as four separate samples. Taken together the four rates differ (Cochran’s Q = 71.9, p < 0.001). Pair by pair, with a Holm correction for the six comparisons:

| Pair | Gap, points (95% interval) | Only first / only second | Holm p |
|---|---|---|---|
| Perplexity vs Gemini | 32.7 points (24.5 to 40.2) | 55 / 3 | < 0.001 |
| Perplexity vs ChatGPT | 20.8 points (13.2 to 28.3) | 38 / 5 | < 0.001 |
| Perplexity vs Google AI Mode | 19.5 points (11.3 to 27.7) | 39 / 8 | < 0.001 |
| Google AI Mode vs Gemini | 13.2 points (7.5 to 19.5) | 24 / 3 | < 0.001 |
| ChatGPT vs Gemini | 11.9 points (5.7 to 18.9) | 24 / 5 | 0.001 |
| Google AI Mode vs ChatGPT | 1.3 points (−5.7 to 8.2) | 15 / 13 | 0.851 |

The gap is the difference in the share of answers with a differing fact. “Only first / only second” counts the businesses where only the first engine’s answer had a differing fact, and those where only the second engine’s did.

### How each fact fared when an engine gave it

| Fact | Given and matching the profile | Given but different | Not given |
|---|---|---|---|
| Website | 98.9% | 1.1% | 0.0% |
| Address | 94.8% | 5.0% | 0.2% |
| Phone number | 90.9% | 8.6% | 0.5% |
| Monday opening hours | 78.6% | 9.4% | 12.0% |

Of the facts that were given, 98.9% of websites, 95.0% of addresses, 91.3% of phone numbers and 89.3% of Monday hours matched the profile. Perplexity declined to give Monday hours far more often (27.2% of its answers) and, when it gave them, matched the profile 70.3% of the time.

### Where the differing phone numbers could be found

| The 55 phone numbers that differed from the Google profile | Answers |
|---|---|
| Published on the business’s own website | 42 |
| Not on the website, but on a page the same answer cited | 6 |
| Website could not be checked; on a page the answer cited | 1 |
| Website could not be checked; not on any readable cited page | 3 |
| Found on neither the website nor any cited page | 3 |

Phone numbers by engine:

| Engine | Matches profile | Profile or website | Found nowhere |
|---|---|---|---|
| Gemini | 99.4% | 100.0% | 0 |
| ChatGPT | 91.2% | 98.7% | 1 |
| Google AI Mode | 90.4% | 97.5% | 2 |
| Perplexity | 84.2% | 95.6% | 0 |
| All engines | 91.3% | 97.9% | 3 |

Shares are of the phone numbers given. “Profile or website” counts a number that matches the profile or appears on the business’s own website; “Found nowhere” counts answers whose number was on neither and on no readable cited page.

For 24.2% of the businesses we could check, the number on the Google profile did not appear anywhere on their own website, so the business publishes at least two numbers. Those businesses account for most of the differing numbers: 30.6% of the phone numbers given for them differed from the profile, against 1.6% for businesses whose website shows the profile number. Call-tracking numbers are one common reason a business shows a different number on Google, but our data does not show which number is the main line.

This does not tell us which source an assistant used. It tells us that when an assistant’s number differed from the profile, the same number was usually published somewhere the business or the answer’s own sources put it: of the 52 differing numbers whose answer cited at least one readable page, 47 appeared on a cited page.

### What the answers cited

The answers contained 4,011 citations. Engines cited very different kinds of source:

| Engine | Any citation | Median | Own site | Google | Other |
|---|---|---|---|---|---|
| ChatGPT | 100.0% | 3 | 96.2% | 0.0% | 67.9% |
| Gemini | 86.2% | 1 | 5.0% | 79.9% | 3.1% |
| Google AI Mode | 99.4% | 2 | 67.3% | 73.0% | 22.6% |
| Perplexity | 100.0% | 19 | 96.2% | 1.3% | 100.0% |

Columns are the share of answers citing anything, the median number of citations, and the share citing the business’s own website, a Google page and any other site.

Gemini mostly cited Google pages, usually its Maps listing, and almost never the business’s website, and it had the highest agreement with the Google profile. Phone numbers given in answers that cited a Google page matched the profile 97.5% of the time, against 88.1% in answers that cited the business’s website and 86.6% in answers that cited another site. This is an association across engines, not a test: citing Google and being Gemini largely coincide, and a citation list does not show which page a number came from. We fetched 2,362 of the 3,230 distinct non-Google pages cited, on 28 September 2026, two days after the answers; a page may have changed in between.

### By country and service

| Country | Businesses | A fact differs (95% interval) | Hours left out |
|---|---|---|---|
| Canada | 40 | 32.5% (22.7% to 42.4%) | 30.0% |
| United Kingdom | 40 | 15.6% (9.7% to 22.6%) | 5.6% |
| Australia | 39 | 13.5% (7.8% to 19.9%) | 6.4% |
| United States | 40 | 13.8% (7.5% to 20.3%) | 11.9% |

“Hours left out” is the same share when Monday hours are not scored.

| Service | Businesses | A fact differs (95% interval) |
|---|---|---|
| Plumber | 32 | 21.9% (12.9% to 31.9%) |
| Physiotherapist | 32 | 21.1% (13.0% to 29.5%) |
| Accountant | 31 | 17.7% (8.9% to 28.2%) |
| Dentist | 32 | 17.2% (9.5% to 26.0%) |
| Personal injury lawyer | 32 | 16.4% (9.1% to 24.3%) |

The countries and services contain different businesses, so we also fitted one model of whether an answer had a differing fact, with engine, country and service together and errors clustered by business. Canada’s higher rate remained after allowing for engine and service (odds ratio 3.48 against the United States, 95% interval 1.61 to 7.54; country overall p = 0.0021). The services did not differ (p = 0.847). In Canada the gap came mostly from Google AI Mode (45.0% of its Canadian answers had a differing fact) and ChatGPT (27.5%), and it remains when hours are left out. With about 40 businesses per country, this is an observed difference in this sample, not an explanation of it; we did not measure what distinguishes the Canadian businesses or their listings.

Whether some engines differ more in some countries could not be fully tested. Gemini had no differing answer in Australia and only one to three elsewhere, which makes its country terms unreliable. Among the other three engines, the model found no clear engine-by-country interaction (p = 0.092).

### Monday hours: what was left out

The 34 businesses excluded from the hours comparison were not a random set. They were 24 of the 32 plumbers and 10 of the 32 personal injury lawyers, and no dentists, accountants or physiotherapists. By country they were 14 of 40 in the United States, 9 of 40 in Canada, 7 of 39 in Australia and 4 of 40 in the United Kingdom. The hours results therefore describe plumbers and lawyers poorly. Two checks show the headline does not depend on this choice: leaving hours out altogether, 13.5% of answers had a differing fact, with the same engine order (Gemini 2.5%, ChatGPT 11.9%, Google AI Mode 15.7%, Perplexity 23.9%). Scoring the 24-hour profiles as they stand gives 20.4% (Gemini 4.4%, ChatGPT 19.5%, Google AI Mode 18.9%, Perplexity 39.0%). The comparison covers Monday only. No profile in the sample had split Monday hours, and two were closed on Mondays.

### ChatGPT: being listed and being described

Every business in this study came from the same Google Maps searches as our ChatGPT local recommendations study, which asked ChatGPT “Who is the best {service} in {city}?” four times per search on 26 and 27 September. That lets us set how often ChatGPT **lists** a business beside how well it **describes** it when asked by name:

| Listed by ChatGPT | Businesses | Every fact matched |
|---|---|---|
| In all four | 30 | 90.0% |
| In one to three | 55 | 85.5% |
| In none | 74 | 79.7% |

“Listed by ChatGPT” counts how many of its four “best {service}” answers named the business; “every fact matched” is from the separate question about the business by name.

53.5% of the businesses were listed at least once. Businesses ChatGPT listed were described slightly more often without a differing fact, but the difference is within chance for a sample this size (p = 0.283). Being recommended and being described consistently are separate questions, and a business can do well on one and poorly on the other.

## Checking the scoring

We checked the scoring rules against a blind reading of the answers. We drew a stratified random sample of 136 scored decisions (38 phone, 36 address, 45 Monday hours and 17 website), oversampling facts the rules had scored as differing. Two independent AI coders (Claude Sonnet and Claude Opus) each read the answer and the profile value without seeing the rules’ label. Where the coders disagreed, we settled the label from the raw answer. No person coded the sample.

- The two coders agreed on 93.4% of the decisions (Cohen’s kappa 0.89).
- The check found one systematic gap. The 1.0 website rule could only score a website as matching or missing, and all 7 answers it counted as giving no website in fact named a different site, such as another clinic or a directory page. We corrected the rule and rescored all 636 answers. Only those 7 website scores changed, which moved the share of answers with a differing fact from 18.6% to 18.9%, Gemini’s from 3.8% to 4.4% and Perplexity’s from 36.5% to 37.1%.
- With the corrected rule, the rules and the settled labels agreed on 128 of 136 decisions (kappa 0.9). Of the 67 sampled facts the rules scored as differing, 64 were confirmed and 3 in fact matched the profile. The other 5 misreads were matching values the rules scored as not given.

The remaining rule errors are few and run in both directions; if anything, the differing rates above are slightly overstated.

## How this compares with other studies

| Source | Sample and date | Figure |
|---|---|---|
| This study | 159 local businesses, 636 answers from 4 engines, September 2026 | 8.6% of answers gave a phone number different from the Google profile; 18.9% had any fact that differed from it |
| Seer Interactive | 178 phone number questions about large brands, 7 AI models, December 2025 | 36% of phone numbers were not on the brands’ customer service pages; Google profiles matched least often |
| Searchable (reported by Search Engine Journal) | 72,000+ questions about UK high street retailers, July 2026 | One in 16 answers incorrect; Perplexity 10%, Gemini 5%, ChatGPT 4% |
| Searchable (reported by Search Engine Journal) | 13,365 questions about 165 London businesses, July 2026 | 93% of businesses had at least one basic fact wrong or missing |

In both our sample and Searchable’s, Perplexity had the highest observed rate of the engines compared. The studies use different reference records: Seer compared numbers with brands’ own customer service pages, we compared them with Google profiles, and each finds that the two records often disagree for some businesses. Rates differ because the studies ask about different businesses and facts and measure against different references.

Sources: [Seer Interactive](https://www.seerinteractive.com/insights/ai-models-provide-incorrect-phone-numbers-36-of-the-time-heres-what-you-can-do); [Search Engine Journal on Searchable](https://www.searchenginejournal.com/ai-answers-about-your-locations-are-often-wrong-check-before-customers-do/582205/); related: [our ChatGPT local recommendations study](https://underneath.agency/research/chatgpt-local-recommendations-study).

## What this means

This section sets out what we take from the results. It is interpretation, not measurement.

- **Publish one set of facts everywhere.** The phone numbers that differed from the profile were concentrated in businesses whose website and Google profile show different numbers. We cannot say which source the assistants read, but a single phone number, address and set of hours, identical on the website, the Google profile and directories, leaves an assistant nothing to choose between.
- **Check each assistant separately.** In this sample the engines differed widely, and one answer on one day is only a snapshot. Asking each assistant about your business from time to time shows what customers are told.
- **Hours need the most care.** Monday hours were the fact most often missing or different. Keep hours current everywhere, including on the website in plain text.

## Methodology

- **Businesses:** 160 drawn at random (seed 20260927) from the Google Maps top 20 for “best {service} in {city}” across 120 searches collected for our ChatGPT local study; required phone, address, website and hours on the profile; multi-city chains and directory listings excluded; 1 excluded after collection (ambiguous shortened name), leaving 159.
- **Engines and date:** as in the table above; one question per business per engine on 26 September 2026, 07:52 to 08:14 UTC.
- **Reference data:** the business’s Google Business Profile fields as returned in Google Maps results on the same day; the business’s own homepage and contact pages, fetched the same day; every page each answer cited, fetched on 28 September 2026 (Google pages were not fetched).
- **Scoring:** the fixed rules described above, implemented in code with unit tests, and checked against a blind model-coded sample of 136 decisions (see “Checking the scoring”).
- **Uncertainty:** 95% intervals from 4,000 resamples of the 159 businesses, each keeping its four answers; engines compared with Cochran’s Q and exact McNemar tests on the same businesses, Holm-corrected; a logistic model with engine, country and service and business-clustered robust errors for the country and service comparisons.
- **Version 1.1 (28 September 2026):** the same 636 answers, reanalyzed after an independent methodological review. The corrected website rule changed 7 website scores and, with them, the headline rate (18.6% to 18.9%), Gemini’s rate (3.8% to 4.4%), Perplexity’s (36.5% to 37.1%), and the United States, dentist and personal injury lawyer rates. All other 1.0 figures are unchanged. See the changelog in the data folder.
- **Update schedule:** quarterly.

## Limitations

- The Google Business Profile is the reference. A fact that differs from it is not necessarily false, and no fact was confirmed with the business.
- One answer per engine per business, from one wording, on one morning. Generative answers vary between runs and with wording, so the engine rates are a snapshot, not a stable ranking.
- Finding a number on the business’s website or a cited page shows that the sources agree. It does not show which source the engine used.
- Cited pages were fetched two days after the answers, and 868 of the 3,230 could not be read.
- The businesses were already in the Google Maps top 20; businesses with a weak Google presence may fare differently.
- Perplexity was tested through its API, which may differ from its app.
- Hours were compared for Monday only, and 34 businesses, mostly plumbers, were left out of the hours comparison.
- About 40 businesses per country and 32 per service: enough to detect large differences, not small ones.
- The scoring rules still misread a few answers: in the validation sample they scored 3 matching facts as differing and missed 5 matching values, so the differing rates may be slightly overstated. The validation was coded by AI models, not by people.

## Data and downloads

- Every answer, scored field by field: [s11_business_facts.csv](https://underneath.agency/research-data/ai-business-facts-accuracy-study/s11_business_facts.csv) and [JSON](https://underneath.agency/research-data/ai-business-facts-accuracy-study/s11_business_facts.json)
- Every phone answer with its source class and what the answer cited: [s11_phone_provenance.csv](https://underneath.agency/research-data/ai-business-facts-accuracy-study/s11_phone_provenance.csv)
- Every citation, its source type and the phone numbers on the cited page: [s11_citations.csv](https://underneath.agency/research-data/ai-business-facts-accuracy-study/s11_citations.csv)
- The validation sample with both coders’ labels: [s11_validation.csv](https://underneath.agency/research-data/ai-business-facts-accuracy-study/s11_validation.csv)
- Every statistic on this page, with intervals and tests: [stats.json](https://underneath.agency/research-data/ai-business-facts-accuracy-study/stats.json)
- Machine-readable methodology: [methodology.json](https://underneath.agency/research-data/ai-business-facts-accuracy-study/methodology.json)

The data is free to reuse with attribution (CC BY 4.0).

To cite: Underneath. (2026). *Do AI answers match a business’s Google profile?* (Version 1.1). Underneath Research. https://underneath.agency/research/ai-business-facts-accuracy-study

## Frequently asked questions

### Do AI assistants give the right phone number for a business?

In our test of 159 local businesses, 91.3% of the phone numbers the engines gave matched the business’s Google profile, and 97.9% matched either the profile or the business’s own website. Only 3 of the 633 numbers given could not be found on the profile, the website or any page the answer cited. We did not call the businesses, so this measures agreement with their published numbers, not whether each number reaches them.

### Which AI assistant agrees most often with a business’s Google profile?

In this sample, Gemini: 4.4% of its answers had a fact that differed from the business’s Google profile, against 37.1% for Perplexity. Most of Gemini’s answers cited a Google page. This comes from one answer per business on one day, so it describes this sample rather than a lasting ranking.

### Why do AI assistants give different opening hours?

Our data shows the answers, not the cause. Hours were the fact most often missing or different: 9.4% of answers gave Monday hours that differed from the Google profile, and 12.0% gave none. One plausible explanation, which this study does not test, is that businesses list different hours in different places.

### How can a business keep AI assistants’ answers about it consistent?

Most phone numbers that differed from the profile were numbers the business publishes itself. Use the same phone number, address and hours on the website, the Google Business Profile and directories, and check what each assistant says about the business regularly.

## Related research

- [ChatGPT local recommendations: stable details, shifting lists](https://underneath.agency/research/chatgpt-local-recommendations-study)
- [Do ChatGPT, Gemini, Perplexity and Claude agree on brands?](https://underneath.agency/research/ai-assistants-brand-agreement-study)
- [How often do AI Overviews appear?](https://underneath.agency/research/ai-overviews-frequency-study)
- [Which Google Maps businesses does ChatGPT recommend?](https://underneath.agency/research/chatgpt-local-picks-google-profile-study)

## Related guides

- [Do AI agents make up facts about my business, or just leave them out?](https://underneath.agency/resources/do-ai-agents-invent-or-omit-business-facts)
- [How do you fix wrong information about your brand in AI answers?](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers)
- [How can a hospital, clinic or practice win new patients from AI search?](https://underneath.agency/resources/healthcare-providers-patients-ai-search)
- [Who in my organization should own AI search visibility?](https://underneath.agency/resources/who-should-own-ai-search-visibility)
- [How can a CPA or tax firm win better clients through AI search?](https://underneath.agency/resources/accounting-firms-clients-ai-search)

---

This is the Markdown twin of https://underneath.agency/research/ai-business-facts-accuracy-study. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Do ChatGPT, Gemini, Perplexity and Claude cite pages that rank?"
description: "8.3% of ChatGPT’s citations rank in Google’s top 10 for the question, Claude’s 25.7%. We trace where the gap opens, stage by stage, over 80 buyer questions."
canonical: "https://underneath.agency/research/ai-citations-google-rankings-study"
published: 2026-09-26
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Research · AI assistants

# Do ChatGPT, Gemini, Perplexity and Claude cite pages that rank?

Search marketing often assumes that ranking on Google is the way into AI answers. That assumption treats an AI answer as one step, when it is several: the assistant decides to search, rewrites the question into its own searches, retrieves candidate pages from an index, and chooses which of them to cite. A page can drop out at any of these stages, and Google rankings only describe one of them.

This study measures how far Google rankings get you through that chain. For 80 US buyer questions on 26 September 2026 we compared every source ChatGPT, Gemini, Perplexity and Claude cited, 2,492 citations in all, with Google’s results for the question as typed. Version 2.0 then asks where the gap opens. We collected Google’s top 100 for the questions and for the 509 searches the assistants ran on them, sorted every citation by the route through which Google could have surfaced it, and modeled which top-ranking pages get cited. It is an observational study: it describes where Google’s view and the assistants’ choices part ways, not what would change them.

## The short version

1. Google’s top 10 for the question explains a minority of citations. 8.3% of the pages ChatGPT cited ranked there (95% interval 5.6% to 11.3%), Gemini 16.6%, Perplexity 14.2% and Claude 25.7%. Google’s own AI Overviews cite top-10 pages 28.7% of the time.
2. The searches the assistants run explain part of the rest. Adding pages in Google’s top 10 for the assistant’s own searches raises Claude from 25.7% to 35.6% and ChatGPT from 8.3% to 18.5% (for ChatGPT, using searches from a separate run as a stand-in).
3. A large share of citations are not in Google’s top 100 for the question or for any of those searches: 68.5% for ChatGPT, 66.8% for Perplexity, 49.8% for Gemini and 36.8% for Claude. For ChatGPT, Google rankings describe less than a third of what it cites.
4. Among the pages Google ranks in its top 10, ranking for the assistant’s own search predicts citation more strongly than ranking for the question. For ChatGPT, once its searches are accounted for, rank for the question no longer predicts citation (odds ratio 0.93, 95% interval 0.55 to 1.57).
5. Engines differ in ways the questions do not explain. Paired on the same questions, ChatGPT’s citations ranked less often than every other engine’s. Gemini cited ranked pages much less often for questions naming a city (15.6 points fewer, 95% interval 4.2 to 25.1). Version 1.0 also reported that Claude cited ranked pages more often for local questions; with intervals, that difference is not distinguishable from zero and is withdrawn.

## Research questions

| Question | Answered here? |
|---|---|
| RQ1. How often do cited pages rank on Google for the question as typed? | Yes |
| RQ2. How much of the gap do the assistants’ own searches account for? | Exactly for Claude; by proxy for ChatGPT and Gemini; not for Perplexity |
| RQ3. Among pages Google ranks, what predicts which get cited? | Yes, with models clustered by question |
| RQ4. Do the engines differ on the same questions? | Yes |
| RQ5. Do local questions behave differently? | Yes, with small subgroups |
| RQ6. Would changing a page change whether it is cited? | No: nothing was changed |
| RQ7. Are the results stable across runs, wordings and dates? | No: one run on one day |

## Where Google rankings sit in an AI answer

| Stage | What happens | What this study can see |
|---|---|---|
| Search activation | The assistant decides whether to search at all | Whether an answer cites anything |
| Query rewriting | The question becomes one or more searches | Claude’s own searches; stand-ins for ChatGPT and Gemini |
| Retrieval | An index returns candidate pages | Google’s ranking, as a stand-in for an index we cannot see |
| Selection and citation | The model picks which pages to cite | Which pages were cited, and in what order |
| Use in the answer | The cited page shapes what the answer says | Not measured |

Google rankings are evidence about the retrieval stage only, and only for Google’s index. Search activation already separates the engines: Claude cited sources in 61 of 80 answers and Gemini in 73, while ChatGPT and Perplexity cited sources in all 80. An answer that does not search cannot cite a page however well it ranks.

## What we measured

We used the 80 buyer questions of our [four-assistant study](https://underneath.agency/research/ai-assistants-brand-agreement-study), ten in each of eight industries, and the sources each assistant cited when answering them on 26 September 2026. ChatGPT and Gemini were queried through their consumer apps, and Perplexity and Claude through their APIs with web search. The primary comparison is with Google’s US top 10 for each question as typed, collected the same day. A citation ranks when its normalized address is one of those results. It shares a site when its domain has some page there.

For version 2.0 we collected Google’s US top 100 on 28 September for the 80 questions and for every distinct search the assistants ran on them, 576 searches in all. For Claude these are the 61 searches recorded in the same answers. For ChatGPT (296 searches) and Gemini (152) they come from a separate API run on the same questions and day, because the consumer apps do not expose their searches; they show the kind of search each assistant runs, not the searches behind these answers. Perplexity’s searches are not available.

Each citation is assigned to the first of four routes that fits, checked in this order.

- **A.** In Google’s top 10 for the question as typed.
- **B.** In Google’s top 10 for one of the assistant’s own searches.
- **C.** Anywhere in Google’s top 100 for the question or one of those searches.
- **D.** None of these: Google does not surface the page for any of these searches.

## Findings

### How often cited pages rank for the question

| Assistant | Citations | Page in top 10 | Site in top 10 | First three |
|---|---|---|---|---|
| ChatGPT | 384 | 8.3% | 22.9% | 8.3% |
| Perplexity | 1,588 | 14.2% | 26.5% | 23.8% |
| Gemini | 259 | 16.6% | 23.6% | 16.3% |
| Claude | 261 | 25.7% | 32.2% | 22.9% |
| AI Overviews | 4,051 | 28.7% | 43.0% | |

“First three” counts only the first three sources of each answer, which removes the effect of how many sources an engine lists. Perplexity lists a median of 20 per answer, the others 3 to 5, and its first three citations rank far more often than the rest (23.8% against 14.2% overall). The AI Overview figures come from 800 Google searches in our [citation study](https://underneath.agency/research/ai-overview-citations-study), so they are a reference point rather than a like-for-like comparison.

Intervals for the page-level share: ChatGPT 5.6% to 11.3%, Perplexity 11.9% to 16.6%, Gemini 11.3% to 22.2%, Claude 19.4% to 32.0%.

### How many answers cite a ranked page

| Assistant | Share of all 80 answers | 95% interval |
|---|---|---|
| ChatGPT | 32.5% | 22.5% to 42.5% |
| Gemini | 36.2% | 26.2% to 47.5% |
| Claude | 46.2% | 35.0% to 57.5% |
| Perplexity | 78.8% | 68.8% to 87.5% |

An answer without citations counts as citing none. Version 1.0 divided by answers that had citations, which gives Claude 60.7% and Gemini 39.7%. Perplexity’s high answer-level figure mostly reflects how many sources it lists: at the citation level it is below Gemini and Claude.

### Where the gap opens

| Assistant | A | B | C | D |
|---|---|---|---|---|
| ChatGPT | 8.3% | 10.2% | 13.0% | 68.5% |
| Gemini | 16.6% | 13.5% | 20.1% | 49.8% |
| Perplexity | 14.2% | n/a | 19.0% | 66.8% |
| Claude | 25.7% | 10.0% | 27.6% | 36.8% |

A is the top 10 for the question, B the top 10 for the assistant’s own search, C the top 100, and D not surfaced by Google. Route B is exact for Claude and a proxy for ChatGPT and Gemini. For Perplexity (n/a), whose searches are unknown, some of its C and D citations would move to B if they were known. Route C is a lower bound: 2.4% of the Google collections returned partial results, and the median collection reached rank 73.5 rather than 100. Route D is the upper bound of what Google cannot explain for these searches.

Three readings follow from the table. For ChatGPT, over two thirds of citations come from pages Google does not surface for the question or for searches of the kind ChatGPT runs, so they must come from another index, another search or a different retrieval process. Claude is the engine most aligned with Google: it has the highest shares in routes A and C and the fewest citations Google does not surface. And the rewritten searches matter, but they are not the main story for any engine: route B adds 10.0 to 13.5 points.

### The searches matter less than their existence

We checked whether an engine’s citations match its own searches better than they match the searches another engine ran on the same question. For Claude they do not. 24.5% of its citations are in Google’s top 10 for its own search, 25.3% for the ChatGPT stand-in searches and 29.1% for the Gemini ones. Claude ran one search per answer, and other engines’ searches on the same question find its sources as often. For ChatGPT, its stand-in searches match more of its citations (13.8%) than Gemini’s (9.4%) or Claude’s (4.2%).

What this suggests is that rewriting a question into more specific searches reaches different pages from the question as typed, and that the precise wording of the rewrite matters less than the rewriting itself. Since ChatGPT’s searches here come from a separate run, that part is suggestive, not measured.

### Which top-ranking pages get cited

Among the pages in Google’s top 10 for each question, we modeled whether each engine cited them (logistic models with errors clustered by question).

| Assistant | Top-10 pages cited | Ranks 1 to 3 | Ranks 4 to 10 |
|---|---|---|---|
| ChatGPT | 4.7% | 6.7% | 3.6% |
| Gemini | 6.9% | 6.8% | 7.0% |
| Claude | 12.8% | 20.2% | 8.8% |
| Perplexity | 33.0% | 40.8% | 28.8% |

Higher rank predicts citation for ChatGPT (odds ratio per doubling of rank 0.63), Claude (0.65) and Perplexity (0.74), but not for Gemini (0.97, 95% interval 0.69 to 1.36). Adding whether the page is also in the top 10 for the assistant’s own search changes the picture.

- **Claude:** a top-10 page that also ranks for Claude’s own search is more likely to be cited (odds ratio 2.25, 95% interval 1.18 to 4.31), and rank for the question still predicts citation (0.69).
- **ChatGPT, stand-in searches:** ranking for the search predicts citation (odds ratio 5.93, 95% interval 2.01 to 17.48), and rank for the question no longer does (0.93).
- **Gemini:** neither predicts citation (1.22 for the search, 95% interval 0.55 to 2.72).

### Engines compared on the same questions

Differences below are paired by question, so they are not caused by one engine answering easier questions. At the citation level, ChatGPT’s citations ranked less often than Gemini’s (8.3 points fewer, 95% interval 2.9 to 14.0), Perplexity’s (5.9 fewer, 2.8 to 8.8) and Claude’s (17.3 fewer, 10.9 to 24.0). Gemini’s ranked less often than Claude’s (9.1 fewer, 0.4 to 17.5). At the answer level, Perplexity cites a ranked page far more often than every other engine, and the other three are not distinguishable from each other.

### National and local questions

Version 2.0 classifies each question by the place it names, listed in the protocol: 19 name a city, 4 a travel destination, 1 a state and 56 none. Version 1.0 used a text rule that counted three questions naming only “the US” as local.

| Assistant | National | Naming a city | Local minus national, 95% interval |
|---|---|---|---|
| ChatGPT | 8.3% | 11.1% | −6.0 to 7.1 points |
| Perplexity | 14.7% | 15.1% | −6.3 to 3.0 points |
| Gemini | 21.6% | 6.9% | −25.1 to −4.2 points |
| Claude | 20.9% | 35.7% | −1.3 to 25.8 points |

Only Gemini shows a difference whose interval excludes zero: for local questions its citations rarely rank. Claude’s higher figure for city questions is within chance for 24 local questions, and it no longer counts as a finding.

### Google moves too

Between 26 and 28 September, Google’s top 10 for the same 80 questions changed substantially (median overlap 0.58, mean 0.44, measured as shared pages over all pages). Measured against the 28 September results, ChatGPT’s page-level match falls from 8.3% to 5.2%. A comparison with one day’s rankings is a snapshot of a moving target.

### What happened to the Bing comparison

We planned to compare the same citations with Bing’s top 10. For all 80 questions, the Bing results returned by our data provider were unrelated to the question (for example, pages about melamine resin for a question about robo-advisors). A retest with different settings gave the same kind of result. We discarded the Bing data. This says nothing about whether any assistant uses Bing.

## Observed, inferred and unknown

**What we observe.** For 80 buyer questions on one day, most of what the four assistants cited was not in Google’s top 10 for the question. Much of it was not in Google’s top 100 for the question or for searches of the kind the assistants run. The share differs by engine, consistently across the same questions, with ChatGPT furthest from Google and Claude closest.

**What we infer.** The gap between Google and AI citations opens at several stages. Some of it comes from query rewriting, some from citing pages deeper in the results, and for ChatGPT, Perplexity and Gemini most of it from retrieval that Google’s rankings do not describe. Google rankings are therefore a partial stand-in for AI visibility, and a weaker one for some engines than for others.

**What remains unknown.** Which index or search each consumer app actually used. Whether the same questions give the same citations on another run, in other wording or on another day. Whether a cited page shaped the answer or was listed without being used. And whether changing a page, its ranking or its structure would change whether it is cited: nothing was changed in this study.

## What this means

The points below are our interpretation. They follow from the findings but were not tested.

- **Treat Google rankings as one stage, not the outcome.** A page that ranks for the question typed is cited by ChatGPT 4.7% of the time. Ranking helps most where the assistant’s own searches find the same page.
- **Track each engine separately.** The same questions produce different overlaps with Google for each engine, and Gemini behaves differently again for local questions.
- **Think in searches, not questions.** The assistants turn a buyer’s question into narrower searches, and the pages that rank for those searches are the ones more likely to be cited.
- **Being a trusted site can matter more than ranking one page.** For ChatGPT, the site-level overlap (22.9%) was nearly three times the page-level overlap (8.3%): it often cited a different page from a site that ranks.

## Where this sits in GEO research

The first generative engine optimization study, by [Aggarwal and colleagues](https://arxiv.org/abs/2311.09735) (KDD 2024), rewrote pages already placed in the model’s context and measured how visible they became in the answer. That isolates the citation stage: a page that is never retrieved cannot benefit. [SAGEO Arena](https://arxiv.org/abs/2602.12187) (Kim and colleagues, 2026) tested such rewrites in a full retrieval, reranking and generation pipeline and found that they often hurt retrieval and reranking even where they helped generation. A [2026 critical survey of the field](https://arxiv.org/abs/2607.14035) (Martinez) argues that GEO has to be measured stage by stage, with repeated runs and clustered statistics.

This study sits at the boundary between retrieval and citation, on the live consumer systems rather than in a simulated pipeline. It shows that the stages can be told apart in live data: ranking for the question and ranking for the assistant’s own search predict citation separately, and differently for each engine. It does not test any intervention. The question that follows is causal: when the same information is added to a page’s body, its structure or its evidence, does it help retrieval, citation or both, does a gain at one stage cost something at another, and does the answer hold across engines, question types, rewordings and repeated runs? A controlled design that measures each stage separately is needed to answer it. Until then, the figures here describe where the gap is, not how to close it.

## Methodology

- **Questions:** the 80 US buyer questions of our [four-assistant study](https://underneath.agency/research/ai-assistants-brand-agreement-study), ten in each of eight industries; locality classified by the place each question names, listed in the protocol.
- **Citations:** sources cited by ChatGPT and Gemini (consumer apps via DataForSEO LLM Scraper, US location), Perplexity (sonar) and Claude (Haiku 4.5) through DataForSEO’s LLM Responses API with web search, all 26 September 2026; addresses normalized (tracking parameters, fragments, “www.” and trailing slash removed).
- **Google, question as typed:** DataForSEO SERP API, Google.com, United States, English, desktop, top 10 organic results, 26 September 2026.
- **Google, top 100:** same settings, depth 100, 28 September 2026, for the 80 questions and every distinct search the assistants ran on them. One search, a “site:” query, returned no results.
- **Assistants’ searches:** Claude’s recorded in the same answers; ChatGPT’s and Gemini’s from a separate API run on the same questions and day (proxy); none for Perplexity.
- **Statistics:** 95% intervals from a cluster bootstrap that resamples questions (10,000 resamples, seed 23); engine differences paired by question; a difference is reported as a finding only when its interval excludes zero. Candidate models are logistic GEE with robust errors clustered by question.
- **Protocol:** written and frozen on 28 September 2026 before the version 2.0 collection; deviations are listed in the methodology file.
- **Bing:** collected with DataForSEO’s Bing SERP API and discarded after inspection.
- **Update schedule:** quarterly.

## Limitations

- One run per assistant per question, on one day; run-to-run variation is not measured here.
- The searches for ChatGPT and Gemini come from API models, not the consumer apps whose citations are analyzed, and Perplexity’s are unknown.
- Google’s top 100 was collected two days after the answers, and Google’s results moved in that time.
- Google is a stand-in for retrieval; the indexes the assistants actually use are not observed.
- Local subgroups are small (24 local questions).
- Observational: no page, ranking or answer was changed, so none of the findings is causal.

## What changed in version 2.0

Version 1.0 (26 September 2026) compared citations with Google’s top 10 for the question as typed. Following an external methodological review, version 2.0 adds Google’s top 100 for the questions and for the assistants’ own searches, the four-route breakdown, models of which top-ranking pages are cited, answer-level figures over all 80 questions, 95% intervals clustered by question, paired engine comparisons and an explicit locality list. The citation-level figures are unchanged. Claude’s local-question difference (33.0% against 20.9% in version 1.0) is withdrawn, because its interval includes zero.

## Data and downloads

- Every citation with its route, Google ranks and searches: [s23b_citations_pathways.csv](https://underneath.agency/research-data/ai-citations-google-rankings-study/s23b_citations_pathways.csv) and [JSON](https://underneath.agency/research-data/ai-citations-google-rankings-study/s23b_citations_pathways.json)
- Every Google top-10 page and whether each engine cited it: [s23b_candidates.csv](https://underneath.agency/research-data/ai-citations-google-rankings-study/s23b_candidates.csv) and [JSON](https://underneath.agency/research-data/ai-citations-google-rankings-study/s23b_candidates.json)
- Every statistic on this page with its interval: [stats.json](https://underneath.agency/research-data/ai-citations-google-rankings-study/stats.json)
- Machine-readable methodology and deviations: [methodology.json](https://underneath.agency/research-data/ai-citations-google-rankings-study/methodology.json)
- Version 1.0 citations and statistics: [s23_citations_vs_google.csv](https://underneath.agency/research-data/ai-citations-google-rankings-study/s23_citations_vs_google.csv) and [stats_v1.0.json](https://underneath.agency/research-data/ai-citations-google-rankings-study/stats_v1.0.json)

The data is free to reuse with attribution (CC BY 4.0).

To cite: Underneath. (2026). *Do ChatGPT, Gemini, Perplexity and Claude cite pages that rank?* (Version 2.0). Underneath Research. https://underneath.agency/research/ai-citations-google-rankings-study

## Frequently asked questions

### Does ranking on Google help you get cited by ChatGPT?

Less than people assume. Only 8.3% of the pages ChatGPT cited for 80 buyer questions ranked in Google’s top 10 for the same question, and 68.5% were not in Google’s top 100 for the question or for searches of the kind ChatGPT runs. Among top-10 pages, those that also rank for ChatGPT’s own kind of search are the ones more likely to be cited.

### Which AI assistant cites pages that rank on Google most often?

Claude, among the four we tested: 25.7% of its citations ranked in Google’s top 10 for the same question, and 36.8% were not surfaced by Google at all. Google’s own AI Overviews cite top-10 pages 28.7% of the time.

### Why don’t AI assistants cite the pages that rank?

Partly because they search with their own rewritten queries, which find different pages. Those searches account for 10.0 to 13.5 points of citations. The larger part of the gap is pages Google does not surface for these searches at all, which points to retrieval that Google’s rankings do not describe.

### Does ChatGPT use Bing results?

We could not test it: the Bing results we collected were unrelated to the questions and were discarded. This study says nothing either way.

### Is this study proof of what causes AI citations?

No. It is observational: it shows where Google rankings and AI citations diverge, stage by stage, but nothing was changed to test a cause. Whether optimizing a page helps one stage and hurts another needs a controlled experiment.

## Related research

- [Do AI Overviews cite the pages that rank?](https://underneath.agency/research/ai-overview-citations-study)
- [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- [How fresh are the pages AI engines cite?](https://underneath.agency/research/ai-source-freshness-study)
- [Do ChatGPT, Gemini, Perplexity and Claude agree on brands?](https://underneath.agency/research/ai-assistants-brand-agreement-study)

## Related guides

- [Do AI search engines cite the same websites as Google?](https://underneath.agency/resources/do-ai-search-engines-cite-the-same-sites-as-google)
- [Is traditional SEO still important for visibility in AI search?](https://underneath.agency/resources/is-seo-still-important-for-ai-search)
- [Do backlinks still matter for visibility in Perplexity?](https://underneath.agency/resources/do-backlinks-matter-for-perplexity-visibility)
- [Does being cited by ChatGPT boost our Google rankings?](https://underneath.agency/resources/do-chatgpt-citations-boost-google-rankings)
- [How many sources does each AI search engine cite per answer?](https://underneath.agency/resources/how-many-sources-ai-search-engines-cite)

---

This is the Markdown twin of https://underneath.agency/research/ai-citations-google-rankings-study. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Which AI crawlers do top websites block? 10,000 sites, 2026"
description: "15.2% of top sites block GPTBot and 7.4% OpenAI’s search crawler. A same-day check shows ChatGPT cited no page its crawler was barred from."
canonical: "https://underneath.agency/research/ai-crawler-blocking-study"
published: 2026-09-26
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Research · AI crawlers

# Which AI crawlers do top websites block? 10,000 sites, 2026

We fetched and parsed the robots.txt file of every domain in the Tranco top 10,000 on 26 September 2026 and evaluated it the way a compliant crawler would, for 20 AI crawlers and three search-engine controls. The result is a dated, crawler-by-crawler record of what these files ask each AI crawler to do. That is the first stage of a longer pipeline that runs from permission to fetching, retrieval, citation, accurate representation and, finally, a reader who acts. Version 1.2 places the study on that pipeline, adds a same-day cross-check against the pages four AI assistants and Google’s AI Overviews actually cited, and sets out which later stages the next edition will test and how.

## The short version

1. 15.2% of the 5,572 top sites with a readable robots.txt block OpenAI’s GPTBot from the whole site (95% interval 14.3% to 16.2%); 14.1% block Anthropic’s ClaudeBot and 16.2% block Common Crawl’s CCBot, the most-blocked AI crawler.
2. AI search crawlers are blocked about half as often as training crawlers: 7.4% block OAI-SearchBot (ChatGPT search) and 7.1% block Claude-SearchBot.
3. 51.7% of sites that block GPTBot still allow OAI-SearchBot. But 82.7% of those sites never mention OAI-SearchBot: they block GPTBot by name and the search crawler is allowed by default. Only 76 sites, 17.3% of the group, name both crawlers and treat them differently.
4. 5.2% of sites block ChatGPT’s search crawler while allowing Googlebot, and most of them (241 of 291) name OAI-SearchBot to do it.
5. The top 1,000 sites block most: 30.0% block at least one AI crawler, against 19.2% of sites ranked 5,001 to 10,000. Below the top 1,000 the rate is roughly flat.
6. On the same day, ChatGPT cited 123 pages on top sites whose file we could check against the page’s own address, and none of them was closed to OAI-SearchBot. Six were closed to GPTBot, the training crawler, and were cited anyway.
7. Perplexity behaved differently: 209 of the 517 pages it cited on those sites (40.4%) were closed to PerplexityBot by the site’s own file, mostly by a rule that names PerplexityBot. Permission stated in robots.txt is not the same thing as exclusion from every AI answer.

## Where this study sits

A page has to pass several stages before it shapes an AI answer, and “AI visibility” can mean any of them. We keep them apart because a site can pass one stage and fail the next.

| Stage | The question | In this study |
|---|---|---|
| 1. Permission | Does robots.txt let the crawler in? | Measured, 5,572 files |
| 2. Fetch | Does the crawler request the page? | Not measured |
| 3. Retrieval | Is the page pulled in for a question? | Not measured |
| 4. Citation | Is the page shown as a source? | Cross-checked, one day |
| 5. Use | How much of the answer comes from it? | Not measured |
| 6. Fidelity | Is it represented correctly? | Not measured |
| 7. Outcome | Does a reader click or act? | Not measured |

Stage 1 is the only stage a site controls completely, and the only one this study measures in full. The cross-check at stage 4 uses citations collected for our other studies on the same day. Stage 6 is the subject of our [business facts accuracy study](https://underneath.agency/research/ai-business-facts-accuracy-study), and the relation between Google rankings and citations is covered in [AI citations and Google rankings](https://underneath.agency/research/ai-citations-google-rankings-study).

## What we measured

A robots.txt file tells crawlers which parts of a site they may fetch. Each AI company publishes the name (the “product token”) its crawlers obey, and most now run separate crawlers for separate jobs:

- **Training crawlers** collect pages to train models: GPTBot (OpenAI), ClaudeBot (Anthropic), Google-Extended (a control token for Gemini training, not a separate crawler), Applebot-Extended, CCBot (Common Crawl), Bytespider (ByteDance), Meta-ExternalAgent, Amazonbot.
- **AI search crawlers** build the index an assistant searches when it answers with sources: OAI-SearchBot, Claude-SearchBot, PerplexityBot, DuckAssistBot.
- **User-triggered fetchers** load a page when a person asks the assistant to read it: ChatGPT-User, Claude-User, Perplexity-User, MistralAI-User, meta-externalfetcher.
- **Older tokens** that many robots.txt files still name: anthropic-ai and Claude-Web (Anthropic) and cohere-ai (Cohere). They count toward the 20 but are not reported separately below.

For every domain we requested /robots.txt once and, for each token, evaluated whether the site root (/) is allowed, using the rules in RFC 9309: the group that names the token applies, otherwise the wildcard (*) group; the longest matching rule wins; allow wins a tie. A site “blocks” a crawler when the root is disallowed for it. Of the 10,000 domains, 5,572 returned a readable plain-text robots.txt; the rest were infrastructure domains without a website, returned an error, or served an HTML page at the address.

Every percentage on this page uses those 5,572 files as the base unless it says otherwise. So 15.2% means 15.2% of top domains with a readable robots.txt, not 15.2% of the top 10,000. The choice of base changes the size of the numbers but not the pattern:

| Base | Domains | GPTBot | OAI-SearchBot | Any AI crawler |
|---|---|---|---|---|
| Readable robots.txt (this page) | 5,572 | 15.2% | 7.4% | 20.7% |
| Readable file or no file | 7,003 | 12.1% | 5.9% | 16.5% |
| Live homepage, readable file or no file | 5,372 | 12.8% | 6.3% | 17.6% |
| All 10,000 domains (lower bound) | 10,000 | 8.5% | 4.2% | 11.5% |

Each column is the share that blocks that crawler from the whole site. A missing file allows everything, so the second and third rows count those domains as not blocking. The last row counts every unreadable file as not blocking, so it is a floor.

## Findings

### Training crawlers are blocked most

| Crawler | Operator | Purpose | Blocks site | Named |
|---|---|---|---|---|
| CCBot | Common Crawl | Training datasets | 16.2% | 15.2% |
| Bytespider | ByteDance | Training | 15.3% | 13.2% |
| GPTBot | OpenAI | Training | 15.2% | 18.0% |
| ClaudeBot | Anthropic | Training | 14.1% | 15.3% |
| Google-Extended | Google | Gemini training (control token) | 12.3% | 13.7% |
| Meta-ExternalAgent | Meta | Training | 12.3% | 10.7% |
| Applebot-Extended | Apple | Training (control token) | 11.6% | 10.2% |
| Amazonbot | Amazon | Training and Alexa answers | 11.6% | 10.3% |

Share of 5,572 sites with a readable robots.txt. “Blocks site” means the whole site is blocked, including by a wildcard rule for every unnamed bot; “Named” counts files that mention the token at all, whether to block or allow it.

### AI search and user-triggered crawlers are blocked about half as often

| Crawler | Operator | Purpose | Blocks the whole site |
|---|---|---|---|
| PerplexityBot | Perplexity | Search index | 10.9% |
| ChatGPT-User | OpenAI | User-triggered | 9.7% |
| meta-externalfetcher | Meta | User-triggered | 8.3% |
| DuckAssistBot | DuckDuckGo | Search answers | 7.8% |
| Perplexity-User | Perplexity | User-triggered | 7.7% |
| OAI-SearchBot | OpenAI | Search index | 7.4% |
| Claude-User | Anthropic | User-triggered | 7.4% |
| Claude-SearchBot | Anthropic | Search index | 7.1% |
| MistralAI-User | Mistral | User-triggered | 7.1% |

The controls show how unusual AI blocking is: only 2.3% of the same sites block Googlebot, 3.1% block Bingbot and 4.4% block Applebot. GPTBot is blocked 6.6 times as often as Googlebot.

### Half the sites that block training still allow AI search, mostly by default

Of the 849 sites that block GPTBot, 439 (51.7%, 95% interval 48.3% to 55.1%) still allow OAI-SearchBot, the crawler OpenAI uses for ChatGPT search. The same split appears at Anthropic: 49.6% of ClaudeBot blockers leave Claude-SearchBot open. Google-Extended works the same way by design: 10.1% of all sites opt out of Gemini training through Google-Extended while still allowing Googlebot.

Whether that split is a choice is a different question, and the files mostly say no. Of the 439 sites, 363 (82.7%) never mention OAI-SearchBot: they block GPTBot by name, and OAI-SearchBot falls through to rules that allow it. Only 76 (17.3%) name OAI-SearchBot and treat it differently from GPTBot. At Anthropic the share is smaller: 48 of 391 (12.3%) name Claude-SearchBot.

| OpenAI crawlers | Sites | Share |
|---|---|---|
| GPTBot allowed, OAI-SearchBot allowed | 4,718 | 84.7% |
| GPTBot blocked, OAI-SearchBot allowed | 439 | 7.9% |
| of which OAI-SearchBot is named in the file | 76 | |
| GPTBot blocked, OAI-SearchBot blocked | 410 | 7.4% |
| GPTBot allowed, OAI-SearchBot blocked | 5 | 0.1% |

So the pattern is consistent with sites separating training from AI search, but for most of them the separation follows from naming only the training crawler. Robots.txt cannot tell us whether the owners meant it.

### A smaller group asks AI search crawlers to stay away but lets Google in

5.2% of sites (95% interval 4.7% to 5.8%) block OAI-SearchBot while allowing Googlebot, and 5.7% block all three AI search crawlers we tested (OAI-SearchBot, Claude-SearchBot and PerplexityBot). This group does look deliberate: 241 of the 291 name OAI-SearchBot in the file. OpenAI’s documentation says sites that opt out of OAI-SearchBot “will not be shown in ChatGPT search answers, though can still appear as navigational links.” We did not test whether these sites appear in ChatGPT or Claude answers.

### Blocks that come from a wildcard rule

A crawler that no group names follows the wildcard (*) group. A file written years ago to keep all bots out therefore also blocks crawlers that did not exist when it was written. The rule source differs by crawler type:

| Crawler | Blocked | Named in the rule | Through the wildcard group |
|---|---|---|---|
| GPTBot | 849 | 676 | 173 (20.4%) |
| ClaudeBot | 788 | 601 | 187 (23.7%) |
| PerplexityBot | 607 | 429 | 178 (29.3%) |
| OAI-SearchBot | 415 | 242 | 173 (41.7%) |
| Claude-SearchBot | 398 | 208 | 190 (47.7%) |

Across all AI search crawlers, 39.8% of blocks come through the wildcard group, against 24.9% for training crawlers. 214 files (3.8%) disallow the root for any unnamed bot; 155 of them (72.4%) name no AI crawler at all, and 32 let OAI-SearchBot in by name.

### Most files restrict paths, not whole sites

Blocking the root is the strongest setting but not the most common. For OAI-SearchBot, 7.4% of files block the whole site, 3.3% have path rules written for it by name, 68.2% apply general path rules from the wildcard group (for example /admin/ or /search), and 21.0% have no disallow rule that applies to it. Googlebot has a similar profile, apart from the full blocks.

To see whether AI crawlers get less of a site than Google, we tested every path each file itself lists. Among 4,009 sites whose root is open to both OAI-SearchBot and Googlebot, 5.7% close at least one of those paths to OAI-SearchBot but not to Googlebot; for GPTBot it is 6.0% of 3,620 sites. We cannot tell which of a site’s pages matter for AI answers, so this is a count of differences, not of their importance.

### The top 1,000 sites block most

| Tranco rank | Readable files | Any AI crawler | 95% interval | GPTBot |
|---|---|---|---|
| 1 to 1,000 | 507 | 30.0% | 26.2% to 34.1% | 20.7% |
| 1,001 to 5,000 | 2,196 | 20.5% | 18.9% to 22.3% | 15.4% |
| 5,001 to 10,000 | 2,869 | 19.2% | 17.8% to 20.7% | 14.2% |

Share of each tier’s readable files that block at least one AI crawler, or GPTBot, from the whole site.

The top 1,000 block at least one AI crawler 10.8 percentage points more often than sites ranked 5,001 to 10,000 (95% interval 6.7 to 15.1), and a trend test across rank bands is clear (p < 0.001). The gap survives two checks: among domains with a live homepage it is 27.2% against 18.3%, and without the files that block every unnamed bot it is 24.8% against 16.6%. But the gradient is not smooth. Split into bands of 1,000 ranks, every band from 1,001 to 9,000 falls between 19.1% and 22.9%; the difference sits in the top 1,000 (and a lower last band, 13.1%). Popular domains differ from the rest in type, size and ownership, so this is an association with rank, not an effect of popularity.

Overall, 20.7% of sites block at least one of the 20 AI crawlers from the whole site, and 24.2% mention at least one AI crawler by name.

### Content Signals are still rare

Cloudflare’s Content Signals proposal adds a line such as “Content-Signal: search=yes, ai-train=no” to robots.txt. 159 sites (2.9%) carry one. Among them, 99.4% say search=yes, 64.2% say ai-train=no and 73.6% allow ai-input (use of content to ground AI answers). The pattern matches the crawler data: yes to being found, often no to training.

### From permission to citation: a same-day cross-check

To see whether stage 1 shows up at stage 4, we matched every source cited in two of our other datasets, both collected on 26 September 2026, against these robots.txt files: the answers ChatGPT, Claude, Gemini and Perplexity gave to 80 buyer questions, and the sources in 481 US AI Overviews. A cited page counts when its domain is in the top 10,000 with a readable file. The engine operator’s own pages (google.com for Google’s answers) are left out. For pages on the domain itself or its www host, we tested the page’s own path and query against the rules for the engine’s search crawler.

| Engine | Search crawler | Cited pages checked | Closed to that crawler |
|---|---|---|---|
| ChatGPT | OAI-SearchBot | 123 | 0 (0.0%) |
| Claude | Claude-SearchBot | 50 | 7 (14.0%) |
| Perplexity | PerplexityBot | 517 | 209 (40.4%) |
| Gemini | Googlebot | 41 | 2 (4.9%) |
| AI Overviews | Googlebot | 979 | 64 (6.5%) |

Pages on subdomains with their own robots.txt are not in the “checked” column. Intervals for the shares are in stats.json.

- **ChatGPT matches its published rules.** None of the 123 checked pages was closed to OAI-SearchBot (95% interval 0.0% to 3.0%), while 6 were closed to GPTBot. Training opt-outs did not keep pages out of ChatGPT’s answers; search opt-outs, in this sample, did.
- **Perplexity cites pages its crawler is told to avoid.** The 209 pages come from 37 domains, led by forbes.com (61 pages), nytimes.com (15) and reddit.com (13). In 170 of them the file names PerplexityBot. 194 of the 209 were open to Googlebot and 74 to Perplexity-User. We cannot tell from answers alone whether Perplexity fetched these pages, took them from another search index or relied on text it already held; the files only show that its declared crawler was asked to stay out.
- **Claude’s seven are one site.** All seven are Yelp search pages whose file names Claude-SearchBot. Claude also cited sites that block ClaudeBot, the training crawler, far more often than their share of top sites would predict (52.7% of its cited pages on top-10,000 domains with a readable file, against 17.8% expected).
- **Google’s exceptions are Reddit.** All 64 AI Overview citations closed to Googlebot are reddit.com pages; Reddit’s file disallows every crawler, and Google has a data licensing agreement with Reddit announced in 2024. Google-Extended, Google’s training opt-out, made no visible difference: 17.8% of cited pages on top-10,000 domains with a readable file came from sites that use it, against 16.8% expected from their ranks.

Expected rates reweight the share of blocking sites to the rank mix of the cited domains, because AI answers cite popular sites more than the top 10,000 as a whole. This is one day, 80 questions and a few hundred pages per assistant, with citations clustered on a small number of domains, so the results describe that sample, not each engine in general.

## How this compares with other studies

Published figures differ mainly because they measure different things. We report them with their definitions rather than as competing answers.

| Source | Sample and date | What was measured | Figure |
|---|---|---|---|
| This study | Tranco top 10,000, 5,572 readable files, September 2026 | GPTBot disallowed at the site root | 15.2% |
| This study | Same | GPTBot named anywhere in robots.txt | 18.0% |
| HTTP Archive, Web Almanac 2025 | Whole web, July 2025 crawl | GPTBot appears in robots.txt | 4.5% of desktop sites |
| Cloudflare | Cloudflare’s top 10,000 domains, 3,816 with robots.txt, June 2025 | Domains disallowing GPTBot | 312 (250 fully, 62 partially) |
| Cloudflare | Cloudflare customers, September 2026 | Dashboard settings, not robots.txt | 17% block training in some way; less than 1% block search bots |

- **Top sites block far more than the web as a whole.** The Web Almanac counts mentions across millions of mostly small sites; our top-10,000 mention rate is four times higher.
- **Training versus search is the same split Cloudflare sees.** Its customer settings show training blocked far more often than search; our robots.txt data shows the same direction (15.2% against 7.4% for OpenAI’s two crawlers).
- **Content Signals:** the only published figure is Cloudflare’s own deployment of the line through its managed robots.txt on over 3.8 million domains, which is not voluntary adoption. Our 2.9% of top sites is, as far as we can find, the first independent count. Google’s John Mueller has said the directive has “no effects whatsoever for any crawler or LLM”.

Sources: [HTTP Archive Web Almanac 2025, SEO chapter](https://almanac.httparchive.org/en/2025/seo); [Cloudflare, June 2025](https://blog.cloudflare.com/from-googlebot-to-gptbot-whos-crawling-your-site-in-2025/); [Cloudflare, September 2026](https://blog.cloudflare.com/accountable-mixed-use-ai-crawlers/); [Cloudflare, Content Signals Policy](https://blog.cloudflare.com/content-signals-policy/); [Search Engine Roundtable](https://www.seroundtable.com/google-cloudflare-content-signals-41631.html).

## What this means

The following is our interpretation of the numbers, not part of the measurement.

- **Blocking training is not the same as leaving AI search.** Each AI company runs its training and search crawlers under separate names, so a site can refuse one and allow the other. Most sites in that position got there by naming only the training crawler. A brand that wants to be eligible for AI search should say so explicitly, with a rule for each search and user crawler, rather than rely on defaults that can change.
- **Check for accidental blocks.** A wildcard rule written years ago for scrapers blocks every new AI crawler too: 3.8% of sites disallow the root for any unnamed bot, and almost three quarters of those files name no AI crawler at all. About two in five blocks on AI search crawlers come through the wildcard group rather than a rule that names them.
- **Robots.txt is only half the story.** Many sites block AI crawlers at the firewall or CDN instead, which this study cannot see. A brand should test what the crawlers actually receive, not only what the file says.

## What this establishes, and what it does not

Each result below supports a narrower claim than the headline might suggest. The last column is the test that would close the gap.

| Result | What it shows | Limit | Next test |
|---|---|---|---|
| 15.2% block GPTBot | What files ask | Not what crawlers do | Server logs |
| Training blocked more than search | Stated preference | Mostly by default | Panel of file changes |
| Wildcard blocks | Old rules reach new bots | Intent unknown | Owner survey |
| Top 1,000 block most | Association with rank | No controls | Matched comparison |
| ChatGPT cites no barred page | Rules and citations agree | One day, 123 pages | Repeated dates |
| Perplexity cites barred pages | Rules do not bind citations | Route unknown | Logs and fetch tests |

Allowing a search crawler does not mean a site is fetched, retrieved or cited, and blocking it does not rule out a citation that arrives through another index or a user-triggered fetch. The files also do not say why sites configure them as they do. Sites that allow AI crawlers differ from sites that block them in size, type and ownership, so any comparison of their AI visibility needs controls for those differences.

### Hypotheses for the next edition

These are hypotheses, not findings. The cross-check above bears on the first two; the rest are untested.

- **H1.** For engines that build answers from their own crawler’s index, permission is close to necessary for citation. For engines that also draw on other indexes, it is not.
- **H2.** Training opt-outs (GPTBot, ClaudeBot, Google-Extended) do not change whether a site is cited.
- **H3.** Among sites that allow the search crawler, permission explains little of the variation in citation; relevance and ranking explain more.
- **H4.** When a site opens a previously blocked search crawler, citations follow only after a delay, and the delay differs by engine.
- **H5.** Citation and accurate representation can diverge: a page can be cited for claims it does not make.

### How the next edition will test them

The robots.txt scan will repeat each quarter on the same 10,000 domains, which turns it into a panel. Sites that change their rules between editions are a natural experiment. We will compare their citations before and after the change with matched sites of similar rank and type that did not change, across the four assistants, a fixed set of questions per industry, repeated runs and at least two dates per edition. A change to an unrelated path serves as a placebo. For the fetch stage, we will use server logs from sites that share them, starting with our own, and test each cited page’s path, not only the root.

Answers vary from run to run, so results will be reported as distributions with intervals, from a model with separate terms for the engine, the question, the domain and the date, rather than as single averages. Cited claims will be checked against the page they cite, so that a rise in citations is not read as a gain when the page is misquoted.

## Methodology

- **Sample:** registrable domains ranked 1 to 10,000 in the Tranco list L5PZ4 (downloaded 26 September 2026), a research ranking that averages several traffic sources.
- **Collection:** one HTTPS request for /robots.txt per domain on 26 September 2026 (04:16 to 04:38 UTC), following redirects, with a user agent identifying Underneath’s research crawler.
- **Inclusion:** 5,572 domains that returned HTTP 200 with a plain-text body. Excluded: 2,467 unreachable domains (mostly infrastructure with no website), 1,431 that returned 4xx (under RFC 9309 a missing file means everything is allowed), 485 that served HTML, 45 server errors.
- **Evaluation:** an RFC 9309 parser (named group, else wildcard group; longest match; allow wins ties), tested against known cases. “Blocked” means the site root is disallowed. A block is “named” when a user-agent line names the token and comes “through the wildcard group” otherwise.
- **Uncertainty:** 95% Wilson intervals for shares, Newcombe intervals for differences between rank tiers, and a Cochran-Armitage test for trend across rank bands, repeated on domains with a live homepage and without files that block every unnamed bot.
- **Crawler purposes:** training, search and user-triggered labels are the operators’ own published descriptions as of 26 September 2026. methodology.json records each token with the page its label came from. Operators rename crawlers and change their jobs; each edition re-checks these pages and reports changed labels as changes.
- **Citation cross-check (version 1.2):** sources cited by ChatGPT, Claude, Gemini and Perplexity for 80 buyer questions and by 481 US AI Overviews, all collected on 26 September 2026 for our other studies. Each cited host is matched to a top-10,000 domain by suffix. Page-level permission is tested with the same parser on the page’s path and query, for pages on the domain or its www host only. Expected rates reweight the population share of blocking sites to the cited domains’ mix of 1,000-rank bands; exact binomial tests are in stats.json.
- **Version 1.2** (28 September 2026) adds the pipeline framing, the citation cross-check, the table of what each result establishes, and the hypotheses and design for the next edition. No robots.txt figure changed.
- **Version 1.1** (28 September 2026) reanalyzes the same fetch after an external methodology review. It adds the intervals, the alternative bases, the rule-source and path-level counts, and the dated crawler list, and revises the reading of the training-versus-search split. No 1.0 figure changed.
- **Update schedule:** quarterly; the next edition will report change against this baseline.

## Limitations

- Tranco ranks domains, not businesses; the top 10,000 includes content delivery and API domains.
- Robots.txt is a request, not enforcement. Crawlers that ignore it, and sites that block at the network level, are not measured.
- Permission is not crawling, indexing or retrieval; none of those was measured. Citation was cross-checked on one day only.
- The cross-check covers a few hundred cited pages per assistant, concentrated on a few domains, and only pages on top-10,000 domains with a readable file. Citations show that a page was named, not how the engine obtained it.
- The files show what sites ask, not why. Rank differences are associations, with no controls for site type or size.
- “Blocked” is measured at the site root. Path rules are summarized, but we do not know which pages matter for AI answers.
- One fetch per domain on one day. A few servers answer an unfamiliar user agent differently from a browser.
- Results describe popular global sites, not small business websites.

## Data and downloads

- Per-domain results for every crawler: [s1_robots_by_domain.csv](https://underneath.agency/research-data/ai-crawler-blocking-study/s1_robots_by_domain.csv) and [JSON](https://underneath.agency/research-data/ai-crawler-blocking-study/s1_robots_by_domain.json)
- Version 1.1 per-domain file with rule source (named or wildcard) and path restrictions for every crawler: [s1_robots_by_domain_v11.csv](https://underneath.agency/research-data/ai-crawler-blocking-study/s1_robots_by_domain_v11.csv)
- Version 1.2 citation cross-check, one row per cited page with its crawler permission: [s1_citation_linkage_v12.csv](https://underneath.agency/research-data/ai-crawler-blocking-study/s1_citation_linkage_v12.csv) and the summary [linkage_v12.json](https://underneath.agency/research-data/ai-crawler-blocking-study/linkage_v12.json)
- Every statistic on this page: [stats.json](https://underneath.agency/research-data/ai-crawler-blocking-study/stats.json)
- Machine-readable methodology: [methodology.json](https://underneath.agency/research-data/ai-crawler-blocking-study/methodology.json)

The data is free to reuse with attribution (CC BY 4.0).

To cite: Underneath. (2026). *Which AI crawlers do top websites block? 10,000 sites, 2026*. Underneath Research. https://underneath.agency/research/ai-crawler-blocking-study

## Frequently asked questions

### What percentage of websites block GPTBot?

In our September 2026 scan of the Tranco top 10,000, 15.2% of the 5,572 sites with a readable robots.txt block GPTBot from the whole site. Among the top 1,000 sites the figure is 20.7%.

### Does blocking GPTBot keep a site out of ChatGPT search?

Not by itself. GPTBot collects training data. ChatGPT search relies on OAI-SearchBot, and page visits requested by a user come from ChatGPT-User. 51.7% of the sites that block GPTBot still allow OAI-SearchBot, most of them because the file does not mention it. Allowing the crawler makes a site eligible; it does not mean ChatGPT will find or cite it, which this study did not measure.

### Can a site be cited by AI if its robots.txt blocks the crawler?

Yes, depending on the engine. In our same-day check, ChatGPT cited none of 123 pages that were closed to OAI-SearchBot, but 40.4% of the pages Perplexity cited on the sites we could check were closed to PerplexityBot, and Google’s AI Overviews cited Reddit pages that Reddit’s file closes to all crawlers. Robots.txt is a request; it is not a guarantee of exclusion.

### Which AI crawler is blocked most often?

Common Crawl’s CCBot, blocked by 16.2% of sites, followed by ByteDance’s Bytespider (15.3%) and OpenAI’s GPTBot (15.2%).

### What is the difference between Google-Extended and Googlebot?

Googlebot crawls for Google Search, including AI Overviews. Google-Extended is a separate token that only controls whether content is used to train and ground Gemini models. 10.1% of sites opt out through Google-Extended while still allowing Googlebot.

### What is a Content-Signal line in robots.txt?

A proposal from Cloudflare that lets a site state, in robots.txt, whether its content may be used for search, as AI input, or for AI training. 2.9% of the sites we checked have one.

## Related research

- [How many websites have an llms.txt file?](https://underneath.agency/research/llms-txt-adoption-study)
- [Do websites serve Markdown to AI agents?](https://underneath.agency/research/agent-readable-web-study)
- [What pages cited by AI Overviews have in common](https://underneath.agency/research/ai-overview-cited-pages-study)
- [How often do AI Overviews appear?](https://underneath.agency/research/ai-overviews-frequency-study)

## Related guides

- [Does blocking Google-Extended in robots.txt reduce our visibility in AI Overviews?](https://underneath.agency/resources/does-blocking-google-extended-hurt-ai-overviews)
- [Can our server logs show what content AI bots are looking for on our site?](https://underneath.agency/resources/ai-bot-server-logs-content-demand)
- [If AI agents can’t read my website, where does their answer about us come from?](https://underneath.agency/resources/when-ai-agents-cant-read-your-site)
- [Do AI agents recommend businesses whose websites they can read more often?](https://underneath.agency/resources/do-ai-agents-recommend-readable-websites)
- [Do off-site mentions help once an AI agent is reading about you?](https://underneath.agency/resources/off-site-mentions-vs-site-readability-ai-agents)

---

This is the Markdown twin of https://underneath.agency/research/ai-crawler-blocking-study. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "The hidden searches AI assistants run before they answer"
description: "ChatGPT ran 3.7 searches per buyer question before answering. When a search named a source such as NerdWallet or Avvo, the answer cited it 44.0% of the time."
canonical: "https://underneath.agency/research/ai-hidden-searches-study"
published: 2026-09-26
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Research · AI assistants

# The hidden searches AI assistants run before they answer

When an AI assistant answers with web search, it does not search for the question as the user typed it. It writes its own searches first, often several, and answers from what those return. These “fan-out” searches are the first step of a longer process: searches, then results, then the sources the assistant selects, then the answer, then the citations. We collected the searches for the 80 buyer questions in our [four-assistant study](https://underneath.agency/research/ai-assistants-brand-agreement-study) from ChatGPT, Gemini and Claude on 26 September 2026. This version of the study also asks whether those searches are linked to what the answer finally cites.

They are. The assistants write their searches in different ways. When a search names a particular source, that source is cited far more often than when it is not named. But named sources explain only a small part of what gets cited. Most cited pages come from ordinary searches that name no source.

## The short version

1. ChatGPT ran a mean of 3.7 searches per answer (95% interval 3.1 to 4.2), Gemini 1.9 and Claude 0.76. None of the 509 searches repeated the user’s question word for word. Every answer that ran a search cited at least one page, and none of the 42 answers without a search cited anything.
2. The assistants search for different things. In 43.8% of its answers ChatGPT ran a search aimed at a named publication, ranking or award, against 26.2% for Gemini and 1.2% for Claude. ChatGPT also looked for reviews in 46.2% of answers and prices in 23.8%. Claude’s searches were almost all plain searches for options.
3. When a search named a source, such as NerdWallet, Avvo or Super Lawyers, the answer cited that source 44.0% of the time (40 of 91). In answers from the same assistant and industry that did not name it, the figure was 8.1%. That is an association, not proof that naming causes citing.
4. Only 4.2% of cited domains (40 of 943) had been named in a search of the same answer. Gemini cited Reddit in 31 answers but mentioned Reddit in a search in only 5.
5. More searches go with more cited domains: with twice as many searches, an answer cited 1.45 times as many domains, after allowing for assistant and industry. At the same number of searches, Gemini cited 3.29 times as many domains as ChatGPT.
6. Version 1.0 reported that 10.5% of ChatGPT’s searches named a past year. Most of those looked for the latest edition of an annual ranking, such as J.D. Power 2025. Only 3 of ChatGPT’s 296 searches used an old year for no apparent reason, and Gemini’s past years always came with 2026.

## Where the hidden searches sit

Getting cited by an AI assistant takes several steps, and a page can drop out at each one.

1. The user asks a question.
2. The assistant writes its own searches.
3. The search engine behind it returns results.
4. The assistant selects some of those pages.
5. It writes the answer, using some of what the pages say.
6. It attaches citations to the answer.

A page is only credited if it is found, then used, then cited. Improving one step does not guarantee the others. A page can be found but ignored, used without being cited, or cited for a claim it does not support. It can also be cited today and gone tomorrow.

This study observes step 2 directly and step 6 at the level of the whole answer. The assistants’ APIs report the searches they ran and the pages they cited, but not the results each search returned, so steps 3 to 5 are not observed. Every link reported below connects a search to the citations of the same answer, not to the results of that particular search.

## Research questions

- **RQ1.** How many searches does each assistant run, and how far do they move from the question?
- **RQ2.** What do the searches try to find: options, current information, named authorities, reviews, prices, comparisons, particular platforms?
- **RQ3.** Do answers that run more searches cite a wider range of websites, for the same assistant and industry?
- **RQ4.** When a search names a source, is that source cited in the same answer more often than when it is not named?
- **RQ5.** Are the searches, and their link to citations, stable across repeated runs, dates and wordings? This is not answered here: we have one run per question and assistant.

## What we analyzed

For each of the 80 buyer questions (10 in each of eight industries, such as “What is the best CRM for a small business?” and “Who is the best divorce lawyer in Houston?”), we asked ChatGPT (GPT-5.4 nano) and Gemini (Gemini 3.5 Flash-Lite) through their APIs with web search on. We recorded the searches each one reported running and the pages it cited. Claude’s searches and citations (Claude Haiku 4.5, also with web search) come from the answers collected for the four-assistant study the same day. That gives 240 answers and 509 searches. Version 1.1 re-analyzes the same answers; no assistant was asked anything new.

Each search was analyzed in three ways. Fixed phrase rules detect a year, review words, a site: operator and similar patterns, as in version 1.0. A model coder labeled what each search is for and the role of any year in it, without being told which assistant ran it. And a sentence-embedding model measured how close each search is in meaning to the original question.

## Finding 1: how many searches, and how far from the question

| Measure | ChatGPT | Gemini | Claude |
|---|---|---|---|
| Answers with a search | 73.8% | 97.5% | 76.2% |
| Mean searches per answer | 3.7 | 1.9 | 0.76 |
| 95% interval | 3.1 to 4.2 | 1.8 to 2.0 | 0.7 to 0.8 |
| Total searches | 296 | 152 | 61 |
| Median words per search | 8 | 6 | 6 |
| Median similarity to question | 0.73 | 0.85 | 0.9 |
| Search functions per answer | 5.76 | 3.35 | 2.36 |

Similarity runs from 0 to 1; 1 means the search says the same thing as the question. ChatGPT moved furthest from the question: it added a median of 5 words that were not in it, Gemini 2 and Claude 1. Claude never ran more than one search per answer. The last row counts the distinct functions (see the next section) covered by an answer’s searches.

Intervals come from resampling the 80 questions, so answers to the same question are not treated as independent.

## Finding 2: what the searches are for

A model coder read every search next to its question and recorded each function the search performs. The coder was not told which assistant ran it.

| Share of answers with a search that… | ChatGPT | Gemini | Claude |
|---|---|---|---|
| Looks for options | 72.5% | 97.5% | 76.2% |
| Asks for current information | 52.5% | 76.2% | 43.8% |
| Looks for a specific feature | 47.5% | 30.0% | 20.0% |
| Looks for reviews or ratings | 46.2% | 30.0% | 5.0% |
| Targets a named authority | 43.8% | 26.2% | 1.2% |
| Restricts to a place | 40.0% | 50.0% | 30.0% |
| Compares options | 33.8% | 0.0% | 0.0% |
| Targets a platform or directory | 31.2% | 10.0% | 0.0% |
| Looks for prices | 23.8% | 6.2% | 3.8% |
| Targets an option’s own site | 23.8% | 0.0% | 0.0% |
| Looks for complaints | 6.2% | 0.0% | 0.0% |

“Named authority” means a publication, ranking or award, such as U.S. News, Wirecutter, J.D. Power or Super Lawyers. “Platform or directory” means a site such as Avvo, Yelp, G2 or Reddit, often searched with a site: operator. A search can do several things at once, so the rows do not add up to 100%.

The phrase rules of version 1.0 give the same picture. With 95% intervals: 37.5% of ChatGPT’s answers had a search naming a publication or directory (27.5% to 47.5%), 41.2% added review words (30.0% to 52.5%) and 10.0% used a site: operator (3.8% to 17.5%).

ChatGPT’s searches split one question into several jobs. Its most common primary purpose was finding options (32.4% of its searches), followed by reaching a named authority (23.0%). For Gemini, 86.8% of searches were primarily about finding options, and for Claude 98.4%.

### What a real set of searches looks like

For “Who is the best divorce lawyer in Houston?”, ChatGPT ran four searches:

- best divorce lawyer in Houston TX AVVO top rated family law attorney
- Houston divorce attorney Martindale-Hubbell top rated family law
- Houston family law divorce attorney U.S. News Best Lawyers Houston divorce
- Super Lawyers Houston family law divorce 2025 2026

Gemini ran two (“best divorce lawyers in Houston Texas” and “top rated family law attorneys Houston TX”). For a question about San Francisco employment law firms, ChatGPT searched inside single directories: “site:lawyers.com employment law San Francisco” and “site:avvo.com employment lawyer San Francisco review”.

## Finding 3: years in the searches

Assistants write the year into their searches: 52.5% of ChatGPT’s answers, 76.2% of Gemini’s and 43.8% of Claude’s included a search with a year. Version 1.0 counted any year before 2026 as out of date. The model coder instead recorded why the year was there.

| Year in the search | ChatGPT | Gemini | Claude |
|---|---|---|---|
| No year | 170 | 81 | 26 |
| 2026 only | 95 | 46 | 35 |
| 2026 and an earlier year | 14 | 25 | 0 |
| Earlier edition of a ranking | 14 | 0 | 0 |
| Earlier year, no reason | 3 | 0 | 0 |

Counts are searches. An “earlier edition” search asks for an annual ranking whose latest edition may carry the earlier year, such as “J.D. Power homeowners insurance customer satisfaction Florida 2025” or “Gartner Magic Quadrant 2024 managed detection and response”. Gemini’s past years always came alongside 2026 (“best boutique hotels Nashville 2025 2026”). Only 3 searches, all ChatGPT’s, used an earlier year with no apparent reason.

## Finding 4: from searches to citations

### A named source is often cited, but most citations are not named

We matched source names in the searches (Avvo, NerdWallet, U.S. News, Reddit and 28 others, listed in the dataset) to their websites. Then we checked whether the answer cited that website.

| Source named in a search | ChatGPT | Gemini | Both |
|---|---|---|---|
| Times named | 68 | 23 | 91 |
| Cited in the same answer | 45.6% | 39.1% | 44.0% |
| Cited when not named | 5.1% | 11.4% | 8.1% |

“Cited when not named” uses answers from the same assistant and the same industry, for the same source, whose searches did not name it. Claude never named a source. In a model that accounts for answers to the same question, the odds of a source being cited were 10.0 times higher when a search named it (95% interval 5.92 to 16.91).

Sources differed. Reddit was cited all 5 times it was named, and NerdWallet 5 times out of 9. U.S. News was named 7 times and never cited, Yelp 6 times and never cited. Of 12 searches inside one site with a site: operator, 5 led to that site being cited.

The reverse view is smaller. Of the 943 cited domains, counted once per answer, 40 (4.2%) had been named in one of that answer’s searches. Gemini cited Reddit in 31 of its 80 answers and mentioned Reddit in a search in 5 of them. Most cited pages arrive through ordinary searches for options, not searches aimed at them.

A broader search purpose did not show the same link. Answers with a search targeting a named authority or platform cited at least one of the listed sources in 55.3% of cases (76 answers), and answers without one in 54.1% (122 answers). The odds ratio was 0.86 (0.39 to 1.9). The link is with the specific source named, not with the general habit of searching for authorities.

### More searches, more cited websites

Within each assistant, answers with more searches cited more distinct domains: Spearman correlation 0.35 for ChatGPT and 0.3 for Gemini. Claude ran at most one search, so it cannot show this. In a model that accounts for answers to the same question and for industry, twice as many searches went with 1.45 times as many cited domains. At the same number of searches, Gemini cited 3.29 times as many domains as ChatGPT (2.41 to 4.49) and Claude 2.98 times as many (1.8 to 4.96).

So the number of searches is not the whole story. ChatGPT searched the most and cited the fewest websites: 2.38 domains per answer, against 6.28 for Gemini and 3.14 for Claude. How many results each assistant keeps from a search matters as much as how many searches it runs.

## Three ways of searching

| Measure | ChatGPT | Gemini | Claude |
|---|---|---|---|
| Mean searches | 3.7 | 1.9 | 0.76 |
| Similarity to question | 0.73 | 0.85 | 0.9 |
| Targets a named authority | 43.8% | 26.2% | 1.2% |
| Targets a platform | 31.2% | 10.0% | 0.0% |
| Looks for reviews | 46.2% | 30.0% | 5.0% |
| Cited domains per answer | 2.38 | 6.28 | 3.14 |

Shares are of answers; similarity is the median per search.

The three assistants follow three different patterns. ChatGPT runs many narrow searches, aimed at named authorities, directories, reviews and prices, and cites few websites. Gemini runs one or two broad searches, almost always with the year, and cites many websites. Claude runs a single search close to the question and cites a moderate number. The patterns describe these API models on one day, not a fixed trait of each company’s product.

## Observed, inferred and unknown

- **Observed.** The searches each API reported, the pages each answer cited, and how the two line up within an answer.
- **Inferred.** Naming a source in a search is associated with citing it. The link is strong and consistent across both assistants that name sources, but engines, questions and searches were not randomized. Questions that lead an assistant to name Avvo may also be questions where Avvo would be cited anyway.
- **Unknown.** Which results each search returned, and therefore whether a page was found and passed over, or never found at all. Whether the answer’s claims match the cited pages. Whether the same question gives the same searches tomorrow. Whether being cited leads to clicks or customers.

A source can be found but unused, used but uncited, cited but unsupported, visible one day and gone the next, or visible without anyone acting on it. These are different outcomes. This study measures the first step and the last, and only for one run.

## What this means

This section is our interpretation.

- **Being in the sources an assistant names is a direct route.** When ChatGPT names Avvo, Super Lawyers or NerdWallet in its own searches, those sites are cited far more often. A business listed there, with a good profile, sits on that route. The route is narrow: most citations come from searches that name no source.
- **Different assistants reward different presence.** ChatGPT’s searches reach for rankings, directories and reviews. Gemini’s broad searches bring back many pages, often from Reddit and YouTube. Claude’s single search returns what a plain search for options returns. One tactic does not cover all three.
- **Freshness is written into the query, often through rankings.** Many searches ask for the current year. ChatGPT often asks for the latest edition of an annual list, which favors pages that are clearly dated and kept current.

## Where this sits in GEO research

Early work on generative engine optimization showed that changing a page can change how often an AI answer uses it, once the page is in front of the model. Later work compared which sources different AI search systems cite. More recent reviews argue that “visibility” mixes several outcomes: being found, being used, being credited, staying visible, and producing value. They also note that much of the evidence covers only the step after a page is found.

This study looks at the step before: the searches that decide what can be found. It links those searches to the final citations, at the level of the answer. The next questions need other designs. The first is to record which results each search returned, to separate “never found” from “found and passed over”. The second is to repeat the same questions over several days to test stability. The third is to check claims against cited pages. The fourth is to add or remove a source under controlled conditions, to test cause rather than association.

## Methodology

- **Questions:** the 80 US buyer questions of our [four-assistant study](https://underneath.agency/research/ai-assistants-brand-agreement-study), 10 per industry.
- **Engines:** ChatGPT (GPT-5.4 nano) and Gemini (Gemini 3.5 Flash-Lite) through DataForSEO’s LLM Responses API with web search, 26 September 2026. Claude (Haiku 4.5) through the same API, collected for the four-assistant study the same day. The searches are those each API reported (“fan-out queries”); cited pages are the answer’s annotations, reduced to registrable domains and counted once per answer.
- **Phrase rules:** fixed rules applied to every search (listed in the analysis code), unchanged from version 1.0.
- **Function coding:** Claude Opus coded every search, with the question shown, the assistant hidden and the order shuffled. It recorded every function from a list of 12, the primary function, and the role of any year. Claude Sonnet coded the same searches independently. They agreed on the primary function for 78.6% of searches (kappa 0.69) and on the year role for 99.4% (kappa 0.99). The median kappa across the 12 functions was 0.9; the weakest were “pins down which company” (0.11) and “looks for a specific feature” (0.66), so read the feature row in Finding 2 as indicative.
- **Similarity:** cosine similarity between each search and its question, using the all-mpnet-base-v2 sentence-embedding model.
- **Named sources:** a fixed list of 32 source websites, matched from source names in the searches, set before the citations were examined.
- **Models:** GEE regression with answers grouped by question: logistic for “named source cited”, with the assistant as a fixed effect; Poisson for distinct cited domains, with assistant, industry and the log of the number of searches.
- **Intervals:** 95% bootstrap resampling questions (2,000 resamples).
- **Update schedule:** quarterly.

## Limitations

- One answer per question and assistant, on one date. Whether the searches, and their link to citations, hold across repeated runs, days and rewordings is not measured.
- The APIs report the searches but not the results of each search. We cannot tell a page that was never found from one that was found and not used.
- The link between searches and citations is an association. Questions, assistants and searches were not assigned at random.
- These are API models, not the consumer apps; the consumer ChatGPT and Gemini may search differently.
- The function and year labels were coded by two AI models from the same family. Their agreement checks consistency, not accuracy. No person coded the searches.
- The named-source list is fixed, so searches naming other sources are not matched. Whether the answers’ claims are supported by the cited pages was not checked.

## What changed in version 1.1

Version 1.0 (26 September 2026) described the searches with phrase rules. Following an external review, version 1.1 re-analyzes the same 240 answers. It places the searches in the steps that lead to a citation and adds the cited pages of each answer. It also adds model-coded search functions and year roles, a similarity measure, intervals and models that group answers by question, and a second coder.

The phrase-rule figures are unchanged. The statement that 10.5% of ChatGPT’s and 16.4% of Gemini’s searches “asked for a past year” is replaced by year roles. Most of those searches asked for the latest edition of an annual ranking or paired the earlier year with 2026. The earlier advice that a business absent from named directories “is less likely to be found” is reworded as an association. The full change log is in the dataset.

## Data and downloads

- Every search with its functions, year role, similarity and named sources: [s19_searches_v11.csv](https://underneath.agency/research-data/ai-hidden-searches-study/s19_searches_v11.csv) and [JSON](https://underneath.agency/research-data/ai-hidden-searches-study/s19_searches_v11.json)
- Every answer with its searches, named sources and cited domains and pages: [s19_answers_v11.csv](https://underneath.agency/research-data/ai-hidden-searches-study/s19_answers_v11.csv) and [JSON](https://underneath.agency/research-data/ai-hidden-searches-study/s19_answers_v11.json)
- Every statistic on this page, including models and agreement: [stats.json](https://underneath.agency/research-data/ai-hidden-searches-study/stats.json)
- Version 1.0 searches and statistics: [s19_fanout_queries.csv](https://underneath.agency/research-data/ai-hidden-searches-study/s19_fanout_queries.csv) and [stats_v1.0.json](https://underneath.agency/research-data/ai-hidden-searches-study/stats_v1.0.json)
- Machine-readable methodology: [methodology.json](https://underneath.agency/research-data/ai-hidden-searches-study/methodology.json)

The data is free to reuse with attribution (CC BY 4.0).

To cite: Underneath. (2026). *The hidden searches AI assistants run before they answer* (Version 1.1). Underneath Research. https://underneath.agency/research/ai-hidden-searches-study

## Frequently asked questions

### What is query fan-out?

When an AI assistant uses web search, it writes its own search queries, often several, instead of searching for the user’s question as typed. Those searches decide which pages it can read and cite. In our test ChatGPT ran a mean of 3.7 searches per buyer question, Gemini 1.9 and Claude 0.76.

### Does a source named in ChatGPT’s search get cited?

Often. When ChatGPT’s search named a source such as NerdWallet, Avvo or Super Lawyers, the answer cited that source 45.6% of the time, against 5.1% in comparable answers that did not name it. But only 4.2% of all cited domains had been named in a search.

### Does ChatGPT search for old years?

Rarely without reason. 52.5% of ChatGPT’s answers included a search with a year, mostly 2026. Searches with an earlier year usually looked for the latest edition of an annual ranking; 3 of its 296 searches used an old year for no clear reason.

### Do more searches mean more sources?

Within an assistant, yes: twice as many searches went with 1.45 times as many cited domains. Between assistants, no: ChatGPT searched most but cited the fewest domains per answer.

### Do AI assistants search Reddit?

Rarely by name: 6.2% of Gemini’s answers included a search mentioning Reddit, and none of ChatGPT’s or Claude’s did. Gemini still cited Reddit in 31 of its 80 answers, found through ordinary searches.

## Related research

- [Do ChatGPT, Gemini, Perplexity and Claude cite pages that rank?](https://underneath.agency/research/ai-citations-google-rankings-study)
- [Do ChatGPT, Gemini, Perplexity and Claude agree on brands?](https://underneath.agency/research/ai-assistants-brand-agreement-study)
- [How fresh are the pages AI engines cite?](https://underneath.agency/research/ai-source-freshness-study)
- [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)

## Related guides

- [How do professional services firms win clients when buyers ask AI who to hire?](https://underneath.agency/resources/professional-services-firms-clients-ai-search)
- [Why do AI search engines cite different sources every time I check?](https://underneath.agency/resources/why-ai-search-citations-change)
- [Who in my organization should own AI search visibility?](https://underneath.agency/resources/who-should-own-ai-search-visibility)
- [How do SIEM vendors get onto enterprise shortlists through AI search?](https://underneath.agency/resources/siem-enterprise-pipeline-ai-search)
- [How many sources does each AI search engine cite per answer?](https://underneath.agency/resources/how-many-sources-ai-search-engines-cite)

---

This is the Markdown twin of https://underneath.agency/research/ai-hidden-searches-study. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "AI Mode vs AI Overviews: how different are the sources?"
description: "Same 400 searches: AI Mode and AI Overviews shared 13.9% of cited URLs, and AI Mode cited 16.3% of top-10 pages to the AI Overview’s 29.7%."
canonical: "https://underneath.agency/research/ai-mode-vs-ai-overviews-study"
published: 2026-09-26
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Research · AI search

# AI Mode vs AI Overviews: how different are the sources?

Google now answers searches with AI in two places: the AI Overview at the top of the normal results page, and AI Mode, a separate conversational tab. We asked both the same 400 US searches on the same day, 26 September 2026, and compared what each one cited. This version treats the two as separate search surfaces rather than one “Google AI search”. It splits visibility into stages (does the surface answer, does it cite, does a ranking page get cited, which source comes first, does a source cited on one surface also appear on the other) and reports how far each difference holds search by search.

## The short version

1. The two surfaces are different environments. Where both cited sources (236 of 400 searches), they shared a mean 13.9% of cited URLs (Jaccard 0.139, 95% interval 0.111 to 0.163) and 24.2% of cited domains excluding Google’s own. 30.5% of these searches had no cited URL in common, and they led with the same source in 20.9%.
2. Ranking on page one carries over to AI Overviews far better than to AI Mode. Of the 2,118 URLs in the organic top 10 on searches that showed both surfaces, the AI Overview cited 29.7% and AI Mode 16.3%. For positions 1 to 3 the figures were 42.3% and 26.0%.
3. AI Overviews cite more forums, social and video: 18.8% of their citations on the matched searches against 11.2% for AI Mode, and the AI Overview had the higher share in 56.4% of searches against 20.3% the other way.
4. Correction to version 1.0: most of AI Mode’s heavy use of Google’s own pages comes from searches that show no AI Overview at all. On those 147 searches, 69.3% of AI Mode’s citations were Google pages. On the matched searches the Google-owned share was 20.4% for AI Mode and 14.8% for AI Overviews, a gap of 1.4 times rather than 2.7, with wide, overlapping intervals.
5. AI Mode answered all 400 searches and cited a source in 94.0%; an AI Overview appeared on 63.2% (95% interval 52.2% to 74.2%). Counting every search where AI Mode cited something, the AI Overview cited at least one of the same URLs in 43.6%.

## Research questions

| Question | Answered here? |
|---|---|
| RQ1. How often does each surface answer and cite, counting every search? | Yes |
| RQ2. How much do the cited sources overlap, and does that depend on the measure? | Yes, with sensitivity checks |
| RQ3. Given that a page ranks in the top 10, how likely is each surface to cite it? | Yes |
| RQ4. Do the two surfaces lead with the same source? | Yes |
| RQ5. Does the gap depend on industry, intent or popularity? | Yes, with a clustered model |
| RQ6. Is the gap stable across runs, dates, wordings and follow-up turns? | No: one run per surface and search |

## Visibility in stages

Being “visible in Google’s AI” is not one number. We measure five stages separately for each surface. Three further stages are named here because they matter, but this study does not measure them.

| Stage | Question | Measured? |
|---|---|---|
| Appearance | Does the surface answer the search? | Yes |
| Citation | Does the answer cite any source? | Yes |
| Ranking to citation | Is a page that ranks in the top 10 cited? | Yes |
| Prominence | Is the source the first one cited? | Yes |
| Transfer | Is a source cited on one surface also cited on the other? | Yes |
| Attribution | Does the cited page support the sentence it is attached to? | No |
| Influence | Did the source shape the answer? | No |
| User outcome | Did anyone click, trust or act on it? | No |

## What we measured

We took 400 keywords from the 800-keyword US sample of our [AI Overview frequency study](https://underneath.agency/research/ai-overviews-frequency-study), 50 per industry (20 of the highest-volume and 30 of the randomly drawn keywords), spread over 32 seed topics. For each one we recorded the regular Google results page (with any AI Overview) and the AI Mode answer, both from Google.com, United States, English, desktop, on 26 September 2026. From each AI answer we collected every cited link in reading order, normalized the URLs and compared the two sets.

Every search stays in the results. A search with no AI Overview, or an answer with no citation, is recorded as an outcome rather than dropped. Keywords that share a seed topic (“pizza”, “car insurance”) are not independent, so every 95% interval on this page comes from resampling the 32 seed topics, not individual keywords.

## Findings

### Every search, by outcome

| Outcome (400 searches) | Searches | Share |
|---|---|---|
| Both shown and both cited | 236 | 59.0% |
| AI Mode only (no AI Overview shown) | 147 | 36.8% |
| Both shown, only the AI Overview cited | 13 | 3.2% |
| Both shown, only AI Mode cited | 4 | 1.0% |

AI Mode answered every search and cited at least one source in 94.0% of them. The AI Overview appeared on 63.2% and cited a source in 62.2%. Of the 376 searches where AI Mode cited something, the AI Overview cited at least one of the same URLs in 43.6% (95% interval 33.2% to 54.8%) and one of the same non-Google domains in 47.6%.

### The two answers rarely cite the same pages

| Measure (236 searches where both cited) | Result | 95% interval |
|---|---|---|
| Mean URL overlap (Jaccard) | 0.139 | 0.111 to 0.163 |
| Searches with no cited URL in common | 30.5% | 24.4% to 38.1% |
| Mean domain overlap, excluding google.com | 0.242 | 0.192 to 0.281 |
| Share of AI Mode’s URLs also in the AI Overview | 28.4% | 23.9% to 32.6% |
| Share of the AI Overview’s URLs also in AI Mode | 22.2% | 17.0% to 27.0% |

The median URL overlap was 0.1, and 23.2% of searches had no cited domain in common.

The finding does not depend on how overlap is measured. The AI Overview cites more sources (median 8 against 4.5 on these searches), which pulls Jaccard down. The overlap coefficient, which divides by the smaller set instead, is 0.332: even then, two-thirds of the shorter list is not in the longer one. Dropping Google’s own URLs gives 0.138, counting google.com in the domain comparison gives 0.272, and comparing only the first three sources each surface cites gives 0.164.

### They lead with different sources

On the 235 searches where both surfaces cited a source inside the answer, we compared the first source each one cited.

| First source | Result | 95% interval |
|---|---|---|
| Same URL | 20.9% | 14.1% to 26.6% |
| Same domain | 34.0% | 24.1% to 44.5% |
| Overview’s first, cited by AI Mode | 41.7% | 32.7% to 49.6% |
| AI Mode’s first, cited by Overview | 42.6% | 34.2% to 49.8% |

A page that is the lead source on one surface is missing entirely from the other surface’s answer more often than not.

### Ranking pages are cited far less in AI Mode

Organic rankings are the one stage before citation we can observe. On the 253 searches that showed both surfaces, we followed each of the 2,118 URLs in the organic top 10.

| Top-10 URLs cited | AI Mode | AI Overviews |
|---|---|---|
| All top-10 URLs | 16.3% | 29.7% |
| Positions 1 to 3 | 26.0% | 42.3% |
| Positions 4 to 10 | 11.0% | 22.7% |

Intervals for all top-10 URLs: AI Mode 12.1% to 21.2%, AI Overviews 26.9% to 32.7%. The intervals do not overlap.

Of those 2,118 ranking URLs, 10.0% were cited by both surfaces, 19.7% only by the AI Overview, 6.3% only by AI Mode and 63.9% by neither. Seen from the citation side, 19.9% of AI Mode’s citations on the matched searches were top-10 pages, against 29.7% for AI Overviews, and the AI Overview had the higher top-10 share in 50.4% of searches against 30.1% the other way.

### Source types on the same searches

On the 236 searches where both cited, AI Mode made 1,728 citations and the AI Overviews 2,013. “Other websites” are brands, publishers and blogs.

| Source type | AI Mode | AI Overviews |
|---|---|---|
| Other websites | 62.8% | 59.0% |
| Google’s own pages | 20.4% | 14.8% |
| Forums, social, video | 11.2% | 18.8% |
| Reviews and directories | 1.9% | 2.6% |
| Government | 1.4% | 0.9% |
| Education | 1.2% | 1.5% |
| Marketplaces, retailers | 0.6% | 1.4% |
| Wikipedia | 0.5% | 0.8% |

The forum, social and video gap holds search by search: the AI Overview had the higher share in 56.4% of searches, AI Mode in 20.3%, and they tied in 23.3%. The Google-owned gap is smaller and less consistent. In 72.5% of searches the two had the same Google share (usually none), AI Mode was higher in 17.8% and the AI Overview in 9.7%. The pooled shares have wide intervals (AI Mode 8.7% to 36.5%, AI Overviews 7.3% to 24.1%) because Google pages cluster in a few local and shopping topics.

AI Mode also spreads its citations more widely. Counting each domain once per answer, its effective number of domains was 93.5 against 58.8 for AI Overviews, even though it drew on fewer distinct domains (512 against 602).

### Where the Google-owned citations come from

On the 147 searches with no AI Overview, AI Mode made 1,155 citations, and 69.3% of them were Google pages, mostly the “searchviewer” listing pages for local businesses and products. Those searches are concentrated in franchises (20.4% of them), healthcare, home services and legal (17.7% each). On the same searches only 7.8% of AI Mode’s cited URLs ranked in the top 10. Version 1.0 compared AI Mode across all 400 searches with AI Overviews on the searches that showed one, so it mixed a surface difference with this difference in search mix.

### The gap depends on the industry, not the intent

| Industry | Searches where both cited | Mean URL overlap |
|---|---|---|
| Hospitality and travel | 33 | 0.221 |
| Financial services and insurance | 44 | 0.184 |
| B2B software and technology | 46 | 0.144 |
| Healthcare and dental | 23 | 0.136 |
| Franchises and multi-location brands | 18 | 0.114 |
| Legal and professional services | 21 | 0.099 |
| Retail and ecommerce | 30 | 0.075 |
| Home and local services | 21 | 0.061 |

In a model of per-search URL overlap with industry, search intent and sample stratum together, and standard errors clustered by seed topic, industry mattered (p < 0.001) but intent (p = 0.624) and highest-volume versus random keywords (p = 0.494) did not. Overlap was highest in travel (95% interval 0.187 to 0.269) and lowest in home and local services (0.039 to 0.094). The model explains little of the variation (R² 0.138): most of the difference is search by search.

## How this compares with other studies

| Source | Sample and date | Figure |
|---|---|---|
| This study | 236 US search pairs, September 2026 | Mean URL overlap 13.9%; domain overlap 24.2% |
| Ahrefs | 540,000 query pairs, US, September 2025 | 13.7% of cited URLs matched |
| SE Ranking | 10,000 US keywords, August 2025 | URL overlap 10.7%; domain overlap 16% |
| SE Ranking | Same | AI Mode against organic top 10: 14% of URLs, 21.9% of domains |
| seoClarity | 1,000 transactional US queries, October 2025 | 19% of AI Mode citations from the top 20 rankings |

A year on, the two products still diverge about as much as Ahrefs and SE Ranking measured in 2025: our URL overlap is within a few points of both, and our domain overlap is higher than SE Ranking’s. Each study defines overlap slightly differently, so small gaps are not meaningful. Our share of AI Mode’s cited URLs that rank in the top 10 across all 400 searches (15.1%) is close to SE Ranking’s 14%.

Sources: [Ahrefs](https://ahrefs.com/blog/ai-overviews-vs-ai-mode/); [SE Ranking](https://seranking.com/blog/ai-mode-research/); [seoClarity](https://www.seoclarity.net/ai-mode-rankings-overlap).

## Observed, inferred and unknown

**What we observe.** For the same search on the same day, AI Mode and the AI Overview cite mostly different pages, lead with different sources, and turn organic rankings into citations at very different rates. AI Overviews lean more on forums, social and video. AI Mode appears on every search, including the local and shopping searches where no AI Overview is shown, and on those it cites Google’s own listings heavily.

**What we infer.** “Google AI search” is not one retrieval environment. A result measured on one surface should not be assumed to hold on the other, and visibility figures that blend them hide which one a brand is winning.

**What remains unknown.** Whether the gap is stable across repeated runs, days and rewordings (part of it may be ordinary run-to-run variation). Whether AI Mode behaves differently as a follow-up in a conversation. Why each surface picks the sources it does: this study observes citations, not the retrieval behind them. Whether the cited pages support the sentences they are attached to, whether they shaped the answer, and whether anyone clicked.

## What this means

The points below are our interpretation. They follow from the findings but were not tested.

- **Track the two surfaces separately.** Being cited in the AI Overview says little about being cited in AI Mode for the same search. Report each on its own.
- **Ranking well is a better route into AI Overviews than into AI Mode.** A top-3 page was cited by the AI Overview about four times in ten and by AI Mode about one time in four.
- **For searches without an AI Overview, Google’s own listings are the front door.** Where AI Mode is the only AI answer, most of its citations point to Google’s listing pages, so the information Google holds about a business (its Business Profile, product feed and reviews) is likely to matter most there.
- **AI Overviews reward video and community presence more.** The forum, social and video share is higher in AI Overviews on most searches, not just on average.

## Where this sits in GEO research

Research on generative engine optimization began by asking whether changing a page changes how often an AI answer uses it. Later work compared the sources that different AI search systems draw on, and argued that visibility should be measured as separate quantities (retrieval, citation, prominence, attribution, stability) rather than one score, and that search surfaces within one engine can differ as much as engines do. This study measures that last point directly for Google: same engine, same searches, same day, two surfaces. It is observational. It shows where the two surfaces differ, not what would change a page’s visibility on either, and it does not yet measure stability. A second collection of the same searches is planned, and we will add run-to-run figures to this page when it is analyzed.

## Methodology

- **Keywords:** 400 of the 800 US keywords in our AI Overview frequency study, 50 per industry (the first 20 of the 40 highest-volume keywords and the first 30 of the 60 randomly drawn keywords, in sample order), from 32 seed topics.
- **Collection:** DataForSEO SERP API, Google organic (with asynchronous AI Overviews) and Google AI Mode, Google.com, United States, English, desktop, 26 September 2026, one request each.
- **Citations:** every cited link in each AI answer (links attached to each paragraph in reading order, then the reference list), normalized (tracking parameters, fragments, “www.” and trailing slash removed) and de-duplicated. The first source is the first link cited inside the answer body.
- **Overlap:** Jaccard similarity (shared divided by all distinct sources) per search, averaged over the 236 searches where both answers cited at least one source; domain overlap excludes google.com unless stated. Sensitivity checks use the overlap coefficient, overlap without Google URLs and the first three sources.
- **Ranking to citation:** the share of organic top-10 URLs that each surface cites, on the 253 searches that showed both.
- **Uncertainty:** 95% bootstrap intervals resampling the 32 seed topics (2,000 resamples, seed 20260926).
- **Moderators:** OLS of per-search URL overlap on industry, keyword intent (DataForSEO’s label) and stratum, with standard errors clustered by seed topic; joint Wald tests.
- **Update schedule:** a second collection date with our week-over-week volatility study; monthly after that.

## Limitations

- One run per surface and search on one date. Both surfaces vary from run to run (see our [consistency study](https://underneath.agency/research/ai-recommendation-consistency-study)), so part of the difference may be ordinary variation.
- Observational only. The organic top 10 is the only stage before citation we can see; a page can be retrieved without ranking there.
- Citation is not attribution or influence. We did not check whether cited pages support the claims attached to them, or measure clicks.
- AI Mode was queried as a first question, not as a follow-up in a conversation.
- Source types use a fixed domain list; intent is DataForSEO’s keyword label. Small industry groups (18 to 46 searches) have wide intervals.

## What changed in version 1.1

Version 1.0 (26 September 2026) compared the two surfaces on overlap and source mix. Following an external review, version 1.1 (28 September 2026) re-analyzes the same 400 searches with no new queries. It keeps every search as an outcome, splits visibility into stages, adds prominence, ranking-to-citation and transfer measures, adds intervals clustered by seed topic, sensitivity checks and a moderator model.

One headline changed. Version 1.0 reported that 39.9% of AI Mode’s 2,894 citations were Google pages, 2.7 times the 14.6% in AI Overviews, and that AI Overviews cited forums, social and video 2.4 times as often (19.2% against 8.1%). Those figures compared AI Mode on all 400 searches with AI Overviews on the searches that showed one. On the same searches the ratios are 1.4 and 1.7. Version 1.0’s overlap figures are unchanged.

## Data and downloads

- Per-search results for version 1.1 (outcome, overlap, first sources, top-10 citations): [s6_searches_v11.csv](https://underneath.agency/research-data/ai-mode-vs-ai-overviews-study/s6_searches_v11.csv) and [JSON](https://underneath.agency/research-data/ai-mode-vs-ai-overviews-study/s6_searches_v11.json)
- Every statistic on this page: [stats.json](https://underneath.agency/research-data/ai-mode-vs-ai-overviews-study/stats.json)
- Machine-readable methodology: [methodology.json](https://underneath.agency/research-data/ai-mode-vs-ai-overviews-study/methodology.json)
- Chart: [top-10 URLs cited by each surface](https://underneath.agency/research-data/ai-mode-vs-ai-overviews-study/ranking-to-citation-ai-mode-vs-overviews.svg)
- Version 1.0 data and statistics: [s6_ai_mode_vs_overview.csv](https://underneath.agency/research-data/ai-mode-vs-ai-overviews-study/s6_ai_mode_vs_overview.csv) and [stats_v1.0.json](https://underneath.agency/research-data/ai-mode-vs-ai-overviews-study/stats_v1.0.json)

The data is free to reuse with attribution (CC BY 4.0).

To cite: Underneath. (2026). *AI Mode vs AI Overviews: how different are the sources?* (Version 1.1). Underneath Research. https://underneath.agency/research/ai-mode-vs-ai-overviews-study

## Frequently asked questions

### Do Google AI Mode and AI Overviews cite the same sources?

Mostly not. For the same search on the same day, they shared a mean 13.9% of cited URLs, 30.5% of searches had no cited URL in common, and they led with the same source in 20.9%.

### Does ranking on page one get you cited in AI Mode?

Less often than in AI Overviews. Of URLs ranking in the organic top 10, AI Mode cited 16.3% and the AI Overview 29.7% on the same searches; for positions 1 to 3, 26.0% against 42.3%.

### Does AI Mode cite Google’s own pages more than AI Overviews?

Mainly on searches that show no AI Overview, such as local and shopping searches, where 69.3% of AI Mode’s citations were Google pages. On searches where both cited sources, the shares were 20.4% and 14.8%.

### Does every search have an AI Mode answer?

In our sample, yes: AI Mode answered all 400 searches, while an AI Overview appeared on 63.2% of them.

## Related research

- [Do AI Overviews cite the pages that rank?](https://underneath.agency/research/ai-overview-citations-study)
- [How often do AI Overviews appear?](https://underneath.agency/research/ai-overviews-frequency-study)
- [Do ChatGPT, Gemini, Perplexity and Claude agree on brands?](https://underneath.agency/research/ai-assistants-brand-agreement-study)
- [What pages cited by AI Overviews have in common](https://underneath.agency/research/ai-overview-cited-pages-study)

## Related guides

- [How much traffic would AI Mode as Google’s default cost you?](https://underneath.agency/resources/ai-mode-default-traffic-loss)
- [If we rank on Google, will we show up in AI Overviews?](https://underneath.agency/resources/does-ranking-on-google-get-you-into-ai-overviews)
- [Do AI search engines cite the same websites as Google?](https://underneath.agency/resources/do-ai-search-engines-cite-the-same-sites-as-google)
- [How much control do we have when Google changes how search results look?](https://underneath.agency/resources/google-search-design-changes-traffic-risk)
- [Does social media content help your brand show up in AI search?](https://underneath.agency/resources/does-social-media-help-ai-search-visibility)

---

This is the Markdown twin of https://underneath.agency/research/ai-mode-vs-ai-overviews-study. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Do AI Overviews cite the pages that rank? 4,051 citations"
description: "Only 28.7% of 4,051 AI Overview citations are page-one Google results. A page-one result’s citation rate falls from 49.5% at position 1 to 15.5% at 9."
canonical: "https://underneath.agency/research/ai-overview-citations-study"
published: 2026-09-26
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Research · AI search

# Do AI Overviews cite the pages that rank? 4,051 citations

Ranking on Google and being cited in the AI Overview above the results are two different outcomes, and this study measures the gap between them. We collected every source cited in 481 US AI Overviews on 26 September 2026, 4,051 citations in all, and compared each one with the organic results on the same page. Organic ranking turned out to be neither irrelevant nor sufficient: a page-one result is far more likely to be cited the higher it ranks, yet most citations come from pages that do not rank on page one at all.

## The short version

1. 28.7% of AI Overview citations were page-one organic results for the same search (95% interval 26.8% to 30.7%). Leaving out Google’s own pages, which can never be organic results, the share is 33.6%.
2. Position matters steadily. The first organic result was cited in 49.5% of AI Overviews, the third in 33.5%, the sixth in 23.2% and the ninth in 15.5%. Each step down the page cuts the odds of being cited by about 17.9%.
3. 78.8% of AI Overviews cited at least one page-one result, but 21.2% cited none, and 15.8% cited nothing that appeared anywhere on the results page, not even in a video or forum box.
4. Of the citations that were not page-one results, the largest group, 38.5% of all citations, came from sites with no page-one organic result for the search.
5. YouTube was cited in 64.2% of AI Overviews. Only 12.5% of its 638 citations were page-one organic results; 60.7% did not appear anywhere on the results page.

## What this study measures, and what it does not

An AI Overview goes through several stages before a source reaches the reader: Google decides to show one, retrieves candidate pages, picks some to cite, and writes an answer that draws more or less on each. From outside, only some of these stages can be seen.

| Stage | Measured here? | Result |
|---|---|---|
| AI Overview shown | Yes | 486 of 800 searches (60.8%) |
| Organic ranking | Yes, page one | 4,016 organic results on 481 pages |
| Candidates retrieved | No | Not visible from outside |
| Source cited | Yes | 4,051 citations |
| Share of the answer from each source | No | Not measured |

So this is a study of how one observable outcome, organic ranking, relates to another, citation. It does not show how Google retrieves or chooses sources, and it does not show that ranking higher causes a page to be cited. Pages that rank well may simply share whatever else the AI Overview looks for.

## Finding 1: most citations are not page-one results

| Measure | Result | 95% interval |
|---|---|---|
| Exact URL is a page-one organic result | 28.7% | 26.8% to 30.7% |
| Same page with a looser URL match | 30.9% | 28.8% to 33.1% |
| URL appears anywhere on the results page | 35.7% | 33.4% to 38.0% |
| Domain has a page-one organic result | 43.0% | |
| Median citations per AI Overview | 8 | |

The looser match treats URLs that differ only by query string, letter case, mobile or AMP versions, or an index or .html ending as the same page. It adds 2.2 points. Counting links that appear in video carousels, image boxes and other results-page features adds 4.8 more. Intervals are clustered by AI Overview, because the 8 or so citations in one AI Overview are not independent of each other.

## Finding 2: the higher a page ranks, the more often it is cited

The version 1.0 page counted cited URLs at each position. That undercounts lower positions, because most results pages do not have ten organic results: the median page had 9, and only 14 of 481 (2.9%) had 10. The rate below divides by the pages that had a result at each position.

| Position | Pages with it | Cited | Rate | 95% interval |
|---|---|---|---|---|
| 1 | 481 | 238 | 49.5% | 45.0% to 53.9% |
| 2 | 481 | 179 | 37.2% | 33.0% to 41.6% |
| 3 | 481 | 161 | 33.5% | 29.4% to 37.8% |
| 4 | 481 | 137 | 28.5% | 24.6% to 32.7% |
| 5 | 480 | 122 | 25.4% | 21.7% to 29.5% |
| 6 | 475 | 110 | 23.2% | 19.6% to 27.2% |
| 7 | 448 | 99 | 22.1% | 18.5% to 26.2% |
| 8 | 404 | 75 | 18.6% | 15.1% to 22.6% |
| 9 | 271 | 42 | 15.5% | 11.7% to 20.3% |
| 10 | 14 | 1 | 7.1% | 1.3% to 31.5% |

By band, 40.1% of results in positions 1 to 3 were cited, 25.7% in positions 4 to 6 and 19.1% in positions 7 to 10. Overall, 1,164 of the 4,016 page-one organic results (29.0%) were cited.

A logistic model of whether each organic result was cited, with position, industry and the number of citations in the AI Overview, and standard errors clustered by AI Overview, gives an odds ratio of 0.821 per position (95% interval 0.798 to 0.845). Its average predicted citation rate is 44.4% at position 1, 35.2% at position 3, 27.0% at position 5 and 17.2% at position 8. The decline is gradual, with no cliff after the top three.

## Finding 3: where the other citations come from

Every citation falls in exactly one of these classes, checked in this order.

| Class | Citations | Share |
|---|---|---|
| Exact page-one organic URL | 1,164 | 28.7% |
| Same page, looser URL match | 89 | 2.2% |
| Elsewhere on the results page | 193 | 4.8% |
| Google’s own pages | 589 | 14.5% |
| Other page on a page-one domain | 455 | 11.2% |
| Domain not on page one | 1,561 | 38.5% |

Of the 193 links found elsewhere on the results page, 155 were in the video carousel. Google’s own pages (shopping and local listings on google.com) cannot be organic results, so they are a separate class; 589 of the 602 google.com citations fall into it.

YouTube shows the gap between ranking and citation most clearly. Of its 638 citations, 80 (12.5%) were page-one organic results, 251 (39.3%) appeared somewhere on the results page, mostly in the video carousel, and 60.7% did not appear on the page at all.

## Finding 4: which domains and source types are cited

| Domain | AI Overviews citing it | Share of 481 |
|---|---|---|
| youtube.com | 309 | 64.2% |
| google.com (Google’s own pages) | 122 | 25.4% |
| reddit.com | 87 | 18.1% |
| nerdwallet.com | 48 | 10.0% |
| libertymutual.com | 36 | 7.5% |
| facebook.com | 34 | 7.1% |
| allstate.com | 33 | 6.9% |
| wikipedia.org | 33 | 6.9% |
| zapier.com | 32 | 6.7% |
| thezebra.com | 29 | 6.0% |
| hubspot.com | 29 | 6.0% |
| salesforce.com | 28 | 5.8% |

YouTube accounted for 638 citations, 15.7% of all 4,051. In some categories, brands’ own sites are among the most-cited sources: salesforce.com, hubspot.com, zapier.com and zoho.com in software; libertymutual.com, allstate.com and progressive.com in insurance.

| Source type | Share of citations |
|---|---|
| Other websites (brands, publishers, blogs) | 57.8% |
| Forums, social and video platforms | 20.0% |
| Google’s own pages | 14.9% |
| Review sites and directories | 2.3% |
| Marketplaces and retailers | 1.7% |
| Government | 1.3% |
| Education | 1.1% |
| Wikipedia | 0.9% |

## Finding 5: industries differ in how far the AI Overview reaches

| Industry | AIOs | Citations | On page one | 95% interval |
|---|---|---|---|---|
| Healthcare and dental | 42 | 358 | 38.0% | 31.0% to 45.2% |
| Finance and insurance | 88 | 644 | 34.8% | 30.0% to 39.6% |
| Legal and professional | 39 | 318 | 33.0% | 25.8% to 41.3% |
| Franchises | 35 | 309 | 32.7% | 25.6% to 40.8% |
| Hospitality and travel | 61 | 447 | 30.9% | 24.4% to 37.7% |
| Home and local services | 53 | 464 | 24.4% | 18.9% to 30.6% |
| B2B software | 96 | 870 | 23.9% | 20.1% to 27.9% |
| Retail and ecommerce | 67 | 641 | 21.7% | 17.2% to 26.3% |

“On page one” is the share of citations that are exact page-one organic URLs. Healthcare and finance searches lean more on pages that already rank than retail and software searches, and their intervals do not overlap. Most neighboring industries’ intervals do overlap, so smaller gaps should not be read as real differences. Google’s own pages make up 32.9% of retail citations and 25.6% of home-services citations but 0.9% in software, which explains part of the gap.

## How this compares with other studies

| Source | Sample and date | Figure |
|---|---|---|
| This study | 4,051 citations, 481 US AI Overviews, September 2026 | 28.7% of cited URLs are page-one results |
| Ahrefs | 863K keyword results pages, March 2026 | 37.9% of cited URLs appear in the first 10 result blocks (about 76% in July 2025) |
| BrightEdge | 9 industries, September 2025 | Overlap with organic rankings grew from 32.3% to 54.5% |
| seoClarity | US keywords, 2025 | 97% of AI Overviews cite at least one source from the top 20 |

The trend is disputed: Ahrefs reports the top-10 share falling from about 76% to 37.9% between July 2025 and March 2026, while BrightEdge reported overlap rising through 2025. Our 28.7% is closer to the Ahrefs direction. Figures also differ because some count citations and others count AI Overviews, and because “top 10” means page-one organic results in our study (a median of 9) but ten result blocks, including features, in others. For most-cited domains, Ahrefs (September 2026) also puts YouTube first and Reddit second among the top cited sources, and Pew Research found Wikipedia, YouTube and Reddit made up 15% of AI summary sources in March 2025.

Sources: [Ahrefs, top-10 citations](https://ahrefs.com/blog/ai-overview-citations-top-10/); [BrightEdge](https://www.brightedge.com/resources/weekly-ai-search-insights/rank-overlap-after-16-months-of-aio); [seoClarity](https://www.seoclarity.net/research/ai-overviews-impact); [Ahrefs, most-cited domains](https://ahrefs.com/blog/most-cited-domains-ai-overviews/); [Pew Research Center](https://www.pewresearch.org/short-reads/2025/07/22/google-users-are-less-likely-to-click-on-links-when-an-ai-summary-appears-in-the-results/).

## What this means

The measured results are in the findings above. What follows separates them from our reading of them.

| Observed | Plausible reading | Not shown by this study |
|---|---|---|
| Citation rate falls steadily with organic position | Ranking and citation draw on overlapping signals | That ranking higher causes a citation |
| 38.5% of citations come from sites with no page-one result | The AI Overview also draws on pages that answer sub-questions of the search | Which pages Google retrieved but did not cite |
| YouTube is cited in 64.2% of AI Overviews, mostly from outside the results page | Video is a large, separate source pool for AI Overviews | That publishing videos would get a brand cited |
| Brands such as insurers and software firms are cited from their own sites | A clear factual page on your own site can be a candidate citation | That such a page raises traffic or sales |

- **Ranking helps, but it is not the whole job.** Half of AI Overviews cite the first result, and the rate keeps falling down the page, yet seven in ten citations are not page-one results.
- **Measure citation separately from ranking.** A rank report cannot tell you whether you are cited; track the AI Overview itself for the searches that matter.
- **Treat any single figure with care.** This is one day and one results page per search.

## Open questions

The limits of this study are also the next questions to test.

- **How stable are citations?** We collected once. Our planned week-over-week re-collection of the same 800 searches will measure how many cited sources persist.
- **What explains citation beyond rank?** Page content, freshness and source type may account for part of the 38.5% cited from off page one; our [cited pages study](https://underneath.agency/research/ai-overview-cited-pages-study) looks at what cited pages have in common.
- **Does citation mean influence?** A cited page may supply one fact or most of the answer. Measuring how much of the answer each source supplies is a separate study.
- **Do the same patterns hold in other engines?** Our [AI Mode comparison](https://underneath.agency/research/ai-mode-vs-ai-overviews-study) and [AI citations and Google rankings study](https://underneath.agency/research/ai-citations-google-rankings-study) test this for AI Mode and chat assistants.

## Methodology

- **Keywords and results pages:** the 800-keyword US sample of our [AI Overview frequency study](https://underneath.agency/research/ai-overviews-frequency-study) (DataForSEO SERP API; Google.com, United States, English, desktop, 26 September 2026). 486 pages showed an AI Overview and 481 of those cited at least one source.
- **Citations:** all URLs in each AI Overview’s reference list, per-paragraph references and inline links, normalized (tracking parameters, fragments and “www.” removed, trailing slash dropped) and de-duplicated within each AI Overview.
- **Ranking match:** normalized citation URL equal to a normalized page-one organic URL. The looser match also ignores the query string, letter case, mobile and AMP versions, and index or .html endings. Domain match is on the registrable domain.
- **Citation rate by position:** cited organic results at a position divided by results pages that had an organic result at that position, with Wilson intervals.
- **Intervals and model:** citation-level shares use a bootstrap over AI Overviews (2,000 resamples). The position model is a logistic regression with position, industry and number of citations, standard errors clustered by AI Overview.
- **Source types:** a fixed domain list (published in the methodology file): forums, social and video platforms; review sites and directories; marketplaces; government (.gov and national equivalents); education (.edu and national equivalents); Wikipedia; Google’s own pages; everything else.
- **Version 1.1 (28 September 2026):** same collection, no new data. Adds the stage framing, the citation rate by position, the looser match, the single class list, clustered intervals, the position model and the observed versus interpretation table, after an external methodological audit.
- **Update schedule:** monthly, same keywords.

## Limitations

- Associations, not causes; retrieval and reranking are not visible from outside.
- One results page per keyword on one day, so run-to-run and week-to-week variation is not measured.
- Page-one organic results only; some cited pages will rank on page two or lower.
- We measure whether a source is cited, not how much of the answer it supplies or whether the answer represents it correctly.
- Source types come from a fixed list; small forums and review sites not on the list are counted as other websites.

## Data and downloads

- Every citation with its match class and organic position (version 1.1): [s5_v11_citations.csv](https://underneath.agency/research-data/ai-overview-citations-study/s5_v11_citations.csv)
- Every citation with its domain and source type (version 1.0): [s5_ai_overview_citations.csv](https://underneath.agency/research-data/ai-overview-citations-study/s5_ai_overview_citations.csv) and [JSON](https://underneath.agency/research-data/ai-overview-citations-study/s5_ai_overview_citations.json)
- Every statistic on this page: [stats.json](https://underneath.agency/research-data/ai-overview-citations-study/stats.json)
- Machine-readable methodology: [methodology.json](https://underneath.agency/research-data/ai-overview-citations-study/methodology.json)
- Chart: [citation rate by organic position](https://underneath.agency/research-data/ai-overview-citations-study/citation-rate-by-organic-position.svg)

The data is free to reuse with attribution (CC BY 4.0).

To cite: Underneath. (2026). *Do AI Overviews cite the pages that rank? 4,051 citations* (Version 1.1). Underneath Research. https://underneath.agency/research/ai-overview-citations-study

## Frequently asked questions

### Do you need to rank on page one to be cited in an AI Overview?

No. In our sample 71.3% of AI Overview citations were not page-one organic results for the same search. Ranking still helps: the first result was cited in 49.5% of AI Overviews and the ninth in 15.5%.

### Does ranking higher get you cited?

Higher-ranked pages are cited more often, but this study shows an association, not a cause. Pages that rank well may share the other qualities the AI Overview selects for.

### What website is cited most in Google AI Overviews?

YouTube. It was cited in 64.2% of the 481 AI Overviews we collected, and accounted for 15.7% of all citations, most of them videos that did not appear on the results page.

### How often does Reddit appear in AI Overviews?

Reddit was cited in 18.1% of AI Overviews in our September 2026 US sample, the third most frequent domain after YouTube and Google’s own pages.

### How many sources does an AI Overview cite?

A median of 8 sources per AI Overview in our sample.

## Related research

- [How often do AI Overviews appear?](https://underneath.agency/research/ai-overviews-frequency-study)
- [AI Mode vs AI Overviews: how different are the sources?](https://underneath.agency/research/ai-mode-vs-ai-overviews-study)
- [What pages cited by AI Overviews have in common](https://underneath.agency/research/ai-overview-cited-pages-study)
- [How many “best of” lists cited by AI rank their own brand first?](https://underneath.agency/research/self-promoting-best-lists-study)

## Related guides

- [If we rank on Google, will we show up in AI Overviews?](https://underneath.agency/resources/does-ranking-on-google-get-you-into-ai-overviews)
- [Is traditional SEO still important for visibility in AI search?](https://underneath.agency/resources/is-seo-still-important-for-ai-search)
- [If AI Overviews cite my site, does that make up for lost clicks?](https://underneath.agency/resources/ai-overview-citations-vs-lost-clicks)
- [Which types of websites do AI search engines rely on most?](https://underneath.agency/resources/which-websites-do-ai-search-engines-cite)
- [How can an insurance carrier or agency win more quotes when shoppers ask AI?](https://underneath.agency/resources/insurance-companies-customers-ai-search)

---

This is the Markdown twin of https://underneath.agency/research/ai-overview-citations-study. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "What pages cited by AI Overviews have in common: 3,096 pages"
description: "3,096 top-10 pages on the same searches: rank predicted AI Overview citation. Compared within each search, only a machine-readable date held up."
canonical: "https://underneath.agency/research/ai-overview-cited-pages-study"
published: 2026-09-26
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Research · AI search

# What pages cited by AI Overviews have in common: 3,096 pages

Most advice about getting cited in Google’s AI Overviews compares cited pages with “the web”. We compared them with the pages they beat: for 486 US searches that showed an AI Overview, we fetched every page ranking in the top 10 and compared the ones the AI Overview cited with the ones it passed over on the same results page. Version 1.1 goes further. It compares each page only with its competitors on the same search, enters all 14 page features at once, and follows cited pages past the citation: whether they were the answer’s first source, and how much of the wording that cites them can be found on the page. Ranking position still came first. Of the on-page features, only a machine-readable date held up, and only on informational searches.

## The short version

1. Position came first: 41.7% of pages ranking 1 to 3 were cited, against 20.1% at positions 7 to 10. Within the same search, position accounted for 6.1% of the variation in which pages were cited, and the 14 page features added 2.5 points more; 91.4% was left unexplained by anything we measured.
2. Compared only with pages on the same search, with all features entered together, one feature survived a correction for testing 14 at once: a machine-readable date (+7.9 points, 95% interval 2.9 to 12.9).
3. Question-form headings, the largest difference in version 1.0 (+9.6 points), fell to +3.5 (interval −0.3 to 7.1) and are no longer distinguishable from zero. About half of the original gap came from which searches have such pages; entering the other features took off a little more.
4. The date association is confined to informational searches: +17.8 points there, against +1.0 on commercial and transactional searches. Article schema shows the same split (+10.9 against +0.1).
5. Organization schema, BreadcrumbList schema and HTML tables showed no meaningful association with citation; their intervals sit inside ±5 points.
6. Being cited is not the end of the story. Only 21.0% of cited pages were the answer’s first source, and in 62.2% of cited pages no passage citing the page had at least 80% of its content words on the page. Pages with a table were 14.7 points more likely to have their passages absorbed, although tables did nothing for citation itself.

## Why this version separates the stages

An AI Overview does not rank pages; it picks some, then writes with them. A page can rank and not be picked, be picked and appear only as the fifth link, or be linked while the sentence it is attached to says something the page does not. Treating “visibility” as one number hides which of these steps a change to a page affects.

Version 1.1 therefore measures three outcomes separately, each among the pages that reached the previous step.

| Stage | Question | Pages |
|---|---|---|
| 1. Citation | Is a top-10 page cited at all? | 3,096 |
| 2. Prominence | Is a cited page the first source, and what share of passages link it? | 870 |
| 3. Absorption | Is the wording that cites it found on the page? | 842 |

Three steps sit outside this data. Whether Google retrieved a page at all (its candidate set, including pages outside the top 10) is not visible. Whether the answer states faithfully what the page says (fidelity) needs a reading of meaning, not word overlap. What users then do is not observed.

A second separation matters as much. This study measures how often pages with a trait are cited. It does not measure what happens when a trait is added to a page, because no page was changed. A 2023 lab study of generative engine optimization (Aggarwal et al.) reported that rewriting a source could raise its visibility by up to 40%, but in that setup the source was already among the documents given to the engine. That result and ours answer different questions, and neither says whether adding a trait moves a real page into a live AI Overview.

## What we measured

We started from the 486 US searches in our [AI Overview frequency study](https://underneath.agency/research/ai-overviews-frequency-study) that showed an AI Overview, and took every organic result in the top 10, leaving out platform pages (YouTube, Reddit, Facebook, Google’s own pages and other social sites) whose structure the site owner does not control. Each page was labeled cited or not cited according to whether its URL appeared in that search’s AI Overview.

We fetched each page directly, and through Firecrawl when the direct request failed, returned an error status or returned too little HTML; Firecrawl supplied 612 of the usable pages. In all, 3,096 of 3,624 pages (85.4%) returned a usable HTML page: 936 cited and 2,160 not cited. Fetch success was nearly identical for the two groups (86.1% and 85.1%).

For each page we recorded its JSON-LD types, its newest machine-readable date, author signals, word count, headings, question-form headings, FAQ sections, tables and lists.

New in version 1.1, we read the text of each AI Overview, which links its sources passage by passage. Of the 936 cited pages, 870 (92.9%) are linked inside the answer text; the rest appear only in the source list. No uncited page appeared in the text, which confirms the version 1.0 labels.

## Stage 1: which ranking pages get cited

### Position

| Organic position | Cited | Not cited | Share cited |
|---|---|---|---|
| 1 to 3 | 459 | 642 | 41.7% |
| 4 to 6 | 291 | 779 | 27.2% |
| 7 to 10 | 186 | 739 | 20.1% |
| All | 936 | 2,160 | 30.2% |

Within the same search and holding the page features constant, a page at position 2 was 12.9 points less likely to be cited than the page at position 1, and a page at position 9 was 34.1 points less likely.

### On-page features, three ways

The table gives the difference in the chance of being cited, in percentage points, between pages with and without each feature. “1.0” is the version 1.0 estimate (cited against uncited within position bands). “Alone” compares pages on the same search at the same position, one feature at a time. “Together” does the same with all 14 features and word count entered at once.

| Feature | 1.0 | Alone | Together | Interval |
|---|---|---|---|---|
| Machine-readable date | +7.0 | +8.0 | +7.9 | 2.9 to 12.9 |
| VideoObject schema | +3.5 | +18.3 | +12.1 | 2.1 to 22.2 |
| Product schema | +2.9 | +15.3 | +10.1 | 1.9 to 18.2 |
| FAQPage schema | +6.0 | +7.5 | +5.9 | 1.2 to 10.5 |
| Question heading | +9.6 | +4.9 | +3.5 | −0.3 to 7.1 |
| BreadcrumbList schema | +3.6 | +4.5 | +1.3 | −2.3 to 4.8 |
| HTML table | +1.9 | +2.0 | +0.6 | −3.3 to 4.1 |
| Organization schema | +5.9 | +3.9 | +0.5 | −3.5 to 4.7 |
| Review schema | +4.5 | +8.3 | +0.4 | −6.3 to 6.6 |
| Article schema | +5.5 | +2.2 | −1.0 | −6.3 to 4.2 |
| Author signal | +4.8 | +2.1 | −2.1 | −7.4 to 2.9 |
| FAQ heading | +4.6 | +2.4 | −2.5 | −6.9 to 1.8 |
| Dated within 90 days | +1.8 | +1.7 | −4.0 | −8.7 to 0.8 |
| LocalBusiness schema | −2.9 | −2.6 | −6.2 | −11.9 to −0.5 |

The 95% interval is for the “Together” estimate. A question heading is an H2 or H3 phrased as a question; review schema includes ratings; LocalBusiness includes its subtypes. “Dated within 90 days” in the “Together” column is the difference beyond simply having a date.

Five of the 14 “Together” intervals exclude zero. After a Holm correction for testing 14 features, only the machine-readable date remains (adjusted p = 0.028). VideoObject and Product schema are the largest estimates but rest on few pages (6.6% and 9.3% of cited pages) and have wide intervals; neither survives the correction.

### Robustness

The same model was run two other ways. Resampling by website instead of by search, because some domains rank on many searches, widened every interval; the date association was the only one still clear of zero (+7.9, interval 1.3 to 13.4). Dropping the 612 pages fetched through Firecrawl left the date (+6.9) and gave question-form headings a narrow positive interval (+4.7, interval 0.6 to 8.7), which again did not survive the correction. It also made recency beyond having a date slightly negative (−5.6 points, interval −10.9 to −0.7).

### Negative results

Three features showed no meaningful association with citation. Their intervals lie inside ±5 points and they are nowhere near significance.

- Organization schema: +0.5 (interval −3.5 to 4.7)
- BreadcrumbList schema: +1.3 (interval −2.3 to 4.8)
- HTML table: +0.6 (interval −3.3 to 4.1)

Word count did not matter once position and search were held constant: no length quartile differed from the shortest by a clear margin. Recency beyond having a date did not help either; pages dated in the last 90 days were, if anything, 4.0 points less likely to be cited than other dated pages (interval −8.7 to 0.8).

## When a feature matters depends on the search

The keyword database labels each search by intent. Informational searches had cited pages more often (39.3% of their top-10 pages) than commercial and transactional ones (27.1%). The associations also differed.

| Feature | Informational | Commercial | Gap | Interval |
|---|---|---|---|
| Date | +17.8 | +1.0 | 16.8 | 8.1 to 25.8 |
| Article schema | +10.9 | +0.1 | 10.8 | 2.3 to 18.9 |
| Organization | +11.6 | +3.5 | 8.1 | −1.4 to 16.4 |
| Question heading | +9.3 | +10.7 | −1.4 | −10.0 to 7.1 |
| FAQPage | +3.1 | +7.8 | −4.6 | −12.1 to 2.5 |

Commercial includes transactional searches; the interval is for the gap. These use the version 1.0 band-adjusted method within each group (133 informational searches, 321 commercial or transactional). A dated article is a common shape for an informational answer and an unusual one for a product or service search, so the date and Article associations may mark the type of page that fits the question rather than the markup itself.

Position did not change the picture much. Adding a term for “feature on a top-3 page” to the model gave −1.5 points for question-form headings (interval −7.7 to 4.8) and +3.3 for dates (interval −3.3 to 9.3): no evidence that either matters more or less near the top.

## Stage 2: first source or one of many

The median AI Overview in this sample linked 8 distinct sources in its text. Among the 870 cited pages linked in the text, 183 (21.0%) were the first source linked. The median cited page was linked from 25.0% of the answer’s cited passages.

| Organic position | First source | Passage share |
|---|---|---|
| 1 to 3 | 27.5% | 31.3% |
| 4 to 6 | 14.9% | 29.1% |
| 7 to 10 | 15.7% | 32.1% |

Both columns are among cited pages: the share that were the first source, and the mean share of the answer’s passages that link them. A cited top-3 page was 1.8 times as likely as a cited page at 7 to 10 to be the first source, but once cited, lower-ranked pages were linked from as many passages. No page feature predicted being the first source after the correction. The largest estimate, a machine-readable date, pointed the other way (−11.8 points, interval −23.7 to 0.1).

## Stage 3: is the cited wording on the page

For each passage in an AI Overview that links a page, we measured the share of its content words (stop words removed) that appear anywhere in that page’s text. This is a lexical proxy for absorption. It shows whether the answer’s wording can be traced to the page, not whether the answer is faithful to it.

- On 800 cited pages from searches that also had uncited pages to compare, 63.8% of the content words in the passages citing a page were on that page. The same passages scored against the uncited top-10 pages of the same search gave 53.5%. The gap, 10.3 points (interval 8.2 to 12.2), shows the measure tracks the cited source rather than the topic alone.
- Across 842 scored cited pages, 27.0% of citing passages, on average, had at least 80% of their content words on the page. In 62.2% of cited pages, no citing passage reached that level.
- The median share of five-word sequences copied verbatim from the page was 0.0%. AI Overviews paraphrase; they rarely lift sentences.

One feature stood out after correction. Cited pages with an HTML table were 14.7 points more likely to have their citing passages absorbed (interval 6.5 to 22.9, adjusted p = 0.014), and were linked from 5.0 points more of the answer’s passages (interval 1.3 to 9.0, not significant after correction). Tables showed no association with being cited at all (+0.6). A machine-readable date went the other way, with less of the wording traceable to the page (−15.1 points, interval −25.9 to −4.2), but that did not survive the correction. A table does not seem to get a page chosen, but once chosen, an answer draws more of its wording from it, plausibly because tables hold the specific figures an answer repeats.

## How this compares with other studies

| Source | Design and date | Finding |
|---|---|---|
| This study, 1.1 | 3,096 top-10 pages, compared within the same search, September 2026 | Position first; only a machine-readable date survives correction, and only on informational searches |
| Aggarwal et al. (GEO) | Lab rewrites of sources already given to the engine, 2023 | Visibility up by as much as 40% when the source is in the engine’s context |
| Ahrefs | 1,885 pages that added schema against 4,000 controls, August 2025 to March 2026 | “Adding schema produced no major uplift in citations on any platform”; −4.6% on AI Overviews |
| Seer Interactive | 8,500 keywords, May 2026 | “FAQ + HowTo schema is irrelevant for AIO” |
| SE Ranking (reported by Search Engine Journal) | 216,524 pages, ChatGPT citations, November 2025 | Pages with FAQ schema averaged 3.6 citations against 4.2 without; question-style headings 3.4 against 4.3 |
| Ahrefs | 16.975M cited URLs, July 2025 | AI Overviews and organic results are the most likely to cite older pages |

Version 1.1 brings this study closer to the before-and-after evidence. Ahrefs’ test, the strongest design here, found that adding schema did not raise citations; our schema associations mostly shrink toward zero once pages are compared within the same search and with each other. The version 1.0 finding on question headings, which ran against SE Ranking’s ChatGPT data, largely came from comparing pages across different kinds of searches.

Sources: [Aggarwal et al., GEO](https://arxiv.org/abs/2311.09735); [Ahrefs, schema test](https://ahrefs.com/blog/schema-ai-citations/); [Seer Interactive](https://www.seerinteractive.com/insights/what-it-takes-to-rank-in-googles-ai-overviews-in-2026-is-not-what-you-think); [Search Engine Journal on SE Ranking](https://www.searchenginejournal.com/new-data-top-factors-influencing-chatgpt-citations/561954/); [Ahrefs, freshness](https://ahrefs.com/blog/do-ai-assistants-prefer-to-cite-fresh-content/).

## What this means

This section is interpretation. Because the study is observational, none of these points is a proven cause.

- **Ranking is still the way in.** Among pages that already rank, position explained more of which were cited than all 14 page features together, and most of the variation was explained by neither.
- **There is no markup shortcut.** After comparing like with like, schema types, author signals and FAQ sections showed no reliable association with citation. Adding them is cheap, but the data gives no reason to expect a lift.
- **Match the page to the kind of question.** On informational searches, dated article pages were cited far more often; on commercial searches they were not. The useful question is what shape of page answers this search, not which tags to add.
- **Being cited and being used are different.** Most cited pages were one source among about eight, and for most of them the answer’s wording could not be traced closely to the page. Specific, quotable content such as a table of figures went with more of the answer coming from the page.

## What a causal test would need

The associations above cannot say whether changing a page would change its citations. A test that could would change one thing at a time on real pages, with the rest held constant.

- Pairs of comparable pages, one left as it is and one given a single change (a date, a question heading, a table), assigned at random.
- The same searches run several times per date, and on several dates, because AI Overviews vary from run to run.
- Searches held out from the design, reworded searches, and more than one AI engine, to test whether an effect generalizes.
- All three stages measured, plus a check that the answer states what the page says.
- A version where competing pages are changed too, to tell an absolute gain from one that only holds while competitors stay the same.

We have not run such a test; this study provides the observational baseline for one.

## Methodology

- **Pages:** all top-10 organic results for the 486 US keywords (from our 800-keyword sample) that showed an AI Overview on 26 September 2026, excluding YouTube, Reddit, Facebook, Google, Instagram, TikTok, X, LinkedIn, Quora and Pinterest pages.
- **Cited:** the page’s normalized URL appears among the AI Overview’s references or links on the same results page.
- **Fetching:** direct HTTP with an identifying research user agent; Firecrawl for pages that returned an error or an empty page. Included if the result was HTML of at least 2,000 characters.
- **Features:** JSON-LD types parsed recursively; dates from JSON-LD dateModified, datePublished and uploadDate, article meta times and HTML time elements (newest valid date); question-form heading = an H2 or H3 ending in a question mark; word count on visible text.
- **Stage 1 model:** linear probability model of cited on the 14 features, organic-position dummies and word-count quartile, with search fixed effects, so each page is compared only with pages on the same results page. Features entered together and one at a time.
- **Stage 2:** first source = the first distinct URL linked in the AI Overview text; passage share = share of the answer’s cited passages that link the page.
- **Stage 3:** share of each citing passage’s content words (stop words removed, three or more letters) found in the page’s visible text; absorbed at 80% or more; baseline = the same passages against uncited top-10 pages of the same search.
- **Intent:** DataForSEO keyword-intent labels.
- **Inference:** 95% intervals from 1,000 bootstrap resamples of whole searches (seed 20260928; version 1.0 used seed 20260926); Holm-adjusted p-values across the 14 features; a negative result requires the interval inside ±5 points and an adjusted p above 0.05.
- **Version 1.0 method:** difference between cited and uncited shares within position bands 1 to 3, 4 to 6 and 7 to 10, averaged with band-size weights.
- **Update schedule:** quarterly.

## Limitations

- Correlation, not causation: the study measures how often pages with a trait are cited, not what happens when the trait is added.
- One results page per search on one day; run-to-run and week-to-week stability are not measured, so small estimates may not repeat.
- Retrieval is not observed: pages outside the top 10, and Google’s candidate set, are not seen.
- Absorption is a word-overlap proxy. It does not check that the answer is faithful to the page, and heavy paraphrase lowers it.
- Site authority, topic and content quality are not controlled beyond comparison within the same search and position.
- Intent groups rest on keyword-database labels; the navigational group (31 searches) was too small to analyze separately.
- Declared dates can be changed without changing content.
- Pages that errored, returned a 400+ status, or gave either fetcher too little HTML (14.6%) are excluded; many of these are behind bot protection.

## Data and downloads

- Every page with its features, position and citation status: [s10_pages.csv](https://underneath.agency/research-data/ai-overview-cited-pages-study/s10_pages.csv) and [JSON](https://underneath.agency/research-data/ai-overview-cited-pages-study/s10_pages.json)
- Every cited page with its prominence and absorption scores: [s10_v11_cited_pages.csv](https://underneath.agency/research-data/ai-overview-cited-pages-study/s10_v11_cited_pages.csv)
- Every statistic on this page, version 1.0 and 1.1: [stats.json](https://underneath.agency/research-data/ai-overview-cited-pages-study/stats.json)
- Machine-readable methodology: [methodology.json](https://underneath.agency/research-data/ai-overview-cited-pages-study/methodology.json)

The data is free to reuse with attribution (CC BY 4.0).

To cite: Underneath. (2026). *What pages cited by AI Overviews have in common: 3,096 pages* (Version 1.1). Underneath Research. https://underneath.agency/research/ai-overview-cited-pages-study

## Frequently asked questions

### Does schema markup help a page get cited in AI Overviews?

Not reliably. Compared with pages on the same search and with other features held constant, no schema type survived a correction for multiple testing, and Organization and BreadcrumbList schema showed no meaningful association at all. A controlled test by Ahrefs also found that adding schema did not raise citations.

### What matters most for being cited in an AI Overview?

Ranking position. 41.7% of pages in positions 1 to 3 were cited, against 20.1% in positions 7 to 10. Among on-page features, only a machine-readable date held up (+7.9 points), and only on informational searches.

### Do question-form headings help?

The evidence is weak. Version 1.0 found a +9.6 point difference, but comparing pages within the same search, with other features held constant, reduced it to +3.5 points with an interval that includes zero.

### Do AI Overviews prefer fresh content?

Having a machine-readable date went with citation (+7.9 points), mainly on informational searches. Being dated within the last 90 days added nothing beyond that (−4.0 points, interval −8.7 to 0.8).

### Does being cited mean the AI Overview used my page?

Not necessarily. Only 21.0% of cited pages were the first source, and in 62.2% of cited pages no passage citing the page had at least 80% of its content words on the page. Pages with tables had more of the answer’s wording traceable to them.

## Related research

- [Do AI Overviews cite the pages that rank?](https://underneath.agency/research/ai-overview-citations-study)
- [How often do AI Overviews appear?](https://underneath.agency/research/ai-overviews-frequency-study)
- [Do websites serve Markdown to AI agents?](https://underneath.agency/research/agent-readable-web-study)
- [AI Mode vs AI Overviews: how different are the sources?](https://underneath.agency/research/ai-mode-vs-ai-overviews-study)

## Related guides

- [What on-page signals are linked to citations in Google AI Overviews and Perplexity?](https://underneath.agency/resources/on-page-signals-linked-to-ai-citations)
- [Can restructuring existing content, without changing what it says, increase AI citations?](https://underneath.agency/resources/does-content-structure-increase-ai-citations)
- [Do FAQ pages help you get cited by AI search engines?](https://underneath.agency/resources/do-faq-pages-help-ai-citations)
- [What does content that AI engines prefer look like?](https://underneath.agency/resources/what-content-do-ai-engines-prefer)
- [What decides whether an AI engine cites my page over a competitor’s?](https://underneath.agency/resources/why-ai-cites-competitor-page-first)

---

This is the Markdown twin of https://underneath.agency/research/ai-overview-cited-pages-study. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "The YouTube videos Google’s AI cites are small | Underneath"
description: "Within the same search, the YouTube videos Google’s AI cited came from smaller channels than the ones it showed and skipped. 800 US searches."
canonical: "https://underneath.agency/research/ai-overview-youtube-videos-study"
published: 2026-09-26
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Research · AI search

# The YouTube videos Google’s AI cites are small

YouTube is the domain Google’s AI Overviews cite most: 63.6% of the AI Overviews in this sample cited at least one YouTube video. Version 1.0 of this study described the 469 videos that Google’s AI Overviews and AI Mode cited for 800 US searches on 26 September 2026. Most had modest views and came from modest channels. That description could not say what separates a cited video from one that is not cited. Version 1.1 adds a comparison group: for the same searches, every YouTube video Google showed on the first results page, cited or not. It also splits visibility into stages, from appearing on the results page to being cited and quoted.

## The short version

1. The AI Overview mostly cites videos that are not on the results page. Of 629 YouTube citations in AI Overviews, only 15.1% (95% interval 10.1% to 19.2%) were videos Google also showed on page one of the same search. Those that were not on the page were much smaller: a median of 2,776 views and 21,400 subscribers, against 34,100 views and 77,000 subscribers for the ones that were.
2. Within the same search, being popular did not help. Where the AI Overview cited one video Google showed and skipped another, the cited video had more views in 39.7% of pairs and a bigger channel in 37.0% (50% would mean no difference). Holding the other measures fixed, a bigger channel lowered the odds of citation (odds ratio 0.72 per standard deviation, 95% interval 0.54 to 0.97) and views made no difference (1.05, 0.78 to 1.42).
3. Shorter and older videos were favored. Holding the other measures fixed, longer videos were less likely to be cited (odds ratio 0.7 per standard deviation) and older ones more likely (1.29). Cited videos had a median length of 266 seconds, against 398 for the shown videos that were not cited.
4. The text Google shows for a cited video comes from what is said in it. For 97.9% of 616 cited videos with a text excerpt, the excerpt was not in the video’s description. 97.3% of those excerpts began in lowercase, like YouTube’s automatic captions, and 99.4% of the videos we could check had captions.
5. The results-page feature matters, the slot does not. The AI Overview cited 29.7% of videos in Google’s video pack, 15.6% of YouTube organic results, 12.1% of short videos and none of the 111 videos in the “What people are saying” row. Within the video pack, the first, second and third slots were cited at similar rates (30.9%, 29.3%, 31.6%).
6. The version 1.0 description still holds: the median cited video had 4,283 views and the median channel 34,900 subscribers.

## Research questions

| Question | Answered here? |
|---|---|
| RQ1. Across all 800 searches, how often is there an AI Overview, a YouTube video on page one, and a cited YouTube video? | Yes |
| RQ2. Given that Google shows a video on page one, how likely is the AI Overview to cite it? | Yes, by results-page feature and slot |
| RQ3. Within the same search, how do cited videos differ from shown videos that were not cited? | Yes, for views, channel size, age, length, word match, description and captions |
| RQ4. Does the pattern hold in AI Mode and across industries? | Partly: AI Mode cites fewer videos, and industries differ |
| RQ5. Does the answer use what the video says, and represent it correctly? | No: not measured |

## Visibility in stages

Being “cited by Google’s AI” is one stage of several. We measure four; three more are named because they matter, but this study does not measure them.

| Stage | Question | Measured? |
|---|---|---|
| AI Overview shown | Does the search show an AI Overview? | Yes |
| Shown on the results page | Does Google show the video on page one? | Yes, the only retrieval step we can see |
| Cited | Does the AI Overview cite the video? | Yes |
| Prominence | Where in the AI Overview’s sources does it sit? | Yes |
| Use | Did the answer take anything from the video? | No |
| Fidelity | Did the answer represent the video correctly? | No |
| User outcome | Did anyone watch or act on it? | No |

## What we measured

We used the 800 US keywords of our [AI Overview frequency study](https://underneath.agency/research/ai-overviews-frequency-study), 100 per industry across eight industries and 32 seed topics, and the Google results pages collected for them on 26 September 2026 (Google.com, United States, English, desktop), plus AI Mode answers for 400 of the same keywords. On each results page we recorded every YouTube video Google showed on page one (organic results, the video pack, the short-videos row and the “What people are saying” row) and every YouTube video the AI Overview cited. Version 1.1 made no new searches.

Because views and subscriber counts change daily, we re-read both the cited and the shown videos from YouTube on the same day, 28 September 2026, so the two groups are measured alike: 914 videos in all. The version 1.0 figures further down use the counts read on 26 September.

Every search stays in the results, including those with no AI Overview. Keywords that share a seed topic are not independent, so every 95% interval comes from resampling the 32 seed topics (2,000 resamples).

## Findings

### Every search, by outcome

| Outcome (800 searches) | Searches | Share |
|---|---|---|
| No AI Overview, no YouTube on page one | 233 | 29.1% |
| AI Overview cites a YouTube video, no YouTube on page one | 187 | 23.4% |
| AI Overview cites no YouTube video, no YouTube on page one | 135 | 16.9% |
| AI Overview cites a YouTube video, YouTube shown on page one | 122 | 15.2% |
| No AI Overview, YouTube shown on page one | 81 | 10.1% |
| AI Overview cites no YouTube video, YouTube shown on page one | 42 | 5.2% |

An AI Overview appeared on 60.8% of searches (95% interval 50.1% to 71.5%) and Google showed a YouTube video on page one of 30.6% (21.0% to 40.4%). The AI Overview cited a YouTube video on 38.6% of all searches. When YouTube was on page one, 74.4% of the AI Overviews cited a YouTube video (61.8% to 83.6%). When it was not, 58.1% still did (46.4% to 68.1%).

### Most cited videos are not on the results page

| Where the cited video was on the page | Share of 629 |
|---|---|
| In the video pack | 12.6% |
| As an organic result | 1.9% |
| In the short-videos row | 0.6% |
| In “What people are saying” | 0.0% |
| Not on page one | 84.9% |

The videos the AI Overview found beyond the results page were far smaller than the ones it took from the page (medians):

| Cited videos | Videos | Views | Subscribers | Age (days) |
|---|---|---|---|---|
| Shown on page one | 95 | 34,100 | 77,000 | 352 |
| Not on page one | 534 | 2,776 | 21,400 | 329 |

For the most part, the AI Overview does not choose from the videos Google ranks for the search. It reaches much deeper into YouTube, which is why most cited videos are small.

### Shown on page one, then cited

On searches with an AI Overview, Google showed 487 YouTube videos on page one, and the AI Overview cited 19.5% of them (95% interval 14.7% to 26.4%).

| Page feature | Shown | Cited | 95% interval |
|---|---|---|---|
| Video pack | 266 | 29.7% | 22.5% to 37.3% |
| Organic result | 77 | 15.6% | 5.3% to 31.1% |
| Short videos | 33 | 12.1% | 0.0% to 25.0% |
| What people are saying | 111 | 0.0% | 0.0% to 0.0% |

Within the video pack, slot made no clear difference: 30.9% of first-slot videos were cited, 29.3% of second and 31.6% of third.

Cited videos rarely lead the list. Of the 629 citations, 3.2% were the AI Overview’s first source and 18.9% among its first three; the median cited video was sixth of a median nine sources.

### Cited against not cited, within the same search

This is the comparison version 1.0 could not make. We took the 111 searches where the AI Overview cited at least one video Google showed on page one and skipped at least one other (23 seed topics, 261 cited and 283 skipped videos). For each search we compared every cited video with every skipped one.

| Measure | Cited higher | 95% interval |
|---|---|---|
| Views | 39.7% | 33.3% to 45.9% |
| Channel subscribers | 37.0% | 30.9% to 44.2% |
| Age | 56.2% | 46.4% to 63.3% |
| Length | 41.6% | 29.7% to 53.7% |
| Title matches the search words | 48.5% | 44.4% to 51.6% |
| Description matches the search words | 49.6% | 42.4% to 55.9% |
| Description length | 42.9% | 33.8% to 53.3% |
| Has chapters | 47.3% | 39.7% to 56.7% |
| Has uploaded (not automatic) English captions | 54.7% | 50.4% to 58.6% |

“Cited higher” is the share of cited-against-skipped pairs in which the cited video had the higher value; 50% means no difference, and ties count as half.

| Same 111 searches | Cited | Skipped |
|---|---|---|
| Views | 12,066 | 15,301 |
| Channel subscribers | 50,050 | 84,500 |
| Age (days) | 399 | 256 |
| Length (seconds) | 266 | 398 |
| Mean share of search words in the title | 72.1% | 74.6% |
| Uploaded English captions | 21.0% | 12.1% |

Views and channel size move together, so we also fitted a model that compares cited with skipped videos search by search (a conditional logit, 535 videos) with all the measures at once. Odds ratios are per standard deviation.

| Measure | Odds ratio | 95% interval | p |
|---|---|---|---|
| Views | 1.05 | 0.78 to 1.42 | 0.736 |
| Channel subscribers | 0.72 | 0.54 to 0.97 | 0.028 |
| Age | 1.29 | 1.04 to 1.61 | 0.021 |
| Length | 0.7 | 0.54 to 0.91 | 0.007 |
| Title matches the search words | 0.77 | 0.59 to 1.01 | 0.063 |
| Description matches the search words | 1.09 | 0.83 to 1.42 | 0.552 |
| Description length | 1.14 | 0.87 to 1.51 | 0.337 |

Once channel size is held fixed, views no longer matter. A bigger channel, a longer video and a newer video each made citation less likely. Matching the search words in the title or description made no difference, perhaps because almost every video Google shows already matches them. Nearly every video had captions of some kind (100.0% of cited and 97.1% of skipped videos we could check), so captions cannot separate them.

### What Google quotes from a cited video

Google shows a short text excerpt under most cited videos. We checked whether the first 40 characters of that excerpt, after the title, appear in the video’s description.

| Excerpt | Share of 616 |
|---|---|
| Found in the description | 2.1% |
| Not in the description | 97.9% |
| Not in the description and starting in lowercase | 97.3% of those |
| Not in the description, video has captions | 99.4% of those we could check |

The excerpts read like spoken words (“so you’re thinking about a tankless water heater…”). They look like YouTube’s automatic captions: what Google’s AI quotes is what is said in the video, not what is written under it. We did not download transcripts, so this is an inference from the excerpts.

### AI Mode

AI Mode cited a YouTube video in 21.8% of its 400 searches (95% interval 12.8% to 32.2%), 173 citations in all. Only 8.1% of those were videos Google showed on page one, and AI Mode cited 4.6% of the videos shown there. 22.5% of AI Mode’s YouTube citations were also cited by the AI Overview for the same search. AI Mode relies even less on the results page than the AI Overview does.

### By industry

| Industry | Cites YouTube | YouTube on page one | Shown videos cited |
|---|---|---|---|
| B2B software and technology | 91.0% | 37.0% | 41.7% |
| Financial services and insurance | 49.0% | 8.0% | 21.4% |
| Retail and ecommerce | 45.0% | 67.0% | 15.0% |
| Hospitality and travel | 42.0% | 24.0% | 9.4% |
| Home and local services | 36.0% | 58.0% | 16.5% |
| Healthcare and dental | 21.0% | 20.0% | 17.9% |
| Legal and professional services | 14.0% | 9.0% | 0.0% |
| Franchises and multi-location brands | 11.0% | 22.0% | 13.0% |

The first two columns are shares of the industry’s 100 searches; the last is the share of page-one videos the AI Overview cited. In B2B software, the AI Overview cited a YouTube video on 91.0% of searches, far more often than Google showed one on page one. In retail and home services, Google showed videos often but the AI Overview cited few of them.

## The cited videos (version 1.0)

These figures describe every distinct video cited in AI Overviews and AI Mode, 469 in all, with counts read on 26 September.

### Views

| Views | Share of cited videos |
|---|---|
| Under 1,000 | 26.1% |
| Under 10,000 | 60.9% |
| 1 million or more | 4.1% |

A quarter of cited videos had 856 views or fewer; three quarters had 35,450 or fewer.

### Channel size

| Channel subscribers | Share of cited videos |
|---|---|
| Under 10,000 | 32.0% |
| Under 100,000 | 68.5% |
| 1 million or more | 6.5% |

The channels cited most often were specialists: OneHourProfessor, Gauging Gadgets and Brennan Valeski had five cited videos each; HR Toolbox, CRM Central, TechnologyAdvice, Eating on a Dime and a dental implant clinic’s channel had four.

### Age and length

| Measure | Result |
|---|---|
| Median age | 377 days |
| Published within the last year | 48.9% |
| Older than 3 years | 15.6% |
| Older than 5 years | 6.8% |
| Median length | 6 minutes 6 seconds |
| One minute or shorter | 11.1% |
| Longer than 20 minutes | 9.0% |

### Category and industry

YouTube classed the cited videos mainly as Education (129), People and Blogs (99), Howto and Style (96) and Science and Technology (60). Median views varied by industry: 21,066 for retail and 16,487 for franchise brands, but 1,403 for healthcare and 921 for legal services.

| Industry | Median views of cited videos |
|---|---|
| Retail and ecommerce | 21,066 |
| Franchises and multi-location brands | 16,487 |
| Home and local services | 7,829 |
| B2B software and technology | 3,794 |
| Hospitality and travel | 2,932 |
| Financial services and insurance | 1,532 |
| Healthcare and dental | 1,403 |
| Legal and professional services | 921 |

## Observed, inferred and unknown

**What we observe.** Google’s AI Overviews cite YouTube on most searches where they appear, but mostly videos that Google does not show on the results page. Where they choose among the videos Google does show, they pick the smaller channel, the shorter video and the older video more often, and views make no difference once channel size is held fixed. The text Google shows for a cited video almost always comes from what is said in it.

**What we infer.** Popularity is not what gets a video cited. The AI appears to search YouTube for the spoken answer to the question and cite the video that gives it, which favors short, focused videos from specialist channels.

**What remains unknown.** Why a given video is picked: we see the results page but not the AI’s own search, so 84.9% of cited videos came from a step we cannot observe. Whether the answer actually uses what the video says, and uses it correctly. Whether the pattern holds from one run to the next. Transcript content was not analyzed.

## What this means

The points below are our interpretation. They follow from the findings but were not tested.

- **You do not need an audience to be cited.** Within the same search, bigger channels were less likely to be cited, and views made no difference.
- **Say the answer out loud, early.** Google’s AI quotes what is spoken in the video. A video that states the answer to a common question plainly gives it something to quote.
- **Keep it short and on one question.** Cited videos were shorter than the videos Google showed and skipped for the same search.
- **Ranking in the video pack is not the route in.** Most cited videos were nowhere on the results page, and the slot in the video pack made no difference.
- **Check your captions.** Automatic captions are what the AI seems to read. Correcting them costs little; uploaded captions were slightly more common among cited videos, though the difference is small.

## Where this sits in GEO research

Research on generative engine optimization began by asking whether changing a page changes how often an AI answer uses it. Later work compared the sources different AI search systems draw on, and argued that visibility should be measured in stages (retrieval, citation, prominence, use, fidelity) rather than as one score. Most of that work studies web pages. This study applies the stage view to video. It is observational: it compares cited videos with videos Google showed for the same search, not with every video the AI could have found, and it does not test whether changing a video changes its chances. Measuring whether cited videos shape the answer, and how accurately, is the next step.

## Methodology

- **Searches:** the 800 US keywords of our AI Overview frequency study, 100 per industry, 32 seed topics; Google results pages with AI Overviews and, for 400 of the keywords, AI Mode answers (DataForSEO SERP API, Google.com, United States, English, desktop, 26 September 2026, one request each).
- **Shown on page one:** YouTube videos in organic results, the video pack, the short-videos row and the “What people are saying” row, each counted once per search at its first appearance.
- **Cited:** YouTube videos among the AI Overview’s sources and in-text links, in reading order, each counted once per search.
- **Video data:** version 1.1 re-read all 914 cited and shown videos on 28 September 2026, from the watch page or, when YouTube limited our requests, from YouTube’s public player data (views, date, length, description, channel size; caption data only from watch pages). Version 1.0 read the 469 cited videos on 26 September.
- **Word match:** the share of the search’s words (lowercased, common words removed) that appear in the title or description.
- **Within-search comparison:** for each search, the share of cited-against-skipped pairs in which the cited video has the higher value, averaged over searches; and a conditional logit with one group per search and standardized measures.
- **Uncertainty:** 95% bootstrap intervals resampling the 32 seed topics (2,000 resamples, seed 20260926); model intervals are the model’s own.
- **Update schedule:** quarterly.

## Limitations

- The comparison group is the videos Google showed on page one, not every video the AI could have found. 84.9% of cited videos were not on the page, so the within-search comparison covers only the step from shown to cited.
- Observational, one run per search on one date. The study shows which videos were cited, not why, and run-to-run variation is not measured.
- Citation is not use. Whether the answer took anything from the video, and represented it correctly, is not measured; nor are views or clicks that came from the citation.
- View and subscriber counts are lifetime figures read two days after the searches; YouTube rounds subscriber counts.
- Word match is a rough measure of relevance, and transcripts were not analyzed. The within-search comparison covers 111 searches in 23 seed topics, so its intervals are wide.

## What changed in version 1.1

Version 1.0 (26 September 2026) described the 469 cited videos. It said plainly that it could not compare them with videos that were not cited. Following an external review, version 1.1 (28 September 2026) adds that comparison using the same searches, with no new search queries: every YouTube video Google showed on page one serves as the comparison group. It also splits visibility into stages, counts every search as an outcome, adds prominence, AI Mode and industry breakdowns, and adds intervals clustered by seed topic. No version 1.0 figure changed; they appear above as the description of the cited videos. The new comparisons use counts re-read on 28 September for both groups.

## Data and downloads

- Every cited or shown video for every search, version 1.1 (results-page feature, slot, cited or not, views, channel size, age, length, captions, word match): [s20_videos_v11.csv](https://underneath.agency/research-data/ai-overview-youtube-videos-study/s20_videos_v11.csv) and [JSON](https://underneath.agency/research-data/ai-overview-youtube-videos-study/s20_videos_v11.json)
- Every cited video, version 1.0: [s20_youtube_videos.csv](https://underneath.agency/research-data/ai-overview-youtube-videos-study/s20_youtube_videos.csv) and [JSON](https://underneath.agency/research-data/ai-overview-youtube-videos-study/s20_youtube_videos.json)
- Every statistic on this page: [stats.json](https://underneath.agency/research-data/ai-overview-youtube-videos-study/stats.json); version 1.0: [stats_v1.0.json](https://underneath.agency/research-data/ai-overview-youtube-videos-study/stats_v1.0.json)
- Machine-readable methodology: [methodology.json](https://underneath.agency/research-data/ai-overview-youtube-videos-study/methodology.json)
- Chart: [shown videos cited, by results-page feature](https://underneath.agency/research-data/ai-overview-youtube-videos-study/youtube-cited-given-shown.svg)

The data is free to reuse with attribution (CC BY 4.0).

To cite: Underneath. (2026). *The YouTube videos Google’s AI cites are small* (Version 1.1). Underneath Research. https://underneath.agency/research/ai-overview-youtube-videos-study

## Frequently asked questions

### Do you need a big YouTube channel to be cited in Google AI Overviews?

No. Among videos Google showed for the same search, the ones from bigger channels were less likely to be cited, and views made no difference once channel size was held fixed. The median cited video came from a channel with 34,900 subscribers.

### Does ranking in Google’s video pack get a video cited in AI Overviews?

Sometimes: the AI Overview cited 29.7% of video-pack videos, whichever slot they were in. But 84.9% of the YouTube videos AI Overviews cited were not on the results page at all.

### What does Google’s AI read from a YouTube video?

Apparently what is said in it. For 97.9% of cited videos, the text Google showed was not in the description, and it read like YouTube’s automatic captions.

### How long are the YouTube videos AI Overviews cite?

Short to mid-length: a median of 266 seconds on the searches where we could compare, against 398 seconds for the videos Google showed but did not cite.

## Related research

- [Do AI Overviews cite the pages that rank?](https://underneath.agency/research/ai-overview-citations-study)
- [Which Reddit threads do AI answers cite?](https://underneath.agency/research/ai-reddit-citations-study)
- [AI Mode vs AI Overviews: how different are the sources?](https://underneath.agency/research/ai-mode-vs-ai-overviews-study)
- [How often do AI Overviews appear?](https://underneath.agency/research/ai-overviews-frequency-study)

## Related guides

- [Can smaller, lower-traffic websites get cited by AI search engines?](https://underneath.agency/resources/can-small-websites-get-cited-by-ai)
- [How do application security vendors win demand when developers and CISOs ask AI?](https://underneath.agency/resources/application-security-demand-ai-search)
- [Will AI steer a plant’s next machine purchase toward us or a rival?](https://underneath.agency/resources/machinery-companies-buyers-ai-search)
- [Is AI search deciding which HR software gets the demo?](https://underneath.agency/resources/hr-software-ai-search)
- [Does social media content help your brand show up in AI search?](https://underneath.agency/resources/does-social-media-help-ai-search-visibility)

---

This is the Markdown twin of https://underneath.agency/research/ai-overview-youtube-videos-study. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "When does Google show an AI Overview? 1,248 US searches"
description: "1,248 US searches: AI Overviews followed the search more than the industry. Local packs cut the rate to 30.7%; rewording a topic lifted it from 59.4% to 93.8%."
canonical: "https://underneath.agency/research/ai-overviews-frequency-study"
published: 2026-09-26
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Research · AI search

# When does Google show an AI Overview? 1,248 US searches

Before a page can be cited in an AI Overview, Google has to show one. This study asks when that happens. We ran 960 US Google searches on 26 September 2026 across the eight industries Underneath works in, then 288 more on 28 September to test the same topics in different wordings. Version 1.1 no longer treats the share of keywords with an AI Overview as the finding. It models which features of a search go with an AI Overview, checks how much of the industry gap is really a difference in the kind of searches each industry has, and then follows the AI Overviews that did appear: which sources they cite, whether those sources rank, whether the cited page supports the sentence, and whether the result holds two days later.

## The short version

1. There is no single AI Overview rate. In our sample of 800 commercial keywords, 60.8% showed one (95% interval 50.1% to 71.5%), but only 40.8% when weighted by search volume, 53.8% among the highest-volume keywords and 98.1% among question-form keywords. A 2026 study of 55,393 trending queries found 13.7%. The figure depends on which searches are counted.
2. The local pack (the map with three businesses) is the strongest signal. Holding industry, intent, “near me”, length and volume constant, a search with a local pack had a 30.7% predicted chance of an AI Overview against 78.6% without one (odds ratio 0.07). The gap appeared in every industry that had local packs.
3. Industry still matters, but less than the raw table suggests. The observed range between industries was 60.0 points; after adjusting for query mix it was 36.7 points. B2B software (83.0% adjusted) and financial services (77.7%) stay highest.
4. Wording changes the outcome for the same topic. We searched 96 keywords three ways: as written, 59.4% showed an AI Overview; as a long non-question form, 87.5%; as a natural question, 93.8%. Most of the lift came with the longer, more specific wording. The question form added 6.3 points over the long form, which was not statistically significant.
5. Once an AI Overview appears, most of what it cites does not rank on page one. It cited a median of 8 sources; 30.2% of cited URLs (44.3% of cited domains) were in the organic top 10 of the same search, and 20.2% of AI Overviews cited no top-10 URL at all.
6. A citation is not always support. In a model-coded pilot of 97 cited sentences, both coders agreed on a judgment for 71: 66.2% were supported by the cited page, 21.1% partly and 12.7% not.
7. The outcome is fairly stable over two days but the sources are not. Re-searching 96 keywords two days later, 83.3% had the same outcome (16 flipped); where both days showed an AI Overview, the median overlap of their cited URLs was 0.43 (Jaccard).

## Research questions

| Question | Answered here? |
|---|---|
| RQ1. Which features of a search go with an AI Overview, and does industry matter once they are held constant? | Yes, with a clustered logistic model |
| RQ2. For the same topic, does the wording change activation? | Yes, with a matched experiment (96 topics, three wordings) |
| RQ3. How much does the headline rate depend on what is counted? | Yes |
| RQ4. Which sources are cited, and do they rank on page one? | Yes |
| RQ5. Does the cited page support the sentence it is attached to? | Pilot only (97 sentences, model-coded) |
| RQ6. Does the same keyword get the same outcome on another day? | Partly: two dates, 96 keywords |

## Visibility in stages

Being visible in Google’s AI answers is not one number. This study measures the first stages separately and names the later ones it does not measure.

| Stage | Question | Measured? |
|---|---|---|
| Activation | Does Google show an AI Overview for this search? | Yes, 1,248 searches |
| Source selection | Which pages are cited, and do they rank in the top 10? | Yes, 643 AI Overviews |
| Grounding | Does the cited page support the sentence? | Pilot, 97 sentences |
| Stability | Is the outcome the same on another day? | Two dates, 96 keywords |
| Retrieval | Which pages did the system consider? | No: only the organic top 10 is visible |
| Influence | Did a source shape the answer? | No |
| User outcome | Did anyone click, trust or act on it? | No |

## What we measured

**Main sample.** From DataForSEO’s US keyword database we took four seed topics per industry (for example “plumber”, “dental implants”, “personal injury lawyer”, “crm software”), 32 in all. For each industry we kept the 40 highest-volume suggestions and 60 more drawn at random from the rest, 800 keywords, plus a separate sample of 160 question-form keywords (20 per industry). Each was searched once on Google.com, United States, English, desktop, on 26 September 2026, with the AI Overview allowed to load. The sample leans toward popular commercial searches (median 165,000 monthly US searches), so it estimates AI Overview exposure among the keywords businesses in these industries track, not across all Google searches.

**Matched wordings.** We drew 96 of the 800 keywords at random (12 per industry, excluding keywords that were already questions or contained “near me”). For each, we wrote a natural question and a long non-question form on the same topic, for example “dentist crowns” → “how much does a dental crown cost” → “dental crown cost with and without insurance”. All three were searched on 28 September 2026 with the same settings. The keyword as written doubles as a second date for those 96 keywords.

**Grounding pilot.** From the AI Overviews of 26 September we drew 120 (15 per industry), picked one sentence with a citation from each, fetched the first page cited for it and kept the passages that shared the most words with the sentence. Two AI models, working independently, judged whether those passages support the sentence.

Keywords that share a seed topic are not independent, so the 95% intervals and the model’s standard errors are computed by seed topic, not by keyword.

## Findings

### How often, by industry and type of search

| Industry | Keywords | With an AI Overview | 95% interval | Question-form keywords with one |
|---|---|---|---|---|
| B2B software and technology | 100 | 96.0% | 85.9% to 100.0% | 90.0% |
| Financial services and insurance | 100 | 88.0% | 79.5% to 92.5% | 100.0% |
| Retail and ecommerce | 100 | 68.0% | 39.5% to 85.5% | 95.0% |
| Hospitality and travel | 100 | 62.0% | 41.4% to 90.9% | 100.0% |
| Home and local services | 100 | 53.0% | 34.0% to 69.8% | 100.0% |
| Healthcare and dental | 100 | 43.0% | 26.9% to 75.8% | 100.0% |
| Legal and professional services | 100 | 40.0% | 25.9% to 56.6% | 100.0% |
| Franchises and multi-location brands | 100 | 36.0% | 18.6% to 64.6% | 100.0% |
| **All industries** | **800** | **60.8%** | **50.1% to 71.5%** | **98.1%** |

Question-form column: 20 keywords per industry. The wide intervals show that within an industry, the result depends heavily on the seed topic.

| Search type | Keywords | With an AI Overview |
|---|---|---|
| Question-form (separate sample) | 160 | 98.1% |
| Informational intent | 181 | 73.5% |
| Commercial intent | 366 | 65.0% |
| Transactional intent | 165 | 50.3% |
| Navigational intent | 88 | 36.4% |
| Results page has a local pack | 275 | 21.8% |
| Contains “near me” | 75 | 17.3% |
| Contains a price or cost word | 31 | 100.0% |

Intent labels are DataForSEO’s classification. Every one of the 31 keywords with a price or cost word (“cost”, “price”, “fees”, “how much”) showed an AI Overview.

### What goes with an AI Overview when everything is held constant

A logistic regression on all 800 keywords, with standard errors clustered by the 32 seed topics:

| Feature | Odds ratio | 95% interval | Change in probability |
|---|---|---|---|
| Results page shows a local pack | 0.07 | 0.04 to 0.13 | −36.4 points |
| Contains “near me” | 0.27 | 0.12 to 0.62 | −18.0 points |
| Each extra word in the search | 1.31 | 1.05 to 1.64 | +3.8 points |
| Monthly volume (per tenfold increase) | 0.89 | 0.63 to 1.26 | −1.6 points |
| Transactional intent (vs commercial) | 0.44 | 0.22 to 0.88 | −11.2 points |
| Informational intent (vs commercial) | 0.61 | 0.30 to 1.24 | −6.8 points |
| Navigational intent (vs commercial) | 0.62 | 0.29 to 1.34 | −6.5 points |
| B2B software (vs home and local services) | 5.17 | 1.14 to 23.44 | +22.7 points |
| Financial services (vs home and local services) | 3.18 | 1.30 to 7.81 | +16.0 points |

Change in probability is the average marginal effect. Industry as a whole remained significant (Wald test, p < 0.001); intent as a whole did not (p = 0.14). Once the other features are held constant, search volume has no clear effect, and neither does whether a keyword came from the highest-volume group or the random draw (odds ratio 0.8, interval 0.41 to 1.56). The effect of length did not differ by intent (p = 0.54).

How much each set of information predicts the outcome, measured by cross-validated AUC (0.5 is a coin toss, 1.0 is perfect; folds split by seed topic):

| Information used | AUC |
|---|---|
| Intent label only | 0.581 |
| Industry only | 0.586 |
| The words of the query (length, “near me”, price word, “best” or “top”) | 0.635 |
| Intent and the words of the query | 0.669 |
| Plus whether the results page has a local pack | 0.832 |
| Everything, including industry and volume | 0.845 |

The query text on its own predicts the outcome only modestly. Knowing whether Google also shows a local pack is what lifts the prediction, which suggests that Google first decides what kind of results page a search deserves, and a map-led page rarely gets an AI Overview.

### Industry, before and after adjusting for query mix

| Industry | Observed | Adjusted | Local pack on the page | Navigational or transactional intent |
|---|---|---|---|---|
| B2B software and technology | 96.0% | 83.0% | 2.0% | 0.0% |
| Financial services and insurance | 88.0% | 77.7% | 18.0% | 6.0% |
| Retail and ecommerce | 68.0% | 64.8% | 20.0% | 60.0% |
| Home and local services | 53.0% | 62.2% | 49.0% | 51.0% |
| Healthcare and dental | 43.0% | 60.5% | 60.0% | 39.0% |
| Legal and professional services | 40.0% | 54.6% | 70.0% | 3.0% |
| Franchises and multi-location brands | 36.0% | 51.8% | 56.0% | 52.0% |
| Hospitality and travel | 62.0% | 46.3% | 0.0% | 42.0% |

Adjusted: the average predicted rate if every keyword in the sample belonged to that industry, holding the other features as they are. The observed range of 60.0 points shrinks to 36.7 points: about 39% of the industry gap is explained by the kind of searches each industry has. Healthcare, legal and franchise keywords show few AI Overviews largely because their results pages are dominated by local packs. Travel moves the other way: it had no local packs at all, so once that is accounted for, its rate is the lowest.

### Local searches are still the map’s territory

| Industry | Local pack: keywords | Local pack: with an AI Overview | No local pack: keywords | No local pack: with an AI Overview |
|---|---|---|---|---|
| Financial services and insurance | 18 | 55.6% | 82 | 95.1% |
| Retail and ecommerce | 20 | 30.0% | 80 | 77.5% |
| Home and local services | 49 | 28.6% | 51 | 76.5% |
| Legal and professional services | 70 | 24.3% | 30 | 76.7% |
| Healthcare and dental | 60 | 13.3% | 40 | 87.5% |
| Franchises and multi-location brands | 56 | 7.1% | 44 | 72.7% |

The gap appears in all six industries with enough local packs to compare, and it does not differ significantly between them (p = 0.099). Across all 800 keywords, 21.8% of searches with a local pack showed an AI Overview against 81.1% without one, 3.7 times as often. After adjustment the predicted rates are 30.7% and 78.6%. The local pack is itself Google’s choice, so this shows that the two rarely appear together, not that one suppresses the other.

### Same topic, three wordings

| Wording (96 topics, 28 September) | Median words | With an AI Overview | With a local pack |
|---|---|---|---|
| Keyword as written | 4 | 59.4% | 35.4% |
| Long non-question form | 9 | 87.5% | 11.5% |
| Natural question | 8 | 93.8% | 9.4% |

Pair by pair, the question form gained an AI Overview on 34 topics and lost one on 1 (exact McNemar test, p < 0.001); the long form gained 28 and lost 1 (p < 0.001). Between the question and the long form, the question gained 10 and the long form 4 (p = 0.18). Of the 39 topics with no AI Overview as written, 87.2% got one as a question and 71.8% as a long form. Of the 34 whose keyword showed a local pack, 88.2% got an AI Overview when asked as a question.

Rewording did two things at once: it made the search longer and more specific, and it mostly removed the local pack. Question form and length move together too closely in this design to separate them cleanly (in a model with both, word count was no longer significant). What the experiment does show is that, for the same topic, a specific, sentence-like search is far more likely to get an AI Overview than a short keyword.

| Industry (12 topics each) | As written | Long form | Question |
|---|---|---|---|
| B2B software and technology | 91.7% | 91.7% | 100.0% |
| Financial services and insurance | 91.7% | 100.0% | 100.0% |
| Hospitality and travel | 66.7% | 83.3% | 91.7% |
| Healthcare and dental | 58.3% | 91.7% | 100.0% |
| Retail and ecommerce | 50.0% | 91.7% | 100.0% |
| Home and local services | 41.7% | 66.7% | 75.0% |
| Franchises and multi-location brands | 41.7% | 100.0% | 83.3% |
| Legal and professional services | 33.3% | 75.0% | 100.0% |

In the version 1.0 question sample, which was drawn separately from question-form suggestions, questions beat the main sample within 25 of the 30 seed topics that had both, by a mean 31.8 points. Those questions were also longer (median 6 words against 4). The matched design above is the fairer comparison.

### How much the denominator matters

| What is counted | With an AI Overview |
|---|---|
| Question-form keywords (160) | 98.1% |
| Keywords without a local pack (525) | 81.1% |
| Randomly drawn keywords (480) | 65.4% |
| All 800 keywords | 60.8% |
| Highest-volume keywords (320) | 53.8% |
| All 800, weighted by search volume | 40.8% |
| Keywords with a local pack (275) | 21.8% |

The volume-weighted figure is 20.0 points lower because a few enormous searches dominate the total: the 10 highest-volume keywords make up 38.4% of all volume, and only 2 of them showed an AI Overview. By volume fifth, the rate fell from 70.5% in the smallest fifth to 50.5% and 56.7% in the two largest, but once local packs and intent are held constant, volume has no clear effect: the biggest searches are disproportionately local and navigational.

Holding everything else about our keywords fixed, the model predicts 78.6% if no results page had a local pack, 66.1% if a quarter did and 53.6% if half did (the sample has 34.4%). The same method on the same day can give any figure in this range, depending on the mix of searches. A published AI Overview rate is only meaningful alongside the population of searches it was measured on.

### Two days later

We searched the 96 matched keywords as written on both 26 and 28 September. 61.5% showed an AI Overview on the first date and 59.4% on the second. 83.3% had the same outcome on both days: 7 gained an AI Overview and 9 lost one. Whether the page showed a local pack was more stable (the same on 96.9%). Where both days showed an AI Overview (50 keywords), the median overlap of cited URLs between the two days was 0.43 and of cited domains 0.58 (Jaccard: shared divided by all distinct). A single search is a sample, not a fixed property of the keyword.

### Which sources an AI Overview cites

Across all 643 AI Overviews (main and question samples), the median number of cited sources was 8 (middle half 6 to 10, range 0 to 29; 0.8% cited none). Navigational searches cited fewer (median 5.5), question searches slightly more (median 9).

| Source type | Share of 5,449 citations | Share on transactional searches | Share that also rank in the top 10 |
|---|---|---|---|
| Other websites (brands, publishers, blogs) | 59.4% | 33.9% | 41.1% |
| Forums, social and video | 21.2% | 16.8% | 11.0% |
| Google’s own pages | 12.0% | 37.9% | 1.5% |
| Review sites and directories | 2.4% | 1.6% | 39.4% |
| Government | 1.5% | 1.2% | 49.4% |
| Marketplaces and retailers | 1.5% | 7.9% | 50.6% |
| Education | 1.2% | 0.0% | 41.2% |
| Wikipedia | 0.8% | 0.8% | |

YouTube was cited in 62.4% of AI Overviews, Reddit in 21.8% and Google’s own pages in 20.7%. On transactional searches, Google’s own pages were the largest single source type.

### Most cited sources do not rank on page one

Of the 5,449 citations in the 638 AI Overviews that had both citations and organic results, 30.2% pointed to a URL in the organic top 10 of the same search (95% interval 26.8% to 33.9%) and 44.3% to a domain in it (40.4% to 48.1%). 20.2% of AI Overviews cited no top-10 URL, and 10.0% no top-10 domain. Forums, social, video and Google’s own pages rarely rank in the top 10 themselves, which explains part of the gap.

| Organic position | Share of top-10 URLs cited by the AI Overview |
|---|---|
| 1 | 50.0% |
| 2 to 3 | 37.6% |
| 4 to 6 | 27.8% |
| 7 to 10 | 21.8% |

Ranking helps: the first result was cited half the time, positions 7 to 10 about one time in five. But AI Overview source selection is clearly not the organic ranking repeated. Transactional searches had the lowest overlap (18.3% of citations in the top 10), informational the highest (34.8%).

### Does the cited page support the sentence?

A pilot on 120 cited sentences. For 23 the cited page could not be read (10 forum, social or video pages, 6 Google pages, 7 others), leaving 97. Two AI models (Claude Sonnet and Claude Opus), working independently, judged each sentence against the passages of the cited page that matched it best. They agreed on 86.6% of items (Cohen’s kappa 0.8).

| Judgment | Coder A | Coder B | Both agreed |
|---|---|---|---|
| Supported | 50.5% | 50.5% | 47 |
| Partly supported (a number, qualifier or name missing or different) | 17.5% | 20.6% | 15 |
| Not supported | 16.5% | 9.3% | 9 |
| Unclear (passages too thin to judge) | 15.5% | 19.6% | 13 |

Of the 71 sentences where both coders agreed and could judge, 66.2% were supported, 21.1% partly and 12.7% not: about one in three cited sentences was not fully backed by the passage we found on the cited page. This is a pilot: it checks only the first source cited for each sentence, against selected passages rather than the whole page, and no human reviewed the labels. A missed passage would count against the AI Overview, so the true support rate may be higher.

### What else was on the page

People Also Ask appeared on 95.0% of results pages, a local pack on 34.4%, video results on 16.0% and a discussions and forums block on 7.2%.

## How this compares with other studies

| Source | Sample and date | Figure |
|---|---|---|
| This study | 800 US desktop keywords, 8 industries, September 2026 | 60.8% of keywords; 40.8% volume-weighted |
| Xu, Iqbal and Montgomery | 55,393 trending queries, 19 categories, 40 days, spring 2026 | 13.7% of queries; 64.7% of question-form queries |
| Pew Research Center | Real searches by 900 US adults, March 2025 | 18% of searches |
| Semrush | 10 million+ keywords, November 2025 | 15.69% of queries |
| seoClarity | 500 million+ US keywords, September 2025 | 30% of US desktop keywords |
| BrightEdge (reported by Search Engine Journal) | Tracked keywords in 9 industries, February 2026 | “nearly half” of tracked queries |

These figures do not contradict each other; they measure different populations. Trending queries (news, sport, people) rarely get AI Overviews; the commercial and research-heavy keywords businesses track often do; Pew counted the searches people actually made, whatever they were. Directional findings agree: Xu and colleagues found question-form queries far more likely to activate, and that nearly 30% of cited domains were absent from the first page of results (we find 55.7% of cited domains outside the top 10 of the same search). They found 11.0% of 98,020 atomic claims unsupported by the cited pages; our smaller pilot, judging whole sentences against one source, found 12.7% not supported and 21.1% partly supported. Pew found an AI summary on 60% of searches beginning with a question word, and Whitespark (May 2025) found AI Overviews on 15% of local-intent searches against 92% of informational ones.

Sources: [Xu, Iqbal and Montgomery (arXiv)](https://arxiv.org/abs/2605.14021); [Pew Research Center](https://www.pewresearch.org/short-reads/2025/07/22/google-users-are-less-likely-to-click-on-links-when-an-ai-summary-appears-in-the-results/); [Semrush](https://www.semrush.com/blog/semrush-ai-overviews-study/); [seoClarity](https://www.seoclarity.net/research/ai-overviews-impact); [Search Engine Journal on BrightEdge](https://www.searchenginejournal.com/google-ai-overviews-surges-across-9-industries/568448/); [Whitespark](https://whitespark.ca/blog/case-study-the-prevalence-of-ai-overviews-in-local-search/).

## Observed, inferred and unknown

**What we observe.** AI Overviews rarely share a results page with a local pack, in every industry where both occur. Longer, more specific wordings of the same topic get far more AI Overviews. Industries differ even after adjusting for query mix, but by less than the raw table suggests. Most cited sources do not rank in the top 10 of the same search. Some cited sentences are not fully supported by the cited page. About one keyword in six changed outcome within two days.

**What we infer.** Google appears to decide what kind of results page a search needs before deciding whether to add an AI Overview: searches it treats as “find a place” get a map, searches it treats as “explain something” get an AI answer. Headline AI Overview rates are mostly a description of the sample they come from.

**What remains unknown.** Why Google shows an AI Overview for one search and not another; this study sees only the outcome. Whether a page’s content changes its chance of being cited (nothing here tests that). How the results vary by device, location and over weeks rather than days. Whether the cited pages shaped the answer, and whether anyone clicked.

## What this means

The points below are our interpretation. They follow from the findings but were not tested.

- **Measure AI Overview exposure on your own keyword set.** A published average says little about a given business: in our data the rate ran from 7.1% for franchise searches with a local pack to 96.9% for B2B software searches without one.
- **Local businesses face two different searches.** Short “service near me” searches mostly get a map, so the Google Business Profile matters there. The longer, specific questions customers ask before they call (“how much does a dental crown cost”) mostly get an AI Overview.
- **Ranking helps but is not enough.** The top organic result was cited half the time, and most citations went to pages outside the top 10, including video and forum pages.
- **Check what AI Overviews say about you, not just whether you are cited.** About one cited sentence in three in our pilot was not fully backed by the passage we found on the cited page.

Version 1.0 advised planning content around questions. That advice is withdrawn: this study shows that question-shaped searches get more AI Overviews, not that publishing question-shaped content earns more citations.

## Where this sits in GEO research

Research on generative engine optimization began by asking whether changing a page changes how often an AI answer uses it. Later work compared the sources that different AI search systems draw on, and a 2026 critical survey argued that visibility should be split into stages (activation, retrieval, citation, prominence, absorption, user outcome) rather than treated as one score. This study measures the first of those stages for Google, activation, and then follows activated answers into citation, grounding and short-term stability. It is observational: it describes when AI Overviews appear and what they cite, not what would change a page’s chance of being cited.

## Methodology

- **Keywords:** DataForSEO Labs keyword suggestions (United States, English) for 32 seed terms, four per industry; per industry the 40 highest-volume keywords plus 60 drawn at random (seed 20260926) from the remaining suggestions with at least 30 monthly searches. Question sample: suggestions beginning with a question word (how, what, why, when, which, who, where, can, does, do, is, are, should) followed by a space, 20 drawn at random per industry; 41 keywords that did not meet this rule were replaced before analysis.
- **Matched wordings:** 96 of the 800 (12 per industry, seed 20260928), excluding keywords that already began with a question word or contained “near me”. For each, a natural question and a long non-question form on the same topic were written by the research team with an AI model; the full list is in the downloads.
- **Results pages:** DataForSEO SERP API, Google.com, United States, English, desktop, top 10 organic results, asynchronous AI Overviews loaded; one search per query on 26 September 2026 (960) and 28 September 2026 (288).
- **Activation model:** logistic regression of AI Overview presence on industry, DataForSEO intent, local pack, “near me”, word count and log volume, with standard errors clustered by the 32 seed topics. Industry before and after adjustment uses average predicted probabilities. Price or cost words are reported descriptively because every such keyword showed an AI Overview. Predictive comparisons use 8-fold cross-validation grouped by seed topic.
- **Matched tests:** exact McNemar tests on the 96 pairs, and a logistic model of wording plus word count clustered by topic.
- **Citations:** all distinct cited links in each AI Overview, normalized (tracking parameters, fragments, “www.” and trailing slash removed). Top-10 overlap compares each cited URL, or its registrable domain, with the organic top 10 of the same results page. Source types use a fixed domain list.
- **Grounding pilot:** 120 AI Overviews (15 per industry, seed 20260928), one cited sentence each; the first cited page fetched; the six passages sharing most words with the sentence (up to 4,500 characters) judged independently by two AI models; figures on items where both agreed. Model-coded, not human-validated.
- **Uncertainty:** 95% bootstrap intervals resampling the 32 seed topics (2,000 resamples, seed 20260926).
- **Update schedule:** a further date for these keywords with our week-over-week volatility study (on or after 3 October 2026); monthly after that.

## Limitations

- The sample is commercial keywords in eight industries, leaning to high volume. It does not estimate the share of all Google searches with an AI Overview.
- Observational. The local pack is chosen by Google, so its link with fewer AI Overviews is an association, not a cause.
- The matched wordings were written by us to keep the topic; they also add detail, so the effect of the question form and of length cannot be fully separated.
- One search per query per date from one US location on desktop. Stability covers two dates and 96 keywords.
- The grounding pilot is small, model-coded, uses the first cited source only and judges selected passages, not the full page; 23 of 120 sources could not be read.
- Intent labels come from DataForSEO’s classifier. Source types do not separate a brand’s own site from independent publishers.

## What changed in version 1.1

Version 1.0 (26 September 2026) reported how often AI Overviews appeared by industry and type of search. Following an external review, version 1.1 (28 September 2026) reframes the study around when AI Overviews activate and what happens next. It adds a clustered logistic model, industry figures before and after adjustment, the local pack effect after adjustment, a predictive comparison, a denominator and sampling-frame analysis, a matched wording experiment with 288 new searches, a second date for 96 keywords, citation overlap with the organic top 10, and a model-coded grounding pilot. All version 1.0 figures are unchanged. The advice to plan content around questions was withdrawn because the study does not test it.

## Data and downloads

- All 960 keywords with features, AI Overview result and top-10 overlap: [s4_keywords_v11.csv](https://underneath.agency/research-data/ai-overviews-frequency-study/s4_keywords_v11.csv) and [JSON](https://underneath.agency/research-data/ai-overviews-frequency-study/s4_keywords_v11.json)
- Matched wordings, 288 searches: [s4_matched_forms.csv](https://underneath.agency/research-data/ai-overviews-frequency-study/s4_matched_forms.csv) and [the wording list](https://underneath.agency/research-data/ai-overviews-frequency-study/s4_matched_forms_design.tsv)
- Grounding pilot with both coders’ labels: [s4_grounding_pilot.json](https://underneath.agency/research-data/ai-overviews-frequency-study/s4_grounding_pilot.json)
- Every statistic on this page: [stats.json](https://underneath.agency/research-data/ai-overviews-frequency-study/stats.json)
- Machine-readable methodology: [methodology.json](https://underneath.agency/research-data/ai-overviews-frequency-study/methodology.json)
- Charts: [industry before and after adjustment](https://underneath.agency/research-data/ai-overviews-frequency-study/ai-overviews-industry-adjusted.svg) and [three wordings](https://underneath.agency/research-data/ai-overviews-frequency-study/ai-overviews-matched-query-forms.svg)
- Version 1.0 data and statistics: [s4_keywords.csv](https://underneath.agency/research-data/ai-overviews-frequency-study/s4_keywords.csv) and [stats_v1.0.json](https://underneath.agency/research-data/ai-overviews-frequency-study/stats_v1.0.json)

The data is free to reuse with attribution (CC BY 4.0).

To cite: Underneath. (2026). *When does Google show an AI Overview? 1,248 US searches* (Version 1.1). Underneath Research. https://underneath.agency/research/ai-overviews-frequency-study

## Frequently asked questions

### What percentage of Google searches show an AI Overview?

There is no single figure. In our sample of 800 US commercial keywords, 60.8% showed one, or 40.8% weighted by search volume. A 2026 study of trending queries found 13.7%. The rate depends mostly on the mix of searches: local, map-led searches rarely get one, and specific questions almost always do.

### Which industries see the most AI Overviews?

B2B software and technology (96.0% of keywords) and financial services and insurance (88.0%). After adjusting for query mix they are still highest (83.0% and 77.7%), but the gap between industries shrinks from 60.0 to 36.7 points.

### Do local searches show AI Overviews?

Rarely. When Google showed a local pack, an AI Overview appeared on 21.8% of searches, against 81.1% without one. The gap held in every industry with local packs and after adjusting for intent, length and volume.

### Do question searches trigger AI Overviews more often?

Yes. When we searched the same 96 topics three ways, 93.8% of the questions showed an AI Overview, against 59.4% of the original keywords. A long non-question wording of the same topic reached 87.5%, so most of the lift comes from a longer, more specific search, not the question word itself.

### Do AI Overviews cite the pages that rank first?

Partly. 30.2% of cited URLs were in the organic top 10 of the same search, and 20.2% of AI Overviews cited none of the top-10 URLs. The first organic result was cited half the time.

### How many sources does an AI Overview cite?

A median of 8, with the middle half between 6 and 10. Only 0.8% cited none.

## Related research

- [Do AI Overviews cite the pages that rank?](https://underneath.agency/research/ai-overview-citations-study)
- [AI Mode vs AI Overviews: how different are the sources?](https://underneath.agency/research/ai-mode-vs-ai-overviews-study)
- [What pages cited by AI Overviews have in common](https://underneath.agency/research/ai-overview-cited-pages-study)
- [How many “best of” lists cited by AI rank their own brand first?](https://underneath.agency/research/self-promoting-best-lists-study)

## Related guides

- [Which Google searches trigger an AI Overview?](https://underneath.agency/resources/which-searches-trigger-ai-overviews)
- [How much traffic would AI Mode as Google’s default cost you?](https://underneath.agency/resources/ai-mode-default-traffic-loss)
- [Are people more likely to stop browsing after seeing an AI Overview?](https://underneath.agency/resources/ai-overviews-end-browsing-sessions)
- [Does the way customers phrase a question change which sources AI search cites?](https://underneath.agency/resources/does-question-phrasing-change-ai-sources)
- [How much control do we have when Google changes how search results look?](https://underneath.agency/resources/google-search-design-changes-traffic-risk)

---

This is the Markdown twin of https://underneath.agency/research/ai-overviews-frequency-study. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How faithfully do AI assistants quote software prices?"
description: "61.9% of 840 plan prices four AI assistants quoted for 45 software products were fully faithful to the official page. Most of the rest were real variants."
canonical: "https://underneath.agency/research/ai-pricing-accuracy-study"
published: 2026-09-26
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Research · AI assistants

# How faithfully do AI assistants quote software prices?

Buyers ask AI assistants what software costs, sometimes instead of opening the pricing page. On 26 September 2026 we asked ChatGPT, Gemini, Perplexity and Google AI Mode “How much does {product} cost? List each plan and its monthly price in US dollars” for 45 software and subscription products, and captured each official pricing page the same day. Version 1.0 of this study checked whether each quoted number appeared on that page. This version asks a different question: when an assistant turns a published price list into a sentence, what survives and what is lost? The number, the plan it belongs to, the billing term it depends on, the source it came from, and the plans left out are all measured separately.

## The short version

1. Of 965 dollar amounts in the answers, 840 were plan prices. The rest (13.0%) were add-ons, yearly totals, other products or ranges, which a page-matching check counts as errors.
2. 61.9% of plan prices were fully faithful to the captured page: right amount, right plan, right billing term. Another 3.8% had the right amount but dropped a condition that changes what a buyer pays, usually by presenting an annual-billing price as the monthly price.
3. 21.9% were not on the captured page but fit a variant the page itself offers (the other side of a monthly/annual toggle, a higher contact tier). These look real; we could not verify them from the capture.
4. 7.6% differed from the page for the same plan and term, and 1.4% put a real price on the wrong plan. Where the size of the gap could be measured, 29 of 35 were underquotes, with a median gap of 13.3%.
5. Most differing prices were not invented. For 39 of 64, the same figure appears on another page of the vendor’s own site, and for 21 more on a third-party page cited for the product. Only 4 were found on no source we could fetch.
6. Once answers are clustered by product, the engines do not differ reliably (joint test p = 0.16), and a simple count of pricing-page complexity does not predict fidelity. Products differ far more than engines do.

## Research questions

- **RQ1.** What is a quoted price, once add-ons, totals and ranges are separated from plan prices?
- **RQ2.** How often is a plan price fully faithful, and when it is not, what is lost: the amount, the plan, or the condition?
- **RQ3.** When a price differs from the pricing page, where else does the same figure appear: the vendor’s other pages, a cited third-party page, or nowhere?
- **RQ4.** How much of each price list do answers cover, and do they say which billing term the prices assume?
- **RQ5.** Do engine, citing the vendor’s own pages, or pricing complexity explain fidelity once products are accounted for?

## From “on the page” to price fidelity

Version 1.0 reported that 63.1% of quoted amounts appeared, to the cent, on the captured pricing page. That figure is unchanged and still in the data, but it measures string agreement with one page state, not accuracy. A price can be on the page and still be wrong for the plan the answer names, and a real price can be missing from the page because the page showed annual billing that day.

This version treats each quoted price as a claim with several parts (amount, plan, billing term, unit, conditions and source) and places it on a ladder:

| Level | What the claim is | Share of 965 amounts |
|---|---|---|
| 4 | Fully faithful | 53.9% |
| 3 | Right amount, condition missing | 3.3% |
| 2 | Plausible variant, not on page | 19.1% |
| 1 | Not supported by the page | 10.8% |
| 0 | Not a plan price | 13.0% |

Level 1 combines prices that differ from the page, real prices put on the wrong plan, and prices that could not be judged. The levels are kept separate in the analysis and are not combined into one score.

## What we measured

We started with 60 software and subscription products with public US pricing pages. Each page was captured on 26 September 2026 with a rendering service; 45 pages loaded and showed at least three prices, and those 45 products are the sample. Each engine was asked once per product: ChatGPT and Gemini through their consumer apps, Google AI Mode through Google, Perplexity through its sonar API, all located in the United States. That gives 180 answers.

Every dollar amount above zero in an answer is one claim. Two AI models coded all 965 claims independently. Claude Opus coded each product with the four answers shuffled and the engine hidden, and Claude Sonnet coded the same material blind to the first. Each claim got a role (plan price, add-on, total, saving, other product, range) and, for plan prices, a verdict against the captured page. Every page each answer cited was fetched on 28 September 2026 to see where prices that were not on the pricing page could be found. Intervals and models account for the fact that the four answers about one product are not independent.

## Findings

### What the quoted prices turned out to be

| Outcome for 840 plan prices | Share | 95% interval |
|---|---|---|
| Fully faithful | 61.9% | 53.7% to 69.7% |
| Plausible variant, not on page | 21.9% | 14.7% to 29.2% |
| Differs from the page | 7.6% | 3.5% to 13.2% |
| Right amount, condition missing | 3.8% | 1.6% to 6.3% |
| Cannot be judged | 3.3% | 0.7% to 6.8% |
| Right amount, wrong plan | 1.4% | 0.6% to 2.4% |

The page-matching rule and the coding disagree in both directions. Of the 609 amounts found on the page, 65 were not plan prices, 28 dropped a condition and 11 were on the wrong plan. The coders judged 41 amounts that the rule could not find to be the page price after all, usually because the answer rendered the number differently, for example cutting off the last digit of Asana’s $10.99.

### What gets lost: conditions, plans and billing terms

A dropped condition was almost always a billing term. In 17 of the 32 such cases the answer stated the price as a plain monthly price. Asana’s Starter plan was quoted at $10.99 a month, which is the annual-billing rate; billed monthly it was $13.49. Buffer’s Essentials plan was quoted at $5 a month without saying the page price assumes paying $60 a year.

Wrong-plan errors take a real number and attach it to the wrong row. One answer gave Semrush’s $117.33 annual-billing price for the SEO plan as the Pro plan price, and another gave NordVPN’s $14.99 Basic monthly price for the Complete plan.

Answers usually do say which billing term they assume: 95.0% mention monthly or annual billing somewhere. They rarely say when the prices are from. 30.0% add a caveat that prices change or point to the official page, ranging from 57.8% for ChatGPT to 2.2% for Perplexity.

### When a price differs, where the figure comes from

We checked every price that was not on the captured pricing page against the other pages each engine cited for that product:

| Amounts not on the pricing page (356) | Count |
|---|---|
| On another vendor-owned page cited for the product | 138 |
| On a third-party page cited for the product | 183 |
| On no page we could fetch | 35 |

For the 64 prices coded as differing from the page, the split is 39 on another vendor page, 21 on a third-party page and 4 nowhere. The vendor pages are FAQ and knowledge-base articles, product pages, and blog or support posts about earlier price changes, for example Semrush’s plan FAQ, Zendesk’s article on its 2023 pricing update and Webflow’s post on its 2026 plan changes. An answer that quotes Semrush’s old Business plan at $499.95, or DocuSign Standard at $25 against the page’s $30, is repeating a figure the vendor’s own site still publishes. In these cases the vendor’s pages disagree with each other, and the assistant picked the other one.

A matching figure on a cited page could be coincidence, because comparison pages list many prices. As a check, we looked for each differing price on pages cited for a different product. It appeared there 20.0% of the time, against 90.0% on the product’s own cited pages. This shows the figures come from the cited material. It does not prove which page the engine used.

Where the size of a differing price could be measured against the same plan on the page (35 prices), 31 were within 5% to 20%, 3 were further off, 1 was within 5%, and 29 were below the page price. Old prices tend to be lower, which fits these being out-of-date figures.

### How much of the price list answers cover

Answers named a price for 84.2% of the paid plans on the captured page on average, and gave the page price for 68.4%. 59.4% of answers priced every paid plan the page showed. The denominator is demanding, because it includes team, student and bundle editions where the page lists them. 44.4% of answers had every plan price fully faithful, and 16.1% had at least one price that differed or sat on the wrong plan.

### Engines

| Engine | Plan prices | Faithful | Right amount | Differs |
|---|---|---|---|---|
| ChatGPT | 188 | 69.7% | 74.5% | 6.9% |
| Google AI Mode | 191 | 64.4% | 69.6% | 6.8% |
| Gemini | 264 | 58.0% | 58.3% | 7.2% |
| Perplexity | 197 | 57.4% | 63.5% | 9.6% |

Faithful means fully faithful; right amount counts faithful prices plus right amounts with a condition missing. The engines also differ in how they source and qualify prices. Google AI Mode cited the official pricing page in 88.9% of answers, Perplexity in 77.8%, ChatGPT in 51.1% and Gemini in 11.1%. ChatGPT (57.8%) and AI Mode (55.6%) usually added a date or change caveat; Gemini (4.4%) and Perplexity (2.2%) rarely did.

The raw shares order the engines, but the ordering does not hold up as a finding. The 95% intervals overlap widely (ChatGPT 58.8% to 79.1%, Perplexity 45.7% to 68.4%). In a model clustered by product that also includes complexity and vendor citations, only Perplexity’s lower odds reach p < 0.05 on its own (odds ratio 0.61), and the joint test of all engine differences does not (p = 0.16). Gemini quotes the most prices and the most amounts that are not plan prices (17.8% of its amounts), and cites the vendor’s own site least often (22.2% of answers against 97.8% or more for the others).

### Citations and complexity

Citing the vendor does not guarantee a faithful price. Plan prices in answers that cited a vendor-owned page were fully faithful 63.9% of the time, against 55.5% in answers citing only third-party pages. The adjusted odds ratio is 1.63, with a 95% interval of 0.7 to 3.83, so the difference is uncertain. 9.6% of prices in answers citing the vendor still differed from the page.

We scored each pricing page on six features that make a price harder to state: a monthly/annual toggle, per-seat pricing, usage or volume tiers, a contact-sales plan, an introductory price, and several products on one page. The score does not predict fidelity: 66.1% fully faithful on the simplest 18 pages, 47.4% on 11 medium pages, 67.8% on 16 of the most complex (adjusted odds ratio 0.90 per feature, 95% interval 0.69 to 1.19). One feature stands out descriptively. Pages with a billing toggle had 58.4% fully faithful prices against 72.9% without, which matches the billing-term errors above. The evidence for an engine-by-complexity interaction is weak (p = 0.056).

### How far the two coders agreed

On the verdicts for all 965 amounts, the two model coders agreed 85.2% of the time (Cohen’s kappa 0.78). They agreed 95.9% on role and 85.9% on the six complexity features. The second coder was somewhat stricter: 57.5% fully faithful and 14.6% differing or on the wrong plan, against 61.9% and 9.0% from the primary coder. The headline shares depend on the coder by a few points; the pattern does not.

## How this compares with other studies

VisibilityStack checked 13,867 queries about 108 software brands across six AI platforms between May and August 2026 and found AI “gets your whole price list right about one in four times” (24%). It also reported that “2 in 3 wrong prices are underquoted”. Its unit is the whole price list, graded against the vendor’s published price. Ours is the single price claim, graded against a same-day capture. Both find that underquoting dominates, which is what out-of-date prices would produce.

Source: [VisibilityStack](https://www.visibilitystack.ai/academy/geo/ai-quotes-your-pricing-wrong-mostly).

## Observed, inferred and unknown

- **Observed:** what each answer said, which pages it cited, what the captured pricing page showed, and where else each differing figure appears on fetchable pages.
- **Inferred:** that prices coded as plausible variants are real (they fit a toggle or tier the page shows), and that differing figures found on older vendor or third-party pages are out of date rather than invented.
- **Unknown:** which page an engine actually used for a given number, whether the same question asked again gives the same price, whether wording changes fidelity, and how long an old price persists after a vendor changes it.

## What this means

These points are interpretation, drawn from the results but not measured.

- **Most price errors start with the vendor’s own web footprint.** Retire or update old plan pages, help-center articles and announcements when prices change, not only the pricing page. An assistant that finds two official prices can quote either.
- **State both billing terms in plain text.** A monthly/annual toggle hides half the price list from any single read of the page, and billing terms are where correct amounts turn into misleading claims.
- **Check what assistants say about your plans.** Check the plan names too, not only the numbers. Outdated plan lineups (Semrush, Webflow and Zendesk tiers that no longer exist) were a common form of differing price.

## Where this sits in GEO research

Research on generative engine optimization first asked whether a source can become visible in an AI answer. Later work asked which sources different engines select and how that varies with engine, wording and intent. More recent work argued that retrieval, citation, prominence and fidelity should be measured separately, with repeated runs and clustered uncertainty. This study takes the next step for one kind of fact. Once an engine has selected information about a product, how faithfully does it turn it into a claim a buyer will act on? Prices suit the question because each claim can be checked on several dimensions. The answer so far is that fidelity is lost more in billing terms and in stale figures on the web than in invented numbers.

Two further phases would test what this snapshot cannot, and both need new data collection. The first is repeatability and wording: 15 to 20 products, each engine asked five times with up to six controlled wordings (direct, explicit monthly, billing-aware, cite the official page, as of today, buyer framing). The second is staleness: pricing pages captured weekly, and after a detected price change the engines asked again at 1, 3, 7, 14 and 30 days.

## Methodology

- **Products:** 60 software and subscription products with public US pricing pages; 45 had a usable captured page (at least three prices other than free plans). Excluded products are listed in stats.json.
- **Reference:** each official pricing page captured on 26 September 2026 with Firecrawl (rendered, US location).
- **Question and engines:** “How much does {product} cost? List each plan and its monthly price in US dollars.” ChatGPT and Gemini consumer apps (DataForSEO LLM Scraper), Google AI Mode (DataForSEO SERP API) and Perplexity sonar (DataForSEO LLM Responses API), US, one run each, 180 answers.
- **Claims:** every dollar amount above zero in an answer (965). The provider’s AI Mode text splits some numbers with spaces (“$3 5” for $35); these were rejoined.
- **Coding:** each claim coded for role, plan, stated billing term, verdict against the captured page and the page price for the same plan, plus six pricing-page features and the plans on the page. Primary coder Claude Opus, second coder Claude Sonnet, both through the Claude Code command-line tool, one call per product with the answers shuffled and the engine hidden. This is model coding; no person coded the data.
- **Sources:** all 932 cited URLs fetched over plain HTTP on 28 September 2026 (702 readable); dollar amounts on each page compared with the claims. Vendor-owned domains identified by the coder from the cited domains.
- **Statistics:** 95% percentile intervals from 2,000 resamples of products. GEE logistic models of plan-price outcomes on engine, complexity score and vendor citation, with exchangeable correlation within product and Wald tests for engine and engine-by-complexity.
- **Code:** cite/pipeline/s15_analyze.py (the 1.0 rule) and s15_v11.py (1.1 coding, source checks, analysis).
- **Update schedule:** quarterly.

## Limitations

- The captured pricing page is the reference, and it shows one billing term, region and promotion. A price coded as a plausible variant or as not supported may still be real, and no price was confirmed with a vendor.
- All coding is by two AI models, not people. Their agreement is reported; accuracy against a human standard is not.
- One run per engine per product, on one day. Run-to-run stability, sensitivity to wording and staleness after a price change are not measured.
- Cited pages were fetched two days after the answers, over plain HTTP. Pages that load prices with scripts or block fetching are missed, so the source counts are lower bounds. A figure on a cited page shows agreement, not which page the engine used.
- Plan coverage counts every paid plan on the page, including team, student and bundle editions, so full coverage is a demanding standard.
- 15 of 60 products were excluded because their page could not serve as a reference. Perplexity was queried through its API, the other engines through their consumer interfaces.

## What changed in version 1.1

Version 1.0 (26 September 2026) reported the share of quoted amounts found on the captured pricing page, and estimated from a check of 25 mismatches that roughly 11.8% of quoted prices were wrong or out of date. Following an external methodological review, version 1.1 (28 September 2026) re-analyzes the same 180 answers with no new queries. The review found that page matching is not accuracy, that claims about one product are not independent, that an engine ranking should not be the main result, and that context and omissions should count as outcomes.

Version 1.1 codes all 965 amounts, replacing the 25-price check and its 11.8% estimate. It adds the fidelity ladder, billing-term and wrong-plan errors, severity and direction, plan coverage, a source check for every price not on the page, citation analysis, pricing complexity, product-clustered intervals and models, and a second coder. The 1.0 figures (63.1% of amounts on the page; ChatGPT 75.9%, Perplexity 66.7%, Google AI Mode 57.5%, Gemini 56.1%) remain in stats.json and the 1.0 dataset. The 1.0 hand check was done by the AI research assistant, not by a person.

## Data and downloads

- Every price claim with its role, verdict, plan, billing term and where else the figure was found (version 1.1): [s15_price_claims.csv](https://underneath.agency/research-data/ai-pricing-accuracy-study/s15_price_claims.csv) and [JSON](https://underneath.agency/research-data/ai-pricing-accuracy-study/s15_price_claims.json)
- Every answer with plan coverage, caveats, citations and the pricing-page features (version 1.1): [s15_answers_v11.csv](https://underneath.agency/research-data/ai-pricing-accuracy-study/s15_answers_v11.csv) and [JSON](https://underneath.agency/research-data/ai-pricing-accuracy-study/s15_answers_v11.json)
- Every answer with its prices and which appeared on the page (version 1.0): [s15_price_answers.csv](https://underneath.agency/research-data/ai-pricing-accuracy-study/s15_price_answers.csv) and [JSON](https://underneath.agency/research-data/ai-pricing-accuracy-study/s15_price_answers.json)
- Every statistic on this page, with intervals, model results and coder agreement: [stats.json](https://underneath.agency/research-data/ai-pricing-accuracy-study/stats.json)
- Machine-readable methodology: [methodology.json](https://underneath.agency/research-data/ai-pricing-accuracy-study/methodology.json)
- Charts: [plan prices by outcome](https://underneath.agency/research-data/ai-pricing-accuracy-study/plan-prices-by-outcome.svg), [by engine](https://underneath.agency/research-data/ai-pricing-accuracy-study/faithful-by-engine.svg), [by pricing complexity](https://underneath.agency/research-data/ai-pricing-accuracy-study/faithful-by-pricing-complexity.svg)

The data is free to reuse with attribution (CC BY 4.0).

To cite: Underneath. (2026). *How faithfully do AI assistants quote software prices?* (Version 1.1). Underneath Research. https://underneath.agency/research/ai-pricing-accuracy-study

## Frequently asked questions

### Can you trust the prices AI assistants give for software?

As a starting point, not as a quote. 61.9% of the plan prices four assistants gave for 45 products were fully faithful to the official pricing page. Another 21.9% were probably real prices for a billing term or tier the page did not show by default. 7.6% differed from the page, usually because an older figure is still published somewhere.

### Which AI assistant is most accurate about prices?

None reliably, in this test. ChatGPT had the highest share of fully faithful plan prices (69.7%) and Perplexity the lowest (57.4%). Once the products were accounted for, the engine differences were not statistically reliable as a group, and the product being asked about mattered more than the engine.

### Why do AI assistants quote wrong prices?

Mostly because the web gives them wrong or outdated figures, and because billing terms get lost. Of 64 prices that differed from the pricing page, 39 appear on another page of the vendor’s own site and 21 on a third-party page cited for the product. Many other errors were right amounts with the billing term dropped, such as an annual-billing price presented as the monthly price.

### How can a company make sure AI quotes its prices correctly?

Update every page that states a price or plan name when prices change, including help-center articles and old plan pages. Show monthly and annual prices in plain text. Check what the assistants say about your plans, not only whether the numbers match.

## Related research

- [Do AI answers match a business’s Google profile?](https://underneath.agency/research/ai-business-facts-accuracy-study)
- [“Is this brand legit?” How AI assistants build a reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study)
- [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- [Do ChatGPT, Gemini, Perplexity and Claude agree on brands?](https://underneath.agency/research/ai-assistants-brand-agreement-study)

## Related guides

- [How do you fix wrong information about your brand in AI answers?](https://underneath.agency/resources/fix-wrong-brand-information-in-ai-answers)
- [How does a payroll provider get named when businesses ask AI?](https://underneath.agency/resources/payroll-software-ai-search)
- [Do AI agents make up facts about my business, or just leave them out?](https://underneath.agency/resources/do-ai-agents-invent-or-omit-business-facts)
- [How do subscription brands win new subscribers through AI search?](https://underneath.agency/resources/subscription-ecommerce-customers-ai-search)
- [How do project management tools win customers from AI answers?](https://underneath.agency/resources/project-management-software-customers-ai-search)

---

This is the Markdown twin of https://underneath.agency/research/ai-pricing-accuracy-study. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Does rewording a question change AI brand recommendations?"
description: "In 1,560 AI answers, a tight budget kept the original’s first brand 15.3% of the time, against 68.0% when the same question was simply asked again."
canonical: "https://underneath.agency/research/ai-prompt-phrasing-study"
published: 2026-09-26
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Research · AI assistants

# Does rewording a question change AI brand recommendations?

Buyers with the same need ask in different words. One says they run a small business, one says money is tight, one asks for a straight answer. This study asks how stable the brands ChatGPT, Gemini and Perplexity recommend are when the question is reworded, and how much of the change comes from the wording rather than from the ordinary variation between two runs of the same question. On 28 September 2026 we asked 20 buyer questions in four wordings, five times each, on all three engines, and three more phrasings of the budget request: 1,560 answers. Every rewording moved the recommendations further than a plain rerun did. A budget moved them most, and an explicit request for an unbiased answer moved them least. That other phrasings change AI recommendations is already established ([arXiv 2605.27440](https://arxiv.org/abs/2605.27440)). What this study adds is a same-batch rerun baseline for every question, a split of the change into its parts, and a test of whether the budget effect comes from the constraint or from the particular words.

## The short version

1. Asking the identical question again kept the same first brand 68.0% of the time. Adding “on a tight budget” kept it 15.3%, adding “I run a small business with about 10 employees” 40.6%, and asking for “an honest, unbiased answer” 54.9%.
2. Every rewording changed the brand list by more than a rerun. Brand overlap with the original (Jaccard) was 0.638 for a rerun, 0.547 for the unbiased request, 0.389 for business context and 0.326 for a budget. The excess over rerun variation was 0.273 for a budget (95% interval 0.220 to 0.322) and 0.058 for the unbiased request (0.036 to 0.080).
3. The wording explained 46.7% of the variation in brand lists between the budget and original answers to the same question, and 21.6% for the unbiased request; the rest is run-to-run variation. By chance alone it would explain about 0.11.
4. The budget effect comes mainly from the constraint, not the exact words. Four different ways of saying “cheap” overlapped with each other at 0.474, far closer to a same-phrasing rerun (0.563) than to the original question (0.341). The words mattered more on Perplexity than on ChatGPT or Gemini.
5. The top brand is not protected. A brand named first in the original survived a rerun 94.3% of the time but a budget rewording only 54.5%. Brands that came in on a budget were given a price or value reason 55.9% of the time (model-coded).

## Research questions

| Question | Answered here? |
|---|---|
| RQ1. Does adding context, a constraint or a meta-instruction change the brands recommended? | Yes |
| RQ2. Is that change larger than rerunning the identical question? | Yes, for every question, engine and wording |
| RQ3. How do the lists change: membership, first brand, top three, order, length? | Yes |
| RQ4. Does sensitivity differ by engine and category? | Yes, with a mixed model |
| RQ5. Which brands hold their place and which appear only under one wording? | Yes |
| RQ6. Does the budget effect come from the constraint or the words? | Yes, with four phrasings |

## Hypotheses and results

| Hypothesis | Result |
|---|---|
| H1. Adding decision information lowers overlap with the original | Supported |
| H2. A budget changes the list more than an unbiased-answer request | Supported |
| H3. Rewordings change the list more than a rerun | Supported for all three |
| H4. Sensitivity differs by engine | Partly |
| H5. Some brands are far more wording-sensitive than others | Supported |
| H6. Rewording removes lower-ranked brands more than top ones | Not for a budget |

H1: business context (0.389) and budget (0.326) against a rerun (0.638). H2: the intervals for budget (0.272 to 0.386) and the unbiased request (0.503 to 0.590) do not overlap. H3: the excess over reruns is above zero for all three rewordings, though small for the unbiased request. H4: the raw amount of change did not differ by engine once wording was accounted for (p = 0.304), but measured against each engine’s own rerun variation Perplexity was the most wording-sensitive (p = 0.006). H5 and H6 are covered under the brand findings below.

## A model of what moves a recommendation

We treat an AI recommendation as the result of several inputs: the buyer’s underlying need (the question), the wording, the engine, the date and the run. This study changes the wording while holding the question, engine and date fixed, and measures run-to-run variation directly by asking each wording five times. A brand therefore does not have one AI visibility score. It has a visibility for a question, in a wording, on an engine, at a time.

The three rewordings are different kinds of change, not three paraphrases.

| Rewording | Kind of change | Example |
|---|---|---|
| Business context | Who is asking | “I run a small business with about 10 employees. What is the best air fryer?” |
| Budget | A constraint | “What is the best air fryer on a tight budget?” |
| Unbiased answer | How to answer | “I’m comparing options and want an honest, unbiased answer, not a sales pitch: what is the best air fryer?” |

## What we measured

We took the 20 national buyer questions used in the first version of this study, drawn from our [four-assistant study](https://underneath.agency/research/ai-assistants-brand-agreement-study), such as “What is the best air fryer?” and “Which email marketing platform is best for ecommerce?”.

- **Main design (1,200 answers):** each question in the original wording and the three rewordings, five runs each, on ChatGPT and Gemini (their consumer apps) and Perplexity (sonar with web search), United States, 28 September 2026.
- **Budget phrasings (360 answers):** the budget request stated three other ways, two runs each on each engine: “What is the best affordable air fryer?”, “What is the best air fryer for someone trying to spend less?” and “Keeping the total price as low as possible is my priority. What is the best air fryer?”.
- **Same batch:** the answers to all wordings of a question on one engine were collected side by side, a median of 0.1 minutes apart. The whole collection took 0.4 hours, so time cannot explain the differences between wordings.

Brands were identified as in the four-assistant study: names in bold or headings, classified as brand or not, name variants merged per question, then matched in each answer in order of first mention. For each pair of answers we measured brand overlap (Jaccard: brands in both divided by brands in either), whether the first brand matched, how many of the first three were shared, rank-biased overlap (an agreement score that weights the top of the list most), the share of the original’s brands kept, the share of new brands, the order of shared brands (Kendall’s tau) and the change in list length. The rerun baseline compares the five runs of the same wording with each other. Every 95% interval comes from resampling the 20 questions, because answers to one question are not independent.

## Findings

### Every rewording moved the list further than a rerun

| Against the original | Overlap | First brand | Excess | Interval |
|---|---|---|---|---|
| Same wording, run again | 0.638 | 68.0% | | |
| Unbiased answer | 0.547 | 54.9% | 0.058 | 0.036 to 0.080 |
| Business context | 0.389 | 40.6% | 0.194 | 0.154 to 0.240 |
| Budget | 0.326 | 15.3% | 0.273 | 0.220 to 0.322 |

Overlap is brand overlap (Jaccard), first brand the share with the same first brand, and interval the 95% interval of the excess. Excess change is the drop in brand overlap beyond what reruns of the two wordings show on their own. The excess was positive in 93.3% of the 60 question-and-engine pairs for a budget, 91.7% for business context and 76.7% for the unbiased request. For the first brand, the budget excess was 48.8 percentage points.

### How much of the variation is the wording

Within each question and engine we split the variation in brand lists into a part explained by the wording and a part left over from run to run (a permutation analysis of variance on brand-list distances).

| Wording | Explained | Significant |
|---|---|---|
| Unbiased answer | 21.6% | 28.3% |
| Business context | 38.6% | 68.3% |
| Budget | 46.7% | 81.7% |

Explained is the share of variation explained by the wording; significant is the share of the 60 question-and-engine pairs where that effect is significant at the 5% level. With two wordings and ten answers, chance alone would give about 0.11. Taking all four wordings together, the wording explained 50.4% of the variation (95% interval 46.6% to 53.7%) and was significant at the 5% level in 95.0% of question-and-engine pairs.

A mixed model of the change per question, engine and rewording, with a random effect for each question, confirms the ordering: relative to the unbiased request, a budget raised the set-change score by 0.221 (0.164 to 0.277) and business context by 0.158 (0.102 to 0.214), and the wording as a whole was highly significant (p < 0.001). Differences between questions accounted for about a quarter of the remaining variance (0.255).

### Engines differ in noise more than in sensitivity

| Brand overlap with the original | Rerun | Unbiased | Context | Budget |
|---|---|---|---|---|
| ChatGPT | 0.642 | 0.555 | 0.408 | 0.363 |
| Gemini | 0.545 | 0.469 | 0.327 | 0.305 |
| Perplexity | 0.727 | 0.616 | 0.431 | 0.310 |

| Same first brand as the original | Rerun | Unbiased | Context | Budget |
|---|---|---|---|---|
| ChatGPT | 74.5% | 57.4% | 48.5% | 21.8% |
| Gemini | 50.5% | 40.4% | 33.3% | 13.6% |
| Perplexity | 79.0% | 66.8% | 40.0% | 10.4% |

Gemini’s lists change the most in absolute terms, but much of that is its own run-to-run variation: two runs of the identical question shared only 0.545 of their brands. Perplexity is the most repeatable (0.727) and so the most wording-sensitive once its low rerun variation is taken into account: the budget excess was 0.414 on Perplexity against 0.226 on ChatGPT and 0.179 on Gemini, and in the model of excess change Perplexity added 0.113 (0.069 to 0.157). The raw amount of change did not depend on the engine and wording together (p = 0.304).

### How the lists change

| Measure (all engines) | Rerun | Unbiased | Context | Budget |
|---|---|---|---|---|
| Brand overlap | 0.638 | 0.547 | 0.389 | 0.326 |
| Top three shared | 75.5% | 66.7% | 51.6% | 33.1% |
| Rank-biased overlap | 0.794 | 0.713 | 0.554 | 0.418 |
| Original brands kept | | 72.0% | 49.2% | 43.7% |
| New brands | | 31.0% | 39.3% | 48.2% |
| Order of shared brands | | 0.495 | 0.361 | 0.100 |
| Change in list length | | +0.24 | −1.00 | −1.12 |

Rank-biased overlap and order are scores from 0 to 1, where 1 is identical. A budget replaced about half the list and scrambled the order of the brands that stayed: the order score of 0.100 (interval −0.03 to 0.235) is not distinguishable from random. Business context and a budget both shortened the list by about one brand; the unbiased request slightly lengthened it.

### The top brand is not protected

For each brand in an original answer we asked how often it appeared in a rerun, and how often in a reworded answer.

| Place in the original | Rerun | Unbiased | Context | Budget |
|---|---|---|---|---|
| First | 94.3% | 87.6% | 69.1% | 54.5% |
| Second or third | 85.0% | 80.3% | 58.4% | 42.8% |
| Fourth or lower | 65.4% | 58.5% | 35.7% | 34.2% |

Based on 293 first-placed, 574 second- or third-placed and 903 lower-placed brand mentions. Lower-placed brands are less stable in every condition, but a budget costs the first brand about as much as the rest of the list (it keeps 54.5% of its 94.3%). We expected rewording to trim mainly the tail of the list (H6); for a budget it does not. Business context does hit the tail harder.

### Which brands hold their place

For each brand, question and engine we counted how often the brand appeared in the five runs of each wording (1,074 brand-question-engine combinations).

| Brand pattern | Share |
|---|---|
| Named in only one or two of about 20 answers | 40.6% |
| Leaves when a budget is added | 15.0% |
| No clear pattern | 14.5% |
| Enters when a budget is added | 9.4% |
| Stable: in at least 60% of runs of every wording | 7.4% |
| Enters with business context | 5.0% |
| Leaves with business context | 4.9% |
| Moves with the unbiased request | 3.2% |

Of the 337 brands named in at least 60% of the original’s runs, 78.6% held that level under the unbiased request, 51.3% under business context and 42.7% under a budget, and 47.8% were named in at most one of the five budget runs. Among the 638 brands named in three or more answers, 88.6% had a visibility that differed by 40 points or more between their best and worst wording. Engine-specific brands, named in half of one engine’s answers and almost never by the other two, were 10.5% of the 143 brands that reached that level anywhere.

### A brand does not have one AI visibility score

Perplexity, “Which email marketing platform is best for ecommerce?”, runs out of five that named each brand.

| Brand | Original | Context | Budget | Unbiased |
|---|---|---|---|---|
| Omnisend | 5 | 5 | 5 | 5 |
| Klaviyo | 5 | 5 | 0 | 5 |
| Mailchimp | 5 | 5 | 0 | 2 |
| Drip | 5 | 0 | 0 | 2 |
| ActiveCampaign | 4 | 0 | 0 | 5 |
| MailerLite | 0 | 3 | 5 | 4 |

Klaviyo was named in every Perplexity answer except the budget ones. On the same budget question ChatGPT named Klaviyo in all five runs and Gemini in one. Asked for the best air fryer, Perplexity named Cosori, TurboBlaze, Ninja and Instant Vortex Plus in all five original runs; on a tight budget it led with Gourmia and the Chefman TurboFry 2 Quart, and named Ninja in none. A single “AI visibility” figure for any of these brands would average over wordings that give 0 and 5 out of five.

### The constraint matters more than the words

To test whether the budget effect comes from the constraint or from the phrase “tight budget”, we asked the same constraint four ways.

| Comparison (all engines) | Brand overlap | Same first brand |
|---|---|---|
| Original wording, run again | 0.638 | 68.0% |
| Same budget phrasing, run again | 0.563 | 60.1% |
| Two different budget phrasings | 0.474 | 48.6% |
| A budget phrasing and the original | 0.341 | |

The phrasings agreed with each other much more than any of them agreed with the original: overlap with the original was 0.326 for “tight budget”, 0.370 for “affordable”, 0.353 for “someone trying to spend less” and 0.338 for “lowest total price is my priority”. Brands that came in with “tight budget” also appeared with each other phrasing 69.6% of the time. The words still mattered a little. The gap between a same-phrasing rerun and a different phrasing was 0.090 across engines (0.061 to 0.117), but it was small on ChatGPT (0.031) and Gemini (0.042) and larger on Perplexity (0.196), whose answers depend more on the exact search terms. A recent study of API models found that the prompt string, more than the buyer’s intent, decided which brands surfaced. In these consumer assistants, the intent (a price limit) did most of the work.

### What changes with the wording: sources and prices

Each engine cites web pages with its answers, so we can see whether a reworded question also draws on different sources.

| Domains | Cites | Rerun | Unbiased | Context | Budget |
|---|---|---|---|---|---|
| ChatGPT | 97.5% | 0.632 | 0.390 | 0.292 | 0.282 |
| Gemini | 95.2% | 0.325 | 0.190 | 0.107 | 0.147 |
| Perplexity | 100.0% | 0.986 | 0.820 | 0.517 | 0.595 |

Overlap of cited domains with the original; “cites” is the share of answers citing any source. Rewording changed the cited sources far more than rerunning did, and the questions whose sources changed most were the ones whose brands changed most (Spearman correlation 0.35 on ChatGPT, 0.344 on Gemini and 0.427 on Perplexity, each p < 0.01). Perplexity cited almost the same sites on every rerun (0.986), so its brand changes follow the changed search. The answers also changed in content: 84.7% of budget answers quoted a dollar figure, against 33.0% of original answers, 37.3% with business context and 43.0% with the unbiased request.

### Why brands came in with a budget

For the 10 questions where a budget changed the list most, we took each brand that appeared in the first budget answer but not in the first original answer (102 brands) and coded the main reason the answer itself gave for it. Two independent Claude model coders agreed on 87.3% (Cohen’s kappa 0.799); a third model pass settled the 13 disagreements. These are model-coded, not human-validated.

| Reason given for a new brand | Share |
|---|---|
| Price or value | 55.9% |
| Features | 21.6% |
| Suited to a type of buyer | 7.8% |
| Quality or reliability | 6.9% |
| Named as an alternative, no reason | 5.9% |
| No reason | 2.0% |

Price or value was the main reason for 71.0% of new brands on Perplexity, 55.9% on ChatGPT and 43.2% on Gemini. The change was mostly explained, not arbitrary: the assistants brought in brands they could justify on price.

### Where rewording matters most

| Category | Questions | Rerun | Context | Budget |
|---|---|---|---|---|
| Consumer products | 5 | 0.555 | 0.386 | 0.194 |
| Business software | 5 | 0.708 | 0.526 | 0.363 |
| Insurance and finance | 7 | 0.610 | 0.289 | 0.341 |
| Health and legal services | 3 | 0.725 | 0.398 | 0.451 |

Brand overlap with the original. Consumer products were the most budget-sensitive: robot vacuums (a set-change score of 0.844, where 1 means no brand in common), noise-cancelling headphones (0.828) and air fryers (0.827). The least sensitive were online LLC formation (0.355), car insurance for young drivers (0.486), whose original question already asks for the cheapest, and robo-advisors (0.502). Questions with more stable reruns were somewhat less budget-sensitive (Spearman 0.29, p = 0.025). The category groups are small, so these differences are descriptive.

### Two days apart, and the first version

Answers to the original wording on 26 September (from the first version of this study) overlapped with the 28 September runs about as much as two same-day runs did: 0.642 against 0.642 on ChatGPT, 0.546 against 0.545 on Gemini, and 0.688 against 0.727 on Perplexity. Over two days, time added little beyond run-to-run variation.

| Average of three engines | v1.0 | v1.1 |
|---|---|---|
| Same first brand, budget | 12.0% | 15.3% |
| Same first brand, business context | 40.0% | 40.6% |
| Same first brand, unbiased | 55.0% | 54.9% |
| Brand overlap, budget | 0.366 | 0.326 |
| Brand overlap, business context | 0.412 | 0.389 |
| Brand overlap, unbiased | 0.574 | 0.547 |

Version 1.0 used one run per wording on 26 September; version 1.1 five runs per wording on 28 September. The first version’s ordering holds on all three engines. One difference: in version 1.0, ChatGPT’s budget answers overlapped with the original more than its business-context answers did (0.454 against 0.402). With five runs, the budget is the largest change on every engine.

## A prompt sensitivity index

Because no single number captures how a list changes, we report sensitivity as three scores, each from 0 (no change) to 1 (complete change), between the original and the reworded answers.

| Wording | Set change | First-brand change | Rank change |
|---|---|---|---|
| Same wording, run again | 0.362 | 0.32 | 0.206 |
| Unbiased answer | 0.453 | 0.451 | 0.287 |
| Business context | 0.611 | 0.594 | 0.446 |
| Budget | 0.674 | 0.847 | 0.582 |

Set change is 1 minus brand overlap, first-brand change is 1 minus the same-first-brand rate, and rank change is 1 minus rank-biased overlap. The first row is the floor that ordinary variation sets. A tracker that reports one prompt run once cannot tell a wording effect from that floor.

## How this compares with other studies

| Study | What changed | Engines | Runs | Main figure |
|---|---|---|---|---|
| Aggarwal et al. 2024 (GEO) | The source pages | Generative engines | Controlled | Page edits change how visible a source is |
| Chen et al. 2025 | Engine, language, vertical, paraphrase | Several AI search engines | Controlled | Cited sources differ by engine and phrasing |
| Jack et al. 2026, persona | A persona prefix | OpenAI, Anthropic APIs | 2,000 | Jaccard lowered by 0.12 to 0.20 |
| Jack et al. 2026, paraphrase | Cosmetic and constraint paraphrases | OpenAI, Anthropic APIs | About 12,000 | 0.288 and 0.135, against 0.50 to 0.61 for reruns |
| Rankshift 2026 | Seven near-synonym CRM prompts | Several | 1,176 each | HubSpot at 98% on every variant |
| Tannenbaum 2026 | None (monitoring data) | GPT, Gemini | 34,960 | Mentions depend on the brand being retrieved |
| This study | Context, constraint, meta-instruction, four budget phrasings | ChatGPT, Gemini, Perplexity apps | 5 per wording | 0.326 for a budget against 0.638 for reruns |

The API study of paraphrases found constraint-adding rewordings overlapping at 0.135 against 0.50 to 0.61 for reruns. Our budget rewording (0.326 against 0.638) points the same way, less extremely. The clearest difference is on cosmetic paraphrase: there, different wordings of the same intent overlapped at 0.288; here, different wordings of the same budget constraint overlapped at 0.474, close to a rerun on ChatGPT and Gemini. Persona conditioning lowered overlap by 0.12 to 0.20, and our business-context wording, a short persona, lowered it by more (0.194 beyond reruns). A study of monitoring data found that brands are mentioned mainly when their own pages appear in what the engine retrieves. Our finding that brand changes follow source changes is consistent with that.

Sources: [arXiv 2311.09735](https://arxiv.org/abs/2311.09735); [arXiv 2509.08919](https://arxiv.org/abs/2509.08919); [arXiv 2605.30207](https://arxiv.org/abs/2605.30207); [arXiv 2605.27440](https://arxiv.org/abs/2605.27440); [Rankshift](https://www.rankshift.ai/blog/impact-of-prompt-phrasing-on-ai-brand-visibility/); [arXiv 2609.23162](https://arxiv.org/abs/2609.23162).

## Observed, inferred and unknown

**What we observe.** Asked the same question in the same batch, the three assistants return lists that vary from run to run, and adding context or a constraint changes the list well beyond that. A budget changes the first brand in most answers, replaces about half the brands and scrambles the order of the rest. Several ways of stating a budget lead to similar lists. The cited sources change along with the brands, and the new brands come with price reasons.

**What we infer.** The assistants treat a budget or a stated situation as a different request and search for it differently, rather than reacting to particular words. AI brand visibility is conditional on how the need is expressed, so it should be measured across wordings and repeated runs, not from one prompt run once.

**What remains unknown.** Whether the new recommendations are better or worse for the buyer: we measured difference, not quality. How much answers drift over weeks rather than two days. Whether logged-in history or personalization adds more variation. Which brand characteristics (price level, size, review volume) predict holding a place across wordings: the brand-level data is published for that analysis, but we have not collected those attributes.

## What this means

The points below are interpretation. They follow from the findings but were not tested.

- **Report visibility for a question in a wording, not one score.** A brand that appears in every answer to “What is the best email marketing platform for ecommerce?” can be absent from every answer that adds a budget. Measure each important wording separately.
- **Run each prompt several times before reading a change.** Two runs of the identical question shared 0.638 of their brands. A difference smaller than that is often noise.
- **Match the buyer’s constraint, not their exact phrase.** Different ways of asking for something cheap led to similar lists, so content that states prices and value plainly serves all of them.
- **Being first in the default answer does not carry over.** The first-named brand kept its place in only about half of budget answers.

## Where this sits in GEO research

Generative engine optimization began by asking whether changing a source page changes how visible it is in AI answers. Later work showed that AI search engines differ by language, vertical and phrasing. A 2026 survey of the field names repeated runs, paraphrase, denominators and causal measurement as unsolved measurement problems. This study takes up two of them: it separates wording effects from rerun variation with a same-batch baseline, and it shows that a brand’s visibility has no meaning without saying for which wording. It does not show how to change that visibility. The next collection repeats the main design on a new date to separate time from run-to-run variation.

## Methodology

- **Questions:** 20 national buyer questions from the four-assistant study (every second national question), each in the original wording and three rewordings; three further budget phrasings.
- **Engines:** ChatGPT and Gemini consumer apps via DataForSEO LLM Scraper; Perplexity (sonar) via DataForSEO LLM Responses API with web search; United States; English.
- **Collection:** 1,560 requests on 28 September 2026 between 08:57 and 09:19 UTC: 1,200 in the main design (five runs per wording) and 360 for the budget phrasings (two runs each). All 1,560 returned; six unfinished ChatGPT answers under 200 characters are kept as outcomes but excluded from brand measures, and 18 answers named no brand. Version 1.0 answers (26 September 2026, one run per wording) are used for the cross-date and version comparisons.
- **Brands:** candidate names from bold text and headings (and ChatGPT’s brand panels); each unique string classified as brand or not (study 7 decisions reused; new strings first filtered by rule for prices, pronouns and long phrases, then classified by Claude model coders, with directive fragments and URLs corrected by rule); name variants merged per question; matched in each answer in order of first mention. Names that are also ordinary words (such as Choice, Fidelity or Travelers) count only where the answer lists them as an item, and review sites cited as sources are not counted as brands.
- **Measures:** Jaccard similarity, same first brand, top-three overlap, extrapolated rank-biased overlap (persistence 0.9), retention, replacement, Kendall’s tau on shared brands (three or more), list length. Rerun baseline from all pairs of the five runs; between-wording figures from all 25 pairs of original and reworded runs.
- **Variance and models:** permutation analysis of variance on Jaccard distance per question and engine (499 permutations); linear mixed model with wording and engine as fixed effects and question as a random intercept; likelihood-ratio tests.
- **Reasons for change:** 102 brands entering in the first budget answer against the first original answer, for the 10 most budget-sensitive questions; two independent model coders, third-pass adjudication.
- **Uncertainty:** 95% bootstrap intervals resampling the 20 questions (2,000 resamples, seed 20260928).
- **Update schedule:** quarterly; next collection repeats the main design on a new date.

## Limitations

- Three engines, 20 questions, the United States and English, with the main design on one day. Other categories, countries and dates may differ.
- ChatGPT and Gemini answers come from their consumer apps through a data provider, without a signed-in history; Perplexity is the API with web search. Personalization is not measured.
- Four wordings and four budget phrasings are a small sample of how buyers ask. The business-context wording was added to every question, including consumer ones where it is less natural.
- Five runs per wording estimate run-to-run variation with some noise, and two runs per budget phrasing is thin.
- Brand identification is automatic and can merge or split names; the reasons for change are model-coded, not human-validated.
- We change the question, not the model or its retrieval, so the source analysis shows association, not cause.
- Different recommendations are not necessarily better or worse ones.

## What changed in version 1.1

Version 1.0 (26 September 2026) asked each rewording once and compared it with the original from our four-assistant study, with a rerun comparison on only 5 questions from a batch collected hours earlier. Following an external review, version 1.1 (28 September 2026) is a new collection of 1,560 answers. It adds five same-batch runs of every wording for all 20 questions, measures of how lists change (not only Jaccard), a variance split and mixed model with intervals clustered by question, brand-level patterns, model-coded reasons for change, a test of four budget phrasings, hypotheses taken from the review before the new data were collected, and a comparison with the literature. Brand matching now ignores ordinary words and review sites that the earlier method counted as brands. The version 1.0 figures are kept in the data files and compared above; its conclusions hold.

## Data and downloads

- Every answer with its brands in order, first brand and cited domains: [s18v11_answers.csv](https://underneath.agency/research-data/ai-prompt-phrasing-study/s18v11_answers.csv) and [JSON](https://underneath.agency/research-data/ai-prompt-phrasing-study/s18v11_answers.json)
- Every question, engine and wording comparison with all measures: [s18v11_cells.csv](https://underneath.agency/research-data/ai-prompt-phrasing-study/s18v11_cells.csv)
- Each brand’s appearance rate by wording: [s18v11_brand_fate.csv](https://underneath.agency/research-data/ai-prompt-phrasing-study/s18v11_brand_fate.csv)
- The budget-phrasing comparison: [s18v11_paraphrase_cells.csv](https://underneath.agency/research-data/ai-prompt-phrasing-study/s18v11_paraphrase_cells.csv)
- Reason codes for brands entering with a budget: [s18v11_reason_codes.csv](https://underneath.agency/research-data/ai-prompt-phrasing-study/s18v11_reason_codes.csv)
- Every statistic on this page: [stats.json](https://underneath.agency/research-data/ai-prompt-phrasing-study/stats.json); version 1.0: [stats_v1.0.json](https://underneath.agency/research-data/ai-prompt-phrasing-study/stats_v1.0.json) and [s18_phrasing_answers.csv](https://underneath.agency/research-data/ai-prompt-phrasing-study/s18_phrasing_answers.csv)
- Machine-readable methodology: [methodology.json](https://underneath.agency/research-data/ai-prompt-phrasing-study/methodology.json)
- Charts: [rewording against rerun](https://underneath.agency/research-data/ai-prompt-phrasing-study/rewording-vs-repeat-overlap.svg); [share of variation from wording](https://underneath.agency/research-data/ai-prompt-phrasing-study/share-of-variation-from-wording.svg)

The data is free to reuse with attribution (CC BY 4.0).

To cite: Underneath. (2026). *Does rewording a question change AI brand recommendations?* (Version 1.1). Underneath Research. https://underneath.agency/research/ai-prompt-phrasing-study

## Frequently asked questions

### Does the way you word a question change which brands ChatGPT recommends?

Yes, beyond ordinary variation. Asked the identical question again, ChatGPT named the same first brand 74.5% of the time; with “on a tight budget” added, 21.8%.

### Which rewording changes AI recommendations the most?

A budget, among the three we tested. Averaged across ChatGPT, Gemini and Perplexity, it kept the original’s first brand 15.3% of the time and shared 0.326 of its brands with the original, against 68.0% and 0.638 for a rerun.

### Does asking an AI for an unbiased answer change its recommendations?

A little. The unbiased request changed the list slightly more than a rerun (an excess of 0.058 in brand overlap), much less than a budget or business context.

### Is the budget effect caused by the exact words?

Mostly not. Four ways of asking for something cheap produced similar lists (overlap 0.474 with each other against 0.341 with the original). The exact words mattered more on Perplexity than on ChatGPT or Gemini.

### How should a brand track its AI visibility across different phrasings?

Track several wordings of each important buyer question, run each several times, and report visibility per wording and engine, because one brand can appear in every answer to one wording and none to another.

## Related research

- [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)
- [Same question, four countries: do AI recommendations change?](https://underneath.agency/research/ai-recommendations-by-country-study)
- [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- [Do ChatGPT, Gemini, Perplexity and Claude agree on brands?](https://underneath.agency/research/ai-assistants-brand-agreement-study)

## Related guides

- [Do AI search engines just agree with however a question is phrased?](https://underneath.agency/resources/do-ai-search-engines-agree-with-leading-questions)
- [How should you design prompts and runs to track AI visibility?](https://underneath.agency/resources/how-to-design-ai-visibility-tracking)
- [Can I trust a one-time AI visibility report for my brand?](https://underneath.agency/resources/one-time-ai-visibility-report)
- [How many prompts do we need to track to measure AI visibility?](https://underneath.agency/resources/how-many-prompts-to-track-ai-visibility)
- [Does the way customers phrase a question change which sources AI search cites?](https://underneath.agency/resources/does-question-phrasing-change-ai-sources)

---

This is the Markdown twin of https://underneath.agency/research/ai-prompt-phrasing-study. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Ask an AI the same question 5 times: do the brands change?"
description: "Same 20 buyer questions, five runs each: a quarter of ChatGPT’s brands appeared every time. AI visibility is a distribution, and one answer is a sample."
canonical: "https://underneath.agency/research/ai-recommendation-consistency-study"
published: 2026-09-26
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Research · AI assistants

# Ask an AI the same question 5 times: do the brands change?

When a report says a brand is “visible in ChatGPT” for a question, it usually means the brand appeared in one answer. But assistants give a different answer each time. We asked ChatGPT, Gemini and Perplexity the same 20 buyer questions five times each on 26 September 2026, with the five runs of a question a median of about 11 minutes apart, and compared all 300 answers.

Version 1.0 of this study showed that the brands change. This version asks what that means for measurement. We treat each brand’s visibility as a frequency: the share of runs that name it. Then we look at how the change is structured (which brands, in what order, which one first, from which sources), what it goes with, and how much a single answer gets wrong. No new questions were asked; the 300 answers are the same, plus the answers to the same questions from our country and rewording studies, collected a few hours later that day.

The short answer: AI visibility is not a yes-or-no property of a brand and a question. It is a distribution with a small core and a long tail, and its shape depends on the assistant, the question and when you ask.

## The short version

1. A quarter of ChatGPT’s brands were named in every run. Of the brands ChatGPT named for a question across five runs, 25.2% appeared in all five (95% interval 15.9% to 36.8%) and 36.6% in only one. For Gemini the figures were 13.7% and 46.7%; for Perplexity 40.9% and 15.8%.
2. One answer shows about half the picture. A single ChatGPT answer showed 57.8% of the brands its five answers named between them; Gemini 48.4%, Perplexity 72.2%. The fifth run still added about one new brand on ChatGPT and Gemini, so five runs had not found the whole tail either.
3. Being named is not the same as being ranked. ChatGPT named the same first brand in two runs of a question 53.8% of the time, and its first pick changed at least once for 80.0% of questions. Of the brands ChatGPT named in every run, 86.4% moved position between runs.
4. Same sources, different brands. In 166 pairs of Perplexity runs the cited URLs were identical, yet the brand list changed in 91.6% of them. Across all three assistants, run pairs that shared more sources shared slightly more brands (about 0.03 more brand overlap per 0.1 more source overlap), but source overlap explained little of the change in brands (correlation 0.357 on ChatGPT, 0.228 on Gemini).
5. Stability depends on the question as much as the assistant. The mean overlap between runs ranged from 0.299 to 0.747 across questions. The question accounted for 30.0% of the variation, the assistant for 25.5%, and the combination of the two for the rest; a question that was stable on one assistant was not reliably stable on another.
6. Time matters within hours. For ChatGPT, an answer to the same question about 4.4 hours later overlapped with the five runs less (0.477) than the runs overlapped with each other (0.588), and an answer for another country less again (0.398). Rewording a question changed Gemini’s brands more than asking again did; for ChatGPT, most of the apparent rewording effect was already there when the same question was asked hours later.
7. Five runs are too few for a precise number. A brand named in 3 of 5 runs has a 95% interval of 23.1% to 88.2%. Pinning a frequency near 50% to within 10 points would take about 97 independent runs; our data cannot yet say how many are needed in practice.

## Research questions

| Question | Answered here? |
|---|---|
| RQ1. How much do the recommended brands change across identical runs? | Yes |
| RQ2. How is the change structured: core and tail, order, first pick? | Yes |
| RQ3. Is the change in brands tied to a change in cited sources? | Yes, as an association |
| RQ4. What does one run miss, and how precise are five? | Partly |
| RQ5. How does repetition compare with hours later, rewording and country? | Partly, on 5 to 10 questions |
| RQ6. Drift over days and weeks; runs needed in practice; retrieval vs generation | No, planned |

## A brand’s visibility is a frequency, not a yes or no

Across five runs, each brand an assistant named for a question gets a frequency: named in 1, 2, 3, 4 or 5 of the runs. We group them as core (5 of 5), stable (4 of 5), intermittent (2 or 3) and tail (1 of 5).

| Share of brands | ChatGPT | Gemini | Perplexity |
|---|---|---|---|
| Core (5 of 5) | 25.2% | 13.7% | 40.9% |
| Stable (4 of 5) | 15.3% | 14.1% | 14.8% |
| Intermittent (2 to 3) | 22.9% | 25.6% | 28.6% |
| Tail (1 of 5) | 36.6% | 46.7% | 15.8% |
| Brand-question pairs | 262 | 227 | 203 |

The 95% intervals, from resampling the 20 questions, are wide for the core share: 15.9% to 36.8% for ChatGPT, 8.3% to 19.8% for Gemini and 35.0% to 47.8% for Perplexity. The ordering of the three assistants holds: Perplexity’s mean run-to-run brand overlap was 0.135 higher than ChatGPT’s (interval 0.034 to 0.247) and 0.260 higher than Gemini’s (0.168 to 0.359), and ChatGPT’s was 0.125 higher than Gemini’s (0.013 to 0.233).

The frequency is an observed share of five runs, not a known probability. With five runs it can only place a brand in broad bands (see “How precise five runs are” below).

## How concentrated the recommendations are

Counting the brands in one answer understates how many brands an assistant spreads its answers over. We measured the effective number of brands per question: the number of equally frequent brands that would give the same spread (the exponential of the entropy of each brand’s share of all mentions across the five runs).

| Per question (median) | ChatGPT | Gemini | Perplexity |
|---|---|---|---|
| Brands in one answer | 6.5 | 5.5 | 6.8 |
| Distinct brands in 5 answers | 11 | 10.5 | 9.5 |
| Effective number of brands | 9.5 | 9.2 | 8.6 |
| Core brands (5 of 5) | 3 | 2 | 4 |
| Tail brands (1 of 5) | 4 | 4 | 1 |

On average Gemini’s effective number of brands was about twice the length of one answer (2.0 times), ChatGPT’s 1.8 times and Perplexity’s 1.3 times. An answer of six brands from ChatGPT typically comes from a pool of nine or ten that it moves between.

## Five kinds of change

Two answers can differ in which brands they name, in the order, in the first pick, in the sources they cite and in the wording. We measured each separately over every pair of runs of the same question (596 pairs). Each figure below is the mean share that changes between two runs, from 0 (identical) to 1 (nothing shared).

| Change between two runs | ChatGPT | Gemini | Perplexity |
|---|---|---|---|
| Brands named | 0.455 | 0.580 | 0.320 |
| Order (rank-biased) | 0.405 | 0.528 | 0.243 |
| Order of shared brands | 0.182 | 0.280 | 0.159 |
| First pick | 0.462 | 0.570 | 0.220 |
| Top 3 | 0.444 | 0.637 | 0.287 |
| Cited domains | 0.525 | 0.674 | 0.013 |
| Cited URLs | 0.637 | 0.779 | 0.020 |
| Words used | 0.588 | 0.689 | 0.551 |

How to read the rows: “brands named” is 1 minus the Jaccard overlap of the two brand lists; “order (rank-biased)” is 1 minus rank-biased overlap, which weighs the top of the list most; “order of shared brands” rescales Kendall’s tau for brands both runs named; “first pick” is the share of run pairs whose first-named brand differs.

Three things stand out. When two runs named the same brands, they mostly kept them in a similar order (the lowest of the brand rows), so most change is in which brands appear, not how they are ranked. The first pick changed about as often as the brand list did. And Perplexity is the outlier on sources: its citations barely moved while its brands and wording still did.

- Chart: [how much changes between two runs, cited URLs vs brands named](https://underneath.agency/research-data/ai-recommendation-consistency-study/sources-vs-brands-change.svg)
- Chart: [share of brands named in every one of five runs](https://underneath.agency/research-data/ai-recommendation-consistency-study/brands-in-every-run.svg)

## The first pick and the top 3

For someone asking once, the brand an answer names first is effectively the recommendation.

| Across 20 questions | ChatGPT | Gemini | Perplexity |
|---|---|---|---|
| First pick changed at least once | 80.0% | 85.0% | 40.0% |
| Mean share of runs with the most common first pick | 71.8% | 63.0% | 87.0% |
| A brand in the top 3 of every run | 60.0% | 40.0% | 90.0% |
| Same top 3 in all runs | 15.0% | 0.0% | 15.0% |

Stable membership did not mean a stable position. Of the brands named in every run, 86.4% moved position at least once on ChatGPT, 87.1% on Gemini and 72.3% on Perplexity; 34.8%, 25.8% and 42.2% of them were in the top 3 every time. A brand can have steady visibility and an unsteady rank.

## Same sources, different brands

If the brands change because the assistant finds different pages each time, runs that cite the same sources should name the same brands. Partly, they do. Within a question, each 0.1 more overlap in cited domains went with 0.03 more brand overlap (95% interval 0.017 to 0.043), with similar slopes for ChatGPT (0.027) and Gemini (0.032).

But Perplexity cited nearly the same sources on every run (mean URL overlap 0.980), and its brands still changed. In the 166 pairs of Perplexity runs with identical cited URLs, the brand lists differed 91.6% of the time (mean overlap 0.675) and the first pick differed 21.1% of the time. The same happened on ChatGPT in the few pairs with identical URLs (9 pairs, 44.4% with a different brand list).

So the evidence an answer shows and the brands it recommends are two different things to measure: evidence can be stable while recommendations are not. Our data measures outputs only, so it cannot say which step inside the assistant produces the change; the cited sources are what the answer shows, not everything the assistant read.

## What makes a question more or less stable

The mean run-to-run overlap ranged from 0.299 for the least stable question to 0.747 for the most stable (averaged over the three assistants). Splitting the variation across the 60 question-assistant cells, the question accounted for 30.0%, the assistant for 25.5% and their combination for 44.5%. That last share matters: a question that was stable on one assistant was not reliably stable on another (correlation 0.301 between ChatGPT and Gemini, around zero between either and Perplexity).

With 20 questions, the patterns by type are exploratory.

- **By industry:** B2B software had the most stable brands (mean overlap 0.708 across assistants), franchises the least (0.434), with retail close behind (0.458).
- **Place-named questions:** on Gemini, the 7 questions naming a city were less stable than the 13 national ones (0.343 vs 0.460); on ChatGPT and Perplexity there was little difference.
- **List length:** on Gemini, questions with longer answers were more stable (correlation 0.573 between brands per answer and overlap); on Perplexity the reverse (−0.431).

## Minutes, hours, rewording and country

The five runs were minutes apart. Our [country study](https://underneath.agency/research/ai-recommendations-by-country-study) and [rewording study](https://underneath.agency/research/ai-prompt-phrasing-study) asked some of the same questions about 4.4 hours later the same day, including a US answer to the unchanged question. That lets us compare, on the same questions, how far an answer moves when you ask again at once, ask again hours later, and change the country or the wording.

| Mean overlap with the five runs | ChatGPT | Gemini |
|---|---|---|
| Another run, same batch (10 questions) | 0.588 | 0.453 |
| Same question, hours later | 0.477 | 0.491 |
| Another country, hours later | 0.398 | 0.363 |
| Another run, same batch (5 questions) | 0.694 | 0.453 |
| Same question, hours later | 0.455 | 0.466 |
| Reworded, hours later | 0.431 | 0.298 |

The first three rows use the 10 questions in the country study; the last three the 5 in the rewording study.

For ChatGPT, time alone moved the answer: the same US question hours later overlapped with the runs less than the runs did with each other. Changing the country moved it further (0.477 to 0.398). Rewording, by contrast, added little beyond time (0.455 to 0.431). An earlier comparison against the minutes-apart runs alone (0.694 to 0.431) would credit that whole gap to the wording.

For Gemini there was no sign of change over the hours (0.453 then 0.491), and both country (0.491 to 0.363) and rewording (0.466 to 0.298) moved the answer. Perplexity was not in the country study and had no hours-later answer; its reworded answers overlapped with the runs 0.462, against 0.663 between runs.

These are single later answers on 5 or 10 questions, so they show the direction, not firm sizes. They do show that short-term repetition, elapsed time and the question itself are separate sources of change, and that a comparison between two conditions needs a baseline collected at the same time.

## What one run gets wrong

A tracker that asks each question once records one draw. Compared with the five runs, a single run did three things.

- It showed 57.8% of the brands the five ChatGPT answers named (95% interval 50.0% to 65.2%), 48.4% for Gemini and 72.2% for Perplexity.
- It left out brands that were not marginal: of the brands missing from a single run, 16.3% on ChatGPT, 15.8% on Gemini and 27.5% on Perplexity were named in at least three of the five.
- It sometimes reversed two brands: for pairs of brands whose five-run frequencies differed by 0.4 or more, a single run named the rarer brand but not the more frequent one 4.2% of the time on ChatGPT, 6.1% on Gemini and 2.0% on Perplexity.

Each extra run kept finding new brands. Averaged over every order of the runs, the second run added 2.3 new brands on ChatGPT, the third 1.4, the fourth 1.1 and the fifth 1.0. On Perplexity the fifth run added 0.3. Five runs do not exhaust the tail on ChatGPT or Gemini.

### How precise five runs are

A frequency from five runs is coarse. Treating runs as independent, the 95% interval for a brand named in 3 of 5 runs is 23.1% to 88.2%; for 5 of 5 it is 56.6% to 100%; for 1 of 5, 3.6% to 62.4%. The median interval in our data spans 0.588. To estimate a frequency near 50% to within 10 points would take about 97 runs, to within 20 points about 25; near 80%, about 62 runs for 10 points. Runs minutes apart may not be independent, and drift over days adds variation, so these are lower bounds on the effort, not a recommendation. Measuring how the estimate settles as runs are added is the next step.

## Observed, inferred and unknown

- **Observed:** brand lists, order, first picks, cited sources and wording across 300 answers, and their change between runs, hours later, across countries and across rewordings on the shared questions.
- **Inferred:** that visibility should be reported as a frequency with an interval; that evidence stability and recommendation stability are distinct; that time and condition effects need time-matched baselines.
- **Unknown:** drift over days and weeks; how many runs a given precision needs in practice; which internal step (retrieval, ranking or writing) produces the change.

## What this means

What follows is our interpretation of the numbers above.

- **Report frequency, not presence.** “Named in 4 of 5 runs” says more than “appeared in ChatGPT”. Give the number of runs and an interval alongside it.
- **Separate membership from position.** A brand can be in every answer and still move between first and sixth. Track the share of runs a brand is named, is in the top 3 and is first, separately.
- **Stable sources do not mean stable recommendations.** Tracking which pages an assistant cites is not a substitute for tracking what it recommends.
- **Compare like with like in time.** A change between two measurements hours or days apart mixes real change with drift and noise; measure the baseline variation at the same time.
- **Aim for the core.** Brands named in every run are the assistant’s default answer for that question. Moving from the tail into the core is a real change; a single appearance is not.

## What comes next

- A longitudinal panel: the same questions on several days over several weeks, to separate minute-scale variation from drift (planned with study 24, on or after 3 October 2026).
- 20 runs on a subset of questions, to measure how the estimated frequency settles as runs are added and how many runs a target precision needs in practice.
- Repeated rewordings: each rewording asked several times, so wording and repetition can be compared on equal terms.
- A fixed-evidence experiment: the same pages given to a model repeatedly, to separate variation in what is found from variation in what is written.

## How this compares with other studies

| Source | Sample and date | Figure |
|---|---|---|
| This study | 20 questions, 5 runs, 3 assistants, runs within minutes, September 2026 | 25.2% of ChatGPT brands in all 5 runs |
| SparkToro and Gumshoe | 2,961 runs of 12 prompts by 600 volunteers, November to December 2025 | Less than a 1 in 100 chance of the same list twice |
| Detailed.com | 70,000+ responses over 28 days, September 2026 | On ChatGPT the core brands changed 13% from one day to the next, the tail 78% |
| Detailed.com | Same | “Only about a quarter of the pages ChatGPT cited for a prompt were cited again the next day” |

The studies agree on the shape: a small, fairly stable core and a long, unstable tail. Ours adds that the variation is there within about a quarter of an hour, that it is mostly in which brands appear rather than their order, and that it happens even when the cited sources are identical. Recent research on measuring AI search visibility makes the same methodological point: visibility should be measured repeatedly and reported as a distribution, and evaluation should look at the full search, ranking and writing pipeline rather than the final answer alone.

Sources: [SparkToro](https://sparktoro.com/blog/new-research-ais-are-highly-inconsistent-when-recommending-brands-or-products-marketers-should-take-care-when-tracking-ai-visibility/); [Detailed.com](https://detailed.com/ai-volatility/); [Don’t Measure Once (arXiv)](https://arxiv.org/abs/2604.07585); [SAGEO Arena (arXiv)](https://doi.org/10.48550/arXiv.2602.12187).

## Methodology

- **Questions:** 20 of the 80 questions in the [four-assistant comparison](https://underneath.agency/research/ai-assistants-brand-agreement-study), at least two per industry, 7 naming a place; listed in the dataset.
- **Collection:** five runs per question per assistant on 26 September 2026; the first run is the answer used in the four-assistant study. ChatGPT and Gemini via DataForSEO LLM Scraper (consumer apps, US location); Perplexity sonar via its API with web search. The five runs of a question were a median of about 11 minutes apart (at most 16). No two answers were word-for-word identical.
- **Valid runs:** answers of at least 200 characters. One ChatGPT run (an unfinished reply) is excluded from the frequencies, so that question has four valid ChatGPT runs.
- **Brands:** candidate names from each answer, classified brand or not once per unique name, clustered per question over every answer to it (including the country and rewording answers), then matched by normalized name. Rank is order of first mention.
- **Measures:** frequency with a Wilson interval; core, stable, intermittent and tail classes; entropy and effective number of brands; over all pairs of runs, Jaccard overlap of brands, top 3, cited domains, cited URLs and words, rank-biased overlap (p = 0.9), Kendall’s tau among shared brands, and same first pick.
- **Models:** brand overlap on domain overlap, pooled with assistant indicators and within question and assistant; split of the variation in mean overlap by question and assistant.
- **Intervals:** 95% bootstrap resampling the 20 questions (2,000 resamples, seed 20260926), paired across assistants.
- **Code:** cite/pipeline/s8_v11.py and s8_v11_package.py; version 1.0 figures are kept in stats.json.

## Limitations

- Five runs over about 11 minutes: frequencies are coarse, and drift over days and weeks is not measured.
- 20 questions and 3 assistants; the analysis by question type is exploratory.
- Outputs only: the data cannot show whether the change comes from retrieval, ranking or writing. Cited sources are what the answer shows, not everything the assistant read.
- The hours-later, country and rewording answers are single answers on 5 or 10 questions.
- Perplexity was queried through its API, not the consumer app; Claude was not included.
- Brand matching by name can miss variants; rank is the order of first mention, not a ranking the assistant stated.

## What changed in version 1.1

Version 1.0 (26 September 2026) reported how much the brands change. Version 1.1 (28 September 2026) re-analyzes the same answers as a measurement problem, after an external audit: per-brand frequencies with intervals, core-to-tail classes, concentration, order and first-pick stability, a five-part profile of change, source-brand coupling including runs with identical sources, question characteristics, a time-matched comparison with the country and rewording answers, and what one run gets wrong.

Two method changes moved the headline figures slightly: one unfinished ChatGPT run is now excluded, and the brand list for each question is shared with the country and rewording answers. Brands in all five runs moved from 23.7% to 25.2% (ChatGPT), 13.6% to 13.7% (Gemini) and 41.2% to 40.9% (Perplexity); mean overlap between runs from 0.530 to 0.545, 0.421 to 0.420 and 0.682 to 0.680. The conclusions are unchanged. Version 1.0 figures remain in stats.json.

## Data and downloads

- Every answer with its brands in order, cited domains and condition: [s8_answers_v11.csv](https://underneath.agency/research-data/ai-recommendation-consistency-study/s8_answers_v11.csv)
- Every brand’s frequency, interval and ranks per question and assistant: [s8_brand_visibility.csv](https://underneath.agency/research-data/ai-recommendation-consistency-study/s8_brand_visibility.csv)
- Every pair of runs with all overlap measures: [s8_run_pairs.csv](https://underneath.agency/research-data/ai-recommendation-consistency-study/s8_run_pairs.csv)
- Stability per question and assistant: [s8_question_engine.csv](https://underneath.agency/research-data/ai-recommendation-consistency-study/s8_question_engine.csv)
- Version 1.0 answer file: [s7_s8_answers.csv](https://underneath.agency/research-data/ai-recommendation-consistency-study/s7_s8_answers.csv)
- Brand classification audit: [brand_classification_audit.json](https://underneath.agency/research-data/ai-recommendation-consistency-study/brand_classification_audit.json)
- Every statistic on this page: [stats.json](https://underneath.agency/research-data/ai-recommendation-consistency-study/stats.json)
- Machine-readable methodology: [methodology.json](https://underneath.agency/research-data/ai-recommendation-consistency-study/methodology.json)

The data is free to reuse with attribution (CC BY 4.0).

To cite: Underneath. (2026). *Ask an AI the same question 5 times: do the brands change?* (Version 1.1). Underneath Research. https://underneath.agency/research/ai-recommendation-consistency-study

## Frequently asked questions

### Does ChatGPT give the same recommendations every time?

No. Across five runs of the same question minutes apart, 25.2% of the brands ChatGPT named appeared in all five answers and 36.6% in only one. Its first pick changed at least once for 80.0% of questions.

### Which AI assistant gives the most consistent recommendations?

Perplexity, in our test: 40.9% of its brands appeared in all five runs and its first pick stayed the same for 60.0% of questions. Gemini was the least consistent. The ranking describes these 20 questions on one day; how stable a question is also depends on the question itself.

### How many times should you ask an AI assistant to measure brand visibility?

More than once, and more than five times for a precise figure. One ChatGPT answer showed 57.8% of the brands five answers named, and the fifth run still added about one new brand. A brand named in 3 of 5 runs has a 95% interval of 23.1% to 88.2%.

### Why do AI answers change between identical questions?

We measured the outputs, not the mechanism. The change is only partly tied to sources: in 166 pairs of Perplexity runs with identical cited URLs, the brand list still changed 91.6% of the time. Language models generate text with some randomness, and assistants that search the web can also retrieve and rank sources differently, but our data cannot separate those steps.

## Related research

- [Do ChatGPT, Gemini, Perplexity and Claude agree on brands?](https://underneath.agency/research/ai-assistants-brand-agreement-study)
- [Does rewording a question change AI brand recommendations?](https://underneath.agency/research/ai-prompt-phrasing-study)
- [Same question, four countries: do AI recommendations change?](https://underneath.agency/research/ai-recommendations-by-country-study)
- [ChatGPT local recommendations vs Google Maps](https://underneath.agency/research/chatgpt-local-recommendations-study)

## Related guides

- [Why does ChatGPT give a different answer about my brand each time?](https://underneath.agency/resources/why-ai-answers-about-your-brand-change)
- [Which AI engine gives the most consistent answers about brands?](https://underneath.agency/resources/most-consistent-ai-engine-for-brands)
- [Can I trust a one-time AI visibility report for my brand?](https://underneath.agency/resources/one-time-ai-visibility-report)
- [How many prompts do we need to track to measure AI visibility?](https://underneath.agency/resources/how-many-prompts-to-track-ai-visibility)
- [Why do AI search engines cite different sources every time I check?](https://underneath.agency/resources/why-ai-search-citations-change)

---

This is the Markdown twin of https://underneath.agency/research/ai-recommendation-consistency-study. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Same question, four countries: do AI recommendations change?"
description: "Five runs per country, 40 questions: ChatGPT’s brand lists overlapped 0.594 within one country but 0.429 across the US, UK, Canada and Australia."
canonical: "https://underneath.agency/research/ai-recommendations-by-country-study"
published: 2026-09-26
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Research · AI assistants

# Same question, four countries: do AI recommendations change?

A brand that ChatGPT recommends in New York may not be the one it recommends in London, Toronto or Sydney. We asked ChatGPT and Gemini the same 40 buyer questions, with no country in the wording, five times each from the United States, United Kingdom, Canada and Australia on 28 September 2026, and asked them again with the country named. Repeating every question in every country lets us separate the country effect from the ordinary variation between two runs of the same question, and see which parts of the answer move: which brands appear, which comes first, and which sources are cited.

## The short version

1. The country changed ChatGPT’s recommendations by more than run-to-run variation. Two ChatGPT answers from the same country shared a mean 0.594 of their brands (Jaccard similarity); two answers from different countries shared 0.429. The gap, 0.164 (95% interval 0.114 to 0.221), held for 85.0% of questions.
2. Gemini responded to the country far less. Its same-country and different-country figures were 0.492 and 0.436, a gap of 0.056 (0.029 to 0.086).
3. The country effect grew with how much the question should depend on the country. On questions such as insurance, lending and tax, the gap was 0.180; on global products such as headphones and CRM software, 0.034.
4. When the country changed the brands, it usually changed the sources too. About half (52.7%) of ChatGPT’s same-country advantage in brand overlap went with differences in the websites it cited. Answers from different countries that cited similar sources named similar brands.
5. Location alone did part of what naming the country did. With only the location set, local-market brands made up 25.4% of the brands ChatGPT named in the UK; with “in the United Kingdom” added to the question, 49.1%. For Gemini the figures were 12.9% and 50.9%.
6. The brand that comes first moved too. ChatGPT put the same brand first in all four countries for 25.4% of questions. Of the brands it put first in most runs in some country, 38.5% were never named in five runs from at least one other country.

## Research questions

| Question | Answered here? |
|---|---|
| RQ1. How much do brands change between countries? | Yes |
| RQ2. Is that more than repeat-run variation? | Yes, five runs per country |
| RQ3. Do ChatGPT and Gemini differ? | Yes |
| RQ4. Do the cited sources change with the brands? | Yes, as an association |
| RQ5. Which questions are most affected? | Yes, by expected dependence |
| RQ6. Does the country change order, or only membership? | Yes |
| RQ7. Does location alone act like naming the country? | Yes |

## How to read “the country effect”

Whether a brand is recommended depends on the question, the engine, the date, the person asking and chance: two runs of the same question rarely match. The person asking includes their language, location, history and account. This study changes one of these, the country, holds the language at English, and keeps the rest as fixed as the data provider allows. Every comparison is made against the variation between repeated runs from one country, so what we report as the country effect is only the change beyond that variation.

We also measure the answer at several points, not as one score. These are presence (is a brand named), first place (which brand leads), rank-weighted overlap (do the leading brands match), the cited websites, and local-market share (how many named brands sell mainly in that country). These can move separately.

## What we measured

The 40 questions are the national buyer questions from our [four-assistant study](https://underneath.agency/research/ai-assistants-brand-agreement-study), such as “What is the best CRM for a small business?” and “Which pet insurance company is the best?”. None names a country. On 28 September 2026 we asked ChatGPT and Gemini each question five times with the location set to each country (1,600 requests), and twice more with the question ending “in the United Kingdom?” (or the country in question) and the location set to that country (640 requests). The single run of 26 September 2026 (version 1.0 of this study) is kept as a pilot and as an earlier date.

Brands were identified as in the four-assistant study. We extracted candidate names from each answer, classified each as a brand or not, grouped name variants per question, and counted a brand where its name appears in the answer. Brand order is the order of first mention. Two answers are compared with Jaccard similarity (brands in both divided by brands in either, 1.0 is identical), whether they put the same brand first, and rank-biased overlap, which weights the top of the list. Cited websites are compared the same way.

Before looking at any answers, each question was coded for how much a good answer should depend on the country. That gave 17 high (insurance, lending, tax, legal and telehealth), 10 medium and 13 low (global products and software). Every brand was coded for its primary market (US, UK, Canada, Australia, or sold in most of the four). Both codings were done by Claude Opus and checked by Claude Sonnet, which agreed on 85.0% of questions and 90.8% of 2,080 brands. No labels are human-coded.

Answers to the same question are not independent, so every 95% interval on this page comes from resampling the 40 questions.

## Findings

### Country effect beyond ordinary variation

| Two answers from | ChatGPT | Gemini |
|---|---|---|
| The same country | 0.594 | 0.492 |
| Different countries | 0.429 | 0.436 |
| Gap | 0.164 | 0.056 |
| 95% interval of gap | 0.114 to 0.221 | 0.029 to 0.086 |

Mean Jaccard similarity of brand sets, 40 questions, 28 September 2026.

A second way to see it is to ask how much of the variation among the 20 answers to one question (five runs in each of four countries) lines up with the country. For ChatGPT the country accounted for 0.388 of it, for Gemini 0.234, against 0.158 expected by chance. In a permutation test the country grouping was significant (p < 0.05) for 72.5% of questions on ChatGPT and 37.5% on Gemini. The difference between the engines, 0.154 (0.100 to 0.209), is itself clear.

| Same-country baseline and country pairs (ChatGPT) | Jaccard |
|---|---|
| Two runs from the US | 0.658 |
| Two runs from Canada | 0.607 |
| Two runs from the UK | 0.566 |
| Two runs from Australia | 0.544 |
| US and Canada | 0.499 |
| US and UK | 0.440 |
| US and Australia | 0.425 |
| Canada and Australia | 0.405 |
| UK and Canada | 0.404 |
| UK and Australia | 0.402 |

Every cross-country pair overlapped less than any pair of runs from one country. The US and Canada were the most alike.

### Which brand comes first

| Measure | ChatGPT | Gemini |
|---|---|---|
| Pairs with the same first brand, same country | 0.631 | 0.452 |
| Pairs with the same first brand, different countries | 0.439 | 0.390 |
| Rank-weighted overlap, same country | 0.731 | 0.619 |
| Rank-weighted overlap, different countries | 0.535 | 0.550 |
| Same first brand in all four countries | 25.4% | 14.0% |

Pair figures are shares of pairs of answers (1.0 means every pair); the last row is the share of questions, one run per country.

The country changed the order as well as the membership. For small business accounting software, ChatGPT put QuickBooks first in every US and Canadian run and Xero first in every UK and Australian run, while naming both everywhere. For air fryers it led with a Cosori model in the US and with Ninja in Australia, where the Cosori model did not appear. Gemini, located in Australia, led with the US Chase Sapphire Preferred card in four of five runs for travel rewards cards.

### Where the country matters most

| Expected country dependence | Questions | Gap, ChatGPT | Gap, Gemini |
|---|---|---|---|
| Low (global products, software) | 13 | 0.053 | 0.015 |
| Medium | 10 | 0.157 | 0.022 |
| High (insurance, lending, tax) | 17 | 0.254 | 0.106 |

Gap = same-country minus different-country brand overlap.

The country’s share of variation rose by 0.087 (0.047 to 0.126) per step from low to high dependence. By industry it was highest for financial services and insurance (0.434) and lowest for business software (0.233). Even for low-dependence questions the country mattered on ChatGPT: the gap there, 0.053 (0.022 to 0.091), is small but above zero, and the country grouping was significant for 38.5% of low-dependence question and engine pairs.

### Sources and brands move together

The websites cited changed with the country as well. Two ChatGPT answers from the same country shared 0.534 of their cited domains, answers from different countries 0.324. The country accounted for 0.430 of the variation in ChatGPT’s cited domains, as much as for its brands.

Two tests link the two changes, and both are associations, not proof of cause.

- **Pairs of answers.** Among cross-country pairs, the quarter whose cited sources overlapped most named brands with a mean overlap of 0.712 for ChatGPT, higher than two same-country runs (0.594). The quarter that shared no sources overlapped by 0.152. Similar sources went with similar brands regardless of country.
- **Decomposition.** Answers from the same country share more sources and more brands. Holding source overlap fixed, the same-country advantage in brand overlap fell from 0.163 to 0.077 for ChatGPT. So 52.7% (39.3% to 66.6%) of it went with the difference in sources, and the rest did not. For Gemini the share was 38.0% (26.8% to 53.9%).

The remaining part of the country effect, with sources held fixed, points to the model using the location when it writes the answer, not only when it searches. Answers were more likely to mention the country, its currency or its regulators when it was set: ChatGPT did so in 68.8% of UK answers, 69.2% of Canadian and 76.8% of Australian answers, though the question named no country. Gemini did so in 20.0%, 30.2% and 19.2%.

### Local sources and local brands

| Location only | Local domains | Local brands | US brands |
|---|---|---|---|
| ChatGPT, UK | 26.1% | 25.4% | 23.2% |
| ChatGPT, Canada | 21.8% | 17.6% | 34.7% |
| ChatGPT, Australia | 36.8% | 22.2% | 25.7% |
| Gemini, UK | 9.9% | 12.9% | 37.2% |
| Gemini, Canada | 11.4% | 19.6% | 33.0% |
| Gemini, Australia | 10.5% | 11.6% | 37.5% |

Local domains = share of cited domains ending in .uk, .ca or .au. Local and US brands = share of named brands whose model-coded primary market is that country or the US.

Outside the US, ChatGPT answers that cited at least one local domain (44.2% of them) had a local-brand share of 45.0%, against 3.1% for answers citing none. Comparing answers to the same question, each additional share of local sources went with more local brands (slope 0.276, 95% interval 0.048 to 0.504). For Gemini, whose answers cited a local domain less often (20.2%), the link was stronger (slope 0.765).

### Location alone vs naming the country

| Local-brand share | Location only | Country named | Location’s share |
|---|---|---|---|
| ChatGPT, UK | 25.4% | 49.1% | 51.7% |
| ChatGPT, Canada | 17.6% | 44.3% | 39.7% |
| ChatGPT, Australia | 22.2% | 48.9% | 45.4% |
| Gemini, UK | 12.9% | 50.9% | 25.3% |
| Gemini, Canada | 19.6% | 50.0% | 39.2% |
| Gemini, Australia | 11.6% | 48.9% | 23.7% |

Location’s share = location-only figure as a percentage of the country-named figure.

Naming the country moved the answers much further than setting the location. Two answers from different countries with the country named overlapped by only 0.205 (ChatGPT) and 0.173 (Gemini), against 0.429 and 0.436 with the location alone. Part of that is the rewording itself: adding “in the United States” to a US question shifted ChatGPT’s answer by 0.103, against 0.227 for the other three countries. For Gemini the US shift was 0.003, against 0.201 elsewhere.

### A check across dates

The pilot run two days earlier overlapped with the new same-country runs by 0.539 for ChatGPT, against 0.594 between runs on the same day, a difference of 0.055 (0.014 to 0.103). For Gemini the difference was 0.025. Two days added a little change, much less than the country did.

## How this compares with other studies

We found no other published study that compares brand recommendations for the same English questions across the US, UK, Canada and Australia with repeated runs. Profound’s analysis of 3.25 billion citations across 14 countries found that the language of a question shapes AI citations more than the country does. Our results show that, with the language held at English, the country still changes which brands and sources an answer uses, much more for ChatGPT than for Gemini. An arXiv study of brand preferences (ChoiceEval) found that the US-developed Gemini and GPT models show a marked preference for American entities. That is consistent with the US-market brands that still made up about a third of Gemini’s answers in the UK, Canada and Australia.

Sources: [Profound](https://www.tryprofound.com/blog/how-query-language-reshapes-ai-citations); [ChoiceEval, arXiv](https://arxiv.org/abs/2603.18300).

## Observed, inferred and unknown

**What we observe.** For the same English question, ChatGPT’s recommended brands, first brand and cited websites change with the country more than they change between repeated runs, and more for questions about national, regulated products. Gemini’s change much less. Sources and brands change together. Location alone produces between about a quarter and a half of the local-brand share that naming the country produces.

**What we infer.** The country is part of what an AI recommendation depends on, like the engine and the date. A measurement taken in one country describes that country. Much of the country effect runs through the sources an engine finds for that country, and part appears to come from the model using the location when writing the answer.

**What remains unknown.** Whether changing the sources would change the brands (we observed sources, we did not alter them). How much a city, a device, an account or a user’s history adds. Whether the same holds in other languages, other engines, other question types and over longer periods. Why Gemini uses the location so much less.

## What this means

The points below are our interpretation. They follow from the findings but were not tested.

- **Measure each market separately, with repeated runs.** In this sample one run per country cannot tell a country effect from ordinary variation, and one country’s results did not describe another’s.
- **Expect the most difference where products are national.** For insurance, lending, tax and similar questions, the brands a buyer sees depend heavily on where they are. For global products the difference is smaller but not zero on ChatGPT.
- **Local sources are part of local visibility.** Answers that cited a country’s own websites named that country’s brands far more often. Coverage in each market’s review sites, comparison sites and publications is a reasonable place to start.
- **A named country is a different question.** People who type the country get far more local answers than people who rely on location. Track both wordings where buyers use both.

## Where this sits in GEO research

Research on generative engine optimization began by asking whether changing a page changes how often an AI answer uses it. Later work compared the sources that different AI search systems use across engines, languages and industries. It argued that visibility should be measured as a distribution over repeated runs and separate stages, not one rank, and that the person asking, including their location, is part of what it depends on. This study isolates one part of that, the country, with the language held constant. It measures the country effect against repeated runs and relates it to the sources cited. It is observational. It shows where and how much the country matters, not what would change a brand’s visibility in a given country.

## Methodology

- **Questions:** 40 national buyer questions from the four-assistant study, with no country in the wording (listed in the dataset with their codes).
- **Engines and locations:** ChatGPT and Gemini consumer apps via DataForSEO LLM Scraper, English, location set to the United States, United Kingdom, Canada and Australia.
- **Design:** 5 runs per question, engine and country on 28 September 2026 with no country in the question (1,600 requests); 2 runs with the country named in the question and the location set to it (640 requests); the 26 September 2026 run (320 answers) as the pilot. All 2,240 requests returned; 21 unfinished answers under 200 characters (18 ChatGPT, 3 Gemini) are excluded.
- **Brands:** candidate names from bold text, headings and ChatGPT’s brand entities. Candidates were classified with the four-assistant study’s labels, and Claude Opus classified names not seen before. Variants were grouped per question and counted where the name appears; order is order of first mention. Two answers that both name no brand count as identical.
- **Pair measures:** Jaccard similarity of brands and of cited registrable domains; same first brand; rank-biased overlap (p = 0.9). Same-country figures average the 10 pairs of runs in each country; different-country figures average the 150 cross-country pairs.
- **Country share of variation:** the share of variation in Jaccard distances among the 20 answers to a question explained by country (PERMANOVA R²), with a 499-permutation test per question.
- **Sources and brands:** brand overlap regressed on source overlap with question fixed effects, and the same-country advantage before and after holding source overlap fixed, with intervals from 300 question resamples.
- **Local sources and brands:** local domains are those ending in .uk, .ca or .au; brand markets and question dependence are model-coded (Claude Opus, checked by Claude Sonnet).
- **Uncertainty:** 95% bootstrap intervals resampling the 40 questions (2,000 resamples, seed 20260928).
- **Update schedule:** quarterly.

## Limitations

- Country-level locations set by the data provider, not a local account, city or device; user history is not varied.
- Two engines, one language, four countries and 40 questions, all asking for the best option in a category.
- Five runs on one day, plus one earlier run; change over weeks is not measured.
- The link between sources and brands is an association. We did not change the sources to test it.
- Country-code domains are a rough measure of local sources (many local sites use .com). Brand markets, brand classification and question codes are model-coded, not checked by people.
- A brand counts when it is named, including in a caution, so a mention is not always a recommendation.

## What changed in version 1.1

Version 1.0 (26 September 2026) asked each question once per country and compared the country differences with five US runs from our consistency study, collected earlier and analyzed with a different brand list, on only 10 questions. Following an external review, version 1.1 (28 September 2026) is a new collection. It uses five runs in every country on the same day, a country-named condition, rank and source measures, questions coded for expected country dependence, a source decomposition and intervals clustered by question.

Two version 1.0 figures change. The matched comparison (0.617 same-country against 0.292 across countries for ChatGPT, 0.453 against 0.328 for Gemini) overstated the country effect, because the two sides came from different collections. The balanced design gives 0.594 against 0.429 for ChatGPT and 0.492 against 0.436 for Gemini: a smaller effect, still clear for ChatGPT. The share of brands named in all four countries (25.1% in version 1.0) was 24.9% for ChatGPT in the new runs, but four runs from one country share only 38.6%, so part of that figure is run-to-run variation. Version 1.0 also said that local sources feed local answers; version 1.1 tests that as an association and finds it holds, with the caveats above.

## Data and downloads

- Every answer in version 1.1 (condition, country, run, brands in order, cited domains): [s17v11_answers.csv](https://underneath.agency/research-data/ai-recommendations-by-country-study/s17v11_answers.csv) and [JSON](https://underneath.agency/research-data/ai-recommendations-by-country-study/s17v11_answers.json)
- Per question and engine (same-country and cross-country measures, country share, tests): [s17v11_cells.csv](https://underneath.agency/research-data/ai-recommendations-by-country-study/s17v11_cells.csv)
- Per brand, question and engine (presence, first place and prominence by country): [s17v11_brands.csv](https://underneath.agency/research-data/ai-recommendations-by-country-study/s17v11_brands.csv)
- Model codes: [brand markets](https://underneath.agency/research-data/ai-recommendations-by-country-study/s17v11_brand_markets.csv) and [question dependence](https://underneath.agency/research-data/ai-recommendations-by-country-study/s17v11_question_codes.csv)
- Every statistic on this page: [stats.json](https://underneath.agency/research-data/ai-recommendations-by-country-study/stats.json)
- Charts: [same country vs different countries](https://underneath.agency/research-data/ai-recommendations-by-country-study/same-vs-different-country.svg), [country share by type of question](https://underneath.agency/research-data/ai-recommendations-by-country-study/country-share-by-question-type.svg), [local brands with location only and country named](https://underneath.agency/research-data/ai-recommendations-by-country-study/local-brands-location-vs-named.svg)
- Machine-readable methodology: [methodology.json](https://underneath.agency/research-data/ai-recommendations-by-country-study/methodology.json)
- Version 1.0 data and statistics: [s17_country_answers.csv](https://underneath.agency/research-data/ai-recommendations-by-country-study/s17_country_answers.csv) and [stats_v1.0.json](https://underneath.agency/research-data/ai-recommendations-by-country-study/stats_v1.0.json)

The data is free to reuse with attribution (CC BY 4.0).

To cite: Underneath. (2026). *Same question, four countries: do AI recommendations change?* (Version 1.1). Underneath Research. https://underneath.agency/research/ai-recommendations-by-country-study

## Frequently asked questions

### Does ChatGPT recommend different brands in different countries?

Yes, beyond ordinary run-to-run variation. Two ChatGPT answers to the same question from one country shared a mean 0.594 of their brands; answers from different countries shared 0.429. The difference was largest for insurance, lending, tax and similar national products.

### Is the difference between countries just random variation?

Not for ChatGPT. We asked each question five times in each country, so the variation between runs could be measured directly; the country added clear change beyond it for 85.0% of questions. For Gemini the country effect was real but small (a gap of 0.056).

### Does ChatGPT use local websites for local answers?

Often, and the local sources go with local brands. In Australia 36.8% of the domains ChatGPT cited ended in .au. Outside the US, ChatGPT answers citing at least one local domain had a local-brand share of 45.0%, against 3.1% for answers citing none.

### Is setting the location the same as naming the country in the question?

No. Naming the country roughly doubled the share of local-market brands in ChatGPT’s answers (for example 25.4% to 49.1% in the UK). For Gemini it rose from 12.9% to 50.9% in the UK and from 11.6% to 48.9% in Australia.

### How should a brand in several countries measure its AI visibility?

Separately in each market, with repeated runs, and with the country both left out of and included in the question. One run per country cannot separate the country effect from ordinary variation.

## Related research

- [Do ChatGPT, Gemini, Perplexity and Claude agree on brands?](https://underneath.agency/research/ai-assistants-brand-agreement-study)
- [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)
- [Does rewording a question change AI brand recommendations?](https://underneath.agency/research/ai-prompt-phrasing-study)
- [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)

## Related guides

- [Do AI engines cite local-language sources in other languages?](https://underneath.agency/resources/do-ai-engines-cite-local-language-sources)
- [Is an English-only AI visibility check enough if you sell abroad?](https://underneath.agency/resources/english-only-ai-visibility-audits)
- [How does a payroll provider get named when businesses ask AI?](https://underneath.agency/resources/payroll-software-ai-search)
- [How should global brands approach GEO across languages?](https://underneath.agency/resources/geo-strategy-for-global-brands-across-languages)
- [Are AI assistants now choosing accounting software for small businesses?](https://underneath.agency/resources/accounting-software-ai-search)

---

This is the Markdown twin of https://underneath.agency/research/ai-recommendations-by-country-study. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "When does a Reddit thread become evidence in Google’s AI?"
description: "Google’s AI cited Reddit in 17.9% of AI Overviews. Cited threads had twice the comments of skipped ones, and 20.8% of cited claims were not supported."
canonical: "https://underneath.agency/research/ai-reddit-citations-study"
published: 2026-09-26
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Research · AI search

# When does a Reddit thread become evidence in Google’s AI?

Reddit is often described as one of AI’s favorite sources. Version 1.0 of this study counted which Reddit threads AI answers cite. This version asks when and why a thread gets cited, and what it is used for. We follow each thread through stages. Google shows it, the AI cites it, the citation is attached to a sentence, and the thread supports that sentence or not. We then compare cited threads with threads that were not cited, and repeat part of the sample on a second date and in other wordings.

The data are the same Google AI Overviews, AI Mode answers and assistant answers as version 1.0 (26 September 2026). We add two collections already made for our other studies (28 September 2026) and the full text, votes and comments of 1,855 Reddit threads from the public Arctic Shift archive. No new searches were bought.

## The short version

1. Reddit is a Google AI source much more than an assistant source. 17.9% of AI Overviews (95% interval 10.7% to 26.2%) and 9.8% of AI Mode answers cited a Reddit thread. Perplexity did in 12.5% of 80 buyer questions, Gemini in 2.5%, ChatGPT and Claude in none.
2. Google’s AI cites Reddit far more often when Google’s own results already show a Reddit thread. That happened in 33.5% of those AI Overviews, against 8.3% when the results page showed none. Still, 53.9% of cited threads in AI Overviews and 79.1% in AI Mode were not on the results page at all. On 505 other searches where we saw Google’s top 100, 17 of the 27 Reddit citations went to threads outside the top 100.
3. Among threads Google showed for the same search, the cited ones had more discussion: a median of 40 comments against 20. Each standard deviation more comments multiplied the odds of being cited by 2.65 (95% interval 1.8 to 3.88), the only thread feature that held after correcting for multiple tests. Question-style titles, matching the search’s words and coming from a specialist community did not reliably predict citation once the search was held fixed.
4. The cited threads are used mainly for facts (25.2% of cited sentences), recommendations (24.4%) and people’s experiences (19.6%). Judged against the post and its top 10 comments, 36.4% of sentences citing Reddit were supported, 42.8% partly supported and 20.8% not supported. Both model coders agreed a sentence was unsupported in 18.8%.
5. Reddit citations are not stable. For the same 96 keywords two days apart, 62.5% of the AI Overviews that cited Reddit still did, and 50.0% of the cited threads were cited again. Rewording a keyword as a question kept only 7.1% of the cited threads.

## Research questions

| Question | Answered here? |
|---|---|
| RQ1. How often do AI answers cite Reddit, by engine, industry and information need? | Yes |
| RQ2. Are cited threads the ones Google already shows, and how deep does the AI reach? | Yes |
| RQ3. What distinguishes cited threads from threads that were not cited? | Yes, observationally |
| RQ4. What does the Reddit thread contribute, and does it support the sentence? | Yes, model-coded |
| RQ5. Are Reddit citations stable across dates and wordings? | Partly: two dates, three wordings, 96 keywords |
| RQ6. Does changing a thread change whether it is cited? | No: no experiment |

## Visibility in stages

“AI cites Reddit” mixes several different events. We measure four of them separately. Two further stages are named because they matter, but this study does not measure them.

| Stage | Question | Measured? |
|---|---|---|
| Shown | Does Google’s results page show the thread? | Yes |
| Cited | Does the AI answer link to the thread? | Yes |
| Attached | Is the link tied to a specific sentence? | Yes |
| Supported | Does the thread support that sentence? | Yes, model-coded |
| Influence | Did the thread change what the answer said? | No |
| User outcome | Did anyone click, trust or act on it? | No |

## What we measured

We took every link to an individual Reddit thread in five sets of answers.

- **Google AI Overviews and AI Mode, 26 September 2026:** the 800 US keywords of our [AI Overview frequency study](https://underneath.agency/research/ai-overviews-frequency-study) (486 AI Overviews) and AI Mode answers for 400 of them, as in version 1.0.
- **ChatGPT, Gemini, Perplexity and Claude, 26 September 2026:** the 80 buyer questions of our [four-assistant study](https://underneath.agency/research/ai-assistants-brand-agreement-study), as in version 1.0.
- **Same keywords, two days later, 28 September 2026:** 96 of the 800 keywords searched again, each in three wordings: the original keyword, a natural question and a longer rephrasing (from the frequency study’s version 1.1).
- **Deep results, 28 September 2026:** 576 searches with Google’s organic top 100 and the AI Overview on the same page (from our [AI citations vs Google rankings study](https://underneath.agency/research/ai-citations-google-rankings-study)); 505 showed an AI Overview.
- **Thread text:** title, post, votes, date and full comment tree of all 1,855 threads involved, from the public Arctic Shift archive of Reddit. Reddit blocks automated access to its own pages, which is why version 1.0 could not measure these.

Every search stays in the results: an answer that cites no Reddit thread counts as an outcome, not a gap. Keywords that share a seed topic are not independent, so intervals for Google rates come from resampling the seed topics. Counts from small groups carry Wilson intervals.

## Findings

### How often each AI cites Reddit

| AI answer | Answers | Citing Reddit | 95% interval |
|---|---|---|---|
| Google AI Overviews | 486 | 17.9% | 10.7% to 26.2% |
| Perplexity | 80 | 12.5% | 6.9% to 21.5% |
| Google AI Mode | 400 | 9.8% | 5.2% to 15.6% |
| Gemini | 80 | 2.5% | 0.7% to 8.7% |
| ChatGPT | 80 | 0.0% | 0.0% to 4.6% |
| Claude | 80 | 0.0% | 0.0% to 4.6% |

The assistants answered 80 buyer questions; Google’s figures cover 800 searches, so the two sets differ. In our [study of “Is this brand legit?” questions](https://underneath.agency/research/is-it-legit-ai-reputation-study), ChatGPT did cite Reddit, in 11.4% of answers: whether an assistant turns to Reddit depends on the question.

### Industry, not information need, carries the difference

| Industry | AI Overviews | Citing Reddit |
|---|---|---|
| Home and local services | 53 | 37.7% |
| B2B software and technology | 96 | 35.4% |
| Hospitality and travel | 62 | 16.1% |
| Legal and professional services | 40 | 15.0% |
| Retail and ecommerce | 68 | 10.3% |
| Healthcare and dental | 43 | 7.0% |
| Financial services and insurance | 88 | 5.7% |
| Franchises and multi-location brands | 36 | 5.6% |

Two model coders sorted every keyword by the need behind it (choosing between options, cost, how-to, facts, a local provider, a brand). They agreed on 84.6% (Cohen’s kappa 0.8). AI Overviews for “choice” searches cited Reddit most often (28.9%), cost searches least (8.9%). Once industry is held fixed, though, the information need adds nothing we can detect (joint test p = 0.61). 77.1% of the software AI Overviews were choice searches, against 11.4% in financial services, so the two cannot be separated in this sample. Keyword intent labels (commercial, informational, transactional, navigational) showed no clear difference either: 16.4% to 20.5%.

### Stage 1 to 2: Google shows a thread, the AI cites it

| AI Overviews, 26 September | Answers | Citing Reddit |
|---|---|---|
| Results page showed a Reddit thread | 185 | 33.5% |
| Results page showed none | 301 | 8.3% |

Of the 243 Reddit threads shown on those results pages, the AI Overview cited 16.9%. Position mattered: 27.3% of threads ranking 1 to 3 were cited, 14.9% at 4 to 10 and 1.8% of threads shown only in the Discussions and forums box. AI Mode cited 5.6% of the shown threads.

Yet most citations did not come from the page. 53.9% of Reddit threads cited in AI Overviews and 79.1% of those cited in AI Mode appeared in neither the organic top 10 nor the forums box for the same search.

### How deep does the AI reach?

On 505 searches with an AI Overview where we recorded Google’s organic top 100, Google ranked 1,142 Reddit threads. The AI Overview cited 5.9% of those at positions 1 to 3, 0.5% at 4 to 10 and none of the 798 at positions 11 to 100. Of the 27 Reddit citations on these searches, 9 were to threads ranking 1 to 3, 1 to a thread at 4 to 10 and 17 to threads outside the top 100.

Google’s AI does not work down its own ranking into Reddit. It either takes a thread from the very top, or reaches one that the ranking does not show at all. That points to a separate retrieval step, but this study cannot see that step directly.

### What distinguishes a cited thread

We compared cited and uncited threads in two ways.

The main comparison holds the search fixed. It covers every Reddit thread that Google showed or its AI cited for the same search, on the 100 searches that had both cited and uncited threads (301 thread appearances, a conditional logistic model with one stratum per search).

| Thread feature | Odds ratio | 95% interval | After Holm correction |
|---|---|---|---|
| More comments (per SD) | 2.65 | 1.8 to 3.88 | p < 0.001 |
| Higher upvote score (per SD) | 1.44 | 1.08 to 1.93 | p = 0.175 |
| Longer post and top comments (per SD) | 1.49 | 1.08 to 2.04 | p = 0.178 |
| Specialist community | 1.83 | 0.73 to 4.55 | p = 1 |
| Title is a question | 1.18 | 0.7 to 1.99 | p = 1 |
| Search words in the post and comments (per SD) | 1.18 | 0.86 to 1.62 | p = 1 |
| Older thread (per SD) | 0.77 | 0.58 to 1.04 | p = 0.984 |

Across all 1,964 appearances, cited threads had a median of 40 comments against 20 for uncited ones. 65.5% of cited threads had question titles, and so did 53.8% of uncited ones. In a joint model with the other features, comments kept their effect (odds ratio 2.99 per SD, 1.94 to 4.6). Nothing about prices, numbers, lists or instructions in the text mattered.

The second comparison matches each of 88 cited threads with up to three threads from the same subreddit whose titles share the search’s words and that no answer in our data cites (264 controls). Here the gaps are much larger: a median of 44.5 comments against 6, and more upvotes, older threads and longer text all predicted citation. These controls, however, are mostly threads Google would not show for the search at all. The larger gap therefore mixes the “shown” stage with the “cited” stage. The title-matching step also makes controls match the search words by design. We report this comparison in the data files but draw conclusions from the same-search comparison.

Version 1.0 noted that a third of cited threads had question titles. That is true, but question titles are nearly as common among the Reddit threads Google showed and the AI skipped, so it is not a signal of citation.

### Which communities

| Community type | Subreddits | Citations | Shown | Cited if shown |
|---|---|---|---|---|
| Specialist practitioners | 111 | 50.8% | 33.5% | 19.3% |
| Consumer advice | 103 | 22.0% | 25.1% | 9.9% |
| Product or tool | 115 | 20.5% | 25.8% | 7.7% |
| Local place | 153 | 4.5% | 9.7% | 10.3% |
| General interest | 66 | 2.3% | 6.0% | 8.3% |

Citations and Shown are each type’s share of Google’s Reddit citations and of the Reddit threads on the results pages; the last column is the share of shown threads that the AI Overview cited.

Two model coders classified all 548 subreddits in the data. They agreed on 89.4% (kappa 0.86). Specialist practitioner communities are those where people who do the work answer questions about it, such as plumbers, project managers and security staff. They supply half of Google’s Reddit citations, and a thread from one that Google showed was cited 19.3% of the time, against 7.7% to 10.3% for other types.

That fits the idea that Google’s AI favors communities that act as specialist knowledge bases rather than social media in general. But when the search is held fixed, the specialist effect is not statistically clear (odds ratio 1.83, 95% interval 0.73 to 4.55). Specialist communities cluster in the industries where Reddit is cited most: they supplied 58.7% of the Reddit threads shown in software searches and 50.0% in home services, against 17.2% in financial services and 2.7% in retail. The idea remains a hypothesis this sample cannot confirm.

The most-cited communities in Google’s answers on 26 September were r/askaplumber (14 citations, each on a different search), r/projectmanagers (9) and r/crmsoftware (7). 51.7% of the 58 cited communities were cited once.

### What the thread is used for, and whether it supports the sentence

96.9% of Reddit citations in Google’s AI answers were attached to a specific sentence, and 24.8% of those sentences named Reddit in the text (“users on Reddit say…”). For each of the 250 sentences, two model coders read the sentence, Google’s snippet and the archived thread (post and top 10 comments). They then recorded what the thread contributed and whether it supported the sentence.

| Role of the thread | Share of 250 sentences | Supported | Not supported |
|---|---|---|---|
| Fact (price, rule, specification) | 25.2% | 46.0% | 12.7% |
| Recommendation of a named option | 24.4% | 27.9% | 23.0% |
| People’s experience or opinion | 19.6% | 34.7% | 4.1% |
| Procedure or fix | 13.2% | 57.6% | 9.1% |
| No recognizable use | 10.0% | 0.0% | 100.0% |
| Comparison of options | 7.6% | 47.4% | 0.0% |

Overall, 36.4% of sentences were supported, 42.8% partly supported (often one commenter’s view stated as what “users” say) and 20.8% not supported. The coders agreed on 74.0% of support labels (kappa 0.61), and both called a sentence unsupported in 18.8%.

Recommendations were the weakest. Many cite a general “which tool do you use?” thread for a sentence praising a product the post and top comments do not mention. The coders saw only the top 10 comments. A text search of the complete comment trees found the option named in the sentence somewhere in the thread for 77.3% of the 132 sentences that name one. Of the 36 such sentences coded unsupported, 27.8% named an option that appears nowhere in the thread. The support figures describe what a reader of the top of the thread would find, not a final verdict on every comment.

### How stable are Reddit citations?

We compared the 96 keywords searched on 26 and 28 September, and on 28 September in three wordings.

| Comparison | Answers | Reddit again | Same thread | Same community |
|---|---|---|---|---|
| Two days apart | 16 | 62.5% | 50.0% | 56.2% |
| As a question | 12 | 33.3% | 7.1% | 14.3% |
| Longer wording | 12 | 50.0% | 0.0% | 14.3% |

Answers are the AI Overviews that cited Reddit in the first search of each pair (the keyword on 26 September, or the plain keyword on 28 September for the two rewordings). The other columns show how many of those citations came back in the second search.

Where both dates showed an AI Overview (50 keywords), they agreed on whether to cite Reddit in 86.0%. The wording mattered more than the date. Question wordings cited Reddit in 35.6% of their AI Overviews and longer wordings in 35.7%, against 21.1% for the plain keywords that day. But they rarely cited the same thread or even the same community. The idea that AI answers are stable at the community level while threads vary is not supported here: both changed with the wording. These comparisons rest on 12 to 16 answers each, so the intervals are wide (for example 38.6% to 81.5% for “still citing Reddit” two days apart).

## What this means

These points are our reading of the data rather than measured results.

- **Google’s AI treats Reddit as a source of specific evidence, not as “social media” in general.** It cites Reddit mostly in home services and software, mostly from communities of practitioners, and mostly for facts, recommendations and experience. Whether the specialist pattern survives with more data is the open question.
- **A cited thread is usually one with a real discussion.** Among threads Google already surfaces, the ones cited have more comments. Nothing on this page shows that adding comments to a thread would get it cited: that is an observational association, and the causal test has not been run.
- **Google’s ranking is a partial guide.** A Reddit thread at the top of Google’s results is often cited, one further down almost never, and many cited threads are not in the results at all. Tracking only the Reddit threads Google ranks misses most of what its AI uses.
- **Treat any single Reddit citation as a snapshot.** Two days or a rewording were enough to change which thread was cited. A monitoring report on Reddit citations needs repeated measurements, not one check.
- **A citation is not an endorsement.** About one in five sentences citing Reddit was not supported by the top of the thread, most often when the AI turned a general question thread into a product recommendation. The legitimate way for a business to be part of these threads is for its staff to take part honestly, as themselves, in the communities where its buyers ask questions. Posing as customers is not, and Reddit’s rules on impersonation and spam forbid it.

## Methodology

- **Answers:** Google AI Overviews (486, from 800 US keywords) and AI Mode (400) via DataForSEO SERP API, Google.com, United States, English, desktop, 26 September 2026. ChatGPT and Gemini consumer apps via DataForSEO LLM Scraper; Perplexity sonar and Claude Haiku 4.5 via DataForSEO LLM Responses API with web search, 26 September 2026. 96 keywords again in three wordings (288 searches) and 576 searches with the organic top 100, 28 September 2026, same settings. One request per search.
- **Reddit citations:** links to individual threads (addresses containing /comments/) from every reference, per-paragraph reference and inline link in the answer. The attached sentence is the answer line that carries the link.
- **Shown by Google:** the thread in the organic results (top 10, or top 100 for the deep set) or the Discussions and forums box of the same results page. AI Mode has no results list of its own, so its “shown” set is the regular results page for the same keyword.
- **Thread data:** Arctic Shift archive (post by ID, full comment tree), retrieved 28 September 2026; comments and votes are as archived, which may differ from the day of the search. 100 of the 1,855 posts were removed or deleted by the time of archiving.
- **Features:** comments (log), upvote score (log), age at the search date, title a question, length of post plus top 10 comments, share of the search’s words in the title and in the text, first-person words, price mentions, numbers, instruction verbs, lists.
- **Comparisons:** conditional logistic regression, one stratum per search (same-search design) or per matched set (same-community design), continuous features standardized, Holm correction across the features tested.
- **Model coding:** community type (548 subreddits), information need (1,568 keywords), and role and support of each attached sentence (250) were each coded independently by two models, Claude Opus and Claude Sonnet. The first coder’s label is used, and agreement is reported. No human coding was done.
- **Intervals:** Google rates resample seed topics (2,000 resamples, seed 20260928); other proportions use Wilson 95% intervals.
- **Update schedule:** quarterly. A week-over-week rerun of the AI Overview sample is planned with our volatility study, and this page will add those figures.

## Limitations

- Observational: every result is an association. Threads were not changed to see whether their citation changed, so “more comments are cited” does not mean “adding comments gets a thread cited”.
- Comments and votes come from the archive, not from the day of the search, and threads keep growing, which may inflate differences for older threads.
- Support was judged by models against the post and top 10 comments, not by people and not against the whole thread; the full-thread text check covers only sentences that name an option.
- Stability rests on 96 keywords over two days, with 12 to 16 Reddit-citing answers per comparison.
- The deep top-100 searches are different queries (questions reformulated by AI assistants), not the 800 keywords.
- The assistants answered 80 questions each, so their rates rest on small counts. This study counts Reddit threads; our [citation study](https://underneath.agency/research/ai-overview-citations-study) counted any reddit.com link and reported Reddit in 18.1% of AI Overviews with sources.

## Data and downloads

- Every Reddit thread appearance (cited or shown), same-community controls, thread features, community type, information need and the coded role and support: [s21_reddit_threads_v11.csv](https://underneath.agency/research-data/ai-reddit-citations-study/s21_reddit_threads_v11.csv) and [JSON](https://underneath.agency/research-data/ai-reddit-citations-study/s21_reddit_threads_v11.json)
- The version 1.0 citation list: [s21_reddit_citations.csv](https://underneath.agency/research-data/ai-reddit-citations-study/s21_reddit_citations.csv)
- Every statistic on this page: [stats.json](https://underneath.agency/research-data/ai-reddit-citations-study/stats.json)
- Machine-readable methodology: [methodology.json](https://underneath.agency/research-data/ai-reddit-citations-study/methodology.json)

The data is free to reuse with attribution (CC BY 4.0).

To cite: Underneath. (2026). *When does a Reddit thread become evidence in Google’s AI?* (Version 1.1). Underneath Research. https://underneath.agency/research/ai-reddit-citations-study

## Frequently asked questions

### How often do Google AI Overviews cite Reddit?

In our September 2026 sample of 800 US searches, 17.9% of AI Overviews cited at least one Reddit thread, and more than a third in home services and B2B software. When Google’s own results showed a Reddit thread, the AI Overview cited Reddit in 33.5% of cases.

### Does ChatGPT cite Reddit?

For the 80 buyer questions we tested (“best X” and “best X in a city”), ChatGPT cited no Reddit threads. For “Is this brand legit?” questions it cited Reddit in 11.4% of answers.

### Are the Reddit threads AI cites the ones that rank on Google?

Sometimes. A Reddit thread in Google’s top three is often cited, one ranked 11 to 100 almost never. 53.9% of the Reddit threads cited in AI Overviews were not on the results page, and on searches where we saw the top 100, 17 of 27 cited threads were outside it.

### What makes a Reddit thread more likely to be cited?

Among threads Google already shows for the same search, cited threads had more discussion (a median of 40 comments against 20). Question titles, matching the search’s words and the kind of community did not reliably predict citation once the search was held fixed. This is an association, not a tested cause.

### Do Reddit citations in AI answers support what the AI says?

Not always. Judged against the post and top 10 comments, 36.4% of sentences citing Reddit were supported, 42.8% partly and 20.8% not supported. Recommendations of named products were the least often supported.

### Which subreddits do AI answers cite most?

In our data, r/askaplumber, r/projectmanagers and r/crmsoftware, followed by r/cybersecurity, r/projectmanagement and r/homeowners. Half of Google’s Reddit citations came from communities of practitioners answering questions about their own work.

## Related research

- [Do AI Overviews cite the pages that rank?](https://underneath.agency/research/ai-overview-citations-study)
- [AI citations vs Google rankings](https://underneath.agency/research/ai-citations-google-rankings-study)
- [The YouTube videos Google’s AI cites are small](https://underneath.agency/research/ai-overview-youtube-videos-study)
- [“Is this brand legit?” How AI assistants build a reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study)
- [How often do AI Overviews appear?](https://underneath.agency/research/ai-overviews-frequency-study)

## Related guides

- [Does Reddit and forum content actually shape Google AI Overview answers?](https://underneath.agency/resources/does-reddit-shape-google-ai-overviews)
- [How do project management tools win customers from AI answers?](https://underneath.agency/resources/project-management-software-customers-ai-search)
- [Does the way customers phrase a question change which sources AI search cites?](https://underneath.agency/resources/does-question-phrasing-change-ai-sources)
- [Which types of websites do AI search engines rely on most?](https://underneath.agency/resources/which-websites-do-ai-search-engines-cite)
- [Why do AI assistants keep naming the same CRMs, and how do we join them?](https://underneath.agency/resources/crm-software-ai-search)

---

This is the Markdown twin of https://underneath.agency/research/ai-reddit-citations-study. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How fresh are the pages AI engines cite? | Underneath"
description: "Held to the same questions, AI assistants cited pages first published half as long ago as Google’s top 10. AI Overviews showed no pull toward recent pages."
canonical: "https://underneath.agency/research/ai-source-freshness-study"
published: 2026-09-26
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Research · AI search

# How fresh are the pages AI engines cite?

Does AI prefer new content? We measured the age of the pages six AI engines cite and compared them with the pages Google ranks. Version 1.0 found that the assistants (ChatGPT, Gemini, Perplexity and Claude) cited clearly fresher pages than Google ranks, while Google’s AI Overviews and AI Mode did not. Version 1.1 asks why. It compares each assistant with Google’s results for the same question, looks at which pages win within the same results, and separates a page that is new from a page that only carries a new date. The main finding holds in a sharper form: the assistants’ edge is mostly pages first published recently, not old pages with recent dates.

## The short version

1. Against Google’s top 10 for the same 80 questions, the four assistants together cited pages first published about half as long ago (ratio 0.50, 95% interval 0.38 to 0.65). Each engine on its own shows the same direction, with intervals below 1.
2. Pages published in the last 90 days made up 17.4% to 22.6% of each assistant’s dated citations, against 6.9% of Google’s top 10. Within the same question, the assistants cited 11.9 points more new pages (interval 6.4 to 16.9).
3. On the newest date a page declares, the version 1.0 measure, the gap is smaller once the question is held fixed. Google’s pages for the assistants’ questions had a median age of 92 days, not 111. Only ChatGPT stays clearly fresher on its own (0.55 times the age, interval 0.36 to 0.80).
4. Within Google’s top 10 for a question, the pages an assistant cited were more often dated within 90 days: +6.2 points (interval 1.0 to 11.6) on a base rate of 15.5%.
5. AI Overviews showed no such pull. On the same results page, a dated page was more likely to be cited than an undated one, but among dated pages a recent date made no difference (−4.1 points, interval −9.6 to 1.3).
6. Most recent dates on the web are old pages with a new modified date: 66.4% of the recently dated pages in Google’s top 10, against 47.2% for ChatGPT.
7. None of this is a causal effect. We changed no pages. The test that could show whether updating a page changes its citations is set out below. We have not run it yet.

## Why this version changes the question

Version 1.0 described what the engines cite. A difference in age between cited pages and Google’s pages has at least three explanations, and they call for different responses.

- **Selection preference.** The engine favors recent pages over older ones that answer the same question. If so, fresher pages should win *within* the same question or results page.
- **Query-driven recency.** The questions asked call for recent information. If so, the gap should shrink when cited pages are compared with Google’s pages for the same question, and it should be larger for time-sensitive queries.
- **Source composition.** The engines cite kinds of pages (articles, dated editorial content) that are recent anyway. If so, the gap should shrink when each engine’s pages are reweighted to Google’s mix.

A fourth question cuts across all three: is a “fresh” page new, or an old page with a new date? A declared date can change without the content changing, so we read each page’s publish date and modified date separately.

Freshness can act at more than one stage. A page has to be found (retrieval), then chosen as a source (citation), then used (absorption) and represented correctly (fidelity). This study sees only the citation stage, with Google’s top 10 as a stand-in for what could have been found. It says nothing about how cited pages were used.

| Analysis | Unit | Compared with | Tests |
|---|---|---|---|
| Same question | Page cited for a question | Google top 10, same question | B |
| Within results | Page in a top 10 | Other pages, same results | A |
| Reweighting | Dated cited page | Google’s page mix | C |
| Date anatomy | Dated page | Publish vs modified date | New or re-dated |

## What we measured

The cited pages are the version 1.0 sample: up to 400 pages per engine drawn at random from the citations collected on 26 September 2026, AI Overviews and AI Mode for 800 US searches across eight industries, and ChatGPT, Gemini, Perplexity and Claude for 80 buyer questions. Platform pages (YouTube, Reddit, social networks, Google, Amazon, Yelp) and PDFs are excluded. New in version 1.1, each page is linked back to the queries that cited it.

Two Google baselines. The first is the version 1.0 baseline, the 3,096 top-10 pages for the searches in our [study of pages cited by AI Overviews](https://underneath.agency/research/ai-overview-cited-pages-study). The second is new: Google’s top 10 for the assistants’ own 80 questions, from our [study of AI citations and Google rankings](https://underneath.agency/research/ai-citations-google-rankings-study) (26 September). We fetched those 547 non-platform pages by plain request on 28 September.

Age is the time from a page’s newest machine-readable date (structured data, article meta tags, time elements) to 26 September, as in version 1.0. Publish age uses the earliest declared publish date. A new page is one published in the last 90 days. Intervals come from resampling whole queries, not single pages, since pages cited for the same question are not independent.

## Finding 1: held to the same questions, the gap narrows on dates and holds on new pages

Google’s top 10 for the assistants’ 80 questions is younger than the version 1.0 baseline: a median of 92 days (interval 67 to 121), against 111 days for the 485 searches of the baseline. Part of the gap in version 1.0 therefore came from the questions, not from the engines.

The table compares each assistant with Google’s pages for the same question. An age ratio below 1 means the assistant’s cited pages are younger; the difference in pages dated within 90 days is in percentage points.

| Engine | Age ratio | 95% interval | Within 90 days |
|---|---|---|---|
| ChatGPT | 0.55 | 0.36 to 0.80 | +13.5 |
| Claude | 0.76 | 0.52 to 1.09 | +7.0 |
| Perplexity | 0.78 | 0.53 to 1.15 | +8.0 |
| Gemini | 0.78 | 0.50 to 1.15 | +6.8 |
| All four | 0.70 | 0.51 to 0.94 | +8.8 |

On the newest declared date, only ChatGPT is clearly fresher on its own; for Claude, Perplexity and Gemini the direction is the same but the intervals include no difference. The within-90-day intervals follow the same pattern (ChatGPT 4.2 to 23.5 points; all four 0.7 to 17.2).

The publish date tells a clearer story. Measured by when a page was first published, every assistant cites younger pages than Google ranks for the same question.

| Engine | Publish-age ratio | 95% interval | New pages |
|---|---|---|---|
| ChatGPT | 0.39 | 0.24 to 0.61 | +16.1 |
| Claude | 0.55 | 0.38 to 0.77 | +9.5 |
| Perplexity | 0.66 | 0.48 to 0.94 | +10.9 |
| Gemini | 0.49 | 0.35 to 0.69 | +12.1 |
| All four | 0.50 | 0.38 to 0.65 | +11.9 |

New pages is the difference, in points, in the share of pages published in the last 90 days. The intervals are 7.0 to 25.8 (ChatGPT), 2.6 to 17.3 (Claude), 4.2 to 16.9 (Perplexity), 3.8 to 20.4 (Gemini) and 6.4 to 16.9 (all four).

Why the two measures differ: Google’s top pages are often old pages with a recent modified date, which makes them look fresh on the newest date. The assistants cite more pages that are actually new.

## Finding 2: new pages or new dates

We sorted every dated page by what its dates say.

| Engine | New pages | 95% interval | Median publish age (days) |
|---|---|---|---|
| Gemini | 22.6% | 14.8 to 31.5 | 284 |
| ChatGPT | 22.2% | 14.8 to 30.3 | 294 |
| Perplexity | 19.2% | 13.7 to 24.0 | 537 |
| Claude | 17.4% | 11.8 to 23.8 | 311 |
| AI Overviews | 10.8% | 6.0 to 16.1 | 575 |
| Google top 10, same questions | 9.0% | 5.7 to 12.5 | 733 |
| AI Mode | 7.8% | 4.0 to 12.6 | 736.5 |
| Google top 10 (baseline) | 6.9% | 5.5 to 8.3 | 801 |

Among pages dated within the last 90 days, the share that are older pages with a recent modified date was 66.2% for AI Overviews, 68.4% for AI Mode and 66.4% for Google’s top 10, against 47.2% for ChatGPT, 54.9% for Perplexity, 56.1% for Gemini and 59.3% for Claude. A further 2.3% to 11.0% of recently dated pages carried a date within a day of our fetch, which usually means the site stamps the current date on every page.

So a recent date on the web usually marks a revision, not a new page, and Google’s results reflect that. The assistants lean further toward pages that are new.

## Finding 3: within the same results, who picks the fresher page

For each of the 80 questions we took the dated pages in Google’s top 10 and asked whether the pages an assistant cited were more often recent than the ones it passed over, with the organic position held fixed. Across 320 engine and question pairs, 15.5% of these pages were cited. A page dated within 90 days was 6.2 points more likely to be cited (interval 1.0 to 11.6). This is the pattern a selection preference would produce, though a recent page may also differ in ways we did not measure.

AI Overviews behave differently. On the same Google results page and at the same position, pages with any machine-readable date were more likely to be cited than pages without one: +5.7 points for a date within 90 days, +8.3 for 91 to 365 days and +12.1 for over a year (intervals 1.0 to 10.3, 2.6 to 13.6 and 7.1 to 17.2). Among the 1,480 dated pages, a recent date made no measurable difference (−4.1 points, interval −9.6 to 1.3). What goes with citation in AI Overviews is declaring a date, not having a recent one; this matches [the date finding in our cited-pages study](https://underneath.agency/research/ai-overview-cited-pages-study).

## Finding 4: the questions matter, but we could test that only on Google

If freshness depends on the question, it should matter most for time-sensitive queries (a year, “latest”, prices, rates) and least for evergreen ones. We classed every query by its wording. Of the 800 Google searches, 105 were time-sensitive, 16 were recommendations and 679 were other searches (informational, navigational or local). Of the 80 assistant questions, 71 were recommendations (“What is the best…”) and only 9 were time-sensitive, too few to test.

On Google’s side the test shows no clear effect. In AI Overviews, a recent date made no measurable difference on time-sensitive searches (+2.0 points, interval −11.0 to 13.0, 86 searches) or on other searches (−5.3, interval −11.7 to 0.9). By Google’s own intent label, recent pages were less likely to be cited on commercial searches (−9.3, interval −16.0 to −2.4) and not clearly more likely on informational ones (+6.3, interval −2.9 to 16.0). The query effect we can see is the one in finding 1: the assistants’ questions have younger Google results to begin with.

## Finding 5: the kind of page does not explain the gap

The assistants cite pages with article markup somewhat more often than Google ranks them (63.6% for ChatGPT to 90.4% for Gemini, against 59.0%), and articles are more often recent. Reweighting each engine’s pages to Google’s mix of page kinds (article markup or not, by reference, review, retail or other site) barely moves the results.

| Engine | Median age | Reweighted | Within 90 days, reweighted |
|---|---|---|---|
| ChatGPT | 50.5 | 43 | 66.5% |
| Claude | 54 | 54 | 60.4% |
| Perplexity | 60 | 60 | 57.8% |
| Gemini | 65 | 65 | 56.4% |
| AI Mode | 111 | 103 | 46.1% |
| AI Overviews | 113 | 117 | 42.9% |

Ages in days. At this level of detail, source composition does not account for the difference. A finer classification (topic, publisher, commercial versus editorial intent) could still find one.

## The three explanations, weighed

- **Selection preference: supported for the assistants, not for Google’s AI.** Within Google’s top 10 for a question, the assistants picked recent pages more often, and they cited younger pages than Google ranks for the same question. AI Overviews showed no preference for recent pages among dated ones.
- **Query-driven recency: part of the version 1.0 gap.** Google’s own results for the assistants’ questions are younger than for the searches in the version 1.0 baseline, which narrows the gap on the newest date. Whether freshness matters more for time-sensitive questions is not settled: the assistants’ questions barely vary, and on Google the intervals are wide.
- **Source composition: not supported at the level measured.** Reweighting to Google’s page mix leaves the ages almost unchanged.

## Version 1.0 results, kept for comparison

The version 1.0 figures are unchanged. The intervals now come from resampling whole queries.

| Engine | Dated | Median (days) | 95% interval | Within 90 days |
|---|---|---|---|---|
| ChatGPT | 162 | 50.5 | 36 to 69.5 | 65.4% |
| Claude | 149 | 54 | 40 to 75.5 | 61.1% |
| Perplexity | 219 | 60 | 42 to 99 | 55.7% |
| Gemini | 115 | 65 | 45 to 102 | 57.4% |
| AI Mode | 167 | 111 | 74.5 to 158 | 45.5% |
| Google top 10 | 1,480 | 111 | 93 to 127 | 46.2% |
| AI Overviews | 166 | 113 | 75 to 157 | 44.6% |

Between 4.0% (Claude) and 7.8% (Perplexity) of the assistants’ dated citations were more than two years old, against 13.9% of Google’s top 10. The share of fetched pages with a machine-readable date ranged from 47.8% (Google’s top 10) to 81.0% (Gemini).

## What this means

These are inferences from observational data, not measured effects.

- **For the assistants, new pages have an edge that re-dated pages do not show as clearly.** The assistants’ advantage is largest on the publish date. A new, substantive page on a topic is the thing the data points to; changing the date on an old one is not.
- **Google’s AI follows Google’s results.** AI Overviews and AI Mode cite pages as old as Google’s top 10, and in AI Overviews a recent date does not help a page that already ranks. Declaring a date, accurately, does go with citation there.
- **Compare like with like.** Much of the freshness advice in circulation compares AI citations with the web at large. Held to the same questions, the gap is smaller and depends on the engine.

## What a causal test would need

Only an experiment can say whether updating a page causes more citations. The design we plan is below; it needs pages we control and budget for repeated runs, and we have not run it.

- Four versions of comparable pages, assigned at random: unchanged, date changed only, a minor edit, and a substantive update (new facts, figures and comparisons).
- Queries in two groups, time-sensitive and evergreen, so the effect can differ by what is asked.
- Several engines, each query run several times per date and on several dates, since answers vary from run to run.
- Outcomes beyond citation: whether the page is cited at all, how prominently, whether its content is used in the answer, and whether the answer states what the page says.
- A condition where competing pages are updated too, to tell an absolute gain from one that only holds while competitors stay still.

The hypotheses, stated before any data: a substantive update raises the chance of citation more than a date change alone; a date change alone has little or no effect; the effect is larger for time-sensitive queries; and it differs by engine.

## Methodology

- **Cited pages:** the version 1.0 sample, up to 400 per engine drawn at random (seed 20260926) from citations collected on 26 September 2026 for our AI Overview frequency, AI Overview citation, AI Mode and four-assistant studies; platform pages and PDFs excluded; linked back to every query that cited them.
- **Baselines:** the 3,096 top-10 pages of our study of pages cited by AI Overviews; and the Google top 10 for the assistants’ 80 questions (26 September), 547 non-platform pages fetched by plain request on 28 September 2026.
- **Dates:** newest valid date among JSON-LD dateModified, datePublished and uploadDate, article meta tags, og:updated_time and time elements; publish date is the earliest declared datePublished, uploadDate or article:published_time; ages counted to 26 September 2026.
- **Models:** linear probability models with question or search fixed effects and dummies for organic positions; comparisons of cited pages with Google pages for the same question use question fixed effects and log age.
- **Query classes:** rules on the query text (time-sensitive, recommendation, other), and DataForSEO’s intent label for Google searches.
- **Intervals:** bootstrap resampling whole queries, 2,000 resamples for descriptive figures and 1,000 for models (seed 20260928).
- **Code:** s22_v11.py (collect, analyze, package), and s22_fresh.py for version 1.0.
- **Update schedule:** quarterly, with a second collection date to measure stability.

## Limitations

- Observational only: nothing was changed on any page, so no difference here is a causal effect of freshness.
- Retrieval is not observed. Google’s top 10 is a stand-in for what could have been found; the assistants run their own searches.
- Declared dates are what a page says, not what changed on it. We did not compare page versions over time.
- One collection date; run-to-run and month-to-month stability are not measured.
- 71 of the 80 assistant questions are recommendations, so query temporality could not be tested for the assistants.
- The same-question Google pages were fetched two days after the cited pages. Any update in between makes them look slightly fresher, which, if anything, narrows the gap.
- 23.0% to 32.5% of sampled cited pages per engine could not be fetched by plain request, and only pages with a machine-readable date are aged.
- Query classes come from simple word rules and were not checked by hand.

## Data and downloads

- Every cited and same-question page with its query, dates, date anatomy and query class: [s22_v11_pages.csv](https://underneath.agency/research-data/ai-source-freshness-study/s22_v11_pages.csv)
- The version 1.0 sample with ages: [s22_cited_page_dates.csv](https://underneath.agency/research-data/ai-source-freshness-study/s22_cited_page_dates.csv) and [JSON](https://underneath.agency/research-data/ai-source-freshness-study/s22_cited_page_dates.json)
- Every statistic on this page, version 1.0 and 1.1: [stats.json](https://underneath.agency/research-data/ai-source-freshness-study/stats.json)
- Machine-readable methodology: [methodology.json](https://underneath.agency/research-data/ai-source-freshness-study/methodology.json)

The data is free to reuse with attribution (CC BY 4.0).

To cite: Underneath. (2026). *How fresh are the pages AI engines cite?* (Version 1.1). Underneath Research. https://underneath.agency/research/ai-source-freshness-study

## Frequently asked questions

### Does ChatGPT prefer recent content?

The pages ChatGPT cited were younger than the pages Google ranks for the same questions: 0.55 times the age on the newest date, and 0.39 times on the publish date. That is a pattern in what it cites, not proof that recency causes citation.

### Do Google AI Overviews prefer fresh content?

Not in our data. Cited pages were as old as Google’s top 10, and among dated pages on the same results page a recent date made no measurable difference. Pages that declared a date at all were cited more often.

### Which AI engine cites the freshest sources?

ChatGPT, by the newest declared date (median 50.5 days). By publish date, ChatGPT and Gemini cite the most new pages (22.2% and 22.6% published in the last 90 days).

### Does updating the date on a page help it get cited by AI?

Our data does not support it. Most recently dated pages on the web are old pages with a new modified date, and the assistants’ edge is in pages that are actually new. Whether a substantive update causes more citations is the experiment described above, which we have not yet run.

## Related research

- [What pages cited by AI Overviews have in common](https://underneath.agency/research/ai-overview-cited-pages-study)
- [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- [Do ChatGPT, Gemini, Perplexity and Claude cite pages that rank?](https://underneath.agency/research/ai-citations-google-rankings-study)
- [How often do AI Overviews appear?](https://underneath.agency/research/ai-overviews-frequency-study)

## Related guides

- [What on-page signals are linked to citations in Google AI Overviews and Perplexity?](https://underneath.agency/resources/on-page-signals-linked-to-ai-citations)
- [Why doesn’t ChatGPT mention our newly launched product?](https://underneath.agency/resources/why-chatgpt-misses-new-products)
- [Do AI search engines cite the same websites as Google?](https://underneath.agency/resources/do-ai-search-engines-cite-the-same-sites-as-google)
- [Will optimizing content for ChatGPT hurt our Google rankings?](https://underneath.agency/resources/chatgpt-optimization-google-rankings)
- [What decides whether an AI engine cites my page over a competitor’s?](https://underneath.agency/resources/why-ai-cites-competitor-page-first)

---

This is the Markdown twin of https://underneath.agency/research/ai-source-freshness-study. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Do Wikipedia and schema make AI assistants recommend a brand?"
description: "Brands on Wikipedia are named by more AI assistants, but within the same question most of the gap goes once brand prominence is taken into account."
canonical: "https://underneath.agency/research/brand-entity-ai-recommendations-study"
published: 2026-09-26
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Research · AI assistants

# Do Wikipedia and schema make AI assistants recommend a brand?

When ChatGPT, Gemini, Perplexity and Claude answer the same buyer question, a few brands are named by all four and most by only one. Version 1.0 of this study found that the brands the assistants agree on are far more likely to have a Wikipedia article, a Wikidata entry and Organization structured data. That leaves the obvious question open: is this the entity record, or just the fact that well-known companies have both a Wikipedia article and a place in AI answers?

This version tests that. We rebuilt the data at the level of each brand, for each question, in each assistant’s answer, added measures of how prominent each brand is and how widely independent sites cover it, and compared brands within the same question. The pattern version 1.0 described is real, but most of it goes away once prominence is accounted for.

## The short version

1. The raw pattern holds: of the options all four assistants named for a question, 60.0% have an English Wikipedia article, against 20.6% of options only one assistant named.
2. Comparing options within the same national question, an article goes with modestly higher odds of being recommended by a given assistant (odds ratio 1.29, 95% interval 1.03 to 1.61).
3. Most of that goes once prominence is accounted for: the odds ratio falls to 1.11 (0.81 to 1.51) after adding the brand website’s traffic rank and how often Wikipedia mentions the brand, and to 0.99 after adding independent coverage in the pages the assistants cited.
4. Independent coverage was the strongest predictor we measured: each tenfold increase in the number of independent sites naming a brand in the cited pages went with 4.7 times the odds of being recommended.
5. Where an article made any difference, it was in the evidence, not the choice: brands with an article were more often named in the pages an assistant cited (74.3% against 59.5%), and once named there, were recommended at the same rate (50.2% against 48.9%).
6. On local questions (a business in a named place), no entity signal made a difference in any model.
7. Of the 110 options all four assistants named for a question, 21 have no Wikipedia article for themselves or their parent brand; 11 of them are local businesses.

## Where this study fits

Research on AI search has moved through three questions, and this study addresses the fourth.

- **Can content change an answer once it is retrieved?** The original generative engine optimization paper ([Aggarwal et al., 2024](https://arxiv.org/abs/2311.09735)) showed that rewriting a source that is already in an AI system’s context can raise its visibility in the answer by up to 40%.
- **Which sources do AI systems use?** Comparative work ([Chen et al., 2025](https://arxiv.org/abs/2509.08919)) found that AI search engines lean heavily on earned, third-party sources rather than brand-owned pages, differ from Google and from each other, and show a “big brand bias”.
- **What remains unknown?** A 2026 critical survey ([Martinez, 2026](https://arxiv.org/abs/2607.14035)) concludes that the evidence covers content already retrieved, not organic discoverability: what makes a brand enter the evidence in the first place is still open.
- **This study** looks at one layer of that question: whether the way a brand is represented as an entity on the public web (Wikipedia, Wikidata, homepage structured data) goes with being recommended by several assistants, and whether that holds once prominence is accounted for.

We use a working model of how a recommendation forms: the public web describes the brand; an assistant retrieves and cites some pages; it selects options from them; it recommends and ranks them; several assistants may or may not agree; and the result may or may not repeat on the next run. This model is proposed, not tested. We observe the entity signals, the cited pages, the recommendations, ranks and agreement, and stability on a subset; we do not observe crawling, indexing or what users do next.

### Research questions

- How consistently do the four assistants recommend the same brands for the same buyer question?
- Are entity signals associated with being recommended, for each brand, question and assistant?
- Does the association remain after accounting for brand prominence and independent coverage?
- Does it differ by assistant, by question wording, and between national and local questions?
- Are brands with an article named more consistently across repeated runs?
- What do the exceptions look like: brands every assistant named without an article, and brands with an article only one assistant named?

## What we analyzed

The answers come from our [four-assistant study](https://underneath.agency/research/ai-assistants-brand-agreement-study): ChatGPT, Gemini, Perplexity and Claude each answered the same 80 US buyer questions (ten in each of eight industries, 24 naming a place) on 26 September 2026, 320 answers in all. No new AI questions were asked for this version.

- **Unit of analysis.** Every option named in those answers was coded by a model (Claude Opus) as recommended or merely mentioned, with its rank. That gives 1,419 question-option pairs and, for each pair, one row per assistant: 5,676 observations. Version 1.0 reduced each brand to one number, the most assistants that named it for any question.
- **Outcomes, from weakest to strongest.** Mentioned; recommended; ranked in the top three; the number of assistants recommending it for the question; and, for 20 questions asked five times, the share of runs that named it.
- **Entity signals.** An English Wikipedia article; a Wikidata entry (and whether it lists social profiles); Organization structured data and sameAs links on the homepage. Version 2.0 finds articles two ways: through a Wikidata match, as in version 1.0, and by looking up the option’s own name as a Wikipedia title. A separate measure also counts an article for the parent brand (QuickBooks for QuickBooks Online).
- **Prominence.** The brand website’s rank in the Tranco list of the top million sites, and the number of English Wikipedia articles that mention the brand’s name. We also tried to count news coverage; the news database refused our requests, so news is not measured.
- **Independent coverage.** How many different websites, other than the brand’s own, name the brand among the readable pages cited anywhere in the study. This sits inside the assistants’ own evidence, so it is partly the same process as the outcome; we report it as a separate step.
- **Comparisons.** Each model compares options within the same question and assistant (question fixed effects), with uncertainty that allows for the same brand appearing several times. National and local questions are analyzed separately.

Version 1.0’s brand list came from an automatic extractor that missed 453 of the 1,419 options the model coder found (Vanguard and several law firms among them). This version gives those options the same signals with the same rules.

## Finding 1: the pattern version 1.0 described is still there

### Share with an English Wikipedia article, by how many assistants named the option

| Named by | National | Local | All |
|---|---|---|---|
| 1 of 4 | 32.9% | 6.7% | 20.6% |
| 2 of 4 | 47.1% | 10.2% | 32.7% |
| 3 of 4 | 46.2% | 18.8% | 39.0% |
| All 4 | 69.9% | 5.9% | 60.0% |

The national rows cover 511, 138, 91 and 93 options; the local rows 449, 88, 32 and 17.

Chart: [share with an article, by assistants naming it](https://underneath.agency/research-data/brand-entity-ai-recommendations-study/wikipedia-by-assistant-count.svg)

### From mention to agreement, national questions

| Outcome | With article | Without |
|---|---|---|
| Mentioned by a given assistant | 50.3% | 37.9% |
| Recommended by a given assistant | 42.0% | 34.2% |
| Recommended, when mentioned | 83.5% | 90.1% |
| In the top 3, when recommended | 58.1% | 47.6% |
| Assistants recommending (mean of 4) | 1.68 | 1.37 |
| Recommended by 3 or 4 assistants | 29.4% | 15.4% |

Options with an article are mentioned more often but are slightly less often recommended once mentioned: large companies and platforms (Google Workspace, Microsoft Teams) are often named in passing rather than put forward as a pick.

## Finding 2: most of the association is prominence

Brands with an article are, above all, better-known brands: the article measure correlates at 0.654 with the website’s Tranco score. So we added prominence to the comparison step by step.

### Odds of being recommended, national questions (odds ratio, 95% interval)

| Signal | Signal only | + prominence | + independent coverage |
|---|---|---|---|
| Own English Wikipedia article | 1.29 (1.03 to 1.61) | 1.11 (0.81 to 1.51) | 0.99 (0.73 to 1.33) |
| Own or parent brand article | 1.43 (1.17 to 1.74) | 1.29 (1.04 to 1.60) | 1.12 (0.90 to 1.39) |
| Wikidata entry | 1.33 (1.09 to 1.63) | 1.00 (0.72 to 1.39) | 0.79 (0.57 to 1.10) |
| Article, version 1.0 method | 1.27 (1.02 to 1.59) | 1.01 (0.73 to 1.40) | 0.93 (0.68 to 1.26) |

An odds ratio of 1 means no difference. Each model compares options within the same question and includes the assistant.

In percentage points, an article went with a 5.6-point higher chance of being recommended by a given assistant before controls, 2.2 points after prominence, and −0.3 points after independent coverage. A mixed model that treats brands and questions as random rather than fixed gives a small positive association after prominence (1.18, 1.06 to 1.33); either way it is small.

The broadest measure, an article for the brand or its parent, holds up after prominence but not after independent coverage. One reading, which this data cannot prove: brands with an established record are more widely written about by independent sites, and it is that coverage the assistants draw on. Independent coverage itself went with 4.7 times the odds of being recommended per tenfold increase (3.4 to 6.5), the largest association in the study.

Chart: [odds of being recommended as controls are added](https://underneath.agency/research-data/brand-entity-ai-recommendations-study/wikipedia-odds-ratio-with-controls.svg)

## Finding 3: evidence, not selection

Using the pages each assistant cited (fetched two days after the answers), we split the recommendation into two stages: was the option named in the pages the assistant itself cited, and, if so, did the assistant recommend it?

| National questions | With article | Without |
|---|---|---|
| Named in the assistant’s cited pages | 74.3% | 59.5% |
| Recommended, if named there | 50.2% | 48.9% |
| Recommended, if not named there | 16.0% | 14.6% |

After prominence controls, an article went with somewhat higher odds of appearing in the cited pages (1.38, 0.94 to 2.04) and no difference in being selected from them (0.95, 0.63 to 1.44). If an article matters at all, it matters for whether a brand is in the evidence, not for how the assistant chooses from it. That fits the earned-media finding in the literature.

## Finding 4: homepage schema, sameAs and entity tiers

For the national brands whose homepage we could fetch (215 brands), Organization schema kept a clearer association than Wikipedia did: 1.54 (1.07 to 2.22) after prominence and 1.39 (0.99 to 1.97) after independent coverage. sameAs links: 1.36 (0.94 to 1.95) after prominence. These rest on a smaller, selected group, and [Ahrefs’ controlled test](https://ahrefs.com/blog/schema-ai-citations/) found that adding schema did not raise AI citations, so we read this as a marker of organizational maturity rather than an effect of the markup.

We also grouped homepage-checked national brands into tiers of entity representation.

| Tier | Brands | Recommended by 3 or 4 | Assistants recommending |
|---|---|---|---|
| Weak: no Wikidata, no article, no schema | 5 | 14.3% | 1.43 |
| Partial | 122 | 24.7% | 1.58 |
| Strong: Wikidata, schema and sameAs | 46 | 20.0% | 1.64 |
| Strong and widely covered | 42 | 39.7% | 1.91 |

“Widely covered” means more independent domains in the cited pages than the median (9). Strong markup alone did no better than partial; the jump comes with independent coverage. The trend across tiers was not significant after prominence (1.22 per tier, 0.96 to 1.54), and the weak tier is only five brands.

## Finding 5: no clear differences by assistant or wording, and none locally

- **By assistant** (national, after prominence): ChatGPT 1.37, Gemini 1.19, Claude 1.42, Perplexity 0.68; none is individually significant, and the four do not differ significantly from each other (p = 0.12).
- **By question wording**: plain “best X” questions 0.88, questions with a constraint (“for a small business”, “under $100,000”) 1.00, and questions about an attribute (“most reliable”, “cheapest”) 2.51 (1.24 to 5.10). The attribute group is only 8 national questions and the overall test for a difference is not significant (p = 0.28), so treat this as a lead to test, not a finding. Every question in the study asks for a recommendation; informational and transactional questions are not covered.
- **Local questions**: an article went with no difference before or after controls (1.07, 0.74 to 1.53; 1.02, 0.65 to 1.60), and the same held for Wikidata, schema and sameAs. Only 46 of the 586 local question-option pairs have an article, and the national and local associations do not differ significantly (p = 0.36), so the data cannot say whether entity signals work differently locally, only that no effect is visible there.

## Finding 6: repeated runs

For 20 of the questions, ChatGPT, Gemini and Perplexity answered five times (from our [consistency study](https://underneath.agency/research/ai-recommendation-consistency-study)). On national questions, brands with an article were named in four or five of the five runs 61.1% of the time, against 43.6% without. After prominence, the difference in the share of runs was 2.7 points (−7.8 to 13.1): brands with articles are more stable mainly because they are more prominent.

## Finding 7: matched pairs

Within each national question, we paired every option that has an article with the most similar option that does not, on traffic rank, Wikipedia mentions and independent coverage. Only 42 of the 340 options with an article had a close enough partner: most brands with an article are simply far more prominent than any brand without one in the same answer. In those 42 pairs, the option with an article was recommended by 0.17 fewer assistants on average (−0.64 to 0.33); it was ahead in 26.2% of pairs, tied in 40.5% and behind in 33.3%. The pairs are few, but they give no sign that an article helps when prominence is equal.

## The exceptions

### Consensus without an article

110 options were named by all four assistants for a question. The automatic checks found no article for 44 of them. We looked each one up.

| On review | Options | Examples |
|---|---|---|
| Own article under another title | 8 | Wix (Wix.com), Jira, Betterment, Chubb |
| Product or program of a parent with an article | 15 | QuickBooks Online, Hilton Honors, Invisalign |
| No article found | 21 | OnPay, GlassesUSA, Navien, local firms |

Version 1.0 reported that 38 of the 96 brands named by all four assistants had no Wikipedia article; much of that was matching error and products of large parents. Real consensus without any encyclopedic record exists, but it is mostly local: 14 of the 21 come from local questions, where the assistants lean on directories, reviews and the business’s own site (the own site was cited for 38.1% of them).

### An article, but only one assistant

198 options with an article were named by only one assistant. These are prominent (83.8% have a website in the Tranco top million; a median of 303 Wikipedia articles mention them), and the one assistant recommended them 63.1% of the time. ChatGPT (69) and Gemini (65) account for most, against Claude (38) and Perplexity (26). Many are adjacent or parent brands named alongside the real answer (Google Workspace, Slack or Zoom for a project management question), or large banks and chains one assistant added to a longer list. An article makes a brand available to be named; it does not make it the answer.

### Identity errors along the way

- 145 options (10.2%) were named differently by different assistants, and 447 were recorded with a parent brand, so a brand’s identity is often split across names.
- Where Wikidata and the AI answers both give a website (51 brands), they agree 84.3% of the time; the rest include a Turkish American Express domain on Wikidata and Fidelity’s UK site for the US company.
- Wikidata search matched the wrong entity in 13 of 119 new matches (a product model, a same-name hotel elsewhere, a drink), and we rejected them on review.

## What this means

This is interpretation. The data shows associations, not causes.

- **A Wikipedia article is mostly a sign of a brand the web already knows well.** Within the same question, once website traffic rank and encyclopedic mentions are accounted for, an article adds little or nothing to the chance of being recommended.
- **Independent coverage is where the difference is.** The brands the assistants agree on are the ones many independent sites write about, and the difference sits in whether a brand appears in the pages an assistant cites. That matches the earned-media bias other researchers found.
- **Keep the entity record accurate anyway.** A correct Wikidata entry, Organization schema with sameAs links to real profiles, and consistent names and websites cost little and prevent the identity errors above. Wikipedia has strict notability and conflict-of-interest rules, so an article is not something a company can simply write for itself.
- **For local businesses, look elsewhere.** No entity signal made a difference locally. Our [ChatGPT local study](https://underneath.agency/research/chatgpt-local-recommendations-study) points to Google Business Profile data and local review sites instead.

## What comes next

- Repeated runs for all four assistants over several weeks, so that agreement and stability can be separated properly.
- News coverage and search demand as prominence measures, and coverage measured outside the assistants’ own citations.
- Controlled tests that change a brand’s entity record (for example, a corrected Wikidata entry) and track recommendations before and after.

## How this compares with other studies

A claim that “entities with Wikipedia pages are 50% more likely to appear in AI-generated top-ten lists” is widely repeated. The source we could trace, an [ALLMO article](https://allmo.ai/articles/what-we-know-about-the-impact-of-wikipedia-on-chatgpt-search-results) citing Semrush research from 2025, says something narrower: that half of the marketing agencies most often cited in AI answers had Wikipedia pages, from 58 questions to four assistants. Neither version accounts for prominence; ours suggests that once it is, most of the Wikipedia association goes. On structured data, Ahrefs’ controlled test of 1,885 pages found adding schema “produced no major uplift in citations on any platform”. [Chen et al.](https://arxiv.org/abs/2509.08919) found AI search leans on earned media and favors big brands, which fits our finding that independent coverage, not the entity record, carries the association. The [critical survey by Martinez](https://arxiv.org/abs/2607.14035) calls for repeated observations, several engines and hierarchical models; this version moves toward that, with the gaps listed below.

## Methodology

- **Answers:** 320 answers (ChatGPT and Gemini consumer apps via DataForSEO, Perplexity sonar and Claude Haiku 4.5 through their APIs with web search; 80 US buyer questions; 26 September 2026) from our four-assistant study. Options coded by Claude Opus, all four answers to a question at once with the assistants hidden: recommended or mentioned, rank, type, parent brand.
- **Panel:** 1,419 question-option pairs × 4 assistants = 5,676 rows; 1,282 distinct brands. 453 pairs added in version 2.0 with the version 1.0 rules.
- **Wikipedia article (primary):** Wikidata match with an English sitelink, or the option’s own name resolves to a Wikipedia page (redirects followed) whose short description fits the industry; a redirect to a different title counts only as a parent article. All 40 title additions and all 119 new Wikidata matches were reviewed by Claude, not by a person; 13 matches and 6 titles were rejected.
- **Homepage signals:** Organization-type JSON-LD and sameAs, fetched 26 to 28 September 2026 with the version 1.0 fetcher.
- **Prominence:** Tranco list 64X3X (downloaded 28 September 2026), scored as 6 minus log10 of the rank, with a flag for sites outside the top million; English Wikipedia full-text mentions of the brand name.
- **Independent coverage:** distinct domains, excluding the brand’s own, among the 1,419 readable pages (of 2,209 cited) that name the brand.
- **Models:** logistic regression with question fixed effects and assistant, standard errors clustered by brand; added in steps (signal only, plus prominence, plus independent coverage); Bayesian mixed model with random intercepts for brand and question as a check; likelihood-ratio tests for interactions. Matched pairs: standardized distance of 0.5 or less, without replacement; intervals from 2,000 bootstrap resamples (seed 20260928).
- **Question wording:** a rule on the wording: attribute (“most reliable”, “cheapest”, “best value”), constrained (“for a small business”, “under $100,000”) or plain.
- **Update schedule:** quarterly, with the next four-assistant run.

## Limitations

- Observational. Prominence, independent coverage and entity signals move together, so their separate contributions come with wide intervals, and few brands with an article have a comparable brand without one.
- Only options at least one assistant named are in the data, so the study compares named brands with each other; brands no assistant named are not observed.
- The prominence measures are partial: website traffic rank and Wikipedia mentions, without news, search demand or revenue.
- Independent coverage is counted in the pages the assistants cited, two days later, so it is partly the same process as the outcome.
- One run per assistant on one date (five runs for 20 questions and three assistants), US English, and recommendation questions only.
- Entity matching still misses some articles under ambiguous names; all reviews were done by a model, not a person.

## What changed in version 2.0

- **Unit of analysis:** from each brand’s maximum number of assistants to every option, question and assistant (5,676 rows), with mention, recommendation, rank, agreement and stability as separate outcomes.
- **Coverage:** 453 options version 1.0’s extractor missed were added, and a second article check by Wikipedia title found articles the Wikidata search missed.
- **Controls:** prominence and independent coverage added; comparisons made within the same question.
- **Headline:** version 1.0 reported that brands with an article were named by three or four assistants 47.5% of the time, against 24.5% without, on national questions. That descriptive gap is real, but most of it reflects prominence. “38 of 96 brands named by all four have no article” is replaced by 21 of 110 after review. Version 1.0 figures are kept in stats.json.

## Data and downloads

- Every option, question and assistant with outcomes and signals: [s12_panel_v2.csv](https://underneath.agency/research-data/brand-entity-ai-recommendations-study/s12_panel_v2.csv) and [JSON](https://underneath.agency/research-data/brand-entity-ai-recommendations-study/s12_panel_v2.json)
- Every brand with its signals, prominence and coverage: [s12_brands_v2.csv](https://underneath.agency/research-data/brand-entity-ai-recommendations-study/s12_brands_v2.csv)
- Matched pairs: [s12_matched_pairs.csv](https://underneath.agency/research-data/brand-entity-ai-recommendations-study/s12_matched_pairs.csv)
- Version 1.0 brand table: [s12_brand_entities.csv](https://underneath.agency/research-data/brand-entity-ai-recommendations-study/s12_brand_entities.csv)
- Every statistic on this page, and the version 1.0 statistics: [stats.json](https://underneath.agency/research-data/brand-entity-ai-recommendations-study/stats.json)
- Machine-readable methodology: [methodology.json](https://underneath.agency/research-data/brand-entity-ai-recommendations-study/methodology.json)

The data is free to reuse with attribution (CC BY 4.0).

To cite: Underneath. (2026). *Do Wikipedia and schema make AI assistants recommend a brand?* (Version 2.0). Underneath Research. https://underneath.agency/research/brand-entity-ai-recommendations-study

## Frequently asked questions

### Does having a Wikipedia page help a brand get recommended by ChatGPT?

Brands with a Wikipedia article are named by more AI assistants, but in our data that is mostly because they are better-known brands. Comparing brands within the same question, the association fell from an odds ratio of 1.29 to 1.11, with an interval that includes no effect, once website traffic rank and Wikipedia mentions were accounted for.

### What predicts whether several AI assistants recommend the same brand?

In our data, the strongest predictor was independent coverage: how many different websites name the brand in the pages the assistants cite. Brands with an article were more often in those pages, but once there, they were recommended at the same rate as brands without one.

### Do local businesses need a Wikipedia page for AI recommendations?

Our data suggests not. Only 46 of the 586 local options have an article, and no entity signal made a difference to how often the assistants recommended local businesses.

### Does Organization schema help AI assistants recommend a brand?

Among national brands whose homepage we could check, Organization schema went with higher odds of being recommended even after prominence (1.54). A controlled test by Ahrefs found adding schema did not raise AI citations, so the difference probably reflects the kind of company that publishes schema.

### Can a brand be recommended by every AI assistant without a Wikipedia page?

Yes. Of the 110 options all four assistants named for a question, 21 had no article for themselves or a parent brand, and most of those were local businesses.

## Related research

- [Do ChatGPT, Gemini, Perplexity and Claude agree on brands?](https://underneath.agency/research/ai-assistants-brand-agreement-study)
- [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)
- [What pages cited by AI Overviews have in common](https://underneath.agency/research/ai-overview-cited-pages-study)
- [Do websites serve Markdown to AI agents?](https://underneath.agency/research/agent-readable-web-study)

## Related guides

- [Why does Wikipedia matter so much for AI search visibility?](https://underneath.agency/resources/why-wikipedia-matters-for-ai-search)
- [How can a small brand get recommended by AI assistants?](https://underneath.agency/resources/how-small-brands-get-recommended-by-ai)
- [Do AI assistants favor big brands over smaller competitors?](https://underneath.agency/resources/do-ai-assistants-favor-big-brands)
- [How do brands build authority that AI search recognizes?](https://underneath.agency/resources/how-brands-build-authority-for-ai-search)
- [Does advertising spend help our brand get recommended by AI, or does something else matter more?](https://underneath.agency/resources/does-ad-spend-help-ai-brand-recommendations)

---

This is the Markdown twin of https://underneath.agency/research/brand-entity-ai-recommendations-study. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Which Google Maps businesses does ChatGPT recommend? | Underneath"
description: "Across 898 ChatGPT answers over two days, more reviews than the local median raised a Maps business’s chance of being listed by about 20 points."
canonical: "https://underneath.agency/research/chatgpt-local-picks-google-profile-study"
published: 2026-09-26
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Research · Local AI search

# Which Google Maps businesses does ChatGPT recommend?

Our [study of ChatGPT’s local recommendations](https://underneath.agency/research/chatgpt-local-recommendations-study) found that the business details ChatGPT shows match Google Maps almost field for field. This study asks the next question: of the businesses Google Maps shows for a search, which ones does ChatGPT pick, and do the same kinds of business get picked when the question is asked again, reworded or asked the next day?

We compared the Google Maps top 20 for 120 local searches in the United States, United Kingdom, Canada and Australia with the businesses ChatGPT listed for the same question: one answer per search on 26 September 2026, and seven more per search on 27 September 2026 (three repeats of the same question and four rewordings). In all, 898 answers showed a business list, giving 16,294 pairs of a Maps business and an answer.

Version 1.1 of this study adds the second day, the repeats and the rewordings, a regression model that adjusts for all signals at once, and measures of how stable ChatGPT’s picks are. One finding from version 1.0 did not hold up: keyword-stuffed business names were not reliably favored once more answers were examined.

## The short version

1. Review volume was the one signal that held everywhere: businesses with more reviews than the median for their search were more likely to be listed at the same Maps rank in all 8 sets of answers (by 15.8 to 22.0 points). In a model that adjusts for rank, every other signal, country, service, wording and day, the difference was 19.5 points (95% interval 15.0 to 23.0).
2. Maps rank matters most. On 26 September ChatGPT listed 67.7% of the businesses ranked 1 to 3 in Maps and 25.4% of those ranked 11 to 20; after adjusting for the other signals, a rank of 11 to 20 instead of 1 to 3 lowered the chance by 34.1 points.
3. A rating of 4.8 or more added 9.9 points in the model; a rating above the local median added 7.1 points, which version 1.0 (one answer per search) could not detect.
4. Keyword-stuffed business names did not replicate: the 10.6-point advantage seen on 26 September was significant in only 2 of the 8 answer sets, and 2.4 points (not significant) in the model. We no longer report it as a finding.
5. ChatGPT’s picks are fairly stable. Two repeat runs of the same question shared 0.74 of their listed Maps businesses on average (Jaccard); a run and a rewording shared 0.71, and the two days shared 0.68. Of the Maps businesses listed at least once in three repeat runs, 60.0% were listed all three times.
6. Stable picks have more reviews. Among businesses listed at least once, 69.2% of those with more reviews than the local median were listed in all three repeats, against 44.1% of those with fewer.

## The questions we tested

We treat each claim as a question the data can answer, rather than a conclusion:

1. Is a business higher in Google Maps more likely to be listed by ChatGPT?
2. At the same Maps rank, are businesses with more reviews, higher ratings, more photos or keyword-stuffed names more likely to be listed?
3. Does each association hold across repeat runs of the same question, four rewordings and a second day?
4. When the question is repeated, how much of the listed set stays the same, and which businesses stay on it?

The data shows which businesses ChatGPT lists, not why: we observe the answer, not the retrieval behind it.

## What we measured

For each of the 120 local questions in our earlier study (“Who is the best {service} in {city}?”, five services, six cities in each of four countries), we took the Google Maps top 20 for “best {service} in {city}”, collected on the same day as the ChatGPT answers. Each Maps business was labeled listed or not listed in each ChatGPT answer that showed a business list.

**Listed** means ChatGPT named the same business at the same location, using the branch-aware matching of our [local recommendations study](https://underneath.agency/research/chatgpt-local-recommendations-study): a listing and a Maps entry that share a website, name or phone number are compared on street address, postcode and phone, and another branch of the same company does not count. Version 1.0 counted every Maps entry that shared a listed business’s website as listed, including other branches of the same chain; with the stricter rule the 26 September listing rate is 40.8% instead of 42.9%.

Reviews, photos and ratings mean different things in different cities, so each business is compared with the others in the same search: above or below that search’s median. We estimate each signal in two ways:

- **Within rank bands** (the version 1.0 method): the difference in listing rate with and without the signal inside Maps rank bands 1 to 3, 4 to 10 and 11 to 20, averaged, separately for each of the 8 answer sets.
- **A pooled model**: a logistic regression on all 16,294 business-answer pairs with every signal, the rank band, country, service, reworded or not, and day, with uncertainty that treats each search as one cluster. We report each signal as an average marginal effect: the change in the predicted chance of being listed when the signal is switched on for every business.

All 95% intervals come from resampling whole searches, so answers and businesses from the same search are never treated as independent.

## Findings

### Maps rank

On 26 September, with one answer per search:

| Google Maps rank | Businesses | Listed by ChatGPT |
|---|---|---|
| 1 to 3 | 282 | 67.7% |
| 4 to 10 | 637 | 50.9% |
| 11 to 20 | 914 | 25.4% |
| All | 1,833 | 40.8% |

The pattern repeated in every answer set on 27 September: between 68.4% and 71.1% of the Maps top 3 were listed, and between 24.5% and 27.5% of ranks 11 to 20.

### Profile signals, all adjusted at once

| Signal | Businesses with it | Change (points) | 95% interval |
|---|---|---|---|
| More reviews than local median | 49.5% | +19.5 | 15.0 to 23.0 |
| Rating 4.8 or higher | 84.0% | +9.9 | 4.7 to 14.7 |
| Rating above local median | 27.1% | +7.1 | 2.5 to 11.3 |
| More photos than local median | 48.7% | +3.2 | 0.5 to 6.3 |
| Keyword-stuffed name | 13.0% | +2.4 | −2.2 to 6.8 |
| Maps rank 11 to 20 (vs 1 to 3) | 49.9% | −34.1 | −40.4 to −28.5 |

“Change” is the model’s estimate of how much the signal changes the chance of being listed, holding the other signals, rank, country, service, wording and day fixed. “Businesses with it” is the share on 26 September.

A claimed profile, a website and a category that matches the service are left out of the model because 98% or more of the businesses have them, which leaves too few without them to compare.

The median listed business on 26 September had 235 reviews against 133.5 for businesses ChatGPT left out; both groups had a median rating of 4.9 and a similar number of photos (25 and 22).

### Does each signal hold across runs, rewordings and days?

The within-band difference, estimated separately in each set of answers:

| Answer set | Answers | Reviews | Rating | Stuffed name |
|---|---|---|---|---|
| 26 Sep, original | 103 | +21.3 | +2.5 | +10.6 |
| 27 Sep, repeat 1 | 112 | +20.4 | +8.5 | +6.0 |
| 27 Sep, repeat 2 | 115 | +20.4 | +4.5 | −0.6 |
| 27 Sep, repeat 3 | 112 | +22.0 | +7.3 | +6.0 |
| 27 Sep, wording A | 113 | +18.5 | +5.6 | +1.7 |
| 27 Sep, wording B | 116 | +15.8 | +7.4 | +4.6 |
| 27 Sep, wording C | 115 | +21.3 | +5.7 | +2.3 |
| 27 Sep, wording D | 112 | +18.6 | +8.2 | +3.6 |
| **Significant in** | | **8 of 8** | **6 of 8** | **2 of 8** |

Differences are in points: more reviews than the local median, a rating above the local median, and a keyword-stuffed name. The wordings were A “Which {service} in {city} would you recommend?”, B “I need a good {service} in {city}. Who should I go to?”, C “List the top {services} in {city}.” and D the original question with “Please cite your sources.”

Review volume is the robust finding. A rating above the local median is a smaller, fairly consistent advantage that a single answer per search was too noisy to detect. The keyword-stuffing advantage came mostly from the first day’s answers and should be treated as chance.

### How stable are ChatGPT’s picks?

We compared the sets of Maps top-20 businesses listed in two answers to the same search with the Jaccard index (shared businesses divided by all businesses listed in either; 1 means identical).

| Comparison | Pairs of answers | Mean overlap | 95% interval |
|---|---|---|---|
| Repeats, 27 Sep | 283 | 0.74 | 0.72 to 0.77 |
| Original vs rewording, 27 Sep | 1,141 | 0.71 | 0.70 to 0.73 |
| 26 Sep vs 27 Sep | 251 | 0.68 | 0.66 to 0.71 |

In the 100 searches where all three repeats showed a list, 939 of 1,829 Maps businesses were listed at least once. Of those, 60.0% were listed in all three repeats, 19.3% in two and 20.8% in one. A top-3 business listed at least once was listed all three times 84.1% of the time; a business ranked 11 to 20, 41.1% of the time.

The Maps top 20 itself barely moved between the two days (mean overlap 0.87). Of the 1,438 businesses in both days’ top 20, 88.0% had the same status on both days (listed on 26 September and in at least two of three repeats on 27 September, or neither).

### How widely ChatGPT spreads its picks

Across the seven answers per search on 27 September, a typical search had 8.7 Maps businesses listed per answer but 12 different ones across all seven. Counting how often each was listed, the picks behaved like 9.9 equally used businesses (median effective number; concentration index 0.101), and the three most-listed businesses took 34.4% of all listing slots. ChatGPT’s picks rotate among a core of businesses rather than settling on a single fixed list.

Google Maps is not the whole picture: 45.1% of the businesses ChatGPT listed on 27 September were not in the Maps top 20 for that search at all (46.8% on 26 September). This study is about the Maps top 20 only.

## How this compares with other studies

| Source | Sample and date | Finding |
|---|---|---|
| This study | 120 searches, 898 ChatGPT answers over two days, 4 countries, September 2026 | More reviews than the local median: +19.5 points in all answers pooled, significant in 8 of 8 answer sets |
| Ibrahim and Zaki (arXiv) | Four service domains in the 100 largest US metros, September 2026 | “a 3-5x review-count premium” but “a rating premium of at most a tenth of a star” |
| SOCi 2026 Local Visibility Index | About 350,000 locations of multi-location brands, January 2026 | ChatGPT recommends 1.2% of locations; “ChatGPT-recommended locations average 4.3 stars” |
| Whitespark 2026 Local Search Ranking Factors | Survey of 47 local search experts, November 2025 | Experts rank presence on curated “best of” lists first for AI visibility, and high Google ratings eighth |

Our results and Ibrahim and Zaki’s point the same way: review volume goes with being recommended by AI. With eight times as many answers, we do see a modest rating advantage within a search, which is consistent with their “at most a tenth of a star” once most businesses are rated 4.8 or more. Whitespark’s ranking is expert opinion rather than measurement, and SOCi’s figures describe multi-location brands, a different population from our independent businesses.

Sources: [Ibrahim and Zaki](https://arxiv.org/abs/2609.18341); [SOCi](https://www.soci.ai/blog/local-visibility-benchmarks/); [Whitespark](https://whitespark.ca/local-search-ranking-factors/); [our ChatGPT local recommendations study](https://underneath.agency/research/chatgpt-local-recommendations-study).

## What this means

This is our interpretation; the data shows associations, not causes.

- **Reviews are the signal a business can most directly influence.** Review volume was the strongest and most consistent profile signal, and a business can grow it honestly by asking every satisfied customer for a review.
- **Ranking in Maps still matters most.** Being in the Maps top 3 made a listing far more likely than any single profile signal, and top-3 businesses were also the most consistently listed.
- **Judge AI visibility over several answers, not one.** About a quarter of the listed set changed between two identical questions, so a single check can mislead in either direction.
- **Don’t stuff keywords into the business name.** The advantage we reported in version 1.0 did not replicate, and [Google’s guidelines](https://support.google.com/business/answer/3038177) do not allow it: a profile that breaks them can be suspended.

## Methodology

- **Searches:** the 120 local questions of our ChatGPT local study (5 services, 24 cities, 4 countries).
- **Answers:** ChatGPT answers collected through DataForSEO’s LLM Scraper: the original question once on 26 September 2026; on 27 September 2026 three repeats and four rewordings (“Which {service} in {city} would you recommend?”, “I need a good {service} in {city}. Who should I go to?”, “List the top {services} in {city}.”, and the original with “Please cite your sources.”). 898 of 960 answers showed a business list.
- **Businesses:** the Google Maps top 20 for “best {service} in {city}” (DataForSEO Google Maps, country-level location), pulled on each collection day: 1,833 businesses on 26 September and 2,173 on 27 September.
- **Listed:** the same business at the same branch appears in the answer’s business list (branch-aware entity resolution from our local recommendations study).
- **Signals:** review count, photo count and rating compared with the median of the same search; rating of 4.8 or more; keyword-stuffed name (a name containing “ - ”, “|” or brackets); claimed profile, website and matching primary category reported but not modeled.
- **Within-band estimate:** listing-rate difference with and without each signal inside Maps rank bands 1 to 3, 4 to 10 and 11 to 20, weighted by band size, per answer set; 95% intervals from 500 resamples of searches.
- **Model:** logistic regression on all business-answer pairs with rank band, the five signals, country, service, reworded question and day; standard errors clustered by search; average marginal effects with 95% intervals from 200 resamples of searches (seed 20260928).
- **Stability:** Jaccard overlap of the listed Maps businesses between answers to the same search; listing frequency across the three repeats; effective number of businesses (inverse of the sum of squared listing shares) across the seven 27 September answers.
- **Code:** cite/pipeline/s14_v11.py (version 1.1) and s14_analyze.py (version 1.0).
- **Update schedule:** quarterly.

## Limitations

- **Associations, not causes.** Businesses with more reviews may differ in other ways, and we cannot see which pages ChatGPT retrieved before answering. Testing whether a signal changes the answer would need a controlled experiment, which this study is not.
- **Two days only.** The repeats show short-term variation, not how picks drift over months.
- **Maps results were collected at country level for a city search**, so the Maps top 20 may differ from what a searcher in that city sees.
- **The Maps top 20 is not the whole field.** 45.1% of the businesses ChatGPT listed on 27 September were outside it and are not analyzed here.
- **Simple signal definitions.** The keyword-stuffing rule is a text pattern and misses stuffed names without separators; signals are split at the local median rather than measured on a scale.
- **Matching was validated by model coders, not people.** The branch-aware matching was checked in our local recommendations study with model-coded, adjudicated samples; no person coded them.

## Data and downloads

- Every business-answer pair (version 1.1) with its signals and listing status: [s14_business_answer_pairs.csv](https://underneath.agency/research-data/chatgpt-local-picks-google-profile-study/s14_business_answer_pairs.csv)
- The version 1.0 businesses: [s14_maps_businesses.csv](https://underneath.agency/research-data/chatgpt-local-picks-google-profile-study/s14_maps_businesses.csv) and [JSON](https://underneath.agency/research-data/chatgpt-local-picks-google-profile-study/s14_maps_businesses.json)
- Every statistic on this page: [stats.json](https://underneath.agency/research-data/chatgpt-local-picks-google-profile-study/stats.json)
- Machine-readable methodology: [methodology.json](https://underneath.agency/research-data/chatgpt-local-picks-google-profile-study/methodology.json)

The data is free to reuse with attribution (CC BY 4.0).

To cite: Underneath. (2026). *Which Google Maps businesses does ChatGPT recommend?* (Version 1.1). Underneath Research. https://underneath.agency/research/chatgpt-local-picks-google-profile-study

## Frequently asked questions

### How does ChatGPT choose which local businesses to recommend?

Our data shows what the businesses it picks have in common, not how it decides. ChatGPT listed about two thirds of the Google Maps top 3, and at the same Maps rank, businesses with more reviews than the local median were about 20 points more likely to be listed, in every set of answers we collected.

### Do Google reviews help a business appear in ChatGPT?

They go together. Businesses with more reviews than others in the same search were more likely to be listed in all 8 sets of answers, and more likely to be listed every time the question was repeated.

### Does a higher star rating matter for ChatGPT recommendations?

Less than review volume. A rating of 4.8 or more added about 10 points, and a rating above the local median about 7 points, after adjusting for review count and Maps rank.

### Does ChatGPT recommend the same businesses every time?

Mostly, but not entirely. Two answers to the same question shared about three quarters of their listed Maps businesses, and 60.0% of the businesses listed at least once were listed in all three repeats.

### Should a business add keywords to its Google Business Profile name?

No. The advantage we saw in our first day of data did not replicate across the other answer sets, and Google’s guidelines for business names do not allow added keywords.

## Related research

- [ChatGPT local recommendations vs Google Maps](https://underneath.agency/research/chatgpt-local-recommendations-study)
- [Do AI answers match a business’s Google profile?](https://underneath.agency/research/ai-business-facts-accuracy-study)
- [Do ChatGPT, Gemini, Perplexity and Claude agree on brands?](https://underneath.agency/research/ai-assistants-brand-agreement-study)
- [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)

## Related guides

- [Does keyword stuffing still work in AI search?](https://underneath.agency/resources/does-keyword-stuffing-work-in-ai-search)
- [What makes ChatGPT or Gemini recommend one hotel over another?](https://underneath.agency/resources/why-ai-recommends-one-hotel-over-another)
- [How can an IT services firm win more projects and support calls from AI search?](https://underneath.agency/resources/it-service-firms-leads-ai-search)
- [Should our reputation priorities for AI assistants differ from those for human customers?](https://underneath.agency/resources/ai-vs-human-reputation-priorities)
- [Can AI search bring a recruiting agency more employer clients?](https://underneath.agency/resources/recruiting-agencies-employer-clients-ai-search)

---

This is the Markdown twin of https://underneath.agency/research/chatgpt-local-picks-google-profile-study. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "ChatGPT local recommendations: stable details, shifting lists"
description: "ChatGPT’s local listings match Google Maps field by field, but which businesses it lists changes from run to run, and cited pages often do not name them."
canonical: "https://underneath.agency/research/chatgpt-local-recommendations-study"
published: 2026-09-26
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Research · Local AI search

# ChatGPT local recommendations: stable details, shifting lists

When people ask ChatGPT for the best dentist, plumber or lawyer in their city, it usually answers with a list of named local businesses, each with an address, a phone number, a star rating and a link. We asked ChatGPT 120 of these questions in the United States, United Kingdom, Canada and Australia, 960 times in all over two days, and compared every business it listed with the Google Maps results collected for the same search on the same day. The name, link, address and phone ChatGPT shows for a business are almost always the same as on Google Maps and the same from one answer to the next. Which businesses make the list is much less stable, and the pages ChatGPT cites often do not name the businesses it lists.

## The short version

1. **The details match Google Maps.** For listings that are the same business at the same location as a Google Maps top-20 entry, 99.7% of names, 98.1% of website links and 99.6% of phone numbers were identical on 26 September. Star ratings matched for 81.4% and review counts for 66.0%.
2. **The details are stable.** When the same business appeared in two answers to the same question on the same day, every displayed field (name, link, address, phone, rating and review count) was identical in 95.7% of pairs; across days it was 89.0%, mostly because review counts change.
3. **The list is not.** Two answers to the same question on the same day had a business-set overlap of 0.65 on a scale where 1 means identical lists. Google Maps’ top 20 for the same searches overlapped 0.87 a day apart.
4. **Overlap with Google Maps depends on the direction you count.** 53.2% of ChatGPT’s listings were in the same-day Maps top 20; ChatGPT listed 67.7% of the Maps top 3 and 40.8% of the Maps top 20 when it showed a list.
5. **Cited pages often do not name the listed businesses.** For 29.3% of listings, a page the same answer cited names the business. For 6.3% no cited page could be read, and for 31.4% only some could, so the share lies between 29.3% and 67.0%.
6. **The four other wordings we tried showed no difference that survived correction for multiple tests.** That holds for those four wordings under this design, not for wording in general.

## What we asked and collected

For five services (dentist, plumber, personal injury lawyer, accountant, and physiotherapist, asked as “physical therapist” in the US) in six cities in each of the four countries, we asked ChatGPT “Who is the best {service} in {city}?” through DataForSEO’s LLM Scraper, which queries the ChatGPT consumer web app from the matching country.

- **26 September 2026:** one run of each question (120 answers) and the Google Maps top 20 for “best {service} in {city}”.
- **27 September 2026:** three more runs of the same question (360 answers), one run of each of four rewordings (480 answers) and the Google Maps top 20 again. The rewordings were “Which {service} in {city} would you recommend?”, “I need a good {service} in {city}. Who should I go to?”, “List the top {services} in {city}.” and “Who is the best {service} in {city}? Please cite your sources.”

All 960 requests returned an answer. ChatGPT showed no business list in 62 of them and cited no source in 31; both stay in every count below.

## What we measured

Every measure has a direction and a denominator, and they differ, so the same “overlap” can give different numbers.

| Measure | Counted over | Question it answers |
|---|---|---|
| Concordance | Listings present in the same-day Maps top 20 | Is each displayed field identical to the Maps entry? |
| ChatGPT to Maps | All ChatGPT listings | Is the listed business in the Maps top 20? |
| Maps to ChatGPT | Maps top-3, top-10 or top-20 entries, per answer | Did the answer list that business? |
| Set stability | Pairs of answers to the same question | How many businesses do two lists share? |
| Attribute stability | The same business in two answers to the same question | Are its displayed fields identical both times? |
| Ordering | Answers with at least 3 present businesses | Does ChatGPT’s order follow Google’s rank? |
| Citation traceability | All listings | Does a page the same answer cites name the business? |

**Present** means the listing is the same business at the same location as a Maps top-20 entry. A listing and a Maps entry that share a website or a name are compared on street address, postcode and phone: a matching street or phone confirms the same branch; a different street or postcode makes it the same company at another location, which is not counted as present. A street or postcode conflict outranks a shared phone number, and a candidate found by phone alone needs the same street. Listings with nothing to compare (4 in all) are counted as present. Of the 7,020 listings that the first version’s matching rule links to Google Maps across both days, 6,699 are the same branch, 317 are another location of the same company (for example another STAR Physical Therapy clinic in Nashville), and 4 have nothing to compare.

**Set overlap** is the Jaccard index: the number of businesses on both lists divided by the number on either list. A business is identified by its Google Maps place when it is present, otherwise by its phone number, otherwise by its name and street.

## Findings

### The details match Google Maps

| Field (listings present in the same-day Maps top 20) | 26 September | 27 September, 3 runs |
|---|---|---|
| Business name, character for character | 99.7% | 99.4% |
| Website link | 98.1% | 98.0% |
| Street number and name | 99.8% | 99.8% |
| Phone number | 99.6% | 99.2% |
| Star rating | 81.4% | 81.7% |
| Review count | 66.0% | 67.3% |

The match extends to details that exist only on Google listings. 95.5% of website links carrying Google tracking tags (such as ?utm_source=gmb) were identical to the Maps link, tags included, and all 115 keyword-stuffed names (such as “Business Name - Emergency Plumber - City”) were reproduced verbatim on 26 September.

Review counts are the exception, and the gap depends on the country. They matched for 99.3% of present listings in the UK and 98.3% in Australia, but only 17.8% in the US and 14.7% in Canada. When the count differed, ChatGPT showed fewer reviews than Google Maps in 95.7% of cases.

The first version of this study reported 95.6% identical names and 93.2% identical links. The difference comes from branches: that version counted a listing as matched when it shared a website with any Maps entry, including another branch of the same chain, whose details naturally differ.

### Overlap with Google Maps, in each direction

| Direction, 26 September | Counted | Share (95% interval) |
|---|---|---|
| ChatGPT listings present in the Maps top 20 | 747 of 1,405 listings | 53.2% (47.8% to 58.5%) |
| ChatGPT listings that are another location of a Maps company | 29 of 1,405 | 2.1% |
| Maps top 3 listed by ChatGPT, answers with a list | 191 of 282 | 67.7% (60.6% to 74.7%) |
| Maps top 3 listed, all answers (no list counts as not listed) | 191 of 333 | 57.4% (50.0% to 65.0%) |
| Maps top 10 listed, answers with a list | 515 of 919 | 56.0% (51.4% to 60.8%) |
| Maps top 20 listed, answers with a list | 746 of 1,829 | 40.8% (37.6% to 44.1%) |

On 27 September the Maps top 3 was listed in 69.4% of cases (answers with a list), within the interval above. Intervals on this page come from resampling the 120 questions, so they allow for answers to the same question being alike.

| Country, 26 September | Answers with a list | Listings present in the Maps top 20 |
|---|---|---|
| United States | 80.0% | 56.3% |
| United Kingdom | 90.0% | 67.7% |
| Canada | 86.7% | 39.2% |
| Australia | 86.7% | 47.2% |

### The same question, a different list

| Measure (original wording) | Value (95% interval) |
|---|---|
| Same business in two answers the same day: every displayed field identical | 95.7% of 3,280 pairs (93.3% to 97.6%) |
| Same business in two answers on different days: every displayed field identical | 89.0% of 2,833 pairs (86.4% to 91.4%) |
| Business listed in one run is listed again in another run the same day | 76.6% (74.0% to 79.0%) |
| Business-set overlap, two runs the same day | 0.65 (0.61 to 0.68), 119 questions |
| Business-set overlap, runs on different days | 0.59 (0.55 to 0.62), 103 questions |
| Google Maps top-20 overlap, one day apart | 0.87 (0.85 to 0.90), 120 questions |

Per question, ChatGPT’s lists a day apart overlapped 0.278 less than Google Maps’ top 20 a day apart (95% interval 0.235 to 0.323). The result does not change when companies are counted by website domain instead of branches: the same-day overlap is still 0.65. The overlaps above use pairs where both answers showed a list; 40 same-day pairs and 66 pairs across days had one answer with no list, and counting those as zero overlap lowers the same-day figure to 0.57.

Whether a list appears at all also varies. 86 of the 120 questions showed a list in every answer and 34 in only some; every question showed a list at least once.

### Order is only weakly related to Google’s ranking

Within an answer, ChatGPT’s order and Google’s Maps rank were weakly and positively related: the mean rank correlation was 0.28 (95% interval 0.19 to 0.37) across the 87 answers on 26 September with at least 3 present businesses. The first business ChatGPT named was in Google’s top 3 in 35.0% of lists (36 of 103). This is an association in the order of two lists; it does not show that Google’s ranking causes ChatGPT’s.

### Can a listed business be traced to a page ChatGPT cites?

We fetched every page ChatGPT cited (1,041 of 1,424 could be read) and checked whether it names each business listed in the same answer, by its full name or its website address. Naming is the only thing checked: a page that names a business is not thereby evidence that the business is good, or that ChatGPT used the page to choose it.

| Listings, 26 September | Share of 1,405 |
|---|---|
| Named on at least one cited page | 29.3% |
| Not named, all cited pages read | 33.0% |
| Not named on the pages read, but some cited pages could not be read | 31.4% |
| No cited page could be read | 6.3% |

So between 29.3% and 67.0% of listings are traceable to a cited page; among listings with at least one readable cited page, the share is 31.3% (25.9% to 36.7%). On the same basis it was higher for businesses present in the Maps top 20 (40.7%) than for those that were not (20.5%). 78.6% of answers with a list had at least one traceable business, and 75.1% of readable cited pages named at least one listed business. On 27 September the figures were similar (30.8% of listings named). Only 1.9% of citations on 26 September pointed to a listed business’s own website.

### Day to day

ChatGPT showed a list more often on 27 September than on 26 September: 94.2% against 85.8% of answers, a difference of 8.3 percentage points (95% interval 1.9 to 15.6). Lists from different days overlapped 0.057 less than lists from the same day (95% interval 0.029 to 0.091). The two days differ in time of day as well as date (04:30 and 16:20 UTC), and the first day has one run, so this does not separate a day effect from ordinary variation.

### Four other wordings

Compared with the original wording on the same day, paired by question, we tested four outcomes for each rewording: whether a list appeared, whether a source was cited, how much the business list overlapped with the original runs (against the overlap between original runs), and the share of identical links. After correcting for the 16 tests, no difference was significant. Before correction, three intervals excluded zero: the lists for “I need a good {service} in {city}” and “List the top {services} in {city}” overlapped the original runs slightly less (by 0.041 and 0.029), and asking for sources gave 4.8 percentage points fewer identical links (95% interval 1.9 to 8.0). We treat these as leads. The Google Maps comparison is not a wording outcome, because the Maps search did not change with the wording.

### How well the matching works

We checked the rules against blind coding by two AI models (Claude Opus and Claude Sonnet) on a targeted sample of 140 items: 25 matches through a website, 20 multi-location businesses, 25 cases the rule first called another location (24 after the postcode fix in the design record), 10 borderline matches, 20 unmatched listings and 40 cited-page decisions. The two coders agreed on all 100 match items and on 39 of 40 page items. Where the coders disagreed with the rule, we settled the case from the raw data and recorded the reason.

| Rule | Checked | Correct |
|---|---|---|
| Listing is the same branch as a Maps entry | 56 | 56 (100%) |
| Listing has no same-branch Maps entry | 44 | 43 (97.7%) |
| Listing is another location of a Maps company | 24 | 23 (95.8%) |
| Cited page does not name the business | 20 | 20 (100%) |
| Cited page names the business | 20 | 16 (80.0%) |

The page rule’s four errors were all chains: the page named another branch of the same brand, which the rule counts because the company name or website domain matches. For chains, “named on a cited page” is therefore a company-level measure. The one missed match was a plumber whose Maps address differs from the address ChatGPT shows but whose phone and branch web page are the same. An earlier round of model coding on stratified random samples of 200 match decisions (first-version matching rule) and 120 page decisions gave agreement of kappa 0.96 and 0.93. This validation is model-coded: no person coded the sample.

## Observed, inferred and unknown

**What we observe.** ChatGPT’s displayed details for a business agree with the same business’s Google Maps entry field by field, including Google-only tracking tags and keyword-stuffed names, and stay the same from answer to answer. Which businesses ChatGPT lists changes much more from run to run than Google’s top 20 does from day to day. The pages ChatGPT cites name only part of the businesses it lists.

**What we infer.** The listing details most likely come from Google business data or a source that mirrors it, since tags and padded names that exist only on Google profiles reappear verbatim. The selection of businesses behaves like a draw that varies from answer to answer, loosely related to Google’s ranking. The cited sources are not a full account of where the list comes from.

**What remains unknown.** Which provider supplies the data; what ChatGPT retrieves or reads before answering (the consumer app shows neither); why it picks one business over another; whether a business named on a cited page was chosen because of that page; and whether the day-to-day difference is a change in ChatGPT or ordinary variation.

## Where this sits in GEO research

Studies of generative engine optimization follow a chain: content → retrieval → selection → citation → prominence → fidelity → stability. Controlled studies test the first links by changing content. This study observes the middle of the chain in real local answers: how selection relates to conventional local search, how faithfully business details are represented, how traceable the list is to citations, and how stable each of these is. It does not test content changes, and it cannot see retrieval.

## How this compares with other studies

The source of ChatGPT’s local data is disputed. A widely repeated claim says 60% to 70% of ChatGPT’s local results come from Foursquare; Steady Demand traced that figure to a small May 2025 study and, in 4,607 ChatGPT runs in August 2026, found Yelp as the grounding source for business cards in 95.83% of runs and Foursquare in 0.00%. Steady Demand also found ChatGPT’s named businesses resembled Google’s set more than Foursquare’s. SOCi’s 2026 Local Visibility Index found ChatGPT recommended 1.2% of the locations of multi-location brands, against a 35.9% appearance rate in Google’s local 3-pack.

Our study measures something different: not which source ChatGPT names, but whether the details it shows agree with Google’s. They do, field by field. Both results can be true at once if the listing card is built from Google-derived data while the text draws on review sites, or if a data provider mirrors Google’s listings; our data cannot say which. None of our answers cited Yelp, Foursquare or TripAdvisor; the most-cited sources were review and ranking sites such as birdeye.com and expertise.com.

Sources: [Steady Demand](https://www.steadydemand.com/chatgpts-local-results-arent-coming-from-foursquare-and-probably-never-really-were/); [SOCi](https://www.soci.ai/news/in-ai-driven-discovery-few-brands-are-chosen-most-disappear/).

## What this means

What follows is interpretation rather than measurement.

- **Check that the Google Business Profile is accurate.** ChatGPT’s details agree with the profile nearly field for field; if the two share a source, an error on the profile would likely appear in ChatGPT too.
- **One answer is one draw.** A business can appear in one answer and not the next; checking visibility needs several runs, not one screenshot.
- **Being on the cited pages is not the whole story.** Between 29.3% and 67.0% of listed businesses were named on a page the answer cited, and for a third no readable cited page named them, so citations alone do not explain who gets listed.

## Methodology

- **Questions:** 120, five services by six cities by four countries (US: Austin, Denver, Charlotte, Phoenix, Columbus, Nashville; UK: Manchester, Bristol, Leeds, Glasgow, Birmingham, Edinburgh; Canada: Toronto, Calgary, Ottawa, Vancouver, Winnipeg, Halifax; Australia: Brisbane, Perth, Adelaide, Melbourne, Sydney, Canberra).
- **ChatGPT:** the consumer web app via DataForSEO LLM Scraper, country-level location; 26 September 2026 at about 04:30 UTC (1 run) and 27 September 2026 at about 16:20 UTC (3 runs and 4 rewordings).
- **Google Maps:** DataForSEO Google Maps SERP, “best {service} in {city}”, same country, top 20, collected on each date; each answer is compared with that day’s results.
- **Same business, same location:** candidates share the listing’s website domain, name or phone; a street or postcode conflict means another location; a matching street, phone or (UK, Canada) postcode confirms the branch. A candidate found by phone alone needs the same street.
- **Cited-page check:** each cited page fetched once (plain request, then a rendering service for failures); a listing counts as named when the page text contains its full normalized name or its website’s domain.
- **Uncertainty:** 95% intervals from 2,000 resamples of the 120 questions; differences are paired by question, with permutation p-values and Holm correction within each family of tests.
- **Design record:** the design file records the rewordings and runs as fixed before the second collection; the file was first committed after that collection and has since been edited, so this cannot be independently verified. The branch-aware matching rule and the revised outcomes were set after a first analysis of these data and refined while inspecting cases; the analysis plan records each change, including a fix found through a validation item (reading the Maps postcode from the full address, because Google’s structured field sometimes holds the street number), with its effect.

## Limitations

- Two dates at different times of day, so day-level change cannot be separated from ordinary variation or from changes in ChatGPT.
- The consumer app shows neither what ChatGPT retrieved nor what it read, so being in the Maps top 20 shows visibility in conventional local search, not retrieval, and the study cannot say why a business was chosen.
- Agreement with Google Maps shows that two displays hold the same data; it does not identify the provider.
- The validation was coded by two AI models, not by people. A sample for human coding was drawn but not coded.
- For chains, the cited-page check works at company level: a page about another branch counts.
- Cited pages were fetched on 27 September and may differ from what ChatGPT read; 383 of 1,424 could not be read.
- Google Maps results were collected at country level for a city search, which may differ from what a searcher in that city sees.
- Only four rewordings were tested; they are not a sample of all the ways people ask.

## Data and downloads

- All 960 answers, with status, wording, run and date: [s9_answers.csv](https://underneath.agency/research-data/chatgpt-local-recommendations-study/s9_answers.csv)
- All 12,271 listings, each with its same-day Google Maps match, match class and field comparisons: [s9_listings_all_days.csv](https://underneath.agency/research-data/chatgpt-local-recommendations-study/s9_listings_all_days.csv)
- Every Google Maps top-20 entry by date: [s9_google_maps.csv](https://underneath.agency/research-data/chatgpt-local-recommendations-study/s9_google_maps.csv)
- Every listing and cited page pair with the check result, and the fetch record of each cited page: [s9_citation_traceability.csv](https://underneath.agency/research-data/chatgpt-local-recommendations-study/s9_citation_traceability.csv), [s9_cited_pages.csv](https://underneath.agency/research-data/chatgpt-local-recommendations-study/s9_cited_pages.csv)
- The targeted validation sample with the rule’s output, both coders’ labels and every adjudication: [s9_validation_targeted.csv](https://underneath.agency/research-data/chatgpt-local-recommendations-study/s9_validation_targeted.csv)
- The design and deviation log: [analysis_plan.txt](https://underneath.agency/research-data/chatgpt-local-recommendations-study/analysis_plan.txt)
- Every statistic on this page: [stats.json](https://underneath.agency/research-data/chatgpt-local-recommendations-study/stats.json)
- Machine-readable methodology: [methodology.json](https://underneath.agency/research-data/chatgpt-local-recommendations-study/methodology.json)
- The first version’s 1,405 listings: [s9_chatgpt_local_listings.csv](https://underneath.agency/research-data/chatgpt-local-recommendations-study/s9_chatgpt_local_listings.csv)

The data is free to reuse with attribution (CC BY 4.0). To cite: Underneath (2026), *ChatGPT local recommendations: stable details, shifting lists*, Underneath Research, https://underneath.agency/research/chatgpt-local-recommendations-study

## Frequently asked questions

### Does ChatGPT copy Google Maps?

Our data shows agreement, not copying. For businesses present in both, ChatGPT’s name, link, street and phone match Google Maps in 98% to 99.8% of cases, including tracking tags that exist only on Google profiles. That points to Google business data or a source that mirrors it, but the study cannot identify the provider.

### Does ChatGPT give the same list every time?

No. Two answers to the same question on the same day shared businesses with an overlap of 0.65 on a scale where 1 means identical lists, and 34 of 120 questions showed a list in some answers but not others. The details of each business, by contrast, were identical in 95.7% of repeat appearances on the same day.

### Does ChatGPT list the businesses at the top of Google Maps?

Often, not always. When it showed a list, ChatGPT listed 67.7% of the Maps top 3 and 40.8% of the Maps top 20. Its order was only weakly related to Google’s ranking.

### Do the sources ChatGPT cites name the businesses it recommends?

Often not. A cited page named the business for 29.3% of listings; because some pages could not be read, the true share is between 29.3% and 67.0%. Naming is not support: the check does not show that a page recommends the business.

### How can a local business appear in ChatGPT’s recommendations?

Our study did not test what makes ChatGPT choose a business. It shows that ChatGPT’s details for a business agree with its Google Business Profile, so checking that the profile is accurate is a sensible step, and that a single answer is one of many possible lists.

## Related research

- [How often do AI Overviews appear?](https://underneath.agency/research/ai-overviews-frequency-study)
- [Do ChatGPT, Gemini, Perplexity and Claude agree on brands?](https://underneath.agency/research/ai-assistants-brand-agreement-study)
- [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)
- [Which Google Maps businesses does ChatGPT recommend?](https://underneath.agency/research/chatgpt-local-picks-google-profile-study)

## Related guides

- [Does keyword stuffing still work in AI search?](https://underneath.agency/resources/does-keyword-stuffing-work-in-ai-search)
- [How can an IT services firm win more projects and support calls from AI search?](https://underneath.agency/resources/it-service-firms-leads-ai-search)
- [How can a CPA or tax firm win better clients through AI search?](https://underneath.agency/resources/accounting-firms-clients-ai-search)
- [Will clients find our engineering firm when they ask AI who can design their project?](https://underneath.agency/resources/engineering-firms-project-inquiries-ai-search)
- [Can AI search bring a recruiting agency more employer clients?](https://underneath.agency/resources/recruiting-agencies-employer-clients-ai-search)

---

This is the Markdown twin of https://underneath.agency/research/chatgpt-local-recommendations-study. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "“Is this brand legit?” How AI assistants build a reputation"
description: "Asked “Is this brand legit?”, four AI engines said yes in every complete answer, and 99.7% raised a problem. How they build a reputation from evidence."
canonical: "https://underneath.agency/research/is-it-legit-ai-reputation-study"
published: 2026-09-26
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Research · AI assistants

# “Is this brand legit?” How AI assistants build a reputation

Before buying, people ask AI assistants whether a company can be trusted. We asked ChatGPT, Gemini, Perplexity and Google AI Mode “Is {brand} legit? What do customers say about it?” for 79 brands across eight industries on 26 September 2026. This version of the study looks past which websites the answers cite. It asks how the engines construct a brand’s reputation from that evidence: which kinds of source they draw on, what positive and negative claims they make, where the negative claims appear, and whether the cited sources support those claims.

Every complete answer said the brand was legitimate, and almost every one then set out problems. Legitimacy and reputation behave as two separate questions: the first always gets a yes, the second never gets an unqualified one. The engines build that reputation from different evidence, and the differences hold after accounting for the brands being asked about.

## The short version

1. **Legitimate, but.** All 312 complete answers affirmed that the brand was legitimate, according to both coders, and 99.7% of them made at least one negative claim about it. No answer painted a purely positive picture: 61.2% gave substantial concerns (95% interval 53.7% to 68.2%) and 38.8% mild ones.
2. **The problems are there, but rarely first.** Of the 5,188 reputation claims in the answers, 35.5% were negative. Negative claims were in the first paragraph of 10.9% of answers: 31.6% for Perplexity, 1.3% for Gemini and for Google AI Mode.
3. **Review platforms carry the negative claims.** 88.0% of answers cited a review or complaint platform. Claims attached only to review platforms were negative 56.5% of the time. Claims attached only to editorial review sites were negative 19.9% of the time, and those attached only to the brand’s own website 6.4%.
4. **The review evidence is concentrated.** Trustpilot and the BBB account for 61.7% of all review-platform citations. Across all sources, 484 domains were cited but the effective number of sources is 30.5.
5. **Each engine has its own evidence regime.** After allowing for brand, brand group and industry, Gemini was far less likely than ChatGPT to cite the BBB (odds ratio 0.05) and Google AI Mode far more likely to cite Reddit (5.76). Perplexity cited the brand’s own site in 94.9% of answers, Google AI Mode in 17.7%.
6. **Citations mostly support the claims they are attached to, when they can be checked.** In a random sample of 240 cited claims, a readable cited page supported the claim fully or in part 72.4% of the time. But 71.4% of claims that said a problem was common rested on evidence that did not show how common it was. Trustpilot, the BBB and ConsumerAffairs block automated reading, so 119 of the 240 claims could not be checked.

## Research questions

| Question | Answered here? |
|---|---|
| RQ1. Which classes of source do the engines cite when judging a brand, and how concentrated are they? | Yes |
| RQ2. Do the engines draw on different evidence, net of brand, brand group and industry? | Yes |
| RQ3. Which positive and negative claims do answers make, and how prominent are the negative ones? | Yes |
| RQ4. Do the cited sources support the claims attached to them? | For a sample, with bounds |
| RQ5. Is the reputation an engine gives a brand stable across runs, wordings and dates? | No: one run per engine and brand |
| RQ6. Are source classes associated with the kind of claims made? | Descriptively only |

## Five parts of an AI reputation

We treat an answer to “Is this brand legit?” as a reputation with five parts, coded separately for every answer, rather than as a single verdict:

- **Legitimacy:** whether the answer says the business is real and legitimate (not whether it is good), and whether its first sentence says so.
- **Overall structure:** positive; legitimate with mild concerns; legitimate with substantial concerns; ambiguous; negative; or undetermined.
- **Complaint profile:** the negative claims, by topic (customer service, billing and fees, product quality, legal and regulatory, ratings), and where in the answer they appear.
- **Evidence:** the sources cited, by class, and which claims each source is attached to.
- **Uncertainty:** whether the answer qualifies the evidence itself, for example by noting a small review sample or that experiences vary by franchise.

This follows the argument that visibility in generative search is a set of distinct measures, not one number. Here the measures are what an engine says about a brand and on what evidence.

## What we analyzed

The data are the 316 answers collected on 26 September 2026 (79 brands, four engines, one run each). No new questions were asked for this version.

- **Brands:** in each of eight industries, five brands the assistants named most widely in our [four-assistant study](https://underneath.agency/research/ai-assistants-brand-agreement-study) and five named by only one assistant (four in home and local services). 21 names were completed or corrected by hand.
- **Claims:** every answer was split into atomic reputation claims, 5,188 in all, each with its polarity, topic, the sources attached to it, its paragraph and how it described prevalence. Each answer was also coded on the five parts above. The coder was Claude Opus, reading the answer with its citations replaced by neutral tags and without being told the engine.
- **Second coder:** Claude Sonnet independently coded the answer-level measures for all 316 answers, and the polarity and topic of the 1,382 claims in 80 randomly drawn answers.
- **Sources:** each of the 484 cited domains was assigned one of eight classes. Domains belonging to the brand or its parent company were identified brand by brand and counted as the brand’s own website.
- **Support:** 60 cited claims per engine were drawn at random. Each was judged against the cited page, fetched on 28 September, or against the passage that Gemini’s link quoted. Two models judged each claim independently.

All coding and judging on this page was done by AI models. No person coded the sample. Answers that were incomplete (4, all ChatGPT) or cited nothing (3 ChatGPT, 3 Gemini) stay in the data as outcomes.

## Study 1: the evidence the engines draw on

### Source classes

| Share of answers citing at least one | All (95% interval) | ChatGPT | Gemini | Perplexity | Google AI Mode |
|---|---|---|---|---|---|
| Review or complaint platform | 88.0% (83.9% to 91.5%) | 92.4% | 68.4% | 97.5% | 93.7% |
| The brand’s own website | 51.6% (47.2% to 56.3%) | 70.9% | 22.8% | 94.9% | 17.7% |
| Editorial review or comparison site | 40.8% (34.2% to 47.2%) | 7.6% | 48.1% | 74.7% | 32.9% |
| Another company’s website | 37.7% (32.0% to 43.4%) | 20.3% | 44.3% | 69.6% | 16.5% |
| Forum, social or video | 28.2% (22.5% to 34.5%) | 11.4% | 26.6% | 31.6% | 43.0% |
| News media | 17.7% (13.9% to 21.5%) | 1.3% | 6.3% | 54.4% | 8.9% |
| Encyclopedia or scholarly reference | 13.9% (10.1% to 17.7%) | 0.0% | 6.3% | 43.0% | 6.3% |
| Government or regulator | 8.2% (5.4% to 11.7%) | 15.2% | 0.0% | 15.2% | 2.5% |
| Consumer advocacy nonprofit | 8.2% (4.4% to 12.7%) | 5.1% | 3.8% | 16.5% | 7.6% |

The individual sites behind these classes are unchanged from the first version: trustpilot.com was cited in 65.8% of answers, bbb.org in 51.3%, reddit.com in 24.7% and consumeraffairs.com in 22.5%.

Counted by citations rather than answers, review platforms made up 56.3% of ChatGPT’s citations and 53.3% of Google AI Mode’s, against 24.3% of Perplexity’s. Perplexity spreads its citations across more classes: a median of 5 source classes per answer, against 2 for each of the other engines.

### Concentration

Counting each domain once per answer, the 316 answers made 1,667 citations to 484 domains. The Herfindahl-Hirschman index, which sums squared shares, is 0.0328, so the answers behave as if they drew evenly on 30.5 sources. The review evidence is far more concentrated. 36 review platforms were cited, but Trustpilot and the BBB account for 61.7% of review-platform citations, and the effective number of review platforms is 4.7.

| Concentration of cited domains | ChatGPT | Gemini | Perplexity | Google AI Mode |
|---|---|---|---|---|
| Distinct domains | 104 | 134 | 376 | 67 |
| Share of the top 3 domains | 48.5% | 21.3% | 18.3% | 48.6% |
| Effective number of sources | 10.4 | 35.4 | 54.9 | 10.6 |

### More citations do not mean more independent evidence

The number of domains an answer cites and the number of source classes it draws on move together (rank correlation 0.84). They are not the same thing. 20.3% of ChatGPT answers drew on a single class of source, and several sites can repeat the same customer reviews. We did not measure whether cited sources are independent of one another, so source classes are only a proxy for diversity of evidence.

### Engine differences, net of the brands asked about

Answers about the same brand are not independent, so we modeled each outcome with brand clusters, adjusting for brand group and industry (odds ratios against ChatGPT, 95% intervals).

| Outcome | Gemini | Perplexity | Google AI Mode |
|---|---|---|---|
| Cites Trustpilot | 0.21 (0.1 to 0.44) | 2.72 (1.25 to 5.91) | 1.5 (0.79 to 2.85) |
| Cites the BBB | 0.05 (0.02 to 0.11) | 1.0 (0.61 to 1.65) | 0.19 (0.1 to 0.36) |
| Cites Reddit | 1.87 (0.79 to 4.42) | 3.15 (1.34 to 7.42) | 5.76 (2.58 to 12.88) |
| Cites the brand’s own website | 0.11 (0.06 to 0.24) | 8.09 (2.72 to 24.11) | 0.08 (0.04 to 0.17) |
| Substantial concerns | 0.21 (0.11 to 0.4) | 1.06 (0.57 to 1.98) | 0.23 (0.12 to 0.43) |
| Negative claim in the first paragraph | 0.13 (0.02 to 0.94) | 5.28 (2.11 to 13.23) | 0.13 (0.02 to 0.94) |
| Qualifies its own evidence | 0.22 (0.1 to 0.47) | 0.28 (0.14 to 0.56) | 0.1 (0.04 to 0.21) |

No Gemini answer cited a government or regulator source, so that outcome is not modeled. A model with a random intercept for each brand gave the same direction for every engine effect. Brand group had no clear effect on any of these outcomes, and no engine-by-group interaction was significant. We read the engines as different evidence-selection regimes, not as better or worse.

## Study 2: what the answers say

### The shape of the answers

| Share of complete answers | All | ChatGPT | Gemini | Perplexity | Google AI Mode |
|---|---|---|---|---|---|
| Legitimate, substantial concerns | 61.2% | 76.0% | 45.6% | 77.2% | 46.8% |
| Legitimate, mild concerns | 38.8% | 24.0% | 54.4% | 22.8% | 53.2% |
| Positive, ambiguous, negative or undetermined | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% |
| First sentence affirms legitimacy, with or without a qualifier (all answers) | 97.2% | 88.6% | 100.0% | 100.0% | 100.0% |
| Qualifies its own evidence | 66.0% | 88.0% | 63.3% | 68.4% | 45.6% |

Perplexity’s first sentence carried a qualifier (“legitimate, but reviews are mixed”) in 31.6% of answers. The ChatGPT answers without a first-sentence verdict (11.4%) opened with a preamble or a list of branches; none disputed that the company was real.

### Legitimate, but

Among the 312 answers that affirmed legitimacy, the median answer gave 33.3% of its claims to negatives. In 20.2% of answers (95% interval 14.6% to 25.6%) the negative claims outnumbered the positive ones. An answer can open “Yes, this is a legitimate company” and spend most of what follows on complaints. A yes to “Is it legit?” says little about the reputation the answer then describes.

### What the claims are about

| Claim topic | Share of all claims | Share of these claims that are negative |
|---|---|---|
| Product or service quality | 22.0% | 37.2% |
| General customer experience | 16.1% | 27.2% |
| Identity and track record | 12.9% | 4.5% |
| Billing, fees, refunds and cancellation | 12.5% | 71.2% |
| Customer service | 12.0% | 69.7% |
| Legitimacy | 11.4% | 0.3% |
| Ratings on named platforms | 8.0% | 36.5% |
| Legal, regulatory or formal complaint record | 2.1% | 81.1% |
| Expert or press assessment | 1.9% | 8.0% |

| Share of complete answers with a negative claim about | All | ChatGPT | Gemini | Perplexity | Google AI Mode |
|---|---|---|---|---|---|
| Customer service | 76.6% | 81.3% | 65.8% | 84.8% | 74.7% |
| Billing, fees or cancellation | 70.2% | 72.0% | 70.9% | 64.6% | 73.4% |
| Product or service quality | 66.3% | 72.0% | 67.1% | 62.0% | 64.6% |
| Legal, regulatory or formal complaint record | 20.8% | 40.0% | 2.5% | 40.5% | 1.3% |

Negative claims were mostly stated as general patterns. 66.0% said a problem was common or frequent without a number, 11.4% gave a number or rating, and 14.1% described individual cases.

### Which sources carry which claims

| Claims whose citations are all from one class | Claims | Negative | Positive |
|---|---|---|---|
| Review or complaint platform | 1,542 | 56.5% | 33.9% |
| Forum, social or video | 140 | 40.0% | 45.7% |
| Another company’s website | 210 | 30.5% | 59.0% |
| Editorial review or comparison site | 351 | 19.9% | 65.8% |
| The brand’s own website | 377 | 6.4% | 66.3% |
| No citation attached | 1,507 | 24.4% | 54.7% |

The engines attached a citation to 71.0% of claims: 93.3% for Perplexity, 56.7% for Google AI Mode. Review platforms supply most of the negative claims that carry a citation. Brand-owned pages supply claims about identity and legitimacy. This is an association between source class and claim type in the answers. It does not show that a source caused a claim.

### Do the cited sources support the claims?

| 240 sampled cited claims (60 per engine) | Claims | Supported or partly supported |
|---|---|---|
| No cited source could be read (mostly Trustpilot, BBB, ConsumerAffairs) | 119 | not judged |
| Readable, but not about the claim’s company or empty | 8 | not judged |
| Judged against fetched cited pages | 58 | 72.4% (62.3% to 82.5%) |
| Judged against fetched pages, every cited source readable | 25 | 84.0% |
| Judged against Gemini’s quoted passages only | 55 | 45.5% (33.3% to 57.1%) |

Against fetched pages, 46.6% of claims were fully supported, 25.9% partly supported and 25.9% not supported; 1.7% were contradicted. When a claim cited several sources and some were blocked, the judge saw only the readable ones, so “not supported” is an upper bound. That is why the rate rises to 84.0% when every cited source could be read. Gemini’s links quote only the start and end of a passage, so its lower figure is partly a limit of the evidence. Across all 240 claims, the supported share lies between 27.9% and 80.8%, depending on how the claims that could not be checked would come out.

The clearest gap concerns prevalence. Of the 63 checkable claims that described a problem as common, 71.4% cited evidence that showed individual reports, or nothing about frequency. The two judges agreed on 81.8% of verdicts (kappa 0.68, supported-or-partly against not).

### Widely named and rarely named brands

Whether a brand was widely named in our earlier study made little difference to how its reputation was built. The difference between the two groups was within the margin of error for citing a review platform, citing the brand’s own site, giving substantial concerns, and putting a negative claim first. The exception was billing: 80.6% of answers about widely named brands made a negative billing claim, against 59.2% for brands named by one assistant (difference 95% interval 7.6 to 35.5 points). This is an association in 79 brands, not an effect of being well known.

### The same brand, four engines

For the 75 brands with four complete answers, all four engines gave the same overall structure for 38.7%. Two engines answering about the same brand shared on average 0.2 of their cited domains (Jaccard index). Different engines give the same brand noticeably different evidence and a different weight of concerns.

## Observed, inferred and unknown

**What we observe.** Every complete answer says the brand is legitimate. Almost every one then makes negative claims, mostly about customer service, billing and product quality, and mostly as general patterns. Review platforms, above all Trustpilot and the BBB, are cited in most answers and carry most of the cited negative claims. The engines differ systematically in which classes of source they cite and in how prominently they place concerns.

**What we infer.** The engines treat “is it legitimate” and “is it good” as separate questions: the first is answered from identity evidence, often the brand’s own pages, and the second from customer-review evidence. The reputation a buyer receives therefore depends on the engine as well as on the brand.

**What remains unknown.** Whether the same engine would give the same reputation on another run, with other wording or on another day. Whether a source that is cited changed the answer. Whether the negative claims reflect the balance of all customer experience or only what review platforms collect. Whether several cited sites repeat the same underlying reviews.

## What this means

The points below are our interpretation. They follow from the findings but were not tested.

- **Being called legitimate is not the finish line.** Every brand in the sample passed that test. What differed was the complaint profile that followed, and that profile draws most heavily on review platforms.
- **Recurring themes on review platforms are the likeliest content of an AI reputation.** Billing, cancellation and support problems appeared in most answers. Fixing them where customers report them addresses the evidence the engines cite.
- **Official pages supply the identity half.** Claims citing the brand’s own site were overwhelmingly about legitimacy and track record. A clear page on who the company is and how it handles complaints gives engines accurate material for that part of the answer, especially for Perplexity and ChatGPT, which cite brand sites most often.
- **Check the engine, not just “AI”.** A brand can get a mostly positive answer from Gemini and a concern-heavy one from Perplexity on the same day.

## Where this sits in GEO research

Research on generative engine optimization began with visibility: whether changing content changes how often it appears in AI answers. Later work compared which sources different AI search systems draw on, and argued that visibility should be measured as several separate quantities (discoverability, citation, claim support, stability) rather than one rank. This study applies that view to a consequential task, judging whether a business can be trusted. It measures source selection, claim construction and claim support in live answers. It does not measure stability over time or causal influence. Those need repeated runs and controlled experiments that add or remove a source.

## Methodology

- **Question and engines:** “Is {brand} legit? What do customers say about it?”; ChatGPT and Gemini consumer apps (DataForSEO LLM Scraper), Google AI Mode (DataForSEO SERP API), Perplexity sonar (DataForSEO LLM Responses API); US location; one run each; 26 September 2026.
- **Brands:** 79 from the national questions of the four-assistant study. Per industry: the five named by the most assistants, and five drawn at random (seed 20260926) from brands named by one assistant.
- **Claim and answer coding:** Claude Opus coded each answer with citations replaced by neutral tags and the engine hidden. It recorded status, first-sentence frame, legitimacy, structure (six categories), uncertainty, and every atomic reputation claim with its polarity, topic, attached sources, paragraph and prevalence wording.
- **Second coder:** Claude Sonnet, blind to the engine and to the first coder’s answer-level codes. Agreement: overall structure 84.2% (kappa 0.68), first-sentence frame 89.9% (kappa 0.63), negative claim in the first paragraph 92.0% (kappa 0.64), uncertainty 72.5% (kappa 0.48), and legitimacy affirmed or not 100.0%. On 1,382 claims, polarity agreed 93.2% (kappa 0.89) and topic 82.8% (kappa 0.8). The uncertainty measure is the least reliable and should be read as indicative.
- **Sources:** registrable domains, each coded into one of eight classes by Claude Opus (the list is in the dataset). Domains owned by the brand or its parent, identified per brand, count as the brand’s own website. The first version matched the brand’s name inside the domain, which found own sites in 39.9% of answers; recognizing parent companies and differently named official sites raises this to 51.6%.
- **Concentration:** each domain counted once per answer; Herfindahl-Hirschman index over domain shares; effective number of sources = 1 / index.
- **Support:** 60 cited claims per engine drawn at random (seed 20260928). Evidence was each cited page fetched over plain HTTP on 28 September 2026, or the passage that Gemini’s link quoted. Claude Opus judged each claim, and Claude Sonnet judged it independently. Claims with no readable source stay in the denominator, and bounds are given. Of the 2,414 cited pages, 38.6% could be read, including 6.6% of review-platform pages.
- **Models:** GEE logistic regression with brand clusters, fixed effects for engine, brand group and industry, and cluster-robust intervals. Check: a logistic model with a random intercept for each brand.
- **Intervals:** 95% bootstrap resampling brands (2,000 resamples).
- **Update schedule:** quarterly.

## Limitations

- One answer per engine and brand, on one date. How stable an engine’s account of a brand is across runs, wordings and days is not measured, and single answers vary.
- Citation is not influence. The data show which sources the engines cite and which claims they attach to them, not that those sources caused the answer.
- All coding and judging was done by two AI models from the same family. Their agreement is a check on consistency, not on accuracy. No person coded the sample.
- The review platforms cited most often block automated reading, so claim support could mostly be checked for other sources. Pages were fetched two days after the answers.
- Whether several cited sources repeat the same underlying reviews was not measured.
- Perplexity was tested through its API. The brand sample leans toward well-known US companies.

## What changed in version 1.1

Version 1.0 (26 September 2026) counted which websites were cited and flagged five kinds of warning with phrase rules. Following an external review, version 1.1 re-analyzes the same 316 answers. It adds claim-level coding, the five-part framework, a source taxonomy with concentration measures, a claim-support check, models that account for answers about the same brand, and a second coder.

The source figures are unchanged. The phrase-rule figures are replaced. The rules caught narrow wordings: they found billing problems in 27.8% of answers and customer service complaints in 30.1%, where claim coding finds negative billing claims in 70.2% and service claims in 76.6%. The complaint rule itself agreed with claim coding on 95.5% of answers. The earlier statement that Trustpilot and BBB profiles shape the answer, and that “what those sites show is what AI repeats”, is withdrawn. The data show that those platforms are a prominent part of the evidence the answers cite, not that they determine the answer. The earlier figure of 81.3% for answers naming a Trustpilot or BBB rating counted any mention of either platform, and is dropped. The full change log is in the dataset.

## Data and downloads

- Every answer with its coded reputation and cited domains: [s16_answers_v11.csv](https://underneath.agency/research-data/is-it-legit-ai-reputation-study/s16_answers_v11.csv) and [JSON](https://underneath.agency/research-data/is-it-legit-ai-reputation-study/s16_answers_v11.json)
- All 5,188 claims with polarity, topic and attached sources: [s16_claims.csv](https://underneath.agency/research-data/is-it-legit-ai-reputation-study/s16_claims.csv) and [JSON](https://underneath.agency/research-data/is-it-legit-ai-reputation-study/s16_claims.json)
- The claim-support sample with both judges’ verdicts: [s16_claim_support.csv](https://underneath.agency/research-data/is-it-legit-ai-reputation-study/s16_claim_support.csv)
- Every cited domain and its source class: [s16_domain_classes.csv](https://underneath.agency/research-data/is-it-legit-ai-reputation-study/s16_domain_classes.csv)
- Every statistic on this page, including models and agreement: [stats.json](https://underneath.agency/research-data/is-it-legit-ai-reputation-study/stats.json)
- Version 1.0 answers and statistics: [s16_legit_answers.csv](https://underneath.agency/research-data/is-it-legit-ai-reputation-study/s16_legit_answers.csv) and [stats_v1.0.json](https://underneath.agency/research-data/is-it-legit-ai-reputation-study/stats_v1.0.json)
- Machine-readable methodology: [methodology.json](https://underneath.agency/research-data/is-it-legit-ai-reputation-study/methodology.json)

The data is free to reuse with attribution (CC BY 4.0).

To cite: Underneath. (2026). *“Is this brand legit?” How AI assistants build a reputation* (Version 1.1). Underneath Research. https://underneath.agency/research/is-it-legit-ai-reputation-study

## Frequently asked questions

### Do AI assistants say a company is legit?

In our test, yes, every time. All 312 complete answers across ChatGPT, Gemini, Perplexity and Google AI Mode said the brand was legitimate. But 99.7% also made at least one negative claim, and 61.2% set out substantial concerns.

### What sources does ChatGPT use to check if a company is legit?

In our test, 92.4% of ChatGPT answers cited a review or complaint platform, most often the Better Business Bureau and Trustpilot. 70.9% cited the company’s own website. It rarely cited editorial review sites (7.6%) or news media (1.3%).

### Does Trustpilot affect what AI says about a business?

Trustpilot was the most-cited source, in 65.8% of answers, and claims citing only review platforms were negative 56.5% of the time. That shows Trustpilot is a prominent part of the evidence AI answers cite. It does not show that Trustpilot causes what the AI says; this study cannot test that.

### Are the complaints AI mentions accurate?

When the cited page could be read, it supported the claim fully or in part in 72.4% of sampled cases. The weaker point is frequency: 71.4% of claims that called a problem common cited evidence that did not show how common it was.

### Which AI assistant cites Reddit most for brand reputation?

Google AI Mode, in 40.5% of its answers, against 11.4% for ChatGPT. Adjusted for the brands asked about, its odds of citing Reddit were 5.76 times ChatGPT’s.

## Related research

- [Do ChatGPT, Gemini, Perplexity and Claude agree on brands?](https://underneath.agency/research/ai-assistants-brand-agreement-study)
- [Which Reddit threads do AI answers cite?](https://underneath.agency/research/ai-reddit-citations-study)
- [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)
- [Do AI answers match a business’s Google profile?](https://underneath.agency/research/ai-business-facts-accuracy-study)

## Related guides

- [Do Google AI Overviews downplay negative content?](https://underneath.agency/resources/do-ai-overviews-downplay-negative-content)
- [How often do AI answers say things their sources do not support?](https://underneath.agency/resources/ai-answers-unsupported-claims)
- [Should our reputation priorities for AI assistants differ from those for human customers?](https://underneath.agency/resources/ai-vs-human-reputation-priorities)
- [Can an AI search engine cite my page for something my page does not say?](https://underneath.agency/resources/ai-citing-pages-for-claims-they-do-not-make)
- [Do ChatGPT and other AI engines cite my own website or third-party reviews?](https://underneath.agency/resources/do-ai-engines-cite-your-own-website)

---

This is the Markdown twin of https://underneath.agency/research/is-it-legit-ai-reputation-study. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How many websites have an llms.txt file? 2026 adoption data"
description: "We checked /llms.txt on 5,902 top websites: 11.5% publish a valid file, 2.1% an llms-full.txt, and 15.7% return an HTML page at the address instead."
canonical: "https://underneath.agency/research/llms-txt-adoption-study"
published: 2026-09-26
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Research · AI crawlers

# How many websites have an llms.txt file? 2026 adoption data

llms.txt is a proposed plain-Markdown file at the root of a website that gives language models a short map of the site. We requested it from every live website in the Tranco top 10,000 on 26 September 2026 and checked whether what came back was a real llms.txt or a web page served in its place. 11.5% of top sites publish a valid file. Even more sites answer the address with an HTML page, which a naive count would score as adoption. A second fetch two days later, with a control address no site publishes, confirmed both figures.

The study measures one thing: whether a site publishes the file. It does not measure whether any AI system requests the file, uses it to find pages, cites the site or uses its content in answers. The section on stages below says what is known about each of those, and from whom.

## The short version

1. 680 of 5,902 live top websites (11.5%) serve a valid llms.txt: HTTP 200, plain text, opening with a Markdown heading as the proposal specifies.
2. 929 sites (15.7%) return an HTML page at /llms.txt, usually a soft 404 or the homepage. For every valid file there are 1.37 of these, so counting any HTTP 200 response more than doubles the apparent adoption.
3. Adoption falls with popularity rank but not by much: 15.2% of the top 1,000 sites, 12.4% of sites ranked 1,001 to 5,000, 10.2% of sites ranked 5,001 to 10,000.
4. Only 2.1% of live sites serve a valid llms-full.txt, the companion file with the full text; 16.3% of llms.txt publishers also publish one.
5. The typical valid llms.txt is 7,706 bytes with 36 links; 80.6% include the one-paragraph summary the proposal recommends, but only 11.9% meet every part of the proposal including links to Markdown pages.
6. The HTML responses are catch-all pages: 96.6% of those sites also answer a made-up address with HTTP 200. None of the 680 valid files is a catch-all response.
7. The measurement is stable: fetched again on 28 September, 99.2% of sites were in the same state and 677 of the 680 valid files were still valid.

## Which stage this study measures

Between a page on a website and a sentence in an AI answer there are several separate steps. A result at one step says nothing about the next, so each claim about llms.txt should name its step.

| Stage | Question | What is known |
|---|---|---|
| Publication | Does the site serve a valid file? | Measured here: 11.5% of live top sites |
| Discovery | Does an AI system request the file? | Not measured here. Ahrefs found most files receive no requests at all |
| Retrieval | Does an AI system use the file to find or read pages? | Not measured by anyone we found |
| Citation | Is the site cited more often because of the file? | Not measured here. SE Ranking found no correlation, which cannot show cause either way |
| Answer use | Does the file’s content change what the answer says? | No evidence either way |

Every finding below is about the first stage.

## What llms.txt is

The llms.txt proposal (llmstxt.org) asks sites to publish a Markdown file at /llms.txt with an H1 title, a short blockquote summary, and sections of links to the pages that matter most, ideally to Markdown versions of those pages. A second file, /llms-full.txt, can hold the full text of the site in one document. The idea is that an AI system can read one small, clean file instead of crawling and parsing a whole site.

## What we measured

For each of the 5,902 domains in the Tranco top 10,000 whose homepage responded with HTTP status below 400, we requested /llms.txt and /llms-full.txt once and classified the response:

| Response at /llms.txt | Sites | Share of live sites |
|---|---|---|
| Valid llms.txt (HTTP 200, text, Markdown H1) | 680 | 11.5% |
| Plain text but no Markdown H1 | 124 | 2.1% |
| HTML page (soft 404, homepage or app shell) | 929 | 15.7% |
| Redirected to a different path | 22 | 0.4% |
| Empty file | 27 | 0.5% |
| Not found or error | 4,120 | 69.8% |

The HTML row is the trap. Many sites answer every unknown path with HTTP 200 and a web page. A study that counted every HTTP 200 response as adoption would report 30.2% (1,782 sites); the real figure is less than half that.

The headline has a 95% interval of 10.7% to 12.4%. Counting all 10,000 domains, including the sites that did not answer us, gives a floor of 6.8%.

## Findings

### Who publishes llms.txt

| Tranco rank | Live sites | Valid llms.txt | Share |
|---|---|---|---|
| 1 to 1,000 | 534 | 81 | 15.2% |
| 1,001 to 5,000 | 2,323 | 289 | 12.4% |
| 5,001 to 10,000 | 3,045 | 310 | 10.2% |

The gap between the top 1,000 and the bottom tier is 5.0 percentage points, with a 95% interval of 2.0 to 8.5, so it is unlikely to be chance. Split into ten bands of 1,000 sites, the decline is uneven: 15.2% in the top band, 9.1% in the next, between 11.1% and 14.4% from rank 2,001 to 8,000, and 6.2% in the last. A trend test over the ten bands confirms a downward slope (z = −4.24), but rank explains only part of who publishes.

The highest-ranked publishers are mostly infrastructure and developer platforms: cloudflare.com, github.com, fastly.net, digicert.com, wordpress.org, adobe.com, opera.com, sentry.io, samsung.com, wordpress.com, dropbox.com, shopify.com, ubuntu.com, stripe.com and hubspot.com are among them. Documentation-heavy companies, whose customers already ask AI assistants how to use their products, are well represented.

### The HTML answers are catch-all pages, not files

On 28 September we also requested /llms-7f3a9c21b8e4d056.txt from every site, an address no site publishes. It works as a control: a site that returns the same kind of response there is answering any address, not serving a file.

| State at /llms.txt on 26 September | Sites | Control address also returned HTTP 200 |
|---|---|---|
| HTML page | 929 | 897 (96.6%) |
| Valid llms.txt | 680 | 75, none with the same content as the file |

So the HTML responses are what they look like: generic pages served for any unknown address. The valid files are real. 605 of the 680 sites (89.0%) returned an error at the control address, and none of the 75 that answered it returned the same text as their llms.txt.

### Two fetches, two days apart

We fetched /llms.txt from all 5,902 sites again on 28 September and classified the response with the same rule. 99.2% of sites were in the same state (Cohen’s kappa 0.983). Of the 680 valid files, 677 were still valid; 3 had gone and 11 new ones appeared, giving 688 (11.7%) on the second date. One fetch was enough for the headline, and the small net change fits the rapid growth other trackers report.

### What the files contain

- Median size 7,706 bytes; 10% of files are larger than 54,044 bytes.
- Median of 36 links.
- 95.0% organize links under H2 sections and 80.6% include the blockquote summary.
- Only 14.6% link to Markdown (.md) versions of pages, the part of the proposal that saves an AI system the most work.

Read as a ladder, each part of the proposal loses sites:

| Meets | Valid files | Share |
|---|---|---|
| Markdown H1 title | 680 | 100% |
| plus a blockquote summary | 548 | 80.6% |
| plus H2 sections of links | 500 | 73.5% |
| plus links to Markdown pages | 81 | 11.9% |

### llms-full.txt is much rarer

123 live sites (2.1%) serve a valid llms-full.txt. Of the 680 sites with a valid llms.txt, 111 (16.3%) also publish the full-text file.

### Some publishers block the crawlers the file is written for

33 sites (4.9% of llms.txt publishers) also block GPTBot, ClaudeBot or OAI-SearchBot from the whole site in robots.txt. The file invites AI systems in; the robots.txt turns some of them away.

## How this compares with other studies

| Source | Sample and date | Figure |
|---|---|---|
| This study | Tranco top 10,000, 5,902 live sites, September 2026, strict validity rule | 11.5% valid |
| Rankability tracker | Tranco top 1,000, checked September 2026 | 9.3% (93 of 1,000) serve llms.txt or llms-full.txt |
| Casey Burridge (HTTP Archive data) | 7,504 crawled sites from the top 10,000, June 2026 | 5.61% valid |
| SE Ranking | About 300,000 domains, November 2025 | 10.13% |
| HTTP Archive, Web Almanac 2025 | Whole web, July 2025 crawl | 2.13% of desktop sites |

Adoption is rising fast: Burridge’s tracking started at 1.04% of the top 10,000 in July 2025. Differences between the figures also come from validity rules and from platform rollouts: Burridge attributes much of the recent growth to Shopify enabling the file across its platform, and the Web Almanac attributes 39.6% of files to one SEO plugin.

**Does anything read it?** Ahrefs examined llms.txt requests across 137,210 sites it monitors and found that “97% of those files received zero traffic in May 2026”; AI retrieval bots made 1.1% of the requests that did arrive. SE Ranking found no correlation between having the file and being cited by AI. Google has said it does not use llms.txt, and in 2025 Google’s John Mueller said none of the AI services had said they were using it. Anthropic publishes its own llms.txt, which is not the same as its crawlers reading yours.

Sources: [Rankability](https://www.rankability.com/blog/llms-txt-adoption/); [Casey Burridge](https://caseyrb.com/blog/state-of-llms-txt-adoption/); [SE Ranking](https://seranking.com/blog/llms-txt/); [HTTP Archive Web Almanac 2025](https://almanac.httparchive.org/en/2025/seo); [Ahrefs](https://ahrefs.com/blog/llmstxt-study/).

## What this means

This study separates three kinds of claim.

**Established by this data.** About one in nine live top sites publishes a valid llms.txt, more among the most-visited. Counting any response at the address roughly doubles that figure, because most of the extra responses are catch-all pages. The count was stable over two days. Most files follow the basic format; few link to Markdown versions of pages.

**Supported interpretation.** Publication is concentrated among developer platforms and documentation-heavy companies, and platform rollouts move the numbers more than individual decisions do. A file is cheap to publish, which fits wide adoption without evidence of benefit.

**Open questions.** Whether AI systems request the file, whether they use it to reach pages, and whether it changes citation or answers. None of these has been tested with a control: a comparison of sites with and without the file, or of pages listed and not listed in it, under otherwise equal conditions.

### What to do with it

- **Treat llms.txt as a small, cheap convenience, not a ranking factor.** Nothing here shows it changes whether an AI system cites you.
- **If you publish one, make it valid.** An HTML page at /llms.txt is worse than nothing for a tool that expects Markdown. Serve text, start with an H1, add the summary, and link to clean Markdown versions of key pages.
- **Keep it consistent with robots.txt.** A file that welcomes AI systems next to rules that block them sends mixed instructions.

## What we will test next

Each open question above can be turned into a test. These are hypotheses, not findings:

1. **Discovery.** AI crawlers request a valid /llms.txt more often than an equally sized control file at another address on the same site. Test: server logs from sites that publish both.
2. **Retrieval.** Pages linked from llms.txt are fetched by AI user agents more often than similar pages on the same site that are not linked. Test: matched pages, logs over several weeks.
3. **Citation.** Adding a file does not by itself change how often a site is cited. Test: repeated AI answers for the same questions before and after publication, against sites that did not publish.
4. **Platform rollouts.** When a platform enables the file for all its customers, any advantage for an individual site shrinks. Test: track cited Shopify stores across the rollout.

## Methodology

- **Sample:** Tranco list L5PZ4, top 10,000 domains, fetched 26 September 2026; 5,902 had a live homepage (status below 400, more than 200 bytes).
- **Collection:** one HTTPS GET each for /llms.txt and /llms-full.txt, following redirects, with an identifying research user agent.
- **Validity rule:** HTTP 200, not HTML (by content type or markup), final path still /llms.txt, at least 20 characters, first non-blank line a Markdown H1. llms-full.txt uses the same rule without the H1 requirement.
- **Control address and second fetch:** on 28 September 2026 we requested /llms.txt again and /llms-7f3a9c21b8e4d056.txt once from each of the 5,902 live sites, with the same collector and rule. A valid file counts as a catch-all when the control address returns the same text.
- **Statistics:** Wilson 95% intervals for shares, a Newcombe interval for the tier difference, and a Cochran-Armitage test for the trend across ten rank bands.
- **Update schedule:** quarterly.
- **Version 1.1** (28 September 2026) follows an external methodology review. It adds the stage table, the control address, the second fetch, the intervals and the conformance ladder, and separates findings from interpretation. No 1.0 figure changed.

## Limitations

- Sites that blocked our user agent at the homepage are outside the denominator.
- Presence is not use: we measured publication, not whether AI crawlers fetch the file.
- Two fetches, two days apart; files are being added quickly, so the next edition may differ materially.
- One user agent. A site may answer AI crawlers’ user agents differently from ours.
- The rank bands are uneven, so the tier figures describe these sites, not a smooth rule about popularity.

## Data and downloads

- Per-domain results: [s2_llms_by_domain.csv](https://underneath.agency/research-data/llms-txt-adoption-study/s2_llms_by_domain.csv) and [JSON](https://underneath.agency/research-data/llms-txt-adoption-study/s2_llms_by_domain.json)
- Version 1.1 per-domain file with both dates, the control address and the conformance ladder: [s2_llms_by_domain_v11.csv](https://underneath.agency/research-data/llms-txt-adoption-study/s2_llms_by_domain_v11.csv)
- Every statistic on this page: [stats.json](https://underneath.agency/research-data/llms-txt-adoption-study/stats.json)
- Machine-readable methodology: [methodology.json](https://underneath.agency/research-data/llms-txt-adoption-study/methodology.json)

The data is free to reuse with attribution (CC BY 4.0).

To cite: Underneath. (2026). *How many websites have an llms.txt file? 2026 adoption data*. Underneath Research. https://underneath.agency/research/llms-txt-adoption-study

## Frequently asked questions

### What percentage of websites have an llms.txt file?

In September 2026, 11.5% of the live websites in the Tranco top 10,000 served a valid llms.txt. Among the top 1,000 sites the figure is 15.2%.

### How do I check whether a site’s llms.txt is valid?

Open /llms.txt in a browser. A valid file is plain text, not a web page, and its first line is a Markdown heading starting with “# ”. If you see the site’s normal design or a “page not found” page, the site does not have one, even if the address loads.

### What is the difference between llms.txt and llms-full.txt?

llms.txt is a short index: a title, a summary and links to the most useful pages. llms-full.txt contains the full text of the site’s key content in one Markdown file. 2.1% of live top sites publish llms-full.txt.

### Does publishing llms.txt get a site cited by AI assistants?

Our data cannot show that. It measures who publishes the file, the first of several stages between a website and an AI answer, not whether AI systems request or use it. Treat llms.txt as a low-cost convenience for AI tools, not as a substitute for crawlable pages and clear facts.

## Related research

- [Which AI crawlers do top websites block?](https://underneath.agency/research/ai-crawler-blocking-study)
- [Do websites serve Markdown to AI agents?](https://underneath.agency/research/agent-readable-web-study)
- [How often do AI Overviews appear?](https://underneath.agency/research/ai-overviews-frequency-study)
- [Do AI Overviews cite the pages that rank?](https://underneath.agency/research/ai-overview-citations-study)

## Related guides

- [Can our server logs show what content AI bots are looking for on our site?](https://underneath.agency/resources/ai-bot-server-logs-content-demand)
- [How much web content is already optimized for AI search?](https://underneath.agency/resources/how-much-web-content-is-optimized-for-ai-search)
- [The AI search glossary, defined plainly.](https://underneath.agency/resources/glossary)
- [How do developer tool companies win users when developers ask AI first?](https://underneath.agency/resources/developer-tools-ai-search)
- [What GEO practices are actually supported by research?](https://underneath.agency/resources/what-geo-practices-does-research-support)

---

This is the Markdown twin of https://underneath.agency/research/llms-txt-adoption-study. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "How many “best of” lists cited by AI rank their own brand first?"
description: "Of 269 AI-cited numbered “best X” lists with an identifiable publisher, 24.2% ranked their publisher first. Such lists were 1.1% of all citations."
canonical: "https://underneath.agency/research/self-promoting-best-lists-study"
published: 2026-09-26
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Research · AI search

# How many “best of” lists cited by AI rank their own brand first?

“Best X” list articles are among the page types ChatGPT cites most, according to Ahrefs. Some are written by companies that sell X, and some of those put themselves at number one. We took every numbered “best” list cited by Google AI Overviews, AI Mode, ChatGPT, Gemini, Perplexity and Claude in our earlier studies (US, collected 26 September 2026), checked whether the publisher ranks itself first, and then asked four further questions: is that measurement reliable, how often do such lists also appear in conventional search results for the same queries, do the AI answers name the list’s top pick, and does the picture hold across repeated runs? About one in four cited numbered lists with an identifiable publisher put that publisher first (65 of 269). That is a share of cited numbered lists, not of citations: these lists were 1.1% of all citations. In this sample, answers named a self-ranking publisher at a rate not statistically distinguishable from the top pick of an independent list.

## The short version

1. Of 269 numbered “best X” lists with an identifiable publisher cited by the six AI surfaces, 65 (24.2%, 95% interval 18.1% to 30.5%) ranked their own publisher first. The denominator is cited numbered lists, not all citations. Our detector misses about one real case in three, so the true share is likely higher.
2. Publishers that included themselves almost always went first: 92.9% of self-including lists put the publisher at number one.
3. Measured against everything the answers cited, self-ranking lists are a small part: 92 of 8,565 citations (1.1%), and 66 of 1,206 answers (5.5%) cited at least one.
4. We found no statistically detectable difference across the six surfaces in this sample (p = 0.41). Shares ranged from 5.9% (ChatGPT, 17 lists) to 34.2% (AI Overviews, 38 lists) with wide intervals, so this does not show that the surfaces behave the same.
5. Self-ranking lists are also present in conventional search results: 29.0% of numbered lists in Google’s top 10 for the AI Overview and AI Mode keywords, and 35.2% for the assistants’ questions. These describe different pools of pages from the 24.2% above, not one common share, and the study cannot see which pages the AI engines retrieved.
6. Within the 55 answers that cited both a self-ranking and an independent list, there was no statistically detectable difference in this sample in how often each list’s #1 entry was named (odds ratio 1.01, interval 0.52 to 2.75). Across all 590 answer–list pairs, which are not independent, the rates were 47.1% (of 136) and 43.4% (of 454). This is a proxy for answer-level representation, not a measure of influence.
7. Within one session, exposure was stable: over five runs minutes apart on ChatGPT, Gemini and Perplexity, in 54 of 60 question-and-surface combinations no self-ranking list was cited in any run, and in 5 one was cited every time. Stability across days was not tested.

## What we measured

We took every page cited by the six AI surfaces in our studies of [AI Overviews](https://underneath.agency/research/ai-overview-citations-study), [AI Mode](https://underneath.agency/research/ai-mode-vs-ai-overviews-study) and the [four assistants](https://underneath.agency/research/ai-assistants-brand-agreement-study), all US and collected on 26 September 2026, and kept those whose title or address marked them as a ranked list (“best …”, “top 10 …”): 1,574 pages. We fetched each page and read its numbered entries.

A list ranks its publisher first when the publisher’s name begins its first numbered entry. The publisher’s name is the name part of its web address (“rippling” for rippling.com) or the site name the page states about itself.

The six surfaces are not the same kind of system. AI Overviews and AI Mode were captured from Google’s results pages, ChatGPT and Gemini from their consumer apps, and Perplexity and Claude through their APIs (Perplexity’s sonar model, and Anthropic’s claude-haiku-4-5 model with web search). We report them side by side but label the difference.

## What this study observes, and what it does not

Visibility in generative search runs through several stages that the research literature treats separately: whether a page is retrieved, what the model is shown, whether the page is cited, how prominently, how much of the answer it shapes, whether the answer represents it accurately, and what the reader does next. This study observes some of these directly, measures others through a proxy, and cannot see the rest.

| Stage | In this study | How |
|---|---|---|
| Retrieval by the AI engine | Not observed | The engines do not expose which pages they retrieved |
| Presence in conventional search results | Observed (Google top 10; Bing not used) | Shows what conventional search offered for the same query; not a measure of AI retrieval |
| What the model was shown, and in what order | Not observed | |
| Citation | Observed | The list’s address appears among the answer’s sources |
| Prominence | Partial proxy | Whether the list’s #1 was the answer’s first recommendation (assistant answers only); citation order used only as a model covariate |
| Answer-level representation | Proxy | The answer names the list’s #1 entry |
| Absorption (how much of the answer’s facts, wording or structure comes from the source) | Not measured | |
| Accuracy of what the answer attributes to the source | Not measured | |
| Exposure across repeated runs | Observed, within one session | Five runs minutes apart on ChatGPT, Gemini and Perplexity |
| Clicks, decisions, conversions | Outside the scope | |

The study therefore measures one source class at three observable points (citation, conventional-search availability and answer-level representation), with a partial proxy for prominence, plus repeated-run exposure. It is not a model of the whole pipeline, and none of its comparisons identify causes.

## How reliable the measurement is

We drew 299 pages across nine groups (self-ranking lists, other numbered lists, pages where parsing rules disagree, unnumbered pages, and cited pages our filter did not flag as lists). Two AI models, Claude Opus and Claude Sonnet, coded each page independently from a saved copy of the page as fetched, without seeing the detector’s result. An AI analyst (Claude) settled their disagreements by reading the page, and each decision is logged. This is model coding, not human coding.

| Check | Result |
|---|---|
| Coders agree the publisher is ranked first | Cohen’s kappa 0.99 |
| Coders agree the page is a ranked list | Cohen’s kappa 0.91 |
| Detector finds real self-first lists (test half) | 64.9% (interval 45.4% to 80.4%) |
| Detector avoids false self-first calls (test half) | 98.0% |
| Self-first calls that are correct (test half) | 93.3% |
| A simpler rule (web address only, every numbered heading) on the same test pages | finds 59.0%, avoids 100.0% |

The detector rarely calls a list self-ranking when it is not, but misses about a third of real cases. In the test half, most misses were pages with more than one numbered section where the detector read the wrong one; the others were a very short brand name and a product name that does not start with the brand (“Ascend by Kindsight”). Of the 65 lists behind the headline, 47 were also in the coded sample, and 4 of those were detector errors, for example two VentureSmarter pages that put ZenBusiness first. We did not adjust the headline for this, because the test sample was too small to estimate the miss rate precisely; read the figures as lower bounds.

Two further checks limit what the data can say. Extending the detector to unnumbered pages (tables, plain lists, structured data) did not work well: it picked the right first entry only 38.7% of the time, so we left those pages out. And of 50 cited pages our title filter did not flag as lists, 5 were ranked lists, 2 of them self-ranking, so some lists are outside the study altogether.

## Findings

### How often the publisher ranks itself first

| Measure | Unit and denominator | Count | Share (95% interval) |
|---|---|---|---|
| Ranks its publisher first | Cited numbered lists with an identifiable publisher | 65 of 269 | 24.2% (18.1% to 30.5%) |
| Includes its publisher | Cited numbered lists with an identifiable publisher | 70 of 269 | 26.0% |
| Puts the publisher first | Lists that include their publisher | 65 of 70 | 92.9% |
| Points to a self-ranking list | All citations in the answers | 92 of 8,565 | 1.1% (0.7% to 1.5%) |
| Cites at least one self-ranking list | All answers, including those with no citations | 66 of 1,206 | 5.5% (3.9% to 7.1%) |

The 269 lists are those with a usable publisher name; 12 more numbered lists had none. The 65 lists came from 55 different publishers (a few of them detector errors, see above); the most frequent were Zendesk and Rippling (3 lists each) and HubSpot, Klaviyo, Huntress, TMetric and CoCountant (2 each); VentureSmarter’s two lists are left out because both were detector errors. Most are software companies writing about their own category; the rest include a roofing company, a law firm, the car rental company Sixt, a heating and cooling contractor and accounting firms.

368 of the 1,574 pages could not be fetched. If the missing pages contain numbered lists at the same rate as the fetched ones, the true share lies between 18.5% and 41.9% in the worst and best cases.

Adding the answers we had already collected for repeat runs, reworded questions and the US arm of our country study gives 358 lists, of which 21.8% ranked their publisher first (16.8% to 27.3%). In the UK, Canada and Australia arms the share was 18.1% of 72 lists.

### By the AI surface that cited the list

| AI surface | How it was collected | Numbered lists cited | Publisher ranks itself first (95% interval) | Answers citing one |
|---|---|---|---|---|
| AI Overviews | Google results page | 38 | 34.2% (21.2% to 50.1%) | 4.7% of 486 |
| Claude | API | 35 | 28.6% (16.3% to 45.1%) | 11.3% of 80 |
| AI Mode | Google results page | 26 | 26.9% (13.7% to 46.1%) | 3.0% of 400 |
| Perplexity | API | 161 | 23.6% (17.7% to 30.7%) | 18.8% of 80 |
| Gemini | Consumer app | 31 | 19.4% (9.2% to 36.3%) | 7.5% of 80 |
| ChatGPT | Consumer app | 17 | 5.9% (1.0% to 27.0%) | 1.3% of 80 |

A list cited by more than one surface counts once for each. A test across all six surfaces (logistic model, errors clustered by publisher) found no statistically detectable difference in this sample (p = 0.41). With 17 to 161 lists per surface, moderate differences could go undetected, so this is not evidence that the surfaces behave the same. Comparisons within the same kind of collection also include zero: AI Overviews minus AI Mode 7.3 points (−10.5 to 25.6), ChatGPT minus Gemini −13.5 points (−31.0 to 5.5), Perplexity minus Claude −5.0 points (−21.9 to 12.0). We therefore do not claim that any surface favors these lists.

### Compared with what search shows for the same questions

We compared the cited lists with the numbered lists that Google showed in its top 10 for the same keywords and questions. The top 10 shows what conventional search offered for the same query. It is not a measure of which pages an AI engine retrieved or could have used, which the study cannot see. We used Google only: the Bing results collected for the assistants’ questions were unrelated to the questions and had already been discarded by the study that collected them.

- **Among lists in the search top 10**, 29.0% (93 lists, AI Overview and AI Mode keywords; interval 16.4% to 43.5%) and 35.2% (54 lists, the assistants’ questions; 22.2% to 48.1%) ranked their publisher first. These are descriptive figures for different source pools from the 24.2% among cited lists, not estimates of a common population share; the formal comparisons are the two below.
- **Overlap was low for both kinds of list.** 18.6% of 97 citations of self-ranking lists and 11.3% of 300 citations of independent lists (counted per question and surface, including repeat runs and the US arm of the country study) pointed to a page in the same question’s search top 10, a difference of 7.2 points computed before rounding (−3.3 to 19.7), not statistically detectable in this sample.
- **Among numbered lists in the same question’s top 10**, there was no statistically detectable association between self-ranking and being cited in this sample (odds ratio 0.37, interval 0.04 to 1.54, from only 18 question-and-surface combinations with both kinds). This describes citation among conventional-search results, not how engines select pages.

Self-ranking lists were present both among conventional-search results and among AI citations, but the study cannot observe the AI engines’ internal retrieval process. These are observational associations with conventional-search availability, not evidence about how the engines retrieve or select pages.

### Answer-level representation: do answers name the list’s top pick?

For every answer that cited a numbered list (590 answer-and-list pairs in 319 answers, from the main answers plus the repeat runs and the US arm of the country study), we checked whether the answer text names the list’s #1 entry. This measure operationalizes an observable proxy for answer-level representation; it does not directly measure absorption in the broader sense used in the research literature (how much of an answer’s facts, wording or structure comes from a source), and it does not show that the list caused the answer to name the company. Names that appear only as a citation label were not counted. On 120 answer-name pairs coded by two AI models, the rule never counted a name the coders said was absent (100.0%) and found 87.1% of real mentions.

| Kind of list | Answer names the list’s #1 entry |
|---|---|
| Publisher ranked itself first | 47.1% of 136 |
| Independent list | 43.4% of 454 |

The pairs are not independent: one answer can cite several lists, and one question has answers from several surfaces and runs. The pooled difference, 3.7 points (−10.9 to 21.5, resampling questions), is not statistically detectable in this sample. The primary comparison uses only the 55 answers that cited both a self-ranking and an independent list and compares the lists within each answer, which holds the question, surface and answer fixed: odds ratio 1.01, with an interval of 0.52 to 2.75 when the 28 questions behind these answers are resampled (0.59 to 1.75 if the answers were treated as independent). In 16 of the 55 answers only the self-ranking list’s #1 was named, in 19 only the independent list’s #1, in 17 both and in 3 neither (an answer counts as naming a kind if it named the #1 of any cited list of that kind; the odds ratio uses all 211 answer–list pairs in these answers). Among other checks, a model adjusting for surface, industry, citation order, whether the entity has a Wikipedia article, how many cited lists put it first and whether another page from the same publisher was cited, with errors clustered by question, gives 1.36 (0.79 to 2.36). Restricting to #1 entries that clearly read as names reverses the direction (47.4% of 135 against 53.1% of 369 pairs; difference −20.4 to 13.3 points). In the ChatGPT, Gemini, Perplexity and Claude answers, the list’s #1 was the answer’s first recommendation 9.8% of the time for self-ranking lists and 18.6% for independent lists (difference −20.4 to 6.6 points).

The pooled figure hides differences by surface in small samples: in AI Overviews and AI Mode answers the self-ranking publisher was named more often (62.5% of 24 and 61.5% of 13, against 30.4% and 25.8% for independent lists), and in Claude answers less often (50.0% of 10 against 72.0% of 25). These splits are small and were not tested separately.

Other answers to the same question that did not cite the list named the same #1 entries in 26.2% (self-ranking) and 33.4% (independent) of cases.

### How stable it is

We had asked 20 questions five times each on ChatGPT, Gemini and Perplexity, minutes apart, so this is stability within one session. In 54 of the 60 question-and-surface combinations no self-ranking list was cited in any run, in 5 one was cited every time, and only 1 changed between runs. Which question it was explained most of the variation (intraclass correlation 0.77). Rewording the question changed whether a self-ranking list was cited in 2 to 5 of 60 pairs, depending on the wording. We have not yet repeated the collection on a different date.

### Robustness

Across the choices below, the share stays between 18.6% and 26.7%.

| Variant | Lists | Share ranking publisher first |
|---|---|---|
| Detector used for the headline | 269 | 24.2% |
| Simpler detector (web address only, every numbered heading) | 285 | 18.6% |
| Without pages fetched on a retry | 265 | 24.2% |
| Including Google question keywords | 272 | 23.9% |
| Leaving out the five most frequent self-ranking publishers (ties broken alphabetically) | 253 | 20.9% |
| One list per publisher | 195 | 26.7% |

Reading unnumbered pages as well gives 12.1% of 810 lists, but that reading failed validation (see above), so it is not used.

## How this compares with other studies

| Source | Sample and date | Figure |
|---|---|---|
| This study | 269 numbered lists with an identifiable publisher cited by 6 AI surfaces, 8 industries, September 2026 | 24.2% ranked their publisher first; 1.1% of all citations |
| Peec AI | 232,000 citations for software-review prompts, December 2025 to February 2026 | “Approximately 1 in 10 citations” came from self-promotional listicles; ChatGPT 3.6%, AI Mode 10.3%, Perplexity 10.4% |
| Ahrefs | 26,283 ChatGPT source URLs, December 2025 | “Best X” lists were 43.8% of cited page types |
| Lily Ray | 100 B2B queries in AI Overviews, April to June 2026 | When a brand’s own list was cited, the brand “was left out of the actual recommendation 69% of the time” |
| Scrunch | About 10,000 cited URLs, May to June 2026 | Reports that brands cited via their own list had a recommendation rate of “roughly 7%”, against “about 4%” |

Peec AI measures self-promotional lists as a share of all citations in software prompts; our comparable figure across eight industries is 1.1%. Lily Ray found that brands cited through their own lists were usually left out of the recommendation, and Scrunch reported a small difference in recommendation rate; both measure recommendation. We measure whether answers named the list’s top pick, and found no statistically detectable difference in this sample. The studies use different samples and measures.

Sources: [Peec AI](https://peec.ai/blog/self-promotional-listicles-analysis-from-232k-citations); [Ahrefs](https://ahrefs.com/blog/best-lists-research/); [Lily Ray](https://lilyraynyc.substack.com/p/why-calling-yourself-the-best-could); [Scrunch](https://scrunch.com/blog/what-happens-when-you-crown-yourself-number-1-listicles-impact-on-ai-answers/) (figures checked via [PPC Land](https://ppc.land/self-ranked-brands-gain-ai-recommendations-from-4-to-7-scrunch-finds/)).

### How this relates to research on generative engine optimization

- **Aggarwal et al. (2024), “GEO: Generative Engine Optimization”** ([arXiv](https://arxiv.org/abs/2311.09735)), showed that rewriting a page’s content changed how much of a generated answer was attributed to it, in a setting where the page was already among the few sources given to the model. It studied content changes mainly with retrieval held fixed (the top five Google results placed in the model’s context), plus a smaller test on Perplexity.
- **Chen et al. (2025), “Generative Engine Optimization: How to Dominate AI Search”** ([arXiv](https://arxiv.org/abs/2509.08919)), found that AI search engines drew on different source ecosystems from Google and from one another, citing earned media more heavily than brand-owned and social sources, with differences across engines, languages and phrasings. The authors note that, in their classification, the line between an earned outlet and a brand’s own blog can sometimes be blurry.
- **Martinez (2026), a critical survey of the field** ([arXiv](https://arxiv.org/abs/2607.14035)), argues that visibility in generative search is a multistage, partly unobservable and variable process, and that retrieval, citation, prominence, absorption, accuracy and downstream behavior should be measured separately.
- **This study** applies that stage-by-stage view to one source class that sits on the blurry line Chen et al. describe: a brand-owned page written as a comparative review. It separates citation, conventional-search availability, answer-level representation and repeated-run exposure, without observing the engines’ retrieval or identifying causes.

## What this means

These points separate what was measured from interpretation, what is not established, practical implications and open questions.

- **Measured:** self-ranking commercial lists are a measurable subset of the sources cited by these six surfaces on 26 September 2026 (US): about a quarter of cited numbered lists with an identifiable publisher, and 1.1% of all citations.
- **Measured:** in this sample, answers named a self-ranking list’s own publisher at a rate not statistically distinguishable from the top pick of an independent list. Ranking oneself first was not associated with a statistically detectable difference in being named in this sample (within-answer odds ratio 1.01, 0.52 to 2.75); the design cannot say whether self-ranking has any effect.
- **Interpretation:** readers of AI answers that cite such lists are sometimes reading a vendor’s ranking of itself.
- **Not established:** whether self-ranking lists improve or harm readers’ decisions.
- **Practical implication (not a measured finding):** a list that states its publisher’s interest lets readers weigh it accordingly.
- **Open question:** in small samples, answers in Google’s AI surfaces named a self-ranking publisher more often than an independent list’s top pick; a larger study should test this directly.

## Methodology

- **Pages:** every page cited by AI Overviews and AI Mode (800 US head and tail keywords) and by ChatGPT, Gemini, Perplexity and Claude (80 US buyer questions), 26 September 2026, whose title or address contains “best”, “top N”, “top rated” or “N best”; platform pages (YouTube, Reddit, social networks, Google, Amazon, Yelp, Wikipedia) excluded; 1,574 pages. Secondary sets: 20 questions repeated five times, three rewordings of 20 questions, and 40 questions in the US, UK, Canada and Australia.
- **Fetching:** plain HTTP, then a rendering service (Firecrawl) for every page that failed; 1,206 of the 1,574 pages fetched.
- **List entries:** the longest run of consecutive numbered H2 or H3 headings (“1.”, “2.”, “3.”), at least three entries.
- **Publisher:** the name part of the registrable domain, plus the site name stated on the page; names under four letters or generic words excluded.
- **Validation:** 299 pages coded by two AI models, split in half to choose and then test the rule; agreement, adjudication log and results in the data files.
- **Search comparison:** Google’s organic top 10 for the same keyword (AI Overviews, AI Mode) and for the same question (assistants); Bing results were not used.
- **Answer mentions:** whole-name match of the list’s #1 entry in the answer text, after removing citation labels.
- **Units of analysis:** the data nest as question → surface and run → answer → cited address → page and list → publisher → list entry. List shares use one row per cited list (clustered by publisher); per-surface shares use one row per list and surface; citation shares use citations and answer shares use answers (both clustered by question); the search comparison uses citations per question and surface, and lists within each question-and-surface pair; answer-level representation uses answer–list pairs, with the primary comparison within answers; stability uses question-and-surface combinations across runs.
- **Statistics:** 95% intervals by bootstrap with 2,000 resamples of publishers (headline shares) or questions (citations, answers), and by Wilson’s method (exact at zero) for per-surface shares; surface differences tested with a logistic model. The analysis was not pre-registered; decision rules were written down before the coding, and every deviation was logged.
- **Update schedule:** a second collection date, with repeated runs and reworded questions, is planned for October 2026.

## Limitations

- The detector misses about a third of real self-ranking lists, and 368 of the 1,574 pages could not be fetched, so the shares are likely undercounts.
- Only numbered lists are measured; unnumbered round-ups could not be read reliably, and the title filter misses some lists.
- The validation sample was coded by two AI models, not by people.
- Google’s top 10 shows what conventional search offered for the same query; it is not a measure of which pages the AI engines retrieved or could have used, which the study cannot see.
- Answer mentions show what answers say, not why they say it, and are only a proxy for representation. The #1 entry of a self-ranking list is a brand name the detector matched, while an independent list’s #1 is read less reliably, so measurement error may differ between the two groups; restricting to #1 entries that clearly read as names reverses the small difference.
- All data comes from one collection date; the repeated runs were minutes apart, and per-surface samples are small.

## Data and downloads

- Every cited list page with its fetch status, class, first entry and self-rank: [s13_list_pages.csv](https://underneath.agency/research-data/self-promoting-best-lists-study/s13_list_pages.csv) and [JSON](https://underneath.agency/research-data/self-promoting-best-lists-study/s13_list_pages.json)
- The 299-page gold standard with both coders’ labels, the codebook and the adjudication log: [gold_standard.csv](https://underneath.agency/research-data/self-promoting-best-lists-study/gold_standard.csv), [codebook](https://underneath.agency/research-data/self-promoting-best-lists-study/codebook.txt), [adjudication.csv](https://underneath.agency/research-data/self-promoting-best-lists-study/adjudication.csv)
- Every statistic on this page: [stats.json](https://underneath.agency/research-data/self-promoting-best-lists-study/stats.json)
- Machine-readable methodology: [methodology.json](https://underneath.agency/research-data/self-promoting-best-lists-study/methodology.json)
- The decision rules written before the coding, with every deviation: [analysis_plan.txt](https://underneath.agency/research-data/self-promoting-best-lists-study/analysis_plan.txt)

The data is free to reuse with attribution (CC BY 4.0). To cite: Underneath (2026), *How many “best of” lists cited by AI rank their own brand first?*, Underneath Research, https://underneath.agency/research/self-promoting-best-lists-study

## Frequently asked questions

### Do AI engines cite self-promotional “best of” lists?

Yes. Of 269 numbered “best X” lists with an identifiable publisher cited by AI Overviews, AI Mode, ChatGPT, Gemini, Perplexity and Claude, 24.2% were published by a company that ranked itself first. They made up 1.1% of all citations.

### Which AI engine cites self-ranking lists most?

We found no statistically detectable difference in this sample. Shares ranged from 5.9% for ChatGPT to 34.2% for AI Overviews, but the samples are small, so this does not mean the surfaces behave the same.

### Are self-ranking publishers named more often in answers that cite their list?

Not detectably in this sample. Within the 55 answers that cited both kinds of list, there was no statistically detectable difference in how often each list’s #1 entry was named (odds ratio 1.01, interval 0.52 to 2.75). Across all answer–list pairs, which are not independent, the rates were 47.1% (self-ranking) and 43.4% (independent). This is an observational comparison of answer-level representation, not a test of influence.

### How common is it for a list publisher to include itself?

26.0% of cited numbered lists with an identifiable publisher included the publisher somewhere, and when they did, 92.9% ranked it first.

## Related research

- Research · AI search

  ### [Do AI Overviews cite the pages that rank?](https://underneath.agency/research/ai-overview-citations-study)

  4,051 citations: only 28.7% rank in the top 10 for the same search, and YouTube appears in 64.2% of AI Overviews.

  Underneath Research
- Research · AI assistants

  ### [Do ChatGPT, Gemini, Perplexity and Claude agree on brands?](https://underneath.agency/research/ai-assistants-brand-agreement-study)

  80 buyer questions: all four assistants recommended the same first pick for 10.0% of questions.

  Underneath Research
- Research · AI assistants

  ### [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)

  About a quarter of the brands ChatGPT named appeared in all five answers to the same question, and one answer showed about half of them.

  Underneath Research
- Research · AI assistants

  ### [“Is this brand legit?” How AI assistants build a reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study)

  79 brands, four engines: every complete answer called the brand legitimate, and 99.7% then raised a problem.

  Underneath Research

## Related guides

- [Which pages should we target to show up in AI recommendations?](https://underneath.agency/resources/best-of-lists-ai-recommendations)
- [Do comparison pages help B2B brands get cited by AI?](https://underneath.agency/resources/do-comparison-pages-help-b2b-ai-citations)
- [Should ads in AI search be clearly labeled, and does labeling help the advertiser?](https://underneath.agency/resources/should-ai-search-ads-be-labeled)
- [Is AI search deciding which HR software gets the demo?](https://underneath.agency/resources/hr-software-ai-search)
- [Are AI search engines citing AI-generated content?](https://underneath.agency/resources/are-ai-search-engines-citing-ai-content)

---

This is the Markdown twin of https://underneath.agency/research/self-promoting-best-lists-study. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.



---
title: "Generative Engine Optimization (GEO) Services | Underneath"
description: "GEO services: get your business understood, cited and recommended in ChatGPT, AI Overviews, AI Mode, Gemini, Perplexity and Copilot answers."
canonical: "https://underneath.agency/services/generative-engine-optimization"
published: 2026-09-24
updated: 2026-09-27
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Get understood, cited and recommended in AI answers

# Generative Engine Optimization Services

Generative engine optimization is the work of making a business easier for AI search and answer systems to understand, retrieve, describe and reference when buyers are deciding what to buy, who to trust or where to go. It is not the same thing as writing content for AI.

We investigate the information underneath the answer: the questions buyers ask, the entities involved, the claims being made, the sources behind them, where that information is missing or contradicts itself, and what the competitors appearing instead have that you don’t. Then we change the parts that are holding you back and measure whether the change worked.

**Diagnose → Architect → Deploy → Compound.** That’s the Underneath Method.

Underneath runs GEO programs for local and multi-location businesses, ecommerce, SaaS and enterprise brands. Depending on the diagnosis, the work covers AI visibility analysis, content and information architecture, technical GEO, entity and source work, measurement and team enablement, on WordPress, Shopify, headless or custom stacks.

[Book a strategy call](https://underneath.agency/contact)[How we work](https://underneath.agency/approach)

Platforms & surfaces

- ChatGPT
- Google AI Overviews
- Google AI Mode
- Gemini
- Perplexity
- Microsoft Copilot
- Claude
- Bing

We measure four layers

- Visibility: what is being said
- Evidence: what supports it
- Behavior: what people did
- Business: whether it mattered

## What is generative engine optimization?

**Make the right information available to the systems that shape the answer.**

Generative engine optimization (GEO) is the practice of improving the information, technical foundation and external presence around a business so that AI-powered search and answer systems can discover, understand and reference it appropriately.

The work builds on SEO. Crawlability, useful content and clear information architecture all still matter, and structured data can still help systems understand what a page represents.

GEO adds a further question: given everything published about this business, on its own site and elsewhere, what will an answer system understand, retrieve and say? That is the layer we investigate.

[Google’s guidance on AI features](https://developers.google.com/search/docs/appearance/ai-features) says SEO best practices still apply to AI Overviews and AI Mode, and that no special optimizations are needed to appear in them. So we don’t treat GEO as a replacement for SEO. We treat it as a broader information problem that sits on top of a sound technical and search foundation.

## Why GEO matters

**A buyer can now get the shortlist before visiting a single website.**

A search results page gives a buyer ten links to investigate. An AI answer may give them a short list of options, already described and compared, with a handful of sources. That changes the visibility problem. Ranking for the question is no longer enough. The information around your business has to be strong enough that, when the question becomes an answer, you are still in it.

Our research looks at [how AI answers cite sources](https://underneath.agency/research/ai-overview-citations-study), [how often assistants agree](https://underneath.agency/research/ai-assistants-brand-agreement-study) and [how consistently recommendations hold up when the same question is asked again](https://underneath.agency/research/ai-recommendation-consistency-study). The underlying point: AI answers are not static rankings. They vary by question, platform, source set and time.

So we measure patterns, and never treat one screenshot as a result.

The Underneath Method

## We reverse-engineer the answer without pretending we can see *inside the model*.

For important customer questions, we work backwards from observable outputs.

1. **The question.** What are buyers actually asking?
2. **The answer.** Who appears, how are they described, and what is being recommended or compared?
3. **The evidence.** Which claims and sources support the answer?
4. **The information system.** How are the business, products, people, services and other entities represented across the web?
5. **The gap.** What is missing, inconsistent, outdated or stronger for a competitor?
6. **The intervention.** What should change?
7. **The measurement.** What would demonstrate that the change mattered?

We don’t publish every detail of how we do this. What a client needs is not the recipe but a stated reason for every change we make.

What’s included

## GEO starts with diagnosis, not a *content calendar*.

The exact work depends on what the diagnosis finds. A typical engagement can include the following.

- ### AI visibility and answer analysis

  We establish how the business currently appears in AI answers and search results for the questions that drive buying decisions.

  We examine more than whether the brand is mentioned. We look at descriptions, recommendations, competing entities, citations, source patterns and how those observations change across related questions.

  **Deliverable:** a baseline and visibility map tied to your priority buyer questions.
- ### Query and buyer-intent research

  We identify the questions behind the buying journey: category, comparison, problem, product, location and high-intent questions. We then examine how related formulations change the answer.

  We are not trying to collect thousands of prompts. We are looking for the question families that expose real visibility gaps.

  **Deliverable:** a prioritized buyer-question set.
- ### Answer and claim analysis

  We decompose important answers into the claims being made about the companies, products or services involved. Then we investigate the information supporting those claims. We look for:

  - Important claims
  - Supporting sources
  - Missing evidence
  - Conflicting information
  - Outdated information
  - Competitor advantages
  - First-party vs third-party descriptions

  **Deliverable:** a prioritized map of information gaps and opportunities.
- ### Content and information architecture

  We improve the pages the diagnosis points to, which can mean rewriting, restructuring, consolidating or, where something is genuinely missing, creating new pages. The aim is clearer information about the questions, entities and claims buyers care about, not more content. That can mean:

  - Stronger definitions
  - Clearer explanations
  - Question-to-answer structure
  - Internal relationships between topics
  - Useful comparison information
  - Better evidence
  - Less duplication

  **Deliverable:** priority pages changed for a defined reason.
- ### Technical GEO

  Important information has to be reachable before it can be understood. We investigate the technical conditions that get in the way. Depending on the site, this can include:

  - Crawl access
  - Rendering
  - Indexability
  - Canonicalization
  - Internal linking
  - URL architecture
  - Structured data
  - JavaScript-dependent content
  - Redirects and status codes
  - Robots directives
  - Bot and security controls

  One example: OpenAI [documents OAI-SearchBot](https://developers.openai.com/api/docs/bots) as the crawler behind ChatGPT search and recommends allowing it in robots.txt for sites that want to appear there. A security rule that blocks it can quietly remove a site from those answers. That is why we handle technical work as part of the visibility problem, not as a checklist of isolated SEO tasks.
- ### Entity understanding

  We investigate how the business is represented as an entity across its own properties and relevant external sources. The work can include:

  - Organization and brand relationships
  - Products and services
  - People and experts
  - Locations
  - Important attributes
  - Consistency across sources
  - Structured relationships
  - Author and expertise signals
  - External entity references

  We are not trying to manufacture a “knowledge graph.” We are making the basic facts about the business clearer and more consistent wherever they appear.
- ### Source and evidence strategy

  AI answers draw on what the rest of the web says about a company, not just its own site. We identify the sources behind the questions and claims we’re investigating, then work out where the business has:

  - Strong first-party evidence
  - Independent corroboration
  - Missing evidence
  - Conflicting information
  - Outdated information
  - Competitor gaps

  Where it is warranted, we help strengthen that footprint through legitimate source development: accurate listings, useful material other sites have reason to reference, and corrections where published information is wrong. We don’t manufacture citations.
- ### Monitoring and experimentation

  We track the priority questions continuously, but we don’t treat every movement as proof. We set a baseline, form a hypothesis, make the change and retest the relevant questions.

  Every finding is reported as **observed** (what the data directly shows), **correlated** (what moved alongside something else), **attributed** (what the available tracking connects to a business outcome) or **hypothesized** (a likely explanation we still need to test).

  AI systems change often. Without those labels, a report can’t tell a result from a coincidence.
- ### Team enablement

  We don’t want your team to depend on an agency forever. We document the principles behind the work and give your content, marketing and development teams what they need to keep it up. Depending on the engagement, that can include:

  - Content guidelines
  - Information architecture principles
  - Technical checklists
  - Entity and source guidelines
  - Measurement routines
  - Team workshops
  - Review processes

Platforms & surfaces

## We work across the answer experiences your buyers actually *use*.

Every program tracks ChatGPT, Claude, Gemini, Google AI Overviews, Google AI Mode, Perplexity, Microsoft Copilot and Bing. Which of them matter most depends on your market and how your customers research.

We don’t assume the platforms behave alike. We measure each one separately and look for the patterns that hold across them, so the program is never built around one platform’s behavior this month.

What we measure

## Visibility is not the same thing as business *impact*.

Our measurement model has four layers.

| Layer | The question | What we look at |
| --- | --- | --- |
| 01 · Visibility | What is being said? | Brand presence, recommendations, descriptions, competitive presence, citations, query coverage |
| 02 · Evidence | What supports it? | Claims, sources, corroboration, conflicts, freshness, information coverage |
| 03 · Behavior | What did people do? | AI referrals, engagement, calls, forms, applications, bookings, conversions |
| 04 · Business | Did it matter? | Qualified leads, pipeline, customers, revenue, acquisition cost, other agreed business outcomes |

We agree the relevant metrics before the engagement begins. We do not collapse everything into a single “AI score.” Every metric and how it is reported is set out in [how we measure](https://underneath.agency/results).

SEO, AEO and GEO

## Different surfaces. Overlapping *foundations*.

SEO, AEO and GEO share most of their technical and content fundamentals. What differs is the outcome being measured. The terms are defined in the [AI search glossary](https://underneath.agency/resources/glossary), and the agency-level comparison is in [GEO agency vs SEO agency](https://underneath.agency/resources/geo-agency-vs-seo-agency).

| | SEO | AEO | GEO |
| --- | --- | --- | --- |
| Primary outcome | Search visibility and clicks | Direct answers | AI-generated answers and recommendations |
| Typical surface | Search results | Featured snippets and direct answers | AI search and answer experiences |
| What matters | Pages, queries, links, technical foundation | Clear answers and structured information | Information, entities, evidence, sources and answer visibility |
| Measurement | Rankings, impressions, clicks, conversions | Answer visibility | Mentions, recommendations, citations, referrals and business outcomes |

A technically weak, poorly structured site is still a problem, and GEO won’t paper over it. But a technically excellent site can still have an information footprint that doesn’t produce the answers the business wants. That’s where GEO begins.

How a GEO engagement works

## Four stages. The last one doesn’t *end*.

1. Step 01 · Weeks 1 to 3

   ### Diagnose

   Establish the baseline. Investigate buyer questions, AI answers, competitors, sources, entities, claims, technical foundations and business measurement.
2. Step 02 · Weeks 2 to 4

   ### Architect

   Turn the findings into a prioritized plan. What needs to change? Where? Why? Who owns it? How will we know?
3. Step 03 · Month 2 onward

   ### Deploy

   Ship the work, on the website and in the sources beyond it.

   - Technical
   - Content
   - Entity
   - Source
   - Evidence
   - Measurement
4. Step 04 · Ongoing

   ### Compound

   Measure what the changes did, test the hypotheses and set the priorities for the next cycle. When the evidence changes, the plan changes with it.

The first 90 days

## From signature to shipped work, without a long *silence*.

| When | Milestone | What happens |
| --- | --- | --- |
| Day 0 | Agreement and access | Goals, markets, competitors, measurement and access agreed. |
| Days 1 to 7 | Audit delivered | Baseline across AI visibility, search, the website, sources, entities, customer language and measurement. |
| By day 30 | Plan agreed | Priorities, pages, sources, interventions, owners and measures. |
| Days 31 to 60 | First work live | Priority changes ship, with the baseline kept so they can be measured. |
| Day 90 | First full review | What changed, what didn’t, what we learned and what happens next. |

Who GEO is for

## GEO matters most when buyers are already asking *questions*.

We work with:

- **Local and multi-location businesses**, where buyers ask AI where to go, who to call or which provider to choose.
- **Ecommerce**, where buyers ask what to buy, which product fits their needs or how products compare.
- **SaaS**, where “best X for Y,” alternatives and category questions influence the shortlist.
- **Enterprise brands**, where long buying cycles, multiple stakeholders and a large information footprint make consistency and discoverability important.

What they share: customers research the decision before they contact you. If you are deciding whether to hire help, start with [Is a GEO agency worth it?](https://underneath.agency/resources/is-a-geo-agency-worth-it) and [How to choose a GEO agency](https://underneath.agency/resources/how-to-choose-a-geo-agency).

Pricing

## Scope determines the *program*.

Underneath’s GEO programs are priced in three bands: $5,000 to $10,000, $11,000 to $20,000, or $20,000+ a month. The band depends on the scope and complexity of the work, not on company size alone.

Every program begins with the Marketing & AI Visibility Audit. The audit can also be bought on its own, and if you continue into a program, its fee is credited toward it. Fixed-scope projects are priced to scope.

FAQ

## Questions about *GEO*

What is generative engine optimization?

It is the work of making sure AI assistants and AI search can find accurate information about your business, understand it and cite it when buyers ask the questions you should be answering: on your own site, in its technical setup and in the sources elsewhere that answers draw on.

Is GEO just SEO for AI?

No. GEO depends on SEO fundamentals, but the problem is broader. Beyond whether a page can rank, we look at how the business is represented across questions, sources, claims, entities and answer experiences.

Can you guarantee that an AI assistant will recommend us?

No. AI systems are dynamic and their answers can vary by question, user context, source availability and platform. We can establish a baseline, improve the information system around the business, measure observable changes and report them honestly.

Do you create lots of AI-generated content?

No. Volume is not the goal. We create, restructure or consolidate pages where the diagnosis shows a need, and often the answer is fewer, clearer pages.

Do you build citations?

Not in the link-building sense. We find the sources that answers draw on for your questions, correct what is wrong there and help create useful, credible information that others have reason to reference.

How long does GEO take?

There is no universal timeline. Changes to your own pages can show up once they are recrawled; work on external sources usually takes longer. We set the baseline first, and the first 90 days are about getting measurable work live against it.

How do you measure results?

In four layers: visibility, evidence, behavior and business. The exact metrics are agreed at the start of each engagement, and [how we measure](https://underneath.agency/results) explains each one.

Can you work with our existing SEO agency?

Yes. GEO usually sits alongside SEO, content, PR, development and in-house marketing. We work with the partners you already have rather than replacing them.

Can you work with any CMS?

Yes. We work on WordPress, Shopify, headless and custom stacks. The implementation changes with the technology; the method doesn’t.

What happens when AI platforms change?

We keep tracking the same questions and sources, so a platform change shows up in the data. When it does, we investigate what changed and adjust the work. Because the strategy isn’t built on one platform’s current behavior, a change on one platform rarely means starting over.

Free strategy call

## Start with the questions your customers are already asking.

On a free 30-minute strategy call, we look at the questions your buyers ask and then send a short written read: what we see, what we would investigate first, and whether we’re the right fit.

We don’t sell the audit on the call.

[Book your strategy call](https://underneath.agency/contact)[How we work](https://underneath.agency/approach)

---

This is the Markdown twin of https://underneath.agency/services/generative-engine-optimization. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.
