Research · AI assistants

How faithfully do AI assistants quote software prices?

Buyers ask AI assistants what software costs, sometimes instead of opening the pricing page. On 26 September 2026 we asked ChatGPT, Gemini, Perplexity and Google AI Mode “How much does {product} cost? List each plan and its monthly price in US dollars” for 45 software and subscription products, and captured each official pricing page the same day. Version 1.0 of this study checked whether each quoted number appeared on that page. This version asks a different question: when an assistant turns a published price list into a sentence, what survives and what is lost? The number, the plan it belongs to, the billing term it depends on, the source it came from, and the plans left out are all measured separately.

The short version

  1. Of 965 dollar amounts in the answers, 840 were plan prices. The rest (13.0%) were add-ons, yearly totals, other products or ranges, which a page-matching check counts as errors.
  2. 61.9% of plan prices were fully faithful to the captured page: right amount, right plan, right billing term. Another 3.8% had the right amount but dropped a condition that changes what a buyer pays, usually by presenting an annual-billing price as the monthly price.
  3. 21.9% were not on the captured page but fit a variant the page itself offers (the other side of a monthly/annual toggle, a higher contact tier). These look real; we could not verify them from the capture.
  4. 7.6% differed from the page for the same plan and term, and 1.4% put a real price on the wrong plan. Where the size of the gap could be measured, 29 of 35 were underquotes, with a median gap of 13.3%.
  5. Most differing prices were not invented. For 39 of 64, the same figure appears on another page of the vendor’s own site, and for 21 more on a third-party page cited for the product. Only 4 were found on no source we could fetch.
  6. Once answers are clustered by product, the engines do not differ reliably (joint test p = 0.16), and a simple count of pricing-page complexity does not predict fidelity. Products differ far more than engines do.

Research questions

  • RQ1. What is a quoted price, once add-ons, totals and ranges are separated from plan prices?
  • RQ2. How often is a plan price fully faithful, and when it is not, what is lost: the amount, the plan, or the condition?
  • RQ3. When a price differs from the pricing page, where else does the same figure appear: the vendor’s other pages, a cited third-party page, or nowhere?
  • RQ4. How much of each price list do answers cover, and do they say which billing term the prices assume?
  • RQ5. Do engine, citing the vendor’s own pages, or pricing complexity explain fidelity once products are accounted for?

From “on the page” to price fidelity

Version 1.0 reported that 63.1% of quoted amounts appeared, to the cent, on the captured pricing page. That figure is unchanged and still in the data, but it measures string agreement with one page state, not accuracy. A price can be on the page and still be wrong for the plan the answer names, and a real price can be missing from the page because the page showed annual billing that day.

This version treats each quoted price as a claim with several parts (amount, plan, billing term, unit, conditions and source) and places it on a ladder:

LevelWhat the claim isShare of 965 amounts
4Fully faithful53.9%
3Right amount, condition missing3.3%
2Plausible variant, not on page19.1%
1Not supported by the page10.8%
0Not a plan price13.0%

Level 1 combines prices that differ from the page, real prices put on the wrong plan, and prices that could not be judged. The levels are kept separate in the analysis and are not combined into one score.

What we measured

We started with 60 software and subscription products with public US pricing pages. Each page was captured on 26 September 2026 with a rendering service; 45 pages loaded and showed at least three prices, and those 45 products are the sample. Each engine was asked once per product: ChatGPT and Gemini through their consumer apps, Google AI Mode through Google, Perplexity through its sonar API, all located in the United States. That gives 180 answers.

Every dollar amount above zero in an answer is one claim. Two AI models coded all 965 claims independently. Claude Opus coded each product with the four answers shuffled and the engine hidden, and Claude Sonnet coded the same material blind to the first. Each claim got a role (plan price, add-on, total, saving, other product, range) and, for plan prices, a verdict against the captured page. Every page each answer cited was fetched on 28 September 2026 to see where prices that were not on the pricing page could be found. Intervals and models account for the fact that the four answers about one product are not independent.

Findings

What the quoted prices turned out to be

Outcome for 840 plan pricesShare95% interval
Fully faithful61.9%53.7% to 69.7%
Plausible variant, not on page21.9%14.7% to 29.2%
Differs from the page7.6%3.5% to 13.2%
Right amount, condition missing3.8%1.6% to 6.3%
Cannot be judged3.3%0.7% to 6.8%
Right amount, wrong plan1.4%0.6% to 2.4%

The page-matching rule and the coding disagree in both directions. Of the 609 amounts found on the page, 65 were not plan prices, 28 dropped a condition and 11 were on the wrong plan. The coders judged 41 amounts that the rule could not find to be the page price after all, usually because the answer rendered the number differently, for example cutting off the last digit of Asana’s $10.99.

What gets lost: conditions, plans and billing terms

A dropped condition was almost always a billing term. In 17 of the 32 such cases the answer stated the price as a plain monthly price. Asana’s Starter plan was quoted at $10.99 a month, which is the annual-billing rate; billed monthly it was $13.49. Buffer’s Essentials plan was quoted at $5 a month without saying the page price assumes paying $60 a year.

Wrong-plan errors take a real number and attach it to the wrong row. One answer gave Semrush’s $117.33 annual-billing price for the SEO plan as the Pro plan price, and another gave NordVPN’s $14.99 Basic monthly price for the Complete plan.

Answers usually do say which billing term they assume: 95.0% mention monthly or annual billing somewhere. They rarely say when the prices are from. 30.0% add a caveat that prices change or point to the official page, ranging from 57.8% for ChatGPT to 2.2% for Perplexity.

When a price differs, where the figure comes from

We checked every price that was not on the captured pricing page against the other pages each engine cited for that product:

Amounts not on the pricing page (356)Count
On another vendor-owned page cited for the product138
On a third-party page cited for the product183
On no page we could fetch35

For the 64 prices coded as differing from the page, the split is 39 on another vendor page, 21 on a third-party page and 4 nowhere. The vendor pages are FAQ and knowledge-base articles, product pages, and blog or support posts about earlier price changes, for example Semrush’s plan FAQ, Zendesk’s article on its 2023 pricing update and Webflow’s post on its 2026 plan changes. An answer that quotes Semrush’s old Business plan at $499.95, or DocuSign Standard at $25 against the page’s $30, is repeating a figure the vendor’s own site still publishes. In these cases the vendor’s pages disagree with each other, and the assistant picked the other one.

A matching figure on a cited page could be coincidence, because comparison pages list many prices. As a check, we looked for each differing price on pages cited for a different product. It appeared there 20.0% of the time, against 90.0% on the product’s own cited pages. This shows the figures come from the cited material. It does not prove which page the engine used.

Where the size of a differing price could be measured against the same plan on the page (35 prices), 31 were within 5% to 20%, 3 were further off, 1 was within 5%, and 29 were below the page price. Old prices tend to be lower, which fits these being out-of-date figures.

How much of the price list answers cover

Answers named a price for 84.2% of the paid plans on the captured page on average, and gave the page price for 68.4%. 59.4% of answers priced every paid plan the page showed. The denominator is demanding, because it includes team, student and bundle editions where the page lists them. 44.4% of answers had every plan price fully faithful, and 16.1% had at least one price that differed or sat on the wrong plan.

Engines

EnginePlan pricesFaithfulRight amountDiffers
ChatGPT18869.7%74.5%6.9%
Google AI Mode19164.4%69.6%6.8%
Gemini26458.0%58.3%7.2%
Perplexity19757.4%63.5%9.6%

Faithful means fully faithful; right amount counts faithful prices plus right amounts with a condition missing. The engines also differ in how they source and qualify prices. Google AI Mode cited the official pricing page in 88.9% of answers, Perplexity in 77.8%, ChatGPT in 51.1% and Gemini in 11.1%. ChatGPT (57.8%) and AI Mode (55.6%) usually added a date or change caveat; Gemini (4.4%) and Perplexity (2.2%) rarely did.

The raw shares order the engines, but the ordering does not hold up as a finding. The 95% intervals overlap widely (ChatGPT 58.8% to 79.1%, Perplexity 45.7% to 68.4%). In a model clustered by product that also includes complexity and vendor citations, only Perplexity’s lower odds reach p < 0.05 on its own (odds ratio 0.61), and the joint test of all engine differences does not (p = 0.16). Gemini quotes the most prices and the most amounts that are not plan prices (17.8% of its amounts), and cites the vendor’s own site least often (22.2% of answers against 97.8% or more for the others).

Citations and complexity

Citing the vendor does not guarantee a faithful price. Plan prices in answers that cited a vendor-owned page were fully faithful 63.9% of the time, against 55.5% in answers citing only third-party pages. The adjusted odds ratio is 1.63, with a 95% interval of 0.7 to 3.83, so the difference is uncertain. 9.6% of prices in answers citing the vendor still differed from the page.

We scored each pricing page on six features that make a price harder to state: a monthly/annual toggle, per-seat pricing, usage or volume tiers, a contact-sales plan, an introductory price, and several products on one page. The score does not predict fidelity: 66.1% fully faithful on the simplest 18 pages, 47.4% on 11 medium pages, 67.8% on 16 of the most complex (adjusted odds ratio 0.90 per feature, 95% interval 0.69 to 1.19). One feature stands out descriptively. Pages with a billing toggle had 58.4% fully faithful prices against 72.9% without, which matches the billing-term errors above. The evidence for an engine-by-complexity interaction is weak (p = 0.056).

How far the two coders agreed

On the verdicts for all 965 amounts, the two model coders agreed 85.2% of the time (Cohen’s kappa 0.78). They agreed 95.9% on role and 85.9% on the six complexity features. The second coder was somewhat stricter: 57.5% fully faithful and 14.6% differing or on the wrong plan, against 61.9% and 9.0% from the primary coder. The headline shares depend on the coder by a few points; the pattern does not.

How this compares with other studies

VisibilityStack checked 13,867 queries about 108 software brands across six AI platforms between May and August 2026 and found AI “gets your whole price list right about one in four times” (24%). It also reported that “2 in 3 wrong prices are underquoted”. Its unit is the whole price list, graded against the vendor’s published price. Ours is the single price claim, graded against a same-day capture. Both find that underquoting dominates, which is what out-of-date prices would produce.

Source: VisibilityStack (opens in a new tab).

Observed, inferred and unknown

  • Observed: what each answer said, which pages it cited, what the captured pricing page showed, and where else each differing figure appears on fetchable pages.
  • Inferred: that prices coded as plausible variants are real (they fit a toggle or tier the page shows), and that differing figures found on older vendor or third-party pages are out of date rather than invented.
  • Unknown: which page an engine actually used for a given number, whether the same question asked again gives the same price, whether wording changes fidelity, and how long an old price persists after a vendor changes it.

What this means

These points are interpretation, drawn from the results but not measured.

  • Most price errors start with the vendor’s own web footprint. Retire or update old plan pages, help-center articles and announcements when prices change, not only the pricing page. An assistant that finds two official prices can quote either.
  • State both billing terms in plain text. A monthly/annual toggle hides half the price list from any single read of the page, and billing terms are where correct amounts turn into misleading claims.
  • Check what assistants say about your plans. Check the plan names too, not only the numbers. Outdated plan lineups (Semrush, Webflow and Zendesk tiers that no longer exist) were a common form of differing price.

Where this sits in GEO research

Research on generative engine optimization first asked whether a source can become visible in an AI answer. Later work asked which sources different engines select and how that varies with engine, wording and intent. More recent work argued that retrieval, citation, prominence and fidelity should be measured separately, with repeated runs and clustered uncertainty. This study takes the next step for one kind of fact. Once an engine has selected information about a product, how faithfully does it turn it into a claim a buyer will act on? Prices suit the question because each claim can be checked on several dimensions. The answer so far is that fidelity is lost more in billing terms and in stale figures on the web than in invented numbers.

Two further phases would test what this snapshot cannot, and both need new data collection. The first is repeatability and wording: 15 to 20 products, each engine asked five times with up to six controlled wordings (direct, explicit monthly, billing-aware, cite the official page, as of today, buyer framing). The second is staleness: pricing pages captured weekly, and after a detected price change the engines asked again at 1, 3, 7, 14 and 30 days.

Methodology

  • Products: 60 software and subscription products with public US pricing pages; 45 had a usable captured page (at least three prices other than free plans). Excluded products are listed in stats.json.
  • Reference: each official pricing page captured on 26 September 2026 with Firecrawl (rendered, US location).
  • Question and engines: “How much does {product} cost? List each plan and its monthly price in US dollars.” ChatGPT and Gemini consumer apps (DataForSEO LLM Scraper), Google AI Mode (DataForSEO SERP API) and Perplexity sonar (DataForSEO LLM Responses API), US, one run each, 180 answers.
  • Claims: every dollar amount above zero in an answer (965). The provider’s AI Mode text splits some numbers with spaces (“$3 5” for $35); these were rejoined.
  • Coding: each claim coded for role, plan, stated billing term, verdict against the captured page and the page price for the same plan, plus six pricing-page features and the plans on the page. Primary coder Claude Opus, second coder Claude Sonnet, both through the Claude Code command-line tool, one call per product with the answers shuffled and the engine hidden. This is model coding; no person coded the data.
  • Sources: all 932 cited URLs fetched over plain HTTP on 28 September 2026 (702 readable); dollar amounts on each page compared with the claims. Vendor-owned domains identified by the coder from the cited domains.
  • Statistics: 95% percentile intervals from 2,000 resamples of products. GEE logistic models of plan-price outcomes on engine, complexity score and vendor citation, with exchangeable correlation within product and Wald tests for engine and engine-by-complexity.
  • Code: cite/pipeline/s15_analyze.py (the 1.0 rule) and s15_v11.py (1.1 coding, source checks, analysis).
  • Update schedule: quarterly.

Limitations

  • The captured pricing page is the reference, and it shows one billing term, region and promotion. A price coded as a plausible variant or as not supported may still be real, and no price was confirmed with a vendor.
  • All coding is by two AI models, not people. Their agreement is reported; accuracy against a human standard is not.
  • One run per engine per product, on one day. Run-to-run stability, sensitivity to wording and staleness after a price change are not measured.
  • Cited pages were fetched two days after the answers, over plain HTTP. Pages that load prices with scripts or block fetching are missed, so the source counts are lower bounds. A figure on a cited page shows agreement, not which page the engine used.
  • Plan coverage counts every paid plan on the page, including team, student and bundle editions, so full coverage is a demanding standard.
  • 15 of 60 products were excluded because their page could not serve as a reference. Perplexity was queried through its API, the other engines through their consumer interfaces.

What changed in version 1.1

Version 1.0 (26 September 2026) reported the share of quoted amounts found on the captured pricing page, and estimated from a check of 25 mismatches that roughly 11.8% of quoted prices were wrong or out of date. Following an external methodological review, version 1.1 (28 September 2026) re-analyzes the same 180 answers with no new queries. The review found that page matching is not accuracy, that claims about one product are not independent, that an engine ranking should not be the main result, and that context and omissions should count as outcomes.

Version 1.1 codes all 965 amounts, replacing the 25-price check and its 11.8% estimate. It adds the fidelity ladder, billing-term and wrong-plan errors, severity and direction, plan coverage, a source check for every price not on the page, citation analysis, pricing complexity, product-clustered intervals and models, and a second coder. The 1.0 figures (63.1% of amounts on the page; ChatGPT 75.9%, Perplexity 66.7%, Google AI Mode 57.5%, Gemini 56.1%) remain in stats.json and the 1.0 dataset. The 1.0 hand check was done by the AI research assistant, not by a person.

Data and downloads

The data is free to reuse with attribution (CC BY 4.0).

To cite: Underneath. (2026). How faithfully do AI assistants quote software prices? (Version 1.1). Underneath Research. https://underneath.agency/research/ai-pricing-accuracy-study

Frequently asked questions

Can you trust the prices AI assistants give for software?

As a starting point, not as a quote. 61.9% of the plan prices four assistants gave for 45 products were fully faithful to the official pricing page. Another 21.9% were probably real prices for a billing term or tier the page did not show by default. 7.6% differed from the page, usually because an older figure is still published somewhere.

Which AI assistant is most accurate about prices?

None reliably, in this test. ChatGPT had the highest share of fully faithful plan prices (69.7%) and Perplexity the lowest (57.4%). Once the products were accounted for, the engine differences were not statistically reliable as a group, and the product being asked about mattered more than the engine.

Why do AI assistants quote wrong prices?

Mostly because the web gives them wrong or outdated figures, and because billing terms get lost. Of 64 prices that differed from the pricing page, 39 appear on another page of the vendor’s own site and 21 on a third-party page cited for the product. Many other errors were right amounts with the billing term dropped, such as an annual-billing price presented as the monthly price.

How can a company make sure AI quotes its prices correctly?

Update every page that states a price or plan name when prices change, including help-center articles and old plan pages. Show monthly and annual prices in plain text. Check what the assistants say about your plans, not only whether the numbers match.

Free strategy call

Want these numbers for your own category?

We run the same measurements for a business’s own buyer questions and competitors. On a free 30-minute call we’ll take a first look and send you a short written read afterward.