The short version
- 18.9% of the 636 answers (95% interval 15.1% to 22.6%) stated at least one fact that differed from the business’s Google profile; in 73.1% every fact matched.
- The engines differed, and the differences hold up when each business is compared with itself across engines. 4.4% of Gemini’s answers had a differing fact, against 16.4% for ChatGPT, 17.6% for Google AI Mode and 37.1% for Perplexity. Perplexity’s rate was 32.7 percentage points above Gemini’s. ChatGPT and Google AI Mode could not be told apart.
- Differing from the profile is not the same as being wrong. Of the 55 phone numbers that differed from the profile, 42 were on the business’s own website and 7 more on a page the same answer cited. Only 3 were found in neither place; for 3 more the website could not be checked. Counting a number on the business’s website as consistent, 97.9% of the phone numbers given matched one of the business’s own sources.
- Phone numbers differed mostly for businesses whose own sources disagree. Where the profile number did not appear on the business’s website, 30.6% of the numbers given differed from the profile. Where it did appear, 1.6% did.
- These are single answers from one day, one wording and one sampling frame (businesses already in the Google Maps top 20). They show how often each engine differed from the profile in this sample. They do not rank the engines’ reliability in general.
What this study measures, and what it does not
This study measures agreement with a reference record: does the fact an assistant states match the business’s Google Business Profile? The profile is a reasonable reference, because it is the record the business controls on the most-used local search surface. It is not proof of the truth. A profile can be out of date or show a call-tracking number, and a business can publish more than one valid number.
We therefore report three things, and keep them apart:
| Measure | Question | Measured here? |
|---|---|---|
| Agreement with the profile | Does the stated fact match the Google Business Profile? | Yes, for all four facts |
| Consistency with the business’s own sources | Does it match the profile or the business’s own website? | Yes, for phone numbers |
| Verified truth | Is it the business’s current, correct fact? | No: nothing was confirmed with the businesses |
Two further limits define the scope. We asked about businesses by name, so the study measures how an assistant describes a business it is asked about, not which businesses it chooses to recommend; our ChatGPT local recommendations study covers that. And matching a number on a website or cited page shows that the sources agree. It does not show where the assistant took the number from.
What we asked and how we scored it
We drew 160 independent businesses at random from the Google Maps top 20 for five services (dentist, plumber, personal injury lawyer, accountant, physiotherapist) in six cities in each of four countries, 8 per service per country. Choosing businesses from Google Maps rather than from any AI answer keeps the test fair to every engine, but it also means every business was already visible in Google local search. We left out chains present in more than one city and directory listings. One business was later excluded because our shortened version of its name was ambiguous, leaving 159.
Each engine was asked once: “What is the address, phone number, website and opening hours of {business} in {city}?”, using the name as the business presents it before any keyword suffix (for example “WT Law”, not “WT Law - Car Accident Lawyers Brisbane”). All 636 answers were collected between 07:52 and 08:14 UTC on 26 September 2026.
| Engine | Queried through | Model reported |
|---|---|---|
| ChatGPT | Consumer app | Not reported for 151; gpt-5-6 for 8 |
| Gemini | Consumer app | 3.5 Flash-Lite |
| Google AI Mode | Google Search, desktop | Not reported |
| Perplexity | Sonar API, web search on | sonar |
ChatGPT and Gemini were reached through DataForSEO’s LLM Scraper and Google AI Mode through its SERP API, all located in the business’s country. Perplexity was reached through DataForSEO’s LLM Responses API, which has no location setting.
Every answer was scored by fixed rules against the Google Business Profile as shown on Google Maps the same day:
- Phone: matches when the last nine digits match. A differing number was then searched for on the business’s own website (homepage and contact pages, fetched the same day) and on every page the answer cited.
- Address: matches when both the street number and the postcode match.
- Website: matches when the business’s own domain appears, and differs when the answer gives another domain as the website. (Version 1.0 could only score a website as matching or missing; see “Checking the scoring”.)
- Monday hours: match when the first opening time and last closing time match. Answers that said the hours were unclear or conflicting count as not given. 34 businesses whose profile says “open 24 hours”, usually a placeholder, were not scored for hours.
Findings
Agreement with the profile, by engine
| Engine | A fact differs (95% interval) | Every fact matches |
|---|---|---|
| Gemini | 4.4% (1.3% to 7.5%) | 94.3% |
| ChatGPT | 16.4% (10.7% to 22.6%) | 81.8% |
| Google AI Mode | 17.6% (11.9% to 23.9%) | 72.3% |
| Perplexity | 37.1% (29.6% to 44.6%) | 44.0% |
Share of each fact that matched the profile:
| Engine | Phone | Address | Website | Monday hours |
|---|---|---|---|---|
| Gemini | 99.4% | 98.7% | 99.4% | 94.4% |
| ChatGPT | 91.2% | 96.2% | 99.4% | 88.8% |
| Google AI Mode | 89.3% | 93.1% | 99.4% | 80.0% |
| Perplexity | 83.6% | 91.2% | 97.5% | 51.2% |
Shares are of all answers, so a fact that was not given counts as not matching. Hours columns cover the 125 businesses with regular hours on their profile; the other columns cover all 159. Intervals come from 4,000 resamples of the 159 businesses, each keeping its four answers.
Are the differences between engines real?
Every business was asked of all four engines, so the engines are compared business by business rather than as four separate samples. Taken together the four rates differ (Cochran’s Q = 71.9, p < 0.001). Pair by pair, with a Holm correction for the six comparisons:
| Pair | Gap, points (95% interval) | Only first / only second | Holm p |
|---|---|---|---|
| Perplexity vs Gemini | 32.7 points (24.5 to 40.2) | 55 / 3 | < 0.001 |
| Perplexity vs ChatGPT | 20.8 points (13.2 to 28.3) | 38 / 5 | < 0.001 |
| Perplexity vs Google AI Mode | 19.5 points (11.3 to 27.7) | 39 / 8 | < 0.001 |
| Google AI Mode vs Gemini | 13.2 points (7.5 to 19.5) | 24 / 3 | < 0.001 |
| ChatGPT vs Gemini | 11.9 points (5.7 to 18.9) | 24 / 5 | 0.001 |
| Google AI Mode vs ChatGPT | 1.3 points (−5.7 to 8.2) | 15 / 13 | 0.851 |
The gap is the difference in the share of answers with a differing fact. “Only first / only second” counts the businesses where only the first engine’s answer had a differing fact, and those where only the second engine’s did.
How each fact fared when an engine gave it
| Fact | Given and matching the profile | Given but different | Not given |
|---|---|---|---|
| Website | 98.9% | 1.1% | 0.0% |
| Address | 94.8% | 5.0% | 0.2% |
| Phone number | 90.9% | 8.6% | 0.5% |
| Monday opening hours | 78.6% | 9.4% | 12.0% |
Of the facts that were given, 98.9% of websites, 95.0% of addresses, 91.3% of phone numbers and 89.3% of Monday hours matched the profile. Perplexity declined to give Monday hours far more often (27.2% of its answers) and, when it gave them, matched the profile 70.3% of the time.
Where the differing phone numbers could be found
| The 55 phone numbers that differed from the Google profile | Answers |
|---|---|
| Published on the business’s own website | 42 |
| Not on the website, but on a page the same answer cited | 6 |
| Website could not be checked; on a page the answer cited | 1 |
| Website could not be checked; not on any readable cited page | 3 |
| Found on neither the website nor any cited page | 3 |
Phone numbers by engine:
| Engine | Matches profile | Profile or website | Found nowhere |
|---|---|---|---|
| Gemini | 99.4% | 100.0% | 0 |
| ChatGPT | 91.2% | 98.7% | 1 |
| Google AI Mode | 90.4% | 97.5% | 2 |
| Perplexity | 84.2% | 95.6% | 0 |
| All engines | 91.3% | 97.9% | 3 |
Shares are of the phone numbers given. “Profile or website” counts a number that matches the profile or appears on the business’s own website; “Found nowhere” counts answers whose number was on neither and on no readable cited page.
For 24.2% of the businesses we could check, the number on the Google profile did not appear anywhere on their own website, so the business publishes at least two numbers. Those businesses account for most of the differing numbers: 30.6% of the phone numbers given for them differed from the profile, against 1.6% for businesses whose website shows the profile number. Call-tracking numbers are one common reason a business shows a different number on Google, but our data does not show which number is the main line.
This does not tell us which source an assistant used. It tells us that when an assistant’s number differed from the profile, the same number was usually published somewhere the business or the answer’s own sources put it: of the 52 differing numbers whose answer cited at least one readable page, 47 appeared on a cited page.
What the answers cited
The answers contained 4,011 citations. Engines cited very different kinds of source:
| Engine | Any citation | Median | Own site | Other | |
|---|---|---|---|---|---|
| ChatGPT | 100.0% | 3 | 96.2% | 0.0% | 67.9% |
| Gemini | 86.2% | 1 | 5.0% | 79.9% | 3.1% |
| Google AI Mode | 99.4% | 2 | 67.3% | 73.0% | 22.6% |
| Perplexity | 100.0% | 19 | 96.2% | 1.3% | 100.0% |
Columns are the share of answers citing anything, the median number of citations, and the share citing the business’s own website, a Google page and any other site.
Gemini mostly cited Google pages, usually its Maps listing, and almost never the business’s website, and it had the highest agreement with the Google profile. Phone numbers given in answers that cited a Google page matched the profile 97.5% of the time, against 88.1% in answers that cited the business’s website and 86.6% in answers that cited another site. This is an association across engines, not a test: citing Google and being Gemini largely coincide, and a citation list does not show which page a number came from. We fetched 2,362 of the 3,230 distinct non-Google pages cited, on 28 September 2026, two days after the answers; a page may have changed in between.
By country and service
| Country | Businesses | A fact differs (95% interval) | Hours left out |
|---|---|---|---|
| Canada | 40 | 32.5% (22.7% to 42.4%) | 30.0% |
| United Kingdom | 40 | 15.6% (9.7% to 22.6%) | 5.6% |
| Australia | 39 | 13.5% (7.8% to 19.9%) | 6.4% |
| United States | 40 | 13.8% (7.5% to 20.3%) | 11.9% |
“Hours left out” is the same share when Monday hours are not scored.
| Service | Businesses | A fact differs (95% interval) |
|---|---|---|
| Plumber | 32 | 21.9% (12.9% to 31.9%) |
| Physiotherapist | 32 | 21.1% (13.0% to 29.5%) |
| Accountant | 31 | 17.7% (8.9% to 28.2%) |
| Dentist | 32 | 17.2% (9.5% to 26.0%) |
| Personal injury lawyer | 32 | 16.4% (9.1% to 24.3%) |
The countries and services contain different businesses, so we also fitted one model of whether an answer had a differing fact, with engine, country and service together and errors clustered by business. Canada’s higher rate remained after allowing for engine and service (odds ratio 3.48 against the United States, 95% interval 1.61 to 7.54; country overall p = 0.0021). The services did not differ (p = 0.847). In Canada the gap came mostly from Google AI Mode (45.0% of its Canadian answers had a differing fact) and ChatGPT (27.5%), and it remains when hours are left out. With about 40 businesses per country, this is an observed difference in this sample, not an explanation of it; we did not measure what distinguishes the Canadian businesses or their listings.
Whether some engines differ more in some countries could not be fully tested. Gemini had no differing answer in Australia and only one to three elsewhere, which makes its country terms unreliable. Among the other three engines, the model found no clear engine-by-country interaction (p = 0.092).
Monday hours: what was left out
The 34 businesses excluded from the hours comparison were not a random set. They were 24 of the 32 plumbers and 10 of the 32 personal injury lawyers, and no dentists, accountants or physiotherapists. By country they were 14 of 40 in the United States, 9 of 40 in Canada, 7 of 39 in Australia and 4 of 40 in the United Kingdom. The hours results therefore describe plumbers and lawyers poorly. Two checks show the headline does not depend on this choice: leaving hours out altogether, 13.5% of answers had a differing fact, with the same engine order (Gemini 2.5%, ChatGPT 11.9%, Google AI Mode 15.7%, Perplexity 23.9%). Scoring the 24-hour profiles as they stand gives 20.4% (Gemini 4.4%, ChatGPT 19.5%, Google AI Mode 18.9%, Perplexity 39.0%). The comparison covers Monday only. No profile in the sample had split Monday hours, and two were closed on Mondays.
ChatGPT: being listed and being described
Every business in this study came from the same Google Maps searches as our ChatGPT local recommendations study, which asked ChatGPT “Who is the best {service} in {city}?” four times per search on 26 and 27 September. That lets us set how often ChatGPT lists a business beside how well it describes it when asked by name:
| Listed by ChatGPT | Businesses | Every fact matched |
|---|---|---|
| In all four | 30 | 90.0% |
| In one to three | 55 | 85.5% |
| In none | 74 | 79.7% |
“Listed by ChatGPT” counts how many of its four “best {service}” answers named the business; “every fact matched” is from the separate question about the business by name.
53.5% of the businesses were listed at least once. Businesses ChatGPT listed were described slightly more often without a differing fact, but the difference is within chance for a sample this size (p = 0.283). Being recommended and being described consistently are separate questions, and a business can do well on one and poorly on the other.
Checking the scoring
We checked the scoring rules against a blind reading of the answers. We drew a stratified random sample of 136 scored decisions (38 phone, 36 address, 45 Monday hours and 17 website), oversampling facts the rules had scored as differing. Two independent AI coders (Claude Sonnet and Claude Opus) each read the answer and the profile value without seeing the rules’ label. Where the coders disagreed, we settled the label from the raw answer. No person coded the sample.
- The two coders agreed on 93.4% of the decisions (Cohen’s kappa 0.89).
- The check found one systematic gap. The 1.0 website rule could only score a website as matching or missing, and all 7 answers it counted as giving no website in fact named a different site, such as another clinic or a directory page. We corrected the rule and rescored all 636 answers. Only those 7 website scores changed, which moved the share of answers with a differing fact from 18.6% to 18.9%, Gemini’s from 3.8% to 4.4% and Perplexity’s from 36.5% to 37.1%.
- With the corrected rule, the rules and the settled labels agreed on 128 of 136 decisions (kappa 0.9). Of the 67 sampled facts the rules scored as differing, 64 were confirmed and 3 in fact matched the profile. The other 5 misreads were matching values the rules scored as not given.
The remaining rule errors are few and run in both directions; if anything, the differing rates above are slightly overstated.
How this compares with other studies
| Source | Sample and date | Figure |
|---|---|---|
| This study | 159 local businesses, 636 answers from 4 engines, September 2026 | 8.6% of answers gave a phone number different from the Google profile; 18.9% had any fact that differed from it |
| Seer Interactive | 178 phone number questions about large brands, 7 AI models, December 2025 | 36% of phone numbers were not on the brands’ customer service pages; Google profiles matched least often |
| Searchable (reported by Search Engine Journal) | 72,000+ questions about UK high street retailers, July 2026 | One in 16 answers incorrect; Perplexity 10%, Gemini 5%, ChatGPT 4% |
| Searchable (reported by Search Engine Journal) | 13,365 questions about 165 London businesses, July 2026 | 93% of businesses had at least one basic fact wrong or missing |
In both our sample and Searchable’s, Perplexity had the highest observed rate of the engines compared. The studies use different reference records: Seer compared numbers with brands’ own customer service pages, we compared them with Google profiles, and each finds that the two records often disagree for some businesses. Rates differ because the studies ask about different businesses and facts and measure against different references.
Sources: Seer Interactive (opens in a new tab); Search Engine Journal on Searchable (opens in a new tab); related: our ChatGPT local recommendations study.
What this means
This section sets out what we take from the results. It is interpretation, not measurement.
- Publish one set of facts everywhere. The phone numbers that differed from the profile were concentrated in businesses whose website and Google profile show different numbers. We cannot say which source the assistants read, but a single phone number, address and set of hours, identical on the website, the Google profile and directories, leaves an assistant nothing to choose between.
- Check each assistant separately. In this sample the engines differed widely, and one answer on one day is only a snapshot. Asking each assistant about your business from time to time shows what customers are told.
- Hours need the most care. Monday hours were the fact most often missing or different. Keep hours current everywhere, including on the website in plain text.
Methodology
- Businesses: 160 drawn at random (seed 20260927) from the Google Maps top 20 for “best {service} in {city}” across 120 searches collected for our ChatGPT local study; required phone, address, website and hours on the profile; multi-city chains and directory listings excluded; 1 excluded after collection (ambiguous shortened name), leaving 159.
- Engines and date: as in the table above; one question per business per engine on 26 September 2026, 07:52 to 08:14 UTC.
- Reference data: the business’s Google Business Profile fields as returned in Google Maps results on the same day; the business’s own homepage and contact pages, fetched the same day; every page each answer cited, fetched on 28 September 2026 (Google pages were not fetched).
- Scoring: the fixed rules described above, implemented in code with unit tests, and checked against a blind model-coded sample of 136 decisions (see “Checking the scoring”).
- Uncertainty: 95% intervals from 4,000 resamples of the 159 businesses, each keeping its four answers; engines compared with Cochran’s Q and exact McNemar tests on the same businesses, Holm-corrected; a logistic model with engine, country and service and business-clustered robust errors for the country and service comparisons.
- Version 1.1 (28 September 2026): the same 636 answers, reanalyzed after an independent methodological review. The corrected website rule changed 7 website scores and, with them, the headline rate (18.6% to 18.9%), Gemini’s rate (3.8% to 4.4%), Perplexity’s (36.5% to 37.1%), and the United States, dentist and personal injury lawyer rates. All other 1.0 figures are unchanged. See the changelog in the data folder.
- Update schedule: quarterly.
Limitations
- The Google Business Profile is the reference. A fact that differs from it is not necessarily false, and no fact was confirmed with the business.
- One answer per engine per business, from one wording, on one morning. Generative answers vary between runs and with wording, so the engine rates are a snapshot, not a stable ranking.
- Finding a number on the business’s website or a cited page shows that the sources agree. It does not show which source the engine used.
- Cited pages were fetched two days after the answers, and 868 of the 3,230 could not be read.
- The businesses were already in the Google Maps top 20; businesses with a weak Google presence may fare differently.
- Perplexity was tested through its API, which may differ from its app.
- Hours were compared for Monday only, and 34 businesses, mostly plumbers, were left out of the hours comparison.
- About 40 businesses per country and 32 per service: enough to detect large differences, not small ones.
- The scoring rules still misread a few answers: in the validation sample they scored 3 matching facts as differing and missed 5 matching values, so the differing rates may be slightly overstated. The validation was coded by AI models, not by people.
Data and downloads
- Every answer, scored field by field: s11_business_facts.csv and JSON
- Every phone answer with its source class and what the answer cited: s11_phone_provenance.csv
- Every citation, its source type and the phone numbers on the cited page: s11_citations.csv
- The validation sample with both coders’ labels: s11_validation.csv
- Every statistic on this page, with intervals and tests: stats.json
- Machine-readable methodology: methodology.json
The data is free to reuse with attribution (CC BY 4.0).
To cite: Underneath. (2026). Do AI answers match a business’s Google profile? (Version 1.1). Underneath Research. https://underneath.agency/research/ai-business-facts-accuracy-study
Frequently asked questions
Do AI assistants give the right phone number for a business?
In our test of 159 local businesses, 91.3% of the phone numbers the engines gave matched the business’s Google profile, and 97.9% matched either the profile or the business’s own website. Only 3 of the 633 numbers given could not be found on the profile, the website or any page the answer cited. We did not call the businesses, so this measures agreement with their published numbers, not whether each number reaches them.
Which AI assistant agrees most often with a business’s Google profile?
In this sample, Gemini: 4.4% of its answers had a fact that differed from the business’s Google profile, against 37.1% for Perplexity. Most of Gemini’s answers cited a Google page. This comes from one answer per business on one day, so it describes this sample rather than a lasting ranking.
Why do AI assistants give different opening hours?
Our data shows the answers, not the cause. Hours were the fact most often missing or different: 9.4% of answers gave Monday hours that differed from the Google profile, and 12.0% gave none. One plausible explanation, which this study does not test, is that businesses list different hours in different places.
How can a business keep AI assistants’ answers about it consistent?
Most phone numbers that differed from the profile were numbers the business publishes itself. Use the same phone number, address and hours on the website, the Google Business Profile and directories, and check what each assistant says about the business regularly.