Research · AI assistants

Same question, four countries: do AI recommendations change?

A brand that ChatGPT recommends in New York may not be the one it recommends in London, Toronto or Sydney. We asked ChatGPT and Gemini the same 40 buyer questions, with no country in the wording, five times each from the United States, United Kingdom, Canada and Australia on 28 September 2026, and asked them again with the country named. Repeating every question in every country lets us separate the country effect from the ordinary variation between two runs of the same question, and see which parts of the answer move: which brands appear, which comes first, and which sources are cited.

The short version

  1. The country changed ChatGPT’s recommendations by more than run-to-run variation. Two ChatGPT answers from the same country shared a mean 0.594 of their brands (Jaccard similarity); two answers from different countries shared 0.429. The gap, 0.164 (95% interval 0.114 to 0.221), held for 85.0% of questions.
  2. Gemini responded to the country far less. Its same-country and different-country figures were 0.492 and 0.436, a gap of 0.056 (0.029 to 0.086).
  3. The country effect grew with how much the question should depend on the country. On questions such as insurance, lending and tax, the gap was 0.180; on global products such as headphones and CRM software, 0.034.
  4. When the country changed the brands, it usually changed the sources too. About half (52.7%) of ChatGPT’s same-country advantage in brand overlap went with differences in the websites it cited. Answers from different countries that cited similar sources named similar brands.
  5. Location alone did part of what naming the country did. With only the location set, local-market brands made up 25.4% of the brands ChatGPT named in the UK; with “in the United Kingdom” added to the question, 49.1%. For Gemini the figures were 12.9% and 50.9%.
  6. The brand that comes first moved too. ChatGPT put the same brand first in all four countries for 25.4% of questions. Of the brands it put first in most runs in some country, 38.5% were never named in five runs from at least one other country.

Research questions

QuestionAnswered here?
RQ1. How much do brands change between countries?Yes
RQ2. Is that more than repeat-run variation?Yes, five runs per country
RQ3. Do ChatGPT and Gemini differ?Yes
RQ4. Do the cited sources change with the brands?Yes, as an association
RQ5. Which questions are most affected?Yes, by expected dependence
RQ6. Does the country change order, or only membership?Yes
RQ7. Does location alone act like naming the country?Yes

How to read “the country effect”

Whether a brand is recommended depends on the question, the engine, the date, the person asking and chance: two runs of the same question rarely match. The person asking includes their language, location, history and account. This study changes one of these, the country, holds the language at English, and keeps the rest as fixed as the data provider allows. Every comparison is made against the variation between repeated runs from one country, so what we report as the country effect is only the change beyond that variation.

We also measure the answer at several points, not as one score. These are presence (is a brand named), first place (which brand leads), rank-weighted overlap (do the leading brands match), the cited websites, and local-market share (how many named brands sell mainly in that country). These can move separately.

What we measured

The 40 questions are the national buyer questions from our four-assistant study, such as “What is the best CRM for a small business?” and “Which pet insurance company is the best?”. None names a country. On 28 September 2026 we asked ChatGPT and Gemini each question five times with the location set to each country (1,600 requests), and twice more with the question ending “in the United Kingdom?” (or the country in question) and the location set to that country (640 requests). The single run of 26 September 2026 (version 1.0 of this study) is kept as a pilot and as an earlier date.

Brands were identified as in the four-assistant study. We extracted candidate names from each answer, classified each as a brand or not, grouped name variants per question, and counted a brand where its name appears in the answer. Brand order is the order of first mention. Two answers are compared with Jaccard similarity (brands in both divided by brands in either, 1.0 is identical), whether they put the same brand first, and rank-biased overlap, which weights the top of the list. Cited websites are compared the same way.

Before looking at any answers, each question was coded for how much a good answer should depend on the country. That gave 17 high (insurance, lending, tax, legal and telehealth), 10 medium and 13 low (global products and software). Every brand was coded for its primary market (US, UK, Canada, Australia, or sold in most of the four). Both codings were done by Claude Opus and checked by Claude Sonnet, which agreed on 85.0% of questions and 90.8% of 2,080 brands. No labels are human-coded.

Answers to the same question are not independent, so every 95% interval on this page comes from resampling the 40 questions.

Findings

Country effect beyond ordinary variation

Two answers fromChatGPTGemini
The same country0.5940.492
Different countries0.4290.436
Gap0.1640.056
95% interval of gap0.114 to 0.2210.029 to 0.086

Mean Jaccard similarity of brand sets, 40 questions, 28 September 2026.

A second way to see it is to ask how much of the variation among the 20 answers to one question (five runs in each of four countries) lines up with the country. For ChatGPT the country accounted for 0.388 of it, for Gemini 0.234, against 0.158 expected by chance. In a permutation test the country grouping was significant (p < 0.05) for 72.5% of questions on ChatGPT and 37.5% on Gemini. The difference between the engines, 0.154 (0.100 to 0.209), is itself clear.

Same-country baseline and country pairs (ChatGPT)Jaccard
Two runs from the US0.658
Two runs from Canada0.607
Two runs from the UK0.566
Two runs from Australia0.544
US and Canada0.499
US and UK0.440
US and Australia0.425
Canada and Australia0.405
UK and Canada0.404
UK and Australia0.402

Every cross-country pair overlapped less than any pair of runs from one country. The US and Canada were the most alike.

Which brand comes first

MeasureChatGPTGemini
Pairs with the same first brand, same country0.6310.452
Pairs with the same first brand, different countries0.4390.390
Rank-weighted overlap, same country0.7310.619
Rank-weighted overlap, different countries0.5350.550
Same first brand in all four countries25.4%14.0%

Pair figures are shares of pairs of answers (1.0 means every pair); the last row is the share of questions, one run per country.

The country changed the order as well as the membership. For small business accounting software, ChatGPT put QuickBooks first in every US and Canadian run and Xero first in every UK and Australian run, while naming both everywhere. For air fryers it led with a Cosori model in the US and with Ninja in Australia, where the Cosori model did not appear. Gemini, located in Australia, led with the US Chase Sapphire Preferred card in four of five runs for travel rewards cards.

Where the country matters most

Expected country dependenceQuestionsGap, ChatGPTGap, Gemini
Low (global products, software)130.0530.015
Medium100.1570.022
High (insurance, lending, tax)170.2540.106

Gap = same-country minus different-country brand overlap.

The country’s share of variation rose by 0.087 (0.047 to 0.126) per step from low to high dependence. By industry it was highest for financial services and insurance (0.434) and lowest for business software (0.233). Even for low-dependence questions the country mattered on ChatGPT: the gap there, 0.053 (0.022 to 0.091), is small but above zero, and the country grouping was significant for 38.5% of low-dependence question and engine pairs.

Sources and brands move together

The websites cited changed with the country as well. Two ChatGPT answers from the same country shared 0.534 of their cited domains, answers from different countries 0.324. The country accounted for 0.430 of the variation in ChatGPT’s cited domains, as much as for its brands.

Two tests link the two changes, and both are associations, not proof of cause.

  • Pairs of answers. Among cross-country pairs, the quarter whose cited sources overlapped most named brands with a mean overlap of 0.712 for ChatGPT, higher than two same-country runs (0.594). The quarter that shared no sources overlapped by 0.152. Similar sources went with similar brands regardless of country.
  • Decomposition. Answers from the same country share more sources and more brands. Holding source overlap fixed, the same-country advantage in brand overlap fell from 0.163 to 0.077 for ChatGPT. So 52.7% (39.3% to 66.6%) of it went with the difference in sources, and the rest did not. For Gemini the share was 38.0% (26.8% to 53.9%).

The remaining part of the country effect, with sources held fixed, points to the model using the location when it writes the answer, not only when it searches. Answers were more likely to mention the country, its currency or its regulators when it was set: ChatGPT did so in 68.8% of UK answers, 69.2% of Canadian and 76.8% of Australian answers, though the question named no country. Gemini did so in 20.0%, 30.2% and 19.2%.

Local sources and local brands

Location onlyLocal domainsLocal brandsUS brands
ChatGPT, UK26.1%25.4%23.2%
ChatGPT, Canada21.8%17.6%34.7%
ChatGPT, Australia36.8%22.2%25.7%
Gemini, UK9.9%12.9%37.2%
Gemini, Canada11.4%19.6%33.0%
Gemini, Australia10.5%11.6%37.5%

Local domains = share of cited domains ending in .uk, .ca or .au. Local and US brands = share of named brands whose model-coded primary market is that country or the US.

Outside the US, ChatGPT answers that cited at least one local domain (44.2% of them) had a local-brand share of 45.0%, against 3.1% for answers citing none. Comparing answers to the same question, each additional share of local sources went with more local brands (slope 0.276, 95% interval 0.048 to 0.504). For Gemini, whose answers cited a local domain less often (20.2%), the link was stronger (slope 0.765).

Location alone vs naming the country

Local-brand shareLocation onlyCountry namedLocation’s share
ChatGPT, UK25.4%49.1%51.7%
ChatGPT, Canada17.6%44.3%39.7%
ChatGPT, Australia22.2%48.9%45.4%
Gemini, UK12.9%50.9%25.3%
Gemini, Canada19.6%50.0%39.2%
Gemini, Australia11.6%48.9%23.7%

Location’s share = location-only figure as a percentage of the country-named figure.

Naming the country moved the answers much further than setting the location. Two answers from different countries with the country named overlapped by only 0.205 (ChatGPT) and 0.173 (Gemini), against 0.429 and 0.436 with the location alone. Part of that is the rewording itself: adding “in the United States” to a US question shifted ChatGPT’s answer by 0.103, against 0.227 for the other three countries. For Gemini the US shift was 0.003, against 0.201 elsewhere.

A check across dates

The pilot run two days earlier overlapped with the new same-country runs by 0.539 for ChatGPT, against 0.594 between runs on the same day, a difference of 0.055 (0.014 to 0.103). For Gemini the difference was 0.025. Two days added a little change, much less than the country did.

How this compares with other studies

We found no other published study that compares brand recommendations for the same English questions across the US, UK, Canada and Australia with repeated runs. Profound’s analysis of 3.25 billion citations across 14 countries found that the language of a question shapes AI citations more than the country does. Our results show that, with the language held at English, the country still changes which brands and sources an answer uses, much more for ChatGPT than for Gemini. An arXiv study of brand preferences (ChoiceEval) found that the US-developed Gemini and GPT models show a marked preference for American entities. That is consistent with the US-market brands that still made up about a third of Gemini’s answers in the UK, Canada and Australia.

Sources: Profound (opens in a new tab); ChoiceEval, arXiv (opens in a new tab).

Observed, inferred and unknown

What we observe. For the same English question, ChatGPT’s recommended brands, first brand and cited websites change with the country more than they change between repeated runs, and more for questions about national, regulated products. Gemini’s change much less. Sources and brands change together. Location alone produces between about a quarter and a half of the local-brand share that naming the country produces.

What we infer. The country is part of what an AI recommendation depends on, like the engine and the date. A measurement taken in one country describes that country. Much of the country effect runs through the sources an engine finds for that country, and part appears to come from the model using the location when writing the answer.

What remains unknown. Whether changing the sources would change the brands (we observed sources, we did not alter them). How much a city, a device, an account or a user’s history adds. Whether the same holds in other languages, other engines, other question types and over longer periods. Why Gemini uses the location so much less.

What this means

The points below are our interpretation. They follow from the findings but were not tested.

  • Measure each market separately, with repeated runs. In this sample one run per country cannot tell a country effect from ordinary variation, and one country’s results did not describe another’s.
  • Expect the most difference where products are national. For insurance, lending, tax and similar questions, the brands a buyer sees depend heavily on where they are. For global products the difference is smaller but not zero on ChatGPT.
  • Local sources are part of local visibility. Answers that cited a country’s own websites named that country’s brands far more often. Coverage in each market’s review sites, comparison sites and publications is a reasonable place to start.
  • A named country is a different question. People who type the country get far more local answers than people who rely on location. Track both wordings where buyers use both.

Where this sits in GEO research

Research on generative engine optimization began by asking whether changing a page changes how often an AI answer uses it. Later work compared the sources that different AI search systems use across engines, languages and industries. It argued that visibility should be measured as a distribution over repeated runs and separate stages, not one rank, and that the person asking, including their location, is part of what it depends on. This study isolates one part of that, the country, with the language held constant. It measures the country effect against repeated runs and relates it to the sources cited. It is observational. It shows where and how much the country matters, not what would change a brand’s visibility in a given country.

Methodology

  • Questions: 40 national buyer questions from the four-assistant study, with no country in the wording (listed in the dataset with their codes).
  • Engines and locations: ChatGPT and Gemini consumer apps via DataForSEO LLM Scraper, English, location set to the United States, United Kingdom, Canada and Australia.
  • Design: 5 runs per question, engine and country on 28 September 2026 with no country in the question (1,600 requests); 2 runs with the country named in the question and the location set to it (640 requests); the 26 September 2026 run (320 answers) as the pilot. All 2,240 requests returned; 21 unfinished answers under 200 characters (18 ChatGPT, 3 Gemini) are excluded.
  • Brands: candidate names from bold text, headings and ChatGPT’s brand entities. Candidates were classified with the four-assistant study’s labels, and Claude Opus classified names not seen before. Variants were grouped per question and counted where the name appears; order is order of first mention. Two answers that both name no brand count as identical.
  • Pair measures: Jaccard similarity of brands and of cited registrable domains; same first brand; rank-biased overlap (p = 0.9). Same-country figures average the 10 pairs of runs in each country; different-country figures average the 150 cross-country pairs.
  • Country share of variation: the share of variation in Jaccard distances among the 20 answers to a question explained by country (PERMANOVA R²), with a 499-permutation test per question.
  • Sources and brands: brand overlap regressed on source overlap with question fixed effects, and the same-country advantage before and after holding source overlap fixed, with intervals from 300 question resamples.
  • Local sources and brands: local domains are those ending in .uk, .ca or .au; brand markets and question dependence are model-coded (Claude Opus, checked by Claude Sonnet).
  • Uncertainty: 95% bootstrap intervals resampling the 40 questions (2,000 resamples, seed 20260928).
  • Update schedule: quarterly.

Limitations

  • Country-level locations set by the data provider, not a local account, city or device; user history is not varied.
  • Two engines, one language, four countries and 40 questions, all asking for the best option in a category.
  • Five runs on one day, plus one earlier run; change over weeks is not measured.
  • The link between sources and brands is an association. We did not change the sources to test it.
  • Country-code domains are a rough measure of local sources (many local sites use .com). Brand markets, brand classification and question codes are model-coded, not checked by people.
  • A brand counts when it is named, including in a caution, so a mention is not always a recommendation.

What changed in version 1.1

Version 1.0 (26 September 2026) asked each question once per country and compared the country differences with five US runs from our consistency study, collected earlier and analyzed with a different brand list, on only 10 questions. Following an external review, version 1.1 (28 September 2026) is a new collection. It uses five runs in every country on the same day, a country-named condition, rank and source measures, questions coded for expected country dependence, a source decomposition and intervals clustered by question.

Two version 1.0 figures change. The matched comparison (0.617 same-country against 0.292 across countries for ChatGPT, 0.453 against 0.328 for Gemini) overstated the country effect, because the two sides came from different collections. The balanced design gives 0.594 against 0.429 for ChatGPT and 0.492 against 0.436 for Gemini: a smaller effect, still clear for ChatGPT. The share of brands named in all four countries (25.1% in version 1.0) was 24.9% for ChatGPT in the new runs, but four runs from one country share only 38.6%, so part of that figure is run-to-run variation. Version 1.0 also said that local sources feed local answers; version 1.1 tests that as an association and finds it holds, with the caveats above.

Data and downloads

The data is free to reuse with attribution (CC BY 4.0).

To cite: Underneath. (2026). Same question, four countries: do AI recommendations change? (Version 1.1). Underneath Research. https://underneath.agency/research/ai-recommendations-by-country-study

Frequently asked questions

Does ChatGPT recommend different brands in different countries?

Yes, beyond ordinary run-to-run variation. Two ChatGPT answers to the same question from one country shared a mean 0.594 of their brands; answers from different countries shared 0.429. The difference was largest for insurance, lending, tax and similar national products.

Is the difference between countries just random variation?

Not for ChatGPT. We asked each question five times in each country, so the variation between runs could be measured directly; the country added clear change beyond it for 85.0% of questions. For Gemini the country effect was real but small (a gap of 0.056).

Does ChatGPT use local websites for local answers?

Often, and the local sources go with local brands. In Australia 36.8% of the domains ChatGPT cited ended in .au. Outside the US, ChatGPT answers citing at least one local domain had a local-brand share of 45.0%, against 3.1% for answers citing none.

Is setting the location the same as naming the country in the question?

No. Naming the country roughly doubled the share of local-market brands in ChatGPT’s answers (for example 25.4% to 49.1% in the UK). For Gemini it rose from 12.9% to 50.9% in the UK and from 11.6% to 48.9% in Australia.

How should a brand in several countries measure its AI visibility?

Separately in each market, with repeated runs, and with the country both left out of and included in the question. One run per country cannot separate the country effect from ordinary variation.

Free strategy call

Want these numbers for your own category?

We run the same measurements for a business’s own buyer questions and competitors. On a free 30-minute call we’ll take a first look and send you a short written read afterward.