The short version
- In a test of 66 European brands in twelve languages, asking in a local champion’s home language raised how often AI named it by 0.80 on a 0 to 1 scale, against 0.15 for global brands (Żatuchin, 2026 (opens in a new tab)).
- In a second study by the same author, the language of the question accounted for 26.5% of the variation in the tone of a single AI answer about a brand, against 1.5% for the brand itself.
- A one-brand test in English and Japanese found two ChatGPT models swapped places: one led by 6.79 points in English, the other by 18.75 points in Japanese.
- Language is not the only blind spot: in our country study, ChatGPT answers from different English-speaking countries shared 0.429 of their brands, against 0.594 from the same country.
Why would the language of a question change what AI recommends?
AI assistants search the web in the language they are asked, and each language surfaces different sources and brands. Żatuchin (opens in a new tab) put the same buyer and reputation questions to GPT-5.4, Gemini 3.1 Pro and Perplexity Sonar Pro. The questions covered 66 brands from eleven Northern, Baltic and Central European markets, asked in twelve languages. The study collected 35,640 answers in April and May 2026, all from assistants that searched the web before answering.
The answers about one brand were not translations of each other. They differed in tone, in the sources behind them and, most of all, in which brands they named. The author notes the tone scores are noisy on long answers, so the direction matters more than the exact size.
Sources shifted at the margin. Wikipedia was the most-cited site in 11 of the 12 languages. In Lithuanian, the national business daily vz.lt edged it out, with 4.38% of citations. The pattern suggests local media matter most in smaller languages. Which engines switch to local sources is covered in our guide to local-language citations.
Which brands does an English-only audit miss?
It misses local champions: brands headquartered in, and strongest in, one market. The study split 65 brands into 24 local champions and 41 global or pan-European brands. It then compared answers to buyer questions such as “Who are the leading [industry] companies in [market]?” in English and in each brand’s home language.
Switching to the home language raised a local champion’s share of answers that named it by 0.80, on a scale where 1 means named every time. For global brands the rise was 0.15. A local champion was rarely named in English answers and named in almost every home-language answer. A global brand was named at similar rates in either language.
The author spells out the business risk. A team auditing only in English will conclude “the model ignores us” for a brand that is the default answer at home. The same audit will give a falsely reassuring picture of a multinational whose home-market coverage matters more.
Does language change how AI describes you, or only whether it names you?
Mostly whether it names you; tone moved far less than visibility. On tone, the home language added 6.2 points for local champions and 9.7 points for global brands on a 0 to 100 scale. That is a mild warming for both groups, with no reversal.
Tone is still not identical across languages. In the main study, 45% of brands showed a tone gap above 0.15 between their English and home-language answers. In a smaller set of 20 brands asked about in several local languages, 90% showed such a gap in at least one language.
So the description of your brand can drift by language, but the bigger commercial effect is on being named at all.
How large is the language effect next to other causes of variation?
In one detailed test, language outweighed every other factor except chance. A follow-up study by the same author (opens in a new tab) used 12,933 answers about 20 Central and Eastern European brands, in eight languages, from three AI models. It split the variation in each answer’s score into its causes.
The language of the question accounted for 26.5% of the variation in a single answer. The brand itself accounted for 1.5%. A brand’s standing against its rivals held roughly steady across AI models and across rewordings of the question, but not across languages. The part tied to a particular brand in a particular language was 8.6%.
The paper’s conclusion is blunt: a measurement that never varies the language is blind to the one factor that most reorders a brand’s score. One caution: the score measured was the tone of each answer, and 91.9% of answers scored as exactly neutral, so this result describes tone more than visibility.
Does the effect show up outside Europe?
Early evidence says yes, but it rests on a single brand. Kato and colleagues (opens in a new tab) asked two OpenAI models 56 product questions, half in English and half in Japanese. Each model answered every question 20 times, giving 2,240 answers in one collection window. The team tracked how often one brand, the web-highlighting tool Glasp, was named.
The ranking of the two models flipped by language. GPT-4o named the brand more often by 6.79 percentage points in English. GPT-5.6 Luna named it more often by 18.75 points in Japanese. A team testing only in English would have drawn the wrong conclusion about which assistant favored them in Japan. Chinese assistants raise a related question, covered in our guide to Chinese and Western AI models.
Location matters even within one language. Our country study kept every question in English and changed only the country. Two ChatGPT answers from the same country shared 0.594 of their brands; two from different countries shared 0.429. In the UK, local brands made up 25.4% of ChatGPT’s picks with the location set, and 49.1% when the question also named the country.
What should you do about it?
Audit in the languages and locations your buyers actually use, starting with the markets where you are strongest. In practice:
- List your priority markets and the language buyers there use for your category. Run your buyer questions in that language, not only in English.
- Set the location to each market as well. Our data show location changes answers even when the language stays the same.
- Translate the buyer’s need, not just the words. A survey of measurement methods (opens in a new tab) warns that a translated prompt may not carry over local availability, vocabulary or regulation. Have someone in each market review the prompts.
- Report results by language and market. Averaging a Lithuanian result into one global number hides the very gap you are looking for.
- Track more than one assistant. In the European study, how repeatable answers were depended far more on the assistant than on the language: Perplexity scored 0.904 on a 0 to 1 scale, against 0.952 for Gemini.
- If you are a multinational, treat a reassuring English result as unconfirmed until you have checked your home markets in their own languages. Our guide to GEO across languages covers the work that follows.
If you want help setting up tracking across languages and markets, see our generative engine optimization service.
What does the research not tell us yet?
The evidence is solid for Northern and Central Europe and thin almost everywhere else. The main gaps:
- The main study covers 66 brands in eleven European markets in one window in April and May 2026. Asian, Latin American and many other languages are untested, apart from one English and Japanese brand test.
- The author of both European studies is affiliated with Rankfor.AI, a company that sells AI brand monitoring, and discloses it. Brands were sorted into local and global groups by hand.
- These studies measure what AI says, not what buyers think or buy. No study yet links a language gap in AI answers to lost revenue.
- A small pilot that crossed language and location had 234 usable runs, and its control category did not repeat the effect. How language and location interact is still open.
- AI models change often. Every figure here is tied to specific model versions and dates.
Frequently asked questions
Do AI assistants answer differently in other languages?
Yes. In 35,640 answers across twelve European languages, the same brand was described and recommended differently depending on the language, and the largest difference was in whether local brands were named at all.
Is English-only AI visibility tracking enough for an international brand?
For a global brand it may be roughly fair; for a local champion it is not. In one large European study, asking in the home language raised a local champion’s naming rate by 0.80 on a 0 to 1 scale, against 0.15 for global brands.
Should we translate our AI tracking prompts?
Yes, but translate the buyer’s need rather than the literal words. Local products, rules and vocabulary can change what a good answer is, so a reviewer from each market should check the prompts.
Does the country setting matter if every question is in English?
Yes. In our study of the US, UK, Canada and Australia, ChatGPT answers from different countries shared 0.429 of their brands, against 0.594 for two answers from the same country.
Sources
- Żatuchin (2026), The Language Blind Spot: How Query Language and Brand Recognition Tier Shape AI-Constructed Brand Reputation Across Twelve European Languages (opens in a new tab), arXiv:2606.23165.
- Żatuchin (2026), Where Does the Noise Come From? A Variance-Components Decomposition of Non-Determinism in LLM Brand Answers (opens in a new tab), arXiv:2607.13304.
- Kato, Honma and Kato (2026), Generative Marketing Mix Modeling: A Causal Inference Framework Linking GEO and GEM to Business Impact (opens in a new tab), arXiv:2609.11915.
- Martinez (2026), Measuring GEO Visibility: Prompt Corpora Define the Answer Market (opens in a new tab), arXiv:2609.06811.
- Underneath (2026), Same question, four countries: do AI recommendations change?