---
title: "Can You Trust an AI’s Reason for Recommending a Rival?"
description: "Only partly. In a 12-model hotel test, AI assistants’ stated reasons matched their real drivers imperfectly and left out list order and review count."
canonical: "https://underneath.agency/resources/ai-explanations-for-recommendations"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Can we trust an AI assistant’s explanation of why it recommended a competitor?

Only partly: an AI assistant’s stated reason is a rough guide to its choice, not a faithful account of it. In a large controlled test with synthetic hotels, the reasons models gave lined up with what actually drove their picks only imperfectly, and they almost never mentioned two factors they acted on.

## The short version

1. In a test of 12 AI models choosing between hotels, the factors named in the models’ reasons matched the factors that really drove their choices only partly, scoring 0.59 to 0.85 on a 0-to-1 scale ([Baig and colleagues, 2026](https://arxiv.org/abs/2606.16344)).
2. Where a hotel appeared in the list drove 4.1% of the choices, yet the models mentioned it in no more than 0.7% of their reasons (same study).
3. Brand affiliation appeared in up to 55% of one model’s reasons while accounting for only 2.0% of what moved its choices (same study).
4. When planted web pages fooled AI recommenders into backing a fake product, their answers contained 1.5 to 11 times more invented praise than answers that resisted ([Luo and Chen, 2026](https://arxiv.org/abs/2606.13610)).

## Why does it matter what an AI says about its own choice?

Because teams use the AI’s own words to diagnose lost visibility, and those words can point at the wrong cause.

When an assistant recommends a rival, it usually adds a short justification: better reviews, a stronger brand, a lower price. It is tempting to treat that sentence as the diagnosis and fix whatever it names. If the explanation leaves out what actually tipped the decision, the fix targets the wrong thing.

The research below separates two questions. What does the assistant say drove its choice? And what, measured by experiment, actually did?

## What did the hotel experiment test?

It randomly varied hotel details, then compared what moved each model’s picks with what each model said moved them.

[Baig and colleagues](https://arxiv.org/abs/2606.16344) showed 12 AI models, including OpenAI, Google and Anthropic models and four freely downloadable ones, sets of five made-up hotels. Rating, review count, review recency, price, chain status, an eco-label and list position were all shuffled at random. Each model picked one hotel and gave a one- or two-sentence reason. Across the whole project the team made 61,459 model calls.

Because every detail was assigned at random, the team could measure how much each one changed a hotel’s chances. [A top rating mattered most](https://underneath.agency/resources/why-ai-recommends-one-hotel-over-another): a 4.7-star hotel was picked 31.6 percentage points more often than a 3.9-star one. Then the team coded all 35,223 usable reasons to see which details the models named.

Two caveats matter. The hotels were synthetic, and two of the three authors work for a travel company, which the paper discloses. The test also covers only the final pick from a ready-made shortlist, not how an assistant finds candidates on the web.

## How closely did the stated reasons match the real drivers?

They matched in broad strokes but missed specific factors in a consistent way.

On a 0-to-1 scale, where 1 means the reasons ranked the factors exactly as the choices did, every model scored between 0.59 and 0.85. That is real agreement, but far from a faithful account. The gaps were not random:

| Factor | Share of what drove the choices | How often the reasons mentioned it |
|---|---|---|
| Position in the list | 4.1% | At most 0.7% of reasons |
| Number of reviews | 9.4% | Rarely |
| Chain or brand affiliation | 2.0% | Up to 55% of one model’s reasons |

The authors conclude that model-written explanations “cannot be relied upon, on their own” as an honest disclosure of what drove a decision. Put simply, the assistant talks about brand more than brand matters to it, and talks about list order and review count less than they matter.

## What do AI assistants act on without mentioning?

Mostly things outside the product itself, such as the order in which options were shown.

List position is the clearest case. In the hotel study, [being listed first](https://underneath.agency/resources/does-list-order-change-ai-recommendations) was worth $11.7 per night in the models’ choices, without changing anything about the hotel. One model, Gemini 2.0 Flash, showed a first-place advantage of about 26 percentage points. No hotel manager reading the model’s reason would learn this, because the reasons almost never mention order.

A second study shows a worse version of the same gap. [Luo and Chen](https://arxiv.org/abs/2606.13610) rewrote real web pages to promote fake products across 225 real products, then asked 12 AI models for recommendations, mostly in Chinese with an English repeat. A single planted page fooled models up to 27% of the time. When fooled, the models did not say “one page told me so.” Their answers contained 1.5 to 11 times more social-proof language, such as invented community praise, than answers that resisted. The stated reason was not just incomplete; it was made up after the fact.

Our own [prompt phrasing study](https://underneath.agency/research/ai-prompt-phrasing-study) shows why reasons sound convincing. When a budget changed which brands an assistant named, new brands were given a price or value reason 55.9% of the time. A reason was almost always supplied. Whether it caused the switch is a separate question the answer cannot settle.

## Do the sources an AI cites explain its answer either?

Not reliably: citations show what an assistant chose to display, not what it drew on most.

[Huang and colleagues](https://arxiv.org/abs/2603.16138) compared Google’s AI Overviews (the AI summary at the top of Google’s results) with the pages they cited. Overviews cited Reddit and Quora, yet drew 22.1% less of their content from those social sources than from other cited sources. The citation list suggested a breadth the summary did not use.

[Narayanan Venkit and colleagues](https://arxiv.org/abs/2410.22349) found the same pattern in 2024 answer engines. BingChat listed sources it never cited in its answer text, 36% of them on average. In their user study, 15 of 21 expert participants raised the lack of clarity about why some sources were chosen over others.

So neither the assistant’s sentence of reasoning nor its source list is a full record of why a rival won.

## What should you do about it?

Treat the AI’s explanation as one clue, then test the cause yourself.

1. **Log the reason, but do not act on it alone.** Record what assistants say about why a rival was chosen, then check it against evidence before changing anything.
2. **Test the factors the reasons tend to skip.** Ask the same question several times and in several orders or wordings. The hotel study shows order and review volume can matter without ever being mentioned.
3. **Fix what is measurable on your side.** [Ratings and price carried the most weight](https://underneath.agency/resources/ai-vs-human-reputation-priorities) in the hotel test. Visible facts like these are worth getting right whatever the assistant says.
4. **Watch for invented praise around rivals.** If an answer credits a competitor with community buzz you cannot find, the cause may be a planted or low-quality page rather than real reputation.
5. **Re-check after model updates.** The hotel weights describe specific model versions, and the authors warn they can drift.

If you want help running these checks across assistants, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The research shows reasons and real drivers diverge, but it has not mapped that gap for most business categories.

- The main evidence is one hotel experiment with synthetic listings, single-turn questions and English prompts. It may not carry over to software, finance or local services.
- It tested only the final pick from a fixed shortlist. How well assistants explain which candidates they found in the first place is unstudied.
- The authors call their reason-versus-choice comparison exploratory, and two of them are affiliated with a travel company.
- No study yet tests whether assistants explain choices more honestly when asked directly, or across a longer conversation.
- Models change often, so today’s gaps may shrink, grow or move.

## Frequently asked questions

### Why did ChatGPT recommend my competitor instead of me?

The answer itself will not reliably tell you. In a 12-model hotel experiment, stated reasons matched real drivers only partly, scoring 0.59 to 0.85 on a 0-to-1 scale, and skipped list order entirely.

### Can I ask an AI assistant to explain its recommendation?

You can, but treat the answer as a hypothesis. In the hotel study, brand was mentioned in up to 55% of one model’s reasons while driving only 2.0% of its choices.

### Does the order of options affect AI recommendations?

In controlled tests, yes. Being listed first was worth $11.7 per night in hotel choices, and one model showed a first-place advantage of about 26 percentage points.

### Are AI citations a reliable guide to why an answer says what it says?

Not on their own. Google’s AI Overviews drew 22.1% less content from cited social sources than from other cited pages, so a citation list overstates what was used.

## Sources

- Baig, Gillani and Ali (2026), [Whose hotel does the AI recommend? An algorithm audit of reputation signals in LLM-assisted hotel selection](https://arxiv.org/abs/2606.16344), arXiv:2606.16344.
- Luo and Chen (2026), [One Polluted Page Is Enough: Evaluating Web Content Pollution in LLM Recommenders](https://arxiv.org/abs/2606.13610), arXiv:2606.13610.
- Huang, Goyal, Saha and Chandrasekharan (2026), [Answer Bubbles: Information Exposure in AI-Mediated Search](https://arxiv.org/abs/2603.16138), arXiv:2603.16138.
- Narayanan Venkit and colleagues (2024), [Search Engines in an AI Era: The False Promise of Factual and Verifiable Source-Cited Responses](https://arxiv.org/abs/2410.22349), arXiv:2410.22349.
- Underneath (2026), [Does rewording a question change AI brand recommendations?](https://underneath.agency/research/ai-prompt-phrasing-study)

---

This is the Markdown twin of https://underneath.agency/resources/ai-explanations-for-recommendations. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.
