---
title: "Do ChatGPT, Gemini, Perplexity and Claude agree on brands?"
description: "80 buyer questions, four AI assistants: a third of picks shared, the same top pick 10% of the time. They mostly choose differently, not read differently."
canonical: "https://underneath.agency/research/ai-assistants-brand-agreement-study"
published: 2026-09-26
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Research · AI assistants

# Do ChatGPT, Gemini, Perplexity and Claude agree on brands?

Ask ChatGPT, Gemini, Perplexity and Claude the same buyer question, such as “What is the best CRM for a small business?” or “Who is the best divorce lawyer in Houston?”, and you get four different shortlists. We put 80 such questions, ten in each of eight industries, to all four assistants on 26 September 2026 and compared the 320 answers.

Version 1.0 of this study measured how much the assistants agree. This version asks where the disagreement comes from. An AI recommendation passes through several stages: the assistant finds pages, cites some of them, selects options from what it read, orders them and explains them. We measured the stages we can observe from the outside, fetched the 2,209 pages the assistants cited, and asked for each option one assistant recommended and another did not: did the other assistant’s own sources name it?

The answer changes the usual explanation. The assistants do cite different websites, but that barely predicts which brands they recommend. On national questions, most of the options one assistant recommended and another left out were named in the pages the second assistant cited. The main difference is what each assistant picks from its evidence, not which evidence it finds. Local questions are the exception: there, most one-sided picks appear in neither assistant’s cited pages.

## The short version

1. **A third of the picks are shared.** For the same question, two assistants’ recommended options overlapped by 0.327 on average (Jaccard, 95% interval 0.281 to 0.375). 66.3% of the options recommended for a question came from one assistant only; all four recommended the same first pick for 10.0% of questions (3.8% to 17.5%).
2. **Place names halve agreement.** Overlap was 0.160 for the 24 questions naming a place and 0.390 for the 56 national ones, a gap of 0.231 (0.166 to 0.294). B2B software had the highest overlap (0.543), home and local services the lowest (0.227).
3. **Different sources barely explain different picks.** Two assistants shared little of what they cited (mean cited-domain overlap 0.079), and 43.8% of question pairs shared no cited domain at all. Yet each 0.1 of extra domain overlap went with only about 0.02 more recommendation overlap, an association that was not distinguishable from zero in most models. Source overlap accounted for 0.9% of the local gap.
4. **Most one-sided picks are selection, not retrieval.** When one assistant recommended an option and another did not, the second assistant’s own cited pages named it 42.1% of the time (34.1% to 50.5%); only the recommender’s pages named it 28.8% of the time; neither did 29.1% of the time. On national questions the selection share was 56.5%; on local questions 54.2% of one-sided picks appeared in neither assistant’s cited pages.
5. **Even with the same evidence, they pick differently.** When both assistants’ cited pages named an option that at least one of them recommended, both recommended it 47.9% of the time (42.5% to 53.4%). Across all options it was 27.8%.
6. **It is not only run-to-run noise.** Across five repeat runs of 20 questions, an assistant’s answers overlapped with its own other runs (0.421 to 0.682) about twice as much as with another assistant’s (0.257 to 0.288). Of the brands one assistant named in at least three of five runs, 50.0% to 56.6% were never named by the other assistant in any of its five.

## Research questions

| Question | Answered here? |
|---|---|
| RQ1. How far do the assistants agree on the options they recommend? | Yes |
| RQ2. Does source overlap predict recommendation overlap, net of question, place, industry and pair? | Yes, as an association |
| RQ3. Did the assistant’s cited pages name the options it recommended, and the options the other assistant recommended? | Yes, for cited pages we could read |
| RQ4. How much of the difference is run-to-run variation? | Partly: 20 questions, 3 assistants |
| RQ5. What reasons for disagreement are visible in the answers? | Yes, model-coded |
| RQ6. With the evidence held constant, how much disagreement remains, and which content changes move which stage? | No: needs controlled experiments (see What comes next) |

## Where the answers diverge

We treat an AI recommendation as a pipeline and report each stage separately rather than folding them into one “AI visibility” number.

| Stage | Observed here? | Result |
|---|---|---|
| Question | Yes | Local questions agree less than half as much as national ones |
| Retrieval | No | The search results each assistant saw are not visible |
| Cited evidence | Yes | Low overlap of cited domains (0.079) |
| Selection | Yes, via cited pages | 42.1% of one-sided picks were named in the other assistant’s cited pages |
| Recommendation | Yes | Overlap 0.327; same first pick 10.0% |
| Justification | No | Whether the sources support each reason given was not checked |
| Stability | Partly | Within an assistant 0.421 to 0.682; across assistants 0.257 to 0.288 |

“Cited evidence” means the pages an assistant linked to. Assistants may read pages they do not cite, so everything we say about evidence is a lower bound on what each one saw.

## What we analyzed

**Questions and answers.** 80 buyer questions, ten per industry, from national product questions (“Which robot vacuum is best?”) to local service questions (“Who are the best roofing contractors in Dallas?”); 24 name a place (23 cities and one state). Each was asked once of each assistant on 26 September 2026.

- **ChatGPT:** the consumer web app, via DataForSEO’s LLM Scraper, US location.
- **Gemini:** the consumer web app, via the same scraper, US location.
- **Perplexity:** the sonar model through its API, with web search.
- **Claude:** Claude Haiku 4.5 through the API, with web search.

**Recommendations, not mentions.** Version 1.0 counted a brand when its exact name appeared anywhere in an answer. That mixes recommendations with passing mentions and splits one firm written two ways into two brands. This version uses a model coder (Claude Opus) that read all four answers to a question at once, under shuffled letters so it could not tell which assistant wrote which. It listed every option a buyer could choose, merged spelling variants, marked whether each answer recommended the option or only mentioned it, and recorded the order of the recommendations. The version 1.0 rule found 70.6% of the options the coder found; 95.6% of the rule’s brands matched a coded option.

**Evidence.** We fetched all 2,209 unique cited URLs on 28 September, two days after the answers, and extracted the main text. 64.2% were readable; the rest blocked automated requests, were not web pages or returned too little text. An answer’s evidence counts as usable when at least half of its cited pages were readable. An option counts as “in” an answer’s evidence when any of its coded names appears in that text.

**Stability.** Our [consistency study](https://underneath.agency/research/ai-recommendation-consistency-study) asked 20 of these questions five times each of ChatGPT, Gemini and Perplexity on the same day. We reuse those answers to compare an assistant with itself and with the others.

**Causes.** Two model coders (Claude Opus and Claude Sonnet), working independently and blind to which assistant wrote which answer, coded every question for the reasons the four answers differ, using a fixed list of nine codes.

## Study 1: how much the assistants agree

### Pairwise agreement on recommended options

| Pair | Overlap | 95% interval |
|---|---|---|
| Gemini and Perplexity | 0.355 | 0.305 to 0.404 |
| ChatGPT and Gemini | 0.346 | 0.283 to 0.406 |
| ChatGPT and Perplexity | 0.336 | 0.285 to 0.391 |
| Gemini and Claude | 0.321 | 0.262 to 0.380 |
| Perplexity and Claude | 0.321 | 0.270 to 0.372 |
| ChatGPT and Claude | 0.284 | 0.230 to 0.342 |

Overlap is the Jaccard similarity of two answers’ recommended options: options both recommended, divided by options either recommended. The intervals overlap heavily, but comparing pairs within the same questions, the pairs do differ (F test, p = 0.0098), driven by ChatGPT and Claude agreeing least.

Of the 1,257 option-and-question combinations, 66.3% were recommended by one assistant only (61.5% to 70.5%) and 8.2% by all four. All four put the same option first for 10.0% of questions, and at least three agreed on the first pick for 36.2%. Counting the first brand named rather than the first recommendation, as version 1.0 did, all four agreed for 11.2% of questions.

### By industry and by place

| Industry | Overlap | 95% interval |
|---|---|---|
| B2B software and technology | 0.543 | 0.462 to 0.637 |
| Hospitality and travel | 0.385 | 0.244 to 0.532 |
| Financial services and insurance | 0.332 | 0.227 to 0.455 |
| Healthcare and dental | 0.294 | 0.192 to 0.412 |
| Retail and ecommerce | 0.289 | 0.191 to 0.393 |
| Franchises and multi-location brands | 0.265 | 0.161 to 0.378 |
| Legal and professional services | 0.236 | 0.127 to 0.368 |
| Home and local services | 0.227 | 0.150 to 0.319 |

Each industry rests on ten questions, so most intervals overlap; only B2B software stands clearly apart. Industry differences are real as a group (F test, p = 0.0087), but much of them is about place: whether a question names a place explains 26.9% of the variation in question-level overlap, and industry and place together 45.4%.

Questions naming a place had a mean overlap of 0.160 (0.124 to 0.197); national questions 0.390 (0.338 to 0.445). In a model with a random intercept for each question and controls for industry, pair, source overlap, number of sources and list length, naming a place still lowered overlap by 0.232.

### Each assistant’s recommendation profile

| Assistant | Median picks | Share of all picks | Picks no other made |
|---|---|---|---|
| ChatGPT | 6 | 52.4% | 39.7% |
| Perplexity | 5 | 45.6% | 31.7% |
| Gemini | 5 | 42.0% | 28.6% |
| Claude | 5 | 41.6% | 36.7% |

“Share of all picks” is the share of the options recommended by any assistant for a question that this assistant recommended. ChatGPT recommended the most options and the most that no other assistant named.

## Study 2: do different sources explain different picks?

The assistants cite little in common. The mean overlap of cited registrable domains between two assistants for the same question was 0.079 (0.068 to 0.092), and 43.8% of pairs shared no cited domain. If each assistant simply recommended what its own sources list, pairs that share more sources should share more picks. They do, but only slightly.

| Model (397 question pairs) | Per 0.1 more domain overlap | 95% interval |
|---|---|---|
| Source overlap alone | +0.021 | −0.014 to +0.056 |
| Plus place, industry, pair, sources, list length | +0.020 | −0.011 to +0.052 |
| Within the same question | +0.025 | −0.004 to +0.055 |
| Question random intercept, with controls | +0.025 | +0.004 to +0.046 |

A bootstrap over questions gave −0.011 to +0.058 for the first slope. Pairs that shared no cited domain still shared 0.291 of their picks; pairs that shared at least one shared 0.355. Source overlap explains 0.6% of the variation on its own.

It also does not explain the local gap. Local questions had somewhat lower domain overlap (0.063 against 0.088), but adding domain overlap to the model moved the local effect from 0.233 to 0.231, 0.9% of the gap.

Cited-domain overlap is a coarse measure: two assistants can cite different sites that list the same firms. Study 3 therefore looks inside the pages.

## Study 3: evidence or selection?

### Were the recommended options in the assistant’s own sources?

| Assistant | Picks in own sources | Picked if listed | Picked if not |
|---|---|---|---|
| Perplexity | 89.1% | 47.7% | 12.6% |
| Claude | 81.2% | 53.2% | 11.7% |
| Gemini | 78.4% | 52.4% | 16.0% |
| ChatGPT | 52.0% | 54.2% | 36.9% |

Answers with usable evidence only. “Picks in own sources” is the share of the assistant’s recommendations that its own cited pages named. The last two columns use every option any assistant named for that question: of those the assistant’s cited pages listed, the share it recommended; of those they did not, the share it recommended.

Perplexity, Claude and Gemini mostly recommend options their cited pages name. ChatGPT is the outlier because of local questions: 26 of its 80 answers included a business list with addresses, and only 27.9% of its local picks appeared in its cited pages, against 75.9% on national questions. Its local picks likely come from a business-listing source it does not cite as a page. Once an option was in its evidence, each assistant recommended it roughly half the time (47.7% to 54.2%): all four select, rather than repeat, what they read.

### Why one assistant recommends an option and the other does not

For each pair of assistants and each option one recommended and the other did not, we checked where the option appeared (1,912 one-sided picks, both answers with usable evidence).

| Where the option was named | All questions | National | Local |
|---|---|---|---|
| In the other assistant’s cited pages (selection) | 42.1% | 56.5% | 19.4% |
| Only in the recommender’s cited pages (evidence) | 28.8% | 30.3% | 26.5% |
| In neither’s cited pages | 29.1% | 13.2% | 54.2% |

On national questions, more than half of the one-sided picks were in pages the other assistant had cited: it had the option in front of it and did not recommend it. On local questions, more than half came from outside either assistant’s cited evidence, which is consistent with local picks coming from business listings, maps or the model’s own knowledge rather than from the pages it links to.

Restricting to answers whose cited pages were all readable (158 one-sided picks) moved the split to 31.6% selection, 39.2% evidence and 29.1% neither. The selection share is sensitive to how much of the evidence can be read; that it is large is not.

A page naming an option is not the same as a page recommending it. The selection share counts options a page named for any reason, including in a comparison or as an also-ran.

## Study 4: is the disagreement just noise?

AI answers change from run to run, so part of the gap between two assistants would appear even between two runs of the same one. On the 20 questions asked five times of three assistants (brand names matched by the version 1.0 rule):

| Measure | ChatGPT | Gemini | Perplexity |
|---|---|---|---|
| Overlap with its own other runs | 0.530 | 0.421 | 0.682 |
| Same first pick as its own other runs | 54.7% | 53.9% | 63.6% |

Between two assistants, over every combination of their five runs, the figures were much lower.

| Pair | Overlap across all 25 run pairs | Same first pick |
|---|---|---|
| ChatGPT and Perplexity | 0.288 | 25.0% |
| Gemini and Perplexity | 0.268 | 16.4% |
| ChatGPT and Gemini | 0.257 | 25.0% |

Cross-assistant overlap was about half of what each assistant’s own run-to-run overlap would allow (0.478 to 0.544 of it). Comparing each assistant’s typical set, the brands it named in at least three of five runs, still gave overlaps of only 0.325 to 0.380. And of the brands one assistant named in at least three of five runs, 50.0% to 56.6% were never named by the other in any of its five. Run-to-run variation is large, but the assistants also differ systematically.

## Study 5: the reasons visible in the answers

The two coders agreed on the main reason for 58.8% of questions (Cohen’s kappa 0.433, moderate). We report a reason only where both coders gave it.

| Reason (both coders) | All questions | Local | National |
|---|---|---|---|
| Many valid options; each picks a different subset | 51.2% | 91.7% | 33.9% |
| The answers substantially agree | 33.8% | 8.3% | 44.6% |
| Each names what its own, different sources list | 25.0% | 54.2% | 12.5% |
| At least one answer names no options | 17.5% | 12.5% | 19.6% |

Among the 40 questions with the lowest overlap, both coders saw a crowded field of valid options in 82.5% and source-driven differences in 45.0%. The coders did not agree reliably on a different reading of the question (kappa 0.195) or on suspect entities (one coder flagged 15.0% of questions, the other none), so we do not report those codes as findings.

The coders’ view fits the numbers: in crowded local categories, each assistant draws a different short list from a long one, and the sources it cites differ too; in national categories with clear leaders, the assistants mostly agree and differ at the margins.

## Observed, inferred and unknown

- **Observed:** which options each answer recommended and in what order; which pages each answer cited; whether those pages, fetched two days later, named each option; how answers varied across five runs for 20 questions.
- **Inferred:** that most national disagreement is selection rather than retrieval rests on cited pages only, and pages the assistants read but did not cite could change it; that ChatGPT’s local picks come from a business listing is our reading of its answer format.
- **Unknown:** what each assistant retrieved but did not cite; whether the reasons the answers give for a pick are supported by the sources; whether any change to a brand’s website would change these outcomes.

## What this means

These are our interpretations, not additional findings.

- **Measure each assistant, and each stage, separately.** A brand can be recommended by one assistant and absent from another. Whether it appears in an assistant’s sources and whether that assistant then picks it are different problems.
- **Being in the sources is necessary but not enough.** For national categories, most one-sided picks were already in the other assistant’s evidence. Being listed on the pages an assistant reads raises the odds of a recommendation (roughly half of listed options were picked, against a small share of unlisted ones) but does not decide it.
- **Local visibility is a different problem.** Local picks often do not trace to any cited page. For local businesses, listings and profiles the assistants draw on without citing may matter more than articles.

## What comes next

This version works only with answers already collected. Three steps would turn these associations into causal estimates, and they are planned for the next versions:

1. Same evidence, different assistants. Give all four the same fixed set of pages for a question, in randomized order, and measure how much disagreement remains when retrieval is held constant.
2. Repeated runs for all four assistants. A stratified subset of questions asked several times on several days, with rewordings, so each option gets a recommendation probability instead of a single yes or no.
3. Content changes. Controlled changes to test pages, with retrieval, citation, recommendation and rank measured separately, to see which stage each change moves.

## How this compares with other studies

| Source | Sample and date | Figure |
|---|---|---|
| This study | 80 questions, 4 assistants, September 2026 | Recommendation overlap 0.327; all four agree on the first pick 10.0% |
| BrightEdge | ChatGPT, AI Overviews and AI Mode, August 2025 | 61.9% of queries get different brand recommendations across platforms; 17% the same brands on all three |
| BrightEdge | Five engines, category-level top-30 brand lists, May 2026 | Pairwise overlap of top-named brands between 36% and 55% |
| BrightEdge | ChatGPT and Gemini, June 2026 | Agree on about 2 of their top 5 brands |

BrightEdge compares category-level brand lists built from many prompts, which smooths out the variation of single answers; our figures compare the answers a buyer sees to one question, which may be why they sit at the low end of BrightEdge’s range. None of the BrightEdge comparisons includes Claude, publishes its prompts or examines the cited pages.

Sources: [BrightEdge, August 2025](https://www.brightedge.com/resources/weekly-ai-search-insights/chatgpt-vs-google-ai-62-brand-recommendation-disagreement); [BrightEdge, May 2026](https://www.brightedge.com/resources/weekly-ai-search-insights/where-ai-engines-agree-on-brands); [BrightEdge, June 2026](https://www.brightedge.com/resources/weekly-ai-search-insights/chatgpt-vs-gemini-same-question-different-brands).

## Methodology

- **Questions:** 80, ten per industry, written as a buyer would ask them; 24 name a place (23 cities, one state). Published in the dataset.
- **Collection:** 26 September 2026, one run per question per assistant, via DataForSEO (LLM Scraper for ChatGPT and Gemini with US location; LLM Responses API for Perplexity sonar and Claude Haiku 4.5 with web search). No new answers were collected for version 1.1.
- **Recommendations:** model-coded by Claude Opus from all four answers at once, engine hidden; options merged across spelling variants; recommended or only mentioned; rank of each recommendation. First pick = rank 1.
- **Agreement:** Jaccard similarity of recommended sets per question and pair, averaged over questions where both answers recommended something. 95% intervals by bootstrap over questions (2,000 resamples). F tests for pair (within questions) and industry.
- **Source models:** ordinary least squares with errors clustered by question, a question fixed-effects model and a question random-intercept model; domain overlap in steps of 0.1.
- **Evidence:** every cited URL fetched on 28 September 2026; readable = HTTP 200 and at least 400 characters of text that is not a block page; usable evidence = at least half of an answer’s cited pages readable; an option is in the evidence if any of its coded names appears in the page text.
- **Stability:** study 8’s five runs of 20 questions for ChatGPT, Gemini and Perplexity, brand names matched by the version 1.0 rule.
- **Reasons:** two model coders, blind to engine and to each other, nine fixed codes, agreement by Cohen’s kappa; codes reported where both agree.
- **Update schedule:** quarterly, same questions.

## Limitations

- One run per assistant for the main sample; stability covers 20 questions and three assistants, not Claude.
- Cited pages are not everything an assistant read, so the evidence and selection shares are estimates from what is visible. 35.8% of cited URLs could not be read, and pages were fetched two days after the answers.
- A page naming an option may name it negatively or in passing; name matching can miss variants.
- Associations only: nothing here changes the evidence an assistant sees, so none of it shows that a source caused a pick.
- All coding is by language models (two Claude models); no person coded this sample.
- Perplexity and Claude were queried through their APIs, which may differ from their consumer apps; Gemini’s scraped answers came from its Flash-Lite model.
- Questions were asked without a user location for the API assistants; city names in the question supply the location.

## What changed in version 1.1

- **The question.** From “do the assistants agree?” to “where do they diverge?”, with separate measures for cited evidence, selection, recommendation and stability.
- **The measure.** Recommended options, coded by a model from the full answers, replace brand names matched anywhere in the text. Headline overlap moves from 0.363 (mentions) to 0.327 (recommendations); the share named by one assistant only from 56.7% to 66.3%; all four agreeing on the first pick from 11.2% to 10.0%. Merging near-duplicate names under the version 1.0 rule would have raised its 0.363 to 0.377.
- **Uncertainty.** 95% intervals for every headline figure, resampling questions.
- **Sources.** Version 1.0 said the assistants “cited almost entirely different sources” and that this “may be one reason” for different picks. That is now stated as low cited-domain overlap (0.079) and tested: it explains little.
- **New measures.** Cited pages fetched and checked; repeated-run comparison; reasons for disagreement coded by two models.

## Data and downloads

- Every answer with its recommended options, first pick and cited domains: [s7_answers_v11.csv](https://underneath.agency/research-data/ai-assistants-brand-agreement-study/s7_answers_v11.csv)
- Coded options per answer (recommended or mentioned, rank): [s7_entities_coded.csv](https://underneath.agency/research-data/ai-assistants-brand-agreement-study/s7_entities_coded.csv)
- Question-by-pair table with recommendation and source overlap: [question_pairs_pipeline.csv](https://underneath.agency/research-data/ai-assistants-brand-agreement-study/question_pairs_pipeline.csv)
- Reasons for disagreement, both coders: [s7_disagreement_causes.csv](https://underneath.agency/research-data/ai-assistants-brand-agreement-study/s7_disagreement_causes.csv)
- Version 1.0 answers, brands and 80 questions: [s7_s8_answers.csv](https://underneath.agency/research-data/ai-assistants-brand-agreement-study/s7_s8_answers.csv)
- Every statistic on this page: [stats.json](https://underneath.agency/research-data/ai-assistants-brand-agreement-study/stats.json) (version 1.0: [stats_v1.0.json](https://underneath.agency/research-data/ai-assistants-brand-agreement-study/stats_v1.0.json))
- Machine-readable methodology: [methodology.json](https://underneath.agency/research-data/ai-assistants-brand-agreement-study/methodology.json)

The data is free to reuse with attribution (CC BY 4.0). Code: cite/pipeline/s7v11_*.py.

To cite: Underneath. (2026). *Do ChatGPT, Gemini, Perplexity and Claude agree on brands?* (Version 1.1). Underneath Research. https://underneath.agency/research/ai-assistants-brand-agreement-study

## Frequently asked questions

### Do ChatGPT and Gemini recommend the same brands?

Partly. For the same buyer question, their recommended options overlapped by 0.346 on a scale from 0 (nothing in common) to 1 (identical). All four assistants put the same option first for 10.0% of questions.

### Why do AI assistants recommend different brands?

Mostly because they choose differently from similar evidence, not because they read completely different sources. On national questions, 56.5% of the options one assistant recommended and another did not were named in the second assistant’s own cited pages. Local questions differ: most one-sided local picks appear in neither assistant’s cited pages.

### Do AI assistants agree more on some industries than others?

Yes. Overlap was highest for B2B software (0.543) and lowest for home and local services (0.227). Questions about a specific place had less than half the agreement of national questions.

### Is the disagreement just randomness?

No. Each assistant’s answers overlapped with its own repeat runs about twice as much as with another assistant’s, and half or more of the brands one assistant named in most runs never appeared in the other’s five runs.

## Related research

- [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)
- [ChatGPT local recommendations vs Google Maps](https://underneath.agency/research/chatgpt-local-recommendations-study)
- [AI Mode vs AI Overviews: how different are the sources?](https://underneath.agency/research/ai-mode-vs-ai-overviews-study)
- [Same question, four countries: do AI recommendations change?](https://underneath.agency/research/ai-recommendations-by-country-study)

## Related guides

- [Is tracking ChatGPT alone enough to measure our AI search visibility?](https://underneath.agency/resources/is-tracking-chatgpt-enough)
- [Do Gemini, GPT and Claude prefer the same kind of content?](https://underneath.agency/resources/do-ai-engines-prefer-same-content)
- [Do ChatGPT, Copilot, Google and Perplexity cite the same sources?](https://underneath.agency/resources/do-ai-search-engines-cite-the-same-sources)
- [Can we combine ChatGPT, Gemini and Perplexity into one AI visibility score?](https://underneath.agency/resources/combine-ai-engines-visibility-score)
- [How can an IT services firm win more projects and support calls from AI search?](https://underneath.agency/resources/it-service-firms-leads-ai-search)

---

This is the Markdown twin of https://underneath.agency/research/ai-assistants-brand-agreement-study. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.
