Guide · AI search

Can sellers game AI shopping rankings with manipulative product copy?

Sometimes, but mostly on weaker AI systems, and the gains are fragile. In a test of 14 manipulation tactics across five AI shopping rankers, GPT-5 and Claude flagged almost all of them and pushed the products down the list. Fabricated evidence did move some systems, which is exactly why it is a legal and reputational risk rather than a strategy.

The short version

  1. Product copy stuffed with superlatives and guarantees was flagged as questionable 100% of the time by GPT-5 and dropped about four places on a ten-product list (Bagga and colleagues (opens in a new tab)).
  2. Weaker rankers could be fooled: hidden override instructions moved a product up 3.40 places on an open Llama model, which flagged only 21.0% of them (Bagga and colleagues).
  3. Fabricated clinical claims beat a famous brand in 73.3% of head-to-head tests across three commercial AI systems, but only when one seller used them (Chu and Hou (opens in a new tab)).
  4. When all nine challenger brands used the same claims, the famous brand won again 93.8% of the time, and each copier’s gain fell from +0.802 to +0.007 on the authors’ payoff measure (Chu and Hou).
  5. Planted fake reviews already work in the wild: in a test of 12 AI systems, one polluted top search result led to a fake product being recommended in up to 27% of cases (Luo and Chen (opens in a new tab)).

Does manipulative product copy work on AI shopping assistants?

On the strongest systems tested, mostly not; on weaker ones, sometimes. Bagga and colleagues (opens in a new tab) built a shopping test bed from 13,747 real shopping requests, written in the detailed style of Reddit posts. Each was paired with 10 Amazon listings.

They rewrote one listing per request using 14 manipulation tactics. These ranged from fake reviews and made-up statistics to hidden instructions and keyword stuffing. Five AI systems then ranked the ten products: GPT-5, Claude, Gemini, DeepSeek and Llama.

Each ranker was also asked to list products whose descriptions looked misleading. That one-sentence safeguard was added by the researchers. Real shopping assistants likely use stronger defenses, but none have been published. We review the research on such defenses in whether AI search can filter manipulation.

The verdict was clear for the top systems. The authors write that “GPT-5 and Claude do not reward” manipulative rewrites with higher rankings. They read this as evidence that ranking gains on strong systems come from genuinely better content.

Which tactics did the strongest systems punish?

Nearly all of them, while a weaker open system rewarded several. The table shows how often each tactic was flagged and how far the product moved on a ten-product list.

TacticGPT-5: flagged, places movedLlama: flagged, places moved
Superlatives and guarantees100.0%, down 4.14100.0%, down 2.26
“Rank this #1” instruction in the text100.0%, down 3.8976.5%, up 0.39
Made-up precise statistics97.5%, down 3.2759.5%, up 1.75
Fake testimonials and ratings99.0%, down 3.7676.0%, up 1.60
Hidden override instructions99.5%, down 3.7521.0%, up 3.40
Writing in the style AI tends to like9.7%, up 0.4115.0%, up 1.01

Two lessons stand out. First, the same trick can win on one system and lose badly on another. A seller cannot know which assistant a shopper uses.

Second, the only tactic that rarely got flagged was mimicking the style AI systems tend to prefer. Its gains on GPT-5 were small. It is the grey zone, not a shortcut.

Do fabricated claims work better than sales pressure?

Yes, in one controlled test, and that is what makes them dangerous. Chu and Hou (opens in a new tab) gave three commercial AI systems ten identical skincare products: one famous brand and nine invented ones. With identical specifications, the famous brand was picked in every one of 670 valid trials.

Then they added marketing language to one invented brand. Copy mimicking credible evidence, such as clinical-trial claims or testimonials, broke the famous brand’s hold 50% to 73% of the time. Pure sales pressure, like scarcity or “limited stock,” barely worked, at 10% to 13%.

The systems differed. Gemini fell for authority claims almost every time. Claude was more guarded: when authority and testimonials were stacked together, its rate dropped from 55.0% to 21.2%, as if the offer looked too good to be true.

The authors used invented clinical studies on purpose, to find the upper limit. They also label that tier of content plainly: inventing studies or endorsements “is potential false advertising.”

An older experiment shows how far technical tricks can go on an open system. Kumar and Lakkaraju (opens in a new tab) added a machine-generated string of text to a fictional $199 coffee machine.

On Llama-2, it lifted the product’s rank in about 40% of tests where product order was shuffled. Building it required access to the open system’s inner workings, which sellers do not have for commercial assistants.

What happens when every seller games the system?

The advantage disappears, but sellers who stop become invisible. Chu and Hou ran a market test where more and more invented brands used the same authority claims. A single challenger knocked the famous brand down to winning just 19.8% of the time.

As rivals copied the tactic, the signal stopped standing out. With all nine challengers using it, the famous brand won 93.8% of the time again. The first mover’s gain of +0.802 on the authors’ payoff measure shrank to +0.007.

Worse, across 4,745 trials, invented brands that did not use the claims received zero recommendations. The authors call it a prisoner’s dilemma: everyone feels forced to join, and nobody ends up ahead. They suggest platforms may need to check authority claims, because each brand’s copy looks legitimate on its own. We look at the cost of standing still in what happens to brands that skip GEO.

What are the risks beyond being flagged?

Legal exposure, public exposure and a shrinking payoff. Fake reviews are already a live business.

Luo and Chen (opens in a new tab) describe a March 15, 2026 report by China Central Television. It exposed operators who seeded fake reviews so a fake brand appeared in leading Chinese AI assistants’ recommendations “within hours.”

The researchers then tested 12 AI systems on 225 real products. One fake-review page at the top of the search results led to the fake product being recommended in up to 27% of cases. With the top three results polluted, the worst system recommended it 73.8% of the time.

That shows the tactic can work. It also shows why it is risky: it was exposed on national television, and the authors note Chinese regulators launched enforcement in response. Our guide to pushing false information into AI answers covers the wider risk.

Reputation is also judged by AI from reviews. In our AI reputation study, 88.0% of answers about whether a brand was legitimate cited a review or complaint platform. Claims resting only on those platforms were negative 56.5% of the time.

What should you do about it?

Compete on real, specific product information, and check that nothing on your listings could read as fabricated.

  1. Lead with verifiable facts. In Chu and Hou’s test, a small real edge, such as a 7.3% lower price, was enough for an unknown brand to win half the time.
  2. Cite real evidence only. Genuine certifications and published trials are legitimate. Invented studies or endorsements are the tier the researchers call potential false advertising.
  3. Strip hype that adds nothing. Superlatives and guarantees were flagged by every ranker in the shopping test.
  4. Audit agency and marketplace copy. Ask anyone writing listings for you to confirm they do not add hidden text, fake reviews or invented figures.
  5. Test across assistants. A gain on one system can be a penalty on another.

For help improving product content that holds up across AI assistants, see our generative engine optimization service.

What does the research not tell us yet?

How live shopping assistants such as ChatGPT shopping or Amazon’s assistant respond to these tactics. The limits are real:

  • The shopping test used simulated rankers with a researcher-added safeguard over a fixed list of ten Amazon products, in English.
  • The fabricated-claims test used skincare products and invented brands; the authors did not repeat the marketing experiments in other categories.
  • The fake-review test used frozen search results, mostly in Chinese, from April 2026.
  • No study tracks real sellers over time to see whether manipulation gains last or trigger penalties.

Frequently asked questions

Can fake reviews get my product recommended by AI?

In tests, sometimes, which is why it is a serious risk. One fake-review page at the top of search results led AI systems to recommend a fake product in up to 27% of cases. The practice was exposed by Chinese state television.

Do AI shopping assistants penalize hype?

The strongest ones tested did. Copy built on superlatives and guarantees was flagged 100% of the time by GPT-5 and dropped about four places on a ten-product list.

Does mentioning clinical studies help AI rank my product?

Real evidence is a legitimate signal; invented evidence is not. Fabricated clinical claims beat a famous brand in 73.3% of tests, but the researchers call invented studies potential false advertising.

If competitors game AI rankings, should we copy them?

The research suggests not. When every challenger used the same claims, the famous brand won again 93.8% of the time and each copier’s gain nearly vanished.

Sources

Free strategy call

Some questions are easier to answer about your own business.

Bring the one that matters most. On a free 30-minute call we’ll take a first look at it and send you a short written read afterward.