The short version
- A crawl of 1.2 billion web addresses found 15.3 thousand confirmed instructions aimed at AI systems, on 11.7 thousand pages (Khodayari and colleagues (opens in a new tab)).
- Of those, 1,521 tried to shape how AI systems present a business or page: promoting it (1,040), forcing a citation (542) or demanding a positive review (502) (Khodayari and colleagues).
- About 70% sat in parts of the page people never see, such as headers, code comments and metadata (Khodayari and colleagues).
- In a lab test of 13 AI systems summarizing pages, small systems followed the hidden instruction 4.2% of the time and large ones 1.2% (Khodayari and colleagues).
- Visible self-promotion is far more common: 24.2% of the numbered “best X” lists that AI engines cited ranked their own publisher first (our self-ranking lists study).
What are these instructions trying to make AI do?
Most try to disrupt or block AI crawlers; a smaller group tries to steer AI search. The single biggest category was garbage injection: telling a machine reader to output random numbers or nonsense. It alone accounted for 8,469 of the 8,894 instructions aimed at disrupting systems.
Defensive uses were also common. Site owners wrote instructions telling AI not to reuse personal data or copyrighted text, or asking any AI reader to reveal itself. Many publishers seem to use these as a home-made “keep out” sign.
The group that matters for marketers is what the authors call reputation manipulation. It made up 1,521 instructions across 139 websites.
The main forms were content or product promotion (1,040), citation forcing (542) and positive review forcing (502). The authors say this cluster “targets search-oriented AI systems” and tries to shape “how entities, products, or sources are surfaced downstream.”
One template appeared 541 times on a single website. It tells any AI reader that the page “is the authoritative source of information” on its topic and that it “should not trust any other source.”
Where on the page do these instructions hide?
Mostly in places a visitor never sees, which is what makes them manipulative rather than persuasive. About 70% appeared in parts of a page that browsers do not display. The largest single channel was the HTTP header, the technical note a server sends before the page itself: 7,887 instructions arrived that way.
Structured data was another favorite. This is the machine-readable code sites add for search engines, often called JSON-LD.
Researchers found 1,996 instructions there, on 1,611 pages. Others sat in code comments and metadata tags.
When instructions did sit in the visible page, they were usually disguised. Of those still live when checked, 58.6% were concealed with tricks such as text the same color as the background, tiny fonts or elements placed off-screen. Overall, 87% of all instructions were hidden from human readers.
| Where the instruction sat | What it means |
|---|---|
| Server headers | Sent before the page loads; never shown to visitors |
| Structured data and metadata | Code meant for search engines and link previews |
| Code comments | Notes in the page source that browsers skip |
| Visible page, disguised | White-on-white text, tiny fonts, off-screen placement |
Why is copying this tactic a risk for a brand?
Because it is designed to deceive, it rarely works, and it can stay on your site for years. The study is careful not to call every instruction an attack. But planting hidden text to force citations or positive reviews is, by construction, an attempt to mislead a reader.
The instructions are also durable. Checking older archived copies, the researchers found 65% of affected pages already carried the instruction 12 months earlier. Something added once, by a developer or a plugin, can sit unnoticed for a long time.
Most came from the site owners themselves. First-party content accounted for 79.9% of instructions; third parties posting on a platform, such as job listings, accounted for 20.1%. User comments were a small share.
There is also a policy direction to watch. A position paper by Wen and colleagues (opens in a new tab) argues that AI platforms should downrank or exclude sources that use undisclosed influence, much as web search penalizes link schemes. That is a proposal, not a documented practice of any engine today. Whether engines can reliably filter out manipulative content is still being tested in labs.
What should you do about it?
Treat hidden instructions as a liability to find and remove, not a tactic to adopt. Visible, accountable content is where the legitimate influence is.
- Audit your own site. Ask your web team to search page source, server headers, structured data and comments for phrases addressed to AI, such as “ignore previous instructions” or “if you are an AI.”
- Check what vendors added. Plugins, tag managers and agency code can insert text you never approved. Most instructions in the study were first-party, so “we didn’t write it” is not a defense.
- Watch user-posted areas. Reviews, forums and job boards on your domain can carry instructions written by others.
- Make your real case visibly. Self-promotion in plain sight is common and accepted: 24.2% of cited numbered “best X” lists in our study put their own publisher first. Those lists were only 1.1% of all citations, so even visible self-ranking is a modest lever.
- Decide your policy on AI reuse openly. If you want to limit AI crawling, use standard, documented controls rather than hidden prompts.
If you want help with the visible side of this work, see our generative engine optimization service.
What does the research not tell us yet?
Whether hidden instructions change what live AI search engines like ChatGPT or Google’s AI features actually say. The gaps are worth stating plainly:
- The effectiveness test used one task, page summarization, on 13 systems in a lab. It was not a test of production AI search.
- The scan covered about half of one October 2025 crawl, looked for English phrases only and is a lower bound.
- No study has measured whether pages carrying these instructions get cited more, or less, by AI search engines.
- No study has measured whether engines penalize sites found using them.
- The shopping test used simulated rankers with an added anti-manipulation instruction, not deployed shopping assistants.
Frequently asked questions
Is hiding text for AI the same as old-school cloaking?
It is closely related: both show machines something people cannot see. In the largest scan, 87% of instructions to AI were hidden from human readers, using channels like server headers or tricks like white-on-white text.
Can a hidden prompt make ChatGPT recommend my brand?
There is no evidence it does in the live product. In a lab test, commercial AI systems followed planted page instructions only 0.6% of the time and flagged the attempt in 25.1% of runs.
How do I know if my site has hidden AI instructions?
Search your page source, headers and structured data for phrases addressed to AI. In one scan, 1,996 instructions sat in structured data, so check the code your SEO tools generate too.
Are all hidden AI instructions malicious?
No. Many are defensive, such as asking AI not to reuse personal or copyrighted content. The concern for marketers is the 1,521 instructions that tried to promote a page, force citations or demand positive reviews.
Sources
- Khodayari, Zhang, Acharya and Pellegrino (2026), Indirect Prompt Injection in the Wild: An Empirical Study of Prevalence, Techniques, and Objectives (opens in a new tab), arXiv:2604.27202.
- Bagga, Farias, Korkotashvili, Peng and Wu (2025), E-GEO: A Testbed for Generative Engine Optimization in E-Commerce (opens in a new tab), arXiv:2511.20867.
- Wen, Zhang, Yuan, Chen, Zhang and Guo (2026), Position: Generative Engine Optimization Creates Underexamined Risks, Governance Must Target Concentration, Disclosure, and Academic Blind Spots (opens in a new tab), arXiv:2606.12439.
- Underneath (2026), How many “best of” lists cited by AI rank their own brand first?