The short version
- Across seven OpenAI releases, the share of an answer built from memory fell from 52% to 14%, in a 2026 test of 90 buyer questions by Finder and colleagues, who work for a company that sells website readiness scores for AI agents.
- In the same team’s main test of four AI agents, training knowledge supplied only 7–10% of the finished answer, whether or not the business’s website was easy for the agent to read.
- Some answers still skip the web: 8.8% of answers from ChatGPT, Copilot, Gemini and Perplexity carried no citation, rising to 11.6% for health questions (Allaham and Diakopoulos, 2026).
- Live search changes who gets found: Perplexity named new Product Hunt startups in 8.29% of discovery questions, against 3.32% for an OpenAI model with no web search (Sharma, 2026).
- In our own test, ChatGPT ran 3.7 searches per buyer question, and none of the answers that skipped searching cited a single page.
How much of an AI answer still comes from memory?
Very little, in the most recent test of AI agents answering buyer questions. Finder and colleagues (opens in a new tab) put 90 buyer questions about real businesses to successive OpenAI releases, with web tools available but optional. A separate AI model then judged how much of each answer came from memory rather than from pages the system retrieved.
The older OpenAI model built about half its answer from memory. Then the share fell steeply, from 52% in gpt-4.1 to 14% in gpt-5.6. In the team’s main experiment, 37,927 agent journeys about 1,056 businesses, training knowledge made up only 7–10% of the finished answer.
| Setup | Share of the answer built from memory |
|---|---|
| gpt-4.1, OpenAI’s last older-style flagship | 52% |
| gpt-5.6 | 14% |
| Four AI agent setups, second half of 2026 | 7–10% |
Read these figures with care. The authors work at ora, which sells the readiness score used in the study, and the memory share was judged by an AI model rather than by people. The release-by-release trend covers OpenAI models only.
How often do AI search engines skip the web entirely?
Sometimes, and the rate depends on the engine and the topic. Allaham and Diakopoulos (opens in a new tab) put 712 real user questions on politics, health and the environment to ChatGPT with search, Copilot, Gemini and Perplexity, through each product’s own interface. Of the 2,848 answers, 8.8% carried no citation at all.
Health questions were the most likely to get an answer with no sources: 11.6%, about one in nine. The figure was 8.2% for environment questions and 5.7% for politics. The authors read this as a sign that the engines still lean on what the model already knows for some questions, which raises doubts about how current those answers are.
Engines also differ sharply. In a mid-2025 study of 55,936 queries by Zhang and colleagues (opens in a new tab), Grok gave no cited website in 82% of its answers and Gemini in 38%. In our study of AI citations and Google rankings, Claude cited sources in 61 of 80 buyer answers and Gemini in 73, while ChatGPT and Perplexity cited sources in all 80.
Does live search change which brands an assistant can find?
Yes, and the difference is largest for new brands. A model without web access cannot know about a product launched after its training ended. Sharma (opens in a new tab) tested 112 startups from the 2025 Product Hunt leaderboard on two AI models in December 2025.
Asked about a product by name, the model without web search recognized it 99.4% of the time. Asked a discovery question, such as which new tools in a category are worth trying, it named the startup in only 3.32% of answers. Perplexity, which searches the web as it answers, did so in 8.29%.
Over the whole test, Perplexity surfaced 31 of the 112 products at least once, 27.7%, against 6 products (5.4%) for the model without search. One caveat matters: the “ChatGPT” in this study is a low-cost OpenAI model called directly by developers with no web search, not the consumer ChatGPT app with search turned on. The same startup test also looked at what predicts visibility in Perplexity.
What do assistants do instead of remembering?
They search, often several times, and they favor recent pages. In our hidden searches study, ChatGPT ran a mean of 3.7 searches per buyer question before answering, Gemini 1.9 and Claude 0.76. Every answer that ran a search cited at least one page, and none of the 42 answers without a search cited anything.
The pages they find tend to be new. In our freshness study, pages published in the last 90 days made up 17.4% to 22.6% of each assistant’s dated citations, against 6.9% of Google’s top 10 for the same questions.
When an agent cannot read a business’s own website, it does not fall back on memory either. Finder and colleagues found that web searches per journey rose from 1.8 on the easiest sites for an agent to read to 4.5 on the hardest. The gap was filled with other people’s pages, not with training knowledge. What that does to accuracy is covered in whether agents invent or omit business facts.
What should you do about it?
Put most of your effort into what AI assistants can read today, not into what they may have memorized. In practice:
- State current facts plainly on your own pages: prices, features, locations and how to get started. That is what an agent fetches when a buyer asks about you.
- Check that AI search crawlers and the fetchers that act for a user can reach those pages. Our crawler blocking study shows how often old robots.txt rules block them by accident. Your server logs show which pages those bots request, as reading AI bot server logs explains.
- Keep earning coverage on other sites. Assistants reach most pages through their own web searches, so what ranks and what others publish about you still shapes the answer.
- Test discovery questions, not just your brand name. Recognition by name, 99.4% in one test, says little about whether an assistant recommends you when a buyer asks for options.
- Watch health and other sensitive topics more closely, because answers with no sources were most common there.
If you want help turning this into a plan, our generative engine optimization service starts with these checks.
What does the research not tell us yet?
The direction is clear, but the evidence has real gaps.
- The sharpest evidence of memory’s decline comes from one vendor study. It covers OpenAI models only and relies on an AI judge.
- No study here measures how memory shapes which brands an assistant searches for in the first place. A search can still be steered by what the model already knows.
- The rates of uncited answers come from a few hundred to a few thousand questions, on specific topics, at one point in time.
- The startup test used a developer version of an OpenAI model with no search, so it does not describe consumer ChatGPT with search on.
- None of these studies changed a website and then tracked what assistants said about it over time.
Frequently asked questions
Does ChatGPT use its training data or search the web?
Both, but for buyer questions it now mostly searches. In one 2026 vendor test, memory supplied 14% of answers from gpt-5.6, down from 52% for gpt-4.1.
Can a brand launched after an AI’s training cutoff appear in its answers?
Yes, but mainly through live search. In a December 2025 test, a model without web search named new startups in 3.32% of discovery questions, against 8.29% for Perplexity with search.
Do AI search answers always cite their sources?
No. In a 2026 audit of four engines, 8.8% of answers had no citation, and in a separate mid-2025 study Grok gave no cited website in 82% of its answers.
Is it still worth trying to influence what AI models learned in training?
It now covers a small part of the answer. Finder and colleagues found training knowledge made up 7–10% of agent answers in their 2026 test, with most of the rest built from pages read at the moment of answering.
Sources
- Finder, Elovic, Shalev and Yosef (2026), AX is the New AEO (opens in a new tab), arXiv:2609.34951.
- Allaham and Diakopoulos (2026), Synthetic Sources?: Auditing Generative Search Engine Citations for Evidence of AI-Generated Sources (opens in a new tab), arXiv:2605.23684.
- Zhang et al. (2025), Source Coverage and Citation Bias in LLM-based vs. Traditional Search Engines (opens in a new tab), arXiv:2512.09483.
- Sharma (2026), The Discovery Gap: How Product Hunt Startups Vanish in LLM Organic Discovery Queries (opens in a new tab), arXiv:2601.00912.
- Underneath (2026), Do ChatGPT, Gemini, Perplexity and Claude cite pages that rank?
- Underneath (2026), The hidden searches AI assistants run before they answer
- Underneath (2026), How fresh are the pages AI engines cite?