The short version
- For the same product question, ChatGPT’s app and API shared on average 12.0% of cited websites and 4.8% of exact pages (Uberti-Bona Marin and colleagues, 2026).
- The split was often total: the app and API shared no website at all in 60.9% of ChatGPT pairs and 43.4% of Gemini pairs.
- The differences had a pattern: YouTube and Reddit were 44.4 and 29.1 percentage points more likely to appear through Gemini’s API than its app.
- Citations are a small slice of what an engine finds: ChatGPT’s API retrieved an average of 37.38 pages per question but cited only 3.22.
- In our own API-based study, ChatGPT’s API reported 3.7 searches per answer, a useful detail, though the app may search differently (our hidden searches study).
Do APIs show the same sources as the ChatGPT and Gemini apps?
No. In the most direct test so far, the app and the API cited mostly different websites.
An API is the developer connection that software uses to send questions to an AI model without opening the chat app. It is attractive for tracking because it can be automated and lets you fix the model and settings. Uberti-Bona Marin and colleagues (opens in a new tab), at universities in Maastricht, Utrecht and Zurich, tested whether it gives the same picture as the app. They took 117 real product questions written by users and sent each one to the logged-out ChatGPT and Gemini apps and to their APIs, from the Netherlands.
Each app request was paired with an API request for the same question. Across 702 pairs, the median gap between the two was 0.17 seconds. The results still diverged sharply.
| Same question, app vs API | ChatGPT | Gemini |
|---|---|---|
| Cited websites shared, on average | 12.0% | 14.8% |
| Exact pages shared, on average | 4.8% | 11.9% |
| Pairs sharing no website at all | 60.9% | 43.4% |
Most of the disagreement was about which websites appeared, not which page on a website. Websites seen on only one side made up 79.8% of all ChatGPT pages observed across the two.
Why do the API and the app give different answers?
They are different setups, even under the same brand. The research has not pinned down exactly which difference matters most.
The logged-out ChatGPT app did not say which model it used, so the researchers picked the closest API model they could. Settings such as location also work differently: the OpenAI API accepts an approximate location, while Gemini’s API does not. The authors stress that they cannot attribute the gap to the access method alone.
Search behavior is part of it. When a tool calls an API, the tool’s settings decide whether and how often the model searches the web. One tracking platform let ChatGPT and Claude search the web only once a week per brand, to control cost. In between, they answered from memory (Kumar (opens in a new tab), who co-founded that platform). A tool built that way measures something different from the app a customer uses. Querying through the API, Sielinski (opens in a new tab) found OpenAI’s search model returned no citations for about 17% of questions on one topic.
APIs also expose different layers. ChatGPT’s API reported both the pages its searches returned and the pages it cited. On average it returned 37.38 pages per question and cited 3.22, and only 27.5% of returned websites made it into the citations. So a tool that reads search results is measuring a different thing from one that reads citations. Counts also vary widely by engine; see how many sources each engine cites.
Do the differences follow a pattern?
Yes. The API and the app leaned toward different kinds of sources, which can bias a tracking report.
For ChatGPT, the app favored technology publishers such as TechRadar and Tom’s Guide, while the API showed more manufacturer websites. For Gemini, YouTube was 44.4 percentage points and Reddit 29.1 points more likely to appear through the API. Gemini’s API also mixed far more types of sources: 69.9% of its answers combined three or more types, against 19.0% in the app.
The APIs also showed sources more often: by 9.7 percentage points for ChatGPT. A brand team reading API-based reports could therefore over- or underestimate how often review sites, forums or its own site reach real customers. Separate engines also cite largely different pages, as our guide to whether AI engines share sources shows.
Is API data useless for tracking AI visibility?
No. It is useful for controlled, repeatable tests, as long as you do not present it as what customers see.
APIs let you fix the model and settings, and some report steps such as the searches an engine ran. In our hidden searches study, ChatGPT’s API reported running a mean of 3.7 searches per answer. And none of the 509 searches logged across the assistants repeated the user’s question word for word. That shows how an engine looks for sources. We note in that study that these are API models, and the consumer apps may search differently.
API results were not reliably steadier either. Across repeated requests, ChatGPT’s API shared more cited websites between runs than the app did (37.9% against 26.0%), but fewer exact pages. Repeats matter on any channel: in the ChatGPT app, the number of different websites seen for a question rose from 2.65 after one request to 5.56 after three. We explain that built-in churn in why AI citations keep changing.
How do researchers capture what consumers actually see?
By automating the consumer apps themselves, logged out, and recording the page a user would see. It is harder but closer to reality.
An audit of AI-generated sources by Allaham and Diakopoulos (opens in a new tab) used the chat interfaces of ChatGPT, Copilot, Gemini and Perplexity for this reason, collecting 26,266 unique cited addresses. Schulte and colleagues (opens in a new tab) kept their ChatGPT analysis to one data source after finding that mixing API and app data would be inconsistent. A spurious image-delivery address had made up 5.8% of ChatGPT citations in part of their data. A small study discussed by Martinez (opens in a new tab) compared an interface and an API directly. Its 234 usable runs showed differences in which providers were named under some conditions, too few to generalize.
Our own studies mix both channels, and we label them. For example, our consistency study collected ChatGPT and Gemini answers from their consumer apps through a data provider, and Perplexity through its API.
What should you do about it?
Ask every vendor or team how the data is collected, and match the channel to the question. In practice:
- Ask whether the numbers come from the app or the API, for each engine, and which model and settings.
- Use app-based data for “what do customers see?” questions, such as share of voice and cited sources.
- Use API data for controlled experiments, where you need fixed settings and repeatable runs, and label it as such.
- Check whether web search was on for every API answer, and how often answers came back without sources.
- Never mix channels in one trend line. A switch from app to API can look like a sudden change in visibility.
- Spot-check the app by hand each month for your most important questions, and compare with the tool.
- Ask each question more than once on either channel, since one request shows only part of the sources.
If you want help setting up monitoring this way, see our generative engine optimization service.
What does the research not tell us yet?
Only one study has compared app and API sources head to head for ChatGPT and Gemini. Open questions:
- Why they differ. Model version, search settings and hidden system instructions were not separated.
- Other places and topics. The test used 117 general product questions, asked from one country, the Netherlands, at one point in time.
- Signed-in users. Logged-out apps may differ from what people see with an account and chat history.
- Brand mentions. The head-to-head comparison focused on cited sources, not on which brands the app and API recommended.
- Change over time. Both apps and APIs update often, so the size of the gap may move.
Frequently asked questions
Do ChatGPT API results match the ChatGPT app?
Not closely, for cited sources. In a 2026 test with 117 product questions, the app and API shared on average 12.0% of cited websites, and nothing at all in 60.9% of pairs.
Why do AI visibility tools use APIs?
Because APIs are easy to automate at scale and let you fix the model and settings, which makes runs comparable. The cost is that the results may not reflect the consumer app your customers use.
Is Gemini’s API closer to the Gemini app than ChatGPT’s?
Somewhat, but still far apart. Gemini’s app and API shared 14.8% of cited websites on average and no website at all in 43.4% of pairs.
Should we stop using API-based tracking?
No, but label it and do not mix it with app data. Use it for controlled tests, and confirm key findings against the consumer apps.
Sources
- Uberti-Bona Marin, L. G., Bertaglia, T., Astante, G., Rijsbosch, B., van Dijck, G., Hannák, A., Spanakis, G. and Kollnig, K. (2026), "If I Had to Buy Just ONE: Galaxy S26 Ultra": Auditing AI-Generated Product Recommendations (opens in a new tab), arXiv:2609.18729.
- Kumar, P. (2026), Generative Engine Optimization at Scale: Measuring Brand Visibility Across AI Search Engines (opens in a new tab), arXiv:2606.20065.
- Sielinski, R. (2026), From Stochastic to Stable: Rank Stability and Structural Sufficiency in AI Visibility Measurement (opens in a new tab), arXiv:2607.10341.
- Allaham, M. and Diakopoulos, N. (2026), Synthetic Sources?: Auditing Generative Search Engine Citations for Evidence of AI-Generated Sources (opens in a new tab), arXiv:2605.23684.
- Schulte, J., Bleeker, M. and Kaufmann, P. (2026), Don’t Measure Once: Measuring Visibility in AI Search (GEO) (opens in a new tab), arXiv:2604.07585.
- Martinez, O. (2026), Measuring GEO Visibility: Prompt Corpora Define the Answer Market (opens in a new tab), arXiv:2609.06811.
- Underneath (2026), The hidden searches AI assistants run before they answer
- Underneath (2026), Ask an AI the same question 5 times: do the brands change?