Gemini app vs Gemini API answers: how to investigate a disagreement
Gemini (app) and Gemini (API) can answer the same prompt differently. Align the question, date and collection context first, then compare answer evidence with a worksheet.
On this page
- In short
- Why would Gemini's app and API give different answers?
- What has to match before you compare two answers?
- What evidence can you compare in each answer?
- The discrepancy worksheet
- Worked example
- Common mistakes and what this cannot tell you
- How should you report the two Gemini engines?
- Frequently asked questions
- Next step
Gemini (app) and Gemini (API) are two different measurements of Google's Gemini, so their answers to the same prompt can legitimately differ. Gemini (app) is the answer a person who is not signed in sees on gemini.google.com. Gemini (API) is a direct call to Google's Gemini API with Google Search grounding. Before you treat a disagreement as a problem, align three things: the exact question and target, the date of the run, and the collection context. Then compare the evidence in each answer (brands named, sources cited, answer features) using one worksheet.
In short
- The two engines are separate in every filter and breakdown. A difference between them is a difference between two collection routes, not proof that one is wrong.
- Align the question, the date and the collection context before you compare any answer text.
- Compare evidence, not impressions: brands named, cited web pages, answer features and searches recorded.
- Some numbers are not comparable across the two, notably average citation position.
- Report the two as separate engines, and say what each one represents.
Why would Gemini's app and API give different answers?
The two engines reach Gemini by different routes, so identical prompts do not guarantee identical answers. This is documented behaviour: in the engines reference, DiscoveredBy states that the two are separate engines "in every filter and breakdown" and that their answers to the same prompt can differ.
Here is how each is collected:
- Gemini (app) is collected by DataForSEO, a licensed third-party data provider. It puts the prompt to gemini.google.com in a session that is not signed in and returns the answer the page shows. The platform labels this collection method "Licensed data".
- Gemini (API) is DiscoveredBy's own call to Google's Gemini API, with Google Search grounding. Grounding means the model is given Google Search results to base its answer on. The collection method is
api.
One concrete difference is the model. In every check DiscoveredBy ran on 2026-09-28, DataForSEO reported the app's model as "3.5 Flash-Lite", which is not the model Gemini (API) calls. Treat that as an observation from those checks, not a permanent fact: the app's model is chosen by Google and can change. DiscoveredBy stores the app's model as gemini-app so a change at Google does not split the engine's history.
The routes also differ in what they can see and send. The two engines are separate from Google AI Overviews and Google AI Mode too, so "Gemini said" in a report can mean up to four different things. The Gemini engine page lists both variants side by side.
What has to match before you compare two answers?
Align the question, the date and the collection context first; most apparent disagreements shrink or vanish once you do. Only after that is any remaining difference worth explaining.
1. The question and its target
A tracked prompt can have several targets: a country, sometimes a city, a persona and a language (see Prompts). Compare only runs of the same prompt with the same target.
Gemini (app) does not receive everything the API receives. Per the docs, the app gets the prompt as written on one line, the country as a setting (or the city's coordinates, with a 20-metre radius), and a language code. It gets no audience block, no response-language instruction and no "Search context" block, so persona variants do not run on it. A Chinese variant does not run, and a prompt over 2,000 characters once encoded does not run either.
So if you are comparing a persona variant, the app will have nothing to compare to. Run history says why, for example "personas are not sent to Gemini's app".
Location is delivered differently too. Gemini (API) is called with no location setting: the country and city reach it only as a "Search context" text block in the prompt. Gemini (app) receives them natively, as a setting of the request. Each answer records how its location was sent, and the prompt detail run history shows it, such as "Country in the prompt; city in the prompt" against "Country sent natively; city sent natively". A city with no coordinates on record runs at country level on the app, and the run says so.
2. The date
Compare runs from the same day, or the same window. An answer engine's response can change between two runs of one engine, so a Gemini (app) answer from Monday and a Gemini (API) answer from Friday are not evidence of anything on their own. Use the filter bar's window (7, 28 or 90 days ending yesterday) so both engines are read over the same period, and look at rates over many runs rather than one pair of answers.
A failed run is not an answer. When DataForSEO cannot read a Gemini (app) answer, the run is recorded as failed, and the day's failed runs are retried automatically at 03:00, 05:00, 09:00 and 17:00 UTC. A run still failed after 17:00 stays failed for that day. If only one engine has an answer on a given day, check for a failed run before you look for a difference in content.
3. The collection context
Confirm you are comparing the two channels, not a blend. In the Explorer, engine is a dimension and each collection channel is its own value, so break down by engine or filter to one. Do not pool them.
Also note how each engine gets to searching. Gemini (API) is called with Google Search grounding. On the app, Gemini decides for itself whether to search, because DiscoveredBy never asks it to, so some app answers list no sources at all.
What evidence can you compare in each answer?
Compare the observable evidence: which brands are named, which web pages are cited, which answer features appear, and what the engine reports about its searches. Each is an observation of one answer, not a window into the model.
| Evidence | Gemini (API) | Gemini (app) |
|---|---|---|
| Brand mentions | Extracted from the answer text | Extracted from the answer text, with bare-domain source links blanked out first |
| Citations | Grounding sources, attributed to the domain the Google redirect lands on | Web pages among the answer's sources, one per page, in order of first appearance |
| Citation rank | Order the engine reported its sources | Order the answer first cites them |
| Google product and Maps place links | Not counted as a separate type | Counted as shopping and local businesses, not citations |
| Searches behind the answer | The searches Google reports | Not reported; fan-out reads "not applicable" |
| Retrieved vs cited channel | Grounding sources only | Cited sources only |
| Answer features counted | Web search | Web search, shopping, local businesses, images, tables |
Three details in that table cause most false alarms.
Citations are built differently. Gemini (API) grounding links are opaque Google redirect URLs, so DiscoveredBy follows the redirect and attributes the citation to the domain it lands on. Gemini (app) citations are the web pages among the answer's sources, and a page the answer quotes for several passages is still one citation. A domain can therefore appear cited on one engine and not the other purely because of how each engine exposes sources.
Ranking is not comparable. The docs say the two Gemini engines rank citations differently, so an average position is not comparable between Gemini (API) and Gemini (app). Compare whether you were cited, not where in the list.
Mentions come from text, and the app text has source links in it. Gemini (app) usually writes a link to the source after a passage it supports, with the site address as the link text. Mention extraction blanks those out, so a site named only in such links does not count as a brand mention. The stored answer keeps them. If you read the raw app answer and see your domain, check whether it was only in a source link.
The discrepancy worksheet
Use this worksheet for every prompt where the two Gemini engines disagree. Copy it into a document or sheet, one block per prompt. The steps are ordered so you rule out the boring explanations first.
GEMINI APP vs API DISCREPANCY WORKSHEET
Prompt (exact text):
Target: country / city / persona / language:
Window or run date compared (same for both):
Analyst / date of review:
STEP 1. Did both engines actually run?
- Gemini (API) completed runs in window: ___
- Gemini (app) completed runs in window: ___
- Gemini (app) skipped, and reason (persona, Chinese, over 2,000 chars): ___
- Any failed runs on either side? (yes/no, dates): ___
STEP 2. Was the location delivered the same way?
- API delivery (from run history): ___
- App delivery (from run history): ___
- City sent at country level on the app? (yes/no): ___
STEP 3. Compare answer evidence (per engine)
API APP
- Brands named (list): ___ ___
- Own brand named? (y/n, counts): ___ ___
- Competitors named (list): ___ ___
- Cited domains (list): ___ ___
- Own domain cited? (y/n, counts): ___ ___
- Sources listed at all? (y/n): ___ ___
- Answer features seen: ___ ___
- Searches recorded (API only): ___ ___
STEP 4. Classify the difference (tick one primary)
[ ] Not a real difference: unequal runs, failed runs or dates
[ ] Context difference: persona, language, location or prompt length
[ ] Source-exposure difference: same brands, different citation detail
[ ] Content difference: different brands named in a like-for-like run
[ ] Unresolved: needs more runs
STEP 5. What we will say in the report
- Wording (name both engines and what each represents): ___
- Next check (more runs, a different prompt, a manual look): ___
Fill Step 1 and Step 2 before you read a single answer. If either step explains the gap, stop.
How do you get the numbers for step 3?
Open the prompt from the Prompts list. The prompt detail page has an engine-by-engine breakdown, a tally of every brand an engine cited for the question, the citations grouped by engine, and a run history that names each run's location and how it was delivered. Filter by engine to read one side at a time, using the same window for both. The run history shows at most the 50 most recent runs in the window.
For a project-wide view, build the same query twice in the Explorer with engine as a breakdown, for example brand visibility and citation rate for the Gemini engines side by side over the same window. Copying the page address gives another project member the same query. The key terms page defines mention, citation and retrieval, which you will want to keep distinct in the write-up.
Worked example
Illustrative example: Quillstone and its competitors are fictional, and the numbers are made up to show the method.
Quillstone sells document-review software to mid-sized legal and compliance teams. Its analyst tracks the prompt "best document review software for a compliance team" with one target, United Kingdom, as written, which runs daily. Competitors are Brieflane and Clausewise. Over the same 28-day window she sees this:
| Measure | Gemini (API) | Gemini (app) |
|---|---|---|
| Completed runs | 28 | 28 |
| Runs naming Quillstone | 14 | 7 |
| Share of runs naming Quillstone | 50% | 25% |
| Runs citing quillstone.example | 10 | 8 |
| Runs naming Brieflane | 12 | 15 |
Her first read is "Gemini is inconsistent." She then works the sheet.
Step 1. Both engines completed 28 runs, one for each day, and no failed runs are listed. The counts are equal, so the gap is not a coverage artefact.
Step 2. The target is a country, as written, so no city or persona is involved. Location arrived in the prompt for the API and natively for the app. That is a real difference in how the question is asked, and she records it rather than dismissing it.
Step 3. Brand mentions: 14 of 28 is 50%, and 7 of 28 is 25%. Citations: the API cites Quillstone in 10 runs, and the app in 8. The app cites the domain in more runs than it names the brand, so she opens the app answers that cite it without naming it. In those she finds the domain only in a source link after a passage, which mention extraction blanks out, so the link counts as a citation but not as a mention.
She also notices that 5 of the 28 app answers list no sources. Gemini decides whether to search on the app, so an answer with no sources cannot cite Quillstone at all.
Step 4. Primary class: context difference and source-exposure difference, not a content contradiction. The question and dates match, but the two routes deliver location and sources differently.
Step 5. Her report gives the two rates on separate lines under "Gemini (API), our own call with Google Search grounding" and "Gemini (app), what a signed-out visitor sees on gemini.google.com". It does not average them, and it does not compare citation position. The next check is a second 28-day window to see whether the 50% and 25% gap holds, as the sample is small.
Note what she did not do. She did not conclude that the model "prefers" Brieflane in the app, or that Quillstone is ranked lower there. The data shows what each answer named and cited, not why.
Common mistakes and what this cannot tell you
- Averaging the two engines into one "Gemini" number. They are separate engines with separate channels. Pooling hides the very difference you are trying to explain.
- Comparing citation position across them. The docs are explicit that it is not comparable between the two.
- Treating the app result as the truth and the API as noise, or the reverse. Gemini (app) is what a signed-out visitor sees, captured through a data provider; Gemini (API) is a direct API output. Neither is what a signed-in person with history and preferences would see.
- Blaming the model for a location effect. A city target on the app arrives as coordinates; on the API it arrives as text. A country-wide local question on the app is answered for wherever the data provider's connection is, not a place DiscoveredBy chooses. In one check, a United States local question was answered for Colorado. Track local prompts as city targets.
- Reading a single pair of answers. One API answer and one app answer tell you very little. Compare rates over many runs.
- Forgetting the app's own limits. No persona variants, no Chinese variant, no very long prompts, and no reported searches.
- Explaining the cause. A citation or mention is an observation of one answer. It does not show why Gemini produced it, and a difference between engines does not say which internal factor drove it.
What this method can tell you is whether a disagreement survives once the runs, dates and delivery are aligned, and where the evidence differs. What it cannot tell you is what any individual user saw, or what Google's models weigh.
How should you report the two Gemini engines?
Report them as two lines with their labels, and say what each represents. If a client asks "what does Gemini say about us", the honest answer has two parts.
A useful template sentence: "On Gemini (app), the answer a signed-out visitor sees on gemini.google.com, Quillstone was named in X of N runs. On Gemini (API), our own call to Google's Gemini API with Search grounding, it was named in Y of M runs." Keep the run counts visible so a reader can see the denominators.
If you already keep a separation between app and API results, see Keep app results and API results separate in your reports. Two related posts are Why your manual ChatGPT check differs from a monitoring result and Keep a measurement change log when an AI engine changes.
Frequently asked questions
Which Gemini result should I trust?
Neither replaces the other; they answer different questions. Gemini (app) shows what a signed-out visitor sees on gemini.google.com. Gemini (API) shows what Google's Gemini API returns with Search grounding. Pick the one that matches the claim you are making, and label it.
Can I add the two together for a single Gemini score?
DiscoveredBy keeps them as separate engines in every filter and breakdown, and that is the safer default for reporting. Adding them mixes two collection routes, two ways of exposing sources and, in DiscoveredBy's checks of 2026-09-28, two different models. If a leadership summary needs one line, show both and let the reader see the split.
Why does the app cite fewer sources than the API, or none?
On Gemini (app), Gemini decides whether to search, and DiscoveredBy never asks it to, so some answers list no sources. Google product and Maps place links are also counted as answer features rather than citations. The API is called with grounding, so its sources are exposed differently.
Is Gemini (app) the same as Google AI Overviews?
No. Gemini (app), Gemini (API), Google AI Overviews and Google AI Mode are four separate engines. AI Overviews is the AI Overview on Google's results page for the prompt, and AI Mode is Google AI Mode's answer. See the glossary entry for Gemini (app) and the AI Overviews entry.
Does it matter that the app is captured signed out?
Yes, for interpretation. The capture is a session that is not signed in, so it cannot reflect a particular person's account, history or personalisation. It is a consistent, repeatable view of what an anonymous visitor gets, which is what makes it comparable over time.
How many runs do I need before I call the difference real?
The docs do not set a threshold, and this post will not invent one. Compare the same window for both engines, look at the counts behind each rate, and repeat over a second window. If the gap moves around, treat it as unresolved.
Next step
Track your key prompts on both Gemini engines, then use the Explorer with engine as a breakdown and the prompt detail run history to fill the worksheet above. The Gemini engine page and the AI visibility feature overview explain what is measured and how, and the Explorer docs cover the breakdowns.
- citations
- reporting
- measurement
- gemini
- engine differences