Reference
Engines and measurement
The eight AI engines we measure, how we query them (the two Google engines, ChatGPT (app) and Gemini (app) through a licensed data provider), which engines run for your project, how a country or city reaches each one, and how each answer is labelled.
The engines we measure
DiscoveredBy measures eight engines: Gemini (API), Perplexity, Google AI
Overviews, Google AI Mode, ChatGPT (app), Gemini (app), Claude and Grok.
Four of them, Gemini (API), Perplexity, Claude and Grok (the chat engines,
on these pages), are queried through their own company's API, never by
scraping a browser session. The other four are collected differently: DataForSEO, a third-party data
provider, fetches the answer a person would see and returns it to us. For
the two Google engines that is Google's own result for the prompt (see
How the Google engines are collected);
for ChatGPT (app) it is chatgpt.com's answer (see
How ChatGPT (app) is collected); for
Gemini (app) it is gemini.google.com's answer (see
How Gemini (app) is collected).
In the platform's own bookkeeping this is recorded as the collection
method: api for the four chat engines, licensed for the four
DataForSEO engines, which the app shows as "Licensed data".
PROVIDER_PROVENANCE in api/services/collection.py holds that value for
each of the eight.
The word the platform uses to track an engine internally is not always the name of the company running it:
| Engine you see | Internal slug | Company |
|---|---|---|
| Gemini (API) | gemini |
|
| Perplexity | perplexity |
Perplexity |
| Google AI Overviews | google_aio |
|
| Google AI Mode | google_ai_mode |
|
| ChatGPT (app) | chatgpt_app |
OpenAI |
| Gemini (app) | gemini_app |
|
| Claude | anthropic |
Anthropic |
| Grok | grok |
xAI |
grok runs on xAI's infrastructure, gemini, google_aio,
google_ai_mode and gemini_app on Google's, and chatgpt_app on
OpenAI's: the slug names what is being called, not
necessarily the brand a reader might expect from the word alone. Gemini
(API), Gemini (app), Google AI Overviews and Google AI Mode are four
separate engines: Gemini (API) is Google's model answering through its
API, Gemini (app) is the answer gemini.google.com shows a person who is not
signed in, Google AI Overviews is the AI Overview on Google's results page
for the prompt, and Google AI Mode is Google AI Mode's answer to it.
ChatGPT is collected from the app only: ChatGPT (app) is the answer
chatgpt.com shows a person who is not signed in. ChatGPT was once also
collected through OpenAI's API as a separate engine, ChatGPT (API), whose
stored answers were deleted when it was removed. Documents kept as they were
sent (weekly reports, shared snapshots and alerts) may still name it.
How Perplexity answers
Perplexity answers through its Sonar API, with the sonar model and a
medium search context size
(api/services/llm.py#run_perplexity, #DEFAULT_PROVIDER_MODELS,
#PERPLEXITY_SEARCH_CONTEXT_SIZE). Its message is built by the same
function as every other engine's and sent as one user message: the prompt
plus any audience, response-language and "Search context" blocks
(api/services/llm.py#render_prompt_with_location). Persona and language
variants run on Perplexity, as on the other chat engines; only the four
DataForSEO engines skip some variants (see
What reaches Google,
What reaches ChatGPT and
What reaches Gemini (app))
(api/services/collection.py#target_runs_on).
Sonar writes a generated answer and marks the sources it used with
numbered markers such as [1], which stay in the saved answer text as
Sonar wrote them. Every entry in Sonar's list of search results is recorded
as a returned search result. Only a source a marker points at becomes a citation, with
the answer text just before that marker, when there is any, saved as the
cited passage; a marker whose number matches no entry in its citation list is ignored
(api/services/llm.py#_sonar_marker_citations). A Perplexity citation's rank
is the order in which the answer first cites it: the first source any marker
points at is 1, the next new source 2, and so on. Sonar's own marker number
counts through its list of sources, not through the answer, so it is not
used as the rank; it is kept in the citation's stored metadata instead. A
cited source that Sonar's search results leave out is still recorded as a
returned search result, after the listed ones
(api/services/llm.py#run_perplexity). That is why Perplexity
has rates on Retrieved vs cited
(api/services/collection.py#RETRIEVAL_VISIBILITY).
How the Google engines are collected
Google AI Overviews and Google AI Mode are not called through an API of
Google's. DiscoveredBy sends each run to DataForSEO, a third-party data
provider, which fetches Google's own result for the prompt and returns it:
for Google AI Overviews, Google's results page, with the AI Overview on it
and the page's organic results; for Google AI Mode, Google AI Mode's answer,
which has no organic results
(api/services/llm.py#run_google_aio, #run_google_ai_mode,
#_run_dataforseo). That is why both carry the collection method
licensed, shown as "Licensed data"
(api/services/collection.py#PROVIDER_PROVENANCE).
What that means for your data: for every Google run, your prompt text, the
target's country (or, for a city target, the city's coordinates) and a
language code are sent to DataForSEO, a third party, which queries Google
with them. Nothing else about your project is sent
(api/services/llm.py#dataforseo_task). Google does not say which model
wrote the answer, so the model recorded on each run names the surface
instead: ai-overview or ai-mode
(api/services/llm.py#DEFAULT_PROVIDER_MODELS). Both engines run only
while DataForSEO credentials are configured on the platform
(api/services/llm.py#provider_has_api_key).
What reaches Google
- The prompt, on one line. Google receives your prompt text as a
search: runs of spaces and line breaks become one space, and the text is
trimmed. Characters such as
+and%are encoded so that Google receives them as typed. The text saved as sent on the run is that one-line version (api/services/google_serp.py#collapse_keyword,#google_keyword,api/services/llm.py#_run_dataforseo). - Nothing added to it. No audience block, no response-language
instruction and no "Search context" block: location and language travel
as separate settings of the request. The results are a desktop search
(
api/services/llm.py#dataforseo_task). - Location. A country is sent as DataForSEO's location code for each of
the 18 countries a project can target, and by its name for any other
country. A city is sent as its latitude and longitude, with a radius for
AI Overviews and a zoom level for AI Mode
(
api/services/google_serp.py#DATAFORSEO_COUNTRY_CODES,api/services/llm.py#dataforseo_location). Every one of the 344 cities we seed has coordinates, filled from Wikidata (api/services/city_coordinates.py#CITY_COORDINATES,api/services/cities.py#seed_cities). A city with no coordinates on record runs at country level on both Google engines, and the run records its city as not sent (api/services/collection.py#COORDINATE_SURFACES). - Language. The variant's language; for an As written target, your
project's default language; otherwise English
(
api/services/llm.py#dataforseo_language). Both Google engines support every language we offer except Chinese, where choosing Simplified or Traditional would be a guess. On AI Mode, Portuguese is sent as Brazilian Portuguese and Tagalog as Filipino (api/services/google_serp.py#AI_MODE_LANGUAGE_CODES,#AI_OVERVIEWS_LANGUAGE_CODES). The language actually sent is shown on each Google run in a prompt's run history, as "Language sent to Google" (api/services/google_serp_reads.py#serp_language_label).
Some variants never run on a Google engine, and each is left out rather
than sent in a changed form
(api/services/collection.py#target_runs_on):
- A persona variant: Google receives only the search text, so there is nowhere to put an audience.
- A language variant in Chinese.
- A prompt longer than 700 characters once encoded for Google. It is
never cut short to fit (
api/services/google_serp.py#google_keyword_fits).
Every other variant still runs on the Google engines the prompt runs on, and every variant still runs on the chat engines. Prompt detail lists each variant a Google engine skips and why (see Prompts).
What counts as the answer
The answer text is the AI Overview's own text, tables and expanded
sections. Product cards, ads and videos inside an Overview are left out of
the text and of the citations (api/services/llm.py#_ai_overview_answer,
#_ANSWER_ELEMENTS, #_EXCLUDED_ELEMENTS). The same rule reads an AI Mode
answer. An Overview's shopping (its product cards and Google's
product-viewer links), ads, videos (its video parts and the YouTube
videos it links) and tables are counted as
answer features instead; Google AI Mode
reports no feature we can count, because its images and product cards are
drawn inside its answer text
(api/services/response_features.py#SUPPORTED_FEATURES).
Its citations are the sources the answer links: each part's references,
then the links in its text, then the Overview's own list of references, in
the order they first appear, so a citation's rank is the order in which the
answer first cites that source. A page cited twice is one citation (a
quote link, utm_ parameters and a YouTube timestamp are removed before
comparing), and Google's own search and shopping pages are not counted as
sources (api/services/llm.py#clean_reference_url, #_is_google_search_url).
A link's anchor text is kept with the citation but never used as the
page's title, since it names what the answer was talking about, not the
page.
Google reports only the sources its answer cites, never a list of pages it
searched, so a Google answer has no returned search results: on
Retrieved vs cited its channel reads
Cited sources only and is not compared, and its fan-out is not
applicable, because Google does not report the searches behind the answer
(api/services/collection.py#RETRIEVAL_VISIBILITY,
#SURFACES_WITHOUT_FAN_OUT). A completed Google answer is analysed for
brand mentions and sentiment like any other answer
(api/tasks.py#enqueue_post_answer_work).
When Google shows no AI answer
Google does not show an AI Overview for every search. When the request
succeeds but there is no AI answer (no Overview on the page, no AI Mode
answer, or an Overview made only of
product cards, ads or videos), the run is stored as its own outcome, no
AI answer shown, with no text and no citations
(api/services/llm.py#_run_dataforseo, #_dataforseo_outcome,
api/services/execution.py#execute_query_for_provider). It is the day's
result for that variant, like a completed run, and is not retried that day
(api/services/execution.py#execution_blocks_same_day_rerun).
A run with no AI answer is not an answer, so it is left out of every
answer-based number: visibility, mentions, share of voice, citation rate
and the answer counts. It is reported instead as the shown rate: the
share of runs on which Google showed an AI answer, per Google engine on
prompt detail and as the Explorer metric AI answer shown rate (see
Metrics defined).
Collection health counts it as a request that succeeded, not as a failure
(api/services/collection_health.py#providers_in_outage). A request that
fails at DataForSEO or Google is a failed run, never "no AI answer shown".
Temporary failures are retried up to three times first, and the run fails
only if the last try fails too: network errors and timeouts, rate limits,
DataForSEO's own server-side errors, an empty response, the
account-verification error DataForSEO can return briefly, and a Google AI
Overviews results page with neither an Overview nor any organic result (an
empty page is not a search Google answered). Any other error, such as an
invalid request, a payment or balance problem or a blocked account, fails
the run at once, without those tries
(api/services/llm.py#_post_dataforseo_with_retries, #_dataforseo_outcome).
Either way, the failed run is then eligible for the same-day retries that
cover every engine (see
A run failed;
api/cli.py#cmd_retry_failed_executions).
Organic rank next to the Overview
For every Google AI Overviews run, whether or not an Overview was shown,
DiscoveredBy also keeps the organic results on the first page of the same
search (a first page holds up to 10, and often fewer): each result's
position among those organic results, its address and its title (api/services/llm.py#_organic_results,
api/services/execution.py#_store_organic_result). Your own site is
matched the way a citation is, by the registrable domain of your project's
site. Prompt detail and the Organic results export match a tracked
competitor by its domain against the competitors active on the project
now (api/services/google_serp_reads.py#organic_panel). Prompt detail
shows each one's organic position next to whether the Overview cited it,
and the Explorer measures it as first-page organic rate and organic
position.
Organic results are only recorded: nothing fetches those pages. AI Mode
has no organic results.
How ChatGPT (app) is collected
ChatGPT (app) is what a person who is not signed in sees on chatgpt.com.
DataForSEO, the data provider that also fetches the two Google engines, asks
chatgpt.com the prompt and returns the answer the page shows, which
DiscoveredBy stores (api/services/llm.py#run_chatgpt_app,
#CHATGPT_APP_URL). It carries the collection method licensed, shown as
"Licensed data" (api/services/collection.py#PROVIDER_PROVENANCE).
It is the only ChatGPT engine.
What that means for your data: for every ChatGPT (app) run, your prompt
text, the target's country and a language code are sent to DataForSEO, a
third party, which puts the prompt to chatgpt.com in a session that is not
signed in. Nothing else about your project is sent
(api/services/llm.py#dataforseo_task). ChatGPT (app) runs whenever
DataForSEO credentials are configured on the platform, exactly like the two
Google engines (api/services/llm.py#provider_has_api_key).
What reaches ChatGPT
- The prompt as written, on one line. Runs of spaces and line breaks
become one space, and
%and+are encoded so they arrive intact (api/services/google_serp.py#google_keyword). A prompt over 2,000 characters once encoded does not run on ChatGPT (app); it is never cut short (api/services/chatgpt_app.py#KEYWORD_LIMIT,#keyword_fits). - Nothing else in the text. No audience block, no response-language
instruction and no "Search context" block, so persona variants do not
run on it (
api/services/collection.py#target_runs_on). - The country, as a setting. DataForSEO's location code for each of
the 18 countries a project can target, and the country's name for any
other, never coordinates (
api/services/llm.py#dataforseo_location). There is no city setting, so a city target does not run on ChatGPT (app), and prompt detail lists it as not run. Where a local question is answered for depends on DataForSEO's connection, not on a place DiscoveredBy chooses: in one test, a "near me" question for India was answered for Siliguri. - The language, as a setting. The variant's language; for an As
written target, your project's default language; otherwise English
(
api/services/llm.py#dataforseo_language). Every language we offer is supported except Chinese: a Chinese variant does not run, and an As written target in a project whose default is Chinese is sent as English (api/services/chatgpt_app.py#LANGUAGE_CODES). Run history shows the language sent on each run, as "Language sent to ChatGPT (app)" (api/services/google_serp_reads.py#serp_language_label). - ChatGPT decides whether to search. DiscoveredBy never asks it to search, so some answers cite no sources; the Explorer metric Answers citing sources shows how often answers cite any.
- Brand studies are the one exception to the text rule. An
objection,
attribute or
fact study question put to ChatGPT (app)
with a study language gets the same response-language line the chat
engines receive, added after the question, because the language setting
alone does not change the language ChatGPT (app) answers in. A study
question is always sent with the United States as its country and
English as its language setting, whatever the study language; an As
written study sends the question alone
(
api/services/brand_study/question.py#study_prompt,#STUDY_APP_COUNTRY,api/services/llm.py#dataforseo_language). Tracked prompts never get the line.
What counts as the ChatGPT (app) answer
- The text the page shows, product lists and local business lists
included, because they name brands a reader sees
(
api/services/llm.py#_chatgpt_app_answer). DataForSEO's copy sometimes repeats the end of an answer after the answer is over; a repeat of 200 characters or more is removed, and a shorter one, such as a repeated closing line, is kept (api/services/chatgpt_app.py#drop_repeated_tail,#MIN_REPEATED_TAIL_CHARS). Sponsored units, if DataForSEO includes them in the text, are part of what is stored. Products, local businesses, ads, images and tables, and whether the answer showed sources, are also counted as answer features. - Citations are the answer's inline sources, in the order they first
appear, so a citation's rank is that order. A page cited twice is one
citation, compared after removing a quote link,
utm_parameters and a YouTube timestamp, and a source whose address cannot be read is skipped (api/services/llm.py#clean_reference_url,#_http_url). Links to a local business's own site, and product and retailer links, are not citations. - The searches it ran and the pages they found, when it searched: the
Fan-out panel lists them, and a cited page is
always among the pages found (
api/services/llm.py#_chatgpt_app_searches). Answers collected before 2026-09-29 read not recorded there. On Retrieved vs cited its channel still reads Cited sources only (api/services/collection.py#RETRIEVAL_VISIBILITY): the pages found are stored but not compared yet, because DataForSEO returned none on 2026-09-28, and the attribution figure needs citation positions this engine does not record. - The model reads
chatgpt-app, whatever model DataForSEO reports, so a model change on chatgpt.com does not split the engine's history (api/services/llm.py#DEFAULT_PROVIDER_MODELS). - When DataForSEO cannot read the answer, the run is a failed run:
ChatGPT (app) has no "no AI answer shown" outcome
(
api/services/llm.py#_dataforseo_outcome). As on every engine, the day's failed runs are retried automatically at 03:00, 05:00, 09:00 and 17:00 UTC, and a run still failed after the 17:00 retry stays failed for that day (see A run failed;api/cli.py#cmd_retry_failed_executions). A response with no answer text, or with an answer that looks cut off, is retried like the temporary failures listed under When Google shows no AI answer, and the run fails if the last try is no better (api/services/llm.py#_post_dataforseo_with_retries). An answer looks cut off when the stored text is a single paragraph shorter than 300 characters and cites no source: DataForSEO can read the page before ChatGPT has finished and return only its opening line (api/services/llm.py#_looks_cut_off,#CUT_OFF_ANSWER_CHARS). The rule does not depend on the language, so a genuinely brief answer of that shape is a failed run too, never an answer that names no brand.
A completed ChatGPT (app) answer is analysed for brand mentions and
sentiment like any other answer (api/tasks.py#enqueue_post_answer_work),
with one difference. After a sentence it supports, the app can write a
small link to the source whose text is the site's address, such as
[rtings.com](https://...). Mention extraction reads the answer with each
such link blanked out, so a site named only in such links is not counted
as a brand mention; the stored answer keeps them
(api/services/brand_mentions.py#mention_text). A link is treated as one
of these source links only when its text is a bare domain and it directly
follows the text it supports: a full stop, !, ?, :, ; or …; the
same marks in other scripts, such as 。, ? and ! (Japanese and
full-width), । (Hindi, Bengali, Punjabi), ۔ (Urdu), ؟ and the Arabic
comma ، (Arabic-script writing often ends a passage with a comma where
English uses a full stop); a closing bracket, such as ), ], ) or
」; a closing quotation mark, such as ”, ’, », or a " or German
“ that follows text; the end of bold text; or another such link. A domain
link that is part of a sentence still counts as a mention: one at the start
of a line, of a list item (after a marker such as -, * or 1.) or of a
quoted line (after >); one right after bold or a quotation mark (", ',
“, ‘ or «) that opens at the start of a line or after a space, as in
1. **[monday.com](https://...)**; and one after an ordinary word, as when
a ChatGPT (app) paragraph opens with [xero.com](https://...) as its
subject. So does a link whose text is a name, such as a product or a place.
In a language written without sentence punctuation, such as Thai, which
ends a sentence with a space, a source link after a passage follows a word,
so it is read like one after an ordinary word: it is not blanked, and the
site it names can count as a mention.
How Gemini (app) is collected
Gemini (app) is what a person who is not signed in sees on
gemini.google.com. DataForSEO, the data provider that also fetches the two
Google engines and ChatGPT (app), asks gemini.google.com the prompt and
returns the answer the page shows, which DiscoveredBy stores
(api/services/llm.py#run_gemini_app, #GEMINI_APP_URL). It carries the
collection method licensed, shown as "Licensed data"
(api/services/collection.py#PROVIDER_PROVENANCE). Gemini (API) is the
other Gemini engine: DiscoveredBy's own call to Google's Gemini API
(api/services/llm.py#run_gemini). The two are separate engines in every
filter and breakdown, and their answers to the same prompt can differ: in
every one of our checks (2026-09-28) DataForSEO reported the app's model as
"3.5 Flash-Lite", which is not the model Gemini (API) calls
(api/services/llm.py#DEFAULT_PROVIDER_MODELS).
What that means for your data: for every Gemini (app) run, your prompt
text, the target's country (or, for a city target, the city's coordinates)
and a language code are sent to DataForSEO, a third party, which puts the
prompt to gemini.google.com in a session that is not signed in. Nothing
else about your project is sent (api/services/llm.py#dataforseo_task).
Gemini (app) runs whenever DataForSEO credentials are configured on the
platform, exactly like the other three DataForSEO engines
(api/services/llm.py#provider_has_api_key).
What reaches Gemini (app)
- The prompt as written, on one line. Runs of spaces and line breaks
become one space, and
%and+are encoded so they arrive intact (api/services/google_serp.py#google_keyword). A prompt over 2,000 characters once encoded does not run on Gemini (app); it is never cut short (api/services/gemini_app.py#KEYWORD_LIMIT). - Nothing else in the text. No audience block, no response-language
instruction and no "Search context" block, so persona variants do not
run on it (
api/services/collection.py#target_runs_on). - The country, as a setting: DataForSEO's location code for each of
the 18 countries a project can target, and the country's name for any
other (
api/services/llm.py#dataforseo_location). - A city, as its coordinates. Unlike ChatGPT (app), Gemini (app) runs
city targets: the city's latitude and longitude are sent with a radius
of 20 metres around that point
(
api/services/llm.py#dataforseo_location,#GEMINI_APP_RADIUS). A city with no coordinates on record runs at country level, and the run records its city as not sent (api/services/collection.py#COORDINATE_SURFACES). - A country-wide local question is answered for wherever DataForSEO's connection is, not for a place DiscoveredBy chooses: in one check, a local question for the United States as a whole was answered for Colorado. To see local answers for a place, track the prompt with a city target.
- The language, as a setting. The variant's language; for an As
written target, your project's default language; otherwise English
(
api/services/llm.py#dataforseo_language). Portuguese is sent as Brazilian Portuguese and Tagalog as Filipino, because those are the codes Gemini (app) lists; a Chinese variant does not run, and an As written target in a project whose default is Chinese is sent as English (api/services/gemini_app.py#LANGUAGE_CODES). Run history shows the language sent on each run, as "Language sent to Gemini (app)" (api/services/google_serp_reads.py#serp_language_label). - Gemini decides whether to search. DiscoveredBy never asks it to search, so some answers list no sources.
What counts as the Gemini (app) answer
- The text the page shows, tables included, as DataForSEO returns it
(
api/services/llm.py#_gemini_app_answer). Nothing is removed from it. - Citations are the web pages among the answer's sources, one per page
however many passages of the answer it supports, in the order they first
appear, so a citation's rank is that order. A page's address is compared
after removing a quote link and
utm_parameters, and a source whose address cannot be read is skipped (api/services/llm.py#clean_reference_url,#_http_url). - Google product links and Google Maps place links are not citations.
Gemini (app) lists them among its sources; they are counted as shopping
and local businesses among the
answer features instead, with web
search, images and tables
(
api/services/response_features.py#SUPPORTED_FEATURES). - Every passage a page supports is recorded. Gemini (app) says which
passage of the answer each source supports. A passage that occurs exactly
once in the answer is recorded for its page, so a
fact check can link a quote in any of them
to its source, and the first one is the citation's evidence for
sentiment (
api/services/llm.py#_gemini_app_answer,api/services/gemini_app.py#MAX_OCCURRENCES). - No list of pages searched, so the answer has no returned search
results: on Retrieved vs cited its
channel reads Cited sources only
(
api/services/collection.py#RETRIEVAL_VISIBILITY), and its fan-out is not applicable (api/services/collection.py#SURFACES_WITHOUT_FAN_OUT). - The model reads
gemini-app, whatever tier DataForSEO reports, so a change of model at Google does not split the engine's history (api/services/llm.py#DEFAULT_PROVIDER_MODELS). - Answers citing sources. An answer whose only sources are Google
product or Google Maps place links (every local answer in our checks) is
left out of the Explorer metric
Answers citing sources:
it lists sources, but none is a web page we count as a citation. It still
counts as an answer that used web search
(
api/services/explorer/metrics.py#_runs,api/services/collection.py#SURFACES_WITH_UNCITED_SOURCES).
Product links listed for another country
Google's product and place links can carry the country they are listed
for (a gl parameter).
When an answer's product links (or place links) are listed for another
country than the target's, the answer is kept as it is, the country is
recorded on the run, and that answer is left out of the shopping rate (or
the local businesses rate) rather than counted
(api/services/gemini_app.py#listed_country_mismatch). In one check, an
answer for the United States listed its products for Malaysia, priced in
ringgit. We do not know whether the rest of such a page was shown for that
other country: the text of that answer read as written for the target.
When the capture fails
When DataForSEO cannot read the answer, the run is a failed run: Gemini
(app) has no "no AI answer shown" outcome
(api/services/llm.py#_dataforseo_outcome). As on every engine, the
day's failed runs are retried automatically at 03:00, 05:00, 09:00 and
17:00 UTC, and a run still failed after the 17:00 retry stays failed for
that day (see A run failed;
api/cli.py#cmd_retry_failed_executions). A response with no
answer text is retried like the temporary failures listed under
When Google shows no AI answer, and the
run fails if the last try is no better
(api/services/llm.py#_post_dataforseo_with_retries). A short answer is an
answer: the cut-off rule of ChatGPT (app) does not apply to Gemini (app),
so a one-line answer such as a sum is stored like any other
(api/services/llm.py#_retry_reason).
Brand mentions and sentiment
A completed Gemini (app) answer is analysed for brand mentions and
sentiment like any other answer (api/tasks.py#enqueue_post_answer_work).
Like ChatGPT (app), Gemini (app) usually writes a link to the source after
a passage it supports, with the site's address as the link text, such as
[example.com](https://...) (of the six answers in our checks that cited
web pages, one had no such links). Mention extraction reads the answer with
each such link blanked out, so a site named only in such links is not
counted as a brand mention; the stored answer keeps them
(api/services/brand_mentions.py#mention_text). The rule is the one
described for ChatGPT (app), including the punctuation of other scripts: a
bare-domain link is blanked only when it directly follows the text it
supports, so a domain link at the start of a line, a list item or a quoted
line, right after opening bold, or after an ordinary word, and a link whose
text is a name, such as a place or a product, are still read. In Thai and
other languages written without sentence punctuation, the links after a
passage follow a word, so they are read and the sites they name can count
as mentions. In our checks, which were English and German answers, every
such Gemini (app) link followed the passage it supports, so all of them
were blanked.
How an answer is labelled
Every prompt execution stamps its own copy of three fields the moment it
runs: platform, surface and collection method (PromptExecution.platform,
.surface and .collection_method in api/models/prompt_execution.py).
They come from the code that calls the engine, not from the provider
record, and they are written when the run starts and again when it
succeeds (api/services/collection.py#provenance_for,
api/services/execution.py#execute_query_for_provider). They are never
looked up live afterward.
That is why an old answer keeps saying how it was actually collected. If the catalogue changes later, a provider renamed, a new surface added, a collection method swapped, none of that can reach back and relabel an execution that already ran: the row already carries its own copy of what happened, not a pointer to a catalogue entry that keeps moving underneath it.
Perplexity answers collected through its earlier Search API, a ranked list
of results rather than a generated answer, keep what was recorded for
them: the "Search results" channel label on the Sentiment screens, a fan-out of not applicable,
no rates on Retrieved vs cited, and, where location delivery was
recorded, a city sent "in the search query"
(api/services/collection.py#SURFACES_WITHOUT_FAN_OUT,
#RETRIEVAL_VISIBILITY,
frontend/src/routes/(app)/sentiment/+page.svelte#channelLabel,
frontend/src/lib/location-delivery.js#deliveryText).
Which engines run for your project
Which of the eight engines actually run for your project comes from two things layered together: what your plan includes, and whether each engine is currently available to run at all.
Your project's plan lists which providers it includes (BillingPlan.providers,
seeded from DEFAULT_PLANS in api/services/seed.py). Different plans
include different subsets of the eight engines. What each plan includes is
covered on Pricing, not repeated here, so there is one place to
keep it correct.
Being on your plan's list is not the same as running today. get_user_provider_ids
in api/services/plan.py filters that list twice more: the provider row has
to be marked active, and provider_is_available (in api/services/llm.py)
has to return true for it, which in turn calls provider_has_api_key to
check whether a key for that engine is currently configured on the platform
(for the four DataForSEO engines, both halves of the DataForSEO sign-in).
An engine can be missing from your project for a reason that has nothing to
do with your plan: if its key is not currently configured, that engine
will not run for anyone, on any plan, until the key is back in place.
If an engine stops running after it has run (its key is removed, for
example), its past answers keep its name, and no new answers arrive. The
dashboard's per-engine figures, the weekly report's engine table and the
engine list the competitors page returns still include it for any period in
which it
answered, so the rows add up to the totals beside them, which count its
answers too
(api/routers/dashboard.py#build_dashboard,
api/routers/pages.py#competitors_page). The comparisons the product makes
for you leave it out while either period holds its answers, because an
engine counts in a period only when it ran on at least 80% of the period's
answered days (api/services/collection_health.py#WINDOW_PRESENCE_MIN_PERCENT):
the Overview's changes (api/services/overview.py#_compare_shared_engines),
the dashboard and weekly report changes
(api/routers/dashboard.py#build_dashboard), alerts
(api/services/watchdog.py#detect_alerts_for_project,
api/services/brand_signals.py#build_population), the competitor trends
(api/routers/pages.py#_trend_engines), social sources
(api/services/social_sources/service.py#social_sources_for_project) and
optimization and citation-gap outcome checks
(api/services/outcome_engines.py#shared_outcome_engines). The comparisons
you set up yourself do not: an Explorer comparison, an Analyst answer, a
saved dashboard and the change column on the Local page compare whatever
engines you select, so with every engine selected an engine that answered
in only one period (an app engine added in the window, for example) counts
in whichever period it answered
(api/services/explorer/query.py#execute_query,
api/services/local_view.py#local_view); choose engines to compare like
for like. Editing a prompt or a
variant template whose engines name it shows that engine as "not
available right now", and saving that engine choice is refused until it is
removed
(api/services/prompt_variants.py#validate_engines).
Fan-out
Some engines don't just answer a prompt directly. Internally, they can expand it into several related queries of their own, the way a person might turn one question into a few separate searches, before composing a final answer. DiscoveredBy captures that expansion where a provider exposes it, and shows it alongside the answer.
Fan-out is detail about a single execution, not a separate collected answer.
However many queries a prompt expands into, it still counts as one completed
execution for that engine on that day: query_executions carries a unique
constraint on prompt target, provider and day
(uq_qe_target_provider_date in api/models/prompt_execution.py). A prompt
target is one prompt tracked in one location (a whole country or one city),
for one audience and one language, so a prompt tracked in three countries
produces three executions per engine per day, not one; fan-out does not
change that count either way. Because the metrics on
Metrics defined that divide by collected or
analysed answers are built from this same execution count, a prompt that fans
out widely does not, by itself, move one of those numbers differently from a
prompt that does not fan out at all. See that page for which metric divides
by which population.
Perplexity is the case to know about. Sonar searches the web for its
answer and returns the sources it found, but not the text of the searches
it ran, so DiscoveredBy records no sub-queries for it rather than
inventing any (api/services/llm.py#run_perplexity). A Perplexity answer
that returned any source therefore reads searches not recorded: a gap
in what the engine reports, not a sign that it answered without searching
(derive_fan_out_state in api/services/fan_out.py).
Google AI Overviews, Google AI Mode and Gemini (app) report no searches at
all, only the sources the answer cites, so their fan-out reads not
applicable
(api/services/collection.py#SURFACES_WITHOUT_FAN_OUT). ChatGPT (app)
reports its searches when it chooses to search, from 2026-09-29.
Where a provider does report sub-queries, DiscoveredBy also has a model read the captured ones, up to a limit on how many one run judges, and file each under a category describing its relationship to the parent prompt, stored in a table of its own rather than as a column on the query record, so a bad classification can never corrupt the record of what the engine actually searched for. See Fan-out for the category list, what happens to a query past that limit, how confident the model has to be before a label is shown, and what a run's fan-out looks like when little or nothing was captured for it.
How location reaches each engine
Every prompt target has a location: a whole country (a country-wide target) or one city in a country (a city target; see Prompts). Each engine accepts a location differently, so DiscoveredBy sends it in one or both of two ways, depending on what that engine supports:
- Natively: in a location field of the engine's own API, or, for the four DataForSEO engines, of the request to DataForSEO.
- In the prompt: as a "Search context" block of text added after the prompt.
| Engine | Country | City | Search context block |
|---|---|---|---|
| Gemini (API) | In the prompt only | In the prompt only | Yes |
| Perplexity | Native, and in the prompt | Native, and in the prompt | Yes |
| Google AI Overviews | Native only | Native only | No |
| Google AI Mode | Native only | Native only | No |
| ChatGPT (app) | Native only | Not sent (city targets do not run) | No |
| Gemini (app) | Native only | Native only (coordinates) | No |
| Claude | Native, and in the prompt | Native, and in the prompt | Yes |
| Grok | Native, and in the prompt | Native, and in the prompt | Yes |
(api/services/collection.py#LOCATION_DELIVERY.)
- Claude gets an approximate
user_locationon its web search tool: the country code, plus the city, its region and its time zone for a city target (api/services/llm.py#_build_anthropic_web_search_tool). Grok gets the same four values through its web search tool's own location fields (api/services/llm.py#_run_grok_sync). A country-wide target sends the country only. - Gemini (API) is called with Google Search grounding and no location
setting, so the text block is the only place its request carries a
location (
api/services/llm.py#_run_gemini_sync,#run_gemini). DiscoveredBy sends Gemini (API) no latitude or longitude and does not use Gemini's Google Maps grounding. - Perplexity gets a
user_locationin its web search options: the country code, plus the city and its region for a city target. Sonar's location has no time zone field, so no time zone is sent. A country-wide target sends the country only (api/services/llm.py#run_perplexity). - Google AI Overviews and Google AI Mode get the location only as
a setting of the request to DataForSEO: the country's location code (or
its name), or for a city target the city's latitude and longitude, with
no region or time zone. A city with no coordinates on record is sent at
country level (see
What reaches Google;
api/services/llm.py#dataforseo_location). - ChatGPT (app) gets the country only, as a setting of the request to DataForSEO: its location code, or its name. A city target does not run on it (see What reaches ChatGPT).
- Gemini (app) gets the location only as a setting of the request to
DataForSEO, like the Google engines: the country's location code (or its
name), or for a city target the city's latitude and longitude, with a
20-metre radius and no region or time zone. A city with no coordinates on
record is sent at country level (see
What reaches Gemini (app);
api/services/llm.py#dataforseo_location).
The chat engines get the text block; the DataForSEO engines never do. It
names the country and, for a city target, the city (api/services/llm.py#render_prompt_with_location):
Search context:
- Country: United Kingdom (GB)
- City: London
Use this location when searching and ranking sources.
The region and time zone come from the city's own record and are sent only
natively, never in the text block, and Perplexity is sent the region only.
Each is sent on its own when the city has it: a city with only a region on
record sends only the region, and one with neither sends neither (api/services/execution.py#execute_query_for_provider,
api/services/llm.py#SearchContext).
A target location is not a customer's location
Every chat engine's answer comes from an API call that DiscoveredBy makes from its own servers. Nothing runs on a phone or in a browser in that city, no IP address from there is used, and no signed-in account is involved. A Google engine's answer comes from DataForSEO, which DiscoveredBy asks for Google's result with the location as a country code or the city's coordinates; how DataForSEO fetches that result is not visible to us. A ChatGPT (app) answer also comes from DataForSEO, with the country only, in a session that is not signed in, and a Gemini (app) answer comes from DataForSEO with the country or the city's coordinates, in a session that is not signed in. The location is information we pass on, and each engine decides how much it weighs it. A city target shows what an engine answers when it is given that city as the location, in whichever of the ways above it supports. That can differ from what a real person there sees, signed in, on their own device.
What each answer records
When an answer runs, the method used for its country and for its city is
stamped on the answer itself, from the table above as it stood at that
moment. A later change to how an engine is called therefore cannot relabel
an older answer. The city of a country-wide target is recorded as "not
sent", and so is the city of a Google or Gemini (app) run whose city had
no coordinates
(api/services/collection.py#location_delivery_for,
api/services/execution.py#execute_query_for_provider).
- Prompt detail: each answer reads, for example, "Country in the
prompt; city in the prompt" (Gemini (API), city target) or "Country sent
natively and in the prompt; no city targeted" (Claude, country-wide
target), or "Country sent natively; city sent natively" (a
Google engine or Gemini (app), city target). A country-wide target's city
reads "no city targeted"; a Google run for a city with no coordinates
reads "city targeted, but sent to Google at country level (no coordinates
for this city)", and a Gemini (app) run reads "city targeted, but sent to
Gemini's app at country level (no coordinates for this city)"
(
frontend/src/lib/location-delivery.js#deliveryText). - Answers export: the
country_deliveryandcity_deliverycolumns (see Exports and activity). - Customer API and MCP:
location_deliveryon each answer (see Customer API and keys).
An answer collected before this stamp existed shows "Location delivery not
recorded" on prompt detail, empty export columns, and null in the API. It
is never guessed from the engine
(frontend/src/lib/location-delivery.js#DELIVERY_NOT_RECORDED).
Related
- Fan-out: the eight categories a captured sub-query can be classified into, and what each capture state means
- Metrics defined: what collected and analysed answers mean, and why fan-out does not change either count
- Prompts: the shown rate, organic rank next to Overview citations, and the variants a Google engine skips, for one prompt
- Pricing: what each plan includes
- Troubleshooting: what to check when an engine looks inactive or a number looks wrong
Last verified 2026-09-29