All documentation

Optimize

Retrieved vs cited

See which sources each AI engine returned while answering your prompts and which ones the answer text actually attributed, one collection channel at a time, with a saved-copy drill-down for every URL.

Retrieved vs cited showing the collection channels, each with its answer count and either the share of returned sources the answer attributed or why it is not compared, down to the Gemini (app) channel's Cited sources only

The ChatGPT (app) and Gemini (app) rows in this screenshot are demo data: fixed fictional responses run through the product's own collection code.

Illustrative development data is shown in this screenshot.

What this screen answers

An AI engine that searches the web usually brings back more pages than its answer credits. Retrieved vs cited shows both sides for each source: how often an engine returned it while answering your prompts, and how often the answer text attributed it. A page that is returned often but rarely attributed was in front of the engine and still lost the citation.

Open Retrieved vs cited from the Retrieved vs cited tab under Sources in the sidebar. All project members, including viewers, can read it on all plans. It reads saved answers only: nothing is collected, fetched or sent to a model when you open it (api/routers/source_retrieval.py#source_retrieval).

Returned and attributed

  • Returned means the engine reported the source among the pages its search or grounding brought back for that answer.
  • Attributed means the answer text credited the source, and a passage or text position was recorded with that credit (api/services/source_retrieval.py#attributed_citation).

A source counts once per answer, however many times it appears in that answer. Several URLs from one domain in one answer count once for the domain (api/services/source_retrieval.py#source_retrieval_for_project).

Attributed counts can be lower than the Citations page on purpose. The Citations page counts every source recorded for an answer. For Claude, every returned search result is recorded as a citation (api/services/llm.py#run_anthropic), and for a Gemini (API) answer without grounding links every grounding source is too (api/services/llm.py#run_gemini). Counting those as attributed would show Claude crediting everything it found. This screen counts only a source the answer text credits. Every Perplexity Sonar citation counts as attributed: only a source its answer marks is recorded as a citation, and each one carries a text position (api/services/llm.py#_sonar_marker_citations).

One collection channel at a time

What an engine reports as returned is not the same across engines, so rates are shown for one collection channel at a time: the recorded provider, platform, surface, collection method and model together. Each channel card is labelled with what its returned sources can show (api/services/collection.py#RETRIEVAL_VISIBILITY):

  • Full search results: Claude reports the results its search tool returned to it, and Perplexity returns every source it retrieved with its answer, marking the ones the answer used (api/services/llm.py#run_perplexity).

  • Grounding sources only: Gemini (API) reports the pages it grounded the answer on, not its full search results, and its grounding links credit nearly all of them whenever it sends any. A rate would measure whether Gemini (API) sent those links rather than which pages it chose, so it is listed with its number of answers and not compared.

  • Partial: browsed and listed sources: Grok reports the pages it opened and the sources it listed; its individual search results are not exposed.

  • Cited sources only: Google AI Overviews, Google AI Mode and Gemini (app) report only the sources their answer cites, with no list of pages they searched, so there is nothing to compare citations against and each of their channels is listed with its number of answers and not compared (api/services/llm.py#_run_dataforseo, #_gemini_app_answer). ChatGPT (app) is listed the same way: since 2026-09-29 it also reports the pages it found, which are stored, but DataForSEO returned none the day before and ChatGPT (app) citations carry no positions for attribution, so it is not compared yet (api/services/llm.py#_chatgpt_app_answer, #_chatgpt_app_searches).

Perplexity answers collected through its earlier Search API are not compared: their channel reads Perplexity's earlier Search API: no generated answer, because those records hold ranked results and no answer to attribute. A surface we do not recognize shows Unknown capture and is not compared. Rates are computed for Claude, Perplexity and Grok (api/services/collection.py#COMPARABLE_RETRIEVAL). The page opens on the first channel that can be compared; select any card to switch (api/services/source_retrieval.py#_select_channel). Each card shows the share of returned source and answer pairs that were attributed, with both counts. Returned sources are what each engine reported, not every page it read internally.

Your domain

For the selected channel, Your domain shows three numbers:

  • Returned in: answers that returned your project domain, divided by all completed answers in the channel and window.
  • Attributed in: answers that attributed your domain, divided by the same answers.
  • Attributed when returned: answers that both returned and attributed your domain, divided by answers that returned it. With no answer returning your domain, it shows a dash, and with fewer than 30 it is marked provisional.

The source table

Group by switches between domains and exact URLs. Sources filters to your domain, competitors or everyone else. A competitor is a domain on the project's active competitor list today, so the label follows your current list rather than the one in place when the answer was collected (api/services/source_retrieval.py#_ownership). The filter narrows the rows; rates still divide by every answer in the channel and window.

Each row shows returned and attributed answers with their rates, Attributed when returned, and Returned, not attributed: the answers where the engine had the page and did not credit it. Order sorts by any of these, with ties broken by returned answers and then name. Results are paginated 25 at a time, and an out-of-range page resolves to the last available page.

Select a domain to list its URLs, and select a URL to open its detail page. The date range chip in the filter bar chooses 7, 28 or 90 days ending yesterday, on this screen and on a URL's detail page; today is still in progress and is excluded. The bar's engine, tag, country, persona and language narrow the answers on both, and with an engine chosen only that engine's channel cards are offered (api/routers/source_retrieval.py#ACCEPTS, api/services/source_retrieval.py#source_retrieval_for_project). Failed and in-flight answers are excluded. Fewer than 30 answers marks a channel, or fewer than 30 returned answers marks a row, provisional (api/services/brand_metrics.py#MIN_OBSERVATIONS). That is a sample-size cue, not a significance test.

Source detail

Source detail showing the saved copy of a page, tracked brand matches and the page text

The detail page opens only for a URL this project's own completed answers returned or recorded as a citation. Any other address reads as not available (api/services/source_retrieval.py#source_retrieval_url_for_project).

Saved copy is the page as our scraper last saved it: fetch time, title, language, word count, canonical URL, the publish and modified dates the page declares, and the full saved text, up to 100,000 characters (api/services/source_mentions.py#MAX_TEXT_CHARACTERS). A copy saved by an earlier successful fetch stays visible, with its fetch time, even if a later fetch failed. It may differ from the page an engine read when it answered. When there is no copy, the page says whether the URL is waiting to be fetched, could not be fetched, had no readable text, or has not been queued (api/services/source_retrieval.py#_saved_page). Opening the page does not fetch it. The publish and update dates an engine reported for the page are shown separately, from the most recent answer that reported any among the answers the engine, tag, country, persona and language chips select, of any date.

Tracked brands in the saved text looks for your project's active brand and competitor names and aliases in that text, and shows how many brands were checked, so an empty result is not read as a verdict when nothing was checked. Only the saved portion is checked.

Answers that returned or attributed this page lists each answer in the window and in the channel you came from (every channel when none is selected), newest first, 25 per page, with its prompt, date, engine, returned position and whether it attributed the page. Up to three recorded passages, each up to 500 characters, are shown for an attributed answer (api/services/source_retrieval.py#_passages). For Grok the passage is the answer sentence that carries the citation, with link markup reduced to the link's text; for Claude it is the passage Claude quoted from the page. For Perplexity it is the answer text just before the numbered marker, back to the start of its sentence, its line or the previous marker, whichever is nearest, without the markers (api/services/llm.py#_sonar_marker_citations, api/services/collection.py#SPAN_IS_CITED_TEXT); a marker with no text between it and the previous marker or the start of its line still counts as attributed, with no passage to show. For a Gemini (API) answer attribution reads not compared, for a Google AI Overviews, AI Mode, ChatGPT (app) or Gemini (app) answer Attribution not compared: only cited sources are reported, for an answer from Perplexity's earlier Search API not applicable, and for an unrecognized surface unknown.

Limits

This is an observation of collected answers. It does not show every page an engine read or why an engine chose a source. The engines it compares are queried through their APIs, so it does not show what a person would see in the engine's own app; ChatGPT (app) and Gemini (app) answers are what chatgpt.com and gemini.google.com show a person who is not signed in, but they are not compared: Gemini (app) reports no pages found, and ChatGPT (app)'s are stored but not yet compared. Values are live and can change as late answers arrive. The customer API and MCP do not include this dataset.

  • Citations: every source recorded for an answer.
  • Source comparison: citation rate changes between two periods.
  • Domains: source reach and consistency over the filter bar's window and filters.
  • Metrics defined: other citation and brand denominators.

Last verified 2026-09-29

Start monitoring your AI visibility.

See how AI search engines talk about your brand.

Free to start. No credit card required.