Your page was retrieved but not cited: what to check next
An AI engine can return your page and still credit someone else. Learn what returned versus attributed means, when you cannot tell, and the checks to run in order.
On this page
- In short
- What does "retrieved but not cited" actually mean?
- Which engines let you see returned versus attributed sources?
- How do you handle a channel that reports citations only?
- A decision tree for a page that was returned but not attributed
- Step 1 to 2: confirm the observation is real
- Step 3: can the engine's side reach and read the page?
- Step 4: does the page answer the prompt as asked?
- Step 5: does the page hold quotable evidence?
- A worked example
- What this cannot tell you
- Frequently asked questions
- Next step
If an AI engine returned your page while it searched but the answer credited someone else, your page was in front of the engine and still did not earn the citation. The cause is not knowable from the outside, so you work through checks in order: first confirm the engine actually reports both sides, then check whether the page is reachable, whether it answers the question the prompt asked, and whether it holds a passage an answer could credit. For several engines you cannot run this test at all, because they report only the sources they cite or ground on. In that case the honest result is "unknown", not "retrieved but not cited".
In short
- Returned means the engine reported a page among the results its search brought back. Attributed means the answer text credited it. A page can be returned without being attributed.
- Only some collection channels report both. Where a channel reports cited or grounding sources only, "retrieved but not cited" is unknown, and you should say so.
- Check in this order: channel, sample size, page access, relevance to the prompt, then evidence in the text.
- A gap between returned and attributed is an observation, not an explanation. Every fix below is a hypothesis to test with a recheck.
What does "retrieved but not cited" actually mean?
It means two separate observations disagree: the engine's own report says your page came back from its search, and the answer text does not credit that page. Definitions matter here, so this post uses these terms throughout.
- Returned: the engine reported the source among the pages its search or grounding brought back for that answer.
- Attributed: the answer text credited the source, and a passage or text position was recorded with that credit.
- Citation: your domain linked as a source for an answer. A related term in key terms is retrieval: an engine pulling one of your pages into its context without linking it in the answer.
A source counts once per answer, however many times it appears in it, and several URLs from one domain in one answer count once for the domain. Returned pages are also only what the engine reported. They are not a complete list of every page it read internally.
Returned says the engine's search found you; attributed says the answer used you. For the wider picture of how mentions, citations and retrieval differ as metrics, see AI visibility, brand mentions, and citations: what each number tells you.
Which engines let you see returned versus attributed sources?
Only three channels in DiscoveredBy can be compared: Claude, Perplexity and Grok. Every other channel either reports cited sources only, reports only grounding sources, or is not compared for a stated reason. The rates on the Retrieved vs cited screen are shown one collection channel at a time, because what an engine reports as returned is not the same across engines. Each channel is labelled with what its returned sources can show.
| Channel | What its returned sources show | Compared? |
|---|---|---|
| Claude | Full search results: the results its search tool returned to it | Yes |
| Perplexity (Sonar) | Full search results: every source it retrieved, with the ones the answer used marked | Yes |
| Grok | Partial: pages it opened and sources it listed; individual search results are not exposed | Yes |
| Gemini (API) | Grounding sources only, and its grounding links credit nearly all of them whenever it sends any | No, listed with its number of answers |
| Google AI Overviews | Cited sources only | No, listed with its number of answers |
| Google AI Mode | Cited sources only | No, listed with its number of answers |
| Gemini (app) | Cited sources only | No, listed with its number of answers |
| ChatGPT (app) | Cited sources only on this screen: the pages it found since 2026-09-29 are stored, but its citations carry no positions for attribution, so it is not compared yet | No, listed with its number of answers |
Perplexity answers from its earlier Search API are not compared (they hold ranked results and no answer to attribute), and an unrecognised surface shows "Unknown capture". For Claude, every returned search result is recorded as a citation, so the Citations page can show more than the answer credited; this screen counts only sources the answer text credits.
How do you handle a channel that reports citations only?
Record the returned side as unknown, and do not call it "retrieved but not cited". Google AI Overviews, AI Mode and Gemini (app) report no list of pages they searched, so they cannot show one that was searched and skipped; their fan-out reads not applicable. Gemini (API) reports grounding sources only, and ChatGPT (app) stores the pages it found but is not compared yet. The answer is the same for all of them: unknown.
What you can still say is limited but useful: whether your domain was cited, and which competing sources were. That is a citation gap question, not a retrieval one. Write: "This channel is not compared, so we cannot tell whether our page was returned."
On comparable channels, watch sample size. A channel is provisional under 30 answers and a row under 30 returned answers. That is a sample-size cue, not a significance test.
A decision tree for a page that was returned but not attributed
Work through these questions in order and stop at the first branch that fits. The diagram below is the visual version; the text version follows and is the one to copy into your notes.
The text version
START: one prompt, one channel, your page returned but not attributed.
1. Does this channel report both returned and attributed sources?
(Comparable today: Claude, Perplexity, Grok.)
- No -> UNKNOWN. Cited sources only, grounding only, or not compared.
Record "cannot tell whether the page was returned". Stop.
- Yes -> go to 2.
2. Is the sample large enough to read?
(Channel: 30 or more answers. Row: 30 or more returned answers.)
- No -> PROVISIONAL. Note it, collect more answers, change nothing yet. Stop.
- Yes -> go to 3.
3. Can the page be reached and read?
Open the source detail: is there a saved copy? Run page diagnostics on the
exact URL: robots policy, HTTP status, complete HTML, body text, indexing
directives, content that only appears after JavaScript.
- No -> ACCESS PROBLEM. Fix it, publish, recheck the same URL. Stop.
- Yes -> go to 4.
4. Does the page answer the prompt as asked?
Read the saved text against the prompt wording, not against your keyword.
- No -> RELEVANCE GAP. Choose: update this page, write a page for that
question, or accept that another kind of page fits. Stop.
- Yes -> go to 5.
5. Compare passages. Does your page hold a specific, checkable passage that
answers the question, comparable to the passage the answers credited on a
cited page?
- No -> EVIDENCE GAP. Add specifics a reader could verify, then recheck.
- Yes -> go to 6.
6. Do the checks so far point to one plausible cause?
- Yes -> Make one change, note the date, recheck later on the same prompts.
- No -> NO CLEAR CAUSE. Record the observation and keep monitoring.
Do not invent a reason.
Step 1 to 2: confirm the observation is real
Open the Retrieved vs cited screen, pick a comparable channel, and set the date range to 7, 28 or 90 days (it ends yesterday). Failed and in-flight answers are excluded.
For your domain the screen shows Returned in (answers that returned your domain, divided by all completed answers in the channel and window), Attributed in (answers that attributed it, over the same answers) and Attributed when returned (answers that did both, divided by answers that returned it; a dash if none did). The source table's Returned, not attributed column covers the answers where the engine had the page and did not credit it. Values are live and can change as late answers arrive.
Step 3: can the engine's side reach and read the page?
Open the page's source detail from the table. The Saved copy is the page as DiscoveredBy's scraper last saved it, with fetch time, title, word count, canonical URL and declared dates. When there is no copy, the page says whether the URL is waiting to be fetched, could not be fetched, had no readable text, or has not been queued. That message is your first access clue.
Then run page diagnostics on the exact URL. Each check returns Check passed, Needs review or Unknown. The relevant ones are declared crawler policies (robots.txt rules for AI crawlers), HTTP delivery (status and final destination), complete HTML and body text (is there extractable text?), indexing directives (noindex or none), and the JavaScript checks (does rendering add substantial text or change the title, H1 or robots directives?).
Read these as a lab observation, not a verdict. They come from one request and one browser load as DiscoveredByBot, do not simulate another crawler, and are not a citation prediction or a grade. A clean result only removes one candidate cause. Running them depends on your plan (see plans and limits). For the same checks across a crawl, see the site audit and An AI search site audit: which issues should you fix first?.
Step 4: does the page answer the prompt as asked?
Read the saved text next to the prompt itself. The prompt is a question you chose to monitor, not a keyword, so the test is whether your page answers it the way a buyer phrased it.
- Does the opening address the same question, or a neighbouring one?
- Does it name the specifics the prompt asks for, such as a comparison, a limit or a step?
- Is it the right kind of page? A product page may be returned for a "how do I" prompt that a guide would fit better.
The source detail also shows Tracked brands in the saved text, with a count of brands checked, so an empty result is not read as a verdict when nothing was checked. If the page is off-question, the choice between updating it and writing a new one is covered in Find the citation gaps that deserve your next content update, and handing the work to a writer in Turn an AI citation gap into a content brief your writer can use.
Step 5: does the page hold quotable evidence?
For an attributed answer, the source detail lists up to three recorded passages, each up to 500 characters. What counts as the passage differs by engine:
- Claude: the passage Claude quoted from the page.
- Grok: the answer sentence that carries the citation.
- Perplexity: the answer text just before the numbered marker, back to the start of its sentence, its line or the previous marker.
Place a credited passage from a cited page beside your own saved text. Is the cited page saying something specific and checkable that yours leaves vague, such as a stated limit, a defined term, a dated figure with its source, or a step sequence? These are hypotheses about what an answer could use, not established citation factors. The saved copy may also differ from the page the engine read when it answered.
A worked example
Illustrative example: Quillstone and its competitors are fictional, and the numbers are made up to show the method.
Quillstone sells document-review software to legal and compliance teams. Its team monitors the prompt "How do I review a contract for compliance gaps before signing?" The Claude channel shows, for a 28-day window:
| Measure | Count | Rate |
|---|---|---|
| Completed answers in the Claude channel | 120 | |
| Answers that returned quillstone.example | 36 | 36 of 120 = 30% (Returned in) |
| Answers that attributed quillstone.example | 9 | 9 of 120 = 7.5% (Attributed in) |
| Attributed when returned | 9 of 36 | 25% |
Both the channel (120 answers) and the domain row (36 returned answers) are above the 30-answer mark, so nothing is provisional. On the same prompt set, a competitor page from Brieflane was returned in 30 answers and attributed in 24.
The team also tracks the prompt on ChatGPT (app), which has 50 completed answers. That channel reads "Cited sources only", so the screen lists its answer count and does not compare it. The team writes down: "Cannot tell whether our page was returned on ChatGPT (app)." They do not run the tree for that channel.
Back on Claude, they open the URL quillstone.example/guides/contract-review-checklist, returned in 24 of those 36 answers and attributed in 3.
- Channel: Claude, comparable. Continue.
- Sample: 120 answers in the channel and 36 returned answers for the domain, both at least 30. Continue.
- Access: the saved copy exists. Page diagnostics passes the crawler policy, HTTP delivery, and body text checks. Nothing to fix.
- Relevance: the guide opens with a general history of contract review and reaches the compliance checks in the seventh section. The prompt asks how to review for compliance gaps before signing. Partial fit.
- Evidence: the recorded Claude passage from the Brieflane page states the specific checks and the order to run them. Quillstone's guide describes the checks in general terms, with no order and no example clause.
- Cause: two plausible candidates (question fit and specificity), both hypotheses.
The team changes one thing: it restructures the guide so the compliance checklist opens the page, with the checks in order and a worked clause example. They record the date and the exact prompt, and schedule a recheck of page diagnostics and of Retrieved vs cited for the same channel and window. If "Attributed when returned" rises, they note an association with the edit. If it does not move, that is a valid result too, and possibly a window with too few new answers to read.
What this cannot tell you
- Why the engine chose another source. The screen shows collected answers, not every page an engine read or its reasons.
- What a person sees in an engine's own app. The engines it compares are queried through their APIs.
- Whether a fix worked. A recheck shows an observed response change, not proof that an edit caused it.
- A bot's real behaviour. Diagnostics come from DiscoveredByBot, not from any engine's crawler.
Common mistakes: treating "unknown" as "not retrieved"; diagnosing from a provisional rate; comparing a Claude rate with a Perplexity rate as if "returned" meant the same thing; changing several things at once; and reading a passing access check as an explanation for the missing citation.
Frequently asked questions
What is the difference between a page being retrieved and being cited?
A retrieved (returned) page is one the engine reported among the results its search or grounding brought back. A cited (attributed) page is one the answer text credited, with a recorded passage or position.
Why can't I see whether my page was retrieved in Google AI Overviews?
Google AI Overviews and AI Mode report only the sources their answer cites, with no list of pages searched. The same holds for Gemini (app) and, currently, ChatGPT (app). Those channels are listed with their answer count and not compared.
Does a low "attributed when returned" figure mean my page is bad?
No. It is an observation that answers returned your page and did not credit it. With fewer than 30 returned answers it is provisional, and even a solid sample does not reveal the engine's reasons.
Can page diagnostics tell me why a page was not cited?
No. They report structure, response and render observations from one lab request as DiscoveredByBot. They can rule out access and delivery problems, which narrows where to look next.
If I fix a page, how do I know it helped?
Recheck the same exact URL, then compare the same prompts, channel and window on Retrieved vs cited. "Resolved on recheck" means a flagged check now passes; it does not prove the edit changed an engine's behaviour. Improvement, no change and too little evidence are all valid outcomes.
Next step
Open Retrieved vs cited in your project, choose a channel that can be compared, and look at the Returned, not attributed column for your own domain. Pick one URL, run the tree above, and log the result. Read the Retrieved vs cited and page diagnostics documentation for the full detail on each screen, or see Citations and Metrics defined for how these numbers relate to the rest of your reporting. To start, sign in or create a project.
- citations
- retrieved vs cited
- page diagnostics
- content optimization