AI crawler logs: how to tell an access problem from a citation problem

Read your logs for AI bot requests, verify who sent them, check what your site served, then compare with cited pages. A copyable checklist, and what a bot request cannot prove.

Kamal 17 min read
A long paper ledger with one line highlighted in coral, and a separate small card beside it holding a single stamped page, suggesting a visit record kept apart from a credit.
On this page
  1. In short
  2. What is an AI crawler log, and what can it tell you?
  3. Step 1: which bots are asking for which pages?
  4. Step 2: did the request really come from the bot?
  5. Step 3: what did your site serve?
  6. Step 4: only now, compare with what AI answers cite
  7. How do I read fetch status and citation together?
  8. The log-review checklist
  9. A worked example
  10. What this method cannot tell you
  11. Frequently asked questions
  12. Next step

To tell an access problem from a citation problem, read your server or CDN logs in four steps: find which AI bots requested each page, check that the requests really came from the bot's operator, look at the HTTP status your site served, and only then compare the pages that were fetched with the pages AI answers cite. A page that bots cannot fetch has an access problem. A page bots fetch fine that never gets cited has a different problem, and the logs cannot say what it is. A bot request is a record of a request, not proof that the page was read, trained on, or used in any particular answer.

In short

  • Two questions, two data sources. Logs answer "could the bots get the page?" Collected AI answers answer "did an engine credit it?" Keep them apart until the last step.
  • A user agent is only a claim. Verify the bot before you count its requests or block it.
  • The status your site served is the fastest access signal: a 4xx or 5xx to a verified bot is a lead worth chasing first.
  • A fetch next to a citation shows the two happened in the same window. It does not show that one caused the other.
  • Missing from the logs is not the same as blocked. Check the log source's coverage before you conclude anything.

What is an AI crawler log, and what can it tell you?

An AI crawler log is the subset of your web server or CDN access log where the user agent names an AI or search bot, such as GPTBot, OAI-SearchBot, ClaudeBot or PerplexityBot. Each line records a time, a path, the HTTP status your site served and, usually, the client address. It tells you what was requested and how your site responded. It does not tell you what the operator did with the page afterwards.

That distinction matters because robots.txt and logs answer different questions. A robots.txt check reports what your file declares for each crawler. The log shows what actually arrived and what your site served in reply. The two can disagree: a file can allow a bot that a firewall rule turns away, or block a bot that keeps knocking anyway.

DiscoveredBy's AI crawlers page reads your own logs, sent through a log source you connect (a webhook, a Cloudflare Worker, Cloudflare Logpush or an uploaded log file). It lists which bots requested your pages, whether each request was verified, what your site served, and which fetched pages are cited in your AI answers. You can also do everything below by hand with grep and a spreadsheet. The checklist is tool-neutral.

Bots also come in different kinds. Following the docs' grouping, training crawlers collect pages for model training, search crawlers index them for AI search, user-triggered fetches happen when someone asks an assistant about a page, and other covers the rest. Read them separately. A search bot's failure and a training bot's absence mean different things, and blocking one on purpose is a legitimate policy.

The logs cover the first three steps. Citation is a separate record.

Step 1: which bots are asking for which pages?

Start by listing, per path, the distinct bots that requested it, how many requests each made, and when the latest was. Pages nobody asked for, and pages one bot asked for once, are your first contrast with the pages that get steady attention.

Group by intent before you group by page. A search bot fetching your pricing page is a different signal from a training crawler doing a sweep of your blog. Also fix the time window, so that two reviews compare like for like. (DiscoveredBy's AI crawlers page offers 7, 30 or 90 days, and keeps individual requests for 30 days and daily totals for 400.)

Two cautions belong here. First, bots can be dormant for reasons that have nothing to do with you; a quiet week is not a finding. Second, a log source only sees what you send it. If your CDN serves a cached copy and the request never reaches your origin, an origin log will not show it. Confirm where in your stack the log is taken before you read absence as a signal.

Step 2: did the request really come from the bot?

Verification means checking the request's address against the ranges the bot's operator publishes, because a user agent string can be typed by anyone. The AI crawlers docs sort each stored request into three outcomes: verified (inside the operator's published ranges), spoofed (ranges held, address outside them) and unverified (the address could not be checked).

Unverified has four causes in the docs: no client address on the line, an invalid address, no published ranges held for that bot, or the operator's ranges not yet loaded. For some bots (the docs name Meta, ByteDance, Amazon and one Mistral crawler) no ranges are held, so a request claiming to be one of them can never be shown as spoofed. Do not read unverified as suspicious or as safe. It means unchecked.

Why this comes second: everything after depends on it. A "GPTBot" that gets a 403 may be an impostor your firewall was right to refuse. Chasing that as an access problem wastes a sprint. Count only verified requests when you ask whether a real operator could reach a page, and keep the spoofed ones as a separate note for whoever owns security. For a bot whose ranges cannot be checked, verified counts will never exist, so its status codes are weaker evidence: say so in your notes.

If you verify by hand, use the operator's own documentation for its published address ranges, and refresh them, since ranges change. The AI crawlers docs describe refetching them daily.

Step 3: what did your site serve?

The HTTP status recorded for a verified bot request is the plainest evidence of access. A 200 means your site returned the page to that request. Redirects mean the bot was sent elsewhere, so check the destination. A 4xx or 5xx means it was not served the page.

Read the codes cautiously, using standard HTTP meanings only:

Status What it says What the log cannot tell you
200 The request was answered with content Whether the content matched what a browser shows (see below)
301 or 302 The bot was redirected Whether it followed, and whether the destination is the page you intended
401 or 403 The request was refused Which rule refused it: your server, CDN, firewall or bot protection
304 The bot's saved copy was reported unchanged Whether the bot ever got a usable copy
404 or 410 Nothing is served at that path Whether the path is a stale link, a typo, or a page you removed
429 The client was asked to slow down The rate limit that applied
5xx Your site failed to answer Whether the failure was brief or repeated

Two traps sit behind a healthy 200. A page can return 200 and still be nearly empty to a bot that runs no JavaScript, because content added by scripts never appears in the raw HTML. The docs' site audit compares raw HTML with a JavaScript-rendered load for exactly this reason, and says plainly that it checks as DiscoveredByBot, not as the AI crawlers. And a bot can get a different response from the one a browser gets, so a page that looks fine to you may not be what the bot received.

The AI crawlers page pulls these out for you: an Error responses headline count of requests answered with a 4xx or 5xx, and a Pages failing for bots list of paths whose latest response to a bot was a 4xx or 5xx. The two differ on purpose. A path that failed once and later succeeded adds to the first and drops out of the second. Error responses counts stored requests of any verification result, so it can include spoofed or unverified ones; the page can also be narrowed by verification result.

Step 4: only now, compare with what AI answers cite

A page that verified bots fetched successfully and that no collected answer cites is a citation question, not an access question. A page that bots cannot fetch is an access question first, whatever the citation data shows.

In the AI crawlers page, a Cited mark on a Top pages row means an answer collected in the same window cited a URL on your domain (subdomains count) with a matching path, after removing the query string, fragment and trailing slash. Three limits apply, all from the docs:

  • The citations are the ones the Citations page counts, which is every source recorded for an answer. For Claude, every search result it returned is recorded as a citation, so a Claude result its answer did not credit can still mark a page as Cited. See Retrieved vs cited for the distinction.
  • The mark sits on Top pages rows, which list the 100 paths bots requested most. A page that is cited but is not among them (no bot requests, or too few to make the top 100) has no row, so look for it on the Citations page instead.
  • A crawler fetch next to a citation shows the two happened in the same window. It does not show that the fetch caused the citation.

That last point is the boundary of the whole exercise. A bot fetch can be for search indexing, training, or one person's question to an assistant. Logs cannot tell you that a page was read, kept, or drawn on for a specific answer, and an AI engine does not have to fetch your page at the moment of the answer to mention you. Treat any fetch-then-citation pairing as a lead to follow, not a mechanism to report. For the timing question, see Are new AI bot visits followed by new citations to a page?

How do I read fetch status and citation together?

Put each important page into one of six cells, using verified requests only. The cell suggests which team owns the next move.

Verified bot fetches, latest status Cited in answers Reading Next move
Served 200 Yes Reachable and credited Keep it stable; watch for regressions after site changes
Served 200 No Reachable, not credited Citation problem: see what to check when a page is retrieved but not cited
Served 4xx or 5xx No Access problem Find the rule or fault serving the error, then recheck
Served 4xx or 5xx Yes Credited despite errors Still fix the errors; the citation may rest on earlier or other retrieval
No verified fetches No Unknown Check log coverage and robots policy before concluding anything
No verified fetches Yes Cited without a fetch in this window Not evidence of a problem; the engine may rely on earlier or other retrieval

A redirect as the latest status belongs in none of these cells until you have checked where it leads.

The log-review checklist

Copy this into your team's tracker and complete one block per priority page.

AI CRAWLER LOG REVIEW
Site / project:            ________   Reviewer: ________   Date: ________
Window reviewed (UTC):     from ________ to ________
Log source(s):             ________ (origin / CDN / uploaded file)
Coverage check:            [ ] Log includes cached responses served at the edge
                           [ ] Log covers the whole window with no gaps
                           [ ] Host filter matches our domain and subdomains

PER PAGE (repeat)
Path:                      ________
Why it matters:           ________ (pricing, integration, comparison, guide...)
1. Bots seen:              name / intent (training, search, user-triggered) / requests / last seen
                           ________
2. Verification:           verified ___  spoofed ___  unverified ___
                           Unverified cause, if known: no IP / no published ranges / other
3. Status served to verified bots: latest ___   any 4xx/5xx in window? Y / N  (which ___)
   Redirect destination checked?  Y / N / n.a.
4. Raw HTML check:         does the raw HTML contain the key content without JavaScript? Y / N / not checked
5. Declared policy:        robots.txt rule for each bot on this path: allow / block / unreadable
   Matches intent?         Y / N (a deliberate block is fine; record it as deliberate)
6. Cited in collected answers this window (same path, normalised)?  Y / N
   Engines / prompts:      ________

DIAGNOSIS (pick one)
[ ] Access problem     -> owner: ________  action: ________  recheck date: ________
[ ] Citation problem   -> owner: ________  action: ________  recheck date: ________
[ ] No evidence        -> what would settle it: ________
[ ] Fine               -> no action

CLAIM STRENGTH (write the weakest that fits)
[ ] Observed: bot X requested path Y and was served status Z
[ ] Correlated: a fetch and a citation share a window
[ ] Not claimed: that the page was read, trained on, or used in an answer
Demo data. Top pages with Cited marks, beside the list of pages failing for bots.

A worked example

Illustrative example: Quillstone and its competitors are fictional, and the numbers are made up to show the method.

Quillstone sells document-review software to legal and compliance teams. Its content lead asks why the SharePoint integration page never appears as a source in answers, while the pricing page does. She reviews 30 days of logs from a Cloudflare feed.

Across the month, the log shows 100 requests from catalogued AI bots:

Bot Intent Requests Verified Spoofed Unverified
GPTBot Training 40 36 4 0
OAI-SearchBot Search 30 30 0 0
ClaudeBot Training 20 18 0 2
PerplexityBot Search 10 10 0 0
Total 100 94 4 2

Step 2 removes noise. Four GPTBot requests came from outside OpenAI's published ranges. She notes them for the security lead and drops them from the access question. Two ClaudeBot lines had no client address, so she leaves them as unchecked.

Step 3 and step 4 by page, using the 94 verified requests only:

Path Verified requests Bots Latest status Cited
/pricing 32 3 200 Yes
/guides/contract-review-checklist 24 2 200 No
/integrations/sharepoint 17 2 503 No
/compare/brieflane 12 2 403 No
/blog/redaction-basics 9 1 200 Yes
Total 94

Reading it cell by cell:

  • /integrations/sharepoint returned a 503 as its latest response, and 13 of its 17 verified requests got a 5xx. That is an access problem. The cited-or-not column is irrelevant until the page reliably answers. Suppose her developer traces it to a backend call timing out under load. Once it is fixed she sets a recheck date.
  • /compare/brieflane served a 403 to 6 of its 12 verified requests, and the latest was a 403. Whether that is a firewall rule aimed at comparison pages or a bot-protection setting, the log cannot say. She files a ticket to find which rule refused verified requests, and records that nobody yet knows whether the block is deliberate.
  • /guides/contract-review-checklist was fetched 24 times and served a 200 every time, yet is not cited. That is a citation question. She opens it in the retrieved-versus-cited review and checks whether the raw HTML actually contains the checklist or builds it with scripts.
  • /pricing and /blog/redaction-basics are reachable and cited. She writes "observed: fetched, served 200, cited in the same window" and does not claim the fetches earned the citations.

Across the five pages, 19 of the 94 verified requests got a 4xx or 5xx (13 from SharePoint plus 6 from the comparison page). Her three findings are an access fix, an access problem with an unknown cause that needs a ticket, and a citation investigation. None required guessing about model internals.

What this method cannot tell you

  • That a page was read or used. The log records a request and your response. It does not show what the operator did next. Do not report "GPTBot trained on our page" or "this fetch produced that citation".
  • That absence means blocked. Missing lines can mean the bot has not come, the request was served from cache and never logged, the log source dropped it, or a limit applied. The AI crawlers docs list several ways lines are ignored, rejected or counted over an allowance, so check the counts before drawing conclusions.
  • Which bots you can verify. Some operators publish no address ranges, so those requests remain unverified whatever their user agent says.
  • Bots the tool does not recognise. DiscoveredBy matches only the bots in its catalogue; a line whose user agent names no catalogued bot is counted as ignored and nothing else about it is kept.
  • What every engine does. Do not carry one crawler's behaviour over to the others. Read each bot's own operator documentation for its purpose, and keep search, training and user-triggered fetches apart.
  • What your robots.txt does in practice. The robots diagnostics reading is a declared policy check. It says allowing a crawler does not establish actual access through a firewall, JavaScript rendering, visits, indexing or citations. Use it alongside the log, not instead of it.

If your robots.txt and llms.txt strategy is the open question, our post on what robots.txt, llms.txt and structured content can actually help with covers those files. For the wider set of technical issues to weigh, see which audit issues to fix first.

A visit record and a credit are two different documents.

Frequently asked questions

How do I know if GPTBot is really crawling my site?

Check the user agent, then verify the address. A request that says "GPTBot" but comes from outside the ranges OpenAI publishes for that bot is spoofed. Only a request from inside those ranges counts as a verified visit. If your log has no client address, the request stays unverified.

Does a bot request mean my page will be cited?

No. A request shows that something asked for the page. Whether the page then appears as a source depends on the engine, the question and what it retrieves, and your logs cannot see that. Compare with collected answers instead, and describe any overlap as a coincidence in time until you have tested more.

Should I block AI crawlers I do not recognise?

First check whether the requests are verified. A spoofed request is an impostor, and a firewall refusing it is doing its job. For a bot whose operator publishes no ranges you cannot verify it either way, so treat it as unchecked. Whether to allow a genuine training or search crawler is a policy decision for you; a deliberate restriction can be appropriate, and search and training are separate controls.

Why does a page return 200 to bots but still look empty to them?

The raw HTML may lack the content that a browser adds with JavaScript. A crawler that runs no scripts sees only the raw response. Compare the raw HTML with the rendered page; the site audit does both for a page, though it checks as DiscoveredByBot and not as the AI crawlers.

How long should I look back in the logs?

Long enough to compare like for like, and no longer than your retention. The AI crawlers docs keep individual requests for 30 days and daily totals for 400. A 30-day window with a 7-day recheck after a fix is a practical rhythm.

What if my important page has no bot requests at all?

Do not conclude it is blocked. Check your log coverage first, then the declared policy for that path, then whether the page is linked and listed in your sitemap. If everything is reachable and no verified bot has requested it, note it as unknown and keep watching.

Next step

If you want the logs and the citations in one place, DiscoveredBy's AI crawlers page shows verified, spoofed and unverified requests, the status your site served, pages failing for bots, and a Cited mark next to each fetched page. Connect a log source, run the checklist above on your ten most important pages, and sign in to start. Availability and monthly request allowances depend on your plan; see plans and limits.

  • citations
  • robots.txt
  • technical geo
  • ai crawlers
  • server logs

Share

Summarize with AI

Start monitoring your AI visibility.

See how AI search engines talk about your brand.

Free to start. No credit card required.