Is your CDN or firewall blocking AI crawlers you meant to allow? A troubleshooting log

A robots.txt that allows AI crawlers does not mean your CDN or firewall does. Use status codes, verified identities and a troubleshooting log to find out, without disabling protections.

Kamal 16 min read
A row of gates in a wall, most open, one shut, with a small coral marker at the shut gate
On this page
  1. In short
  2. Why can robots.txt allow a crawler and your firewall still block it?
  3. What do the status codes tell you?
  4. How do you know a request really came from the bot?
  5. How do you read the AI crawlers page for blocks?
  6. How do you check one page from the outside?
  7. The deliverable: an access-troubleshooting log
  8. Worked example: Quillstone and its pricing page
  9. How do you fix a block without weakening protection?
  10. What can this method not tell you?
  11. Common mistakes
  12. Frequently asked questions
  13. Next step

If your robots.txt allows an AI crawler but your CDN or firewall still refuses its requests, the crawler never gets your page, and robots.txt will not show you that. The way to find out is to read what your site actually answered: the HTTP status for each request from each bot, and whether the request really came from the bot's operator. A verified bot receiving 403, 429 or a challenge page on a page you want public is a block you did not intend. A spoofed bot receiving the same is your protection working. This post gives you a method and a troubleshooting log to tell the two apart without turning protections off.

In short

  • Robots.txt is a request to well-behaved crawlers. Your CDN, web application firewall (WAF) or bot-management layer is a separate control that can refuse a request before your server or your robots.txt matters.
  • Read the status your site served to each bot, then split the results by verification: verified, spoofed or unverified. Only verified refusals are candidates for repair.
  • Fix by adding a narrow, identity-based exception for the specific bot and path, never by switching the protection off.
  • Recheck after each change and log it, so you can tell a fix from a coincidence.

Why can robots.txt allow a crawler and your firewall still block it?

Robots.txt and a firewall answer different questions. Robots.txt declares which paths you would like a crawler to fetch. A CDN, WAF or bot-management rule decides whether a request is served at all, usually from signals such as request rate, IP reputation, a missing header or a user agent it treats as automation.

A WAF (web application firewall) is a rule layer in front of your site that filters requests, and bot management is the part of a CDN or WAF that tries to separate automated traffic from people. Neither reads your robots.txt to decide. So a site can have a perfectly permissive robots.txt and still return errors to crawlers you wanted, because a default rule set, a rate limit or a "block known bots" toggle was enabled for another reason.

This is why the declared policy and the observed access need to be checked separately. DiscoveredBy's Robots diagnostics shows what your robots.txt declares for AI crawlers; the AI crawlers page shows what your own logs say happened. If you are about to change robots.txt itself, review the change before release first; this post starts where that one ends, at the request that was already refused.

What do the status codes tell you?

The status code is your first clue to which layer answered. These are general HTTP meanings, and your vendor's behaviour will vary, so treat the last column as a hypothesis to confirm.

Status General meaning What to suspect in a bot log
200 Served Fine, but check the page body is the real page and not a challenge or block notice served with a 200
301 or 302 Redirect A redirect to a login, consent or challenge URL; follow it to its destination
401 Authentication required A rule or origin setting that expects credentials from a bot
403 Forbidden The classic firewall or bot-management refusal; also origin permission rules
404 Not found Not usually a block, but check that a geo or bot rule is not returning it deliberately
429 Too many requests A rate limit; the response may carry a Retry-After header saying when to try again
5xx Server or gateway error An origin overload, a timeout, or a protection that fails closed

Two cautions. First, some bot-management products answer suspected automation with a challenge or interstitial page, and the status that accompanies it depends on the product and its configuration. A crawler that does not solve the challenge sees a page that is not yours. Check your vendor's documentation for what it returns and for how to log it.

Second, the request may never reach the place you are logging. Depending on where a log is taken, a request refused at the edge may not appear in an origin log at all. Know which layer each of your log sources describes before you conclude that "no requests" means "no visits". For error pages and timeouts specifically, see prioritising the repair when AI crawler logs show errors.

How do you know a request really came from the bot?

A user agent is only a claim: anyone can send a request that says "GPTBot". So before you loosen any rule for a bot, establish whether the request actually came from the bot's operator.

The DiscoveredBy docs describe the check it applies. Each stored request is compared with the address ranges the bot's operator publishes, and then the address is dropped. The result is one of three labels (AI crawlers):

  • Verified: the request came from an address inside the operator's published ranges.
  • Spoofed: ranges are held for that bot and the address is outside them. It claims the bot but did not come from its operator.
  • Unverified: the address could not be checked. The line carried no client address, the address was not valid, no published ranges are held for that bot, or the ranges had not loaded yet.

That third label matters for troubleshooting. For some bots the operator publishes no machine-readable address list, so a request claiming to be one of them can never be shown as spoofed, and a refusal of it cannot be classified either way from this data alone. Do not treat unverified refusals as proof of a problem, and do not treat them as proof that everything is fine.

Outside a monitoring tool, the same principle applies: use the operator's own published guidance for verifying its crawler, such as its documented IP ranges or reverse DNS instructions, and confirm it from the operator's current documentation rather than from memory.

Verification decides whether a refusal is a problem or a protection working.

How do you read the AI crawlers page for blocks?

Start with the numbers that summarise failure, then move to the rows that name the culprit. The AI crawlers page reads your own server or CDN logs, sent through a log source you set up, and shows requests per bot with the status your site served (AI crawlers).

  1. Set the window. Pick 7, 30 or 90 days. Use a window that includes a period before the suspected change if you can.
  2. Read Error responses. This headline count is requests answered with a 4xx or 5xx status. It is a starting point, not a diagnosis, and it does not count a challenge page served with a 200 or a redirect to one.
  3. Filter by verification result. Narrow to verified requests first. Those are the bots you most likely want served.
  4. Open Pages failing for bots. It lists paths whose latest response to a bot was a 4xx or 5xx, most requested first, up to 100. A path that failed once and later succeeded adds to Error responses but does not appear here, so the two counts will not match.
  5. Read Recent requests. The 100 newest requests show time, bot, path, status and verification result, which is where you spot a pattern such as one bot failing on every path.
  6. Check the Bots table. It gives each bot's verified, spoofed and unverified split across all of its requests, not only the failed ones, so you can see whether a bot you are chasing is mostly genuine or mostly imitation.

Two limits to keep in view. The page only sees traffic that reaches the source you connected, and it only recognises catalogued bots. And a Cited mark on a page in Top pages shows the page was cited in the same window, not that the fetch caused it. A blocked page can still be cited from elsewhere, and a cited page can be fetched fine.

How do you check one page from the outside?

Once a log points to a path, request that exact URL yourself. Saved page diagnostics fetches a URL on your project's domain and records the HTTP status returned, the final destination after redirects, and whether complete HTML was readable, with history and a recheck comparison.

The limit is stated plainly in the docs: the request is made as DiscoveredByBot, so the delivery check "does not simulate another crawler or establish WAF access for it". That makes it useful for two things: confirming that your page is reachable to a well-identified lab request at all, and showing the redirects and delivered HTML. It cannot tell you whether GPTBot or PerplexityBot is refused by a rule that targets them. For that you need the logs.

A site audit runs the same checks across a crawl, and it is explicit about its own scope: it does not record whether real AI crawlers actually visited those pages. Use it to find pages that misbehave for everyone, and the logs to find pages that misbehave for specific bots.

The deliverable: an access-troubleshooting log

Keep one row per suspected block. The point of the log is discipline: you record evidence before you change a rule, and you record the recheck after. Copy this into a spreadsheet or ticket.

ACCESS TROUBLESHOOTING LOG

Row ID:
Date opened / owner:
Page or path pattern:
Is this page meant to be public and fetchable? (yes / no / unsure)

1. OBSERVATION
   Bot name (as it appears in logs):
   Verification result: verified / spoofed / unverified
   Status served: 403 / 429 / 5xx / challenge / redirect / other
   Number of requests and date window:
   Log source and layer it describes (edge, CDN, origin):
   Other bots hitting the same path with a 200? (yes / no)

2. DECLARED POLICY
   Does robots.txt allow this bot on this path? (yes / no / not checked)
   Checked on (date):

3. HYPOTHESIS (one at a time)
   Suspected layer: CDN rule / WAF managed rules / bot management / rate limit / origin server / other
   Suspected rule or setting (from the vendor's logs or docs):
   What evidence supports it:

4. CHANGE (narrowest possible)
   Exception scope: this bot + this path + verified identity only
   Protection left on for everything else? (yes)
   Change made by / date / ticket:
   Rollback step:

5. RECHECK
   New requests after the change (date window):
   Status now served to the verified bot:
   Spoofed requests still refused? (yes / no)
   Outcome: resolved / not resolved / cannot tell
   Next step:

Use one hypothesis per row. If you change three rules at once and the bot gets through, you will not know which mattered, and you will have widened your exposure for no reason.

A decision table for what to do next

Verification Status Page meant to be public? Action
Verified 403, 429 or challenge Yes Investigate the rule; add a narrow exception for that bot and path
Verified 5xx Yes Look at origin capacity and timeouts before touching security rules
Spoofed Any refusal Any Leave it; the protection is doing its job
Unverified Any refusal Yes Verify through the operator's documentation before changing anything
Any 200 with unexpected content Yes Compare the body served to the bot with the real page
Any Refused No (private, staging, admin) Leave it; that is the intended outcome

Worked example: Quillstone and its pricing page

Illustrative example: Quillstone and its competitors are fictional, and the numbers are made up to show the method.

Quillstone sells document-review software to legal and compliance teams. Its robots.txt allows OAI-SearchBot, PerplexityBot and ClaudeBot everywhere. A content lead notices that AI answers about "document review software" describe rival Brieflane's pricing but not Quillstone's, and asks whether the pricing page is even reachable.

In a 30-day window the AI crawlers page shows 200 bot requests and 24 error responses. Pages failing for bots lists /pricing and some blog paths. The page has no filter by status, so the team tallies the 24 error requests by bot, path and verification result from Recent requests and its raw logs, and gets this:

Bot Verification Status Path Requests
OAI-SearchBot Verified 403 /pricing 9
PerplexityBot Verified 429 /blog/ pages 7
GPTBot Spoofed 403 /pricing 5
Bytespider Unverified 403 /pricing 3
Total 24

Reading it row by row: the 5 spoofed GPTBot requests are refused, which is the WAF doing its job, so no action. The 3 Bytespider requests are unverified because no published ranges are held for that bot, so Quillstone cannot classify them and leaves the rule as it is. The 9 verified OAI-SearchBot refusals on a page Quillstone wants public are the problem to investigate. The 7 verified PerplexityBot 429s look like a rate limit, a different rule with a different fix.

Two log rows are opened. Row 1 hypothesises that a managed WAF rule is refusing OAI-SearchBot on /pricing. The team confirms the rule name in its CDN's own security event log, adds an exception scoped to that bot's verified identity and that path, leaves the rule on for everything else, and notes a rollback step. Row 2 hypothesises that the blog rate limit is too tight for a crawler that fetches several pages at once, and raises the threshold for that path only.

A week later the recheck shows the verified OAI-SearchBot requests to /pricing served with 200 and the spoofed GPTBot requests still refused. Quillstone records the outcome and does not claim that the fix will change what AI answers say: it now knows the bot can fetch the page, which was a precondition, not a guarantee.

Demo data. Pages failing for bots beside the verification split points at the rows worth investigating.

How do you fix a block without weakening protection?

Make the smallest change that lets the verified bot through, and leave everything else as it was. The aim is a targeted exception, not a lowered guard.

  • Scope it by identity, not only by name. An exception based only on a user agent string can be used by anyone who copies it. Where your vendor supports it, key the exception to the operator's verified identity or published ranges, as well as the path.
  • Scope it by path. Allow the bot on the public content you want fetched, not on the whole site, and keep admin, staging and account areas closed.
  • Prefer tuning over disabling. If a rate limit is the cause, adjust the threshold for that path rather than removing rate limiting.
  • Keep spoofed traffic refused. A working exception should not change what happens to imitation bots.
  • Write the rollback first. If the change misbehaves you should be able to undo it in minutes.

Broadly disabling a protection to "see if that fixes it" is the wrong test. It removes your defence against exactly the imitation traffic the verification labels warn you about, and if the bot then gets through you still do not know which rule was responsible.

What can this method not tell you?

  • It cannot prove why a rule fired. The log shows the status served, not the reasoning inside your vendor's system. Confirm the rule in your vendor's own event log.
  • It sees only what reaches your log source. A request refused before the layer you log at may be invisible. Know where each source sits.
  • It cannot generalise across bots. One bot getting a 200 says nothing about another; each operator and each bot can be treated differently by your rules.
  • It cannot tell you unverified traffic is genuine. For bots with no published ranges, the label stays unverified.
  • Access is not citation. Fixing a block makes a page fetchable. It does not promise that any engine will use, mention or cite the page. Treat later changes in citations as observations to compare, not as proof that the fix worked.
  • A lab fetch is not the crawler. Saved page diagnostics fetch as DiscoveredByBot and say so.

Common mistakes

  • Assuming a permissive robots.txt means access is open.
  • Reading "Error responses" as a list of problems without splitting by verification.
  • Loosening a rule for a user agent string that turns out to be spoofed.
  • Changing several rules at once and then not knowing which one mattered.
  • Treating unverified as either safe or hostile.
  • Forgetting the recheck, so the log never records whether the fix held.
  • Missing that pages loaded only after JavaScript runs are a separate visibility problem; see whether a crawler can see your JavaScript-heavy page.

Frequently asked questions

Can a CDN block AI crawlers even if robots.txt allows them?

Yes. Robots.txt is a declaration that well-behaved crawlers choose to follow; your CDN, WAF or bot-management layer decides separately whether each request is served. A default or managed rule can refuse a crawler your robots.txt welcomes. Check the status your site actually served, not only the declared policy.

Should I allowlist every AI crawler by user agent?

No. A user agent is only a claim, and an exception keyed to the name alone can be used by anyone who sends it. Allow only the bots you have decided you want, prefer verification by the operator's published ranges where your vendor supports it, and limit the exception to the paths that should be public.

Is a 403 to a bot always a mistake?

No. If the request is spoofed, or the page is private by design, a 403 is the intended outcome. It is a problem when a verified bot is refused on a page you want fetched. That is why the verification label and your intent for the page decide the action, not the status alone.

Does fixing a block make AI engines cite my page?

Not on its own. It removes one obstacle to the page being fetched. Whether an engine uses or cites the page depends on many things this data cannot show, and a crawler fetch next to a citation shows they happened in the same window, not that one caused the other.

Why do my logs show fewer bot requests than my CDN dashboard?

They may describe different layers, windows or filters. A monitoring source only sees what you send it, only recognises catalogued bots, and stores bot requests only up to your plan's monthly allowance; requests refused earlier in the chain may never reach it. Compare like with like, and note the layer in your log row.

Should I turn the firewall off briefly to test?

Avoid it. It exposes the site to the traffic your protections exist to stop, and a positive result does not identify the rule at fault. Use your vendor's event log to find the specific rule instead, then test a narrow exception.

Next step

If you do not yet see which AI crawlers your site serves and refuses, connect a log source (availability and monthly request limits depend on your plan; see plans and limits) and read the AI crawlers page with the verification filter on, then use the log above for each verified refusal. To check that a specific page is reachable and see its redirects and delivered HTML, run saved page diagnostics. Sign in at app.discoveredby.ai to get started.

  • robots.txt
  • technical geo
  • ai crawlers
  • cdn
  • firewall

Share

Summarize with AI

Start monitoring your AI visibility.

See how AI search engines talk about your brand.

Free to start. No credit card required.