Optimize
Robots diagnostics
Check declared robots.txt rules for five AI crawlers on a project URL, with matching rules and unknown fetch failures.

Use this screen for an unsaved policy-only quick check. For policy checks saved alongside HTTP and HTML evidence, open Saved page diagnostics. That workflow adds retained run history and recheck comparisons to the page issue queue, under the same feature entitlement. Site audit runs the same declared-policy, HTTP, HTML and browser-render checks across a whole crawl of your site, discovered from your robots.txt, sitemaps and internal links, under the same feature entitlement to start a crawl.
Check the declared policy for a page
Robots diagnostics checks the robots.txt
rules for a URL on your project's domain, including its subdomains. Enter the
full http:// or https:// page URL and choose Check robots.txt. Its path,
capitalization, query parameters and percent escapes are preserved when the
rules are evaluated. The page itself is not fetched
(api/services/robots/service.py#project_url, #check_robots).
Owners and editors can run a check when the project owner's plan includes
robots diagnostics. The standard Starter, Growth and Pro plans include it;
Free and Trial do not. Custom plans follow their saved feature settings
(api/routers/robots.py#check_project_robots). Results stay on the screen for
the current project and are not saved as audit history.
Reading the results
Each crawler gets its own result, with its purpose beside its name:
- OAI-SearchBot: ChatGPT search.
- GPTBot: OpenAI model training.
- Claude-SearchBot: Claude search.
- ClaudeBot: Anthropic model training.
- PerplexityBot: Perplexity search.
Search and training are separate controls. The names and purposes are based on the providers' OpenAI, Anthropic, and Perplexity crawler documentation, reviewed on September 18, 2026.
Allowed by policy means no applicable rule blocks this URL. Blocked by
policy means the winning rule is a Disallow rule. Matching results include
the rule, its original line number and the applicable User-agent group.
Explicit crawler groups take precedence over the wildcard group; repeated
groups for the same crawler are combined. The most specific matching path
wins, with Allow winning equal-specificity ties. Wildcards, end anchors,
comments and percent-encoded paths are supported
(api/services/robots/policy.py#evaluate_policy).
Unknown means the file could not be safely read. This includes timeouts,
redirect failures, responses other than HTTP 200, invalid UTF-8, HTML challenge
pages and files over 500 KiB. Unknown does not mean allowed. A readable empty
file has no blocking rules. The screen shows the HTTP status when available,
fetch URL and check time (api/services/robots/fetch.py#fetch_robots).
A robots.txt is read whatever content type it is served with, including no content type at all, which is common on object storage. Only a response declared or shaped as HTML is refused, because that is the challenge-page case worth catching.
What the check can establish
This is a reading of the declared policy, using RFC 9309 matching rules. It is not a simulation of every crawler's behavior. For example, the standard permits crawling after some missing-file responses and prescribes conservative handling for server failures; this diagnostic reports the policy as unknown when the file could not be read.
The file is fetched as DiscoveredByBot. A server can return different content to another client. Allowing a crawler does not establish actual access through a firewall, JavaScript rendering, visits, indexing or citations. No bot traffic or rendering measurement is included. Changes to robots.txt must be published on your own site, then checked again.
The fetch accepts public web addresses on ports 80 and 443 without embedded
credentials. Each redirect is checked again, and connections use validated
public IP addresses. The check follows at most five redirects and has a
30-second overall deadline (api/services/robots/fetch.py#fetch_robots).
An internationalized domain is converted to its punycode form before the fetch,
so the reported fetch URL for such a site reads as xn-- labels.
Related
- Saved page diagnostics: save policy, HTTP and HTML evidence for one URL, with recheck history.
- Site audit: crawl your site and run the same checks across every page it finds.
- AI crawlers: which AI crawlers actually requested your pages, read from your own server logs
- llms.txt Advisor: review a site's curated AI reading guide
- Plans and limits: feature access follows the project owner's plan
- Domains: measure sources in the answers you collect
Last verified 2026-09-18