All documentation

Optimize

Site audit

Crawl your own site as DiscoveredByBot from robots.txt, sitemaps and internal links, and compare raw HTML with a browser render.

Open Site health → Site audit to crawl your own site as DiscoveredByBot, check each page it finds, and compare the raw HTML response with a browser render of the same page.

Site audit showing illustrative audit progress, counters, filter tabs and a page results table

What a site audit finds

Starting an audit gives it three starting points on your project's domain: robots.txt, a guessed /sitemap.xml, and the audited page itself (api/services/site_audit/service.py#seed_rows). From there it keeps discovering candidate pages by:

  • reading every Sitemap: line in robots.txt (api/services/site_audit/frontier.py#robots_sitemaps);
  • following a sitemap index to the sitemap files it lists, and reading a sitemap's page entries, including a gzip-compressed sitemap (api/services/site_audit/frontier.py#parse_sitemap);
  • following the internal links found in each checked page's raw and rendered HTML.

Sitemap files can be found anywhere on your registrable domain, including a different subdomain. The pages a sitemap or a page links to are kept only when they share the audited page's exact host, so a link to a different subdomain is not queued for a check (api/services/page_diagnostics/extended.py#same_host_links). If the audited page itself redirects to a different host on the same domain, for example an apex domain to its www subdomain, the audit adopts that host as its origin and looks for robots.txt and sitemap.xml there as well.

A discovered page is checked, unless one of three things is true: DiscoveredByBot's declared policy for that exact path is blocked, the response is not declared as HTML, or robots.txt itself could not be read: a server error (HTTP 500 and above), a network failure, a redirect loop, an unreadable file, or any other response that is neither a clean 200 nor an HTTP 400 to 499 status. A missing robots.txt, one that returns any HTTP 400 to 499 status, does not skip the page: under RFC 9309 an unavailable robots.txt places no restrictions on the fetch, so DiscoveredByBot is treated as allowed there. The five AI crawler policies checked below still read unknown in that case, since their own declared policy still cannot be read from a file that is not there. Either way a skipped page is recorded as skipped, not checked (api/services/page_diagnostics/collect.py#inspect).

Start an audit, and follow it

Choose how many pages to check, from 1 to 1,000, 100 by default, then start the audit. Only one audit runs per project at a time, and a project can start a new audit at most once a minute (api/services/site_audit/service.py#start_audit, #START_COOLDOWN_SECONDS). Starting an audit keeps only the newest 10 audits for the project; older audits and their saved pages are removed (api/services/site_audit/service.py#RETAIN_AUDITS).

Pages are checked one at a time, a few seconds apart, in the background; the screen refreshes on its own while an audit is queued or running (api/services/site_audit/service.py#STEP_DELAY_SECONDS). Discovery stops at 5,000 discovered page rows and 20 sitemap files; checking stops once as many pages as you chose have finished, completed, skipped or failed, even if more were discovered. Each of these three limits shows its own note on the audit when it is reached (api/services/site_audit/service.py#MAX_DISCOVERED, #MAX_SITEMAPS, api/services/site_audit/runner.py#refresh_counters, #_insert_discovered).

If a step is interrupted, for example by a worker restart, the audit resumes automatically within about ten to fifteen minutes: an audit with no live lease becomes eligible once its last update is ten minutes old, and a periodic check for resumable audits runs every five minutes, so the wait is that ten minutes plus up to another five for the next check to run (api/services/site_audit/service.py#LEASE_MINUTES, api/services/site_audit/runner.py#resumable_audit_ids, #RESUME_IDLE_SECONDS). Resuming also swaps in a fresh dispatch token for the audit, so a step from before the interruption can never run alongside the resumed one (api/services/site_audit/runner.py#claim_for_resume). The page that was being checked when the interruption happened is not retried: it is recorded as failed, and the audit moves on to its next page (api/services/site_audit/runner.py#_ORPHANED_ROW_ERROR). Cancelling an audit only needs owner or editor project access; it never requires the plan feature that starting an audit does, so anyone who can stop other project work can stop an audit too (api/services/site_audit/service.py#_authorize_cancel). Cancelling keeps every page already finished; only pages still pending are left unchecked (api/services/site_audit/runner.py#run_step).

Only your project's newest completed or cancelled audit feeds the Page issue queue; a running or a failed audit is skipped when choosing that audit, so it never hides the previous finished audit's saved findings. Each inspected page from that audit, including a skipped or a failed one, appears there as one Site audit checks item linking back to that page's evidence here (api/services/page_issues.py#inventory_statement).

What is checked

Each inspected page runs the same checks as Saved page diagnostics: declared crawler policy, delivered HTML, structure and metadata, response size and timing, and a browser render. Every check is Check passed, Needs review or Unknown, with saved evidence and, where the same exact URL has a saved result from an earlier completed or cancelled audit, a comparison against the newest such result using the same first observation, unchanged, new finding, still open, resolved and not comparable outcomes (api/services/page_diagnostics/checks.py#compare, api/services/site_audit/runner.py#run_step). The structure, response and render checks below are disclosed heuristics for review, from one lab request or one lab browser load as DiscoveredByBot: they are not Core Web Vitals, accessibility certification, a citation prediction or a grade (api/services/page_diagnostics/extended.py).

Declared policy and delivered HTML

Check Review rule
Declared crawler policies Needs review when the matched robots.txt rule for a search or training crawler, or DiscoveredByBot, is a Disallow rule at that redirect hop. Unknown when the policy could not be read, except DiscoveredByBot's own check on a robots.txt that is missing outright (HTTP 400 to 499): that passes instead, since RFC 9309 treats an unavailable file as no restrictions.
HTTP delivery Needs review at HTTP 400 or above. Unknown when no status was reached, including an unresolved redirect (missing redirect target or too many redirects) or any other status that is neither 200 nor 400 or above.
Complete HTML response Unknown unless a bounded UTF-8 HTML response was read; this check never shows Needs review on its own.

Structure and metadata

Check Review rule
Page title, Main heading, Meta description, Text in delivered HTML Needs review when the delivered HTML has none.
Indexing directives Needs review when a robots meta tag or X-Robots-Tag carries noindex or none.
Canonical URL Needs review when no canonical link is present, when several different canonical URLs are declared, or when the one declared URL points off the project domain. Declaring the same URL more than once is fine.
Document language Needs review when the page has no html lang attribute.
Heading structure Needs review when there is not exactly one H1 (zero or several) or a heading level is skipped.
Image alternative text Needs review when any image has no alt attribute.
Structured data (JSON-LD) Needs review when any JSON-LD block fails to parse, or none is present.
Reading complexity Unknown under 100 extractable words. Otherwise needs review when the average sentence exceeds 25 words or the Flesch reading ease score is under 30. An English-language heuristic and a review prompt, not a grade (api/services/page_diagnostics/extended.py#structure_checks).

Response size, timing and render

Check Review rule
HTML response time and size Unknown with no complete response. Otherwise needs review when the full HTML response takes over 1,500 ms or the HTML is over 500,000 bytes (api/services/page_diagnostics/extended.py#response_check, #SLOW_RESPONSE_MS, #LARGE_HTML_BYTES).
Rendered page Unknown when the render was not attempted, left the project domain, returned no HTML, or the rendered HTML could not be parsed. Otherwise needs review unless the browser's HTTP status is 200.
Content that needs JavaScript Needs review when rendering adds at least 200 extra characters of text and at least doubles the raw HTML's text.
Metadata changed by JavaScript Needs review when the title, H1 or robots directives differ between the raw and the rendered HTML.
Links that need JavaScript Needs review when 5 or more same-host links appear only after rendering.
Render load and resources Needs review when the render takes over 5,000 ms, or the load event was not reached, loads more than 150 resources, or the render's proxy refused a connection (other than a blocked analytics request) or reached its transfer budget (api/services/page_diagnostics/extended.py#render_checks).

Illustrative audit page evidence scrolled to the readability, response and JavaScript render check cards

Raw HTML versus rendered

A site audit fetches each page's raw HTML for its content and structure checks, then separately loads the same page once in a browser with JavaScript enabled, and compares the two. This matters because crawlers that do not run JavaScript only ever see the raw HTML response. Content, links, or title and heading text that only appear after scripts run are invisible to a crawler that never runs those scripts, even though a person opening the page in a browser sees them (api/services/browser/render.py#render).

Safety

Every browser request the render step makes, including every resource the rendered page loads, goes through a forward proxy that only allows the public internet on ports 80 and 443. A request to a private, loopback or link-local address is refused rather than followed, and refused connections, with a sample of the hosts, appear in the render evidence for that page, alongside how many bytes were transferred (api/services/browser/egress.py#ALLOWED_PORTS, #public_addresses). The browser also disables WebRTC's non-proxied UDP path for every render, so a page cannot use it to reveal a real IP address or reach the network outside the proxy's checks (api/services/browser/render.py#render). The non-render fetches (robots.txt, sitemaps and each page's raw HTML) use the same public-address check, on ports 80 and 443, that Robots diagnostics and Saved page diagnostics use. Page scripts run during the render, and requests to common analytics and tag-manager services (for example Google Analytics, Google Tag Manager, Meta Pixel, Hotjar, Microsoft Clarity and Segment) are blocked, so an audit does not add page views to those tools; the render evidence notes how many analytics requests were blocked. First-party or self-hosted analytics, for example a server-side tag manager or analytics proxied through your own domain, are not blocked: only the specific third-party hostnames above are refused (api/services/browser/render.py#ANALYTICS_HOSTS, api/services/browser/egress.py#host_matches). DiscoveredBy does not edit your site.

A page's raw HTML is bounded to 2,000,000 bytes (api/services/page_diagnostics/collect.py#PAGE_MAX_BYTES); a sitemap file is bounded to 10,000,000 bytes, including after it is decompressed from gzip (api/services/site_audit/frontier.py#SITEMAP_MAX_BYTES). Each page's render is bounded to 200 proxy connections and 15,000,000 downloaded bytes (api/services/browser/egress.py#MAX_CONNECTIONS, #MAX_UPSTREAM_BYTES).

Access and limits

Any current project member can read a saved audit and its page evidence. Starting a new audit needs owner or editor project access, plus the project owner's robots diagnostics feature: the standard Starter, Growth and Pro plans include it, Free and Trial do not; custom plans follow their saved feature setting (api/services/site_audit/service.py#_can_start, api/services/entitlements.py#standard_entitlements). Cancelling a running audit needs only owner or editor project access, described above.

Access is also re-checked before every page it inspects, not just when the audit starts: a running audit stops with a failed status if the requester loses owner or editor project access, the project owner loses the robots diagnostics feature, or the project's domain changes while the audit is running (api/services/site_audit/runner.py#run_step).

This audit does not record whether real AI crawlers actually visited these pages: it checks declared robots policy and delivered content as DiscoveredByBot only, the same way Robots diagnostics and Saved page diagnostics do. It also does not establish complete coverage of your site: discovery is bounded, and a page you never link to and that is missing from every sitemap this audit reads will not be found.

  • Saved page diagnostics: check one exact URL on demand and keep its recheck history.
  • Robots diagnostics: an unsaved, policy-only quick check for one URL.
  • AI crawlers: which AI crawlers actually visited your pages, read from your own server logs.
  • Page issue queue: known owned pages, gaps and optimization work, plus your newest finished audit's inspected pages.
  • Plans and limits: feature access follows the project owner's plan.

Last verified 2026-09-23

Start monitoring your AI visibility.

See how AI search engines talk about your brand.

Free to start. No credit card required.