An AI search site audit: which issues should you fix first?
A site audit returns a long list of findings. Sort them into four tiers by evidence, from blocked access to weak signals, and give each one a page owner.
On this page
- In short
- Why does an AI search site audit need its own priority order?
- What does a site audit actually check, and what does it miss?
- Which findings are blocked access?
- Which findings are failed pages?
- Which findings are content deficiencies?
- Which findings have little supporting evidence?
- The severity-and-evidence checklist
- Who should own each finding?
- A worked example
- Common mistakes and what this cannot tell you
- Frequently asked questions
- Next step
An AI search site audit returns findings of very different weight, so fix them in order of evidence, not in the order the list shows them. Start with blocked access (a robots rule or noindex directive that stops a page from being fetched or used), then pages that failed to load, then content that is missing or only appears after JavaScript runs. Leave findings with thin evidence, such as reading complexity or missing alt text, for last, and treat them as review prompts rather than defects. Give every finding one named page owner.
In short
- Sort audit findings into four tiers: blocked access, failed pages, content deficiencies, and weak-evidence findings.
- A finding you cannot verify outside the audit is a question for its owner, not a confirmed defect.
- Blocked pages can hide other problems, because a page the audit skipped was never checked.
- A passed check is not a clean bill of health, and an audit does not show whether real AI crawlers visited a page.
- Give every finding a page owner and a way to confirm the fix on a recheck.
Why does an AI search site audit need its own priority order?
Because the labels describe the state of a check, not how much it matters. A site audit in DiscoveredBy marks each check as Check passed, Needs review or Unknown, and nothing more. It does not rank the findings, and the site audit docs say its structure, response and render checks are review prompts, not a grade or a citation prediction.
The four tiers and the owner roles in this post are our editorial judgement, not a product feature and not a measured ranking of what affects AI answers. They put findings that stop a page from being reached or used ahead of findings that only make a page harder to read.
A site audit is a crawl of your own site that checks each page it finds and compares the raw HTML with a browser-rendered version. Delivered HTML is what your server sends before any script runs, which is all a crawler without JavaScript sees.
What does a site audit actually check, and what does it miss?
It checks declared crawler policy, the HTTP response, the delivered HTML, structure and metadata, response size and timing, and a single browser render. It does not observe real crawler traffic, and it does not cover pages it cannot discover.
The audit starts from your robots.txt, a guessed /sitemap.xml and the audited page, then follows sitemap entries and internal links. Discovery is bounded, so a page that no sitemap lists and no page links to will not be found. The audit also runs as DiscoveredByBot, so it reports the response that bot received. A firewall can return something different to another client.
Read these limits before you trust a clean result. The structure, response and render checks come from one lab request and one lab browser load, so they are not Core Web Vitals. The audit does not record whether real AI crawlers visited a page; for that, use your server logs (see AI crawlers). A passed check does not establish indexing, citation eligibility or complete page health.
Which findings are blocked access?
These are findings where the page may not be fetched, or may be excluded from a search index, no matter how good it is. Fix them first, but confirm the intent before you change anything, because some blocks are deliberate.
There are four things to look for:
- Declared crawler policies. The check needs review when the matching
robots.txtrule for a search crawler, a training crawler or DiscoveredByBot is a Disallow rule. The robots diagnostics docs list five AI crawlers and separate search from training: OAI-SearchBot (ChatGPT search), GPTBot (OpenAI model training), Claude-SearchBot (Claude search), ClaudeBot (Anthropic model training) and PerplexityBot (Perplexity search). - Indexing directives. The check needs review when a robots meta tag or an
X-Robots-Tagheader carriesnoindexornone. - An unreadable
robots.txt. If the file returns a server error, a network failure or a redirect loop, the audit skips pages and the crawler policies read Unknown. Unknown does not mean allowed. - Skipped pages. A page the audit skipped because DiscoveredByBot's policy for that path is blocked is recorded as skipped, not checked. Nothing else about that page has been examined.
The first judgement is whether the block is intended. A team may block a training crawler on purpose while allowing search crawlers, and the docs note that a deliberate restriction can be appropriate. So the question for the owner is "did you mean this?", not "remove this rule".
Allowing a crawler in robots.txt also does not prove it can get in. The docs are explicit that an allow rule does not establish access through a firewall, rendering, visits, indexing or citations. Confirm real access in your server logs.
Which findings are failed pages?
These are pages the audit reached or tried to reach but could not read. Treat them as the second tier, because a page that returns an error cannot be used by anyone.
- HTTP delivery needs review at HTTP 400 or above.
- Rendered page needs review unless the browser's HTTP status is 200.
- A page can also be recorded as failed to load. In the Page issue queue it reads Needs action if any check on it needs review, and Diagnostics unknown otherwise.
- An interrupted audit records the page it was checking as failed, without retrying it. Recheck it before you file a ticket.
Separate persistent failures from one-off ones; a page that failed because a worker restarted is not a site defect. Run a saved page check on that exact URL and see whether the failure repeats. As saved page diagnostics explains, a finding counts as resolved only after a later check passes.
Which findings are content deficiencies?
These are pages that load but deliver too little, or deliver the important content only after JavaScript runs. They come third, because they are common and fixable, but they only matter once the page can be reached.
The audit's rules are specific, which makes them easy to hand to a developer:
- Page title, main heading and text in delivered HTML each need review when the delivered HTML has none.
- Content that needs JavaScript needs review when rendering adds at least 200 extra characters of text and at least doubles the raw HTML's text.
- Metadata changed by JavaScript needs review when the title, H1 or robots directives differ between raw and rendered HTML.
- Links that need JavaScript needs review when 5 or more same-host links appear only after rendering.
- Canonical URL needs review when there is no canonical link, when several different canonical URLs are declared, or when the one declared URL points off your project domain.
The raw-versus-rendered comparison is the most useful part of this tier. A person looking at the page in a browser sees everything, so the page feels fine. A crawler that does not run scripts sees only the raw response. The audit shows you the gap, but not which crawlers behave which way, so do not generalise it to every engine.
Which findings have little supporting evidence?
These checks are real observations, but they are the weakest basis for an urgent fix. Do them when the higher tiers are clear, and never report them as the reason a page was not cited.
- Meta description needs review when the delivered HTML has none. The docs call this a review prompt, not proof of a citation problem.
- Reading complexity is an English-language heuristic and a review prompt, not a grade. It reads Unknown under 100 extractable words.
- Heading structure flags anything other than exactly one H1, or a skipped level.
- Image alternative text flags any image without an
altattribute. - Document language flags a missing
html langattribute. - Structured data (JSON-LD) flags a block that fails to parse, or none present.
- HTML response time and size flags a full response over 1,500 ms or HTML over 500,000 bytes, from a single request.
- Render load and resources flags a render over 5,000 ms or more than 150 resources, from a single browser load.
Structured data deserves a careful sentence. A JSON-LD block that fails to parse is worth fixing because it is broken. Its absence is a much weaker finding, and the docs say these structural observations do not predict lift or establish citation eligibility. This post makes no claim that structured data or an llms.txt file affects rankings or citations; the audit does not test for that. See the glossary entries for schema markup and llms.txt, and the llms.txt Advisor if you want a review of yours.
Unknown results belong here too. Unknown means the audit could not read the evidence. It never counts as a pass, and it never resolves a finding on recheck.
The severity-and-evidence checklist
Copy this table into your tracker. The tier is our editorial ordering, the evidence column tells you what to open in the audit, and the last column names who should confirm it. Owners are suggestions; use whichever roles your team has.
| Tier | Finding (audit label) | Evidence to open | Confirm outside the audit | Suggested owner | Done when |
|---|---|---|---|---|---|
| 1. Blocked | Declared crawler policies: Disallow | Matched rule, line number and User-agent group | Was the block intended? Check server logs for real crawler requests | Site or web owner | Rule is intended, or removed and recheck shows it allowed |
| 1. Blocked | Indexing directives: noindex or none | Meta tag or X-Robots-Tag value | Is the directive deliberate for this page? | Page owner, or CMS admin | Directive matches intent; recheck shows Check passed |
| 1. Blocked | robots.txt unreadable (Unknown policies, skipped pages) | HTTP status and fetch URL for robots.txt | Load the file yourself; check the host and redirects | Developer or infrastructure | File returns a clean 200; a new audit checks the pages |
| 1. Blocked | Page skipped: DiscoveredByBot blocked | Skipped status on the page item | Is this path meant to be private? | Site owner | Path is intentionally private, or unblocked and audited |
| 2. Failed | HTTP delivery: 400 or above | Status and final destination | Open the URL; check for redirects or removed pages | Developer, or page owner if removed | URL returns 200, or is redirected or retired on purpose |
| 2. Failed | Failed to load or render status not 200 | Render evidence and refused connections | Recheck once to rule out a one-off | Developer | Recheck completes with HTTP 200 |
| 3. Content | No title, H1 or text in delivered HTML | Delivered-HTML checks | View the page source, not the rendered page | Developer, then content owner | Delivered HTML contains the text |
| 3. Content | Content or links that need JavaScript | Text added by render; link count | Compare page source with the rendered page | Developer | Key text and links appear in delivered HTML, or the gap is accepted |
| 3. Content | Metadata changed by JavaScript | Raw versus rendered title, H1, robots values | Check which value should be authoritative | Developer or SEO owner | Raw and rendered values agree |
| 3. Content | Canonical missing, conflicting or off-domain | Declared canonical URLs | Decide the intended canonical URL | SEO owner | One on-domain canonical is declared |
| 4. Weak evidence | Description, alt text, language, headings, reading complexity | Individual check evidence | Judge relevance to the page's purpose | Content owner | Fixed, or recorded as accepted |
| 4. Weak evidence | Structured data fails to parse or is absent | Parse result | Valid markup relevant to the page? | Developer or SEO owner | Blocks parse; absence is a decision, not a defect |
| 4. Weak evidence | Slow or large response; heavy render | Timing, size, resource count | Repeat the request from another place | Developer | Number is acceptable to the team |
| Unknown | Any check Unknown | Reason shown on the check | Retry, or investigate why it could not be read | Whoever owns the affected check's tier | Recheck yields Check passed or Needs review |
How to use the checklist
- Work down by tier. Do not start tier 3 while a tier 1 finding is unexplained.
- For each row, record the URL, the tier, the owner, the date opened and the audit it came from.
- Ask the owner to confirm the finding outside the audit before they change code or copy.
- Record the fix, then recheck the exact URL. A finding counts as resolved only when a later check passes.
Who should own each finding?
Give every finding a page owner: the person who can decide whether the page should look the way the audit found it. This is a role, not a job title, and it changes by tier.
Blocked access usually belongs to whoever runs the site's robots.txt and headers, and failed pages to a developer. Content deficiencies split between the developer (what is missing from delivered HTML) and the content owner (what is missing from the copy). Weak-evidence items belong to the editorial owner, and many should end as "accepted".
In DiscoveredBy, the Page issue queue helps with the page part of this. Each inspected page from your newest finished audit appears as one Site audit checks item, marked with a "Site audit" badge and linking back to that page's saved evidence. The item reads Needs action when any check on the page needs review (including a page skipped because DiscoveredByBot's own policy is blocked), otherwise Diagnostics unknown or Diagnostic checks passed. The docs describe no owner assignment in the queue, so keep owners in your own tracker, alongside the URL. The roles above are suggestions, not product fields.
The queue is known inventory, not a complete site crawl, and "No saved issues" does not mean a page is healthy.
A worked example
Illustrative example: Quillstone and its competitors are fictional, and the numbers are made up to show the method.
Quillstone, a document-review software company, runs a site audit on 40 pages. The audit finishes with 2 pages skipped, 2 pages failed to load and 36 pages completed (2 + 2 + 36 = 40).
Tier 1. The 2 skipped pages are both under /portal/. DiscoveredByBot's declared policy blocks that path. The web lead confirms the block is intended, because those are customer-only pages. The finding is closed as "accepted, private by design". Nothing was changed.
Tier 2. The 2 failed pages are /security and /integrations/import. A recheck of /security completes with HTTP 200, so the first failure was a one-off. The recheck of /integrations/import returns a 404 for the same URL. The developer confirms a migration removed it, adds a redirect to the new page, and the next recheck passes.
Tier 3. Of the 36 completed pages, 5 have content that needs JavaScript: the pricing page and 4 feature pages. On the pricing page the plan comparison table appears only after rendering. The developer moves the table into the delivered HTML. On the other 4 pages the added text is a cookie-preference widget, so the owner marks them accepted. Separately, 3 pages have no canonical link; the SEO owner adds one to each.
Tier 4. Among the same 36 pages, 7 lack a meta description, 9 have an image without alt text and 4 are flagged for reading complexity. The content owner batches these into the next editorial pass and does not link them to any citation goal.
Result: one access decision confirmed, one error page fixed, one important table made visible to crawlers that do not run scripts, and 20 low-evidence items queued as routine editing (7 + 9 + 4 = 20). The team did not start with the 20.
Common mistakes and what this cannot tell you
- Fixing the longest list first. Tier 4 is usually the biggest and least informative.
- Unblocking a training crawler by reflex. Search and training controls are separate. Decide the policy, then change the rule.
- Treating Unknown as fine. Unknown means the evidence was not readable. It is a task, not a pass.
- Reporting the fix as the cause. An audit recheck shows a response changed. It does not show an edit caused any change in AI answers. If you want to evaluate that, use a before-and-after review.
- Stopping at the audit. A page can be reachable and well formed and still not be cited. The audit cannot explain why an engine used one source over another; see your page was retrieved but not cited for that question.
Frequently asked questions
How often should I run an AI search site audit?
The docs describe no schedule; an audit runs when you start it, and a project can start one at most once a minute. A sensible rhythm is after any change to robots.txt, redirects, a template or a migration, plus a periodic pass you choose. Only your newest completed or cancelled audit feeds the Page issue queue.
Does allowing GPTBot or ClaudeBot make my site appear in AI answers?
No. Those two are listed as training crawlers, and search and training are separate controls. Allowing a crawler in robots.txt does not establish access, indexing or citation. Decide each policy on its own merits.
Is a missing meta description or alt text hurting my AI visibility?
The audit cannot say. The docs call an absent meta description a review prompt, not proof of a citation problem. Fix these when they help human readers, and do not expect a measured effect.
Should I add structured data or an llms.txt file first?
Neither is the first step. Access and delivered content come first, and this post does not claim that either structured data or llms.txt influences rankings or citations. A JSON-LD block that fails to parse is a fair bug to fix.
Can the audit find my whole site?
No. It discovers pages from robots.txt, sitemaps and internal links, within bounded limits. A page that no sitemap lists and no page links to will not be found.
Next step
Run a site audit on your project, sort the findings into the four tiers, and put an owner beside each row of the checklist. Site audits, saved page diagnostics and robots checks are available depending on your plan; see plans and limits. The optimize features overview shows where the audit sits with the rest of the workflow. To start, sign in to DiscoveredBy.
- site audit
- robots.txt
- technical geo
- page issues