Can a crawler see the content on your JavaScript-heavy page? A raw HTML vs rendered check

Some crawlers read only the raw HTML your server sends. A five-step check compares that HTML with the rendered page, so you know what a non-rendering crawler can see.

Kamal 13 min read
A sealed envelope beside an opened one, where the opened envelope shows several extra sheets marked in coral, suggesting content that only appears after opening.
On this page
  1. In short
  2. Why does JavaScript rendering matter for AI search?
  3. Which crawlers run JavaScript?
  4. How do you compare raw and rendered content?
  5. How do DiscoveredBy's checks do the same comparison?
  6. The deliverable: a content-availability log
  7. Worked example: Quillstone's pricing tabs
  8. Common mistakes and limits
  9. Frequently asked questions
  10. Next step

A crawler can only see what it receives and processes. If your page ships its content in the first HTML response, any crawler that receives that response can read it. If the content is assembled by JavaScript in the browser, a crawler that does not run JavaScript sees the empty shell instead. To find out which situation you are in, fetch the raw HTML, load the same URL in a real browser, and compare the text, links, title and main heading. Then check what the specific crawler you care about documents, because crawlers differ and one crawler's behaviour tells you nothing about another's.

In short

  • Raw HTML is the response your server sends before any script runs. Rendered HTML is the page after a browser has run its scripts. The gap between them is the content that depends on JavaScript.
  • Some crawlers read only the raw HTML; others run scripts. Check each crawler's own documentation, and never assume that one crawler's behaviour applies to the rest.
  • A five-step check (fetch, view, compare, record, recheck) gives you a reproducible answer for any page. The copyable log below turns it into evidence.
  • DiscoveredBy's page diagnostics and site audit run the raw-versus-rendered comparison for you, but they observe the page as DiscoveredByBot, not as any AI engine's crawler.
  • A gap is a reason to review, not proof of a lost citation.

It matters because your page can look complete to a person and nearly empty to a crawler. A person opens the page in a browser, the browser runs the scripts, and the text appears. A crawler that does not run scripts requests the same URL, gets the HTML the server sent, and stops there.

Two terms make the rest of this post easier to follow. Client-side rendering means the browser builds the page content with JavaScript after loading a mostly empty HTML document. Server-side rendering (and static generation) means the server sends HTML that already contains the content, so scripts only add interactivity. Frameworks can do either, and many sites mix both, which is why you have to test a page rather than assume from the framework name.

For search work this affects three things: whether your body text is available to be read, whether the links to your other pages can be followed, and whether the title and main heading a crawler records match what visitors see. If you are still building the wider picture of how content becomes citable, the key terms page defines mention, citation and related words used across this series.

The same URL, two views. The gap is the JavaScript-dependent content.

Which crawlers run JavaScript?

There is no single answer, and this post will not give you one. Whether a crawler runs scripts is a fact about that crawler, published (or not) by its operator, and it can change over time. Treat it as something to look up, not something to infer.

Here is a safe way to handle it. For each crawler that matters to you, find the operator's own documentation for that crawler and note whether it says the crawler renders pages. Record the date you checked. If the documentation is silent, write "not documented" in your log rather than guessing. The AI engines DiscoveredBy tracks are listed in the engines reference, but that page describes how answers are collected for monitoring, not how each vendor's own crawlers fetch pages, so it cannot settle this question for you.

Because a crawler's behaviour is uncertain and the safe option is cheap, the practical rule is: put the content you most need to be read (main text, key facts, links to important pages, title and headline) in the initial HTML. Then the answer to "does this crawler run JavaScript?" matters much less for that content.

How do you compare raw and rendered content?

Fetch the raw HTML, load the same URL in a browser, and compare the two side by side. Do this for the exact URL, because paths, trailing slashes and query parameters can change what a server returns.

Step 1: capture the raw HTML

Request the URL without running any scripts. A command-line fetch is the simplest reproducible way:

curl -sL -o raw.html -D headers.txt "https://www.example.com/your-page"

The -L follows redirects, -D saves the response headers, and raw.html holds the body. Open headers.txt first: note the final status code and any X-Robots-Tag header. Some servers return different HTML to different user agents, so if you suspect that, repeat the fetch with the user agent string the crawler's operator documents and note the difference.

In a browser you can get the same thing from "view source", which shows the HTML as delivered. It does not show what scripts added afterwards.

Step 2: capture the rendered page

Open the URL in a normal browser, wait for the page to settle, and copy the text a visitor sees. Browser developer tools show the current document (after scripts have run), which is the rendered HTML. Save the visible text, the list of links, the page title and the main heading.

Step 3: compare four things

Compare the raw and rendered versions on the same four items, in this order:

  1. Body text. Is the passage a reader would need (the answer, the specification, the price, the definition) present in the raw HTML?
  2. Links. Are the links to your important pages present as real anchor links in the raw HTML, or do they only appear after scripts run?
  3. Title and main heading. Do they match between raw and rendered?
  4. Indexing directives. Does the robots meta tag differ between the two? A directive that changes after rendering is easy to miss, so look for it deliberately. (An X-Robots-Tag header comes from the server, so read it in the headers you saved in step 1.)

Judge by meaning, not by character counts alone. A page can carry the same wording in both and still differ in markup; what matters is whether the sentences you want quoted are in the raw response.

Step 4: record the result

Write down what you saw, with the date, in a log (template below). Reproducibility is the point: someone else should be able to repeat the fetch and get the same comparison.

Step 5: recheck after changes

If your team changes how the page is delivered, repeat steps 1 to 4 on the same URL and compare with the earlier entry. An observed change in delivery is a fact; it is not proof that the change caused any change in citations.

How do DiscoveredBy's checks do the same comparison?

DiscoveredBy's page diagnostics and site audit automate the comparison, and they are the closest built-in match to the steps above. Each run that gets a successful HTTP 200 HTML response also loads the page once in a browser with JavaScript enabled and compares it with the raw HTML response.

The render checks flag a page for review under these rules, taken from the docs:

Check Flagged for review when
Content that needs JavaScript Rendering adds at least 200 extra characters of text and at least doubles the raw HTML's text
Metadata changed by JavaScript The title, H1 or robots directives differ between raw and rendered HTML
Links that need JavaScript 5 or more same-host links appear only after rendering
Rendered page The browser's HTTP status is not 200
Render load and resources The render takes over 5,000 ms, the load event is not reached, more than 150 resources load, or the render's proxy refuses a connection (other than a blocked analytics request) or reaches its transfer budget

Each check comes back as Check passed, Needs review or Unknown. Unknown appears when, for example, the render was not attempted or the rendered HTML could not be parsed. The site audit runs the same checks across a crawl of your site (you choose how many pages, from 1 to 1,000) and saves the render checks with their evidence for each page it inspects. The docs describe how the audit discovers pages from robots.txt, sitemaps and internal links.

Demo data. Site audit page evidence with the render checks.

Three limits deserve plain statement:

  • It is DiscoveredByBot's view. The docs say the checks come from one lab request or one lab browser load as DiscoveredByBot. They do not simulate another crawler and do not establish that any AI engine's crawler runs scripts or not.
  • Thresholds are heuristics. A page can carry important JavaScript-only content that stays under the 200-character or doubling thresholds, so a pass is not a clean bill of health. The body-text check measures text in delivered HTML, and the docs note it does not measure text visible after JavaScript runs.
  • It does not record real visits. Neither feature records whether real AI crawlers visited the page. For that, the AI crawlers page reads your own server or CDN logs once you connect a log source. It shows requests and what your site served; it does not tell you whether a given crawler then ran your scripts. For the access-versus-citation question, see the related post on AI crawler logs.

The deliverable: a content-availability log

Copy this table into a spreadsheet, one row per page and per crawler question. The table records what you observe; the finding block below it holds the judgements, which you label as such.

URL (exact, with path and query):
Date checked:
Checked by:

| Item                          | Raw HTML (Y/N/partial) | Rendered (Y/N/partial) | Evidence (quote or note) |
| ----------------------------- | ---------------------- | ---------------------- | ------------------------ |
| Title matches                 |                        |                        |                          |
| Main heading (H1) matches     |                        |                        |                          |
| Key answer paragraph present  |                        |                        |                          |
| Key facts (price, specs, etc) |                        |                        |                          |
| Links to priority pages       |                        |                        |                          |
| Robots meta / X-Robots-Tag    |                        |                        |                          |
| FAQ or tab content            |                        |                        |                          |
| Structured data (JSON-LD)     |                        |                        |                          |

Crawler question:
- Crawler name and operator:
- Operator documentation URL and date read:
- Documented as rendering JavaScript? (yes / no / not documented):

Finding (state at the strength the evidence supports):
- Observed: (what differs between raw and rendered)
- Not established: (what this check cannot show)
- Hypothesis to test: (a change you could try, and how you will recheck)

Fix owner / change made / recheck date:

The "Not established" line matters. It stops a log from saying "crawler X cannot see our pricing" when all you observed is that the raw HTML lacks the pricing.

Worked example: Quillstone's pricing tabs

Quillstone sells document-review software to legal and compliance teams. Its marketing team notices that AI answers about "document review software pricing" describe competitors such as Brieflane and Clausewise in detail but say little about Quillstone. They suspect the pricing page.

They run the check on the exact URL https://www.quillstone.example/pricing.

Steps 1 and 2. The raw fetch returns status 200. Removing scripts and styles, the raw HTML contains 1,100 characters of text: a headline, a one-line intro and a "Compare plans" button. The rendered page shows 3,400 characters, including three plan cards, a feature table and an FAQ, all inserted after a script fetches plan data.

Step 3. The comparison table:

Item Raw HTML Rendered
Title "Quillstone" "Quillstone: Document review for legal teams"
H1 Present ("Pricing") Present ("Pricing")
Plan details and feature table No Yes
FAQ answers No Yes
Links to product pages 2 9

Applying the rules in the table above to these made-up numbers: rendering adds 2,300 characters (3,400 minus 1,100), which is at least 200, and the rendered text is more than double the raw text (3,400 is about 3.1 times 1,100), so Content that needs JavaScript would be flagged for review. The title differs, so Metadata changed by JavaScript would be flagged. Seven links (9 minus 2) appear only after rendering, which is 5 or more, so Links that need JavaScript would be flagged too.

Step 4, the finding. Observed: the plan details, the feature table, the FAQ and seven links exist only after scripts run. Not established: whether any particular AI crawler visited the page, and whether the missing pricing text is why Quillstone is described thinly. Hypothesis to test: delivering the plan cards and FAQ in the initial HTML would make them available to a crawler that does not run scripts.

Step 5. The team ships the change, repeats the fetch and sees the raw HTML now holds 3,300 characters and all nine links. They record the date and keep watching the same prompts in their monitoring tool for changes in how the answers describe the pricing. Any shift in those answers is an observation to interpret carefully, not proof the edit caused it.

Common mistakes and limits

  • Assuming your framework tells you the answer. The same framework can serve content in the first response on one route and build it on the client on another. Test the URL.
  • Testing the home page only. Templates differ. Check at least one page from each template that matters: pricing, product, article, docs.
  • Reading "view source" as "what the crawler sees". It shows what your server sent to you. A firewall or CDN rule may treat other visitors differently; see the related post on CDN and firewall blocking.
  • Generalising one crawler to all engines. Rendering behaviour belongs to each crawler. This check tells you what is in the raw HTML, which is the safe common denominator, not which crawlers render.
  • Hiding content behind interactions. Text that appears only after a click, a scroll or a tab switch may never be in the rendered page either. Check that the content is present without interaction.
  • Treating a pass as full health. The diagnostics are single lab observations, not a citation prediction or a grade.
  • Forgetting discovery. Content in the raw HTML still has to be found. If a page is missing from your discovery workflow, see the new-page discovery checklist.

Frequently asked questions

Do AI crawlers run JavaScript?

It depends on the crawler, and the answer belongs to the operator of that crawler. Look for the operator's own documentation and record what it says and when you read it. If it is not documented, treat the crawler as one that may read only the raw HTML, and keep your important content in the initial response.

Is server-side rendering always the fix?

Not always, and it is not the only option. What you need is for the text and links that matter to be in the first HTML response. Server-side rendering and static generation both do that; a small amount of client-side enhancement on top is usually harmless for content that is already present.

Does a raw-versus-rendered gap mean I will not be cited?

No. A gap tells you which content depends on scripts. Citations are observations of answers, and whether any engine can use your page depends on much more than this one factor. Use the gap as a reason to review, then recheck.

How is this different from a robots.txt problem?

A robots.txt rule decides whether a crawler is asked to fetch a path at all. A rendering gap concerns what is in the response after the fetch is allowed. A page can be fully allowed and still hand a non-rendering crawler an empty shell. Robots diagnostics covers the first; this check covers the second.

How often should I run the check?

Run it when a template changes, when a front-end framework or deployment setup changes, and when a key page is redesigned. Saved diagnostics keep the latest 20 runs per exact URL, so a recheck after each release is practical.

Next step

Pick the three pages you most want cited, run the five-step check on each, and fill in the log. If you would rather not do the browser comparison by hand, run page diagnostics on one exact URL, or a site audit across your site, and read the JavaScript render checks alongside the optimisation features. Availability of these checks depends on your plan; see plans and limits. You can start at DiscoveredBy.

  • site audit
  • technical geo
  • javascript rendering
  • crawlers
  • raw html

Share

Summarize with AI

Start monitoring your AI visibility.

See how AI search engines talk about your brand.

Free to start. No credit card required.