Why your AI visibility score changed when your website did not: a diagnostic checklist

A drop or jump in AI visibility often has nothing to do with your site. Work through a ten-step checklist covering windows, filters, engines, failed runs and prompt edits.

Kamal 15 min read
A balance scale with its needle off centre and three small unmarked weights beside it, one of them outlined in coral.
On this page
  1. In short
  2. What does an AI visibility score actually measure?
  3. How should you investigate a change in your score?
  4. The ten-step diagnostic checklist
  5. A change log you can copy
  6. Worked example: Quillstone's visibility falls ten points
  7. Common mistakes and limits
  8. Frequently asked questions
  9. Next step

Your AI visibility score can change when your website has not, because the score is the share of analysed answers that name your brand, and both sides of that ratio can move without your site being involved. The date window can shift, a filter can differ, an engine can join or drop out, runs can fail, prompts can be added or paused, or the engines can simply answer differently on a given day. Check those causes in a fixed order, cheapest first, before you conclude anything about your content.

In short

  • Brand visibility is the share of analysed answers that name your brand. Anything that changes which answers are in the pool changes the score, with no change to your site.
  • Check mechanical causes first: date window, filter chips, then how many answers are analysed.
  • Then check what changed in the pool: engines, failed runs, prompts, competitors and brand names.
  • Small samples and ordinary answer variability can produce swings by themselves. Read the trend, not one day.
  • A checklist tells you which explanations are ruled out. It cannot tell you why an engine produced a particular answer.

What does an AI visibility score actually measure?

In this post, "AI visibility score" means the Visibility figure on Overview, which is brand visibility: analysed answers naming your brand, divided by analysed answers. It is one ratio, not a composite of several metrics. An analysed answer is a completed run whose text has been read and its named brands recorded. A collected answer is any completed run, analysed or not. The metrics reference lists which metric divides by which population.

That definition is the whole reason scores move on their own. The numerator is answers that name you. The denominator is every analysed answer in the window, from whichever prompts, engines, countries and languages you are running. Change the mix in the denominator and the ratio moves, even when each individual engine answers exactly as before.

The score is a ratio; most explanations for a move sit in the denominator.

How should you investigate a change in your score?

Work from the causes that need no judgement to the ones that need the most. Mechanical checks take minutes and eliminate the most, while variability is what remains when everything else is ruled out.

Two ground rules keep the investigation honest. First, write down the old and new value together with the window and filters behind each, because a number without its window cannot be compared. Second, treat every finding as a description of how the measurement changed, not as proof of why an engine produced its answers.

The ten-step diagnostic checklist

Use this table in order. Each row names what to check, where to look, and what a positive finding means.

Step Check Where to look If you find it
1 Are the two windows what you think? Date range chip; the window ends yesterday An answer collected today counts from tomorrow, so a comparison made today can differ from one made tomorrow.
2 Are the filters identical? Engine, tag, country, persona and language chips Chips follow you from screen to screen and can be left set. Compare like with like before anything else.
3 How many answers are analysed? The notice "Brand figures use N of M answers" on Overview Visibility divides by analysed answers, so a lagging analysis step can change the base.
4 Did the engine mix change? Explorer broken down by engine; plan; the note "Changes compare only the engines measured in both periods" A new engine, a removed one, or one that ran on few days changes the pooled number.
5 Did runs fail? The "Failed runs in this window" notice; a prompt's run history Failed runs are absent, not counted as misses, but fewer answers means a smaller sample.
6 Did Google show no AI answer? Run history on the Google engines Those runs are not answers and sit in neither side of the ratio.
7 Did the prompt set or its targets change? Prompts list; paused count; countries, cities, personas, languages The pool now contains different questions.
8 Did the competitor roster or brand names change? Competitors; Settings, Brands Share of voice moves with the roster; re-attribution can move past mentions.
9 Is the sample small? "Low sample" labels, hollow dots on the trend Below 30 observations a number is provisional.
10 Is it ordinary variability? Trend across the whole window Engines can answer the same question differently from one run to the next.

The sections below explain each step and what to conclude from it.

Demo data. The filter bar and the change notes on Overview are the first two places to look.

Step 1: Are the two windows what you think they are?

A 7, 28 or 90 day window ends yesterday, and each change is measured against the same number of days immediately before it. An answer collected today does not count until tomorrow. If you compare a screenshot taken last Tuesday with a screen you opened this morning, the two windows overlap only partly, and they cover different dates from what you remember.

Step 2: Are the filters identical?

The filter bar holds six chips: date range, engine, tag, country, persona and language. Your selection carries from screen to screen, so a chip set during an earlier investigation can quietly narrow today's numbers. A chip also applies only on screens where every number honours it, so two screens can legitimately differ. Click Reset, then re-apply only the filters you intend, before comparing two values. The first results guide lists which screens honour which chips.

Step 3: How many answers are analysed?

Mention extraction runs as a separate step after collection, so the analysed answers in a window are always a subset of the collected ones. When some answers have not been analysed yet, Overview says so: "Brand figures use N of M answers; the rest are still being analysed." Visibility uses the smaller base, while the Cited rate uses the collected answers. Two numbers on one screen can therefore disagree without either being wrong; the metrics reference works this through with 100 collected answers, 80 analysed, and a 50 percent visibility next to a 10 percent citation rate.

Step 4: Did the engine mix change?

An engine joining, leaving or running on only a few days changes the pool. Which engines run for you depends on your plan (see plans and limits) and on whether each engine is available at all; an engine that is unavailable on the platform does not run for anyone until it is back. Past answers keep the engine's name, but no new ones arrive.

DiscoveredBy guards its own change figures against this. An engine counts as measured in a period only when it ran on at least 80% of that period's answered days. When an engine was not measured in both periods, the change compares only the engines that were, while the number itself still covers every engine. Overview says so in a note under the strip's header. That guard does not protect a comparison you build yourself: in the Explorer, a comparison uses whichever engines you select, so break the result down by engine or filter to one engine to compare like for like. Engine collection details are on the engines reference.

Step 5: Did any runs fail?

A failed run is absent from every answer metric, not counted as a miss. Failed runs are retried later the same day at 03:00, 05:00, 09:00 and 17:00 UTC; one that still fails after the last retry stays failed for that day. So failures do not push your percentage down directly. They shrink the sample, and if one engine failed for several days, the mix across engines shifts. Overview lists each engine with failed runs in the window as its failed runs out of its runs, for example "Perplexity 6 of 315".

Step 6: Did Google show no AI answer?

Google AI Overviews and Google AI Mode can return a successful request in which Google showed no AI answer. That run is not a failure and not an answer: it enters no visibility, position, share of voice or citation calculation, on either side. Instead it counts toward the AI answer shown rate. A Google engine whose runs in a window all showed no AI answer has no value there, not 0%.

Step 7: Did the prompt set or its targets change?

Every prompt runs once a day for each active target, where a target is one location, one audience and one language. Adding a country, a city, a persona or a language adds targets. Pausing a prompt stops future runs but keeps its history. Either way the collection of questions behind the score changes, and the score is a share pooled over whichever questions ran in the window. Some variants do not run on some engines; a persona variant, for example, never runs on the two Google engines, ChatGPT (app) or Gemini (app), so adding personas changes the engine balance too. Adding many easier or harder questions moves the pooled share with no change in how any single question is answered.

Step 8: Did the competitor roster or brand names change?

Adding a tracked competitor changes share of voice while the answers themselves stay the same, because the denominator is a sum across every active tracked brand family. Pausing one has the opposite effect. Separately, when you add an alias or a sub-brand, the platform checks past untracked mentions for an exact, case-insensitive match and moves the matching rows onto the new name. That means a past window can read differently after you edit brand names. The brands and sub-brands page explains what moves and what does not.

Step 9: Is the sample small?

Below 30 observations a number is shown but treated as provisional. On the trend chart, a point with fewer than 30 observations is drawn hollow, and Overview shows "Low sample (n = 12)" under such a number. If your visibility rests on a few dozen answers, one extra or missing answer moves the percentage by a visible amount. Do not read a small-sample swing as a change in perception.

Step 10: Is it ordinary variability?

Engines do not answer the same way every time. Each tracked prompt runs again every day, and an engine's answer to an unchanged question can differ from one run to the next even when nothing about your site, your competitors or the prompt has changed. A single day's swing is not, by itself, evidence that anything is different. Read the trend across the whole window; a gap in a line means no measured value, never a drop to zero.

A change log you can copy

Keep one entry per investigation. It forces the window and filters to be recorded next to the numbers.

AI visibility change log

Metric:                     (brand visibility / share of voice / position / cited)
Old value, window, filters: 
New value, window, filters: 
Same filters both times?    (yes / no; if no, reset and re-read)
Analysed vs collected:      (N of M answers, old and new)
Engines in each period:     (added / removed / run on few days)
Failed runs or no AI answer: (engine, count)
Prompts and targets changed: (added / paused / new countries, cities, personas, languages)
Competitors or brand names changed: (added / paused / aliases / sub-brands)
Sample size (old / new):    
Explained by:               (list steps 1 to 9 that apply)
Left unexplained:           (size in points)
Next step:                  (wait a window / re-read by engine / act on content)

Worked example: Quillstone's visibility falls ten points

Illustrative example: Quillstone and its competitors are fictional, and the numbers are made up to show the method.

Quillstone sells document-review software to legal and compliance teams. On Overview, its visibility reads 40.0% for the latest 28 days, and the previous window's report said 50.0%. It looks like a ten point fall, and nothing changed on quillstone.example. The team works down the checklist.

Steps 1 and 2 pass: both windows are 28 days, and all chips are reset. Step 3 shows all answers analysed. Step 5 shows no failed runs of note. Step 4 is where the difference appears: Gemini (app) was added to the plan partway through the latest window. Breaking the result down by engine in the Explorer gives this.

Engine Previous: named / analysed Previous rate Latest: named / analysed Latest rate
ChatGPT (app) 44 / 80 55.0% 40 / 80 50.0%
Perplexity 36 / 80 45.0% 32 / 80 40.0%
Gemini (app) not run none 8 / 40 20.0%
All engines 80 / 160 50.0% 80 / 200 40.0%

The checks on the totals: previously 44 + 36 = 80 named of 80 + 80 = 160 answers, which is 50.0%. In the latest window, 40 + 32 + 8 = 80 named of 80 + 80 + 40 = 200 answers, which is 40.0%.

Now compare only the engines measured in both periods. Latest, ChatGPT (app) and Perplexity together: 72 named of 160 answers, which is 45.0%. Previous, the same two engines: 80 of 160, which is 50.0%. That is a 5.0 point drop, and it is the change Overview would show beside the 40.0%, because it compares only the shared engines. The other 5.0 points of the 10 point gap come from the added engine, whose rate is lower and which now makes up a fifth of the pool (40 of 200).

So half of the apparent fall is the engine mix. The remaining 5.0 points on shared engines is the part worth further attention. With 80 answers per engine per window, the sample is modest, and step 10 applies. The team records both parts in the change log, leaves the 5.0 points as unexplained, and re-reads after the next window before touching any page.

Note what the team did not do. They did not conclude that Gemini (app) "ranks Quillstone lower", because a rate on an engine describes what that engine's answers named, not why. And they did not rewrite content in response to a fall that was partly a change of measurement.

Common mistakes and limits

  • Comparing without the window and filters. The commonest error. A percentage without its window and chip settings cannot be checked.
  • Treating a checklist result as a cause. Ruling out steps 1 to 9 leaves variability as a candidate, not a finding. Engine answers are observations, and this checklist cannot show what drove a given one.
  • Reading one day. A daily figure moves on its own. Judge the trend, and remember a day with too few observations is drawn hollow.
  • Trusting a pooled number without the engine breakdown. A pooled figure hides mix changes like the one above. Sentiment is never averaged across engines for a related reason.
  • Comparing different metrics. Visibility divides by analysed answers, citation rate by collected answers, and observation counts from different metrics are not the same denominator.
  • Editing the setup and the content in the same week. If you change the prompt set and publish an update together, you will not be able to separate them. To judge a deliberate content change, see did your content update help.
  • Overstating the fix. Keeping windows, engines and prompts stable makes changes easier to read. It does not make any single engine's answers repeatable.

Frequently asked questions

Can my AI visibility score change even if my website is unchanged?

Yes. The score is the share of analysed answers that name your brand, so anything that changes which answers are in the window changes it: engines, prompts, filters, failed runs and the date window. Engines can also answer the same prompt differently on different days.

Why does my visibility differ from my citation rate on the same screen?

They divide by different populations. Visibility divides by analysed answers and citation rate by collected answers, and mentions and citations are different things: a brand can be named without a link, and a domain can be linked without being named. The difference between mentions and citations is covered separately.

Does a failed run lower my score?

No. A failed run never reaches any answer metric, so it is absent rather than counted as a miss. It can still reduce your sample size, and repeated failures on one engine can change the engine mix.

Why did my share of voice move when I added a competitor?

Share of voice divides by a sum across every active tracked brand family, so a new competitor adds its own count to the denominator. Nothing in the answers changed; the roster that sets the denominator got bigger. Pausing a competitor works the other way.

How many answers do I need before I trust a change?

Below 30 observations a number is treated as provisional and labelled a low sample. Thirty is a floor for reading a number as settled, not a guarantee. A larger sample reduces noise but does not remove the run-to-run variability of the engines.

Could a missing day in the data cause a change?

A day with no measured value is a gap in the trend, never a drop to zero, and an engine counts as measured in a period only if it ran on at least 80% of the answered days. Missing days therefore affect which engines are compared.

Next step

Before you act on a move in your score, run the checklist against it and record the result in the change log. In DiscoveredBy, Overview, the Explorer and each prompt's run history hold the evidence for steps 1 to 9. See how the platform tracks brand visibility in AI visibility, then sign in to review your latest window.

  • prompt tracking
  • measurement
  • ai visibility score
  • troubleshooting
  • engine coverage

Share

Summarize with AI

Start monitoring your AI visibility.

See how AI search engines talk about your brand.

Free to start. No credit card required.