Measure change honestly
Compare like with like: fixed prompt sets, stated denominators, complete collection, and before-and-after reviews that allow "no clear change".
Chapter 6 of 7: contents
To measure change in AI search honestly, compare like with like. Hold the prompts, engines and metric fixed, state what each rate divides by, check that collection was complete in both periods, and write your decision rule before you open the numbers. Then report what you saw as improved, worse, no clear change, or not enough evidence. Each of those is a valid result. What a comparison cannot do is tell you why a number moved: engines answer differently from run to run, competitors publish, and your own setup changes. So the report describes the answers you collected, and any explanation stays a hypothesis to test.
In short
- A rate is a share of answers, so a change in which prompts, engines or days are in the pool moves it with no change in how any engine answered. Freeze a prompt cohort for comparisons over time.
- Write the numerator, the denominator and the window beside every number. Four correct calculations of "citation rate" can give four different answers from the same data.
- A failed run is absent, not a zero. Gaps shrink a sample, and uneven gaps change its mix, so count what was collected before you read any trend.
- Share of voice is relative to the competitors you track; brand visibility is not. Report the two side by side, with the roster.
- A before and after review needs equal windows, a decision rule set in advance and a log of everything else that changed. "No clear change" is a finding to publish, not a failure to hide.
Why does a fixed set of prompts matter?
Because a rate is only comparable over time when its denominator is made of the same questions. Add prompts, and a rate such as brand visibility can fall while nothing about your brand got worse: you asked different, perhaps harder, questions.
Brand visibility is analysed answers naming your brand, divided by analysed answers (Metrics defined). Every new prompt adds answers to the bottom of that fraction. Pausing prompts, adding a country, persona or language, or a change in which engines run has the same effect. Your list should still grow, as choosing the prompts to track explains; the problem is reading a growing list as if it were fixed.
The fix is a prompt cohort: a fixed list of prompts, written down on a known date, that you report on every period and never edit. Beside it runs the live set, every prompt you currently track. The cohort answers "did anything change on the same questions?" The live set answers "how are we doing on everything we now track?" Report both, each with its prompt count. Because what actually runs is a prompt target (one prompt in one location, for one audience and one language), the cohort's register also records the engines, locations, personas and languages it was frozen with.
Illustrative example: Quillstone and its competitors are fictional, and the numbers are made up to show the method. Quillstone sells document-review software to mid-sized legal and compliance teams, and every Quillstone figure in this chapter belongs to this illustration.
On 31 August, Quillstone freezes a cohort of 12 prompts on one engine, ChatGPT (app). Assume every target ran once a day and every answer was analysed. August gives 372 answers, 93 naming Quillstone: 25.0%. In September the team adds 12 prompts about contract redlining, where Quillstone is rarely named. The live set now gives 720 answers, 144 naming Quillstone: 20.0%. The cohort alone gives 360 answers, 108 naming Quillstone: 30.0%. The new prompts give 360 answers, 36 naming Quillstone: 10.0%.
A live-set report says visibility fell five points; a cohort report says it rose five. Both are right, and they answer different questions. The honest report gives both, plus the split, and treats the cohort's rise as an observation, not a cause.
AugustLive list: 12 prompts, brand visibility 25.0%
Cohort, 12 prompts
25.0%
SeptemberLive list: 24 prompts, brand visibility 20.0%
Cohort, 12 prompts
30.0%
Added this month, 12 prompts
10.0%
The live list falls from 25.0% to 20.0%; the cohort rises from 25.0% to 30.0%. Compare the cohort month to month. Illustrative numbers.
Figure 6.2 shows those two months: the cohort holds the same 12 prompts (25.0%, then 30.0%) while the live list doubles to 24 (25.0%, then 20.0%), so the two can move apart for a legitimate reason.
Four rules keep a cohort useful. Choose it before you look at results, across the buying journey, not from the prompts where you already look good. Make it big enough: a rate built from fewer than 30 observations is provisional (Metrics defined), so aim for at least 30 analysed answers per engine per period. Pause cohort prompts, never delete them: the Explorer counts paused prompts that have answers in the window, but a deleted prompt's answers are gone (Explorer). And replace a cohort on a schedule, with a new name and an overlap period, never by editing it in place.
A dated tag, filtered in the Explorer's Tag dimension, is one practical way to mark a cohort. Tags are read as they are now, including for past answers, so removing the tag from a prompt drops it from every past month of the cohort (Explorer), and a tag cannot be renamed once it exists (Tags and keywords). Keep a written register as the source of truth. Use a fixed prompt cohort for month-to-month comparisons has the register and a monthly log.
What is the denominator?
The denominator is the population a rate divides by, and it is the part people most often forget to state. Name its unit (answers, citations, pages or domains) before you compare two figures, because the same answers divided different ways give different numbers that are all correct.
The core metrics do not share a denominator (Metrics defined). A collected answer is a completed run: the engine ran and returned a result. An analysed answer is a collected one whose brand mentions have also been extracted, which happens in a later step, so the analysed population is always a subset of the collected one. Brand visibility divides by analysed answers. Citation rate is collected answers that cited your domain, divided by collected answers. Share of voice divides by a sum across every active tracked brand family. So visibility can read 50 percent and citation rate 10 percent on one screen for one window, and both be right.
Citations add a trap of their own, because one answer can hold many links. Citation rate denominators explained compares four ways to divide the same evidence:
- Answer level: answers citing your domain, divided by collected answers. An answer counts once however many of your pages it links. This is the product's citation rate.
- Citation level: your citations, divided by all citations in view.
- URL level: your distinct cited pages, divided by all distinct cited pages.
- Domain level: your domain, counted once, divided by all distinct cited domains.
Illustrative numbers from one window of fictional Quillstone data: the same answers, four denominators.
12 answers citing your domain
40 collected answers
30%
How often an answer cites you. This is the citation rate.
18 citations to your domain
120 citations in all
15%
Your share of all the links answers gave.
6 of your pages cited
70 distinct pages cited
8.6%
Your share of the distinct pages cited.
1 domain (yours)
25 distinct domains cited
4%
Your domain as one of many cited domains.
Figure 6.1 shows what each of those rates counts above the line and below it. In the Quillstone illustration, one 28-day window holds 40 collected answers, 12 of which cite quillstone.io. Across them sit 120 citations (18 to Quillstone), 70 distinct cited pages (6 of them Quillstone's) and 25 distinct cited domains. The four fractions are 12 of 40 (30%), 18 of 120 (15%), 6 of 70 (8.6%) and 1 of 25 (4%). Nobody made an arithmetic error; the four figures measure four different things.
The habit that settles this is a definition line beside every number: the metric's name, its numerator and denominator in words, the window and the engines. Keep one fixed sentence per metric and paste it unedited under every report, so a change of method shows up as an edit, not a jump in the line. And never write a zero for a period with nothing to measure: a metric whose denominator is empty reads as no data, not 0%.
What do missing days do to a trend?
A missing day does not lower your rate: a failed run is absent from both sides of the division, not counted as a miss. The danger is subtler. Gaps shrink the sample, and when they fall unevenly on some prompts or engines, the periods you compare are no longer made of the same answers.
Every answer metric counts completed runs only, so a failed run never contributes an answer, a citation or a mention (Troubleshooting). Failed runs are retried the same day at 03:00, 05:00, 09:00 and 17:00 UTC, and one still failed after the 17:00 retry stays failed for that day, so judge a gap only after the last retry. A Google AI Overviews or AI Mode run where Google showed no AI answer is not a gap either: it is a real result that counts toward the shown rate and stays out of answer metrics.
Think of a pooled rate as a weighted average of its prompt groups. In the Quillstone illustration, the team tracks 10 prompts on one engine: five generic "best document review software" prompts (group A) and five "Quillstone versus Brieflane" prompts (group B). In week one every run completes. Group A gives 7 of 35 answers naming Quillstone (20.0%), group B 21 of 35 (60.0%), and the pool 28 of 70 (40.0%). In week two, a provider error on three days hits only group B, and those runs stay failed. Group A is unchanged at 7 of 35. Group B has 12 of 20 (60.0%). The pool is 19 of 55: 34.5%.
The pool shows a fall of 5.5 points, yet neither group changed. The pool fell because week two held fewer answers from the higher-rate group. The supportable report filters both weeks to group A (20.0% and 20.0%) and reports group B separately, marked provisional, since 20 answers is under the floor. What missing collection days do to an AI visibility trend turns this into a checklist with four decisions: compare as is, compare by engine, restrict to the intact prompts, or shorten the window and hold.
To find gaps, count what was collected before reading any rate. In the Explorer, choose Answers collected with a daily grain and an engine breakdown, and compare each day's count with what you expected. Then open the affected prompt's run history, which lists failed runs with a reason drawn from a closed set: timeout, rate limited, provider error or internal error (Troubleshooting).
The expected count is yours to build. A planned run is one active prompt target, on one engine that runs it, for one day; leave out variants that never run on an engine by design, such as persona variants on the Google engines. AI monitoring completeness then defines collection completeness as runs that completed or showed no AI answer, divided by planned runs, with failed and missing runs as their own rates, and analysis completeness as analysed answers over completed ones. There is no documented standard for "complete enough", so set your floor before you see the numbers, and read it per engine: a high blended figure can hide one engine with a whole day missing.
One warning about smoothing. The Explorer's 7-day average pools seven days' numerators over seven days' denominators, so a missing day simply leaves a rate's pool; for answer counts, days with no answers count as zero, so a smoothed count dips across a gap (Explorer). Hunt for gaps on raw daily counts. Smoothing never repairs an uneven mix.
Why can share of voice rise while mentions stay flat?
Because share of voice is a ratio against the competitors you track, not a count of how often you are named. If a rival is named less often, or you pause one, the denominator shrinks and your share grows while the number of answers naming you stays where it was.
Share of voice is distinct answers naming your brand family, divided by the sum, over every active tracked brand family, of that family's own distinct-answer count (Metrics defined). Brand visibility never looks at your roster: it is your count against the analysed total. Visibility has one moving part; share of voice has as many as there are families on the roster.
In the Quillstone illustration, 200 answers were analysed in each of two weeks, and Quillstone was named in 40 both times: visibility of 20.0% in both. In week one the four tracked families (Quillstone, Brieflane, Clausewise and Docket North) were named in 40, 60, 50 and 30 answers, a sum of 180. In week two Brieflane fell to 40 and Clausewise to 30, so the sum fell to 140. Quillstone's share of voice went from 40 of 180 (22.2%) to 40 of 140 (28.6%). Share of voice rose by about six points; visibility did not move.
The accurate report line is not "our share of voice grew". It is: "we were named in the same number of answers; two rivals were named in fewer, so our slice of a smaller total is bigger." That leads to a useful question (why did two rivals drop?) instead of a claim of progress.
The roster moves the number on its own. Adding a tracked competitor adds its count to the sum and can lower your share with no change in any answer (Troubleshooting); pausing one does the opposite. Untracked brands are left out entirely, so a market leader missing from your roster makes your share look healthier than the answers suggest. Report the two metrics as a pair, with the roster for each period and the dates of any roster change, and lead with visibility because it does not shift when you edit the list. Why share of voice can rise while your brand mentions stay flat has a reading table for six common combinations of movements.
How do you read a before and after?
Compare one population of answers before an intentional edit with the same population after it: the same prompts, engines and metric, over two windows of equal length that meet at the date the edit went live. Write the decision rule first, then classify the result as improved, worse, no clear change or insufficient evidence.
Record the plan before the edit ships, because a review written afterwards tends to pick whichever metric looked best. Write down the page and what changed, a hypothesis that could turn out wrong, the prompts the page is meant to win, the engines, the metric with its denominator, the windows, and anything else changing at the same time. The edit date is the day the changed page was live, not the day the draft was finished.
In the Explorer, set a custom window starting on the edit date and turn on Compare with: Previous period, the same number of days immediately before. Windows end yesterday or earlier, so an edit cannot be judged on the day it ships, and a custom window stays on its dates, which is what a log needs. A rate's change is shown in percentage points. Filter to the page's prompts and break the result down by engine.
Before window
Same prompts, engines, length
Edit published
date logged
After window
Same length as before
Edit date. Written down the day the edit went live. Both windows are the same length.
Missing collection days. They add no answers, but can change which prompts and engines make up a window. Compare windows collected the same way.
Engine change. A dated log entry inside the comparison: annotate it or split it.
Figure 6.3 shows the shape of the review, with two things marked that commonly break it. A collection gap shrinks one window's denominator, so compare answer counts in both windows before you compare rates. An engine change matters because the Explorer compares whatever engines you select in both periods, so an engine that started answering partway through counts on one side only (Explorer). Break down by engine, or filter to the engines present in both windows.
In the Quillstone illustration, the team adds a step-by-step comparison table to its contract review checklist page on 2 March, expecting the page to be cited more often for three contract-review prompts. The metric is citation rate, on ChatGPT (app) and Perplexity. With one run per prompt per engine per day, each 28-day window holds 3 x 2 x 28 = 168 answers. The rule, written first and chosen by the team rather than set by any tool: at least 5 points pooled, the same direction on both engines, at least 30 answers behind every figure.
Before the edit, 21 of 168 answers cited the page (12.5%); after it, 39 of 168 (23.2%), a change of 10.7 points. ChatGPT (app) went from 9 of 84 (10.7%) to 19 of 84 (22.6%), and Perplexity from 12 of 84 (14.3%) to 20 of 84 (23.8%). Under the team's rule the outcome is improved. What Quillstone can say: over these prompts and engines, the page's citation rate was higher in the four weeks after the edit than in the four weeks before. What it cannot say: that the table caused it. A competitor might have taken its own checklist page down that month, and the log's "other changes" block is where that goes.
Three habits make a coincidence harder to miss. Run the same query over the same dates for prompts tied to pages you did not touch; if they moved by a similar amount, the edit is a weaker explanation. Log concurrent changes. And review again after a second window, since a bump that vanishes is what variation looks like. On a new project, first check that you have a baseline at all: How much history does a new project need separates an initial observation from a baseline. Did your content update help? has the full method, with worked entries for improved, no clear change and insufficient evidence.
What do you log when an engine changes?
Log a dated entry whenever something that could change your numbers changes: the engine or how it is collected, your own prompts and competitors, or your reporting rules. Each entry carries a source and the comparisons it touches. When a number moves and no entry explains it, write "cause unknown", not a guess.
Keep a measurement change log when an AI engine changes sorts entries into five kinds: engine or provider, collection, availability, your setup, and reporting. Each records the date the change took effect, not the date you noticed it, and a confidence grade: documented, observed in records, or suspected. Suspected entries stay out of client reports. Movements with no matching entry go on a separate list of open questions, so guesses do not harden into history.
The product guards its own comparisons against one kind of engine change. An engine counts as measured in a period only when it ran on at least 80% of that period's answered days, and the comparisons the product makes for you, such as the Overview's changes and the optimization outcome checks, use only engines measured in both periods (Engines and measurement). Comparisons you build yourself, such as an Explorer comparison or a saved dashboard, use whatever engines you select.
In a separate Quillstone scenario, visibility reads 40.0% for the latest 28 days against 50.0% before, while the Quillstone site was unchanged. Broken down by engine, the earlier window had ChatGPT (app) naming Quillstone in 44 of 80 answers and Perplexity in 36 of 80: 80 of 160, or 50.0%. The latest window had 40 of 80 and 32 of 80, plus Gemini (app), added partway through, at 8 of 40: 80 of 200, or 40.0%. On the two engines present in both windows, the figures are 50.0% before and 72 of 160 (45.0%) after. Half of the ten-point fall is the engine mix. The other five points go on the open list until the next window.
Two record fields help separate an engine change from your own. Every run stamps its platform, surface and collection method when it runs, so a later catalogue change cannot relabel an old answer (Engines and measurement). The Explorer offers a Model dimension, shown as "Unrecorded" when empty (Explorer), and the Daily metrics export splits a day into more than one row when the model identifier changes, as it does on every model upgrade (Exports).
That signal has a limit. ChatGPT (app) records its model as chatgpt-app and Gemini (app) as gemini-app, whatever model the data provider reports, and the Google engines record the surface (ai-overview or ai-mode) because Google does not say which model answered (Engines and measurement). An unchanged value on those engines proves nothing. A swing in visibility is no evidence of a new model either; it stays unexplained until a dated source says otherwise. For ruling out causes in order, see why a visibility score can move on its own.
How do you chart change without exaggerating it?
Keep the count behind each point visible, draw missing periods as gaps rather than zeros, mark small samples as provisional, start bar axes at zero, and put a marker wherever the collection setup changed. Each habit stops the drawing from claiming more than the answers support.
A visibility figure built from 60 answers moves by more than a point with one extra mention, so show the observation count as a label or a small series beneath the line, and never put two metrics with different denominators on one shared axis. The Explorer draws a group with nothing in its denominator as no value, an en dash, never 0, and marks a rate from fewer than 30 observations as provisional: a hollow dot on a line, a hatched bar or matrix cell, or the word "provisional" in a table (Explorer). A chart rebuilt elsewhere should do the same, and a real 0% (answers collected, brand not named) must look different from a gap.
A bar's length encodes its value, so a bar chart of a rate starts at zero; a zoomed line needs its range in the caption. Avoid dual axes. The Explorer itself draws several metrics as one panel per metric, each captioned with its unit (Explorer). And say whether a change is in points or percent: 15.0% to 20.0% is 5.0 points and also a one-third relative rise.
In the Quillstone illustration, a weekly chart sent to the board reads "Visibility up a third in six weeks." Brand visibility ran at 9, 10 and 9 of 60 answers in weeks one to three (15.0%, 16.7%, 15.0%), had no completed answers in week four, then 18 of 90 in weeks five and six (20.0%), after a second engine was added in week five. The original engine stayed at 9 of 60 (15.0%); the new one gave 9 of 30 (30.0%). The chart plotted week four as 0%, started its axis at 14%, pooled the new engine without a marker, and called five points "up a third" without saying it was relative. The corrected chart has one line per engine, a labelled break at week four, a marker at week five and a zero-based axis, and its caption says the pooled rise reflects an added engine. Before any chart goes out, run the chart review checklist, including a one-sentence "what this does not show" under the chart.
How do you publish "no clear change"?
Write it up like any other result: the hypothesis as written before the test, what changed and how you measured it, the size of the difference next to the number of answers behind it, what you could not rule out, and what you will do next. "No clear change" is a finding about one edit, one prompt set and one window, not proof that the edit does nothing.
A hidden null result leaves the next person to rerun the same test, and a record of only the experiments that moved makes every edit look effective. A flat chart can come from at least four situations, and the write-up says which it can rule out: the edit did nothing engines respond to; the window was too short; too few answers to see a small shift; or something else moved at the same time, such as a competitor launch, a prompt change, an engine change or incomplete collection.
In the Quillstone illustration, the team expected a feature-comparison table on its "Quillstone vs Brieflane" page to raise how often answers to six comparison prompts named Quillstone. It fixed the rule in advance: five points or more counts; anything smaller is no clear change. Each 30-day window held 60 analysed answers across two engines. Quillstone was named in 21 of 60 before (35.0%) and 23 of 60 after (38.3%), a difference of 3.3 points, below the rule. The write-up said so in its first paragraph, listed what it could not rule out (a short window, a competitor page published mid-test, a modest sample), and ended with a decision: extend the window by 30 days and review again. It neither called the table useless nor credited the small rise to anything.
The product's own outcome check models a rule fixed in advance. Optimizations compares visibility for a fix's prompts with the 30 days before you marked it applied. Under a week it always reads "pending more data", and until the month mark it stays pending unless visibility has already risen ten points or more, which reads "positive" early. At the month mark, five points or more up reads "positive", five or more down "negative", and anything smaller "neutral", meaning nothing definitive happened yet, not necessarily that the fix failed. Your thresholds can differ, but choose them before you look.
Avoid explaining a flat result away, promoting a secondary measure that happened to move, and naming a cause for a small wobble. How to publish an AI search experiment when the result is no clear change has a write-up template, and run it as a program covers where results go in a regular reporting rhythm.
Measurement plan and change log
Fill in the plan half before the edit goes live or the period starts, and the rest when you review. One sheet per question.
MEASUREMENT PLAN AND CHANGE LOG
ID: (e.g. MP-001)
Owner: ____________ Written on: YYYY-MM-DD
1. QUESTION (fill in before looking at any result)
Question: (one sentence, e.g. "Was the checklist page cited
more often after the 2 March edit?")
Hypothesis: (falsifiable, one sentence)
Edit or event, if any: (page URL, what changed, date it went live)
2. COHORT
Cohort name or tag: ____________ Frozen on: YYYY-MM-DD
Prompts (ids and exact wording, or the register file): ____________
Engines (each named, app or API): ____________
Locations, personas, languages held constant: ____________
Competitor roster (needed if share of voice is reported): ____________
Unedited comparison set (prompts for pages not touched): ____________
3. METRIC AND ITS DENOMINATOR
Primary metric: ____________
Numerator (words): ____________
Denominator (words): ____________ (analysed answers / collected
answers / sum across tracked families / other)
Unit counted: answers / citations / URLs / domains
Secondary metrics (labelled as secondary): ____________
Definition line to print beside the number:
"[Metric]: [numerator] divided by [denominator], [window], [engines]."
4. WINDOW
Before window: YYYY-MM-DD to YYYY-MM-DD (N days)
After window: YYYY-MM-DD to YYYY-MM-DD (same N, ends
yesterday or earlier)
Review date: YYYY-MM-DD
5. DECISION RULE (fixed now, not after the numbers)
"Improved" means: (e.g. at least 5 points pooled, same direction on
every engine, at least 30 answers per figure)
"Worse" means: (the mirror of the above)
Anything smaller: no clear change
Fewer than 30 answers in any figure, or completeness below our floor:
insufficient evidence
Completeness floor: ____ % of planned runs, per engine
6. CHANGES DURING THE WINDOW (one line each, dated when effective)
ID Date Kind (engine / collection / availability / our setup /
reporting) What changed Source Confidence
(documented / observed / suspected) Metrics touched
L-001 __________ ____________________________________________________
L-002 __________ ____________________________________________________
Other edits to this page or its links: ____________
Prompts added, paused or reworded: ____________
Competitors added or paused: ____________
Engines added, stopped or relabelled: ____________
7. COLLECTION CHECK
Planned runs per engine: ____
Completed / no AI answer / failed / missing: ____ / ____ / ____ / ____
Collection completeness, per engine: ____
Analysed vs collected answers: ____ of ____
Gap days (date, engine, cause): ____________
8. RESULT
Before (n / N, %) After (n / N, %) Change (points)
Pooled: ________________ ________________ ______
Engine 1: ________________ ________________ ______
Engine 2: ________________ ________________ ______
Comparison set: ________________ ________________ ______
9. OUTCOME (pick one): Improved / Worse / No clear change /
Insufficient evidence
What we observed: (observation only)
What we cannot say: (cause)
Explanations not ruled out: ____________
Next action: keep / extend the window / revisit / try a
different change / stop
Recheck on: YYYY-MM-DD
10. UNEXPLAINED MOVEMENTS (separate list)
ID Date range What moved (metric, engine, counts) Entries checked
U-001 __________ _______________________________________ L-___
Status: open / explained by L-___ / within normal variation /
closed, cause unknown
In DiscoveredBy
Metrics defined gives each metric's numerator and denominator, which ones divide by analysed answers and which by collected answers, and the rules that a number under 30 observations is provisional and that an empty denominator reads as no data rather than zero. The Explorer is where you build the comparison: a custom window that ends yesterday or earlier, Compare with: Previous period, a Tag or Prompt filter for your cohort, and an Engine or Model breakdown so a pooled figure cannot hide a change in the mix. Troubleshooting explains that a failed run is absent from every answer metric rather than counted as a miss, when failed runs are retried, and how each prompt's run history lists failed runs with their reason.
Go deeper
- Did your content update help? A practical before-and-after review
Compare the same prompts, engines and window before and after an edit.
- Why your AI visibility score changed when your website did not
A ten-step checklist covering windows, filters, engines, failed runs and prompt edits.
- Citation rate denominators explained: why two teams get different numbers
Answer, citation, URL and domain denominators, with the definition beside every number.
- Use a fixed prompt cohort for month-to-month comparisons
Freeze a comparison cohort, keep growing the live set, and report both.
- AI monitoring completeness: what percentage of your planned prompt runs actually ran?
Measure completeness before you read any visibility number.
- How to publish an AI search experiment when the result is no clear change
A null result is still a finding, with a copyable write-up template.