Did your content update help? A practical before-and-after review

Log the edit date, then compare the same prompts, engines and window before and after. Improvement, no change and too little evidence are all valid results.

Kamal 15 min read
A hand-drawn timeline with a single coral marker at the edit date, and two matching bracketed windows on either side of it.
On this page
  1. In short
  2. What does a before-and-after review actually compare?
  3. What do you record before you make the edit?
  4. How do you choose the windows, prompts and engines?
  5. What does the experiment log look like?
  6. What does an improvement look like?
  7. What does "no clear change" look like?
  8. What does "insufficient evidence" look like?
  9. What can make a review insufficient even with plenty of answers?
  10. How do you stop a coincidence from looking like a win?
  11. Common mistakes
  12. What this cannot tell you
  13. Frequently asked questions
  14. Next step

To find out whether a content update helped, write down the date the edit went live, then compare the same prompts, on the same engines, over two windows of equal length: one before the edit and one after. Report what you see as improved, worse, not clearly changed, or not enough evidence to say. Each is an honest result. What you cannot do is treat a change as proof that the edit caused it, because other things move at the same time. This post gives you an experiment log you can copy, and three worked entries showing each outcome.

In short

  • Fix four things before you look at any result: the prompts, the engines (collection channels), the metric, and two equal-length windows either side of the edit date.
  • A change in a number is an observation. Whether the edit caused it is a separate claim, and a before-and-after comparison alone cannot settle it.
  • "No clear change" and "not enough evidence" are results, and they are not the same result. Log both.
  • Write the decision rule in the log before you open the numbers. The rule and the log in this post are our suggested method, not a built-in product feature.

What does a before-and-after review actually compare?

It compares one measured population of answers before an intentional edit with the same population after it. "Same population" is the whole idea: the same tracked prompts, the same engines, the same metric, and windows of equal length, so that a difference cannot be explained by having quietly measured something else.

A prompt is a question DiscoveredBy asks AI engines on your behalf, chosen because it is the kind of thing your buyers ask; it is not a keyword (see key terms). An engine here means one collection channel, for example ChatGPT (app) or Perplexity; a provider's app and API channels are separate engines (see the engines reference). An edit date is the day the changed page was live on your site, not the day someone finished the draft.

This post is about evaluating an edit you made on purpose. If you are still deciding which page to edit, start with finding the citation gaps that deserve your next update. If you want the definition of each metric first, see what each AI visibility number tells you.

What do you record before you make the edit?

Record the edit itself, the hypothesis, and the measurement plan, before the edit goes live. A review written afterwards tends to pick whichever metric and window happened to look best.

At minimum:

  • The page URL and a one-line description of what changed.
  • The hypothesis, phrased as something that could turn out wrong. "Adding a step-by-step comparison table will get this page cited for the three prompts below" is a hypothesis. "Improve the page" is not.
  • The exact prompts the page is meant to win.
  • The engines you will judge it on.
  • The metric, with its denominator.
  • The date the edit went live, and the two windows.
  • Anything else you are changing at the same time.

If you use Optimizations, some of this is captured for you (it is not available on every plan; see plans and limits). Marking a fix applied records the trailing month of visibility for the fix's prompts, and later checks measure the same prompts against a baseline taken from the 30 days before the day you marked it applied. That makes the day you click Mark applied the boundary between "before" and "after". Click it when the edit goes live, not when you remember to.

Note that this is a different measurement from the Explorer windows below. Optimizations always takes its baseline from the 30 days before you marked the fix applied and reads a firm verdict at the one-month mark. The Explorer's presets are 7, 28 or 90 days, and a custom window can be any length up to 366 days. The 28-day windows in this post's examples are our choice for a manual review, not a product setting, and the two methods will not give identical figures.

Two equal windows meet at the edit date; the prompts, engines and metric stay fixed.

How do you choose the windows, prompts and engines?

Use two windows of equal length that meet at the edit date, keep the prompt list frozen, and select engines explicitly.

Windows. In the Explorer, a preset window (7, 28 or 90 days) always ends yesterday, so the same saved query moves forward every day. A custom window stays on its dates. For a log, use custom dates so the entry still shows what you measured in six months. The Explorer offers Compare with: Previous period, which is the same number of days immediately before the window. If the post window starts on the edit date, the previous period is exactly your "before".

Windows must end yesterday or earlier, so you cannot judge an edit on the day it ships. Length is your call; the examples below use four weeks either side.

Prompts. Filter the Explorer by prompt (the Prompt dimension) and pick only the prompts the page is meant to win. Do not delete or rewrite those prompts during the review. A deleted prompt's answers are gone, and rewording one means you are no longer measuring the same question.

Engines. Choose the engines in the filter, then break the result down by engine as well. The Explorer compares whatever engines you select in both periods. If you pool every engine and one of them started answering partway through (as an app engine does in the window it is added), it counts on one side only. Breaking down by engine, or filtering to one, compares like for like. See the engines reference for how each channel is collected.

Metric. Name one. Citation rate (collected answers that cite your domain, divided by collected answers) fits an edit meant to earn a citation. Brand visibility (analysed answers naming your brand, divided by analysed answers) fits an edit meant to get the brand named. They have different denominators, so do not mix them in one entry. The metrics page has both definitions.

What does the experiment log look like?

One row per edit, filled in before and after. Copy this into a spreadsheet or your tracker.

EXPERIMENT LOG: one entry per edit

ID:                     (e.g. EXP-001)
Page URL:
What changed:           (one line)
Hypothesis:             (falsifiable, one sentence)
Edit live on:           (date)
Logged by / reviewed on:

MEASUREMENT PLAN (fill in BEFORE looking at results)
Prompts (frozen):       1.  2.  3.
Engines (channels):     (each one named)
Metric + denominator:   (e.g. citation rate = answers citing our domain / collected answers)
Before window:          (start to end, N days)
After window:           (start to end, N days, same N)
Decision rule:          (e.g. "improved" = pooled change of at least 5 points, same direction
                        on every engine, and at least 30 answers per window)

OTHER CHANGES IN THE WINDOWS
Other edits to this page or its links:
Prompt, competitor, engine or plan changes:
Known collection gaps:

RESULT
                        Before        After        Change
Pooled  (n / N, %):
Engine 1 (n / N, %):
Engine 2 (n / N, %):
Fewer than 30 answers in any figure?   yes / no

OUTCOME (pick one):     Improved / No clear change / Worse / Insufficient evidence
What we can say:        (observation only)
What we cannot say:     (cause)
Next action:            (keep, extend the window, revisit, or try a different change)

The "Other changes" block matters more than it looks. It is where you admit that the pricing page was also rewritten that month.

What does an improvement look like?

It looks like a change large enough to clear your own rule, in the same direction on every engine, with enough answers behind it that a couple of chance answers cannot explain it.

Quillstone sells document-review software to legal and compliance teams. On 2 March it added a step-by-step comparison table to its contract review checklist page. The hypothesis: the page will be cited more often for three prompts about reviewing contracts. Metric: citation rate. Engines: ChatGPT (app) and Perplexity. Assume each prompt runs once per engine per day (a simplification; your own run counts can differ, for example with persona or city targets), so each 28-day window holds 3 prompts x 2 engines x 28 days = 168 answers. The decision rule, written down first, is our suggested example rule, not a product threshold: at least 5 points pooled, same direction on both engines, at least 30 answers per figure. "Worse" is the mirror image: a fall of at least 5 points, down on both engines.

Before (2 Feb to 1 Mar) After (2 Mar to 29 Mar) Change
Pooled 21 of 168 (12.5%) 39 of 168 (23.2%) +10.7 points
ChatGPT (app) 9 of 84 (10.7%) 19 of 84 (22.6%) +11.9 points
Perplexity 12 of 84 (14.3%) 20 of 84 (23.8%) +9.5 points

The engines add up to the pooled figures (9 + 12 = 21 and 19 + 20 = 39; 84 + 84 = 168). Every figure stands on at least 84 answers, so none is provisional. The Explorer reports the change in percentage points.

Outcome: improved. What Quillstone can say: over these three prompts and two engines, the page's citation rate was higher in the four weeks after the edit than in the four weeks before, on both engines. What it cannot say: that the table caused it. Perhaps Brieflane, a fictional competitor, took its own checklist page down that month. The log says so under "Other changes".

Demo data. Citation rate for three prompts during September 2–29, 2026, compared with the previous 28-day period and broken down by engine.
Improved, no clear change, insufficient evidence: three valid entries in the same log.

What does "no clear change" look like?

It looks like a movement smaller than your rule, and it is a real result: the edit did not visibly shift this metric on these prompts in this period.

Quillstone edited its integrations page on 16 March, adding supported actions and setup limits. Same design: 3 prompts, 2 engines, 28 days a side, 168 answers per window.

Before (16 Feb to 15 Mar) After (16 Mar to 12 Apr) Change
Pooled 15 of 168 (8.9%) 16 of 168 (9.5%) +0.6 points

A gain of one citing answer, 15 to 16, is +0.6 points. That is well under our 5-point example rule, so the outcome is no clear change.

This is not "the edit failed". The page may still be more accurate for the buyers who read it. Optimizations makes a similar distinction with its own thresholds: at the one-month mark a shift smaller than five points either way reads "neutral", which means nothing definitive happened yet, not necessarily that the fix failed.

A flat result is still worth logging: it tells you not to repeat that kind of edit across ten more pages without a better reason. The next action might be a longer window or a different change.

What does "insufficient evidence" look like?

It looks like a number that moved a lot on very few answers. The right response is to keep the log entry open, not to declare a winner.

Quillstone edited one page on 30 March and, five days later, asked whether it worked. One prompt, one engine (Perplexity), one answer a day. Baseline: 28 days, 28 answers. After: 5 days, 5 answers.

Before (2 Mar to 29 Mar) After (30 Mar to 3 Apr) Change
Perplexity 3 of 28 (10.7%) 2 of 5 (40.0%) +29.3 points

+29.3 points looks dramatic. It rests on two answers. If one of those had not cited the page, the after figure would have been 1 of 5, or 20.0%; if a third had, 3 of 5, or 60.0%. Each answer is worth 20 points here, so the figure can swing enormously on luck alone. Both figures are also under 30 observations, which the Explorer marks provisional: a hollow dot on a line, a hatched bar or cell, or the word "provisional" in a table.

Outcome: insufficient evidence. Say that, then extend the after window to a length that matches the before window and review again.

The Optimizations check shows similar restraint, on its own rules. Under a week since applying, it always reads "pending more data". From a week to a month, it stays pending unless visibility has already jumped by ten points or more. With no completed answers for the fix's prompts in the period measured, it records no score and stays pending at any age, rather than reading as a drop.

What can make a review insufficient even with plenty of answers?

A mismatch between the two windows can, even when both are large. Check these before you trust any before-and-after figure:

  • An engine was measured in only one window. Optimizations counts an engine only if it ran on the fix's prompts on at least 80% of the days those prompts have any answer, in each of the two periods, and uses only the engines measured in both. In the Explorer, do the same by hand: break down by engine.
  • The prompt list changed. Added, deleted or reworded prompts mean the two windows ask different questions.
  • Collection gaps. Days with failed runs shrink a window's denominator. Compare answer counts first.
  • Today's labels. Tags and classifications are read as they are now, even for past answers, so filter by the prompts themselves.

When in doubt, the metrics guide on reading any number covers what "no data" and "provisional" mean.

How do you stop a coincidence from looking like a win?

You cannot fully stop it, but you can make a coincidence harder to miss. Three habits help.

  1. Keep an unedited comparison set (our suggested method). In the Explorer, run the same query for prompts tied to pages you did not touch, over the same dates. If their citation rate moved by a similar amount over the same dates, the edit is a weaker explanation. This is a manual comparison you choose to run; the product does not pair edited and unedited pages for you.
  2. Log concurrent changes. Other edits to the page, new internal links, a prompt change, a competitor launch, a new engine channel.
  3. Review again after a second window. A one-window bump that vanishes is what variation looks like.

A citation is an observation of one answer, not proof of why the model produced it. "Cited more often after the edit" describes your collected answers; "the edit made engines cite it" claims something neither the Explorer nor Optimizations measures.

Common mistakes

  • Choosing the metric after seeing the results.
  • Comparing windows of different lengths, or pooling engines that were not all running in both.
  • Reading a provisional figure as settled.
  • Editing the page again mid-window. Start a new entry with a new edit date.
  • Calling a mixed result (one engine up, another down) "improved". Under the rule above it is no clear change, so log the split.
  • Reporting only the entries that worked.

What this cannot tell you

A before-and-after comparison shows that two periods differ. It does not show why. Engines can answer the same prompt differently from one day to the next, competitors publish, and your own pages change. A window can also be too short or too thin to show an effect that exists. Keep the wording of your report on the observation: "the citation rate was higher after the edit", not "the edit raised the citation rate".

Frequently asked questions

How long should I wait before judging a content update?

There is no universal number in the product docs. The practical rule is to compare windows of equal length and to wait until each figure rests on enough answers that a single answer cannot swing it. As a reference point, Optimizations treats under a week as too early to say, and its firm read comes at the one-month mark. Those are its own rules, not a general standard.

Should I use Optimizations or the Explorer for the review?

Use Optimizations when the edit came from one of its recommendations: it records the applied date and reads the outcome as pending, positive, negative or neutral. Use the Explorer when you want to pick the metric, engines and windows yourself, or when the edit did not start there. They can agree, and your log should name which one produced each number.

Can I run this review without Optimizations?

Yes. The Explorer is on every plan, and every project member can build and run a query. Choose the metric, filter to the prompts and engines, set a custom window, and turn on Compare with: Previous period. The experiment log then holds the plan and the decision that the tool does not record for you.

What if I edited several pages at once?

Then you have one bundle of changes, not several experiments. Log it as a single entry, list every page under "What changed", and do not credit any one page with the movement. If you need to learn which edit matters, change one page at a time.

Can I save the comparison so my colleagues can reopen it?

Yes, in two ways. The query is in the page's address, so copying it gives another project member the same query (they need to be signed in). Owners and editors can also save a query as a view. Use a custom window so the saved comparison keeps its dates.

Next step

Pick one page you plan to edit next. Fill in the plan half of the log first, publish the edit, and set a review date one full window later. If you track your prompts in DiscoveredBy, the Optimizations workflow records the applied date, and the Explorer lets you build the before-and-after comparison on your own metric and engines. Sign in or start here to set up the comparison. For a first-month plan around this, see your first 30 days of AI search optimization.

  • measurement
  • content updates
  • before and after
  • experiment log
  • citation rate

Share

Summarize with AI

Start monitoring your AI visibility.

See how AI search engines talk about your brand.

Free to start. No credit card required.