How to publish an AI search experiment when the result is no clear change
A null result is still a finding. Learn how to write up an AI search experiment that showed no clear change: hypothesis, method, uncertainty and learning, with a copyable template.
On this page
- In short
- Why should you publish a result that shows nothing?
- What does "no clear change" actually mean?
- What must you decide before the experiment starts?
- How do you collect the evidence?
- How do you present uncertainty without hedging everything?
- The deliverable: a null-result article template
- A worked example: a comparison table that changed nothing we could see
- Common mistakes and what this cannot tell you
- Frequently asked questions
- Next step
To publish an AI search experiment that showed no clear change, write it up like any other result: state the hypothesis before the outcome, describe exactly what changed and how you measured it, report the size of the difference alongside how much data sat behind it, and say what you would do next. "No clear change" is a finding about one edit, one set of prompts and one time window. It is not proof that the edit does nothing, and the write-up should say so plainly, without dressing a flat result up as a win or a failure.
In short
- A null result is a finding, but only if the test could have shown a change. Say what would have counted as one before you look.
- Report the method in enough detail that another team could rerun it: page, edit, prompts, engines, dates, and the comparison window.
- State uncertainty in words and numbers: how many answers, how big the difference, and what you could not control.
- Separate what you observed from what you infer. Observed: visibility moved by three points. Inferred: the edit had no effect. Only the first is a fact.
- End with a decision: repeat, extend, change the edit, or move on. A write-up without a next step is a diary entry.
Why should you publish a result that shows nothing?
Because a hidden null result leaves the next person to repeat the same test, and because a public record of flat results keeps the wins honest. If a team only writes up experiments that moved, the reader sees a biased sample and starts to believe every page edit works.
Internally the case is stronger still. A stakeholder who approved a content change will ask what happened. "Nothing we could measure, and here is how sure we are" is a far better answer than silence, and it stops the same idea coming back next quarter with no context.
There is a second reason specific to AI search. Answers are generated, and the same prompt can produce a different answer each time it runs. A small movement in a visibility figure can be ordinary run-to-run variation. Writing down the size of the difference next to the number of answers behind it is how you show readers you knew that.
What does "no clear change" actually mean?
It means the difference you measured was too small, or the data too thin, to tell an effect from ordinary variation. That is not the same as "no effect", and the write-up has to keep the two apart.
There are at least four different situations that all look like a flat chart:
- The edit did nothing that engines respond to. Possible, but you cannot conclude it from one test.
- The edit did something, but the window was too short. Engines may not have re-read the page yet, or your measurement stopped early.
- The measurement was too coarse. Too few prompts or answers to see a small shift.
- Something else moved at the same time. A competitor launched, a prompt set changed, an engine changed its behaviour, or collection was incomplete.
Your write-up should say which of these you can rule out and which you cannot. That list is the difference between a useful null result and an unhelpful shrug. For the wider habit of reading a change against a baseline, the metrics reference opens with a short guide to reading any number, including the point that a metric which reports an observation count is provisional below 30 observations.
What must you decide before the experiment starts?
Write down the hypothesis, the decision rule and the analysis plan before you look at any results. Choices made after seeing the numbers tend to bend toward whatever looks interesting, and a reader cannot tell the difference.
At minimum, fix these:
- The hypothesis, as a sentence that could be wrong. "Adding a comparison table to the pricing-comparison page will increase how often engines cite that page for these prompts." Not "improve our AI visibility".
- The single change. One edit to one page, or one clearly bounded set. If you changed the title, the table and the FAQ together, you can only test the bundle.
- The prompts and engines. The same set before and after. Adding or removing prompts mid-test invalidates the comparison.
- The measure. Choose one primary measure (for example, the share of answers that cite the page) and name it. Extra measures are secondary and should be labelled that way.
- The windows. A baseline period before the change and a comparison period after it, of stated length.
- What would count as a change. A threshold agreed in advance. Below it, you report "no clear change".
DiscoveredBy's own outcome check is a useful model here, because its rules are fixed in advance and written down. It compares visibility for the fix's prompts against a baseline measured over the 30 days before you marked the fix applied. Under a week after applying, it always reads "pending more data". From a week to a month it stays pending unless visibility has already risen ten points or more, which reads "positive" early. At the month mark, five points or more up reads "positive", five or more down reads "negative", and anything smaller reads "neutral". Its docs are explicit that "neutral" means nothing definitive happened yet, not necessarily that the fix failed. Those are that feature's thresholds, described on the Optimizations page; yours can differ, but you need some rule, chosen early.
How do you collect the evidence?
Use the same prompts on the same engines across both windows, keep a dated log of everything that changed, and export the underlying rows so your numbers can be checked by someone else.
A monitoring tool that runs your prompts on a schedule gives you the before and after series. Depending on your plan (see plans and limits), the Exports screen in DiscoveredBy offers thirteen datasets, each downloadable as CSV or JSON, and the Answers, Citations and Daily metrics datasets are the natural ones for an experiment write-up. Three details from that page matter for method:
- Windowed datasets follow the date range you pick. Datasets 2 through 13 are windowed; Prompts is the only one that is not, because it is your current configuration. Export the Prompts dataset on the change date and keep it, so you can show the prompt set did not drift.
- Daily metrics counts every attempted response, including failed and still-running ones, while brand mentions, citations, cited URLs, domains and answers count completed answers only. If you mix them in one calculation, you are dividing by different denominators.
- Answers is capped more tightly than the rest. It holds up to five thousand rows and cuts any response over eight thousand characters, recording the true length in
response_chars. A download over a dataset's cap is refused, not truncated, so narrow the date range if you hit it. - Read CSV by header name, not by column position. The docs note that columns have been inserted mid-row when features shipped, so a script that reads by position can silently misalign.
Also record the collection context: which engines ran, and whether they were app or API collections. The engines reference is the source for that, and the collection completeness post covers how to check that the monitoring you intended actually ran.
If the experiment is an on-page edit, the Optimizations screen keeps a record of when a fix was marked applied and a downloadable task packet of what was proposed. That packet is a copy taken at download time, so save it with the write-up. For the surrounding question of whether an edit helped, see Did your content update help?, which covers reviewing an intentional edit before and after; this post is about what to publish once you have done that review and the answer is flat.
How do you present uncertainty without hedging everything?
Give the reader three things in one place: how many answers each window contains, the size of the difference, and the list of things you could not control. Then make one plain statement about what the data supports.
A workable pattern is to put a small table at the top of the results section:
| Measure | Baseline | Comparison | Difference | Answers behind it |
|---|---|---|---|---|
| Primary measure | value | value | points | count per window |
| Secondary measure (label it) | value | value | points | count per window |
Then write one sentence per row in this form: "The difference was X points on N answers per window, which is below the threshold we set in advance, so we report no clear change."
Avoid three habits that undermine a null-result article:
- Explaining the flat result away. "It would have worked if the window were longer" is a hypothesis, not a finding, unless you plan to test it.
- Promoting a secondary measure. If the primary measure was flat and a secondary one moved, report the secondary as exploratory and say it was not the planned test.
- Claiming a cause for a wobble. If the figure moved a little, do not name the reason. A citation or a stated reason inside an answer is an observation of that answer, not proof of what produced it.
The deliverable: a null-result article template
Copy this template, replace the bracketed parts, and keep the headings. It is designed so the first paragraph can stand alone if it is quoted.
TITLE: We tested [change] on [page or page set]. Here is what we found.
OPENING ANSWER (2 to 3 sentences, stands alone):
We [made this change] to [page] and compared [primary measure] for
[N] prompts on [engines] over [baseline dates] and [comparison dates].
We saw no clear change: the difference was [X points] on [N] answers per
window, below our pre-set threshold of [Y points]. This does not show the
change has no effect; it shows this test did not detect one.
IN SHORT:
- Hypothesis: [one sentence]
- Result: no clear change on [primary measure]
- Confidence: [what limits it, in a few words]
- Decision: [repeat / extend / change the edit / stop]
## The hypothesis
[The sentence you wrote before the test. Date it. Say who wrote it.]
## What we changed
[The single edit. Before and after text or a link to the saved version.
Publish date. Anything else that changed on the page in the same period.]
## How we measured
- Prompts: [count, how chosen, unchanged during the test: yes/no]
- Engines and collection: [which, app or API]
- Primary measure: [definition, in one sentence]
- Windows: baseline [dates], comparison [dates]
- Threshold for "a change": [value, set on date]
- Data: [exports used, saved where]
## What we found
[Table: measure, baseline, comparison, difference, answers behind it.]
[One sentence per row.]
## What could explain a flat result
- Edit had no effect that engines respond to: [can we rule out? no]
- Window too short for engines to re-read the page: [evidence]
- Too few answers to see a small shift: [count, provisional or not]
- Something else changed: [competitor, prompts, engine, collection gaps]
## What we did not test
[Other edits, other pages, other engines, longer windows.]
## What we will do next
[One decision, one owner, one date. Or: why we are stopping.]
## Limits
- These are observations of generated answers, not proof of why an
engine produced them.
- One page, one edit, one period. Do not generalise.
A worked example: a comparison table that changed nothing we could see
Illustrative example: Quillstone and its competitors are fictional, and the numbers are made up to show the method.
Quillstone, a document-review software company, hypothesised that adding a feature-comparison table to its "Quillstone vs Brieflane" page would raise how often AI answers to six comparison prompts named Quillstone. It fixed the rule in advance: a change of five points or more in the share of analysed answers naming Quillstone would count; anything smaller would be reported as no clear change.
The team kept the six prompts and the two engines unchanged, and ran a 30-day baseline and a 30-day comparison after publishing the table. Each prompt ran on each engine five times per window, so each window held 60 analysed answers (6 x 2 x 5).
| Measure | Baseline | Comparison | Difference | Answers behind it |
|---|---|---|---|---|
| Share of answers naming Quillstone (primary) | 35% (21 of 60) | 38% (23 of 60) | +3 points | 60 per window |
| Share of answers citing the comparison page (secondary) | 8% (5 of 60) | 10% (6 of 60) | +2 points | 60 per window |
The arithmetic: 21 of 60 is 35%, 23 of 60 is 38.3%, and the difference is 3.3 points, below the five-point rule. Five of 60 is 8.3% and six of 60 is 10%, a difference of 1.7 points.
The write-up opened: "We added a comparison table to the Quillstone vs Brieflane page and saw no clear change in how often six comparison prompts named Quillstone. The difference was about three points on 60 answers per window, below our five-point threshold. This test cannot tell us the table has no effect."
Its uncertainty section listed what it could not rule out: 30 days may be too short for the page to be re-read, one competitor had published a new page during the comparison window, and 60 answers per window can hide a small shift. The decision was to extend the comparison window by 30 days and add one prompt variant, and to write a second short post whatever the result. Notice what the article did not say: it did not claim the table was useless, and it did not credit the two-point citation change to anything.
Common mistakes and what this cannot tell you
- Moving the goalposts. Choosing the threshold after seeing a three-point rise, then calling it a win.
- Changing the prompt set mid-test. Any change to the prompts breaks the comparison; start a new experiment instead.
- Treating a small sample as settled. Figures built on fewer than 30 observations are provisional; say so.
- Ignoring incomplete collection. If some runs failed or engines started or stopped answering in the window, the two windows may not be comparable. The Optimizations outcome check handles this by comparing only engines measured in both windows; do the same by hand.
- Writing a cause into a flat result. A stated reason or cited source in one answer describes that answer, not why the model behaved as it did.
- Generalising. One page on one set of prompts does not tell you what a table does on other pages. For a different question, see Does updating a page change its AI citation performance?.
A published null result also cannot tell readers whether a change would help their site. It is evidence about one test, and that is all it should claim.
Frequently asked questions
Is "no clear change" the same as "it didn't work"?
No. It means the test did not detect a difference larger than your pre-set threshold. The edit could still have a small effect, an effect that appears later, or an effect on prompts you did not track. Say which of those you can and cannot rule out.
How long should I run an AI search experiment before calling it flat?
There is no universal number, so set it before you start and write it down. Consider that DiscoveredBy's own outcome check treats under a week as too little to say anything, calls only a large lift (ten points or more) early, and otherwise waits for the month mark before giving a verdict; use that as one reference point for your own rule, not as a standard.
Should I publish flat results publicly or keep them internal?
Internal by default; public when the method is worth sharing and the result contains nothing confidential. Either way, use the same structure. If you publish, leave out client-specific and other confidential data.
What if the visibility number moved but I don't trust it?
Check the denominators and the collection first. Look at how many answers sit behind each figure, whether any engine started or stopped answering, and whether the prompt set changed. The visibility score changed post covers diagnosing a movement when your website did not change.
Can I combine several small experiments to find an effect?
Only if you planned to. Pooling flat tests after the fact, or picking the ones that moved, is the same goalpost problem at a larger scale. If you expect to run several similar edits, define the pooled analysis up front.
Next step
If you want the before-and-after series and the exports for your next write-up, DiscoveredBy tracks your prompts across AI engines, records when a page fix was marked applied, and, depending on your plan, lets you download the underlying datasets as CSV or JSON. Start by exporting your prompt set on the day you make the change, and use the template above to write the result, whichever way it goes. For definitions of the terms used here, see key terms and the glossary entry on on-page optimization.
- optimization
- reporting
- measurement
- experiments