How to keep a measurement change log for AI engine changes

A copyable change log for the moments an AI engine, its collection method or your own setup changes, so you know which before-and-after comparisons need a note.

Kamal 17 min read
A hand-drawn timeline with small dated flags along it, one flag in coral, and a ruler laid across the line where it changes direction
On this page
  1. In short
  2. What is a measurement change log?
  3. What kinds of change belong in the log?
  4. Which changes has DiscoveredBy itself documented?
  5. How do you tell an engine change from your own setup change?
  6. Which comparisons need an annotation?
  7. The measurement change log template
  8. Where do you pull the evidence from?
  9. Worked example: Quillstone sees a drop
  10. What can a change log not tell you?
  11. Common mistakes
  12. Frequently asked questions
  13. Next step

To keep a measurement change log for AI engines, write one dated line every time something that could change your numbers changes: the engine or how it is collected, your own prompts and competitors, or the reporting rules. Then, before you compare two periods, check whether a log entry sits between them. If one does, annotate the comparison or split it. Never fill the log by guessing: a swing in visibility is not evidence that an engine shipped a new model, and the log should say "cause unknown" until you have a record that says otherwise.

In short

  • A change log is a dated list of known changes, kept next to your reporting, so a before-and-after comparison is never read across a break nobody wrote down.
  • Log three kinds of change: engine and collection changes, changes to your own setup, and changes to how you report.
  • Record what you know, where you learned it, and which metrics and windows it touches. Leave the cause blank when it is unknown.
  • A visibility swing with no matching entry is an open question, not a model release.
  • The template and the annotation rules below are copyable.

What is a measurement change log?

A measurement change log is a running record of changes that could alter the numbers you report, each with a date, a source and the comparisons it affects. It is not a diary of results. It records causes you can name, and it marks the comparisons that need a caveat.

Two definitions matter for the rest of this post. An engine is an AI answer surface whose answers you track, such as ChatGPT (app), Perplexity or Google AI Overviews (see what an AI engine is). A collection method is how the answer reaches your monitoring: through the company's own API, or through a licensed data provider that fetches the answer a person would see. In DiscoveredBy's engines reference, four engines are collected through their own company's API and four through DataForSEO, a third-party data provider, and the app labels the second group "Licensed data" (engines reference).

Why bother? Because the same chart can be broken in ways the chart cannot show. If an engine's answers changed at the source, or your tool's way of collecting them changed, or you edited your prompt set, the line still looks continuous. The log is what tells the reader it is not.

What kinds of change belong in the log?

Log any change to the engine, to the collection route, to your own setup, or to your reporting rules. The engine is the one people think of first, but your own setup and reporting rules are just as able to move a number, and they are the two you can prove.

Kind Examples Where you would learn of it
Engine or provider An engine is added, removed or relabelled; a provider says a new model is live Provider announcements, your monitoring tool's docs and change notes
Collection A new capture is added (for example fan-out recording), the way location or language is sent changes, a cut-off answer rule is added Your tool's engine documentation and release notes
Availability An engine stops running because its access is removed, or restarts Run history and your tool's engine status
Your setup Prompts added or removed, a persona or language variant added, a competitor added, a city target added Your own project history
Reporting A metric definition changes, a filter default changes, you change the window or the engine set in a saved view Your own report notes

The last two rows are yours to control, so they are the cheapest entries to keep complete. The first three you can only record once you hear about them, which is why every entry carries a source. DiscoveredBy's changelog lists product changes by month, so it is a source for an entry, not an exact effective date.

Which changes has DiscoveredBy itself documented?

The engines reference documents several changes and differences, some with dates, and they show what a real entry looks like. None of them is a claim about how any engine behaves today; they are examples of things a careful log would hold.

  • ChatGPT (API) was removed. ChatGPT is now collected from the app only, and the stored answers of the old API engine were deleted when it was removed. Documents kept as they were sent, such as weekly reports, shared snapshots and alerts, may still name it (engines reference).
  • ChatGPT (app) searches have been captured since 2026-09-29. Answers collected before that date read "not recorded" in the fan-out panel, so any fan-out comparison across that date crosses a collection change (engines reference; fan-out).
  • Gemini (app) is a different engine from Gemini (API). In the documentation's checks on 2026-09-28, DataForSEO reported the app's model as "3.5 Flash-Lite", which is not the model the API engine calls, so the two can answer the same prompt differently (engines reference).
  • Perplexity answers collected through its earlier Search API keep what was recorded for them: a search-results channel label, a fan-out of "not applicable" and no rates on Retrieved vs cited (engines reference).
  • Location delivery was stamped from a certain point. An answer collected before the stamp existed shows "Location delivery not recorded" and empty export columns, never a guess (engines reference).
  • Export columns moved. When language variants shipped, a language column was inserted mid-row in three datasets, so a CSV read by column position rather than header name would have misaligned (exports).

Each of these becomes a log line the day you learn of it. The value is not the history lesson; it is that a future reader of your report can see the break.

How do you tell an engine change from your own setup change?

Check your own setup first, because you can prove it. Then check the collection record, which the product stores per answer. Only then consider that the engine changed.

DiscoveredBy stamps three fields on every prompt execution the moment it runs: platform, surface and collection method. They are written when the run starts and again when it succeeds, and they are never looked up live afterwards, so a later catalogue change cannot relabel an old answer (engines reference). That means an old answer still says how it was collected. The Explorer keeps two channels of one provider as separate engine values, so a channel change does not disappear into a pooled row (Explorer).

Two record fields are worth pulling into any investigation.

  • The model recorded on the answer. The Explorer has a Model dimension, shown as recorded and as "Unrecorded" when empty, and the Daily metrics export carries the model as one of the six columns that define a row (Explorer; exports). The exports page notes that a model identifier changes on every model upgrade, which produces extra rows for the same day, provider and surface.
  • A limit on that signal. For some engines the recorded model is deliberately fixed. ChatGPT (app) reads chatgpt-app whatever model DataForSEO reports, and Gemini (app) reads gemini-app, so a model change at the source does not split the engine's history (engines reference). The model column therefore cannot tell you that an app engine changed underneath you. A stable model value is not proof of a stable engine.

The Google engines record the surface instead of a model, because Google does not say which model wrote the answer: the value reads ai-overview or ai-mode (engines reference).

Which comparisons need an annotation?

Annotate any comparison whose two periods sit on opposite sides of a log entry that touches one of its metrics. The table below is the decision rule, with the documented reason for each row.

If the comparison spans... Annotate because... Where the documentation says so
An engine that started or stopped answering mid-window Built-in comparisons count an engine in a period only when it ran on at least 80% of that period's answered days. Comparisons you build yourself (an Explorer comparison, a saved dashboard) use whatever engines you select, so an engine present in only one period counts there engines reference
A change in the model value, on an engine that records it Daily metrics splits rows by model, so counts for one day can sit in more than one row exports
The date an engine's searches began to be captured Earlier answers read "not recorded", which is a gap in capture and not evidence of no searching fan-out
Answers collected before location delivery was stamped The delivery columns are empty, not guessed engines reference
A prompt added or removed, or a variant added The population of answers changed; a deleted prompt's answers are gone, while paused prompts count if they have answers in the window Explorer
A competitor added Share of voice sums across every active tracked brand family, so the denominator grows without any answer changing troubleshooting
A change in which engines a saved view selects A comparison uses the engines you select in both periods, so you changed the population yourself Explorer
A day of failed runs Failed runs never count as answers, so they leave a gap and not a miss troubleshooting

If none of these rows applies and the number still moved, write "no known change" in the log. That is a valid entry, and it is the honest one. Engines do not answer the same way every time, and the troubleshooting page states that a single day's swing is not, by itself, evidence that anything is different on your site (troubleshooting). Related: why your AI visibility score changed when your website did not.

A comparison that crosses a log entry needs a note; a movement with no entry stays open.

The measurement change log template

Copy this into a spreadsheet or a shared doc. One row per change. The columns after "Source" are the ones that make the log usable for annotation.

MEASUREMENT CHANGE LOG
Project: ____________   Owner of the log: ____________   Started: ____-__-__

ID:              L-001
Date effective:  YYYY-MM-DD   (date the change took effect, not the date you noticed)
Date logged:     YYYY-MM-DD
Kind:            Engine / Collection / Availability / Our setup / Reporting
Engine(s):       (name each; use the labels your tool shows, e.g. "ChatGPT (app)")
What changed:    One sentence, factual. No cause you cannot show.
Source:          Docs page, release note, run history, or "own edit". Include a link.
Confidence:      Documented / Observed in records / Suspected
Metrics touched: e.g. brand visibility, share of voice, citation rate, fan-out, position
Windows touched: Which saved comparisons cross this date
Action:          Annotate / Split the series / No action / Recheck on YYYY-MM-DD
Notes:           Anything the next reader needs

UNEXPLAINED MOVEMENTS (separate list)
ID:              U-001
Date range:      YYYY-MM-DD to YYYY-MM-DD
What moved:      Metric, engine, size of the move, and the counts behind it
Log entries checked and found not to apply:  L-___, L-___
Status:          Open / Explained by L-___ / Within normal variation / Closed, cause unknown
Recheck on:      YYYY-MM-DD

Three habits keep it honest.

  • Use "Date effective", not the date you noticed. A change discovered on Friday may have started on Tuesday.
  • Grade confidence. "Documented" means a source says so. "Observed in records" means your run history or export shows it. "Suspected" means neither, and a suspected entry should not appear in a client report as a fact.
  • Keep unexplained movements separate. They are questions. Mixing them into the change list turns guesses into history.

Where do you pull the evidence from?

Use the export of the affected dataset, plus the Explorer with Engine and Model as breakdowns, over the same dates. The exports carry the fields that show what was collected and how.

  • Daily metrics gives one row per day, platform, surface, collection method, provider and model, counting every attempted response in that slice, whether it completed, failed, is still running or (on the Google engines) showed no AI answer (exports). A new model value, or a run of failed responses, shows up in these rows.
  • Answers carries the response text, the prompt as sent, persona, language, country, city, how the location was delivered and how the answer was collected (exports). A request over five thousand rows is refused rather than truncated, and responses over eight thousand characters are shortened with the true length recorded, so use narrow windows.
  • Explorer lets you break a metric down by Engine and by Model, or filter to one of them (Explorer).
  • Prompts exports the project's current configuration and ignores the date range, so a saved copy dated before an edit documents your setup at that point (exports).
Demo data. The Explorer broken down by Engine and Model, one place to look for a change in what was recorded.

Read exported CSV files by header name and not by column position. The exports page explains that a column added mid-row would silently shift a script that reads by position (exports). For a JSON export, the columns array names each field.

Worked example: Quillstone sees a drop

Illustrative example: Quillstone and its competitors are fictional, and the numbers are made up to show the method.

Quillstone sells document-review software to legal and compliance teams. It tracks 20 prompts in one country on ChatGPT (app), each collected daily, so each week holds 140 analysed answers. Brand visibility (analysed answers naming Quillstone, divided by analysed answers) was 63 of 140 (45%) in the week of 5 October and 35 of 140 (25%) in the week of 12 October. A leadership reader asks whether ChatGPT changed its model.

The analyst does not answer the question. She works the log.

  1. Check own setup entries. L-004 records that on 8 October the team added a competitor, Brieflane. That touches share of voice, but the figure here is brand visibility, which divides answers naming Quillstone by analysed answers, so adding a competitor does not change it.
  2. Check prompts. The log shows no prompt added or removed in either week. Both weeks have 140 analysed answers, so the population matches, and the run history shows no failed runs in either week.
  3. Check collection entries. No engine, collection or availability entry falls in the window.
  4. Check the model field. For ChatGPT (app) the recorded model is fixed at chatgpt-app, so an unchanged value proves nothing about the engine.
  5. Check the answers. She exports Answers for both weeks and reads five of each. The wording of several answers shifted, and two of the five in the later week name Brieflane where earlier ones named Quillstone. That is a description of what the answers say, not of why.
  6. Write the entry.
ID:              U-001
Date range:      2026-10-05 to 2026-10-18
What moved:      Brand visibility on ChatGPT (app): 63/140 answers (45%) week of 10-05, 35/140 (25%) week of 10-12
Log entries checked and found not to apply:  L-001 to L-004 (L-004 affects share of voice only)
Status:          Open. Cause unknown.
Recheck on:      2026-11-02
Notes:           No engine or collection entry in range. Model value is fixed for this engine,
                 so it cannot confirm or rule out a change at the source.
                 Two weeks is a short comparison; recheck before acting.

The report line reads: "Quillstone's visibility on ChatGPT (app) fell from 45% to 25% between two consecutive weeks of 140 analysed answers each. We have found no change in our setup or collection that accounts for it and will recheck on 2 November." That is accurate, careful and quotable, and it does not name a cause nobody can show.

What can a change log not tell you?

It cannot tell you why an engine answered as it did. A log records changes you know about. Absence of an entry is not evidence that nothing changed, since an engine or provider can change without telling you.

Not every movement has an entry; a blank line is a valid answer.
  • It cannot prove a model release. A visibility swing that lines up with a rumoured update is a coincidence you can note, not a finding. Write it as "suspected" at most.
  • It cannot see inside licensed collection. The engines reference says how DataForSEO fetches a Google result is not visible to DiscoveredBy (engines reference), so your log can hold what the documentation says and what the records show, not the provider's internals.
  • It cannot repair a broken series. Annotation warns the reader; it does not make two periods comparable. Where a collection change is clean (a date after which a field is captured), split the series at that date.
  • It is not a substitute for sampling sense. Engines do not answer the same way every time, so a figure can move without any cause worth logging. DiscoveredBy marks a metric provisional when it rests on fewer than 30 observations (metrics); treat a short window with the same caution even above that floor.

Common mistakes

  • Logging the date you noticed a change, not the date it took effect.
  • Writing a cause into the "What changed" line before you have a source.
  • Treating an unchanged model value as proof that an app engine is unchanged.
  • Comparing an engine present in one period with the same engine absent in the other and calling the difference a trend.
  • Reading exported CSV files by column position.
  • Letting the log go stale. A log updated only when someone is worried is missing the calm-period entries that later matter.
  • Overstating in the other direction: assuming every movement is an engine change and never checking your own edits.

Frequently asked questions

How often should I review the change log?

Check it every time you build a before-and-after comparison, and review it as a whole at your regular reporting cadence, such as monthly. The first check takes seconds: look for any entry dated between the two periods you are comparing.

Do I need a log if I only track a few prompts?

Yes, in a lighter form. With few prompts, one added or removed prompt is a large share of your population, so your own setup changes matter more, not less. Keep the same columns and log only what applies.

Can I tell from my results that ChatGPT or Gemini released a new model?

Not from a visibility swing alone. A swing is consistent with many causes, including ordinary run-to-run variation, and the recorded model value is fixed for the app engines. Where an engine does record a model, a changed value is an observation you can log, but it does not explain a movement. Log the swing as unexplained and, if a provider publishes a dated release note, log that separately with its source.

Should I show the change log to clients or leadership?

Show the annotations that affect the numbers they are reading, not the whole log. Each annotated comparison should say what changed, the date and that the periods are not directly comparable. Keep suspected entries out of external reports.

What should I do when an engine stops running in my project?

Log the date it stopped. Its past answers keep its name, and per-engine figures such as the dashboard's include it for any period in which it answered, while the comparisons the product makes for you leave it out unless it ran on enough of the period (engines reference). Annotate any total you built yourself that crosses the date.

Next step

The log is only as good as the records behind it. In DiscoveredBy, every answer carries its collection details, the Explorer breaks metrics down by engine and model, and the exports let you keep a dated copy of a period. Exports are available depending on your plan (plans and limits). If you want that record behind your own change log, read the engines reference, then sign in or start tracking. For related interpretation work, see why your manual ChatGPT check differs from a monitoring result, how to investigate a disagreement between Gemini app and API answers and how to plan separate measurements for Google AI Overviews, AI Mode and Gemini.

  • reporting
  • measurement
  • data quality
  • ai engines
  • change log

Share

Summarize with AI

Start monitoring your AI visibility.

See how AI search engines talk about your brand.

Free to start. No credit card required.