How much history does a new AI visibility project need before you interpret it?

There is no universal number of days. A readiness checklist shows when a new project's results move from a first observation to a provisional read to a baseline you can report.

Kamal 14 min read
A row of small glass jars filling one by one with coral liquid, the first jar nearly empty and the last one full, suggesting a sample building up over time.
On this page
  1. In short
  2. What is the difference between an initial observation and a baseline?
  3. Why can't a single day of results tell you much?
  4. How long until data appears at all?
  5. What does the product itself tell you about readiness?
  6. The readiness checklist
  7. A worked example
  8. What should you do if your data is thin for a long time?
  9. Common mistakes and what this cannot tell you
  10. Frequently asked questions
  11. Next step

A new AI visibility project needs enough completed, analysed answers, spread across enough days, for each number you plan to report. There is no universal number of days. The platform gives you one fixed threshold: a metric built from fewer than 30 observations is provisional, meaning it is shown but the sample is too small to read as settled. Everything beyond that is a judgement you should make before you look at results: which numbers you will report, on which engines, over which window, and what would make you wait longer. This post gives you a readiness checklist for making that call.

In short

  • A first result is an initial observation, not a baseline. It tells you the pipeline works and what one round of answers looked like.
  • In DiscoveredBy, a number built from fewer than 30 observations is marked as a low sample. Treat that as the minimum, not as proof the number is stable.
  • Count observations for each segment you will report (each engine, country or prompt group), not only for the project total.
  • A baseline needs a full comparison window and, for change over time, a full window before it, collected with the same setup.
  • Decide your readiness criteria in advance, then apply them the same way to every project.

What is the difference between an initial observation and a baseline?

An initial observation is what the first few collection runs show. A baseline is a reference period, collected with a stable setup and a sufficient sample, that you later compare changes against. The two look identical on a chart, which is why they get confused.

A provisional number is one the tool shows but flags as resting on too little data. In DiscoveredBy, metrics that report an observation count are provisional below 30 observations (the platform-wide floor is MIN_OBSERVATIONS = 30). The Overview marks them: a headline reads "Low sample: n = 12 analysed answers.", a headline number carries "Low sample (n = 12)", and a trend point with fewer than 30 observations is drawn as a hollow dot.

In this post, three labels describe how far along a project is. They are reporting labels for your own documents, not states the product assigns:

Label What it means What you can say
Initial observation First runs; some numbers have no data yet or a low sample "Monitoring is running and here is what it returned."
Provisional read Sample floor met for the numbers in question, but no comparison period "Visibility is about X on these engines so far; we will confirm."
Baseline Stable setup, sufficient sample per segment, complete window collected "This is our reference. Future changes are measured against it."

Why can't a single day of results tell you much?

One day of results is a small number of answers, and those answers come from a handful of prompts asked once each. Two things limit what you can read from it.

First, the sample is small. If you track five prompts on two engines, a full day produces at most 10 answers. A single answer that changes moves your visibility by ten points.

Second, the observations are not independent. Seventy answers made from five prompts asked daily on two engines for a week are not seventy separate pieces of evidence about your brand. They are five questions asked repeatedly. The 30-observation floor counts observations (analysed answers, for brand visibility); it does not know how many distinct questions sit behind them. That is our own reasoning about sampling, not a rule the product applies, but it is a good reason to keep your prompt set broad enough to represent the buying questions you care about (see how to choose AI tracking prompts).

Answers can also differ between runs of the same prompt, so a first reading may be one draw from a range. You learn the range by collecting more draws.

How long until data appears at all?

Your first scan happens at the next run of the daily job after your prompts exist, which is at most a day away. Per the first scan docs, every active prompt runs against every engine on your plan once a day, and you do not trigger it. Accepting a suggested prompt queues one immediate run on a single engine per country the prompt tracks; that is fast feedback, not a full scan.

Two timing details affect what you can read early:

  • Citation and retrieval numbers arrive first. The engine's citations are recorded when the answer is collected. Brand visibility, position and share of voice need a separate mention-extraction step, so they lag. Until it catches up, expect no data rather than zero.
  • A 7, 28 or 90 day window ends yesterday. An answer collected today counts from tomorrow, so the Overview will not show today's runs in its headline numbers.

The Overview also tells you when extraction has not finished: "Brand figures use N of M answers; the rest are still being analysed" means the brand numbers rest on fewer answers than the citation rate.

What does the product itself tell you about readiness?

The docs describe several built-in signals. Use them as inputs to your checklist.

  • No data is not zero. A metric with an empty denominator reads as an en dash with "Not measured" under it, and a trend shows a gap. Do not report a gap as a 0% result.
  • Low sample below 30. Applies per number, so the project total can pass while one engine or one filter selection does not. The count is in each metric's own unit: brand position counts only the answers that named you, and share of voice counts mentions, so one metric can be provisional while another is not. Do not compare observation counts across metrics.
  • Trends need at least five periods. With fewer than five days (or weeks) that have any measured value, the chart is replaced by "Not enough history for a trend yet".
  • Changes need a comparison period. Each change is measured against the same number of days immediately before the window. A first period with answers on only a few days is measured against those days, so a new project gets changes as soon as it has a prior period. That comparison can be thin, so treat a change on week two as a hint, not a result.
  • Engine measurement rule. An engine counts as measured in a period only when it ran on at least 80% of that period's answered days. If an engine was measured in only one of the two periods, the change compares only the engines measured in both, and a note says so. The Overview trend chart and comparisons you build yourself in the Explorer do not apply this rule.

For share of voice there is an extra reason to wait. Its denominator is relative to the competitors you added and have not paused, so adding or pausing a competitor can change your share without any answer changing. Finish setting up your competitor list before you treat share of voice as a baseline. See competitors.

The readiness checklist

Copy this into your project notes. Fill it in before you look at trends, and again before you present a baseline.

BASELINE READINESS CHECKLIST

Project: ____________________   Date checked: ____________
Window I plan to call the baseline: ____ to ____  (length: ___ days)
Numbers I will report (circle): visibility / position / share of voice / citation rate / sentiment
Segments I will report (engines, countries, personas, languages): ______________________

1. SETUP FROZEN
   [ ] Prompt list final for this window (no prompts added, paused or reworded)
   [ ] Competitor list final (share of voice depends on it)
   [ ] Brand names and aliases set up before the window started
   [ ] Engines on the plan did not change during the window

2. COLLECTION COMPLETE
   [ ] Every intended engine ran on most days of the window
   [ ] Failed runs reviewed (Overview lists failed runs by engine)
   [ ] Days with no answers noted, not silently averaged over

3. ANALYSIS CAUGHT UP
   [ ] No "Brand figures use N of M answers" notice, or N and M are close
   [ ] Newest days excluded if extraction has not finished

4. SAMPLE SUFFICIENT, PER SEGMENT
   [ ] Each number I report has 30 or more observations in EACH segment I report
   [ ] No hollow "low sample" points in the lines I quote
   [ ] Distinct prompts behind the answers: ____ (enough to represent buyer questions?)

5. HISTORY SUFFICIENT
   [ ] At least 5 periods with data (the trend chart appears)
   [ ] One full comparison window collected
   [ ] For change over time: a full earlier window with the same setup
   [ ] The two halves of the window agree closely enough for our purpose (my threshold: ____)

6. WORDING
   [ ] Report says "initial observation", "provisional" or "baseline" honestly
   [ ] Gaps shown as "no data", not 0%
   [ ] Stated what would make me revise this baseline: ______________________

DECISION:  [ ] Wait   [ ] Report as provisional   [ ] Adopt as baseline
Next check date: ____________

Two entries deserve explanation. Item 4 makes the sample check per segment because the floor applies to each number as displayed. If you report ChatGPT (app) and Gemini (app) separately, each needs its own count. Item 5's last line, "the two halves agree", is a stability test of our own suggestion: split the window in two and compare. If the halves disagree by more than you would accept in a report, the window is telling you the number is still moving, so keep collecting. You choose the threshold before you look; the product does not set one.

A worked example

Quillstone sells document-review software to legal and compliance teams. It tracks five prompts on two engines, ChatGPT (app) and Gemini (app), with competitors Brieflane and Clausewise added before the first run. All 10 daily runs complete every day, and all answers are analysed.

That gives 10 analysed answers a day. Here is how brand visibility (answers naming Quillstone divided by analysed answers) reads as the sample builds up:

Days of data Analysed answers Answers naming Quillstone Visibility Status
1 10 6 60.0% Low sample (below 30)
3 30 10 33.3% Floor met, only 3 days, no trend
5 50 18 36.0% Trend chart can appear
7 70 26 37.1% Full first window
14 140 57 40.7% Two full windows

Day one read 60%. By day three it was 33.3%, on a sample of 30 that had just reached the floor. Nothing about Quillstone changed; the sample grew. A report written on day one would have promised a number that never came back.

Now split the fortnight into two 7-day windows and look at each engine, because each has 35 answers per week (5 prompts times 7 days), above the floor:

Segment Week 1 Week 2
ChatGPT (app) 15 of 35 = 42.9% 17 of 35 = 48.6%
Gemini (app) 11 of 35 = 31.4% 14 of 35 = 40.0%
Both engines 26 of 70 = 37.1% 31 of 70 = 44.3%

Visibility rose by about 7 points from week one to week two, and both engines moved the same way. That is encouraging, but Quillstone's team applies its own checklist. Each week rests on only five distinct prompts, and each individual prompt has just 14 answers per week across the two engines, below the floor, so no per-prompt conclusion is allowed. The team labels the fortnight a provisional read, extends collection to four weeks, and writes down its criterion in advance: adopt week-one-to-four as the baseline if no week differs from the four-week figure by more than 8 points (a threshold they chose, not a standard).

Notice what they did not do. They did not credit a content change for the rise, and they did not call 44.3% "the baseline". Evaluating an intentional edit is a separate exercise (see Did your content update help?).

From first runs to a baseline, each step is gated by checks, not by a fixed number of days.

What should you do if your data is thin for a long time?

If a segment never reaches 30 observations, widen it or stop reporting it. A low-volume engine, a narrow country or a single persona may need a longer window, more prompts, or to be rolled into a broader group. The Explorer lets you break metrics down or filter them by engine, country, persona, language and other dimensions. Cells built from fewer than 30 observations are marked provisional, so you can see which segments are still thin before you quote them.

Do not lower the bar to get a cleaner story. If the checklist says wait, report the result as provisional and say when you will look again.

Common mistakes and what this cannot tell you

  • Treating a number of days as the rule. Fourteen days of a broad prompt set can be stronger than thirty days of three prompts. Judge the sample, the setup and the completeness, not the calendar.
  • Reading day one as the baseline. The earliest readings are the noisiest and the least complete, especially while mention extraction lags.
  • Changing the setup mid-window. Adding prompts, competitors or engines changes the population you are measuring. The engine measurement rule exists for this reason; the same care applies to your own edits. If you must change the setup, start a new baseline period and annotate the chart.
  • Reporting a gap as zero. No data and 0% are different states.
  • Confusing volume with independence. More answers from the same few prompts narrows the noise from repeated runs, not the bias of a narrow prompt set.
  • Assuming reaching 30 means "stable". The floor marks when the product stops flagging a low sample. It does not prove the number will hold next week.

This approach also cannot tell you why a number sits where it does. A baseline records what the answers contained. It is an observation of answers, not proof of why a model produced them. Nor does it tell you how much history is "enough" for your decisions; that depends on how costly a wrong call would be.

Demo data. This persona segment has 27 analysed answers and is marked low sample.

Frequently asked questions

How many days of data do I need before AI visibility results are reliable?

There is no fixed number. Check that each number you report has at least 30 observations in each segment, that your prompts and competitors stayed the same, that collection ran on most days, and that you have a full window plus an earlier one for comparison. Then decide whether the halves of your window agree closely enough for your purpose.

Why does my visibility score swing so much in the first week?

Small samples swing widely, and early answers come from a few prompts. As the count grows, the number settles toward what your prompt set actually returns. Also check whether analysis has caught up, because brand figures can rest on fewer answers than the citation rate.

Does reaching 30 observations mean the result is statistically significant?

No. In the product, 30 is the point at which a metric stops being flagged as a low sample. It is a display floor, not a significance test, and the answers are not fully independent, since the same prompts repeat daily.

Why is there no trend chart on my new project?

The chart needs at least five days (or weeks) with a measured value. Until then it says "Not enough history for a trend yet". With a chip other than the date range set, it may instead say there are not enough answers matching the filters.

Should I compare my first month to a month with a different prompt list?

Not without noting the change. A different prompt list changes what is being measured. Keep a fixed prompt cohort for month-to-month comparisons so the comparison is like for like.

What if some engines failed on some days?

Review the failed runs on Overview and record which days are affected. On Overview, the engine measurement rule already limits the changes it shows to engines measured in both periods, but you should still say which days are missing. See what missing collection days do to a trend.

Next step

Set up your prompts, competitors and brand names first, let the first scans run, then work through the checklist above before you present anything as a baseline. DiscoveredBy marks low-sample numbers, shows "not measured" rather than zero, and limits Overview changes to engines measured in both periods, which covers part of several checklist items for you. You still choose the criteria. Start tracking your AI visibility, and read Reading your first results when the first data lands.

  • baseline
  • measurement
  • sample size
  • provisional results

Share

Summarize with AI

Start monitoring your AI visibility.

See how AI search engines talk about your brand.

Free to start. No credit card required.