AI monitoring completeness: what percentage of your planned prompt runs actually ran?
Compare the runs you planned with the answers you collected and analysed. A worksheet for measuring AI monitoring completeness before you read any visibility number.
On this page
- In short
- Why does collection completeness matter before you read a metric?
- What exactly is a "planned run"?
- What are the four outcomes to count?
- Where do you get the numbers?
- The completeness worksheet
- Worked example: a week of Quillstone monitoring
- Common mistakes and what this cannot tell you
- Frequently asked questions
- Next step
Your AI monitoring completeness is the share of the runs you planned that produced a usable result. To measure it, count the runs your setup should have made (active prompt targets times the engines that run them), then count how many came back as a completed answer, how many came back as "no AI answer shown", how many failed, and how many never appeared at all. Divide the resolved runs by the planned runs. Do this before you read any visibility trend, because a metric computed from a partial collection can move for reasons the engines never had.
In short
- Completeness is a ratio of runs, not of prompts: planned runs are prompt targets multiplied by the engines that run them.
- Separate four outcomes: completed, no AI answer shown, failed, and missing (no record at all). They have different causes and different fixes.
- A second ratio matters too: how many collected answers have been analysed for brand mentions. Mention-based metrics only count analysed answers.
- Exclude variants that never run on an engine by design; otherwise you will report a gap that is really a rule.
- This is a data-quality check on collection. It is not the same as asking whether your prompts overlap, which is a separate audit.
Why does collection completeness matter before you read a metric?
Most AI visibility metrics divide by a population of answers, so if the population is partial, the metric describes the partial population. A drop in brand visibility that coincides with an engine going quiet for a day may be a change in what was collected, not a change in what the engines said.
There are two populations to keep straight. A collected answer is a completed prompt execution: the engine ran and returned a result. An analysed answer is a collected one whose brand mentions have also been extracted. Brand visibility divides by analysed answers; citation rate divides by collected answers (Metrics defined sets out the denominators). So two gaps can distort a chart: runs that never became collected answers, and collected answers that were never analysed.
A reassuring fact first: a failed run does not drag your numbers down. Every answer metric is built from completed executions only, so a failure is absent from the count, not counted as a miss (A run failed explains this). That is exactly why completeness needs its own check. Because failures vanish quietly, a visibility chart does not show you that a third of a day's runs are missing.
What exactly is a "planned run"?
A planned run is one active prompt target on one engine that the target runs on, for one day. It is a count you build from your setup, not a number the product hands you, so define it carefully.
A prompt target is one location, one audience (General or a persona) and one language for a prompt. Each target is run independently. A prompt tracked in 3 countries as General plus 2 personas, all "As written", produces 3 × 3 × 1 = 9 targets (Prompts covers this). Once a day, a scheduled job enqueues one run per active prompt target on every engine available on your plan that the target runs on. It requires the project, the prompt and the target all to be active.
So the planned runs for a day are:
planned runs per day = sum over active prompt targets of
(engines that run that target)
Two things shrink the engine count for a target, and neither is a failure.
Rules that keep some variants off some engines. According to the prompts documentation, these variants never run on these engines:
| Engine | Variants that never run on it |
|---|---|
| Google AI Overviews, Google AI Mode | Persona variants, Chinese variants, prompts over 700 characters (and languages the engine does not support) |
| ChatGPT (app) | Persona variants, city variants, Chinese variants, prompts over 2,000 characters |
| Gemini (app) | Persona variants, Chinese variants, prompts over 2,000 characters |
The prompt's own page lists each skipped variant with its reason under the run history, and the Add and Edit prompt previews count the same skips. Take these out of your planned count, or you will report a permanent gap that is really a rule.
Engine choices and availability. A prompt can be limited to a subset of engines. And an engine can be on your plan yet not running: the engine must be active and its key configured on the platform. If an engine stops running, its past answers keep its name and no new ones arrive. Which engines run for your project is described in Engines and measurement.
Exclude one more category: runs you started yourself. Accepting a suggestion queues one fast run outside the daily cycle, and a run started on demand is separate too. They are not part of the daily plan, so leave them out of the denominator.
What are the four outcomes to count?
Every planned run ends in one of four states. Counting them separately tells you what to fix.
- Completed. The engine ran and returned an answer. This is a collected answer.
- No AI answer shown. On Google AI Overviews and Google AI Mode only, the request succeeded but Google showed no AI answer. The run is stored as its own outcome with no text and no citations, and it is not retried that day. It is a real result about that search, and it feeds the shown rate. It is not an answer, so it is left out of every answer-based metric. For completeness purposes it counts as resolved: the run happened and returned a definite outcome.
- Failed. The request errored. A plain-language reason is shown from a closed set (timeout, rate limited, provider error or internal error). Failed runs are retried automatically at 03:00, 05:00, 09:00 and 17:00 UTC; a retry that succeeds replaces the failed run. A run still failed after 17:00 stays failed for that day. If a request fails at the data provider or Google, that is a failed run, never a "no AI answer".
- Missing. There is no run record for a target and engine you expected. The documentation does not enumerate every cause, so treat this as a category to investigate rather than diagnose from a table. Candidates the docs point to: the prompt or target was paused partway through the period, an engine that is on your plan was not running for your project (its key was not configured, for example), or you have miscounted the plan (for example, a skipped variant you forgot to exclude).
Then define the two headline ratios:
collection completeness = (completed + no AI answer) / planned
failure rate = failed / planned
missing rate = missing / planned
analysis completeness = analysed / completed
end-to-end yield = analysed / planned
Once no run is still in flight, completeness, failure rate and missing rate sum to 100 percent of planned runs. The last two describe the second gap: how much of what was collected has been read for mentions.
Where do you get the numbers?
You can get the counts from the product and its exports, but no single screen shows planned-versus-collected for you, so you assemble them. Whether exports are available depends on your plan (see Plans and limits).
- Planned: the Prompts screen and the Prompts export give you the current configuration. The Add and Edit previews count targets and the variants that will not run on each engine. Note that the Prompts export is not windowed: it is the project's configuration as it is now, so it cannot tell you what was active on a past day.
- Attempted, failed and no-answer: the Daily metrics export. It has one row per day, platform, surface, collection method, provider and model, and counts every attempted response in its slice whether it completed, failed, is still running, or showed no AI answer. The last column,
responses_no_answer, counts the Google runs with no AI answer. Read the header row of your download for the other count columns rather than relying on column position, since column order can change. - Completed: the Answers export has one row per completed answer, and its
brand_mentionedcolumn is empty rather than false when brand extraction has not run on that answer, which is how you can see the analysis gap. Answers takes a five-thousand-row ceiling per download, so use a short date window and download in parts if you need it (the docs note that a request over the cap is refused, not truncated). - Per-run detail: each prompt's run history lists the runs in the window, failed ones included, up to the 50 most recent.
- Collection window: to see whether an engine ran at all on most days, note that the product's own comparisons count an engine in a period only when it ran on at least 80% of that period's answered days. That is a useful threshold to borrow, but it is the product's rule for its comparisons, not a completeness standard.
The completeness worksheet
Fill one worksheet per reporting period, using days that have finished. Retries run until 17:00 UTC and analysis lags collection, so today's numbers will look worse than they end up.
COLLECTION COMPLETENESS WORKSHEET
Project: ______________ Period: ____ to ____ (finished days only)
Prepared by: __________ Date prepared: ________
1. PLAN
Active prompt targets on each day (note any mid-period pauses or additions):
Engines running the project (name each, note any engine unavailable for a day):
Variants excluded because the engine never runs them (list with reason):
Runs you started yourself, excluded:
Planned runs per engine per day = targets that run there: ____
PLANNED TOTAL (all engines, all days): ____
2. OUTCOMES (from Daily metrics, Answers export and run history)
Attempted (a record exists): ____
Completed: ____
No AI answer shown (Google only): ____
Failed (after retries): ____
Missing = planned minus attempted: ____
Check: completed + no answer + failed = attempted? yes / no
(If no, the difference is runs still in flight; wait, or drop that day.)
3. RATIOS
Collection completeness = (completed + no answer) / planned = ____ %
Failure rate = failed / planned = ____ %
Missing rate = missing / planned = ____ %
Analysis completeness = analysed / completed = ____ %
End-to-end yield = analysed / planned = ____ %
4. BY ENGINE (repeat the ratios per engine)
Engine | Planned | Completed | No answer | Failed | Missing | Completeness
5. BY DAY (look for a single bad day or a slow slide)
Day | Planned | Resolved | Failed | Missing
6. DECISION
Is this period fit to interpret? yes / partly / no
Engines to exclude or caveat:
Days to caveat on any chart:
Follow-up owner and date:
A practical decision rule is yours to set, since the docs give no threshold for what is "complete enough": agree one in advance (for example, "we do not write commentary on an engine whose completeness is below our floor") and write it in the report, so the standard is not chosen after the numbers are seen.
Worked example: a week of Quillstone monitoring
Illustrative example: Quillstone and its competitors are fictional, and the numbers are made up to show the method.
Quillstone sells document-review software to legal and compliance teams. For one finished week (7 days) it tracks 12 active prompt targets, all General audience, As written, country-wide. Five engines run them: ChatGPT (app), Gemini (app), Google AI Overviews, Perplexity and Grok. No target needs excluding, so each engine plans 12 runs a day.
Planned per engine: 12 × 7 = 84. Planned total: 84 × 5 = 420.
| Engine | Planned | Completed | No AI answer | Failed | Missing | Completeness |
|---|---|---|---|---|---|---|
| ChatGPT (app) | 84 | 80 | 0 | 4 | 0 | 95.2% |
| Gemini (app) | 84 | 76 | 0 | 8 | 0 | 90.5% |
| Google AI Overviews | 84 | 58 | 21 | 5 | 0 | 94.0% |
| Perplexity | 84 | 81 | 0 | 3 | 0 | 96.4% |
| Grok | 84 | 65 | 0 | 7 | 12 | 77.4% |
| All engines | 420 | 360 | 21 | 27 | 12 | 90.7% |
Check the arithmetic: completed 360 plus no answer 21 plus failed 27 equals 408 attempted, and 420 planned minus 408 attempted leaves 12 missing. Resolved runs are 360 + 21 = 381, and 381 / 420 = 90.7%. Failure rate is 27 / 420 = 6.4%, and the missing rate is 12 / 420 = 2.9%. The three parts, 90.7 + 6.4 + 2.9, sum to 100.0.
Of the 360 completed answers, 342 have been analysed for mentions. Analysis completeness is 342 / 360 = 95.0%, and end-to-end yield is 342 / 420 = 81.4%.
What the team reads from this:
- The blended 90.7% hides Grok. Its 12 missing runs equal exactly one day of Quillstone's 12 targets, which is consistent with one day when the engine did not run at all, not scattered failures. The team checks whether the engine was running for the project that day and notes that Grok ran on 6 of 7 days, above the 80% presence threshold the product uses for its own comparisons, so it stays in period comparisons but gets a caveat on that day.
- Google AI Overviews looks partly empty (21 runs with no AI answer) but is fully resolved: 94.0% counts those as real outcomes. The team reports the shown rate separately rather than folding those 21 runs into "failures". They do not appear as answers in visibility or citation metrics.
- Gemini (app) has the highest failure count. The team looks at the failure reasons in run history before deciding whether this is a pattern.
- Analysis completeness of 95.0% means brand visibility this week is computed from 342 answers, not 360. The team notes the sample sizes beside the metric.
Common mistakes and what this cannot tell you
- Using today's configuration for a past window. The Prompts export is a snapshot of now. If you paused or added prompts mid-period, your planned count is wrong unless you rebuild it day by day.
- Counting skipped variants as gaps. A persona variant on Google, or a city variant on ChatGPT (app), never runs. That is a rule, and the run history lists it with its reason.
- Reading an unfinished day. In-flight runs count among attempts in Daily metrics. Retries run through 17:00 UTC. Measure finished days.
- Treating "no AI answer" as a failure. It is a legitimate result for that search. Count it as resolved and report the shown rate separately.
- Confusing completeness with quality. A complete collection can still be a poor sample if your prompts are near-duplicates or miss what buyers ask. That is a different audit; see Audit your prompt list before adding more prompts. Completeness asks only whether the runs you meant to make were made.
- Assuming a full collection makes a trend safe. It removes one explanation for a movement. It does not tell you why an engine's answers changed.
- Diagnosing missing runs from a table. The docs do not list every cause of a missing record, so treat each missing run as something to investigate in your monitoring tool, not something to explain by guesswork.
Frequently asked questions
What is a good AI monitoring completeness percentage?
There is no documented standard. Set your own floor in advance and record it in the report so the bar is not moved after you see the numbers. Just as important, look at the spread: a high blended figure can hide one engine with a large hole.
Should failed runs be counted against my visibility?
No. Failed runs are absent from every answer metric rather than counted as a miss, so they do not lower your visibility. They do lower your sample, which is why you track them here and note sample sizes beside any metric.
Is a run with no AI answer part of a complete collection?
Yes, for the two Google engines. The request succeeded and Google showed nothing to collect, so the run is resolved. Read it through the shown rate, not through answer metrics. See Zero mentions, no answer, or failed collection? for reading the difference on a single result.
Why do brand visibility and citation rate divide by different counts?
Citation data is recorded during collection, so citation rate divides by collected answers. Brand mentions are extracted afterwards, so brand visibility divides by analysed answers, which is always a subset of collected ones. If analysis lags, the two can describe different populations on the same screen.
How do I keep missing days from misleading a trend chart?
Mark the affected days on the chart and compare like for like. For the effect of gaps on a trend, see What missing collection days do to an AI visibility trend, and for stable month-to-month comparisons, Use a fixed prompt cohort for month-to-month comparisons.
Does DiscoveredBy calculate completeness for me?
Not as a single figure, as far as the documentation describes. You assemble it from the plan, the run history and the Daily metrics and Answers exports, as the worksheet above does.
Next step
Run the worksheet on your last finished week before your next report, and attach the ratios to the page you send. The pieces live in the product: the Prompts screen for your targets and skipped variants, the Exports for the counts, and Engines and measurement for what each engine can return. Start a project and pull your first Daily metrics export at app.discoveredby.ai.
- reporting
- exports
- data quality
- collection completeness
- monitoring