How to design an AI visibility chart that does not exaggerate change
A chart-review checklist for AI visibility reports: denominators, missing periods, collection changes and axis choices, so a small move is not drawn as a big one.
On this page
- In short
- Why do AI visibility charts exaggerate change so easily?
- What is a denominator, and why does the chart need it?
- How should a chart show a missing period?
- When should a point be marked provisional?
- What happens to the line when you add an engine or change the prompts?
- Which axis choices exaggerate change?
- How do you review a chart before it goes out?
- How do the product's saved views and exports help?
- Worked example: one chart, two readings
- What can a well-designed chart still not tell you?
- Frequently asked questions
- Next step
To design an AI visibility chart that does not exaggerate change, draw every point with its denominator in view, show missing periods as gaps instead of zeros, mark small samples as provisional, start bar axes at zero (and label the range when you zoom a line), and never join two different collection setups into one unbroken line. A chart is a claim about change. If the number of answers behind a point moved, a day went uncollected or an engine was added, the line can rise or fall without your brand doing anything. The checklist below lets you review any chart, exported or saved, before someone reads a trend into it.
In short
- Every point on an AI visibility chart is a ratio; the count underneath it decides how much the point can be trusted.
- A missing period is not a zero. Draw a gap, and say why it exists.
- Adding an engine, changing a model or editing the prompt set changes the population, so the line needs a marker there.
- Bar axes should start at zero; a zoomed line axis needs its range labelled on the chart.
- Put the metric definition and the window beside the chart, not in a footnote nobody opens.
Why do AI visibility charts exaggerate change so easily?
Because the underlying numbers can be small ratios built from a few dozen answers, and small ratios swing. A brand visibility figure is the share of analysed answers that name your brand (defined in the metrics reference). With 60 answers in a week, one extra mention moves the figure by more than a percentage point.
Three other forces stack on top. Collection can have gaps. The population of answers changes when engines or prompts are added. And many chart tools default to whatever axis makes the line fill the frame, which turns a small move into a cliff.
None of this means the data is bad. It means the chart has to carry the context the number cannot. If you are still deciding what a number means, start with the metrics reference.
What is a denominator, and why does the chart need it?
A denominator is the count of items a rate is divided by; the chart needs it because the same rate means different things on 12 answers and on 300. Every rate-shaped metric in DiscoveredBy is a numerator over a denominator, and a group with nothing in its denominator has no value rather than a zero (Metrics reference).
The denominators are not interchangeable. Brand visibility divides by analysed answers (completed answers whose brand extraction has finished). Citation rate divides by collected answers (completed answers). Brand position divides by only the answers that named you, and share of voice divides by the answer-and-brand-family pairs summed across every active tracked brand family. The metrics reference warns that an observations count from one metric must never be compared against another metric's count.
Two practical rules follow:
- Put the observation count on the chart, as a label, a tooltip or a small second series of bars underneath.
- Never put two metrics with different denominators on one shared axis without saying so.
The same question, "citation rate," can also be computed several ways by different teams. If your readers may disagree about the calculation, see why two teams can calculate different citation rates from the same answers.
How should a chart show a missing period?
Show it as a break in the line, labelled with the reason. A day with no collection has no rate, and the Explorer draws a group with no denominator as no value, never as 0 (Explorer). Your chart should do the same.
There are several different reasons a period is empty, and they read differently to a client:
- Nothing was scheduled to run that day (for example, every prompt was paused).
- Runs were attempted but failed.
- Google showed no AI answer for a run, which is not a collected answer at all (Metrics reference).
- The brand was simply not named. This one is a real zero-mention result, not a gap.
Treating the first three as a drop in visibility is the most common way a chart lies. A line that dives to 0% and recovers the next week tells the reader the brand vanished; a line with a labelled break tells them the data did.
Smoothing needs the same care. The Explorer's 7-day average pools the seven days' numerators over the seven days' denominators rather than averaging seven daily rates, and days with no answers count as zero in an answer-count total, so a count near the start of a project's history reads low (Explorer). If your tool smooths, state the method in the caption. For more on gaps, see What missing collection days do to an AI visibility trend, and for telling the empty states apart see Zero mentions, no answer, or failed collection? Read the difference.
When should a point be marked provisional?
Mark a point provisional when the rate behind it comes from fewer than 30 observations. That is the platform-wide floor, MIN_OBSERVATIONS = 30 in the metrics reference, and a metric below it is still shown but is not settled.
The Explorer draws the distinction for you: a hollow dot on a line, a hatched bar or matrix cell, or the word "provisional" in a table or tile. An answer count is exact and is never provisional (Explorer). When you rebuild a chart from an Explorer export, you lose that styling unless you recreate it; the export carries a provisional column, so use it to format those points differently.
The threshold is a labelling rule, not a guarantee. A point at 31 observations is not safe and a point at 29 is not wrong; the marker tells the reader to weigh the two differently. Also resist slicing so finely that every segment is provisional: a breakdown by engine, persona, country and tag can turn a healthy total into a grid of hollow dots.
What happens to the line when you add an engine or change the prompts?
The line stops being comparable across the change, because the population of answers is different on each side. Pooled over every engine, a comparison includes an engine that started answering during the window on one side only, and the Explorer's guidance is to break the result down by engine, or filter to one, to compare like for like (Explorer).
Keep app and API results apart as well. Each run stamps its own platform, surface and collection method when it runs, and a catalogue change later cannot relabel it (Engines). That provenance is what lets you split a series by how it was collected. The Daily metrics dataset on the Exports screen also emits the day, platform, surface, collection method, provider and model as columns, and warns to expect more than one row per day when any of the last three changes, as a model identifier does on every model upgrade. It counts every attempted response, including failed ones and Google runs that showed no AI answer, which helps when you need the reason behind a gap.
So a chart needs an annotation whenever any of these change:
- an engine is added or removed
- a model or collection method changes
- prompts are added, paused or changed
- a competitor is added or paused (this moves share of voice with no change in any answer, as the metrics reference notes)
Where you can, draw one line per engine instead of a pooled line. For the reporting convention, see Keep app results and API results separate in your reports.
Which axis choices exaggerate change?
Three do: a truncated rate axis, a dual axis, and a different scale on each panel. Bars and lines behave differently, so treat them separately.
A bar's length encodes its value, so a bar chart on a rate should start at zero; a bar that starts at 14% makes a change from 15% to 20% look like a sixfold jump. A line only encodes position, so a zoomed axis is more defensible, but it still magnifies. If you zoom, label the axis range in the caption.
A dual axis lets you make any two lines cross wherever you like. Avoid it for two metrics; use two stacked panels that share the same dates instead. The Explorer takes this approach, drawing several metrics on a line or bar chart as one panel per metric, each captioned with its unit (Explorer).
Keep the same scale when you compare panels, or say plainly that you did not. And when a comparison is on, remember that the change is in percentage points for a rate, not a percentage: 15.0% to 20.0% is 5.0 points, which is a one-third relative increase. Report which one you mean.
How do you review a chart before it goes out?
Run the checklist below on every chart, whether it is a saved view on a dashboard or a figure you rebuilt from a file. It is short on purpose: each line maps to one way a chart overstates change.
AI VISIBILITY CHART REVIEW
Chart: ______________ Reviewer: ______ Date: ______
DEFINITION
[ ] Metric named exactly as the product names it (e.g. brand visibility)
[ ] Denominator stated in words (analysed answers / collected answers / answers naming you)
[ ] Window and end date stated (e.g. 28 days ending 2026-09-28)
[ ] Comparison period stated, and whether the change is in points or percent
DENOMINATORS AND SAMPLES
[ ] Observation count visible for each point, or in an adjacent table
[ ] Points under 30 observations marked provisional
[ ] No two metrics with different denominators on one shared axis
[ ] Breakdown cut (top 25 values, top 10 lines) disclosed if it applies
MISSING DATA
[ ] Days with no collection drawn as gaps, not zeros
[ ] Reason for each gap given: not scheduled / failed / no AI answer shown
[ ] A real zero (answers collected, brand not named) is distinguishable from a gap
[ ] Smoothing method stated, if smoothing is on
CHANGES IN COLLECTION
[ ] Engines added or removed inside the window annotated
[ ] Model, collection method or app/API channel changes annotated
[ ] Prompt additions, pauses or changes annotated
[ ] Competitor roster changes annotated if share of voice is shown
[ ] Pooled line replaced by per-engine lines where the mix changed
AXES AND FORM
[ ] Bars start at zero
[ ] Zoomed line axes labelled with their range
[ ] No dual axis; panels share dates and, where compared, scale
[ ] Change reported as points for a rate, with any percent figure labelled
SHARING
[ ] Anyone reading the chart alone can see the caption and notes
[ ] Snapshot date shown if the chart is frozen
[ ] One-sentence "what this does not show" written under the chart
How do the product's saved views and exports help?
They keep the context attached to the chart, if you use them. The Explorer lists the result's notes under the chart: the window and any comparison window, any breakdown cut to 25 values, partial weeks, the per-day reading of a smoothed count, and other notes the query returns. On a dashboard tile those notes fold under an "About these numbers" section, though the sentence about a breakdown cut to 25 values stays in view (Explorer).
Dashboard tiles have limits worth knowing before you screenshot one. A line chart draws its 10 busiest lines. A dashboard tile shows no full row table, except that the table appears right under the line whenever the line leaves something out (Dashboards and sharing, Explorer).
A share link is a frozen snapshot of the results as they were when it was made, never a live query, and it shows the project name, the view or dashboard name, when the snapshot was taken and when the link expires, plus each view's filters, chart and notes. Like a dashboard tile, it draws each chart without the full row table beneath it, unless a line leaves something out (Dashboards and sharing, Explorer). Creating share links and exporting depend on your plan; see plans and limits. That makes your caption and the "what this does not show" line more important, not less.
When you export a result, the file is long format with one line per result row and metric, and columns for period, partial, metric, value, observations, unit, provisional, prior_value, prior_observations, delta, compare_start and compare_end. Those columns give you the sample size, the provisional flag and the comparison dates for each point; you add the collection-change markers yourself. The CSV has no place for the notes, so keep them alongside it, while the JSON file carries compare_window, truncated and notes at the top level (Dashboards and sharing). Read a CSV by its header row, not by column position, because dataset columns have been inserted mid-row before (Exports).
Worked example: one chart, two readings
Illustrative example: Quillstone and its competitors are fictional, and the numbers are made up to show the method.
Quillstone, a document-review software company, tracks brand visibility weekly. Its content lead sends a chart to the board titled "Visibility up a third in six weeks." Here is the data behind it.
| Week | Analysed answers | Naming Quillstone | Brand visibility | Note |
|---|---|---|---|---|
| 1 | 60 | 9 | 15.0% | |
| 2 | 60 | 10 | 16.7% | |
| 3 | 60 | 9 | 15.0% | |
| 4 | 0 | 0 | no value | Collection outage; nothing completed |
| 5 | 90 | 18 | 20.0% | Second engine added this week (30 answers) |
| 6 | 90 | 18 | 20.0% |
The original chart made four mistakes. It plotted week 4 as 0%, so the line dived and recovered. It started its axis at 14%, so 15.0% to 20.0% was drawn at six times the height. It pooled the new engine into the line without a marker. And it called a 5.0-point move "up a third" without saying that was a relative figure (5.0 points on a base of 15.0).
The new engine explains the rise. In each of weeks 5 and 6 the original engine contributed 60 answers with 9 naming Quillstone (15.0%), and the added engine contributed 30 answers with 9 naming Quillstone (30.0%). Pooled, that is 18 of 90, or 20.0%. Like for like, the original engine sits at 15.0%, inside its earlier range of 15.0% to 16.7%. The added engine's 30 answers sit right at the 30-observation floor, so its 30.0% is a thin read on its own.
The corrected chart shows one line per engine (the added engine's line starts at week 5), a labelled break at week 4 ("collection outage, no value"), a vertical marker at week 5 ("second engine added"), a zero-based axis, and observation counts under each point. The caption reads: "Visibility on the original engine stayed between 15.0% and 16.7%. The pooled rise reflects an added engine, not a change in how the existing engine answers." It is a less exciting chart and a more defensible one.
What can a well-designed chart still not tell you?
It cannot tell you why the line moved. A chart records what was observed in answers; the mentions, citations and reasons in them are observations, not proof of what caused a model to write what it wrote. Do not add a causal headline to a well-drawn line.
It also cannot rescue a thin sample. A carefully annotated point on 12 answers is still a point on 12 answers. And it cannot make two populations comparable after the fact; if the collection setup changed, the honest answer is often two lines, not one. If you are evaluating a change you made on purpose, compare before and after with a fixed set of prompts instead of eyeballing a trend.
Common mistakes to avoid:
- Reporting a relative percentage change on a small base as if it were a large effect.
- Showing only the total when a breakdown by engine would show the mix shifted.
- Reusing a chart from last quarter without re-checking the definitions and the roster.
- Adding a smoothing window to hide a gap instead of labelling the gap.
- Putting the caveat in an appendix, where the person forwarding the chart will not copy it.
Frequently asked questions
Should an AI visibility chart ever start above zero?
A line chart can, if the axis range is labelled and the caption says the view is zoomed. A bar chart should not, because bar length encodes the value. When in doubt, draw both and let the reader see the difference.
How many data points do I need before a trend line means anything?
There is no universal minimum. Check the observation count on each point against the 30-observation provisional floor, and look at whether the collection setup stayed the same across the series.
Is a 0% point the same as a missing point?
No. A 0% point means answers were collected and none named the brand, which is a real result. A missing point means there was nothing in the denominator, and the product shows no value rather than 0. Charts should keep the two visually distinct.
Can I rebuild a dashboard chart in a spreadsheet from an export?
Yes, using the long-format Explorer export, which carries observations, provisional, partial and the comparison dates. Keep the notes with the file, and recreate the provisional styling yourself, since a CSV carries the flag but not the drawing.
What should a client-facing chart caption include?
The metric and its denominator, the window, the comparison basis, any gap or collection change inside the window, and one sentence on what the chart does not show.
Next step
Open one chart you already send to stakeholders and run the review checklist on it, starting with the denominator and the missing periods. If you have a project in DiscoveredBy, build the same query in the Explorer, turn on the comparison, break it down by engine, and check whether the trend survives. To start, sign in or create an account.
- ai visibility
- reporting
- exports
- charts
- data quality