Explain AI visibility to leadership without promising a revenue number: a decision memo
Separate what AI engines say about you, what visits you can measure, and what you cannot, then use a copyable decision memo to ask for a funded first program.
On this page
- In short
- What can you honestly tell leadership about AI visibility?
- What exactly is tier one, and how far can you trust it?
- What is tier two, and what does it leave out?
- Why not just forecast revenue?
- The decision memo: how to ask for funding
- Worked example: Quillstone asks for a first program
- How do you keep reporting after the memo is approved?
- Common mistakes
- Frequently asked questions
- Next step
To explain AI visibility to leadership without promising revenue, split what you know into three tiers and never blend them. Tier one is observed answer presence: how often AI engines name your brand or cite your domain for the questions you track. Tier two is measured referral outcomes: visits and conversions from AI engines that your analytics can attribute. Tier three is everything else, which you cannot measure and should say so. Then ask for a bounded, time-boxed first program with decision criteria written in advance, not a return on investment forecast.
In short
- Executives do not need a revenue prediction; they need to know what is observed, what is measured, and what is unknown.
- Tier one (presence) and tier two (referral sessions and conversions) come from different systems and answer different questions. Report them side by side, never summed.
- A funding memo asks for a small program with a fixed window, a stated cost, and stop or continue rules agreed before any data arrives.
- Say what you cannot see: visits from Google's AI features look like ordinary Google search visits, and people who read an answer and never click leave no referral.
- Small samples are provisional. A number with no data is not a zero.
What can you honestly tell leadership about AI visibility?
You can tell them three different kinds of things, and each has a different level of certainty. Mixing them is what turns a sensible update into an overpromise.
AI visibility here means how often AI answer engines name your brand or cite your domain when asked the questions your buyers ask. It is a measurement of answers, not of buyers. Definitions of the terms used below (mention, citation, retrieval) are on the key terms page.
| Tier | What it is | Where it comes from | Certainty | What you may say |
|---|---|---|---|---|
| 1. Observed answer presence | Whether engines name you or cite your domain for tracked prompts | Answers collected from AI engines for your prompt list | Observed, for the prompts and dates you chose | "For these 40 questions, engines named us in X of Y analysed answers." |
| 2. Measured referral outcomes | Sessions and conversions arriving from an AI engine's own referrer | Your Google Analytics 4 property | Measured, but only for referrals that carry an identifiable source | "GA4 recorded X sessions from known AI referrers, with Y conversions." |
| 3. Unmeasured business effects | Everything that happens without an attributable click | Nowhere | Unknown | "We cannot see this, and here is why." |
The reason to keep tiers apart is that a citation measures the answer, not what the reader did next. The AI traffic docs draw the same line: citations record that an engine named your domain, and AI Traffic is the separate view of whether a person then arrived and converted.
What exactly is tier one, and how far can you trust it?
Tier one is a set of rates over answers your monitoring tool collected, so it is only as broad as your prompt list and only as recent as your window. It tells you how you appear, not why, and not what it is worth.
In DiscoveredBy the metrics reference defines the numbers leadership will see. Four are worth putting in front of executives:
- Brand visibility: analysed answers naming your brand, divided by analysed answers.
- Citation rate: collected answers that cited your domain, divided by collected answers.
- Share of voice: distinct answers naming your brand family, divided by the sum of each tracked brand family's distinct-answer count.
- Brand position: the mean earliest mention order across the answers that named you, where lower is better.
Three rules from the docs keep these honest in a leadership setting:
- Denominators differ. A collected answer is a completed prompt execution; an analysed answer also has its mentions extracted. Brand visibility divides by analysed answers and citation rate by collected ones, so 50 percent visibility next to 10 percent citation rate can both be correct on one screen.
- Share of voice depends on your roster. It is relative to the competitors you chose to track and have not paused, so adding a competitor can lower it with no change in the answers. Say which competitors are in the roster whenever you quote it.
- Small samples are provisional. Below 30 observations a metric still displays, but the sample is too small to read as settled. A metric with no data reads as no data, never as zero.
Also state what tier one cannot say. Mentions, citations and stated reasons are observations of an answer. They do not reveal why a model produced it, and a rise is not proof that anything you did caused it. If leadership will ask whether a specific change worked, that is a separate before-and-after question; see Did your content update help?.
What is tier two, and what does it leave out?
Tier two is sessions and conversions that Google Analytics 4 attributes to an AI engine's referrer. It is real, countable evidence, and it is a floor rather than a total.
AI Traffic reads a connected GA4 property and keeps only sessions whose source matches one of 21 known AI referrer hosts, resolving to ten engines. A session whose source is not on that list, including one with no referrer at all, is not counted, whatever channel it actually came from. Setup needs GA4 connected and a property mapped to the project; see integrations. The screen shows total sessions for a 7, 28 or 90 day window, a breakdown by engine, conversions, and top landing pages. Sessions, engine rows and landing pages are compared with the equal-length window before; the conversion comparison exists for the window total only, not per engine.
Two limits matter for leadership:
- Google's AI features are not separable. Google AI Overviews and Google AI Mode have no referrer of their own. A visit from either reaches GA4 as ordinary Google search traffic, which cannot be told apart from any other Google click, so neither appears in AI Traffic. Your answer-presence data covers those engines; your referral data does not.
- Conversions are a mixture. The docs describe the conversion figure as GA4's own count and revenue on rows where none of your named conversion events fired, and those events' own count and value where they did. Name your conversion events on purpose before you report a number, and say which ones they are.
The glossary entry on AI referral traffic gives the short definition if your audience needs one. For a step-by-step measurement setup, see How to measure AI referral traffic and conversions with GA4.
Why not just forecast revenue?
Because no step connects tier one to tier two to revenue with a rate you have measured. A forecast would need a conversion from "named in an answer" to "visited", and from "visited" to "bought", and neither has a stable, known value.
Consider what breaks the chain. An answer that names you may carry no link. A reader may see the answer and search your brand later, arrive as direct or organic traffic, and be counted in another channel. Answers vary by engine, prompt wording and date. And ranking first in Google is a different achievement from being cited, so your existing search numbers do not transfer.
Any percentage or dollar figure you attach to that chain would be invented. If an executive asks for one, the honest answer is: "We can report what engines say, what visits GA4 attributes, and what we do not see. We will not multiply them together."
The decision memo: how to ask for funding
Ask leadership to fund a bounded first program, judged against criteria you set before the data arrives. The memo below is the deliverable of this post. Copy it, fill the brackets, and keep it to one page.
DECISION MEMO: AI visibility, initial program
To: [decision maker] From: [you] Date: [date]
1. THE DECISION
Approve [budget or staff time] for a [12]-week initial program to
measure and improve how AI answer engines present [brand].
Not requested: a revenue target.
2. WHY NOW (one paragraph, observed facts only)
[What we saw in a baseline run: e.g. "For our [N] tracked questions,
engines named us in X of Y analysed answers; [competitor] appeared in
Z."] Window: [dates]. Engines covered: [list].
3. WHAT WE KNOW, IN THREE TIERS
Tier 1, observed presence: [brand visibility, citation rate,
share of voice, with sample sizes]
Tier 2, measured referrals: [AI-referred sessions, conversions,
conversion events counted, window]
Tier 3, not measured: [Google AI features look like ordinary
Google traffic; no-click influence;
later branded visits]
4. WHAT WE WILL DO
- Fix the prompt list at [N] questions covering [products/markets].
- Baseline already run (section 2); reviews in weeks [6] and [12].
- Actions limited to: [3 named changes, e.g. pages to update].
5. HOW WE WILL JUDGE IT (agreed before we start)
Continue if: [e.g. presence on tracked questions rises across the
window AND the change is larger than the week-to-week
swing we saw in the baseline].
Stop or rethink if: [e.g. no change beyond baseline swing at week 12,
or GA4 attribution cannot be set up].
Unknown by design: revenue effect. We will report referral sessions
and conversions as measured, not as a forecast.
6. RISKS AND LIMITS
- Samples under 30 observations are provisional.
- Share of voice depends on the competitor roster: [roster].
- Answers vary by engine and date; results are not guarantees.
7. THE ASK
[Approve / approve with changes / decline] by [date].
Three choices in that memo do most of the work. Section 5 fixes the decision rule in advance, so nobody argues about what counts as success after the fact. Section 3 forces the tiers to stay separate. And "Not requested: a revenue target" removes the question that would otherwise dominate the meeting.
Worked example: Quillstone asks for a first program
Illustrative example: Quillstone and its competitors are fictional, and the numbers are made up to show the method.
Quillstone sells document-review software to mid-sized legal and compliance teams. Its marketing lead wants a 12-week program and needs sign-off from a chief financial officer who dislikes forecasts. She runs a 28-day baseline on 40 tracked questions against fictional competitors Brieflane and Clausewise.
| Measure | Count | Result |
|---|---|---|
| Collected answers | 120 | (reference) |
| Analysed answers | 96 | 80% of collected |
| Analysed answers naming Quillstone | 24 | Brand visibility: 24 / 96 = 25% |
| Collected answers citing quillstone.example | 12 | Citation rate: 12 / 120 = 10% |
| GA4 sessions in the same window | 4,200 | (reference) |
| Sessions from known AI referrers | 84 | 84 / 4,200 = 2% |
| Conversions on those AI sessions | 3 | (as counted with named events) |
Her memo reports the table as three lines, not one story. Tier one: "Quillstone was named in 24 of 96 analysed answers (25 percent) and its domain cited in 12 of 120 collected answers (10 percent). The two rates use different denominators." Tier two: "GA4 attributed 84 of 4,200 sessions to known AI referrers, with 3 conversions on them." Tier three: "Visits from Google's AI features are not separable from other Google traffic, so this figure understates AI-influenced visits by an amount we cannot state."
The 84 and 3 are small, so she says so, and she does not project them. Her stop or continue rule: continue at week 12 if visibility on the 40 questions is higher than the baseline by more than the largest week-to-week swing in her 28-day baseline, and if GA4 attribution is working. The CFO approves the 12 weeks because the downside is a fixed cost and the decision rule is written down, not because a return was promised.
How do you keep reporting after the memo is approved?
Use the same three tiers on a schedule. Report tier one and tier two as separate lines each time, restate what tier three excludes, and compare like with like: the same prompt list, the same window length, the same engines.
For the recurring version of this conversation, see Turn a weekly analyst report into a monthly leadership update. If you first need to read the numbers themselves, AI visibility, brand mentions, and citations: what each number tells you covers the definitions.
Common mistakes
- Adding tier one and tier two. Answer presence and referral sessions are different units. Present them in separate lines.
- Quoting share of voice without the roster. Leaders will assume it measures the whole market. It measures your chosen competitors.
- Treating a citation as a visit. A citation is an observation about an answer; a visit is an event in analytics.
- Reporting a zero that is really no data. An empty window or a metric with no observations is not a decline.
- Overstating the blind spot. Google's AI features not appearing in AI Traffic does not mean AI has no effect on your business; it means this one view cannot see it. Say exactly that.
- Defining success after the results. Write the continue and stop rules into the memo first.
- Promising outcomes. Suggested content changes are hypotheses to test, not commitments.
Frequently asked questions
Can I show ROI for AI visibility?
Not honestly from these measurements alone. You can show what engines say about you and what GA4 attributes to AI referrers, but the chain from an answer to revenue has no measured conversion rate. Report the tiers, name the gap, and let leadership judge the size of the bet.
What if leadership insists on a revenue number?
Offer the measured part: attributed sessions and conversions, with the conversion events named and the window stated. Then state that this excludes Google's AI features and any visit with no identifiable referrer. If a number must go in a plan, label it a target the program will test, not a forecast.
How long should the initial program run?
Long enough to see more than one reading. The example above uses 12 weeks with reviews at weeks 6 and 12, but the right length depends on your prompt count and how many answers each metric rests on. Remember that metrics built from fewer than 30 observations are provisional.
Which engines should the memo cover?
The ones your buyers are likely to use, and the ones you can measure. Name them in the memo. DiscoveredBy measures eight engines, and its engines reference explains how each is collected, which is worth knowing before you compare them.
Do I need Google Analytics 4 to make the case?
No, but tier two disappears without it. You can still report tier one and say plainly that referral outcomes are not yet measured. Connecting GA4 and mapping a property is what turns that line from "unknown" into a number.
Next step
Run a baseline before you write the memo, so section 2 holds real observed facts. In DiscoveredBy, connect your Google Analytics 4 property (see integrations) so tier two is measured from the start, add your tracked questions, and read the first results with the first results guide. How many prompts you can track depends on your plan; see plans and limits. The analytics feature page shows where the traffic views sit.
- ai visibility
- ai referral traffic
- leadership reporting
- business case