Set quarterly AI search goals your team can actually evaluate
Split quarterly AI search goals into deliverables you control and outcomes you only observe, then pin each to a population, baseline and decision rule. A copyable goal sheet included.
On this page
- In short
- Why can't a team just set a visibility target?
- What is the difference between a deliverable and an outcome goal?
- Which metrics make sensible outcome goals?
- How do you pin an outcome goal to a population?
- How much baseline history do you need?
- What should a decision rule say?
- The deliverable: a quarterly goal sheet
- How do you turn goals into weekly work?
- Worked example: Quillstone plans its quarter
- What can a goal sheet not tell you?
- Frequently asked questions
- Next step: freeze your population and take the baseline
A quarterly AI search goal is evaluable when it names two things separately: the work your team will deliver, which you fully control, and the change you hope to observe in AI answers, which you do not. The deliverable is judged as done or not done. The outcome is judged against a baseline measured on a stated population of answers, over a stated window, with a decision rule agreed before the quarter starts. Skip either half and the review turns into an argument about what the numbers meant.
In short
- Write two kinds of goal: deliverables (shipped or not) and observed outcomes (moved, or did not move, by an agreed amount).
- Never promise an outcome as if it were a deliverable. Engines decide what they say; your team decides what it publishes.
- Every outcome goal needs a metric, a population, a window and a baseline, all written down before work starts.
- Agree the decision rule in advance: what result means continue, change approach, or stop.
- Plan for "no clear change" as a normal result, and write down what you will do when it happens.
Why can't a team just set a visibility target?
A bare target such as "raise AI visibility to 40%" fails because the number depends on choices that are easy to change by accident. Change the prompt list, the engines, the competitors you track or the date window, and the figure moves without a single answer changing.
The DiscoveredBy metrics reference is blunt about this. Brand visibility divides by analysed answers, citation rate divides by collected answers, and the two observations counts on different metrics are different denominators that should never be compared. Share of voice is relative to the competitors you chose to track and have not paused, so adding a competitor can lower it with nothing else changing. A target set on any of these without naming the population is a target nobody can later defend.
The second failure is treating an outcome as something the team can order. Your team can rewrite a page. It cannot make an engine mention you, and a citation is an observation of one answer, not a lever. Goals that blur this line set teams up to be judged on things they never controlled.
What is the difference between a deliverable and an outcome goal?
A deliverable goal is a piece of work your team completes and can prove it completed. An outcome goal is a change in an observed measurement that you hope the work contributes to. The first is binary. The second is a comparison with a baseline, and it never comes with a guarantee.
Some examples of each:
- Deliverable: "Publish revised pricing and integrations pages, reviewed by product, by week 6." Done or not done.
- Deliverable: "Audit the prompt list and freeze the final set by week 2." Done or not done.
- Outcome: "Brand visibility on the fixed prompt set is at least 5 points above the baseline at quarter end." Compared, not commanded.
- Outcome: "Citation rate for our domain on the same set is at least 3 points above baseline."
There is a third useful category, the coverage goal, which is about the quality of your measurement rather than the answers: for example, that the monitored prompts actually ran on at least the share of days you expect. It is in your control in the sense that you can fix a broken setup, and it makes every outcome goal more trustworthy. Without it you can miss a target because collection was patchy, not because anything changed.
Which metrics make sensible outcome goals?
Pick outcome metrics whose definition you can state in one sentence and whose denominator you can name. The first five rows below use the definitions in the DiscoveredBy metrics reference, and they are a fair menu for any tool that exposes similar numbers. The last row comes from the AI Traffic docs.
| Metric | What it counts | Denominator to write down | Watch out for |
|---|---|---|---|
| Brand visibility | Analysed answers naming your brand | Analysed answers | Analysis lags collection, so recent days are incomplete |
| Brand position | Mean of the earliest mention order, over answers that named you | Answers that named you | Lower is better; it is computed only over answers that named you |
| Share of voice (mention-based) | Distinct answers naming your brand family | Sum, over every active tracked brand family, of each family's own count | Changes when you add or pause a competitor |
| Citation rate | Collected answers that cited your domain | Collected answers | Different denominator from brand visibility |
| Domain coverage | Collected answers that retrieved your domain | Collected answers | Retrieval is not citation; the Google engines, ChatGPT (app) and Gemini (app) have no value for it |
| AI referral sessions | GA4 sessions from known AI referrers | None: it is a count, not a rate | Google AI Overviews and AI Mode have no referrer of their own |
The AI Traffic docs say that visits from Google AI Overviews and AI Mode reach GA4 as ordinary Google traffic and cannot be told apart from other Google clicks. A referral goal therefore covers only the engines that send an identifiable referrer. Say so in the goal. Also note that "share of voice" is the mention-based figure here; the Competitors screen shows a second, citation-based figure with the same name, so write down which one you mean.
Pick two or three outcome metrics, not six. Each additional metric is another chance to find a movement by luck, and another number to defend in the review.
How do you pin an outcome goal to a population?
Write the population down as five fixed choices, and do not change them mid-quarter. If any must change, record the date and restart the baseline.
- Prompt set. The exact list of prompts, ideally a fixed cohort. The sibling post on using a fixed prompt cohort for month-to-month comparisons goes further on this.
- Engines. Which engines and whether each is an app or an API collection. The engines reference lists what each one exposes. Keep app and API results apart, and note that which engines run for a project depends on your plan and on whether each engine is currently available.
- Competitor roster. The tracked competitors, because share of voice depends on them.
- Window. The date range for baseline and final reading, ideally the same length (for example, the last 28 days).
- Filters. Country, language, persona or city, if you use them. The Explorer can break the metrics down by all of these, so a goal can be scoped to any of them.
Then add a sixth line: the minimum sample. For a metric that reports an observation count, the metrics reference treats fewer than 30 observations as provisional. A goal whose final reading rests on a handful of answers is not evaluable, however tidy the percentage looks.
How much baseline history do you need?
Enough to see how much the number moves when nothing changes. One reading is a point, not a baseline. If you can, take several consecutive windows before the quarter starts and note the spread; a movement smaller than that spread is not evidence of anything. See also the sibling post on how much history a new project needs. If you have no history, say so in the sheet and set the first weeks of the quarter as the baseline period, with no outcome judgement until it ends.
What should a decision rule say?
A decision rule turns a reading into an action before you know the reading. It has three parts: a threshold, a confidence condition and a consequence. Without the third, a goal review ends with everyone nodding and nothing changing.
A workable pattern uses three bands:
- Clear improvement: the metric is above baseline by at least your threshold and the sample meets the minimum. Consequence: keep the approach and consider widening it.
- No clear change: the difference is smaller than your threshold, or smaller than the normal spread. Consequence: the hypothesis is neither supported nor refuted; decide whether to extend the test, or change the work.
- Clear decline: the metric is below baseline by at least the threshold. Consequence: investigate before changing anything, and check first whether the population, the roster or the engines changed.
The thresholds are your team's judgement, not a law. Choose them by asking what movement would make you spend the next quarter differently, and set them larger than the noise you observed.
DiscoveredBy itself uses a rule of this shape for citation gaps. Once you mark a gap opportunity done, its visibility is checked again at 7, 14 and 30 days after that. It reads as won if visibility for its prompt cluster has risen at least 10 points from a baseline measured over the 30 days before, or reads 20% or higher on its own; as no lift if neither is true by the 30-day mark; and as monitoring while neither is true yet (see citation gaps). That is one product's rule for one screen, not a standard, but it shows the structure: a baseline, thresholds and named outcomes. Even there, the Actions docs describe a result as a saved observation, not evidence that the change caused it.
The deliverable: a quarterly goal sheet
Copy this into your planning doc. Complete the Population block once, freeze it, then add one row per goal. A goal without a baseline and a decision rule is a wish.
QUARTERLY AI SEARCH GOAL SHEET
Quarter: ____ Owner: ____ Review date: ____
POPULATION (frozen; record any change with a date)
Project / domain: ____
Prompt set (name, count): ____ List saved at: ____
Engines (app vs API): ____
Tracked competitors: ____
Countries / languages / personas: ____
Baseline window (dates): ____ Final window (dates): ____
Minimum sample per reading: ____ (metrics reference: under 30 observations is provisional)
A. DELIVERABLE GOALS (done / not done)
# | Deliverable | Owner | Due | Evidence of done | Status
B. COVERAGE GOALS (is our measurement sound?)
# | Check | Target | Baseline | Final
C. OUTCOME GOALS (observed, never guaranteed)
# | Metric | Denominator | Baseline (n=) | Threshold for "clear change" | Final (n=)
| Linked deliverables | Hypothesis (one sentence)
D. DECISION RULES (agreed before work starts)
Clear improvement -> ____
No clear change -> ____
Clear decline -> ____
E. KNOWN LIMITS
What this sheet cannot tell us: ____
Things that could distort it (roster change, engine change, missing days): ____
F. REVIEW RECORD (fill at quarter end)
What was delivered: ____
What was observed: ____
Decision taken and why: ____
Carry forward / stop / change: ____
Section A is where accountability lives, because only deliverables can be judged as owed. Section C carries the hypothesis, the one-sentence link between a deliverable and an outcome ("we expect revised comparison content to raise citation rate on the pricing prompts"). Writing it down makes the review honest: you evaluate the hypothesis, not the people.
How do you turn goals into weekly work?
Break each deliverable into tasks, and let your monitoring tool supply the queue. In DiscoveredBy, the Actions To do list gathers open work from seven sources (fact checks, alerts, optimizations, citation gaps, Search Console items, earned sources and prompt suggestions) into one prioritised list, banded High, Medium and Low. Use it as a source of candidate tasks for your Section A deliverables. The docs are explicit that a result line beside an item is a saved observation, not evidence that the change caused it, and the same caution applies to your goal sheet.
For reporting through the quarter, the client reports overview shows saved weekly readings of brand visibility, citation rate, citation share and average citation position for each project, depending on your plan. They are useful as a running log, but the docs stress that these are separate project readings, not a ranking: prompts, countries and provider mixes may differ between projects. Your goal sheet should compare a project only with its own frozen baseline.
Worked example: Quillstone plans its quarter
Illustrative example: Quillstone and its competitors are fictional, and the numbers are made up to show the method.
Quillstone sells document-review software to legal and compliance teams. Its competitors are Brieflane, Clausewise and Docket North. The marketing lead sets a quarter with a frozen population: four buyer prompts, each tracked in one country, on one engine (ChatGPT (app)), the three named competitors, and a 28-day window for both baseline and final readings. Each prompt runs once a day, so a window can hold at most 4 x 28 = 112 answers. The minimum sample is 30 answers on the relevant denominator.
Baseline (28 days before the quarter):
- Brand visibility: 24 of 80 analysed answers named Quillstone, which is 30%.
- Citation rate: 9 of 100 collected answers cited quillstone.example, which is 9%.
- Collection: 100 of the 112 possible runs completed, which is 89%.
- Normal spread: two earlier 28-day windows read 29% and 31% for visibility, so the number moves by about two points when nothing is changed on purpose.
Note that these sit on different denominators (80 analysed, 100 collected), exactly as the metrics reference describes. The sheet records both.
Goals:
| Type | Goal | How judged |
|---|---|---|
| Deliverable | Revise the pricing page and the integrations page, with product review, by week 6 | Published, with reviewer sign-off |
| Deliverable | Audit the prompt list by week 2 and freeze it | List saved and dated |
| Coverage | At least 90% of the 112 expected runs complete in the final window | Checked at quarter end |
| Outcome | Brand visibility at least 5 points above 30%, comfortably beyond the two-point spread | Final window with at least 30 analysed answers |
| Outcome | Citation rate at least 3 points above 9% | Final window with at least 30 collected answers |
Decision rules: at least +5 points on visibility means continue and widen the page work; a smaller change means extend one more quarter with a sharper hypothesis; a fall of 5 points or more means investigate before any rewrite. Citation rate follows the same shape with its 3-point threshold.
Quarter-end reading (28-day final window):
- Brand visibility: 33 of 90 analysed answers, which is 36.7%, a rise of 6.7 points. That meets the threshold and the sample of 90 exceeds the minimum of 30.
- Citation rate: 12 of 110 collected answers, which is 10.9%, a rise of 1.9 points. That is below the 3-point threshold.
- Collection: 110 of 112 possible runs completed, which is 98%. The coverage goal is met.
Review record: both deliverables shipped and collection was sound, so those goals are met regardless of the outcome. Visibility cleared its bar; citation rate did not. The team records "hypothesis supported for visibility, not for citations", takes the visibility rule (continue), and plans a specific test for citations next quarter. It does not claim the page revisions caused the rise, because a single before-and-after comparison cannot rule out other explanations. For a closer look at testing that, see the sibling post did your content update help?
What can a goal sheet not tell you?
A goal sheet makes a quarter reviewable. It does not make it causal, and it does not make results predictable. Keep these limits in view:
- Correlation, not cause. A metric that rose after a page update may have risen for other reasons. Treat every result as a hypothesis supported or not, never a proven effect.
- Answers vary. Engines can answer the same prompt differently on different days, which is why a threshold should exceed normal spread.
- Engines and models change. A change in an engine's behaviour mid-quarter can move your numbers with no action from you. Log the date if you notice one.
- Engine availability changes. An engine can stop or start running for your project mid-quarter. Compare the baseline and final windows over the engines that ran in both.
- A rising share can hide a falling count. Ratios move with their denominators; always keep the counts beside the percentages.
- Referral traffic is partial. AI referral sessions capture only visits with an identifiable AI referrer, and say nothing about people who saw your brand in an answer and never clicked.
- Business impact is separate. Observed visibility does not equal revenue. Say so plainly when you report the quarter to leadership.
Common mistakes to avoid:
- Setting an outcome as a deliverable ("get cited by ChatGPT").
- Changing the prompt list mid-quarter and comparing to the old baseline.
- Adding a competitor mid-quarter and reading the share-of-voice drop as decline.
- Judging on a sample below the minimum.
- Choosing thresholds after seeing the result.
Frequently asked questions
How many goals should one quarter have?
Few. Three to five deliverables and two or three outcome metrics is usually as much as a small team can review honestly. More goals dilute attention and multiply the chances of a spurious movement.
Should I set a target number for AI visibility?
You can, if it is tied to a named population and a baseline, and treated as a hypothesis. A bare percentage without a prompt set, engines and window cannot be evaluated later. Prefer a threshold relative to your own baseline over a number borrowed from elsewhere.
What if I have no baseline yet?
Use the first weeks of the quarter as a baseline period and set only deliverable and coverage goals for that stretch. Make outcome judgements in the following quarter, against the baseline you have by then.
What if the result is "no clear change"?
That is a legitimate result. Record it, note whether the sample was large enough to detect the threshold you set, and decide from your rule whether to extend, change the work or stop. Do not lower the threshold afterwards to declare a win.
Can I use the same goal sheet for a client?
The structure works for anyone reviewing a period of work. If you report to a client, keep each client's population and baseline separate, and compare each client only with its own baseline.
Should traffic and conversions be goals?
They can be outcome goals with their own caveats. In DiscoveredBy, AI Traffic reads sessions from known AI referrers in your GA4 property, so the goal covers only engines that send a referrer. If total conversions are zero, the screen prompts you to name your conversion events instead of showing zeros, so set that up before the baseline period.
Next step: freeze your population and take the baseline
Before the quarter starts, use Prompt coverage to see which running prompts you could pause while everything they see regularly is still seen by a prompt you keep, then freeze the prompt list you keep and record the baseline readings from your metrics. Paste them into the goal sheet above. To do this in DiscoveredBy, sign in or create a project.
- ai visibility
- metrics
- goal setting
- baselines
- quarterly planning