Run it as a program
Turn the work into a routine: a 30-minute weekly review, reports for clients and leadership, quarterly goals and an evidence-based refresh plan.
Chapter 7 of 7: contents
Running AI search work as a program means replacing one-off audits with a fixed rhythm and written rules. Each week, hold a 30-minute review that reads the completed week, triages alerts, checks the evidence, decides and assigns owners. Each month, condense four weekly reports into one page for leadership that leads with a decision and keeps answer presence, measured referral traffic and what you cannot see on separate lines. Each quarter, review the work with the client and set goals that split deliverables you control from outcomes you only observe, each with a population, a baseline and a decision rule. Refreshes and AI-written drafts follow the same discipline: evidence first, named reviewers, no promised lift.
In short
- Run the weekly review in five timed blocks (read the week, alerts, evidence check, decisions, assignments), and close every item as do, defer, drop or watch.
- Keep one source of evidence and write three reports from it: a weekly report for the team, a monthly page for leadership and a quarterly review for the client.
- Never multiply answer presence by traffic to reach revenue. Report observed presence, measured referrals and the unmeasured as three separate lines.
- A quarterly goal is evaluable only if it names its type, measure, population, baseline and decision rule before the quarter starts.
- Schedule refreshes by evidence (wrong facts, changed sources, citation trends, business changes) and send every AI-drafted claim through a named reviewer.
What does a 30-minute weekly review cover?
A weekly review covers five fixed blocks: five minutes to read the week, five on alerts, seven on checking the evidence, eight on decisions and five on assignments. Every item leaves the meeting as a decision, a follow-up with a named owner, or an explicit "watch".
Figure 7.1 shows the 30 minutes block by block, with the output each block must produce.
0–5 min
Read the week
A one-sentence summary, with caveats.
5–10 min
Alerts
Each alert marked act, watch or dismiss.
10–17 min
Evidence check
A claim-strength label per finding.
17–25 min
Decisions
Do, defer or drop for each open item.
25–30 min
Assignments
An owner and a date in your tracker.
Without a fixed agenda the meeting drifts toward whichever chart is on screen, and because AI answers can vary from day to day, that movement turns into debate.
Prepare alone, meet together. One person spends about ten minutes opening the latest weekly report, noting alerts since the last meeting and skimming the open-work list. Everyone else arrives cold. Put last week's actions at the top of the page, one line each, marked done, not done or dropped.
Minutes 0 to 5: read the week. Say the summary in one sentence, such as "brand visibility flat, two prompts down, one competitor up", and read the caveats before the numbers. In DiscoveredBy the weekly report covers the last complete Monday to Sunday week in UTC and is built on Mondays, so meet after it exists and never review a partial week. When an engine was not measured in both weeks, changes compare only the engines measured in both, and a change with no comparable prior week is blank, not zero.
Minutes 5 to 10: alerts. Alerts come mostly from a daily comparison of adjacent seven-day windows, so they are a shortlist of changes that crossed a threshold, plus fact contradictions and any source watchlist you opted into. Mark each act (it goes to decisions), watch (see whether next week confirms it) or dismiss. Owners and editors can acknowledge, dismiss or snooze an alert with a shared team note, so record the conclusion there. Adding a country or persona audience to a tracked prompt changes the population the checks compare, so be wary of an alert that follows your own change.
Minutes 10 to 17: evidence check. Open one or two saved answers behind the change and read who was named and which sources were cited. Then give each finding a claim strength: 1, a number moved and nobody has read an answer; 2, the change is visible in the answers; 3, a plausible reason is visible in the answer or its sources; 4, a change you made was followed by a measured change, which is still not proof of cause. This is the habit chapter 4 builds around a single answer.
Minutes 17 to 25: decisions. Work from the open-work list, not memory. In DiscoveredBy that is Actions, To do, which bands open records from seven sources as High, Medium or Low; take High first and stop when time runs out. Decide do, defer or drop, and write each "do" as a hypothesis to test. Chapter 5 covers choosing between kinds of fix.
Minutes 25 to 30: assignments. Every "do" leaves with an owner, a due date and a check-back week. Keep owners and dates in your team's tracker. For a page change, hand the editor a page-fix task they can finish without another meeting.
When time runs short, cut decisions first and never the evidence check. In a thin week, "not enough to conclude" is a legitimate result. The template and a worked meeting are in run a weekly AI search review meeting in 30 minutes.
How do you report to clients?
Report to clients at two levels: a weekly report that a client contact can act on, and a quarterly review with the person who pays for the work and decides whether it continues. Both draw on the records the team already uses each week.
Figure 7.2 shows the reporting ladder: who reads each of the three reports, what each contains, and the one source of evidence (saved answers, exports and dated logs) that all three are built from.
Weekly
Report
For: The team, or a client who wants the detail
What was measured, what changed, what the answers said, what was done, what happens next.
Monthly
Leadership update
For: Executives
One page: what to decide, what changed, what is still uncertain, what is needed. The weekly detail goes in an appendix.
Quarterly
Client review
For: The client
Work completed, what was observed, what is unresolved, and named options for next quarter.
One source of evidence
Saved answers, exports and dated logs. Every report is built from these, not from memory.
The weekly client report answers five questions in order: what we measured, what changed, what the answers said, what we did, and what happens next. The population section prints the prompt count, the engines measured in both weeks and both denominators, since brand visibility divides by analysed answers and citation rate by collected answers. The changes section stops a wobble being sold as a trend. Two or three saved answers let the client check what an engine said. The work log shows effort without claiming it caused anything, and dated next actions, each with an owner and a way to judge it, turn the report into a decision. A flat week is still a full report: say why you believe it is flat, show the prompts that gained and slipped underneath, and date the pages too new to judge. How to write an AI visibility report a client can act on has the template.
If you share the product's digest as well, know its limits. A client report link is a fixed digest of one generated week that opens without an account and expires after 1, 7 or 30 days. Anyone it is forwarded to can open it, so treat it as a credential. It leaves out raw answers, prompt lists and historical deltas, so the changes table and the narrative are still yours.
The quarterly client review separates work completed, what you observed, what is unresolved and options for next quarter. Build it a few days ahead from saved records: the weekly reports for the first and last completed weeks, an export for the window, the status of every fix, a snapshot of open items and your own log of scope changes. In DiscoveredBy, Client reports lets you choose one completed week and open each project's exact report; a missing week shows as missing and is never replaced by another, so check both endpoints early.
List every way the two weeks differ: prompts added or paused, engines that started or stopped answering, slices under 30 observations. Show dismissed and stale drafts beside applied ones, because a deck of completed work only invites the question of what you left out. Report movement as an observation ("visibility on the four affected prompts rose six points in the month after the edit"), not as an effect of the edit. End with two or three options, such as deepen, widen, or hold and monitor, each with what it would and would not change. How to review a quarter of AI search work with a client has a 12-slide outline.
For agencies, two more rules. Scope the engagement around what you control (the prompt set, the cadence, the content deliverables, the reporting pack) and state that a week with missing collection is reported as a gap; scope an AI visibility engagement around deliverables you can control has a checklist. And never rank clients against each other: the Client reports cards are separate readings, not an agency average, because prompts, countries and engine mixes may differ between clients. A combined agency report keeps each client's metrics separate, but anyone with the link can read every included client.
How do you report to leadership without promising a revenue number?
Split what you know into three tiers and never add them together: observed answer presence, measured referral outcomes, and business effects you cannot measure. Then ask for a bounded program judged by rules written before the data arrives, and keep leadership informed with a one-page monthly update that leads with a decision.
Tier one is observed answer presence: how often engines name your brand or cite your domain for the prompts you track. Quote it with its denominator, name the competitor roster whenever you quote share of voice (adding a competitor can lower it with no change in the answers), and treat anything under 30 observations as provisional.
Tier two is measured referral outcomes: sessions and conversions your analytics attributes to an AI engine's referrer. In DiscoveredBy, AI Traffic reads your Google Analytics 4 property and keeps only sessions from known AI referrer hosts. It is a floor, not a total: Google AI Overviews and Google AI Mode have no referrer of their own, so their visits arrive as ordinary Google search traffic.
Tier three is everything else: people who read an answer and never click, later branded visits, influence through other channels. Say you cannot see it, without overstating the blind spot: this view cannot see an effect, which is not the same as there being none.
No measured rate connects the tiers. An answer that names you may carry no link, and a reader may search your brand later and arrive as direct traffic. Any figure that multiplies presence into visits and visits into revenue would be invented, so offer the measured part instead, with the conversion events named.
To start a program, write a one-page decision memo: the decision, why now (observed baseline facts only), what you know in three tiers, what you will do, how you will judge it, the risks and the ask. Its key line is "Not requested: a revenue target", and its key section fixes the continue and stop rules before work starts. Explain AI visibility to leadership without promising a revenue number has the memo.
Once the program runs, report monthly on one page: the decision or request, the month in four weekly readings, what changed that matters, what we did, what we still do not know, and resource requests. Write the decision section last; if nothing needs deciding, say so in one line. Show the four weekly readings side by side with their denominators instead of averaging them, because a blended figure can hide a mid-month change in what was measured. Include a change only if it held in at least two of the four weeks, or if it is a single event leadership must know about. Each unknown says what would resolve it and when you will look, and each resource request names what it buys and how you will judge it. The weekly reports stay as the appendix; turn a weekly analyst report into a monthly leadership update has the template.
How do you set quarterly goals you can evaluate?
Write two kinds of goal and keep them apart: deliverables your team controls, judged done or not done, and outcomes you only observe, judged against a baseline measured on a stated population, with a decision rule agreed before the quarter starts. Figure 7.3 shows one outcome goal written this way, taken from the Quillstone example at the end of this chapter.
- Goal
- Answers to Quillstone's tracked buyer prompts name Quillstone more often.
- Type
- An outcome we observe, not a deliverable we control.
- Measure
- Brand visibility: analysed answers naming Quillstone, divided by analysed answers.
- Population
- A frozen set of 10 prompts; United States; ChatGPT (app) and Gemini (app); competitors Brieflane and Clausewise; 28-day windows; at least 30 analysed answers per reading.
- Baseline
- 25.0% (120 of 480 analysed answers). Two earlier windows read 24.0% and 26.0%.
- Decision rule
- 5 points or more above baseline: Continue and widen.
- Less than 5 points either way: No clear change: extend one quarter with a sharper hypothesis.
- 5 points or more below: Investigate population, roster and engines first.
Illustrative numbers.
A bare target such as "raise AI visibility to 40%" fails twice: the number moves whenever the prompt list, engines, roster or window changes, and it blurs what the team can order. A team can rewrite a page; it cannot make an engine mention the brand.
Deliverable goals are work you can prove you finished, such as "audit the prompt list and freeze the set by week 2". Outcome goals are changes in an observed measurement that you hope the work contributes to, never guaranteed. A third kind, the coverage goal, is about the measurement itself, for example that an agreed share of expected runs completed; it stops patchy collection from deciding an outcome goal.
Every outcome goal needs five things written down before work starts.
- Measure. One metric with a one-sentence definition and a named denominator. Two or three outcome metrics per quarter is enough; each extra one is another chance to find a movement by luck.
- Population. The prompt set (ideally a fixed cohort), the engines and how each is collected, the competitor roster, the window, any filters and a minimum sample. Chapter 6 explains why each one changes the number. If any must change mid-quarter, record the date and restart the baseline.
- Baseline. Several consecutive windows, so you can see how much the number moves when nothing changes. With no history, make the first weeks of the quarter the baseline period and set only deliverable and coverage goals for that stretch.
- Threshold. The movement that would make you spend next quarter differently, larger than the spread you saw. Choose it before the result and never lower it afterwards.
- Decision rule. Three bands with consequences. Clear improvement: continue and consider widening. No clear change: the hypothesis is neither supported nor refuted, so extend the test or change the work. Clear decline: investigate first, starting with whether the population, roster or engines changed.
Link each outcome to its deliverables with a one-sentence hypothesis, so the quarter-end review evaluates the hypothesis, not the people. Set quarterly AI search goals your team can actually evaluate has the full goal sheet.
How do you plan refreshes from evidence rather than age?
Rank candidate pages by evidence that something is wrong or has shifted, multiply by how much the page matters to buyers, and fill a fixed number of slots from the top. A page with no evidence scores zero however old it is.
Age says when a page was last edited, not whether it is wrong, still cited or important to buyers, so keep it as a tie-breaker only.
Four kinds of evidence justify a refresh:
- A fact problem (F): an engine contradicts something you know to be true about your brand. In DiscoveredBy, Fact check compares answers with the facts you approved and marks each fact against the previous comparable study, so "Newly wrong" is a strong reason to find the pages that state that fact. Finding those pages is your job.
- An internal change (X): you changed a price, feature or policy and the page still describes the old state.
- A source change (S): a page engines cite for your topic now says something different. Source mention history compares saved snapshots of the most-cited pages and labels gained or lost matches for a brand. That is a reason to read the page, not a finding: it does not date the change, and a page leaving the window is not a lost mention.
- A citation trend (C): engines cite your URL more or less often than in an equal earlier window with the same engines and prompts. The Sources tab in Citations groups citations by URL, but from the newest 500 in the window, so narrow the filter on a busy project. Decide in advance how large a fall counts.
An open citation gap or optimization on the page (G) is a weaker fifth signal. One workable rubric scores F and X at 3 points, S and C at 2 and G at 1, then multiplies the sum by buyer importance from 1 to 3, where 3 is a page buyers meet at the decision point. The weights are a planning device, not a product metric, so keep yours written down and stable.
Then cap the calendar: fill a fixed number of slots per sprint from the top, re-score at the start of each cycle, give every deferred row a written "not doing because", and give pages with no evidence a light annual read instead of a rewrite. Each refresh is a hypothesis with a recheck date, judged later by a practical before-and-after review. Build a content refresh calendar from evidence, not article age has the register.
How do you review AI-written content?
Split the review into four jobs, each with a named person: a claim verifier checks every factual claim against a source, an evidence checker confirms each source really says it, an editor settles disputes, and an approver alone decides on publication. The approver is never the person who drafted or prompted the piece.
Review claims, not paragraphs. A draft reads fluently whether or not its facts are right, so an editor reading for flow will pass an invented price. A claim is any statement a reader could check: a number, a date, a capability, a comparison, a policy or a named source. Mark each one pass, fix or remove against five tests:
- Traceable: there is a named source.
- Supported: the source says it and you would stand behind it. Read the source, not the quote the draft offers.
- Current: the source is up to date.
- Ours to say: a claim about your own product, pricing or policy matches your approved source of truth.
- Plain: no superlatives or guarantees you cannot back.
Know exactly what a tool covers. DiscoveredBy's Articles labels each claim the writer reports, and a claim earns "verified by source" only when its quoted evidence is at least four words long, appears in the scraped text of the named source page and shares real terms with the claim. That does not make the source right or current. The test reads only the claims list the writer returned, so statements elsewhere in the prose go unchecked, and nothing is removed from the draft. Read the prose once more and log anything missing from the list.
Decide in advance how disputes end. The owner of a product, pricing or policy fact decides it; claims about other companies need a primary source; anything unresolved after one round comes out, rather than being softened with "reportedly". Keep a sign-off log per piece, because it is the only record that a review happened. It proves named people checked the listed claims, not that the piece is correct or that an engine will cite it.
After publication, Fact check reads engine answers about your brand, not your drafts, so it shows which approved facts engines get wrong; it is not a pre-publication check. Build a human review process for AI-written marketing content has the rubric and log.
What does a first quarter look like?
A first quarter sets up the routine first and judges an outcome only if a baseline already exists: scope and population in the first two weeks, one documented change in month one, then the weekly, monthly and quarterly rhythm.
Weeks 1 and 2: intake and population. Before configuring anything, collect the business goal, approved facts, every name the brand goes by, the competitors you lose to and the decisions buyers make, then turn only those decisions into prompts; onboard a client into AI visibility monitoring without a sprawling prompt list has the questionnaire and chapter 3 the choices. Freeze the population and set goals. With no history, the first weeks are the baseline period, the goals for that stretch are deliverables and coverage only, and outcome judgements wait for the following quarter.
Month 1: one clean loop. Measure in week 1, choose one gap, page and hypothesis in week 2, make one documented change in week 3, and review like for like in week 4. A mid-month change cannot get a firm verdict by day 30, so treat day 30 as a first read and schedule the firm review a month after the change; your first 30 days of AI search optimization walks through it.
Then the rhythm. Hold the 30-minute review every week from the first completed week. Send the leadership page at the end of each month, even if month one reads "no decision needed". In month 2, build the refresh register and send every AI-drafted piece through the review log. A new starter prompt list deserves a review about 30 days after first results; if that review changes a frozen population, record the date and restart the baseline. In week 13, fill in the goal sheet's review record and hold the quarterly review.
Worked example: Quillstone's first program quarter
Illustrative example: Quillstone, Northfold Digital and the competitors are fictional, and the numbers are made up to show the method.
Northfold Digital, the agency for Quillstone (document-review software for legal and compliance teams), has monitored ten buyer prompts for twelve weeks, in the United States, on ChatGPT (app) and Gemini (app), against Brieflane and Clausewise. That history gives three consecutive 28-day windows.
The population. Ten prompts on two engines over 28 days allow at most 10 x 2 x 28 = 560 answers per window. In the baseline window 500 runs completed and 480 answers were analysed for brand mentions. Quillstone was named in 120 of the 480, so brand visibility was 25.0%. The two earlier windows read 24.0% and 26.0%, so the number moves about two points when nothing is changed on purpose.
The goal card. Northfold writes the outcome goal shown in Figure 7.3:
- Goal: answers to Quillstone's tracked buyer prompts name Quillstone more often.
- Type: an outcome we observe, not a deliverable we control.
- Measure: brand visibility, analysed answers naming Quillstone divided by analysed answers.
- Population: the frozen set of ten prompts, United States, ChatGPT (app) and Gemini (app), competitors Brieflane and Clausewise, 28-day windows, at least 30 analysed answers per reading.
- Baseline: 25.0% (120 of 480 analysed answers), with earlier windows at 24.0% and 26.0%.
- Decision rule: 5 points or more above baseline means continue and widen the page work; less than 5 points either way means no clear change, so extend one quarter with a sharper hypothesis; 5 points or more below means investigate population, roster and engines before any rewrite.
The 5-point threshold is more than twice the two-point spread. Beside the card sit two deliverable goals (confirm and freeze the ten-prompt set by week 2; refresh the top three pages in the register by week 8, with any AI-drafted copy through the review log) and one coverage goal (at least 90% of the 560 expected runs complete in the final window, which is 504 runs).
The refresh register. Scored with the rubric above, /pricing carries a Fact check finding and a price change four weeks old: (3 + 3) x 3 = 18. The redlining feature page has an open citation gap and falling citations: (1 + 2) x 3 = 9. The SharePoint integrations page still says "coming soon" two months after launch, and an open optimization targets it: (3 + 1) x 2 = 8. A five-year-old year-in-review post has no signals and scores 0 x 1 = 0, so it is not scheduled.
During the quarter. Northfold holds 12 of 13 weekly reviews and records the skipped week as skipped. In week 4 a prompt drop is marked act; two of three saved answers show a Clausewise comparison page cited where Quillstone's used to be, which is claim strength 2 with the cause unknown. The first monthly update asks leadership for one decision: approval of the corrected pricing copy. The AI-assisted redlining section goes through the review log with eight claims: five pass, two are fixed after the product owner corrects them, and one superlative is removed (5 + 2 + 1 = 8).
Week 13. The prompt set was frozen unchanged in week 2, so the baseline stands. Two of the three refreshes shipped; the SharePoint page is waiting on product review and is reported as not done, with a blocker and an owner. In the final window 520 of 560 runs completed (92.9%), so the coverage goal is met. Quillstone was named in 160 of 500 analysed answers, which is 32.0%, a rise of 7.0 points on a sample well above 30. That clears the threshold, so the rule says continue and widen. The quarterly review reports this as an observation on the frozen population, lists what else moved during the quarter, and does not claim the refreshes caused the rise. Slide 11 offers three options: deepen on the open queue, widen to the United Kingdom (which would change what later comparisons mean), or hold and monitor.
Weekly review agenda and quarterly goal template
Copy this into your notes tool. Use the first part every week, one page per meeting, and the second part once per quarter, with one block per goal.
WEEKLY AI SEARCH REVIEW (30 minutes)
Project: ____________ Report week (Mon-Sun, UTC): ________ to ________
Facilitator: ________ Note-taker: ________ Attendees: ______________
LAST WEEK'S ACTIONS (one line each: done / not done / dropped)
- ______________________________________________________________
1. READ THE WEEK (minutes 0 to 5)
Headline reading and change vs prior week: ____________________
One-sentence summary: ________________________________________
Caveats (engines compared, blank changes, new prompts/countries): ____
Not measured this week: ______________________________________
2. ALERTS (minutes 5 to 10)
Alert / kind / prompt or brand | Act | Watch | Dismiss | Team note written?
_______________________________|_____|_______|_________|__________________
3. EVIDENCE CHECK (minutes 10 to 17) Never cut this block
Finding | Answers opened | Claim strength (1-4) | Still unknown
________|________________|______________________|______________
1 = a number moved (no answer read yet)
2 = the change is visible in the answers we read
3 = a plausible reason is visible in the answer or its sources
4 = a change we made was followed by a measured change (not proof of cause)
4. DECISIONS (minutes 17 to 25) Open-work list, highest band first
Item | Do / Defer / Drop | Hypothesis to test | How we will check it
_____|___________________|____________________|_____________________
5. ASSIGNMENTS (minutes 25 to 30) Owners live in our task tracker
Task | Owner | Due | Tracker link | Check-back week
_____|_______|_____|______________|________________
WATCH LIST (carry to next week): ________________________________
"Not enough to conclude" items: _________________________________
QUARTERLY GOAL TEMPLATE
Quarter: ________ Owner: ________ Review date: ________
POPULATION (frozen; record any change with its date and restart the baseline)
Project / domain: ______________________________________________
Prompt set (name, count, date saved): ___________________________
Engines (and how each is collected): ____________________________
Competitor roster: _____________________________________________
Countries / languages / personas: ______________________________
Baseline window: ________ Final window (same length): ________
Minimum sample per reading (under 30 observations is provisional): ____
GOAL CARD (one per goal)
Goal: __________________________________________________________
Type: [ ] Deliverable we control [ ] Outcome we observe [ ] Coverage
Measure (definition and denominator): ___________________________
Population: as above, or narrower: _____________________________
Baseline (value, n, earlier windows for spread): ________________
Threshold for "clear change": __________________________________
Linked deliverables and one-sentence hypothesis: ________________
Decision rule (agreed before work starts):
Clear improvement -> __________________________________________
No clear change -> __________________________________________
Clear decline -> investigate population, roster and engines first, then ____
Owner: ________ Due or final reading date: ________
KNOWN LIMITS
What these goals cannot tell us: ________________________________
What could distort them (roster change, engine change, missing days): ____
REVIEW RECORD (fill at quarter end)
What was delivered (done / not done, with blockers): ____________
What was observed (value, n, band): _____________________________
Decision taken and why: ________________________________________
Carry forward / stop / change: _________________________________
In DiscoveredBy
The weekly report is built every Monday, depending on your plan, for the last complete Monday to Sunday week in UTC, with no language model involved; it shows the prompts that moved, a "Do this next" checklist and past weeks from a dropdown, which makes it the input to the weekly review. Alerts supply the changes for the alerts block, and owners and editors can acknowledge, dismiss or snooze each one with a shared team note. Client reports shows the same completed week across your projects as separate readings, never a ranking, and opens each project's exact report. Client report links share a fixed digest of one week that opens without an account, expires after 1, 7 or 30 days and can be revoked. If you would rather have this program run for you, Enterprise is a managed engagement built on this guide, with an onboarding audit in the first 2 weeks and a monthly strategy call.
Go deeper
- Run a weekly AI search review meeting in 30 minutes
Read the week, check the alerts, test the evidence, decide, assign.
- Monthly AI visibility update for leadership: a one-page template
Turn four weekly reports into one page of decisions, uncertainties and requests.
- Explain AI visibility to leadership without promising a revenue number
A decision memo that asks for a funded first program.
- Set quarterly AI search goals your team can actually evaluate
Split goals into deliverables you control and outcomes you only observe.
- How to write an AI visibility report a client can act on
A five-section weekly report template, with an example where the numbers are flat.
- Build a content refresh calendar from evidence, not article age
Rank refreshes by outdated facts, source changes, citation trends and buyer importance.