Choose the prompts to track
Choose a small set of buying questions to track, control how wording changes the audience, and spend a limited prompt budget where decisions happen.
Chapter 3 of 7: contents
To choose the prompts to track, start from the decisions your buyers make, not from keywords or from every question you could think of. Write each decision as the full question a buyer would put to an AI engine, label it by type and buying stage, decide its audience, language and place, and keep it only if a result would change something you do. Every combination of location, audience and language is a separate tracked variant with its own cost, so cost the list before you fill it. Keep seasonal questions in their own segment, and audit for overlap before you add more. The result is an editorial sample of your market, not a measure of what people ask.
In short
- A tracked prompt is a full buyer question run against AI engines on a schedule. It is not a keyword, and a list of them is your chosen sample, not a count of demand.
- Cover the whole path to a decision: problem and category, shortlist, comparison, evaluation, and decision and objection questions, across the awareness, consideration and decision stages.
- A small change in wording can turn a prompt into a different question. Track a new wording beside the original instead of editing it in place.
- Personas, languages and places multiply a prompt's cost, and some of those variants do not run on every engine. Add them where the buyer, language or market really changes the question.
- Spend a fixed budget on the decisions closest to a purchase first, keep seasonal prompts apart, and audit what you have before you add more.
Why can't you track every prompt?
Because every prompt costs something and means something. Each one uses prompt slots, adds answers to every average you report, and has to be explained when a number moves, so a list that tries to cover everything ends up measuring nothing in particular.
A tracking prompt is a complete question that a monitoring tool runs against AI engines on a schedule, so you can see whether your brand is named or cited in the answer. "document review software" is a keyword. "What is the best document review software for a small compliance team?" is a prompt, because it has a buyer and a decision inside it. In DiscoveredBy a prompt is stored as its own row and is what actually runs, while a keyword is only a label you can attach to file it (Prompts). The glossary entry for prompt has the short definition.
Cost is the first reason to be selective. In DiscoveredBy a prompt is tracked as one or more targets, each one combination of location, audience and language, and each target uses one prompt slot. A prompt tracked in three countries, as the plain General audience plus two personas, in one language, produces 3 x 3 x 1 = 9 targets (Prompts). Slots pool across every project you own, so more projects do not mean more slots (Plans and limits). Choosing which engines run a prompt costs no slots.
Meaning is the second. A visibility rate is a share of answers, so it reflects which questions you asked as much as how engines answered them. Every prompt you add changes that mix and makes month-to-month comparisons harder to explain.
Honesty is the third. Your prompt set is editorial: it reflects your judgement about your buyers. Nothing in DiscoveredBy measures how often a question is put to ChatGPT, Gemini or any other assistant (Prompts). With Search Console connected, each prompt can show Google demand: the impressions Google recorded over the last 28 days for the queries that prompt covers. That counts what people typed into Google, not what anyone asked an AI assistant, but it helps you rank your own prompts against each other. Its 1 to 5 score is relative to your own project, and a dash means no query matched, which is not the same as no demand. Use it to order a list, not to veto a buying question you know matters.
So start with a set you can review by hand, and grow it only when a new prompt tests something the current set does not. How to choose AI tracking prompts that reflect buying decisions walks through building that starter set.
What types of prompt are there?
Sort prompts by where the buyer is on the way to a decision, and there are five types: problem and category questions, shortlist questions, comparison questions, evaluation questions, and decision and objection questions. Label each prompt with one of three buying stages (awareness, consideration or decision) so you can see whether the list covers the whole path or only one moment of it.
The five types follow the order in which a buyer works:
- Problem and category questions. What is the problem, and what kind of product solves it?
- Shortlist questions. Which products should I look at, and what are the alternatives?
- Comparison questions. How do these two or three options differ on the things I care about?
- Evaluation questions. What does it cost, how hard is it to implement, what does it integrate with, what are its limits?
- Decision and objection questions. Is it right for a team like mine, and what would make me hesitate?
A type says what the question asks; a stage says how close the asker is to choosing. A list of only "best X for Y" questions covers one type at one stage and tells you nothing about the rest.
To make this concrete, this chapter follows one company throughout. Illustrative example: Quillstone and its competitors are fictional, and the numbers are made up to show the method. Quillstone sells document-review software to mid-sized legal and compliance teams, and its competitors in the example are Brieflane, Clausewise and Docket North. Its first draft list placed its questions like this:
- Problem and category, awareness: "What does document review software do for a compliance team?"
- Shortlist, consideration: "What is the best document review software for a mid-sized legal team?" and "What are alternatives to Brieflane for contract review?"
- Comparison, decision: "How does Quillstone compare with Clausewise for regulatory document review?"
- Evaluation, consideration: "How long does it take to roll out document review software to a legal team?"
- Evaluation, decision: "How much does document review software cost per user?"
- Decision and objection, decision: "Is Quillstone good for legal teams?"
Figure 3.1 places the five prompt types against the three buying stages, with Quillstone's draft questions where the team put them. A type does not belong to one stage: the two evaluation questions landed in different stages. The point of the grid is to show the empty cells.
Scroll sideways to see the whole table.
| Type / stage | awareness | consideration | decision |
|---|---|---|---|
| Problem and category | What does document review software do for a compliance team? | None yet | None yet |
| Shortlist | None yet | What is the best document review software for a mid-sized legal team? What are alternatives to Brieflane for contract review? | None yet |
| Comparison | None yet | None yet | How does Quillstone compare with Clausewise for regulatory document review? |
| Evaluation | None yet | How long does it take to roll out document review software to a legal team? | How much does document review software cost per user? |
| Decision and objection | None yet | None yet | Is Quillstone good for legal teams? |
DiscoveredBy uses a closed vocabulary for the same idea, so your labels can map straight onto fields in the tool. Each prompt can carry an intent (informational, navigational, commercial, comparison or troubleshooting), a buyer stage (awareness, consideration or decision), a theme (competitor alternative, pricing, use case, implementation, troubleshooting, regional opportunity, comparison or other) and a branding label, branded or unbranded (Prompts). All four are optional. A prompt with no classification is normal, and it is not the same as one classified as "other".
A branded prompt tells you how an engine describes a brand when asked directly; an unbranded one tells you whether you appear when nobody asks for you. Keep some of each, and report them separately.
Then run every draft through the inclusion rules. A prompt earns a place only if it passes all of them:
- It is a full question a buyer would plausibly ask. If you cannot say who asks it and when, cut it.
- It maps to a type, a stage and a theme. A prompt you cannot classify is usually too vague to interpret.
- It tests something no other prompt in the set tests. Near-duplicates measure the same thing twice and use up slots.
- A result would change something you do: update a page, brief a writer, investigate a source.
- You know what a correct answer looks like. If you cannot state the fact, you cannot judge the response; Get your brand ready covers writing those facts down.
- Its locations, audiences and languages are justified, because every combination is its own target.
- It is neutral enough to be a fair test.
Quillstone's draft also held "Best document review software", which failed the first rule (it is a fragment, and a near-duplicate of the shortlist question), and "Tell me why Quillstone is the best", which failed the last (it is leading, and tests nothing new). Both were cut.
How does wording change the audience you measure?
Wording decides which question you are measuring. Rephrasing that keeps every condition can share a trend line with the original once you have checked the answers look alike; adding a constraint such as "for a small team", narrowing the scope, or saying who is asking creates a new question, which you track beside the original rather than in its place.
Most edits fall into one of four kinds:
- Pure phrasing. "best document review software for legal teams" becomes "What is the best document review software for legal teams?" Usually the same question, pending a check.
- Scope change. "for legal teams" becomes "for in-house legal teams". A new question.
- Buyer constraint. Adding "for a small team", "on a budget" or "with SOC 2". A new question.
- Who is asking. Adding "as a procurement lead" or "I am a solo founder". A new question, or an audience, depending on how you track it.
The first kind can quietly become the third. Adding "small" feels like polish, but an engine answers the whole message, and a stated team size is information it can use: it may recommend different products, add caveats about pricing, or do almost nothing with the phrase. You cannot tell which from the wording alone.
Figure 3.2 shows one of Quillstone's base prompts, "What is the best document review software for legal teams?", with the modifiers that change who it measures. The phrasing-only version sits on the same-question side; the scope change, the small-team constraint and the "as a procurement lead" version sit on the new-question side.
Base prompt
What is the best document review software for legal teams?
Same question, pending a check
Can share the original's trend line.
- Pure phrasing: best document review software for legal teams
New question
Track beside the original as a new prompt, with its own trend line.
- Scope change: for in-house legal teams
- Buyer constraint: for a small team, on a budget, with SOC 2
- Who is asking: as a procurement lead (a new question, or an audience if tracked as a persona)
So never edit a tracked prompt in place: its chart would then mix two questions you cannot separate afterwards. Add the new wording as a sibling, keep the original running, and compare the two over several days on the same engines and locations. Compare four things: which brands were named, how they were framed, which reasons were stated, and which sources were credited. Then decide whether the sibling is the same question in practice, a different one, or mixed, and log the decision with its date.
Quillstone ran this test. After five daily runs on one engine and one country, the original and the phrasing-only sibling each named Quillstone in 3 of 5 answers, while the small-team version named it in 1 of 5. The team kept watching the phrasing change as the same question and gave the small-team version its own trend line. It did not conclude that its product is worse for small teams: five answers each cannot show that the phrase caused the difference.
Text in the prompt travels to every engine the prompt runs on, while a persona (next section) reaches only the chat engines. Google AI Overviews, Google AI Mode, ChatGPT (app) and Gemini (app) receive only the prompt text, so wording is the only way to put "small team" in front of them (Personas).
Extra clauses are not free: a prompt over 700 characters once encoded does not run on the two Google engines, and one over 2,000 does not run on ChatGPT (app) or Gemini (app) (Languages and templates). Test one or two prompts where a phrase keeps coming up in sales calls rather than adding it everywhere. Does adding "for a small team" change the question you are measuring? gives the full decision tree and a variant log.
How do personas, languages and places fit in?
They are three more dimensions of the same prompt, and each one multiplies its cost: a tracked variant is one location, one audience and one language. Add one only where the buyer, the language or the market really changes the question, and keep a plain baseline beside it so the variant can be read.
Personas
A persona is a named buyer profile, a name and a short description, that a tracked prompt can run as, alongside or instead of the plain General audience (Personas). When a persona variant runs on a chat engine, the message gets an audience block after the prompt text asking the engine to answer for that person. Nothing else separates it from a General run: it is not a signed-in session or a saved profile. It does not run on Google AI Overviews, Google AI Mode, ChatGPT (app) or Gemini (app), yet still uses its slot.
Three habits keep persona results readable. Keep General selected on every prompt you give a persona, because General can be deselected and without it you cannot tell a persona effect from ordinary variation. Compare a persona with General only on the same prompts, engines, countries and window, leaving out the four engines a persona never runs on. And pick two or three personas that match real buying roles, applied to your most decision-relevant prompts. Editing a persona's description changes only future runs, so record the date of every edit.
Persona results describe the answer when a question is asked on behalf of that kind of buyer, not what real people in that role see. Measure AI visibility for different buyer personas without mixing the results covers the comparison in full.
Languages
A language variant does not translate your prompt. The exact text you wrote is sent unchanged: chat engines also get an instruction to answer in that language, while the Google engines and the two app engines get the language only as a setting of the request (Languages and templates). To track what a French-speaking buyer types, you write the prompt in French yourself and set its language to French.
Hold the intent fixed and let the wording vary. Write the buyer's decision as one plain sentence, have a fluent speaker phrase the prompt as a buyer in that market would type it, and have a second fluent reader confirm it implies the same decision. Check the local terms for the category and job titles. A machine-translated prompt may measure how an engine handles awkward text rather than your buyer's experience.
Watch the grid. Two countries and two languages on one prompt give every combination, including German in France, so one prompt per language and market is usually cleaner. Coverage differs too: a Chinese variant does not run on the Google engines, ChatGPT (app) or Gemini (app), and an As written target on those four uses the project's default language, or English. Report each language as its own row with its own count of answers. Track AI visibility across languages without losing comparability has a setup checklist.
Places
Every prompt tracks at least one location: a whole country, or a city (Prompts). A city is a location of its own, so it uses a slot for each audience and language it runs as. You pick cities from a fixed list of 344 across the 18 countries a project can target; nobody types a city name. A city target does not run on ChatGPT (app), while Gemini (app) and the two Google engines receive a city as coordinates.
A target location is not a customer's location. Answers are collected from servers, not from a device in that place, and each engine decides how much weight to give the location (Engines and measurement). Pair every city target with a country-wide target of the same prompt, so only the location differs. Put a place name in the wording only when a real buyer would, and choose questions where place plausibly matters, such as local providers or regional rules. CSV or Excel import and accepting a suggestion add country-wide targets only, so city targets come from the prompt forms or a saved template. Local AI visibility: how to compare city and country results responsibly covers reading the gap.
What about seasonal questions?
Track them as their own segment. A seasonal prompt is a question that matters only, or far more, during part of the year, such as a compliance deadline, a budgeting round or a procurement cycle; give each season its own tag and dates, compare it only with the same season or a pre-season reading, and keep it out of your always-on numbers.
A useful test: would you still want the question answered in the quietest month of the year? If yes, it is probably always-on. Keep a borderline prompt in the always-on set with a note, because moving prompts between segments later changes both populations.
The reason is composition. With counts kept small to show the arithmetic, suppose Quillstone's always-on prompts named it in 16 of 40 analysed answers in both an off-season month and the peak month: 40.0% each time. In the peak month a seasonal set about year-end audit review added 30 analysed answers, 6 naming Quillstone, or 20.0%. Pooled, the peak month reads 22 of 70, or 31.4%: a fall of 8.6 points while the always-on figure did not move.
Set the segment up before the season opens:
- Create one tag per season, not one "seasonal" tag for everything, and attach it to the seasonal prompts only.
- Write down the pre-season start, the season start and end, and the review date.
- Start collecting before the season opens, so you have a pre-season reading of the same prompts.
- Pause the prompts when the season ends, and record the date.
Pausing stops future scheduled runs while keeping a prompt's history, and resuming checks that the project owner has free prompt slots (Prompts). If the season's slots are spent elsewhere when it comes round, the prompts cannot resume, so plan them in advance. Compare a season with the same season a year earlier, or else with a pre-season window fixed in advance, never with whichever month came before. Count the answers before you read a rate. Seasonal prompts and AI visibility has a copyable seasonal calendar.
How do you split a limited budget?
Count the cost of every candidate prompt in slots, rank the candidates by the buying decision they cover, and fund down the ranked list until you reach what you can spend, holding a small reserve back. Fund decisions first, then markets, then variants, because depth on the questions where a shortlist forms usually tells you more than thin coverage of everything.
A workable order:
- Decision-stage and comparison prompts for your main product in your main market.
- The same for your second product, so you can compare products fairly.
- Consideration prompts, one or two per product.
- Market variants of your best prompts, added one market at a time.
- Persona and language variants, only for prompts where the audience or language changes what you would do.
Awareness prompts come last here because they say least about your product; a team creating a new category may reverse that. Before adding a market, ask whether you can act on a finding there, whether it is really a different question, and whether you can compare its results with the rest.
Keep a reserve, because accepting a suggested prompt or resuming a paused one needs free slots. Write down every prompt you chose not to fund, with a reason: those rows are your measurement gaps. Every product and market you report on gets at least one funded prompt, or is written off as not measured.
Quillstone sells two products, Quillstone Review and Quillstone Audit, mainly in the US and UK, with a small new pipeline in Germany. It decided to spend 44 slots on this project, held 4 in reserve, and allocated 40:
- Core product questions: 16 slots. Eight Quillstone Review prompts, mostly decision and comparison questions plus two consideration ones, each in the US and the UK, General, As written: 8 x 2 x 1 x 1 = 16.
- Second product: 8 slots. Four Quillstone Audit decision and comparison prompts in the same two countries: 4 x 2 x 1 x 1 = 8.
- New market: 4 slots. Four Review comparison prompts written natively in German, tracked in Germany in German only: 4 x 1 x 1 x 1 = 4.
- Personas or languages: 6 slots. A "Compliance manager" persona added to three of the core prompts, in both countries: 3 x 2 x 1 x 1 = 6 more targets.
- Seasonal: 6 slots. Three year-end audit prompts in the US and the UK, General: 3 x 2 x 1 x 1 = 6, earmarked for the season.
Figure 3.3 shows that split. The arithmetic: 16 + 8 + 4 + 6 + 6 = 40 allocated, plus 4 in reserve, makes 44. The 40 slots cover 19 prompts (8 + 4 + 4 + 3), because the persona segment adds targets to existing prompts rather than new ones.
- Core product questions: 16 slots. 8 prompts × 2 countries
- Second product: 8 slots. 4 prompts × 2 countries
- New market: 4 slots. 4 German-language prompts in Germany
- Persona: 6 slots. 1 persona on 3 core prompts × 2 countries
- Seasonal: 6 slots. 3 prompts × 2, kept free for the season
- Reserve: 4 slots. Free for accepting a suggested prompt or resuming a paused one
44 slots in all. Illustrative numbers.
Three decisions came out of it. The German block uses one natively written language per prompt rather than As written plus German, which would have doubled its cost. The persona block is a targeted test, because persona variants never run on the two Google engines or the two app engines. And two candidates went on the "later" list with reasons: a London city target on two core prompts (2 slots) and consideration versions of the Audit prompts. Off-season the 6 seasonal slots are free, but the team keeps them unspent so the seasonal prompts can resume on the pre-season date.
None of this says the allocation is right. It says the trade-offs were written down. Allocate a limited prompt budget across products and markets offers a slot ledger with an optional ranking score, a planning heuristic that nothing in the product computes.
How do you audit a list before adding to it?
Judge each prompt by what its answers contain, not by how its wording looks. Two prompts are redundant when everything one of them shows regularly is also seen by another prompt you keep; check that over a fixed window, protect any prompt other work depends on, preview what a pause would remove, and only then pause.
Begin with a manual pass for suspects: pairs that ask the same question in different words, prompts nobody can explain, and prompts that are narrower versions of another. That shortlist is not a verdict, because differently worded questions can draw on the same sources and brands while similar-sounding ones may not.
DiscoveredBy's Prompt coverage screen compares answers. It reads the completed answers from the 28 days ending yesterday and treats two kinds of item as coverage: the domains cited and the brands named. An item is regular for a prompt when it appears in at least 10% of that prompt's answers (and at least 2 of them) on at least 2 different days, and seen when it appears in at least 2.5% (and at least 1). With 140 answers, regular means at least 14 and seen at least 4. The screen recommends a set you could pause together while every item those prompts see regularly is still seen by a prompt you keep. A prompt is judged only once it has analysed answers on at least 14 different days in the window, so a prompt you added last week is never recommended.
Some prompts should never be paused on redundancy alone, and the screen marks them Protected: a prompt chosen for fact checks, one named by an open or unsettled optimization, one linked to a citation gap in progress, and the last running prompt for a country, city, persona, language or engine.
Before pausing, preview the selection: slots freed, answers removed, regular items no remaining prompt would see, occasional items that would go too, any value that would lose its last prompt, and your brand's rate now and after. At Quillstone, two core prompts looked alike: "Best document review software for legal teams" and "Top contract review tools for compliance teams". The preview for pausing the first showed nothing lost. Pausing the second would have dropped Docket North, which only it saw, occasionally. The team paused the first and noted Docket North as a name to watch.
A paused prompt keeps its history, and resuming it needs free slots. A pause also changes the mix behind every rate, so log the date and the prompts paused; Measure change honestly covers fixed cohorts. Audit once a quarter or before adding a batch. Audit your prompt list before adding more prompts has the full checklist.
Prompt list template
Copy this into a spreadsheet or a doc. Fill one row per prompt, funded or not, so the "later" rows stay visible as gaps. The two example rows come from the Quillstone example above.
PROMPT LIST
Project: ______________ Owner of this list: ______________ Date: __________
Slots I can spend: ____ Reserve held back: ____ Allocatable: ____
Slots in use: ____ Prompts tracked: ____ Last audit date: __________
COLUMNS (one row per prompt; each location x audience x language = 1 slot)
ID
Prompt the full question, in the buyer's words
Type problem and category / shortlist / comparison /
evaluation / decision and objection
Stage awareness / consideration / decision
Intent informational / navigational / commercial /
comparison / troubleshooting
Theme competitor alternative / pricing / use case /
implementation / troubleshooting / regional
opportunity / comparison / other
Branded? branded / unbranded
Persona General, or the named persona(s); General kept beside
any persona? (Y/N)
Language As written, or the language it was written natively in;
native wording reviewed by: ______
Place countries; cities only where place changes the answer
(each city paired with its country-wide target)
Segment core / second product / new market /
persona or language / seasonal (tag: ______)
Slots locations x audiences x languages = ____
Engines skipped persona, city or language variants that some
engines do not run
Why it is tracked one line: the decision a result would inform
Owner who reads it and acts on it
Status funded / later / cut (reason)
Date added / date changed
INCLUSION CHECK (every funded row passes all seven)
[ ] A full question a buyer would plausibly ask
[ ] Maps to a type, a stage and a theme
[ ] Tests something no other prompt in the list tests
[ ] A result would change something we do
[ ] We can state what a correct answer says
[ ] Each location, audience and language is justified
[ ] Neutral enough to be a fair test
WORDING CHANGES (never edit a tracked prompt in place)
Original | Sibling | Edit kind (phrasing / scope / constraint /
who is asking) | Date started | Decision (same / different / mixed)
EXAMPLE ROWS
Q1 | How does Quillstone compare with Clausewise for regulatory
document review? | comparison | decision | comparison
| comparison | branded | General | As written | US, GB | core
| 2 | none
| Shows how we are framed against our closest rival at the point
of choice | Content lead | funded | 2026-10-01
Q2 | What is the best document review software for a mid-sized legal
team? | shortlist | consideration | commercial | use case
| unbranded | General | As written | US, GB; London, GB later
| core | 2 now, 3 with London
| ChatGPT (app) would not run the London target
| Shows whether we make the shortlist; London would test whether
city answers differ from UK-wide ones | Content lead
| funded; London later (budget) | 2026-10-01
FOOTER FOR ANY REPORT USING THIS LIST
"Measured on __ prompts across __ markets and __ products.
This is a chosen sample, not the whole market."
In DiscoveredBy
The Prompts screen holds your list: each row shows its locations, audiences, languages and engines, the four classifications can be set by hand or filtered on, and Google demand appears once Search Console is connected. Personas covers defining buyer profiles, attaching them to prompts beside General, and asking each persona to suggest buyer-intent questions for review. Languages and templates covers language variants and a preview that states the variants and new slots a change would use before you save, plus templates that reuse one combination of countries, cities, audience, languages and engines. Prompt coverage recommends prompts you could pause while what they see regularly stays covered, and previews any selection first. When the list is set, Run your first audit is the next step.
Go deeper
- How to choose AI tracking prompts that reflect buying decisions
A starter prompt worksheet by buyer question, intent, stage, audience and location.
- Audit your prompt list before adding more prompts
Find prompts that repeat each other and preview what pausing them would cost.
- Does adding "for a small team" change the question you are measuring?
A decision tree separates harmless rewording from a new constraint or audience.
- Allocate a limited prompt budget across products and markets
A slot-ledger worksheet for spending a fixed prompt budget on the decisions that matter.
- Measure AI visibility for different buyer personas without mixing the results
Track personas separately from a General baseline and compare like with like.
- Track AI visibility across languages without losing comparability
Write each prompt natively, keep the setup fixed, and report each language on its own.