Use a fixed prompt cohort for month-to-month comparisons
Adding prompts changes what your AI visibility rate measures. Freeze a comparison cohort, keep growing the live set, and use a copyable cohort register to report both honestly.
On this page
- In short
- Why does adding prompts break a month-to-month comparison?
- What is a prompt cohort, exactly?
- How do you choose which prompts go into the cohort?
- How do you mark the cohort inside DiscoveredBy?
- How do you compute the cohort trend?
- The deliverable: a cohort register
- A worked example
- When should you retire or replace a cohort?
- Common mistakes and what a cohort cannot tell you
- Frequently asked questions
- Next step: freeze your first cohort
To compare your AI visibility from one month to the next, freeze a prompt cohort: a fixed list of prompts, written down on a known date, that you report on every month and never edit. Keep adding prompts to your live tracking set as your questions grow, but compute the month-to-month trend only on the cohort. Otherwise a rate such as brand visibility moves because the questions changed, not because the answers did. The cohort needs a register that records which prompts belong to it, when it was frozen, and what else stays constant: engines, countries, personas and languages.
In short
- A rate like "brand visibility" is a share of answers. If you add prompts, you change the population it is a share of, and the number can fall while nothing about your brand got worse.
- A prompt cohort is a frozen list of prompts. It answers "did anything change on the same questions?" The live prompt set answers "how are we doing on everything we now track?" Report both, labelled.
- Freeze the cohort with a register: prompt ids, exact wording, frozen date, engines, locations, personas, languages and the rule for retiring a prompt.
- The DiscoveredBy docs describe no dedicated cohort feature, so this post builds one from tags, the Explorer and exports, each of which the docs do describe.
- A cohort is a measuring instrument, not a goal. Do not write prompts to flatter it, and say so when it goes stale.
Why does adding prompts break a month-to-month comparison?
Adding prompts changes the denominator, and a rate is only comparable across months when the denominator is made of the same questions. Brand visibility, as the metrics reference defines it, is analysed answers naming your brand divided by analysed answers. Every new prompt adds analysed answers to the bottom of that fraction.
Suppose you add a batch of questions about a use case where you are rarely named. Your visibility rate drops, and the chart says you lost ground. In truth, you asked harder questions.
Pausing prompts, adding a country or a persona, or a change in which engines run has the same effect. The fix: decide what is held constant, write it down, and report the constant part separately.
If your rate moved and your website did not, that is a different diagnostic; see Why your AI visibility score changed when your website did not. This post covers the case where your own prompt list is the moving part.
What is a prompt cohort, exactly?
A prompt cohort is a fixed set of tracked prompts chosen on a specific date and kept unchanged for a comparison period. The unit that has to stay fixed is not only the prompt text. In DiscoveredBy, a prompt is a full question stored as its own row, and what actually runs is a prompt target: the prompt paired with one location. The prompts documentation explains that a prompt tracked in three countries produces three targets, each checked independently against the engines on your plan. Persona and language variants are further axes.
So a cohort is a list of prompts plus the conditions they ran under. Write those conditions down, because any of them can change the answer population without a prompt being edited.
We use three terms consistently below:
- Live set: every prompt you currently run, including the ones you add this month.
- Cohort: the frozen subset you use for month-to-month comparison.
- Freeze date: the day you wrote the register. Earlier answers exist, so you can compute history for the cohort, but membership is decided at the freeze date and never re-decided in hindsight.
How do you choose which prompts go into the cohort?
Pick prompts you expect to still matter in a year, and choose them before looking at how you perform. Picking the prompts where you look good produces a cohort that flatters.
Some practical selection rules:
- Include prompts across the buying journey, not only the ones where you are strong.
- Prefer stable wording. A prompt with a date or a promotion in it will be out of date soon. Seasonal questions belong in their own segment; see Separate seasonal buying questions from your always-on prompt set.
- Remove redundant near-duplicates first. A cohort of twelve distinct questions is more informative than a cohort of twelve rewordings of two questions.
- Choose a size that yields enough answers. The metrics reference marks a rate built from fewer than 30 observations as provisional, so check that your cohort produces at least 30 analysed answers per engine per month before you read the rate as settled. If it produces fewer, lengthen the comparison period or enlarge the cohort rather than interpreting noise.
How do you mark the cohort inside DiscoveredBy?
Use a tag. A tag is a short label attached to a prompt, scoped to the project (see Tags and keywords). Create a tag such as cohort-2026-09 and attach it to each cohort prompt from the prompt form, or in bulk: the prompts documentation says bulk changes can add, remove or replace tags on a selection, and they are atomic, so an invalid change rejects the whole batch.
Tags let you filter the Prompts list to the cohort, and the Explorer has a Tag dimension that works as both a filter and a breakdown, so you can restrict brand metrics to the cohort over any window you choose.
These limits affect whether the cohort stays frozen:
- Tags are today's tags. The Explorer reads tags (like intent, buying stage, theme and branding) from each prompt as it is now, including for its past answers, and every result carries a note saying so. If someone removes the tag from a prompt, the prompt drops out of the cohort for every past month too. Tags can be changed by anyone with write access to the project (owners and editors, not viewers), so agree that nobody touches the cohort tag, and keep the register as the source of truth.
- The reverse also holds. Attaching the tag to a new prompt pulls that prompt into every past month of the cohort. Watch the count of tagged prompts against the register.
- Prompt wording is also today's. The Explorer shows a prompt by its current text, and an edited prompt keeps its tag. The register's exact wording is how you notice an edit; the Answers export carries the prompt as it was sent.
- The tag tells you membership, not the freeze date. The register, not the tag, records when and why.
A tag cannot be renamed once it exists (the docs say the tags router only creates and deletes), so put the year and month in the name and give a later cohort a new tag.
How do you compute the cohort trend?
You have two routes: the Explorer for a quick read, and exports for a record you own.
In the Explorer, set a filter on Tag equal to your cohort tag, choose the metric, and run the same query over two custom windows, one per month (a custom window ends yesterday or earlier and spans at most 366 days). Compare like with like: the same engines (engines are pooled unless engine is a breakdown or filter), the same country, and the same persona and language. The Explorer counts completed answers, and includes paused prompts and targets that have answers in the window. A deleted prompt's answers are gone, which is why the register rule below is "pause, never delete".
There is no dedicated cohort view; save the tag-filtered query as an ordinary Explorer view, or repeat it each month.
With exports, download the Answers dataset for each month from the Exports screen, then filter to your cohort in a spreadsheet or script. Exports are available depending on your plan; see Plans and limits. Answers has a brand_mentioned column that is empty, not false, when brand extraction has not run on that answer, so leave empty rows out of your denominator. Three facts from the exports docs matter here:
- The Prompts dataset is not windowed. It shows the project's current configuration, so a date range does nothing to it. That makes it a snapshot of prompts, wording and locations, so download it on the freeze date and save the file with the register. The docs do not list a tag column for it, so the register's prompt ids, not the file, record who was in the cohort.
- Answer features carries a
prompt_idthat joins to the Prompts export, so use it to match answers to prompts by id. The Answers dataset carries the prompt as it was sent and, in the docs' column description, no tag or prompt id, so match on the wording in your register. - Answers is capped at five thousand rows per download, counted across the whole project, not only your cohort. A download over the cap is refused rather than truncated, and the docs advise narrowing the date window. A month per download may or may not fit, so check the row count.
Read CSV files by header name, not column position; the exports docs warn that columns have been inserted mid-row.
Whichever route you use, compute the rate from the same definition every month, because a rate depends on its denominator. See also Why two teams can calculate different citation rates from the same answers.
The deliverable: a cohort register
Keep one register per cohort, stored next to the Prompts export you saved on the freeze date. It has a header block of conditions and a table of members.
COHORT REGISTER
Cohort name / tag: cohort-YYYY-MM
Freeze date: YYYY-MM-DD
Owner (only editor): name
Reason for this cohort: one sentence
Engines held constant: e.g. ChatGPT (app), Perplexity
Collection channel: app / API per engine (do not mix channels)
Locations held constant: countries and cities
Personas held constant: General only / list
Languages held constant: As written / list
Metrics reported: e.g. brand visibility, citation rate (with definition)
Competitors tracked: list (share of voice depends on it)
Comparison window: calendar month, exact start and end dates
Minimum answers per cell: 30 (below this, label "provisional")
Prompts file saved: filename of the Prompts export from the freeze date
Retirement rule: pause, never delete; a retired prompt stays listed
Next review date: YYYY-MM-DD (cohort replaced, not edited)
| # | Prompt id | Exact wording at freeze | Buying stage | Locations | Added to cohort | Status | Notes |
|---|---|---|---|---|---|---|---|
| 1 | YYYY-MM-DD | active | |||||
| 2 | YYYY-MM-DD | active | |||||
| 3 | YYYY-MM-DD | paused, reason |
And a monthly comparison log, so each month's number is recorded with its conditions:
MONTH: YYYY-MM
Cohort: cohort-YYYY-MM (n prompts, m targets)
Cohort prompts running: n of n (list any paused or failed)
Cohort answers analysed: count
Cohort metric value: x.x% (numerator / denominator)
Live-set answers analysed: count
Live-set metric value: x.x% (numerator / denominator)
Prompts added this month: count, with tag or theme
Changes that could break the comparison:
engine list / channel / model change: yes / no
location, persona or language change: yes / no
missing collection days: yes / no
Reading (one sentence, hypothesis not cause):
A worked example
Illustrative example: Quillstone and its competitors are fictional, and the numbers are made up to show the method.
Quillstone sells document-review software to mid-sized legal and compliance teams. Its competitors in this example are Brieflane and Clausewise. On 31 August the team freezes a cohort of 12 prompts in the tag cohort-2026-08, tracked in one country on one engine, ChatGPT (app), with the General audience and As written language.
For simplicity, assume every prompt target ran once a day and every answer was analysed. August has 31 days, so the cohort produced 12 × 31 = 372 answers. September has 30 days, giving 12 × 30 = 360.
In September the team adds 12 new prompts about contract-redlining, a topic where Quillstone is rarely named. The live set is now 24 prompts.
| Month | Set | Prompts | Analysed answers | Answers naming Quillstone | Brand visibility |
|---|---|---|---|---|---|
| August | Live set (same as cohort) | 12 | 372 | 93 | 25.0% |
| September | Live set | 24 | 720 | 144 | 20.0% |
| September | Cohort only | 12 | 360 | 108 | 30.0% |
| September | New 12 prompts only | 12 | 360 | 36 | 10.0% |
Check the arithmetic: 93 ÷ 372 = 25.0%; 108 + 36 = 144, and 360 + 360 = 720, so 144 ÷ 720 = 20.0%; 108 ÷ 360 = 30.0%; 36 ÷ 360 = 10.0%.
A live-set-only report says visibility fell five points, from 25.0% to 20.0%. A cohort-only report says it rose five points, from 25.0% to 30.0%. Neither is wrong; they answer different questions. The honest report states both, plus the split: on the questions Quillstone tracked in August, its share of answers naming it rose, and the new redlining questions give a 10.0% baseline to track from now on.
The five-point rise on the cohort is an observation, not a cause. If Quillstone published a page in August, the team can hypothesise a link and test it, as a before-and-after review would, but the cohort number alone does not prove it.
When should you retire or replace a cohort?
Replace a cohort on a schedule you set in advance, not when the numbers disappoint. Choose a review interval, such as every six or twelve months. At the review, create a new cohort with a new tag and run both in parallel for at least one comparison period. The overlap shows how the new cohort's rate relates to the old one on the same days.
Do not edit a cohort in place; you could no longer tell whether a number changed because of the answers or the membership.
Two events force a decision before the scheduled review:
- A prompt becomes unusable. A product is discontinued, or a question no longer makes sense. Pause it, mark it retired in the register with the date and reason, and report the cohort without it from that date, noting that the population shrank.
- You need to free prompt slots. The Coverage tab on Prompts recommends prompts you could pause while what they see regularly is still covered (see Prompt coverage). Check whether any cohort prompt is in the selection before you confirm a pause. Coverage looks at the sources and brands your prompts see, not at your comparison, so it will not protect a cohort for you.
Common mistakes and what a cohort cannot tell you
- Building the cohort after seeing the results. Choose prompts that already look good and the trend will look good. Register before you look.
- Holding the prompts constant but not the engine, channel or location. A change in the engine list, or moving from an API result to an app result, alters the population as surely as a new prompt does. Engines can also stop or start running without any action of yours (a missing key on the platform, or a plan change), so an unfiltered, pooled comparison can include an engine in one month only. Filter to the engines in your register.
- Ignoring missing days. A cohort that lost a week of collection has a smaller denominator that month. See What missing collection days do to an AI visibility trend.
- Comparing counts across months of different length. Compare rates, and keep numerator and denominator visible. 372 answers in a 31-day month and 360 in a 30-day month are not a change in the questions.
- Reading a trend as a cause. The cohort tells you the answer share on a fixed set of questions moved, not why the engines said what they said, or that a page change caused the shift.
Frequently asked questions
How many prompts should a cohort contain?
There is no standard number. Aim for enough distinct questions that each engine you report on produces at least 30 analysed answers per comparison period, since below that the metrics reference labels the rate provisional. Beyond that, size it by how much you can keep stable.
Can I add a prompt to the cohort later?
Not to the same cohort. Create a new cohort that includes the extra prompt and run it alongside the old one for at least one period. Adding to an existing cohort mixes two populations and breaks the line you were protecting.
Does DiscoveredBy have a built-in cohort feature?
The DiscoveredBy docs describe tags, the Explorer's Tag filter and window controls, and the exports, but no dedicated cohort feature. The approach here builds a cohort from those parts and keeps the register outside the product. Check the current docs, since features change.
What if I need to delete a prompt in the cohort?
Pause it instead. The Explorer still counts paused prompts and targets that have answers in the window, whereas a deleted prompt's answers are gone, taking history you may need.
Should I still report the live set?
Yes. The live set shows how you are doing across everything you now care about. Report it next to the cohort, with each one's prompt count, so a reader can see why the two lines differ.
Does this work for share of voice and citation rate as well?
The same logic applies to any rate, because each has a denominator that depends on which prompts ran. Share of voice also depends on which competitors you track, so hold that list constant in the register too.
Next step: freeze your first cohort
Pick twelve or so prompts you would keep for a year, create a dated tag, attach it, download the Prompts export the same day, and fill in the register. Then run the Explorer with the tag filter for last month and this month, and record both in the log.
If you already track prompts in DiscoveredBy, you can do this from your project: sign in and tag your cohort.
- baseline
- ai visibility
- reporting
- prompts
- prompt cohort