Does adding “for a small team” change the question you are measuring? A prompt-edit decision tree
Adding “for a small team” to a tracked prompt can quietly turn it into a different question. A decision tree separates harmless rewording from a new constraint or audience.
On this page
- In short
- What counts as a wording change, and what is a new question?
- Why does “for a small team” matter to an engine?
- How does wording differ from adding an audience?
- The decision tree: keep the trend or start a new one?
- How do you test whether the phrase changed the answers?
- The variant log
- Worked example: Quillstone tries “for a small team”
- What this method cannot tell you
- Common mistakes
- Frequently asked questions
- Next step
Adding “for a small team” to a tracked prompt usually changes the question you are measuring, not just its wording. The phrase adds a constraint: the engine may now weigh price, setup effort or team size in its answer, and the set of brands it names can shift. Rewording that keeps every constraint (“best tool for X” versus “what is the best tool for X”) is a different matter. So the safe habit is to classify each edit first, then either keep the trend line or start a new one. This post gives you a decision tree and a variant log to do that.
In short
- A wording change and a new constraint are different edits. Only the first can safely share a trend line with the original prompt.
- “For a small team” names a buyer condition. Treat it as a new question unless you have evidence that answers to both versions are equivalent for your purposes.
- A persona is a separate mechanism from wording: it adds an audience block to the message on some engines, and does not run on others. Prompt wording, by contrast, reaches every engine the prompt runs on.
- Track a variant beside the original rather than replacing it, then compare over several days, not from one answer each.
- Keep conclusions modest: two answers that differ do not show the phrase caused the difference, and two answers that match do not show they are interchangeable.
What counts as a wording change, and what is a new question?
A wording change keeps the buyer's need, constraints and scope the same and alters only how it is phrased. A new question adds, removes or narrows a condition that an answer could reasonably respond to.
Most prompt edits sit somewhere between. It helps to sort the words you are adding into four kinds:
| Kind of edit | Example | Usually treated as |
|---|---|---|
| Pure phrasing | “best contract review software” to “what is the best contract review software?” | Same question, pending a check |
| Scope change | “for legal teams” to “for in-house legal teams” | New question |
| Buyer constraint | adding “for a small team”, “on a budget”, “with SOC 2” | New question |
| Who is asking | “as a procurement lead” or “I am a solo founder” | New question or audience, depending on how you track it |
The trouble is that the first row can quietly become the second. Adding “small” feels like polish, but a buyer asking for a small-team option is asking for something an answer can honour or ignore. The engine's answer is where you find out, and only across several runs.
DiscoveredBy defines a prompt as a full question, stored as its own row, that is run against the AI engines you track. That framing is useful here: each row is a measurement of one question, so a row whose meaning drifts stops being a clean measurement. For a plain-language definition see the prompt glossary entry.
Why does “for a small team” matter to an engine?
An engine answers the whole message, and a stated team size is information it can use. It may recommend different products, add caveats about pricing tiers, or explain trade-offs for larger organisations. It may also do almost nothing with the phrase. You cannot tell which from the wording alone.
What you can observe is the answer: which brands were named, how they were framed, which reasons were given and which sources were credited. Those are observations of one answer. They do not reveal why the model produced it, so a difference between the original and the variant is a lead to examine, not proof the phrase caused it.
There is also ordinary run-to-run variation. Asking the same question twice can return different answers, so one answer per version tells you very little. Compare a handful of runs of each before you say anything about the phrase.
How does wording differ from adding an audience?
Wording lives inside the prompt text and travels to every engine. An audience (a persona) is a separate setting that DiscoveredBy applies to a prompt you already track, and it reaches only some engines.
Your personas are named buyer profiles, a name and a short description, that a prompt can run as alongside or instead of the plain “General” audience. When a persona variant runs on a chat engine, the message sent gets an audience block after the prompt text, naming the persona and asking the engine to answer for that person. The persona docs are explicit that this is one block of text in a single message: not a logged-in consumer session, not a saved profile the engine remembers, and not a separate account.
The two mechanisms differ in reach, which matters when you choose between them:
| Add “for a small team” to the text | Add a “small team” persona | |
|---|---|---|
| Where it lives | In the prompt itself | An audience block added to the message |
| Engines that receive it | Every engine the prompt runs on | Chat engines only |
| Google AI Overviews, Google AI Mode, ChatGPT (app), Gemini (app) | Receive the prompt text, including your phrase | Persona variant does not run there |
| Prompt slots used | A new prompt uses one slot per country, audience and language it runs in | One extra target, so one extra slot, for each country and language combination |
| Original prompt kept clean | Only if you add a new prompt | Yes, as long as General stays selected beside the persona |
The engines docs say the two Google engines and the two app engines receive the prompt text only, with no audience block, so wording is the only way to put “small team” in front of them. That is a real reason to use wording. It is also a reason for care: on those engines the phrase becomes part of the question itself, with nothing separating it from the original.
Using a persona variant does not turn the result into evidence about actual people. For the workflow of designing and interpreting persona tracking, see Measure AI visibility for different buyer personas without mixing the results. This post covers a narrower question: what to do when the change is a few words in the prompt.
The decision tree: keep the trend or start a new one?
Work through the tree for every edit before you save it. It ends in one of three actions: add a sibling prompt and keep both, add a sibling and retire the old wording after a check, or use a persona.
PROMPT-EDIT DECISION TREE
Start: I want to change the wording of a tracked prompt.
1. Does the edit add, remove or narrow any condition a buyer could
care about (team size, budget, region, industry, compliance,
integration, urgency)?
- YES -> go to 3.
- NO -> go to 2.
2. Is the edit only phrasing (word order, question form, filler,
spelling, punctuation)?
- YES -> Treat it as the same question, but do not assume.
Add the new wording as a sibling for a few days and
compare (step 5). If the answers look alike, retire
the old wording deliberately and note the date.
- NO -> go to 3.
3. Does the edit describe the PERSON asking ("as a procurement lead",
"I am a solo founder") rather than a REQUIREMENT the answer should
meet ("for a small team", "on a budget", "with SOC 2")?
- YES -> go to 4.
- NO -> NEW QUESTION. Add it as a sibling prompt. Keep the
original running. Never merge the two trend lines.
(If you really mean "our buyers are a small team",
treat it as describing the asker and go to 4.)
4. Do you need this on Google AI Overviews, Google AI Mode,
ChatGPT (app) or Gemini (app)?
- YES -> Wording is the only route. Add a sibling prompt with
the phrase in the text and label it as an audience-
constrained question.
- NO -> Prefer a persona on the existing prompt, so General
stays clean and the variant is reported by audience.
5. Compare sibling answers over several days, not one answer:
- Same brands named, same framing, same kinds of sources?
-> Looks equivalent for this engine and period.
- Different brands, framing, reasons or sources?
-> Different question in practice. Report separately.
- Mixed?
-> Report separately and say the evidence is mixed.
6. Write the decision in the variant log (below).
How do you test whether the phrase changed the answers?
Run the original and the variant side by side, as separate prompts, on the same engines, over the same days. Then compare the same four things in both: which brands were named, how they were framed, which reasons were stated, and which sources were credited.
Here is a plain way to run the comparison:
- Add the variant as a new prompt. Do not overwrite the original; keep it running so both accumulate history over the same period. Each country, audience and language combination of a prompt uses one prompt slot, so budget for that, and see Plans and limits for what your plan includes.
- Track both on the same engines and in the same locations. If the original runs in three countries, run the variant in the same three, or your comparison will confound wording with place.
- Let both run for several scheduled runs. The docs describe a single scheduled job that runs each active prompt target once a day per engine, so a week gives you about seven answers per engine per prompt, if nothing failed.
- Open the answers for each version (the answer drawer shows one answer in full) or use an export, and answer the four questions above.
- Decide, using the decision tree. Write the decision down.
Give the two versions the same tag so they are easy to find together. The filter bar narrows results by tag, engine, country, persona and language, and the Explorer can break results down by prompt.
If you are worried about wearing out your prompt allowance, the answer is to be selective. A short comparison of a few prompts is worth more than adding “for a small team” variants to your whole list. Related reading: How to choose AI tracking prompts that reflect buying decisions and Audit your prompt list before adding more prompts.
The variant log
Copy this table into your tracker and add a row for every wording experiment.
| Field | What to write |
|---|---|
| Date started | The day the variant began running |
| Original prompt | Exact text, unchanged |
| Variant prompt | Exact text, including the added phrase |
| Edit type | Phrasing / scope / buyer constraint / who is asking |
| Engines and locations | Must match the original |
| Runs compared | Number of answers per version per engine |
| Brands named (own) | Count of answers naming your brand, each version |
| Brands named (others) | Notable differences in competitors named |
| Framing and stated reasons | What changed, quoted briefly |
| Sources credited | New or dropped source types or domains |
| Decision | Same question / different question / mixed |
| Action taken | Retired old wording / kept both / used persona |
| Caveat | One sentence on what this cannot show |
Worked example: Quillstone tries “for a small team”
Illustrative example: Quillstone and its competitors are fictional, and the numbers are made up to show the method.
Quillstone sells document-review software to mid-sized legal and compliance teams. Its content lead tracks the prompt “What is the best document review software for legal teams?” on one chat engine and wonders whether to add “for a small legal team”, because the sales team hears that phrase often.
Applying the tree: the phrase adds a buyer condition (step 1 yes), and it states a requirement the answer should meet rather than describing the asker (step 3 no), so it is a new question. Quillstone adds it as a sibling and keeps the original. They also add a pure-phrasing sibling, “best document review software for legal teams”, to check whether question form matters.
After five daily runs on the same engine and country, the counts look like this:
| Prompt version | Answers naming Quillstone | Answers naming Brieflane | Answers naming Clausewise |
|---|---|---|---|
| Original: “What is the best document review software for legal teams?” | 3 of 5 | 4 of 5 | 2 of 5 |
| Phrasing only: “best document review software for legal teams” | 3 of 5 | 4 of 5 | 2 of 5 |
| Added phrase: “What is the best document review software for a small legal team?” | 1 of 5 | 5 of 5 | 3 of 5 |
Quillstone also reads the stated reasons. In some of the four small-team answers that omit Quillstone, the answer cites lower setup effort as a reason to prefer another vendor. That is what the answer said, not why the model produced it.
What the log records: the phrasing-only sibling matched the original on all three brands across five runs, so Quillstone treats it as the same question for now and keeps watching. The small-team version differs on all three counts, so it becomes its own trend line. The difference (Quillstone 3 of 5 versus 1 of 5) is a lead. With five answers each, it does not establish that the phrase caused it, and Quillstone does not conclude that its product is worse for small teams. It writes a follow-up: review the stated reasons in those answers and check which sources were credited.
What this method cannot tell you
A match between two versions proves less than it seems, and a difference proves less than it feels like.
- Equal counts are not equal answers. Both versions can name the same brands and still frame them differently. Read the framing, not only the counts.
- A few runs are a few runs. Answer variation means five runs can look different by chance. Use more runs before acting on a small gap, and avoid quoting a percentage from five answers.
- The phrase is not the only thing that changed. Different days, different engines or a changed location will each muddy a comparison. Hold them constant.
- Engines differ. Some engines will honour a team-size constraint and others may not. A result on one engine does not transfer to another. For how each engine is collected, see the engines reference.
- A persona is a simulation. It adds one block of text to a message. It says nothing about how real small-team buyers are answered when signed in on their own devices.
- Length limits apply. Long prompts have limits on some engines: the docs say a prompt over 700 characters (once encoded) does not run on the two Google engines, and one over 2,000 characters does not run on ChatGPT (app) or Gemini (app). Extra clauses are not free.
Common mistakes
- Editing the original prompt in place and continuing the same chart. Your before-and-after line now mixes two questions, and you cannot separate them afterwards. See also Use a fixed prompt cohort for month-to-month comparisons.
- Treating “for a small team” as a persona. It is a constraint on the question unless you deliberately track it as an audience, and the two are reported differently.
- Adding the phrase to every prompt at once. You lose the original baseline and multiply slot usage.
- Judging from one answer per version.
- Reporting “AI prefers competitor X for small teams” from a handful of answers. Say what was observed, on which engine, over which days.
Frequently asked questions
Is adding “for a small team” the same as tracking a small-team persona?
No. Wording changes the text of the prompt, which reaches every engine the prompt runs on. A persona adds an audience block to the message and runs only on chat engines, not on the two Google engines or the two app engines. Use whichever matches the engines you need, and report them separately.
Can I just edit my existing prompt instead of adding a new one?
In most tools you can, but you should not if the edit adds a constraint: the trend line for that prompt would then mix two questions. Add the variant as a sibling and keep the original running, so each has a clean history.
How many runs do I need before I decide?
The docs do not set a number, and there is no honest universal threshold. Use several scheduled runs of each version on the same engine and location, and treat a small difference as a lead. If a decision depends on it, gather more runs first.
Does a difference between versions mean the engine ranks me lower for small teams?
Not necessarily. You are comparing two answers, and stated reasons describe what an answer said, not the internal cause of it. Read the framing and sources in the answers that differ, and test a change to your content rather than assuming the cause.
What if I only want to fix a typo or reorder words?
Treat it as a phrasing change, run a short sibling comparison to confirm, and then retire the old wording deliberately with a note of the date. Record it in your variant log so anyone reading the trend later knows where the wording changed.
Next step
Pick one prompt on your list where the buyer's team size, budget or industry keeps coming up in sales calls. Add the constrained version as a sibling, track it beside the original for a week, and fill in one row of the variant log. Countries, audiences and languages are all set per prompt, as described in the Prompts docs, and the AI visibility features page gives the overview. To start tracking a paired prompt, sign in or create an account.
- measurement
- tracked prompts
- prompt wording
- personas