Why your manual ChatGPT check differs from a monitoring result: a discrepancy log

You asked ChatGPT and saw your brand. Your monitoring report says otherwise. Record seven variables, then compare like with like using this discrepancy log.

Kamal 14 min read
Two hands holding two slightly different copies of the same printed page, set side by side on a desk, with a single coral highlight on the differing line.
On this page
  1. In short
  2. Why do two ChatGPT checks give different answers?
  3. What are the seven variables that separate a manual check from a monitoring result?
  4. How do you compare a manual check fairly?
  5. The discrepancy log
  6. A worked example
  7. What a discrepancy log cannot tell you
  8. Common mistakes
  9. Frequently asked questions
  10. Next step

Your manual ChatGPT check and a monitoring result differ because they are usually two different observations, not one observation measured twice. You are probably signed in, on your own device, in your own location, with a question you worded on the spot, at one moment. A monitoring tool asks a fixed question on a schedule, in a defined country and language, and stores the answer it got. Before you decide which one is wrong, write down the conditions of both. This post gives you a discrepancy log to do that.

In short

  • A single manual check is one sample under conditions nobody recorded; a monitoring result is one sample under conditions that are recorded. Both are observations, and neither is "the truth".
  • Seven variables are worth logging whenever the two disagree: session, wording, timing, location, language, collection method and what counted as the answer.
  • Nobody can reproduce your personal, signed-in session from outside it. Compare the tool's result with a check that matches its conditions instead.
  • Log the differences first. Only then decide whether the gap is expected, worth a wording change, or a real problem to investigate.

Why do two ChatGPT checks give different answers?

Two checks disagree when any input to the answer differs: who is asking, what exactly was asked, when, from where, in which language, and through which channel the answer was collected. Each of those can change what an answer says, and a hand-typed check usually changes several at once without you noticing.

It helps to name the pieces. A manual check is you typing a question into ChatGPT yourself and reading the result. A monitoring result is an answer that a tool collected on a schedule for a saved question and stored with its date, country and collection method. A prompt, in DiscoveredBy's vocabulary, is a full question you chose to track because your buyers ask something like it; it is not a keyword (see Key terms).

In DiscoveredBy, the ChatGPT engine is called ChatGPT (app). The docs describe it as what a person who is not signed in sees on chatgpt.com. DataForSEO, a third-party data provider, puts the prompt to chatgpt.com in a session that is not signed in and returns the answer the page shows (How ChatGPT (app) is collected). That is a specific, documented setup, and it is the setup your manual check has to match if you want a fair comparison.

The same question, asked under two sets of conditions.

What are the seven variables that separate a manual check from a monitoring result?

The seven variables are session, wording, timing, location, language, collection method and what you counted as the answer. Each one is a place where your two observations can quietly diverge, so each gets a line in the log.

1. Session: who was asking

Your manual check happens in whatever session you are in. If you are signed in, your account and its settings are part of the situation, and you cannot see all of them from the outside. The monitoring result for ChatGPT (app) comes from a session that is not signed in, so it carries no account of yours. That difference is not a defect in either one; it means they answer slightly different questions: "what does ChatGPT say to someone using it in your setup" versus "what does chatgpt.com show a visitor who is not signed in".

Do not try to fix this by guessing what your account contributes. Record whether you were signed in, and treat a signed-out check as the closer match.

2. Wording: the exact text sent

People rarely retype a question identically. "Best document review software for legal teams" and "what's the best tool for reviewing contracts at a law firm" are different questions, and the difference can change which brands appear. DiscoveredBy sends a tracked prompt as written, on one line: runs of spaces and line breaks become one space (What reaches ChatGPT). For the ChatGPT (app) engine it adds nothing else to the text, so no audience block and no search-context block.

The practical rule: copy the exact prompt text from the prompt in your monitoring tool into your manual check, character for character. See also Does adding "for a small team" change the question you are measuring?

3. Timing: when the answer was produced

A monitoring tool runs on a schedule. In DiscoveredBy, a single scheduled job once a day enqueues one run per active prompt target, on every engine currently available on your plan that the target runs on (Prompts). Accepting a prompt from Suggestions queues a fast first run outside that cycle, but otherwise the daily job is the pattern. Your manual check is a different moment, often a different day.

Timing matters for two reasons. The answer you see today may not be the answer stored for yesterday. And one answer on one day is a single draw. How much answers move from day to day is an open question you can study for yourself; see How stable are AI recommendations from one day to the next?

4. Location: where the question came from

A manual check inherits your real location. A tracked prompt has a target location, which is a country or a city. For ChatGPT (app), DiscoveredBy sends the country only, as a setting of the request to DataForSEO. There is no city setting, so a city target does not run on ChatGPT (app) at all (How location reaches each engine).

The docs are direct about the limit: a target location is information passed on, and each engine decides how much it weighs it. It "can differ from what a real person there sees, signed in, on their own device" (A target location is not a customer's location). The docs also record one test in which a "near me" question for India was answered for Siliguri, so for local questions in particular, a manual check from a given city may not match.

5. Language: the setting versus the words

The language sent for a tracked prompt is a setting on the request, not part of the text. It is the variant's language; for an "As written" target, your project's default language; otherwise English. Chinese variants do not run on ChatGPT (app), and an "As written" target in a project whose default is Chinese is sent as English. Each run in the history says which language was sent, shown as "Language sent to ChatGPT (app)".

If your manual check was in a different language from the one sent, that can change the answer. Check the run history before comparing.

6. Collection method: the channel

Channel is the variable people forget. In DiscoveredBy's records, every answer carries its own platform, surface and collection method, stamped when it runs and never looked up afterwards (How an answer is labelled). ChatGPT (app) carries the method "licensed", shown as "Licensed data". The four chat engines (Gemini (API), Perplexity, Claude and Grok) are queried through their company's API instead.

The label tells you what kind of observation you have. A licensed-data capture of the page is not an API call to a model, and neither is your own browser session. Keep app results and API results separate in any comparison (see Keep app results and API results separate in your reports).

7. What counted as the answer

Finally, check what was stored. For ChatGPT (app), the stored answer is the text the page shows, product lists and local business lists included. Citations are the answer's inline sources in the order they first appear; a page cited twice is one citation, and links to a local business's own site and product or retailer links are not citations (What counts as the ChatGPT (app) answer).

Two more details bite in practice. ChatGPT decides whether to search, and DiscoveredBy never asks it to, so some answers cite no sources. And a small source link whose text is only a site address, such as a bare domain, is blanked out before brand mentions are extracted when it directly follows the passage it supports. A site named only in such links is not counted as a mention, although the stored answer keeps them. If you counted a brand you saw in a source link, you and the tool may have counted different things.

How do you compare a manual check fairly?

To compare fairly, change your check to match the tool's conditions, not the other way round. You cannot rebuild the tool's collection method by hand, so the goal is to remove every avoidable difference and record the ones left over.

Do these in order:

  1. Open the prompt in your monitoring tool and copy its exact text, country and language.
  2. Open a fresh, signed-out browser session, and note that you did.
  3. Paste the text unchanged, ask once, and save the full answer with the date and time.
  4. Find the tool's stored answer for the nearest date and read it against yours.
  5. Fill in the log below, one row per variable, before writing any conclusion.

A caution: a signed-out check from your own machine is still not the same request path as the one the data provider uses, and it still comes from your location. It removes some differences; it does not make the two identical. Say "closer match", not "reproduction".

The discrepancy log

Copy this into a spreadsheet or document. Fill in the "Manual check" and "Monitoring result" columns from what you can see, then write the verdict last.

DISCREPANCY LOG: manual check vs monitoring result

PROMPT (paste the tracked text exactly):
ENGINE: ChatGPT (app)      TRACKED PROMPT TARGET (country / city / persona / language):

VARIABLE            | MANUAL CHECK                | MONITORING RESULT            | SAME? | NOTE
--------------------+-----------------------------+------------------------------+-------+-----------------
Session             | signed in / signed out?     | not signed in (per docs)     |       |
Wording             | exact text typed            | exact tracked text           |       |
Timing              | date + time                 | date of stored answer        |       |
Location            | real location / VPN?        | country sent (no city)       |       |
Language            | language typed              | "Language sent to ChatGPT"   |       |
Collection method   | my browser, by hand         | Licensed data                |       |
What counted        | what I decided to count     | stored text + inline sources |       |

OUTCOMES COMPARED
Brand named?            manual: yes / no        monitoring: yes / no
Order of brands named:  manual:                 monitoring:
Sources cited:          manual:                 monitoring:
Answer cited sources at all?   manual: yes / no        monitoring: yes / no

DIFFERENCES I CANNOT CONTROL (list them; do not delete them):

VERDICT (pick one)
[ ] Expected: the variables differ enough to explain the gap
[ ] Re-check: repeat with matched wording, language and signed-out session, then re-log
[ ] Investigate: variables match and the outcomes still differ; look at more days
[ ] Fix the tracked prompt: the tool's wording does not match how buyers ask

NEXT ACTION AND OWNER:

When the verdict is "Investigate", do not draw a conclusion from one more manual check. Gather more days of tracked answers first. In DiscoveredBy, the prompt's run history lists the runs in your window with their status, up to the 50 most recent. Depending on your plan, the Answers export carries the response text itself alongside the prompt as it was sent, its country and how it was collected; it is capped at 5,000 rows, and a response over 8,000 characters is shortened, with response_truncated saying so (Exports and activity).

A worked example

Illustrative example: Quillstone and its competitors are fictional, and the numbers are made up to show the method.

Quillstone sells document-review software to mid-sized legal and compliance teams. Its marketing lead, Priya, asks ChatGPT "what are the best document review tools for legal teams?" while signed in and sees Quillstone listed second, after a rival called Brieflane. Her monitoring dashboard says Quillstone's ChatGPT (app) results are weak. She suspects the tool is broken.

She runs the log. Her real question was worded slightly differently from the tracked prompt ("best document review software for in-house legal teams"), she was signed in, and she was in the UK while the tracked target is the United States. Four variables differ: session, wording, location and timing (her check is a different moment from any stored answer). Language matched, English in both, and the collection method differs by nature, since she used a browser by hand.

She then does the fair comparison: exact tracked wording, signed-out session, one ask, saved. Quillstone does not appear in that answer. She reads the last seven stored answers for the prompt.

Day Quillstone named in stored answer?
1 No
2 Yes
3 No
4 No
5 Yes
6 No
7 Yes

That is 3 of 7 completed answers, about 43%. Her signed-in check was one draw under different conditions; the tracked series shows Quillstone appears in some answers and not others. The verdict for the first gap is "Expected". The more interesting question, why Quillstone appears on some days and not others, gets "Investigate", with more days of answers. Notice what she does not do: she does not claim the tool is wrong, and she does not claim her own check is meaningless. She also does not say why the answers vary, because the log records what was observed, not the model's reasons.

Demo data. The run history shows what each collected answer was sent, which is the column to compare against your own check.

What a discrepancy log cannot tell you

The log explains a gap; it does not prove which observation is "right". Keep these limits in view.

  • It does not reveal why an answer changed. Stated sources and names are observations of one answer, not proof of what a model weighed. Do not write "ChatGPT dropped us because...".
  • It does not tell you what your customers see. A signed-out capture from a data provider is a documented baseline. It is not a sample of your buyers' own sessions.
  • One matched check is still one sample. Repeating the manual check three times on one afternoon tells you about that afternoon.
  • A gap can come from your counting. If you counted a brand seen only in a source link, or a brand inside a product card, you may have counted more than the tool did.
  • A failed or cut-off run is not a miss. A run that could not be read is a failed run, not an answer that named no brand, and failed runs are absent from the metrics, not counted as zeros (A run failed).

Common mistakes

  • Comparing your signed-in result to a signed-out capture and calling it a tie-breaker. Record the session first.
  • Retyping the prompt from memory. Paste it.
  • Reading a city-specific check against a country-level target. ChatGPT (app) is sent the country only.
  • Deciding on one day. Look at the series of stored answers before changing anything on your site.
  • Merging app and API observations. They are different collection channels and belong in separate columns.
  • Treating the tool's number as a manual-check forecast. A visibility figure is a rate over many answers; a single check is one answer. See Metrics defined for the denominators.

Frequently asked questions

Why does ChatGPT give me a different answer each time I ask?

Answers can differ from one asking to the next, and your wording, account and location can add to the difference. That is why a single check should not be treated as a measurement. Repeat under recorded, matched conditions, or read a series of dated answers instead.

Is a monitoring result less real than what I see in ChatGPT?

No. It is a different, documented observation: fixed wording, a defined country and language, a recorded date and a labelled collection method. What it cannot do is stand in for a specific signed-in user's personal session.

Can I make my manual check match the monitoring result exactly?

Not exactly. You can match the wording, language and signed-out session, and you can note your location. You cannot reproduce the data provider's request path or another person's account, so log the residual differences rather than hiding them.

Why is my brand missing from the tool but present when I ask?

Check the log: a reworded question, a signed-in account, a different country or a differently counted mention (for example, a brand named only in a source link) are the usual suspects. If the variables match and the outcomes still differ, gather more days of answers before concluding anything.

Does ChatGPT (app) run for city targets or personas?

No. The docs list both as not run on ChatGPT (app): cities are not sent to ChatGPT's app, and neither are personas. The General, country-level variant of the prompt still runs. If your manual check used a city or an audience in the wording, expect it to differ.

Next step

If you track ChatGPT (app) in DiscoveredBy, open the prompt you are comparing, copy its exact text, and use the run history to find the run, its country and its language before you fill in the log. To see how the engine is collected and what it reports, read Engines and measurement, the ChatGPT engine page and the ChatGPT (app) glossary entry. If you are setting up tracking for the first time, AI visibility explains what gets collected each day.

Sign in or create your account and start with one prompt: compare its stored answer with a matched, signed-out check and log the differences.

  • prompt tracking
  • measurement
  • chatgpt
  • ai monitoring
  • data collection

Share

Summarize with AI

Start monitoring your AI visibility.

See how AI search engines talk about your brand.

Free to start. No credit card required.