Track AI visibility across languages without losing comparability

How to track AI visibility in several languages: write each prompt natively, review it, keep the setup fixed, and report every language as its own population.

Kamal 15 min read
A row of matching paper cards with different abstract scripts laid out side by side, each with a small coral tab, showing one question asked in several languages.
On this page
  1. In short
  2. What does "comparable" mean across languages?
  3. What does a language variant actually do?
  4. How do you write the same question in another language?
  5. How many prompt slots does a language use?
  6. Which engines run which languages?
  7. The multilingual prompt checklist
  8. How do you report languages without mixing them?
  9. A worked example
  10. Common mistakes and limits
  11. Frequently asked questions
  12. Next step

To track AI visibility across languages and keep the results comparable, write each prompt natively in its own language, have a fluent person review it for meaning, keep the same question intent in every language, and report each language as its own population with its own denominator. Do not machine-translate your prompts, do not assume English findings carry over, and do not pool languages into one score. A gap between two languages only means something if the prompts, engines, countries and dates behind both numbers were the same kind of thing.

In short

  • A language variant tells the engine which language to answer in, by an added instruction or a request setting. It does not translate your prompt; you write and review every wording yourself.
  • Comparable means same intent, same engines, same period and a stated denominator. Identical wording is not possible across languages, so intent is what you hold fixed.
  • Which engines run a language, and how the language reaches them, differs by engine. Report the population behind every number.
  • Do not assume English results describe other languages. Measure each one on its own.
  • Use the checklist and the comparability log below before you add a second language.

What does "comparable" mean across languages?

Two language results are comparable when the same buying question was asked, on the same engines, over the same dates, and each number states how many answers it is built from. The wording cannot be identical, because a natural French question is not a word-for-word copy of the English one. So you hold the intent fixed and let the wording vary.

A few terms, defined once:

  • Prompt: the question you track, sent to AI engines on a schedule.
  • Language variant: a tracked prompt paired with one language. The prompt text is sent unchanged, and the engine is told which language to answer in, either by an added instruction or by a setting on the request.
  • As written: the default value for a target that carries no language setting. The prompt goes out as typed.
  • Population: the set of answers a number is calculated from. In the metrics reference, brand visibility divides by analysed answers, so "brand named in 6 of 24 analysed German answers" describes a population of 24.
  • Intent: the decision the asker is trying to make, independent of the words used to ask.

The rest of this post is about protecting intent and population from the ways they quietly drift.

What does a language variant actually do?

A language variant sends your exact prompt text and adds an instruction to answer in that language; it never translates anything. According to the languages and templates documentation, if you want to track what a French-speaking buyer would type, you write the prompt in French yourself and set its language to French.

The mechanism differs by engine, and that matters for comparability:

Engine group How the language is applied
Gemini (API), Perplexity, Claude, Grok A response-language instruction is appended to the message ("Write the whole answer in French (fr)")
Google AI Overviews, Google AI Mode The language goes as a setting of the search request; the prompt text gets no added instruction
ChatGPT (app), Gemini (app) The language goes as a setting of the request; the prompt is sent as written

So the same "French" label covers two different things: an instruction the model reads, and a request setting that applies to the search or data request. Neither is a bug, but a difference between engines is not necessarily a difference in how each model treats French. The engines reference explains how each engine is collected.

The same language label reaches different engines in different ways.

The languages are 48 fixed two-letter codes, and a code never carries a region, so there is no separate variant for Brazilian versus European Portuguese. Country already carries the region a prompt runs in.

How do you write the same question in another language?

Start from the intent, not the English sentence. Write down what the buyer is trying to decide, then have someone who searches in that language phrase it the way a buyer would, and have a second fluent reader confirm the meaning matches.

A workable process:

  1. Write the intent in one plain sentence, in your working language.
  2. Ask a fluent speaker to write the prompt as a buyer in that market would type it, not as a translation.
  3. Ask a second fluent reader to read the new prompt cold and say what decision it implies. Compare that against your intent sentence.
  4. Check terminology: the product category, job titles and regulatory terms a local buyer actually uses. A literal translation of your English category name may not be the local term.
  5. Keep the prompt short. The Google engines skip a prompt longer than 700 characters once encoded, and ChatGPT (app) and Gemini (app) skip one longer than 2,000. A skipped variant produces no answers on that engine.
  6. Record the final wording, the reviewer and the date in your log, so a later wording change is a deliberate, dated decision.

If you are still deciding which questions are worth tracking at all, start with How to choose AI tracking prompts that reflect buying decisions before multiplying the list by languages.

How many prompt slots does a language use?

Each combination of location, audience and language is a separate target, and each target uses a prompt slot, so adding a language to a prompt multiplies its targets. A prompt tracked in 2 countries as General, in As written plus French, becomes 2 x 1 x 2 = 4 targets. The docs give this example, and a single prompt tops out at 100 active variants.

The cross-product is also a trap. If you tick Germany, France, German and French on one prompt, you get German in France and French in Germany, which may be meaningless for your buyers. For a multilingual program, it is usually cleaner to make one prompt per language and market, each with its own native wording, and keep the grid small. A variant template saves a combination of countries, languages and engines so you can reuse or bulk-apply it, which helps you keep setups identical.

Language variants use the same prompt-slot budget as countries and personas, and languages, templates and engine choice need no separate entitlement. Read plans and limits for what your plan includes. If you also split by audience, see Measure AI visibility for different buyer personas without mixing the results, because the same population discipline applies.

Demo data. The preview shows the variant count and slots before you save.

Which engines run which languages?

Every chat engine a prompt runs on runs every language variant, while the Google engines and the two app engines skip some. The exact rules come from the engine docs, and they change what your population looks like:

  • Chinese is the one language the Google AI Overviews, Google AI Mode, ChatGPT (app) and Gemini (app) engines do not run. A Chinese variant is left out rather than sent in English, and prompt detail lists it as not supported. It still runs on the chat engines.
  • An As written target whose project default language is Chinese runs on those four engines in English.
  • For Portuguese and Tagalog, the Google AI Mode and Gemini (app) requests are sent as Brazilian Portuguese and Filipino.
  • The language actually sent is shown on each run of those four engines in prompt detail, so you can check it.

Two consequences follow. First, a Chinese result is built from the chat engines only, while a French one can also include the four engines above, so the two are different populations even before you look at any number. Second, an "As written" target is not necessarily English. On those four engines, an As written target uses the project's default language, or English if none is set, so check that setting before you call a baseline "English".

Nothing detects a project's default language from your website. You set it yourself in project settings, and changing it never rewrites an existing target.

The multilingual prompt checklist

Copy this before adding a language. Every line should be answerable in writing.

MULTILINGUAL PROMPT CHECKLIST

Intent
[ ] The buyer decision is written as one plain sentence
[ ] Each language prompt maps to that sentence, not to the English wording
[ ] The same decision stage is covered in every language

Wording
[ ] Written natively by a fluent speaker, not machine-translated
[ ] Reviewed cold by a second fluent reader
[ ] Local category and job terms confirmed
[ ] Prompt is short enough for every engine that should run it
[ ] Final wording, reviewer and date recorded

Setup
[ ] Language set on the prompt (not left as written by accident)
[ ] Project default language checked
[ ] Country matches the market, region carried by country not language
[ ] Engines chosen deliberately and the same across languages (except where an engine skips a language)
[ ] Slot count checked; no accidental cross-product targets

Comparability
[ ] Same competitors and brand names tracked in every language
[ ] Local spellings and transliterations of brand names considered
[ ] Same period compared
[ ] Every number reported with its answer count
[ ] Engines that skipped a language noted next to the result

Interpretation
[ ] No English result is used to describe another language
[ ] Cross-language gaps described as observations, not causes

How do you report languages without mixing them?

Report each language as its own row with its own denominator, and only compare rows that share engines and dates. The product supports this directly: the language filter narrows the screens that take it to As written or one language, Explorer can break down or filter any of its metrics by language, and the Answers and Brand mentions exports carry a language column that is empty for As written.

Two details protect you from misreading:

  • A language breakdown in Explorer keeps every chat engine a prompt runs on, but a Chinese row holds no answers from Google AI Overviews, Google AI Mode, ChatGPT (app) or Gemini (app).
  • Sentiment trends and brand reasons follow the language chip, but there is no separate "by language" view for them, so filter one language at a time.

Brand visibility and the other rates are defined in the metrics reference. Use the same definition in every language.

The comparability log

Keep one row per language and engine group per period. This is the reporting deliverable.

Language Country Prompts Engines that ran Engines that skipped (why) Analysed answers Answers naming the brand Notes on wording changes

Add a header line above the table naming the period, the project default language and the date of the last wording review.

A worked example

Illustrative example: Quillstone and its competitors are fictional, and the numbers are made up to show the method.

Quillstone sells document-review software to legal and compliance teams. It is entering Germany and France and wants to know whether it shows up in local-language answers the way it does in English. Its competitors are Brieflane and Clausewise.

The team writes one intent sentence: "A compliance lead is choosing a document-review tool for contract checks." Then it has each market's prompts written natively and reviewed by a second reader. Each prompt targets one country in one language, so each uses one slot:

Prompt set Country Language Prompts Slots used
English wording United States English 5 5
German wording Germany German 3 3
French wording France French 2 2

That is 10 prompts and 10 slots. The team deliberately avoids ticking both countries and both languages on one prompt, which would have made 2 x 2 = 4 targets per prompt, including German in France.

Hold the intent fixed, write each prompt natively, report each language separately.

After one collection round across the same dates, the log looks like this (made-up counts, one analysed answer per prompt per engine):

Language Prompts Engines that ran Analysed answers Answers naming the brand Rate
English 5 8 40 12 30%
German 3 8 24 6 25%
French 2 8 16 5 31%

The arithmetic: 5 x 8 = 40, 3 x 8 = 24 and 2 x 8 = 16 answers; 12/40 = 30%, 6/24 = 25%, 5/16 = 31.25%, shown as 31%. The answer counts differ only because the prompt sets differ in size.

What the team can say: French and English look similar and German is lower, but the German and French populations are small, so a difference of a few answers moves the rate a lot (one more German answer naming Quillstone would make it 7 of 24, or 29%). What it cannot say: that German engines "prefer" competitors, or that the English rate predicts anything about German. The next step is to read the German answers themselves, see who was named and why, and confirm the German wording is what a buyer would type. If the German prompts were reworded mid-period, the log's wording column is where that shows up.

Common mistakes and limits

  • Treating a language variant as a translation. It sends your text unchanged and asks for an answer in that language. Prompt discovery and research also suggest prompt text the same way regardless of language, so review any suggestion before you tag it with a language.
  • Expecting import to set languages. CSV and Excel import carry no language column; a country new to the prompt lands in the project's current default language.
  • Comparing rates from different populations. A 25% from 24 answers and a 30% from 40 answers are not equally well supported.
  • Ignoring skipped engines. A Chinese variant runs on fewer engines than a French one. Say so next to the number.
  • Changing wording and comparing before and after as if nothing changed. Log wording changes with a date.
  • Assuming English results generalise. Measure each language on its own; an English result does not tell you about another language.
  • Forgetting brand-name spelling. Your brand may appear in a transliterated or localised form in some scripts. Decide up front how you will check for it.

What this cannot tell you:

  • Answer analysis (mentions, sentiment and reasons) runs on non-English answers with the same models used for English ones, and the docs state that accuracy per language has not been separately measured. Spot-check a sample of answers in each language by reading them.
  • Chat-engine answers are collected through APIs, not the consumer apps, so an app's interface-language setting is not reproduced for them.
  • A stated reason or a mention is an observation of an answer, not proof of why an engine produced it.
  • Engine choice is per prompt, not per variant, so one prompt cannot run its French targets on one engine set and its As written targets on another.

For a study-style version of the question, see Does the same buying question get different answers in different languages?. If your markets differ by city as well as language, compare city and country results responsibly.

Frequently asked questions

Does DiscoveredBy translate my prompts?

No. The prompt text is sent exactly as you wrote it, with an added instruction or a request setting that names the language. To track what a German buyer types, write the prompt in German and set its language to German.

Can I compare my English visibility score to my French one?

You can put them side by side, but only as separate populations. State the answer count behind each, confirm the same engines ran, and treat a difference as something to investigate in the answers, not as a finding on its own.

How many languages can a project track?

The product supports 48 language codes. A single prompt can name 1 to 10 languages, and a prompt is capped at 100 active variants across locations, audiences and languages. Your prompt-slot budget, which depends on your plan, is the practical limit.

What happens if a language is not supported on an engine?

That variant is left out on that engine rather than sent in a changed form, and prompt detail says why. Chinese is the documented case for the Google and app engines.

Should I use one prompt with many languages or one prompt per language?

For a multilingual program, one prompt per language and market is usually cleaner, because each gets natively written wording and no accidental combinations. A single prompt with several languages sends the same text in each, which only makes sense if the text is meaningful in all of them.

Can I see which language an answer was collected in?

Yes. The Overview answer drawer tags an answer with its language, and the Answers and Brand mentions exports carry a language column that is empty for As written.

Next step

Pick one market, write its prompts natively, get them reviewed, set the language on each, and log the first period before you add a second language. Set up your first language variants in your DiscoveredBy project, then read how language variants work for the full rules.

  • prompt tracking
  • ai visibility
  • multilingual
  • language variants
  • international

Share

Summarize with AI

Start monitoring your AI visibility.

See how AI search engines talk about your brand.

Free to start. No credit card required.