Klairia Blog
How to Measure Your Brand's Visibility in ChatGPT, Claude, and Gemini
A practical methodology for measuring how AI assistants describe and recommend your brand: first-mention level, progressive prompting, model and language coverage, and weekly tracking.
Most teams know they should care about how AI assistants represent their brand. Very few know how to measure it without guessing.
The reflex is to open ChatGPT, type “best tool for X”, and see what comes back. If your brand is mentioned, you celebrate. If it is not, you panic. Then you try again two days later with a slightly different prompt, get a different answer, and stop knowing what to think.
That is not measurement. That is anecdote.
If you want a reliable signal — one you can track, compare, and act on — you need a method. This article describes the one we use at Klairia, and that you can apply manually too.
What you are actually trying to measure
AI visibility is not a single number. A brand can be:
- Mentioned directly in the answer body.
- Cited as a source in a footnote or link.
- Included in a shortlist of recommended tools.
- Described accurately versus described vaguely or incorrectly.
- Compared favorably against competitors, or not compared at all.
These are distinct outcomes. Conflating them hides what is actually happening.
The most useful starting metric is what we call first-mention level: at what depth of prompting does the model name your brand on its own, without being asked about you?
1. Why one prompt is never enough
A single prompt only tells you what one user, in one moment, in one phrasing, would see.
Real buyers do not arrive at “best CRM for early-stage SaaS founders in Europe” in one shot. They start broad (“best CRM”), narrow down by use case, then by constraints, then ask for a shortlist, then dig into a specific vendor.
A brand that never appears at the broad level might still appear at level 2 or 3. That distinction matters: surfacing at level 1 means the model considers you a category default; surfacing only at level 4 means the model knows you exist but does not reach for you spontaneously.
This is why we use progressive prompting: a sequence of prompts going from broad category to specific buyer intent, recording at which level your brand first appears.
What to do:
- Define 5 to 10 progressive prompt sequences that match how your buyers actually narrow down.
- Record the first level your brand appears at, in every sequence.
- Distinguish unprompted mentions from mentions that only appear after you name yourself.
A brand mentioned only when explicitly asked about is not visible. It is just known to exist.
2. The metric that matters: first-mention level
Once you have progressive prompts, the cleanest signal is the depth at which your brand first appears.
- Level 1 — appears in broad category prompts (“best [category] tools”). You are a default.
- Level 2 — appears once a use case is specified. You are a contextual recommendation.
- Level 3 — appears with specific buyer constraints. You are a fit-based recommendation.
- Level 4 — appears only after competitors are named. You are an alternative.
- Not mentioned — invisible at this prompt sequence.
Tracking this over time is far more informative than counting “mentions per week”. The same number of mentions can mean very different things depending on which level they happen at.
3. One model is not the market
ChatGPT, Claude, Gemini, Perplexity, Mistral, and Copilot do not retrieve, weight, and summarize information the same way. A brand can be a level-1 default in Perplexity, a level-3 alternative in Claude, and invisible in Gemini.
Measuring on only one model gives a distorted view. Most teams optimize for whichever assistant they personally use, and miss large gaps elsewhere.
The same applies to language. AI assistants do not retrieve the same sources in French as in English. A brand strong in English can disappear entirely in Spanish or German, simply because there is less third-party content describing it in those languages.
What to do:
- Pick a set of models that match where your buyers actually research. For B2B SaaS today, that usually means ChatGPT, Claude, Gemini, and Perplexity at minimum.
- Pick languages that match your target markets, not just your home market.
- Run the same prompt sequences across every model and language combination.
- Treat each cell as its own data point. Do not average them into a single score that hides the gaps.
The output is a matrix, not a number.
4. Why weekly cadence beats one-shot audits
Two things change constantly: the web your brand is described in, and the models doing the describing. A single audit captures a moment. By the time you act on it, the picture has already shifted.
A weekly cadence is the smallest interval that catches meaningful change without producing noise. Daily measurement adds variance without adding signal — model outputs fluctuate enough day-to-day that short windows are dominated by randomness.
Weekly tracking lets you:
- See trends, not snapshots.
- Attribute changes to specific actions (a new comparison page, a press mention, a competitor launch).
- Catch silent regressions — cases where your visibility drops without any obvious external cause.
- Compare across models and languages with a consistent baseline.
What you want is a feedback loop, not a report.
5. What to record, every week
For each prompt sequence, model, and language, record at minimum:
- First-mention level (1 to 4, or not mentioned).
- Whether the mention is in the answer body, the shortlist, or only the citations.
- How the brand is described — the exact phrasing matters and changes over time.
- Which competitors are mentioned before you, and how often.
- Which sources are cited when your brand appears or is missing.
The last point is underrated. The cited sources are a map of what the model trusts about your category. If your brand is missing from the answer, it is almost always missing from the sources too. Fixing visibility starts with fixing source presence.
6. Common mistakes when measuring
A few patterns that make the data useless:
- Asking the model about itself (“do you know Klairia?”). This is a recall test, not a visibility test.
- Using leading prompts that name your category in your wording. Use the buyer’s wording.
- Mixing logged-in and logged-out sessions, or letting personalization leak in.
- Running prompts only once. Models are non-deterministic — repeat 3 to 5 times and look at frequency.
- Comparing across weeks with different prompt sets. Trends require a stable baseline.
- Looking only at the answer text and ignoring citations.
If the methodology shifts, the trend line is meaningless.
Putting it together
A usable measurement system has four parts:
- A stable set of progressive prompt sequences that match real buyer journeys.
- A defined set of models and languages that match your real market.
- A weekly cadence with consistent methodology.
- A first-mention level as the primary metric, supported by descriptions, competitor mentions, and cited sources.
That is enough to stop guessing. It will not tell you everything about your GEO performance, but it will tell you, week after week, whether the work you are doing on positioning, content, and external mentions is actually landing inside AI answers.
This is the loop we built Klairia around: weekly progressive-prompt scans across the models and languages you care about, with first-mention level, competitor gaps, and cited sources tracked over time — so you stop asking “did it work?” and start seeing the answer on a curve.