The AI Visibility Measurement Standard
Published . Free to use, including against us.

Six metrics, six formulas
Each one is a single ratio. The numerator is what everyone quotes; the denominator is where the number gets made flattering. Both are stated for all six.
Prompt coverage
How much of the buying conversation you are measuring at all. Every other metric is meaningless without it, because a metric measured on ten flattering prompts is not a measurement.
Define the set of questions a buyer in your category actually asks, including the follow-ups. Prompt coverage is the share of that set you are testing.
What counts
- Questions a real buyer would type, in their words, including brand-free category questions
- Follow-up turns, counted as their own prompts, because visibility routinely collapses at turn three
- Comparison and alternatives questions naming competitors
- Qualifying questions: pricing structure, integrations, company-size fit, migration effort
What does not
- Prompts constructed to include your brand name, which test recall rather than discovery
- Keyword strings rather than questions, which is not how anyone talks to an assistant
- Prompts added after seeing the results, which is selection, not measurement
How this metric gets inflated
Reporting a high citation rate against a small, self-selected prompt set. Always ask for the denominator and how it was chosen before you look at the percentage.
Citation rate
How often you appear as a linked, attributed source in answers to your buying set.
Of every answer generated for your prompt set, the proportion that links to a page on your domain.
What counts
- A linked citation to a page on your own domain
- One citation per answer, regardless of how many of your pages are cited in it
- Citations in follow-up turns, counted against that turn's answer
What does not
- Being named in the answer text without a link, which is brand mention, measured separately
- Citations to third-party pages that talk about you, which is source-of-citation data, measured separately
- Repeated runs of the same prompt counted as separate answers unless the sampling design says so
How this metric gets inflated
Blending brand mention into citation rate. It roughly doubles the number and is the most common inflation in the category. If a vendor cannot tell you which one they are reporting, they are reporting the flattering one.
Brand mention rate
How often you are named in the answer at all, cited or not. On Gemini and other assistants that hide sources, this matters more than citation rate.
Of every answer generated for your prompt set, the proportion that names you in the text.
What counts
- Your brand named anywhere in the answer text
- Named as one of several options, not only as the recommendation
What does not
- A generic description of your category that matches you without naming you
- Mentions in a list the model then dismisses, which are recorded but reported separately with sentiment
How this metric gets inflated
Reporting mention rate as though it were recommendation. Being listed fourth of six is a mention, and reporting it without position is misleading.
Brand accuracy
Whether what the assistant says about you is true. Being visible and wrong is worse than being invisible.
Of the answers that mention you, the proportion that describe you without a material error.
What counts
- Material errors: wrong pricing, wrong category, a feature you do not have, a missing feature you do have, a deprecated integration presented as current
- Errors of omission where the omission changes the buying decision
What does not
- Wording we would not have chosen, where the substance is correct
- Subjective evaluation, which is measured as sentiment rather than accuracy
How this metric gets inflated
Judging accuracy without a pre-agreed fact sheet. What counts as material must be written down before the answers are read, or the assessment moves to fit the result.
Source concentration
Which domains own the citations in your category. This is what tells you whether the work belongs on your site or off it.
Count every source cited across all answers, group by domain, and rank them. The shape of that distribution is the brief for the citation sources programme.
What counts
- Every cited domain, including your own
- Each citation once per answer per domain
What does not
- Sources shown in the interface but not used to build the answer, where the platform distinguishes the two
How this metric gets inflated
Not measuring it at all, which is the norm, and which is why so many programmes optimise the site while the citations sit on Reddit and G2.
The rules that make the metrics mean anything
Model outputs are non-deterministic, so a single run proves nothing. Without these rules the definitions above are still gameable.
- 1
Every metric is reported with its n and its window
A percentage without a denominator and a date range is not a measurement. If we cannot state both, we do not publish the number.
- 2
Prompts are declared before measurement, not after
The prompt set and the competitor set are written down and frozen before the first run. Adding a prompt because the results were disappointing starts a new series.
- 3
Each prompt is run multiple times
Model outputs are non-deterministic. A single run of a prompt tells you almost nothing. We run each prompt repeatedly within a window and report the rate across runs, not a snapshot.
- 4
Platforms are reported separately, never averaged
Averaging ChatGPT with Perplexity produces a number that describes nothing. A blended AI visibility score hides which platform is failing, which is usually the thing you needed to know.
- 5
Follow-up turns are measured, not just first answers
Single-prompt measurement systematically overstates performance, because visibility commonly collapses by the third turn. Conversations are tested as conversations.
- 6
Model and interface versions are recorded
Answers change between model releases without notice. A series that does not record which version produced it cannot explain its own discontinuities.
- 7
A change in method starts a new series
Restating history to fit a new method is how a reporting line becomes fiction. Old series are kept and labelled.
What we will not report, and why
A single blended AI visibility score
It is the most saleable number in this category and the least useful. Platforms behave differently enough that averaging them destroys the information you needed.
Traffic attributed to AI without a referrer
Most assistant surfaces produce no referrer, so anything presented as AI-attributed traffic beyond genuinely referred clicks is modelled. Where we model, we say so and show the model.
Projected visibility
Forecasting citation rates on a platform whose behaviour changes without notice is guessing with a chart attached.