Version 1.0°

The AI Visibility Measurement Standard

Citation rate, share of voice, prompt coverage and brand accuracy are used to mean materially different things by different vendors. That makes cross-vendor comparison meaningless and makes it easy to report a flattering number without lying. These are our definitions, with the formulas, what counts, what does not, and how each one is usually inflated.

Published . Free to use, including against us.

A dashboard showing AI visibility metrics for a SaaS brand
Definitions

Six metrics, six formulas

01

Prompt coverage

How much of the buying conversation you are measuring at all. Every other metric is meaningless without it, because a metric measured on ten flattering prompts is not a measurement.

prompt coverage = prompts tested / prompts in the defined buying set

Define the set of questions a buyer in your category actually asks, including the follow-ups. Prompt coverage is the share of that set you are testing.

What counts

  • Questions a real buyer would type, in their words, including brand-free category questions
  • Follow-up turns, counted as their own prompts, because visibility routinely collapses at turn three
  • Comparison and alternatives questions naming competitors
  • Qualifying questions: pricing structure, integrations, company-size fit, migration effort

What does not

  • Prompts constructed to include your brand name, which test recall rather than discovery
  • Keyword strings rather than questions, which is not how anyone talks to an assistant
  • Prompts added after seeing the results, which is selection, not measurement

How this metric gets inflated

Reporting a high citation rate against a small, self-selected prompt set. Always ask for the denominator and how it was chosen before you look at the percentage.

02

Citation rate

How often you appear as a linked, attributed source in answers to your buying set.

citation rate = answers citing your domain / total answers returned for the tested prompt set

Of every answer generated for your prompt set, the proportion that links to a page on your domain.

What counts

  • A linked citation to a page on your own domain
  • One citation per answer, regardless of how many of your pages are cited in it
  • Citations in follow-up turns, counted against that turn's answer

What does not

  • Being named in the answer text without a link, which is brand mention, measured separately
  • Citations to third-party pages that talk about you, which is source-of-citation data, measured separately
  • Repeated runs of the same prompt counted as separate answers unless the sampling design says so

How this metric gets inflated

Blending brand mention into citation rate. It roughly doubles the number and is the most common inflation in the category. If a vendor cannot tell you which one they are reporting, they are reporting the flattering one.

03

Brand mention rate

How often you are named in the answer at all, cited or not. On Gemini and other assistants that hide sources, this matters more than citation rate.

brand mention rate = answers naming your brand / total answers returned for the tested prompt set

Of every answer generated for your prompt set, the proportion that names you in the text.

What counts

  • Your brand named anywhere in the answer text
  • Named as one of several options, not only as the recommendation

What does not

  • A generic description of your category that matches you without naming you
  • Mentions in a list the model then dismisses, which are recorded but reported separately with sentiment

How this metric gets inflated

Reporting mention rate as though it were recommendation. Being listed fourth of six is a mention, and reporting it without position is misleading.

04

Share of voice

Your visibility relative to the competitors you are actually being compared against, which is the only version of this number that means anything.

share of voice = your brand mentions / total brand mentions across the defined competitor set, over the same prompt set and window

Across every answer to your prompt set, count how many times each competitor is named. Your share of that total is your share of voice.

What counts

  • Mentions of any brand in the pre-declared competitor set, including yours
  • Multiple distinct brands in one answer, each counted once for that answer

What does not

  • Brands added to or removed from the competitor set mid-period, which invalidates the series
  • Categories of tool rather than named products

How this metric gets inflated

Choosing a favourable competitor set. The set must be declared before measurement and held constant, and a change to it starts a new series rather than continuing the old one.

05

Brand accuracy

Whether what the assistant says about you is true. Being visible and wrong is worse than being invisible.

brand accuracy = answers with no material factual error about your product / answers naming your brand

Of the answers that mention you, the proportion that describe you without a material error.

What counts

  • Material errors: wrong pricing, wrong category, a feature you do not have, a missing feature you do have, a deprecated integration presented as current
  • Errors of omission where the omission changes the buying decision

What does not

  • Wording we would not have chosen, where the substance is correct
  • Subjective evaluation, which is measured as sentiment rather than accuracy

How this metric gets inflated

Judging accuracy without a pre-agreed fact sheet. What counts as material must be written down before the answers are read, or the assessment moves to fit the result.

06

Source concentration

Which domains own the citations in your category. This is what tells you whether the work belongs on your site or off it.

source concentration = citations to a given domain / total citations across the tested prompt set

Count every source cited across all answers, group by domain, and rank them. The shape of that distribution is the brief for the citation sources programme.

What counts

  • Every cited domain, including your own
  • Each citation once per answer per domain

What does not

  • Sources shown in the interface but not used to build the answer, where the platform distinguishes the two

How this metric gets inflated

Not measuring it at all, which is the norm, and which is why so many programmes optimise the site while the citations sit on Reddit and G2.

Method

The rules that make the metrics mean anything

Model outputs are non-deterministic, so a single run proves nothing. Without these rules the definitions above are still gameable.

  1. 1

    Every metric is reported with its n and its window

    A percentage without a denominator and a date range is not a measurement. If we cannot state both, we do not publish the number.

  2. 2

    Prompts are declared before measurement, not after

    The prompt set and the competitor set are written down and frozen before the first run. Adding a prompt because the results were disappointing starts a new series.

  3. 3

    Each prompt is run multiple times

    Model outputs are non-deterministic. A single run of a prompt tells you almost nothing. We run each prompt repeatedly within a window and report the rate across runs, not a snapshot.

  4. 4

    Platforms are reported separately, never averaged

    Averaging ChatGPT with Perplexity produces a number that describes nothing. A blended AI visibility score hides which platform is failing, which is usually the thing you needed to know.

  5. 5

    Follow-up turns are measured, not just first answers

    Single-prompt measurement systematically overstates performance, because visibility commonly collapses by the third turn. Conversations are tested as conversations.

  6. 6

    Model and interface versions are recorded

    Answers change between model releases without notice. A series that does not record which version produced it cannot explain its own discontinuities.

  7. 7

    A change in method starts a new series

    Restating history to fit a new method is how a reporting line becomes fiction. Old series are kept and labelled.

Deliberate omissions

What we will not report, and why

A single blended AI visibility score

It is the most saleable number in this category and the least useful. Platforms behave differently enough that averaging them destroys the information you needed.

Traffic attributed to AI without a referrer

Most assistant surfaces produce no referrer, so anything presented as AI-attributed traffic beyond genuinely referred clicks is modelled. Where we model, we say so and show the model.

Projected visibility

Forecasting citation rates on a platform whose behaviour changes without notice is guessing with a chart attached.

About this standard