Citation rate, share of voice, prompt coverage and brand accuracy are used to mean materially different things by different vendors. That makes cross-vendor comparison meaningless and makes it easy to report a flattering number without lying. These are our definitions, with the formulas, what counts, what does not, and how each one is usually inflated.
How much of the buying conversation you are measuring at all. Every other metric is meaningless without it, because a metric measured on ten flattering prompts is not a measurement.
prompt coverage = prompts tested / prompts in the defined buying set
Define the set of questions a buyer in your category actually asks, including the follow-ups. Prompt coverage is the share of that set you are testing.
What counts
Questions a real buyer would type, in their words, including brand-free category questions
Follow-up turns, counted as their own prompts, because visibility routinely collapses at turn three
Comparison and alternatives questions naming competitors
Qualifying questions: pricing structure, integrations, company-size fit, migration effort
What does not
Prompts constructed to include your brand name, which test recall rather than discovery
Keyword strings rather than questions, which is not how anyone talks to an assistant
Prompts added after seeing the results, which is selection, not measurement
How this metric gets inflated
Reporting a high citation rate against a small, self-selected prompt set. Always ask for the denominator and how it was chosen before you look at the percentage.
02
Citation rate
How often you appear as a linked, attributed source in answers to your buying set.
citation rate = answers citing your domain / total answers returned for the tested prompt set
Of every answer generated for your prompt set, the proportion that links to a page on your domain.
What counts
A linked citation to a page on your own domain
One citation per answer, regardless of how many of your pages are cited in it
Citations in follow-up turns, counted against that turn's answer
What does not
Being named in the answer text without a link, which is brand mention, measured separately
Citations to third-party pages that talk about you, which is source-of-citation data, measured separately
Repeated runs of the same prompt counted as separate answers unless the sampling design says so
How this metric gets inflated
Blending brand mention into citation rate. It roughly doubles the number and is the most common inflation in the category. If a vendor cannot tell you which one they are reporting, they are reporting the flattering one.
03
Brand mention rate
How often you are named in the answer at all, cited or not. On Gemini and other assistants that hide sources, this matters more than citation rate.
brand mention rate = answers naming your brand / total answers returned for the tested prompt set
Of every answer generated for your prompt set, the proportion that names you in the text.
What counts
Your brand named anywhere in the answer text
Named as one of several options, not only as the recommendation
What does not
A generic description of your category that matches you without naming you
Mentions in a list the model then dismisses, which are recorded but reported separately with sentiment
How this metric gets inflated
Reporting mention rate as though it were recommendation. Being listed fourth of six is a mention, and reporting it without position is misleading.
04
Share of voice
Your visibility relative to the competitors you are actually being compared against, which is the only version of this number that means anything.
share of voice = your brand mentions / total brand mentions across the defined competitor set, over the same prompt set and window
Across every answer to your prompt set, count how many times each competitor is named. Your share of that total is your share of voice.
What counts
Mentions of any brand in the pre-declared competitor set, including yours
Multiple distinct brands in one answer, each counted once for that answer
What does not
Brands added to or removed from the competitor set mid-period, which invalidates the series
Categories of tool rather than named products
How this metric gets inflated
Choosing a favourable competitor set. The set must be declared before measurement and held constant, and a change to it starts a new series rather than continuing the old one.
05
Brand accuracy
Whether what the assistant says about you is true. Being visible and wrong is worse than being invisible.
brand accuracy = answers with no material factual error about your product / answers naming your brand
Of the answers that mention you, the proportion that describe you without a material error.
What counts
Material errors: wrong pricing, wrong category, a feature you do not have, a missing feature you do have, a deprecated integration presented as current
Errors of omission where the omission changes the buying decision
What does not
Wording we would not have chosen, where the substance is correct
Subjective evaluation, which is measured as sentiment rather than accuracy
How this metric gets inflated
Judging accuracy without a pre-agreed fact sheet. What counts as material must be written down before the answers are read, or the assessment moves to fit the result.
06
Source concentration
Which domains own the citations in your category. This is what tells you whether the work belongs on your site or off it.
source concentration = citations to a given domain / total citations across the tested prompt set
Count every source cited across all answers, group by domain, and rank them. The shape of that distribution is the brief for the citation sources programme.
What counts
Every cited domain, including your own
Each citation once per answer per domain
What does not
Sources shown in the interface but not used to build the answer, where the platform distinguishes the two
How this metric gets inflated
Not measuring it at all, which is the norm, and which is why so many programmes optimise the site while the citations sit on Reddit and G2.
Surge45
Method
The rules that make the metrics mean anything
Model outputs are non-deterministic, so a single run proves nothing. Without these rules the definitions above are still gameable.
1
Every metric is reported with its n and its window
A percentage without a denominator and a date range is not a measurement. If we cannot state both, we do not publish the number.
2
Prompts are declared before measurement, not after
The prompt set and the competitor set are written down and frozen before the first run. Adding a prompt because the results were disappointing starts a new series.
3
Each prompt is run multiple times
Model outputs are non-deterministic. A single run of a prompt tells you almost nothing. We run each prompt repeatedly within a window and report the rate across runs, not a snapshot.
4
Platforms are reported separately, never averaged
Averaging ChatGPT with Perplexity produces a number that describes nothing. A blended AI visibility score hides which platform is failing, which is usually the thing you needed to know.
5
Follow-up turns are measured, not just first answers
Single-prompt measurement systematically overstates performance, because visibility commonly collapses by the third turn. Conversations are tested as conversations.
6
Model and interface versions are recorded
Answers change between model releases without notice. A series that does not record which version produced it cannot explain its own discontinuities.
7
A change in method starts a new series
Restating history to fit a new method is how a reporting line becomes fiction. Old series are kept and labelled.
Deliberate omissions
What we will not report, and why
A single blended AI visibility score
It is the most saleable number in this category and the least useful. Platforms behave differently enough that averaging them destroys the information you needed.
Traffic attributed to AI without a referrer
Most assistant surfaces produce no referrer, so anything presented as AI-attributed traffic beyond genuinely referred clicks is modelled. Where we model, we say so and show the model.
Projected visibility
Forecasting citation rates on a platform whose behaviour changes without notice is guessing with a chart attached.