How-to°

How to monitor AI search visibility

· 8 min read ·
How to monitor AI search visibility

Build a prompt set of 30 to 50 buyer questions and freeze it. Run each at least five times a month, per platform, recording the answer, the sources and the model version. Report citation rate and mention rate separately.

Build the prompt set, then freeze it

The prompt set is the denominator for every number you will produce, so it has to be written down and fixed before you measure anything. Adding a prompt after seeing disappointing results is not measurement, it is selection.

A workable set is 30 to 50 prompts across four stages, and the fourth is the one everyone forgets.

  • Category discovery

    The open question a buyer starts with, with no brand named. "What are the best X tools for Y?" This is where you find out whether you exist to the model at all.

  • Comparison

    Head-to-head against named competitors, and "alternatives to X" queries. High intent, and the ones where third-party sources dominate.

  • Qualification

    Pricing structure, integrations, company-size fit, migration effort. The narrow questions a buyer uses to shortlist.

  • Follow-up turns

    The second and third question in a conversation, counted as their own prompts. Visibility routinely collapses by turn three, and single-prompt testing never sees it.

Decide how many runs you need

Model outputs are non-deterministic. The same prompt, asked twice, can produce different answers citing different sources. A single run tells you almost nothing, and a monitoring programme built on single runs will report noise as trend.

Five runs per prompt per window is the practical minimum we use. Below three, run-to-run variance swamps any real change. Above ten, you are paying for precision you will not act on.

Record the model version with every run. Answers change between releases without notice, and a series that cannot explain its own discontinuities is not much use in a board meeting.

Record the right things

What to capture per run
FieldWhy it matters
Platform and model versionAnswers change between releases. Without this you cannot explain a step change.
Full answer textSo accuracy can be assessed later against a fact sheet, rather than judged in the moment.
Every source cited, in orderSource concentration is the finding that most often redirects the whole programme.
Whether your brand was namedBrand mention rate. Kept separate from citation rate, always.
Whether your domain was linkedCitation rate. This is the stricter measure and the one usually inflated.
Which competitors were namedShare of voice, against a competitor set declared in advance.
Any factual error about youBrand accuracy, and the source it came from, which is usually a third-party listing.

Three reporting mistakes to avoid

  • Blending citation rate and brand mention rate

    This roughly doubles the headline number and is the most common inflation in the category. If a tool cannot tell you which it reports, assume the flattering one.

  • Averaging platforms into one score

    A blended AI visibility score is the most saleable number here and the least useful. Platforms behave differently enough that averaging destroys the information you needed, which is usually which one is failing.

  • Presenting modelled traffic as measured

    Most assistant surfaces send no referrer, so anything beyond genuinely referred clicks is a model. Show the model or do not report the number.

Tooling, and what it will not do

Platforms like Profound, Ahrefs Brand Radar and WriteWorks handle the coverage and cadence problem: many engines, refreshed frequently, across regions, which no human can do by hand at a useful interval. That is a real and substantial saving.

What none of them do is decide which questions are worth asking, judge whether an answer is materially wrong about your product, or work out that most of your accuracy failures trace to one stale third-party listing. That judgement is the work, and a dashboard does not contain it.

WriteWorks is built by Surge45, which we mention wherever we mention the tool.

Our position

Run the tools against your own published definitions rather than accepting theirs. A number you cannot reproduce by hand is not a measurement, it is a subscription.

Related questions

Our measurement definitions, published in full

How to monitor AI search visibility | Surge45