How to monitor AI search visibility
A method you can run yourself: how to build a prompt set, how many runs you need, what to record, and the three reporting mistakes that make a monitoring programme worthless.
The short answer
Build a prompt set of 30 to 50 buyer questions and freeze it. Run each at least five times a month, per platform, recording the answer, the sources and the model version. Report citation rate and mention rate separately.
Build the prompt set, then freeze it
The prompt set is the denominator for every number you will produce, so it has to be written down and fixed before you measure anything. Adding a prompt after seeing disappointing results is not measurement, it is selection.
A workable set is 30 to 50 prompts across four stages, and the fourth is the one everyone forgets.
Category discovery
The open question a buyer starts with, with no brand named. "What are the best X tools for Y?" This is where you find out whether you exist to the model at all, and it is the shape our public assistant citation benchmark uses.
Comparison
Head-to-head against named competitors, and "alternatives to X" queries. High intent, and the ones where third-party sources dominate.
Qualification
Pricing structure, integrations, company-size fit, migration effort. The narrow questions a buyer uses to shortlist.
Follow-up turns
The second and third question in a conversation, counted as their own prompts. Visibility routinely collapses by turn three, and single-prompt testing never sees it.
Decide how many runs you need
Model outputs are non-deterministic. The same prompt, asked twice, can produce different answers citing different sources. A single run tells you almost nothing, and a monitoring programme built on single runs will report noise as trend.
Five runs per prompt per window is the practical minimum we use. Below three, run-to-run variance swamps any real change. Above ten, you are paying for precision you will not act on.
Record the model version with every run. Answers change between releases without notice, and a series that cannot explain its own discontinuities is not much use in a board report.
Record the right things
| Field | Why it matters |
|---|---|
| Platform and model version | Answers change between releases. Without this you cannot explain a step change. |
| Full answer text | So accuracy can be assessed later against a fact sheet, rather than judged in the moment. |
| Every source cited, in order | Source concentration is the finding that most often redirects the whole programme. |
| Whether your brand was named | Brand mention rate. Kept separate from citation rate, always. |
| Whether your domain was linked | Citation rate. This is the stricter measure and the one usually inflated. |
| Which competitors were named | Share of voice, against a competitor set declared in advance. |
| Any factual error about you | Brand accuracy, and the source it came from, which is usually a third-party listing. |
Three reporting mistakes to avoid
Blending citation rate and brand mention rate
This roughly doubles the headline number and is the most common inflation in the category. If a tool cannot tell you which it reports, assume the flattering one, and check it against our published definitions.
Averaging platforms into one score
A blended AI visibility score is the most saleable number here and the least useful. Platforms behave differently enough that averaging destroys the information you needed, which is usually which one is failing.
Presenting modelled traffic as measured
Most assistant surfaces send no referrer, so anything beyond genuinely referred clicks is a model. Show the model or do not report the number.
Tooling, and what it will not do
Platforms like Profound, Ahrefs Brand Radar and WriteWorks handle the coverage and cadence problem: many engines, refreshed frequently, across regions, which no human can do by hand at a useful interval. That is a real and substantial saving.
What none of them do is decide which questions are worth asking, judge whether an answer is materially wrong about your product, or work out that most of your accuracy failures trace to one stale third-party listing. That judgement is the work, and a dashboard does not contain it. Our worked audit shows what it looks like applied.
WriteWorks is built by Surge45, which we mention wherever we mention the tool.
Our position
Run the tools against your own published definitions rather than accepting theirs. A number you cannot reproduce by hand is not a measurement, it is a subscription.
Frequently asked questions
Can we do this without a tool?
Yes, and it is worth doing manually once before you buy anything. Thirty prompts, five runs, five platforms is 750 queries: tedious over a fortnight, and it teaches you more about your category than any dashboard will. The short version takes about an hour.
How often should we report?
Monthly. Weekly produces variance you will over-interpret. The exception is a re-baseline after a major model release, which is worth doing whenever one lands.
What is a good citation rate?
There is no universal benchmark and anyone quoting one is selling something. Rates vary enormously by category and by how much third-party coverage exists. What is comparable is your own series over time and your share of voice against a fixed competitor set.
Sources
- OpenAI, overview of OpenAI crawlers. OpenAI's own documentation of GPTBot, OAI-SearchBot and ChatGPT-User.
About Surge45 Team
AI Search & GEO Specialists
Surge45 is the digital discovery and growth strategic advisory for SaaS. We help software companies become the answer across Google, AI search and communities, then turn discovery into pipeline. We also build WriteWorks, our content engineering platform for AI search.
Related Articles
Get More SaaS Marketing Insights
Join 10,000+ SaaS marketing leaders who receive our weekly newsletter with actionable strategies.
Subscribe to Newsletter