Assistant · OpenAI°

ChatGPT visibility for SaaS

The largest assistant by a distance, and the one whose citation behaviour changes most between a model answering from memory and a model running a search.
51.3%of US assistant usage

Source: First Page Sage, Top Generative AI Chatbots by Market Share. July 2026, US monthly-active-user estimates.

How ChatGPT retrieves and cites sources
Retrieval mechanism

Where a ChatGPT answer comes from

ChatGPT answers a question about your category in one of two modes, and they have almost nothing in common. In the first, the model answers from what it absorbed during training. There is no live fetch, no citation, and no way to influence the answer this quarter. In the second, it runs a search, retrieves pages, and builds an answer from what it just read, with links attached.

Almost everything that can be done for a brand in ChatGPT is work on the second mode. The first mode is a long game measured in training cycles: what shapes it is how widely and how consistently your brand is described across the open web, not anything you can change on your own site this month.

The practical consequence is that the two modes fail differently. If ChatGPT recommends a competitor without citing anything, you have a training-data and entity problem. If it runs a search and cites four sources that are not you, you have a retrieval problem, and retrieval problems are the ones that respond to work inside a quarter.

Crawler policy

Which crawler does what, and what blocking it costs you

Training crawlers and retrieval crawlers are different things. Blocking a retrieval crawler removes you from the answers. Blocking a training crawler does not. Getting the two confused is the most common self-inflicted AI visibility problem we find.

User agentPurposeWhat blocking it costs you
OAI-SearchBotRetrievalBuilds the index behind ChatGPT search results. Blocking it removes you from ChatGPT citations. This is the one that matters and the one most often blocked by accident.
ChatGPT-UserUser-triggered fetchFetches a specific page when a user or a ChatGPT action follows a link. Blocking it breaks the experience for someone who has already found you.
GPTBotTrainingGathers training data. Blocking it is a legitimate commercial decision and does not remove you from search citations. It is a different question from the two above, and should be decided separately.

Our own policy is published at surge45.com/robots.txt, with a line per agent and the reasoning in the file. If we are going to advise on crawler policy, our own should be readable.

Citation behaviour

How ChatGPT chooses what to cite

When ChatGPT searches, it retrieves a small set of pages and writes an answer from them, attaching links to the sources it drew on. The set is small: a handful of sources, not a page of ten blue links. The gap between being retrieved and being cited is narrow, but the gap between the top few results and everything else is enormous.

The sources it retrieves skew towards user-generated content more than the Google-side assistants do. Community discussion and review platforms carry disproportionate weight for software-recommendation queries, which is why an on-site-only programme underperforms here even when the on-site work is good.

Answers are also conversational and multi-turn. A buyer rarely asks one question. They ask for options, then narrow by price, then by integration, then by company size. Being cited on the first turn and absent by the fourth is a common and invisible failure, and it is only detectable if you test the whole sequence rather than a single prompt.

What moves it

Observed to change ChatGPT visibility

  • Presence on the sources ChatGPT retrieves

    Reddit threads, G2 and Capterra listings, and comparison content on publisher domains. These are not your properties, which is exactly why most competitors are not working on them.

  • Unambiguous entity definition

    A model has to be able to say what you are in one clause. Categories that are described differently on your homepage, your G2 listing and your LinkedIn page produce answers that hedge or substitute a competitor with a clearer description.

  • Answer-shaped content that survives extraction

    Content that answers the question in the first two sentences of a section, with the qualifying detail after. A model quoting you should be able to lift one paragraph and have it stand alone.

  • Freshness on comparison and pricing pages

    Comparison queries retrieve heavily and a stale page loses to a current one. This is the least glamorous work on the list and reliably the most effective.

What does not

Sold for ChatGPT, and not worth buying

  • Blocking GPTBot to force citations

    Training and retrieval are separate crawlers. Blocking the training crawler does not change search citations, and blocking the search crawler removes them entirely. This is the most common self-inflicted wound we find.

  • Keyword density and volume plays

    There is no ranking function to game. Retrieval selects passages that answer the question, then the model writes prose. More pages saying the same thing dilutes rather than compounds.

Working together

What we do on ChatGPT

  1. 1Build the real prompt set: the questions your buyers actually ask, including the follow-ups, not the head terms.
  2. 2Run them across the sequence and record which sources are cited at each turn.
  3. 3Separate the training-mode failures from the retrieval-mode failures, because they need different work on different timescales.
  4. 4Audit the OpenAI crawler policy in robots.txt, and fix the accidental blocks first.
  5. 5Prioritise the off-site sources that are being cited instead of you, then the on-site passages that would survive extraction.

ChatGPT questions we get asked