Primary research°

Which sources does each assistant actually cite?

A quarterly study of where ChatGPT, Gemini, Claude, Perplexity and Copilot get their answers from when a SaaS buyer asks a buying question. Reddit or G2? Vendor docs or publisher round-ups? The same across platforms, or completely different?
A technical team analysing which sources AI assistants cite

Status: Methodology published, first run pending

The methodology, the frozen question set and the categories are published on this page as of , before any measurement. The first report covers Q4 2026. There are no results here yet, and we would rather say so than show a chart of something we have not run.

Why this exists

The category has almost no original data in it

Nearly every claim published about AI search is an assertion. So is nearly every claim published by every competitor. The consequence is that the same small number of third-party studies get cited over and over, including the one we lean on ourselves below.

That existing work found the ten most-cited domains carry 35%+ of B2B SaaS citations, with Reddit first and G2 second. It is genuinely useful and it does not break the result down by assistant, which is the question a marketing team actually needs answered: if we can only work on one source type this quarter, which one, and for which platform?

So we are running it. Original data is the fastest route to citations because it is the one thing a model cannot get anywhere else, and it retro-fits credibility onto everything else we publish.

Source: Goodie, The Most Cited B2B SaaS Domains in AI Search. 5.7 million citation links analysed across B2B SaaS prompts.

Method

Eight rules, published before the first run

  1. 1

    Questions published before the first run

    Every question template and every category is on this page now, before any measurement. A benchmark whose questions were chosen after the answers were seen is an advertisement.

  2. 2

    Source types classified before scoring

    Community, review platform, publisher, analyst, video, vendor and other. The taxonomy is fixed in advance so a surprising result cannot be reclassified into a tidier one.

  3. 3

    Five runs per question per platform

    Model outputs are non-deterministic. A single run tells you nothing, and a benchmark built on single runs would report sampling noise as a finding.

  4. 4

    Logged out, no memory, no personalisation

    Personalisation would make the result a description of our own browsing rather than of the platform.

  5. 5

    Platforms reported separately

    The entire point is that they differ. Averaging them would destroy the only finding worth having.

  6. 6

    Model versions recorded

    Answers change between releases without notice. A series that cannot explain its own discontinuities is not usable.

  7. 7

    Raw responses published with every report

    The complete response set, so anyone can reclassify the sources themselves and disagree with our scoring. A benchmark whose working is private is a claim.

  8. 8

    No brand we advise appears in the reported examples

    The benchmark measures source types, not brands, so client work cannot influence the finding. Where a client's brand appears in a raw response, it is left in the data and not used as an illustration.

Platforms

  • ChatGPTSearch mode, logged out, no memory
  • Google GeminiThe assistant app, tested separately from AI Overviews and AI Mode
  • ClaudeWith web search enabled
  • PerplexityDefault model, source list recorded in order
  • Microsoft CopilotWeb answers, outside a Microsoft 365 tenant

Categories

  • CRM
  • Project management
  • Observability and monitoring
  • HR and payroll
  • Customer support and helpdesk
  • Analytics and BI
  • Marketing automation
  • Security and compliance
The question set

Four question types, templated across eight categories

The benchmark measures which source types each assistant draws on, so the questions are templated rather than written per category. Where each type is expected to behave differently, that expectation is stated in advance so the result can contradict it.

Open category discovery

The question a buyer starts with, before any brand is in mind. The purest test of which sources an assistant reaches for when it has nothing to anchor on.

  • What is the best [category] software?
  • What [category] tools should a mid-sized company consider?
  • What are the leading [category] platforms in 2026?

Comparison

Head-to-head and alternatives queries. We expect review platforms and community threads to dominate here, and the benchmark exists partly to test that expectation.

  • [Brand A] vs [Brand B]: which is better?
  • What are the alternatives to [Brand A]?
  • Which [category] tool is best for a team of 30?

Qualification

The narrow questions a buyer shortlists with. We expect vendor domains to perform relatively better here than anywhere else in the set.

  • How much does [category] software typically cost?
  • Which [category] tools integrate with [common platform]?
  • Which [category] tool is easiest to migrate to?

Problem-led

A buyer describing a problem rather than naming a category. Included because it is how a large share of real queries are phrased, and because it tends to surface a different source mix.

  • Our team keeps losing track of [problem]. What should we use?
  • How do other companies handle [problem]?
  • What is the cheapest way to solve [problem]?
What it will report

Stated before we have anything to report

  • Share of citations by source type, per assistant, across all categories
  • How that mix changes between question types, which is where we expect the sharpest differences
  • The most-cited individual domains per assistant
  • How much the source mix varies between the eight categories
  • Vendor-domain citation share, which is the number a SaaS marketing team will care about most
  • Run-to-run variance, so readers can see how stable any of it actually is

Cadence: Quarterly. Metric definitions follow the published measurement standard.

About the benchmark