Brand share of voice in AI answers: ChatGPT vs Perplexity
Illustration: AI-generated
More and more B2B decision makers start their research in ChatGPT or Perplexity rather than Google. According to the G2 AI Search Insight Report 2026, which surveyed 1,076 B2B decision makers, 51 percent of B2B software buyers now start their vendor research inside an AI chatbot.
If you are not mentioned there, you often never make the shortlist in the first place. So the decisive question is no longer only "do we rank on Google?", it is: how often and how prominently do we appear in the models answers compared to our competitors? That is what AI share of voice measures.
What is AI share of voice?
AI share of voice is the proportion of AI answers in which your brand is mentioned or recommended for a defined set of buyer questions, relative to your competitors. There are two common calculations, and they answer different questions:
- Mention rate: answers containing your brand divided by all tested answers, times 100. It tells you how present you are at all.
- Competitive share of voice: your mentions divided by all brand mentions in the test set, times 100. It tells you how much of the pie you get.
The second number is the more meaningful one. A high mention rate paired with a low competitive share of voice means you do get named, but the big players still dominate the answer.
And the most important limitation up front: there is no official metric. Neither OpenAI nor Perplexity publishes anything like it. Every measurement is a snapshot and depends on prompt selection, region, account, timing and model version. That is exactly why you need a reproducible method rather than a gut feeling.
Why you have to measure ChatGPT and Perplexity separately
The two systems pull brands into their answers in fundamentally different ways. Several independent analyses show just how far apart they are, and they agree surprisingly closely:
Only around 11 percent of cited domains appear in both systems. Averi found this across 680 million citations, Whitehat across 118,000 answers and Profound across 100,000 prompts, independently of each other. Perplexity cites 21.9 sources per answer on average, ChatGPT only 10.4. And with ChatGPT, Wikipedia alone accounts for roughly 47.9 percent of the top-10 source share.
These are effectively two separate source worlds. What works on one system does not automatically carry over to the other.
| ChatGPT | Perplexity | |
|---|---|---|
| Main source | training knowledge plus selective web search | live retrieval on almost every query |
| Citations | rarer, more of a narrative recommendation | frequent and transparent, with source links |
| Preferred sources | Wikipedia and established authorities | fresh, easily citable pages and communities |
| Freshness | depends on training cut-off and search mode | very high, close to real time |
| What it means for you | build entity signals and authority | deliver citable, current content |
In practice: measure separately, report separately, optimise separately. A single combined number across both systems blurs exactly the difference that matters.
How do you actually measure share of voice?
You do not need a tool for a solid first picture. A spreadsheet and some discipline are enough.
Step 1, the prompt set. Collect 15 to 25 real buyer questions your audience would ask. Crucially: no brand names inside the prompts. Otherwise you measure recognition instead of recommendation. Cover awareness, comparison and decision questions, and freeze the exact wording, because the prompt set is your yardstick. Examples for a B2B service company:
- Best agencies for AI visibility and GEO in Germany 2026
- How do I measure brand visibility in ChatGPT and Perplexity?
- Which tools help with tracking AI share of voice?
- What do B2B buyers look for when choosing an AI marketing agency?
- ChatGPT or Perplexity: which is better for vendor research?
Step 2, the runs. Measure at least in ChatGPT with web search enabled and in Perplexity. Run every prompt three to five times per system, because the answers vary. Use clean sessions, so a new chat or a window without history, so that stored context does not skew the result. Keep region and language constant.
Step 3, the log. For every answer note: was the brand mentioned, in which position, in what way (recommended, merely mentioned, compared), which competitors appear, which sources are cited, what the tone is, and with which date and model version you measured. Without date and model version the measurement cannot be reproduced later.
Step 4, the calculation. Always calculate per system, never across both together. A worked example: 20 prompts times 5 runs gives 100 answers per system. If your brand appears in 14 of them, the mention rate is 14 percent. If you count 93 brand mentions in total across the same data set, your competitive share of voice is 15 percent. Optionally add the citation rate, meaning how often your own domain shows up as a source.
Step 5, the repeat. Run the same prompt set again every one to two weeks. Individual outliers are normal; what is meaningful is the trend across several measurements.
Where the measurement hits its limits
- Results vary between runs, regions and accounts. A single measurement is worthless.
- There is no official share-of-voice figure from OpenAI or Perplexity to calibrate against.
- A single prompt says almost nothing. Only the distribution across many questions and several runs produces a picture.
- Branded questions measure recognition, not competitive position. What counts for the shortlist are the unbranded buying questions.
- Tools automate this and scale better. For getting started and understanding the logic, manual measurement is still unbeatable.
What you do with the result
The number alone changes nothing. Its value is that it shows where the gap sits, and the gap looks completely different depending on the system.
A low score on Perplexity is usually a content problem: fresh, clearly structured and easily citable pages are missing, along with presence on community and review platforms. That can be fixed comparatively quickly, because Perplexity reads in real time.
A low score on ChatGPT is usually an authority problem: entity signals are missing, as are mentions in established sources and a consistent description of the brand across many pages. That takes longer, but it also lasts longer.
So the question of which measure to take matters more than the number itself. A weak position on Perplexity calls for content work. A weak position on ChatGPT calls for authority building. Confusing the two costs months.
How we set up a measurement
We follow exactly this method: an unbranded prompt set of real buying questions, separate measurement in both systems, several runs against the variance, a log with date and model version. But the part clients bring us in for is not the counting. It is the interpretation: which finding points to which lever, and in what order the work pays off.
And one limitation we state openly: our WhatsApp check is a fast first assessment, not a full measurement. In a few minutes it answers whether your brand shows up at all for typical buying questions, and who it shows up next to. For a solid number over time you need the full prompt set and repeated runs, as described above. We deliberately keep the two apart.
Conclusion
As long as an 11 percent overlap between two systems is the norm, a single visibility number is worthless. What counts is a fixed prompt set, separate measurement per system, enough runs to beat the noise, and the honesty to treat the result as a snapshot.
Getting started costs about an hour: ten to fifteen of your most important buyer questions, two systems, one spreadsheet. After that you know for the first time whether the models know you at all, and whether they recommend you.
Key Takeaways
- 1Only around 11 percent of cited domains overlap between ChatGPT and Perplexity, which makes a single combined visibility number worthless.
- 2Measure with a fixed, unbranded prompt set of 15 to 25 real buying questions. Brand names inside the prompt measure recognition, not recommendation.
- 3Three to five runs per prompt and system, otherwise you are measuring noise rather than position.
- 4Competitive share of voice says more than the mention rate: you can be named and still stand in the shadow of the big players.
- 5A low score on Perplexity usually means a content problem, a low score on ChatGPT usually an authority problem. The remedy is fundamentally different.
Sources (selection): G2 AI Search Insight Report 2026 (1,076 B2B decision makers surveyed, March 2026) for the share of buyers starting their research inside an AI chatbot; Averi (680 million citations), Whitehat SEO (118,000 answers) and Profound (100,000 prompts) for the overlap of cited domains, the split per system and the number of sources per answer. All figures quoted are snapshots from 2026 and change with model versions and provider product decisions.