Ask four AI engines the same buying question and they will name, on average, five brands each. You might reasonably expect a decent overlap between those lists. In our measurement, the average overlap between any two engines was 34.8%, and the worst pair managed 24.8%. Two engines, the same question, the same afternoon, and three-quarters of the brands were different.
That is the problem an AI visibility score exists to solve, and also the reason most of them are wrong. A score is a single number for how often AI answer engines name you, how prominently, and how favourably. The number is useful. The way it is usually taken, once, from one engine, is not.
Here is the formula, the benchmarks to judge your own number against, and the two mistakes that make most scores unfalsifiable.
Where these numbers come from
Every figure below marked as ours comes from one measurement run on 13 July 2026: 20 buyer-intent prompts, each run 3 times, across four engines via API (gpt-5-mini, gemini-2.5-flash, claude-haiku-4-5 and sonar). That is 240 attempted answers, of which 234 returned successfully, six Gemini calls having failed. Where a figure comes from someone else's research, it is attributed in the sentence.
The formula
An AI visibility score worth the name has three terms, because there are three separate ways to be invisible: never named, named last, or named badly.
| Term | What it measures | Why it has to be there |
|---|---|---|
| Mention rate | Share of answers that name you at all | A brand absent from the answer cannot be chosen, whatever else is true |
| Position weight | Where in the list you land, averaged over runs | Being named fifth of five is not the same product recommendation as being named first |
| Sentiment | Whether the framing is positive, neutral or negative | Being named as the expensive one is a mention and a loss |
Multiply, do not add. A brand mentioned in every answer, always last, always described as the budget compromise, should not score the same as one mentioned half as often but named first and well. Addition lets a strong term hide a term that is at zero, which is exactly the failure a score is supposed to surface.
Benchmark 1: what a normal mention rate looks like
The base rate is set by how many brands an engine names at all. Across 234 answers, the four engines averaged 5.04 brands per answer, and they were strikingly consistent about it.
| Engine | Brands named per answer | Answers analysed |
|---|---|---|
| Perplexity | 5.18 | 60 |
| Claude | 5.12 | 60 |
| ChatGPT | 5.05 | 60 |
| Gemini | 4.80 | 54 |
| All four | 5.04 | 234 |
Five slots. That is the whole shelf, and it gives you a usable yardstick: a brand named in every single answer for its category holds roughly 20% share of voice (the proportion of all brand mentions that are yours). Twenty per cent is therefore not a modest result, it is close to the ceiling. If a tool reports your share of voice at 60%, be suspicious of the denominator rather than delighted with the number.
It also explains why the mention-rate term dominates the other two in practice. There is no long tail to climb into. You are in the five or you are nowhere, and no amount of good sentiment rescues a brand that never appears.
Benchmark 2: why one run is not a measurement
Now the uncomfortable part. We ran each prompt three times in a row, unchanged, and recorded how often the engine's top recommendation was a different brand on the next run.
Top-pick change rate
How often the top recommendation changes between identical runs
Claude was steadiest at 27.5%, Gemini the least stable at 44.1%. Read that again: on Gemini, the single most valuable position in the answer changed hands in nearly half of consecutive identical runs. Nothing about the question changed. Nothing about the brands changed. The answer changed anyway.
The full brand list behaves the same way. Brand-set stability, meaning the share of brands that survived from one run to the next, ranged from 42.4% on ChatGPT to 67.4% on Claude. Even the steadiest engine reshuffled a third of the list between identical runs.
| Engine | Top pick changes between identical runs | Brand list stable between runs |
|---|---|---|
| Claude | 27.5% | 67.4% |
| ChatGPT | 35.0% | 42.4% |
| Perplexity | 42.5% | 61.1% |
| Gemini | 44.1% | 53.9% |
This matches what others have found at larger scale. SparkToro's work put it more starkly: the same query changes its answer roughly 70% of the time, and two identical queries return the same brand list less than once in a hundred attempts. Our three-run sample is small enough to be conservative by comparison.
The practical consequence is that a score taken from one run is not a low-precision measurement, it is a sample of size one from a distribution you have not looked at. Averaging over repeated runs is not a refinement you add later when you have budget. It is the difference between a metric and an anecdote, which is the case we made at greater length in our piece on why spot-checking a chatbot by hand fails.
Benchmark 3: why one engine is not a score
The second common mistake is quoting a score from whichever engine was to hand. The engines do not agree enough for that to work. Below is the mean overlap between the brand lists each pair of engines produced for the same prompts, alongside how often the pair picked the same brand first.
| Engine pair | Brand list overlap | Same brand ranked first |
|---|---|---|
| Gemini and Perplexity | 47.4% | 36.8% |
| Gemini and Claude | 43.8% | 52.6% |
| Claude and Perplexity | 36.8% | 35.0% |
| ChatGPT and Gemini | 28.8% | 47.4% |
| ChatGPT and Claude | 27.1% | 40.0% |
| ChatGPT and Perplexity | 24.8% | 20.0% |
ChatGPT and Perplexity, the two engines most people would name if asked to list AI search products, agreed on a quarter of the brands and on the top pick one time in five. A brand with a strong ChatGPT score and no Perplexity data does not have an AI visibility score. It has a ChatGPT score, and it should say so.
Two honest caveats on this table. It reflects one model per engine at one point in time, and consumer apps may behave differently from the APIs we measured. The direction of the finding is robust, the decimal places are not.
The sentiment term, and the mention you do not get credit for
Sentiment is the term teams skip, usually because it feels soft next to a rate. It is not soft. It is the difference between the engine saying you are the best-supported option and the engine saying you are the one to pick if the budget is tight.
There is also a subtler leak worth building into any honest score. Semrush found that 62% of AI citations never name the brand at all, meaning your page can be doing the work of informing the answer while your name stays out of it. A score built purely on citations of your domain will read higher than your actual visibility to a buyer, who only ever sees the text. Score the mention, not the citation, and treat the citation as a separate diagnostic.
Which brings up the obvious question of what actually moves the number. Ahrefs' analysis found AI visibility correlates most strongly with third-party mentions and video, rather than with on-page work, and Reddit alone accounts for 40.1% of all AI citations, the single most-cited source (Semrush). The mechanics of that selection are worth understanding before you spend anything, and we walked through them in our explainer on how AI models decide which brands to recommend.
Putting it together
A defensible AI visibility score, then, is: your mention rate weighted by average position and adjusted for sentiment, measured across several engines, averaged over repeated runs, for a fixed set of prompts your buyers would plausibly type.
That last clause matters more than it looks. Change the prompt set and you change the score, so a score is only comparable against itself over time. Anyone quoting a cross-industry AI visibility benchmark without publishing the prompts is quoting a number with no denominator.
Judge your own result against these, from the run above:
- Roughly 5 brands are named per answer, so a mention rate near 100% caps out around 20% share of voice. Treat 20% as excellent, not as a starting point.
- Expect the top pick to move between 27.5% and 44.1% of the time depending on engine, so any week-on-week change smaller than that is noise, not progress.
- Expect the engines to disagree, with a mean cross-engine brand overlap of 34.8%. A score from one engine tells you about one engine.
- Score mentions, not citations, given that 62% of citations never name the brand (Semrush).

None of this requires a tool. You can run a fixed prompt set by hand across four engines, three times each, and compute the three terms in a spreadsheet. It is genuinely tedious rather than genuinely difficult. Honeyb, which is our product and so treat this paragraph accordingly, automates exactly that loop: the same prompts, on a schedule, across engines, tracking mention rate, position, share of voice and sentiment so the volatility averages out into something you can act on.
The honest summary is that the score is easy and the sampling is hard. A number taken once, from one engine, on one afternoon, is a coin toss wearing a decimal point. The same number taken properly is the only way to tell whether last quarter's coverage push changed anything. You can get a first reading for your own brand with the free AI visibility checker, then decide whether it is worth measuring on a schedule.














