All Articles
    AI SearchPublished July 31, 20269 min read

    AI Visibility Score: How to Calculate Yours (With Benchmarks)

    The three-term formula behind an AI visibility score, the benchmark numbers from our 234-answer measurement, and why a score taken from one engine on one day is closer to a coin toss than a metric.

    Matiss Katanenko

    Matiss Katanenko

    Co-founder, Honeyb

    AI Visibility Score: How to Calculate Yours (With Benchmarks)

    Ask four AI engines the same buying question and they will name, on average, five brands each. You might reasonably expect a decent overlap between those lists. In our measurement, the average overlap between any two engines was 34.8%, and the worst pair managed 24.8%. Two engines, the same question, the same afternoon, and three-quarters of the brands were different.

    That is the problem an AI visibility score exists to solve, and also the reason most of them are wrong. A score is a single number for how often AI answer engines name you, how prominently, and how favourably. The number is useful. The way it is usually taken, once, from one engine, is not.

    Here is the formula, the benchmarks to judge your own number against, and the two mistakes that make most scores unfalsifiable.

    Where these numbers come from

    Every figure below marked as ours comes from one measurement run on 13 July 2026: 20 buyer-intent prompts, each run 3 times, across four engines via API (gpt-5-mini, gemini-2.5-flash, claude-haiku-4-5 and sonar). That is 240 attempted answers, of which 234 returned successfully, six Gemini calls having failed. Where a figure comes from someone else's research, it is attributed in the sentence.

    The formula

    An AI visibility score worth the name has three terms, because there are three separate ways to be invisible: never named, named last, or named badly.

    TermWhat it measuresWhy it has to be there
    Mention rateShare of answers that name you at allA brand absent from the answer cannot be chosen, whatever else is true
    Position weightWhere in the list you land, averaged over runsBeing named fifth of five is not the same product recommendation as being named first
    SentimentWhether the framing is positive, neutral or negativeBeing named as the expensive one is a mention and a loss

    Multiply, do not add. A brand mentioned in every answer, always last, always described as the budget compromise, should not score the same as one mentioned half as often but named first and well. Addition lets a strong term hide a term that is at zero, which is exactly the failure a score is supposed to surface.

    Benchmark 1: what a normal mention rate looks like

    The base rate is set by how many brands an engine names at all. Across 234 answers, the four engines averaged 5.04 brands per answer, and they were strikingly consistent about it.

    EngineBrands named per answerAnswers analysed
    Perplexity5.1860
    Claude5.1260
    ChatGPT5.0560
    Gemini4.8054
    All four5.04234

    Five slots. That is the whole shelf, and it gives you a usable yardstick: a brand named in every single answer for its category holds roughly 20% share of voice (the proportion of all brand mentions that are yours). Twenty per cent is therefore not a modest result, it is close to the ceiling. If a tool reports your share of voice at 60%, be suspicious of the denominator rather than delighted with the number.

    It also explains why the mention-rate term dominates the other two in practice. There is no long tail to climb into. You are in the five or you are nowhere, and no amount of good sentiment rescues a brand that never appears.

    Benchmark 2: why one run is not a measurement

    Now the uncomfortable part. We ran each prompt three times in a row, unchanged, and recorded how often the engine's top recommendation was a different brand on the next run.

    Top-pick change rate

    How often the top recommendation changes between identical runs

    Share of consecutive identical prompt runs where the engine's number-one recommended brand changed: Gemini 44%, Perplexity 43%, ChatGPT 35%, Claude 28%. Honeyb measurement, 13 July 2026: 20 buyer-intent prompts, 3 runs each, via API (gpt-5-mini, gemini-2.5-flash, claude-haiku-4-5, sonar).

    Claude was steadiest at 27.5%, Gemini the least stable at 44.1%. Read that again: on Gemini, the single most valuable position in the answer changed hands in nearly half of consecutive identical runs. Nothing about the question changed. Nothing about the brands changed. The answer changed anyway.

    The full brand list behaves the same way. Brand-set stability, meaning the share of brands that survived from one run to the next, ranged from 42.4% on ChatGPT to 67.4% on Claude. Even the steadiest engine reshuffled a third of the list between identical runs.

    EngineTop pick changes between identical runsBrand list stable between runs
    Claude27.5%67.4%
    ChatGPT35.0%42.4%
    Perplexity42.5%61.1%
    Gemini44.1%53.9%

    Want to get recommended by AI?

    Check your AI search visibility, then let the Honeyb agent write, fix, and earn what gets you recommended. Free to start.

    Free AI visibility checker

    This matches what others have found at larger scale. SparkToro's work put it more starkly: the same query changes its answer roughly 70% of the time, and two identical queries return the same brand list less than once in a hundred attempts. Our three-run sample is small enough to be conservative by comparison.

    The practical consequence is that a score taken from one run is not a low-precision measurement, it is a sample of size one from a distribution you have not looked at. Averaging over repeated runs is not a refinement you add later when you have budget. It is the difference between a metric and an anecdote, which is the case we made at greater length in our piece on why spot-checking a chatbot by hand fails.

    Benchmark 3: why one engine is not a score

    The second common mistake is quoting a score from whichever engine was to hand. The engines do not agree enough for that to work. Below is the mean overlap between the brand lists each pair of engines produced for the same prompts, alongside how often the pair picked the same brand first.

    Engine pairBrand list overlapSame brand ranked first
    Gemini and Perplexity47.4%36.8%
    Gemini and Claude43.8%52.6%
    Claude and Perplexity36.8%35.0%
    ChatGPT and Gemini28.8%47.4%
    ChatGPT and Claude27.1%40.0%
    ChatGPT and Perplexity24.8%20.0%

    ChatGPT and Perplexity, the two engines most people would name if asked to list AI search products, agreed on a quarter of the brands and on the top pick one time in five. A brand with a strong ChatGPT score and no Perplexity data does not have an AI visibility score. It has a ChatGPT score, and it should say so.

    Two honest caveats on this table. It reflects one model per engine at one point in time, and consumer apps may behave differently from the APIs we measured. The direction of the finding is robust, the decimal places are not.

    The sentiment term, and the mention you do not get credit for

    Sentiment is the term teams skip, usually because it feels soft next to a rate. It is not soft. It is the difference between the engine saying you are the best-supported option and the engine saying you are the one to pick if the budget is tight.

    There is also a subtler leak worth building into any honest score. Semrush found that 62% of AI citations never name the brand at all, meaning your page can be doing the work of informing the answer while your name stays out of it. A score built purely on citations of your domain will read higher than your actual visibility to a buyer, who only ever sees the text. Score the mention, not the citation, and treat the citation as a separate diagnostic.

    Which brings up the obvious question of what actually moves the number. Ahrefs' analysis found AI visibility correlates most strongly with third-party mentions and video, rather than with on-page work, and Reddit alone accounts for 40.1% of all AI citations, the single most-cited source (Semrush). The mechanics of that selection are worth understanding before you spend anything, and we walked through them in our explainer on how AI models decide which brands to recommend.

    Putting it together

    A defensible AI visibility score, then, is: your mention rate weighted by average position and adjusted for sentiment, measured across several engines, averaged over repeated runs, for a fixed set of prompts your buyers would plausibly type.

    That last clause matters more than it looks. Change the prompt set and you change the score, so a score is only comparable against itself over time. Anyone quoting a cross-industry AI visibility benchmark without publishing the prompts is quoting a number with no denominator.

    Judge your own result against these, from the run above:

    • Roughly 5 brands are named per answer, so a mention rate near 100% caps out around 20% share of voice. Treat 20% as excellent, not as a starting point.
    • Expect the top pick to move between 27.5% and 44.1% of the time depending on engine, so any week-on-week change smaller than that is noise, not progress.
    • Expect the engines to disagree, with a mean cross-engine brand overlap of 34.8%. A score from one engine tells you about one engine.
    • Score mentions, not citations, given that 62% of citations never name the brand (Semrush).
    Honeyb AI visibility tracking
    Scheduled measurement across engines, which is what turns a volatile signal into a trend line.

    None of this requires a tool. You can run a fixed prompt set by hand across four engines, three times each, and compute the three terms in a spreadsheet. It is genuinely tedious rather than genuinely difficult. Honeyb, which is our product and so treat this paragraph accordingly, automates exactly that loop: the same prompts, on a schedule, across engines, tracking mention rate, position, share of voice and sentiment so the volatility averages out into something you can act on.

    The honest summary is that the score is easy and the sampling is hard. A number taken once, from one engine, on one afternoon, is a coin toss wearing a decimal point. The same number taken properly is the only way to tell whether last quarter's coverage push changed anything. You can get a first reading for your own brand with the free AI visibility checker, then decide whether it is worth measuring on a schedule.

    Frequently asked questions

    What is an AI visibility score?

    It is a single number summarising how often AI answer engines name your brand, how prominently, and how favourably. A sound version multiplies three terms: mention rate (the share of answers naming you at all), position weight (where you land in the list, averaged across runs) and sentiment (whether the framing helps or hurts). Multiplying rather than adding matters, because it stops a strong term from masking one that is effectively zero.

    What is a good AI visibility score?

    Judge it against how many brands get named at all. Across 234 answers we measured on 13 July 2026, the four major engines named 5.04 brands per answer on average, so a brand appearing in every answer for its category tops out near 20% share of voice. That makes 20% close to a ceiling rather than a modest result. Scores are also only comparable against themselves over time, because changing the prompt set changes the number.

    How do I calculate AI visibility manually?

    Fix a set of prompts your buyers would actually type, run each one several times across ChatGPT, Gemini, Claude and Perplexity, and record for every answer whether you were named, in which position, and in what tone. Then average across runs and engines. The method is tedious rather than technically hard. The part people skip is the repetition, and that is the part that determines whether the result means anything.

    Why does my AI visibility score change every time I check it?

    Because the engines themselves are unstable. In our measurement, the top recommended brand changed between two identical consecutive runs 27.5% of the time on Claude and 44.1% of the time on Gemini, with no change to the prompt. SparkToro found the same query changes its answer roughly 70% of the time at larger scale. A single check is a sample of size one, so any movement smaller than those rates is noise rather than progress.

    Can I measure AI visibility on just one engine?

    You can, but you should call it what it is. Across the same prompts, the mean brand-list overlap between any two engines was 34.8%, and ChatGPT and Perplexity agreed on only 24.8% of brands and on the top pick just 20% of the time. A strong result on one engine says very little about the others, so a single-engine figure is a ChatGPT score or a Perplexity score, not an AI visibility score.

    Matiss Katanenko

    About the author

    Matiss Katanenko

    Co-founder, Honeyb

    My name is Matiss Katanenko and I co-founded Honeyb, the AI visibility platform that tracks how ChatGPT, Gemini, Claude, Perplexity and the other major AI engines talk about brands. I'm based in Riga, Latvia. Before Honeyb I spent years on the agency side running SEO and content programs for fast-growing brands across the US and Europe. That work is where I watched AI search start to compress the entire discovery channel into a four-brand short list, and decided to build the tool I wished agencies had. In my free time I'm in the sauna, on a padel court, or behind a drum kit.

    Connect on LinkedIn

    Free to start

    Get recommended by AI search models.

    Run a free AI search visibility check, then let the Honeyb agent do the work that gets you into the answers.

    ChatGPTClaudeGeminiPerplexity