All Articles
    AI Visibility
    Published July 20, 20268 min read

    AI Search Visibility: How to Actually Put a Number on It

    Everyone defines AI search visibility. Almost nobody shows the measurement. Here is how the number is built, why most reported figures are unstable, and what a tracker has to capture to be trusted.

    Matiss Katanenko

    Matiss Katanenko

    Co-founder, Honeyb

    AI Search Visibility: How to Actually Put a Number on It

    AI search visibility is how often AI assistants name your brand in the answers they give to buyer questions. It is measured as a rate, not a position: you run a fixed set of prompts repeatedly across ChatGPT, Gemini, Claude and Perplexity, count the answers that name you, and divide. A single check tells you almost nothing, because the same query changes its answer roughly 70% of the time (SparkToro).

    That last point is why this metric is harder than it looks. Most published visibility numbers come from one run of one prompt on one engine, which is a coin toss dressed as a measurement. The rest of this page covers what the number is made of, what separates it from two things it gets confused with, and what an AI search tracker has to capture before its output is worth acting on.

    Three different metrics people call the same thing

    Visibility, citations and referral traffic are measured differently, break for different reasons, and are fixed by different work. Conflating them is the most common reason a team spends a quarter on the wrong problem.

    MetricQuestion it answersWhere it is measuredTypical fix when it is low
    AI search visibilityIs your brand named in the answer text?Prompt runs across engines, repeatedThird-party mentions, reviews, comparison coverage
    AI citationsIs your domain listed as a source?Citation lists on answers that expose themContent that answers the question directly, crawlable and current
    AI referral trafficDid anyone actually click through?GA4 and server logsNothing on its own; it is downstream of the other two

    The gap between the first two is not theoretical. Semrush found that 62% of AI citations never name the brand being cited. Your domain can be the source an engine reads while a competitor is the name the reader walks away with. Track citations only and you will report progress that no buyer experienced. Track mentions only and you will not know which pages are doing the work.

    Referral traffic is the weakest of the three as a primary metric, because AI answers frequently resolve the question without a click. It still belongs in the stack as confirmation, and AI search analytics in GA4 is the practical way to isolate it. Just do not treat a flat referral line as evidence of flat visibility.

    How the number is actually built

    A defensible AI search visibility score has four inputs. Drop any one and the number stops being comparable over time.

    A fixed prompt set. Twenty to fifty questions a buyer would genuinely type, frozen. Change the prompts and you have changed the instrument, so any movement in the score is meaningless.

    Repetition. Each prompt runs multiple times per engine, on the same day. This is the input almost every free checker skips.

    Multiple engines, scored separately. A blended cross-engine average hides the fact that engines disagree with each other more than they disagree with themselves.

    A mention definition you write down. Named in the answer body, named in a list, named in a citation label and named in a follow-up suggestion are four different events. Pick which ones count before you start.

    Honeyb (our product) ran this shape of measurement on 13 July 2026: 20 buyer prompts, three runs each, four engines via API, 240 answers total. Two results set the floor for what any tracker has to handle.

    First, engines named 4.8 to 5.2 brands per answer. That is the real shape of the opportunity. AI answers are not a ranked list with a winner, they are a shortlist of roughly five, and being on it is the outcome that matters. Chasing the top slot is chasing noise.

    Second, the top-ranked brand changed between two identical runs 28% to 44% of the time, depending on the engine.

    Top-pick change rate

    How often the top recommendation changes between identical runs

    Share of consecutive identical prompt runs where the engine's number-one recommended brand changed: Gemini 44%, Perplexity 43%, ChatGPT 35%, Claude 28%. Honeyb measurement, 13 July 2026: 20 buyer-intent prompts, 3 runs each, via API (gpt-5-mini, gemini-2.5-flash, claude-haiku-4-5, sonar).

    Brand-set overlap between runs told the same story from the other side: Claude 67%, Perplexity 61%, Gemini 54%, ChatGPT 42%. On ChatGPT, less than half the named brands survived a repeat of the identical prompt. A tracker reporting your rank on a single run is reporting a sample of one from a distribution that wide.

    Cross-engine disagreement is larger still. The same prompt produced the same top brand on 20% of ChatGPT and Perplexity pairs, and 53% of Gemini and Claude pairs. There is no single answer to "where do I rank in AI search", because there is no single ranking. We wrote up the method in more detail in how to measure AI share of voice.

    Want to see this in action?

    See how every major AI model talks about your brand. Free to start.

    Free AI Check

    What an AI search tracker must capture

    The word tracker is doing a lot of work in this market. "AI search tracker" gets 320 searches a month in the US at a keyword difficulty of 34 (DataForSEO, July 2026), and the tools ranking for it vary enormously in what they actually record. Use this as the checklist when you evaluate one.

    CapabilityWhy it mattersWhat failure looks like
    Repeat runs per promptAnswers vary run to run by designA rank that moves every login with no cause
    Per-engine breakdownEngines disagree more than they self-disagreeOne blended score that never explains itself
    Answer text stored verbatimLets you audit why a score movedA number with nothing behind it
    Mentions and citations tracked separately62% of citations do not name the brand (Semrush)Credit for sources nobody read as your name
    Competitor set in the same runYour score only means something relativeAbsolute percentages that drift with engine mood
    Sentiment or context of the mentionBeing named as the cautionary example still counts as a mentionRising visibility, falling pipeline
    Dated method notesEngines and models change under youHistoric data silently incomparable
    Prompt-level detailTells you which topics you ownA single score you cannot act on

    The last two are the ones teams regret skipping. When a model version changes, a tracker without dated method notes gives you a step change you will spend a week attributing to your own work. We covered that failure mode in why did my AI visibility drop, and it is worth interrogating a vendor's data accuracy before you sign anything annual.

    Why citation tracking and mention tracking disagree

    Run both and the two numbers will not match. That is correct behaviour, not a bug in either tool.

    Citation coverage is uneven by engine. ChatGPT and Gemini often hide their citations, while Perplexity and Claude expose theirs. In the July measurement, ChatGPT cited 445 distinct domains, Claude 194 and Perplexity 142. Any citation-only tracker is therefore measuring the engines that are willing to be measured, and inferring the rest.

    The sources themselves are also not the ones most SEO plans assume. Reddit is 40.1% of all AI citations, the single most-cited source (Semrush), and in our own run it was 71 of Perplexity's 498 citations, with YouTube at 40. Forbes was the only domain that appeared in all four engines' top citation lists. Ahrefs found that AI visibility correlates most strongly with third-party mentions and video, not on-page work, which is consistent with all of the above.

    So the practical position is: mention tracking tells you the outcome, citation tracking tells you the mechanism, and you need both to know whether a change in one caused a change in the other. If you are working the mechanism side, how to get cited by AI covers what actually moves it.

    One caution on citation share as a headline metric. Reddit's ChatGPT-citation share fell from roughly 60% to roughly 10% in a fortnight in late 2025 (Semrush). Source mixes move that fast. Build a strategy on one platform's citation share and you have built it on sand.

    What the tooling costs

    Pricing on published pages, for orientation rather than a recommendation. Otterly is $29 a month. SE Ranking is around $55. Peec is around $89. AthenaHQ is around $295. Profound is around $399 and is demo-gated, with API and white-label options. Ahrefs Brand Radar is included on Ahrefs plans, Semrush AI Visibility is an add-on with a free checker, and Scrunch is custom priced. Honeyb (our product) offers a free check.

    The price range does not map to measurement quality in any reliable way. What separates the tools is the checklist above, not the tier. A $29 tool that stores answer text and repeats runs will serve you better than a $300 one that reports a single daily rank. Our side-by-side of the category is at AI visibility tools.

    Where to start

    Freeze twenty buyer prompts. Run each of them three times on at least two engines. Record, per prompt, whether you were named, which competitors were named, and whether your domain appeared in the citations where the engine shows them. That baseline takes an afternoon and is more useful than any dashboard you buy before you have it.

    Then repeat monthly, not daily. At the variance levels above, daily readings are mostly noise, which is the case laid out in why spot-checking fails. If you want the vocabulary nailed down first, start with what is AI visibility, and if you would rather work from a structured pass over your own site, a structured AI visibility audit walks through it.

    You can get the first data point now: run a free check on your domain at /tools/ai-visibility-checker and see which prompts already name you.

    Frequently asked questions

    What counts as good AI search visibility?

    There is no universal benchmark, because scores depend entirely on your prompt set. The useful frame is relative: engines name 4.8 to 5.2 brands per answer (Honeyb measurement, 13 July 2026, 240 answers), so appearing in a decent share of answers for your category's prompts is the goal. Compare yourself to the named competitors in the same runs, not to an absolute percentage.

    How often should I measure AI search visibility?

    Monthly for reporting, with the same frozen prompt set each time. Answers vary enough run to run that daily readings mostly capture noise rather than change. The exception is around a known event, such as a large content launch or a model version update, where a before-and-after with repeated runs is worth the extra cost.

    Can I track AI search visibility without a paid tool?

    Yes, for a baseline. Run twenty fixed prompts three times each across two engines and log whether you and each competitor were named. It is a few hours of work and gives you a real number. Paid trackers earn their cost on repetition at scale, stored answer text and history, not on any measurement you cannot reproduce by hand.

    Why does my visibility score differ between two tools?

    Different prompt sets, different engines, different run counts and different definitions of a mention. Two tools can both be correct and report different numbers. Ask any vendor what their prompts are, how many times each one runs, which engines are included and whether a citation without a brand name counts as a mention.

    Is AI search visibility the same as being cited by AI?

    No. Visibility means your brand name appears in the answer a reader sees. Citation means your domain is listed as a source. Semrush found 62% of AI citations never name the brand being cited, so a page can be the source an engine reads while a competitor gets named. Track both, since they move for different reasons.

    Matiss Katanenko

    About the author

    Matiss Katanenko

    Co-founder, Honeyb

    My name is Matiss Katanenko and I co-founded Honeyb, the AI visibility platform that tracks how ChatGPT, Gemini, Claude, Perplexity and the other major AI engines talk about brands. I'm based in Riga, Latvia. Before Honeyb I spent years on the agency side running SEO and content programs for fast-growing brands across the US and Europe. That work is where I watched AI search start to compress the entire discovery channel into a four-brand short list, and decided to build the tool I wished agencies had. In my free time I'm in the sauna, on a padel court, or behind a drum kit.

    Connect on LinkedIn
    Honeyb

    Free to start

    See your brand through every major AI model.

    Run a free check in 30 seconds. The picture is usually different than you'd expect.

    ChatGPTChatGPT
    ClaudeClaude
    GeminiGemini
    PerplexityPerplexity