AI search visibility is how often AI assistants name your brand in the answers they give to buyer questions. It is measured as a rate, not a position: you run a fixed set of prompts repeatedly across ChatGPT, Gemini, Claude and Perplexity, count the answers that name you, and divide. A single check tells you almost nothing, because the same query changes its answer roughly 70% of the time (SparkToro).
That last point is why this metric is harder than it looks. Most published visibility numbers come from one run of one prompt on one engine, which is a coin toss dressed as a measurement. The rest of this page covers what the number is made of, what separates it from two things it gets confused with, and what an AI search tracker has to capture before its output is worth acting on.
Three different metrics people call the same thing
Visibility, citations and referral traffic are measured differently, break for different reasons, and are fixed by different work. Conflating them is the most common reason a team spends a quarter on the wrong problem.
| Metric | Question it answers | Where it is measured | Typical fix when it is low |
|---|---|---|---|
| AI search visibility | Is your brand named in the answer text? | Prompt runs across engines, repeated | Third-party mentions, reviews, comparison coverage |
| AI citations | Is your domain listed as a source? | Citation lists on answers that expose them | Content that answers the question directly, crawlable and current |
| AI referral traffic | Did anyone actually click through? | GA4 and server logs | Nothing on its own; it is downstream of the other two |
The gap between the first two is not theoretical. Semrush found that 62% of AI citations never name the brand being cited. Your domain can be the source an engine reads while a competitor is the name the reader walks away with. Track citations only and you will report progress that no buyer experienced. Track mentions only and you will not know which pages are doing the work.
Referral traffic is the weakest of the three as a primary metric, because AI answers frequently resolve the question without a click. It still belongs in the stack as confirmation, and AI search analytics in GA4 is the practical way to isolate it. Just do not treat a flat referral line as evidence of flat visibility.
How the number is actually built
A defensible AI search visibility score has four inputs. Drop any one and the number stops being comparable over time.
A fixed prompt set. Twenty to fifty questions a buyer would genuinely type, frozen. Change the prompts and you have changed the instrument, so any movement in the score is meaningless.
Repetition. Each prompt runs multiple times per engine, on the same day. This is the input almost every free checker skips.
Multiple engines, scored separately. A blended cross-engine average hides the fact that engines disagree with each other more than they disagree with themselves.
A mention definition you write down. Named in the answer body, named in a list, named in a citation label and named in a follow-up suggestion are four different events. Pick which ones count before you start.
Honeyb (our product) ran this shape of measurement on 13 July 2026: 20 buyer prompts, three runs each, four engines via API, 240 answers total. Two results set the floor for what any tracker has to handle.
First, engines named 4.8 to 5.2 brands per answer. That is the real shape of the opportunity. AI answers are not a ranked list with a winner, they are a shortlist of roughly five, and being on it is the outcome that matters. Chasing the top slot is chasing noise.
Second, the top-ranked brand changed between two identical runs 28% to 44% of the time, depending on the engine.
Top-pick change rate
How often the top recommendation changes between identical runs
Brand-set overlap between runs told the same story from the other side: Claude 67%, Perplexity 61%, Gemini 54%, ChatGPT 42%. On ChatGPT, less than half the named brands survived a repeat of the identical prompt. A tracker reporting your rank on a single run is reporting a sample of one from a distribution that wide.
Cross-engine disagreement is larger still. The same prompt produced the same top brand on 20% of ChatGPT and Perplexity pairs, and 53% of Gemini and Claude pairs. There is no single answer to "where do I rank in AI search", because there is no single ranking. We wrote up the method in more detail in how to measure AI share of voice.
Want to see this in action?
See how every major AI model talks about your brand. Free to start.
What an AI search tracker must capture
The word tracker is doing a lot of work in this market. "AI search tracker" gets 320 searches a month in the US at a keyword difficulty of 34 (DataForSEO, July 2026), and the tools ranking for it vary enormously in what they actually record. Use this as the checklist when you evaluate one.
| Capability | Why it matters | What failure looks like |
|---|---|---|
| Repeat runs per prompt | Answers vary run to run by design | A rank that moves every login with no cause |
| Per-engine breakdown | Engines disagree more than they self-disagree | One blended score that never explains itself |
| Answer text stored verbatim | Lets you audit why a score moved | A number with nothing behind it |
| Mentions and citations tracked separately | 62% of citations do not name the brand (Semrush) | Credit for sources nobody read as your name |
| Competitor set in the same run | Your score only means something relative | Absolute percentages that drift with engine mood |
| Sentiment or context of the mention | Being named as the cautionary example still counts as a mention | Rising visibility, falling pipeline |
| Dated method notes | Engines and models change under you | Historic data silently incomparable |
| Prompt-level detail | Tells you which topics you own | A single score you cannot act on |
The last two are the ones teams regret skipping. When a model version changes, a tracker without dated method notes gives you a step change you will spend a week attributing to your own work. We covered that failure mode in why did my AI visibility drop, and it is worth interrogating a vendor's data accuracy before you sign anything annual.
Why citation tracking and mention tracking disagree
Run both and the two numbers will not match. That is correct behaviour, not a bug in either tool.
Citation coverage is uneven by engine. ChatGPT and Gemini often hide their citations, while Perplexity and Claude expose theirs. In the July measurement, ChatGPT cited 445 distinct domains, Claude 194 and Perplexity 142. Any citation-only tracker is therefore measuring the engines that are willing to be measured, and inferring the rest.
The sources themselves are also not the ones most SEO plans assume. Reddit is 40.1% of all AI citations, the single most-cited source (Semrush), and in our own run it was 71 of Perplexity's 498 citations, with YouTube at 40. Forbes was the only domain that appeared in all four engines' top citation lists. Ahrefs found that AI visibility correlates most strongly with third-party mentions and video, not on-page work, which is consistent with all of the above.
So the practical position is: mention tracking tells you the outcome, citation tracking tells you the mechanism, and you need both to know whether a change in one caused a change in the other. If you are working the mechanism side, how to get cited by AI covers what actually moves it.
One caution on citation share as a headline metric. Reddit's ChatGPT-citation share fell from roughly 60% to roughly 10% in a fortnight in late 2025 (Semrush). Source mixes move that fast. Build a strategy on one platform's citation share and you have built it on sand.
What the tooling costs
Pricing on published pages, for orientation rather than a recommendation. Otterly is $29 a month. SE Ranking is around $55. Peec is around $89. AthenaHQ is around $295. Profound is around $399 and is demo-gated, with API and white-label options. Ahrefs Brand Radar is included on Ahrefs plans, Semrush AI Visibility is an add-on with a free checker, and Scrunch is custom priced. Honeyb (our product) offers a free check.
The price range does not map to measurement quality in any reliable way. What separates the tools is the checklist above, not the tier. A $29 tool that stores answer text and repeats runs will serve you better than a $300 one that reports a single daily rank. Our side-by-side of the category is at AI visibility tools.
Where to start
Freeze twenty buyer prompts. Run each of them three times on at least two engines. Record, per prompt, whether you were named, which competitors were named, and whether your domain appeared in the citations where the engine shows them. That baseline takes an afternoon and is more useful than any dashboard you buy before you have it.
Then repeat monthly, not daily. At the variance levels above, daily readings are mostly noise, which is the case laid out in why spot-checking fails. If you want the vocabulary nailed down first, start with what is AI visibility, and if you would rather work from a structured pass over your own site, a structured AI visibility audit walks through it.
You can get the first data point now: run a free check on your domain at /tools/ai-visibility-checker and see which prompts already name you.





