An AI visibility audit measures how often AI assistants name your brand when buyers ask the questions that lead to a purchase. You run a fixed set of buyer prompts across several engines, repeat each prompt multiple times, record which brands get named and which domains get cited, then compute your share of answers. The whole procedure takes a marketer about half a day and costs nothing but time.
The reason it needs a procedure rather than a few spot checks is variance. The same AI query changes its answer roughly 70% of the time (SparkToro). Anything you measure once, you have not measured.
What the audit produces
Three numbers, and they are the only three that matter at the end of a first audit.
Share of answers. Of all the answers your prompt set generated, what percentage named your brand at all. This is the headline. It is the AI-era equivalent of share of voice, and it is covered in more depth in how to measure AI share of voice.
Average position. When you are named, are you first in the list or seventh. Engines named 4.8 to 5.2 brands per answer in Honeyb's (our product) July 2026 measurement, so being named is not the same as being recommended.
Cited domains. Which third-party pages the engine leaned on to build its answer. This is your action list, because AI visibility correlates most strongly with third-party mentions and video, not on-page work (Ahrefs).
Step 1: Build the prompt set
Write 20 to 30 prompts a real buyer would type. Not keywords. Full questions, in the phrasing a person actually uses.
Cover four intent bands, roughly evenly:
- Category discovery. "What is the best tool for X", "who are the leading providers of X" - Constrained comparison. "Best X for a team of five under $100 a month", "X for regulated industries" - Alternatives. "Alternatives to [the market leader]", "what should I use instead of X" - Problem-first. The symptom your product solves, with no category word in it at all
Include your brand name in exactly zero of them. A prompt that names you tests whether the engine knows you exist, not whether it recommends you. If you want a starting library, 21 buyer questions AI gets asked is a usable base to adapt.
Freeze the list before you run anything. Editing prompts mid-audit destroys comparability with your next audit.
Step 2: Choose engines
Four is the practical set: ChatGPT, Gemini, Perplexity and Claude. They behave differently enough that three of them will not stand in for the fourth. In Honeyb's (our product) measurement of 13 July 2026, the same prompt produced the same top brand on only 20% of ChatGPT and Perplexity pairs, and 53% of Gemini and Claude pairs.
Citation transparency also varies. ChatGPT and Gemini often hide their citations, while Perplexity and Claude expose theirs. If your budget of effort only stretches to two engines, run one from each group so you still get a citation list.
Use fresh sessions or logged-out windows. Personalisation and chat memory will quietly contaminate the sample otherwise.
Step 3: Decide the run count
This is where most manual audits fail. A single run per prompt is noise, not a reading.
In the 13 July 2026 Honeyb (our product) measurement, 20 buyer prompts were run three times each across four engines via API, producing 240 answers. The top-ranked brand changed between identical runs on Gemini 44% of the time, Perplexity 43%, ChatGPT 35%, and Claude 28%. Brand-set overlap between runs was Claude 67%, Perplexity 61%, Gemini 54%, ChatGPT 42%.
Top-pick change rate
How often the top recommendation changes between identical runs
Read that Gemini figure literally. If you run a prompt once on Gemini and record the winner, there is close to a coin-flip chance a second run would have given you a different winner. One run tells you what happened once.
The practical guidance:
| Runs per prompt per engine | What it supports | What it cannot support |
|---|---|---|
| 1 | Nothing. A single sample from a distribution that moves | Any claim at all |
| 3 | A first read on whether you appear at all | Ranking yourself against a close competitor |
| 5 | A share-of-answers number you can report internally | Detecting a change smaller than about 20 points |
| 10+ | Tracking movement over time on a small prompt set | Nothing much beyond this for a manual audit |
Three runs is the floor for a first audit: 25 prompts, 4 engines, 3 runs is 300 answers, which is a long but finishable day. If that is too much, cut prompts before you cut runs. Fifteen prompts run three times beats forty-five prompts run once, every time. The reasoning is set out in why spot-checking fails.
Want to see this in action?
See how every major AI model talks about your brand. Free to start.
Step 4: Record every run in a fixed schema
Record as you go, in one sheet, one row per run. Retrofitting structure onto screenshots afterwards does not work.
| Field | Type | Example | Why it is there |
|---|---|---|---|
| run_id | integer | 147 | Unique key, makes duplicates visible |
| date | ISO date | 2026-07-20 | Answers drift; undated rows are worthless later |
| prompt_id | string | P07 | Ties runs of the same prompt together |
| prompt_text | string | "best CRM for a 5 person agency" | The frozen wording, verbatim |
| intent_band | enum | comparison | Lets you see which intents you lose |
| engine | enum | perplexity | The four-engine split |
| run_number | integer | 2 | 1 to N, so variance is computable |
| brands_named | list | Acme, Bolt, Corvid | In order of appearance |
| brand_count | integer | 5 | Sanity check against the 4.8 to 5.2 norm |
| our_position | integer or null | 3 | null means not named, which is data |
| domains_cited | list | forbes.com, reddit.com | Blank where the engine hides them |
| answer_text | string | full text | For re-reading claims about you later |
Two fields earn their keep more than the rest. `our_position` as null rather than blank forces you to distinguish "not named" from "forgot to check". And `answer_text` is what you go back to when the audit says your share dropped and you need to know what changed, a diagnosis walked through in why AI visibility drops.
Step 5: Compute share of answers
Share of answers is your named-count divided by total answers, per engine and overall.
Compute it three ways. Overall gives you the headline. Per engine tells you where the gap is, and the gaps are usually wide. Per intent band tells you what kind of buyer never hears about you, which is the most actionable cut of the three.
Then do the same for your three closest competitors from the same rows. You already recorded every brand named, so this costs nothing extra and turns an absolute number into a ranking.
Step 6: Log the cited domains
Pull every domain from `domains_cited` and count frequency. This is your earned-media target list.
Expect Reddit to dominate. Reddit is 40.1% of all AI citations and the single most-cited source (Semrush). In the Honeyb (our product) measurement, Reddit was 71 of Perplexity's 498 citations at 14%, and YouTube 40 at 8%. Citation breadth varied sharply: ChatGPT cited 445 distinct domains, Claude 194, Perplexity 142. Forbes was the only domain that appeared in all four engines' top citation lists.
Also note that 62% of AI citations never name the brand being cited (Semrush). A page can feed the answer without your name surviving into it, which is why the cited-domain list and the brands-named list have to be tracked separately. How to get cited by AI covers what to do with the list once you have it.
Step 7: Benchmark the result
Share of answers has no universal pass mark, but the bands below are how to read a number in a category with five or so credible players.
| Share of answers | Reading | Sensible next move |
|---|---|---|
| 0% | Invisible. Engines do not associate you with the category | Third-party coverage and category pages; nothing on-page will fix this |
| 1 to 15% | Present but incidental. You appear when the list runs long | Target the intent bands where you scored zero |
| 16 to 35% | Established. You are a normal member of the consideration set | Work on position, not presence |
| 36 to 60% | Strong. You are a default answer for most of your prompt set | Defend it; measure monthly |
| Above 60% | Category-defining, or your prompt set is too narrow | Re-check the prompts for wording that favours you |
The last row is not a joke. The commonest way to get a flattering audit is to write prompts that describe your product rather than the buyer's problem.
Step 8: Set the re-measure cadence
Monthly, on the same frozen prompt set, same engines, same run count. Quarterly is too slow for a channel that moves this fast: Reddit's ChatGPT-citation share fell from around 60% to around 10% inside a fortnight in late 2025 (Semrush).
Only change the prompt set once a year, and when you do, run both the old and new sets once so you can bridge the series.
When to automate instead
The manual audit is the right first move. It is free, it teaches you how the engines actually behave, and it produces a defensible baseline.
It stops being the right move at the point where the arithmetic turns against you. Twenty-five prompts across four engines at five runs each is 500 answers a month, by hand, forever, with the run-to-run variance above meaning a shortcut month invalidates the series. That is where a tool is worth paying for, and the entry prices are modest: Otterly at $29/mo, SE Ranking at around $55/mo, Peec at around $89/mo, AthenaHQ at around $295/mo, Profound at around $399/mo, Scrunch on custom pricing, plus Ahrefs Brand Radar on existing Ahrefs plans and Semrush AI Visibility as an add-on with a free checker.
Honeyb (our product) sits in this set and runs the procedure above on a schedule, with a free check to start. We would rather you ran the manual audit first, because a team that has seen the variance with its own eyes reads any tool's dashboard more sceptically, and that is the correct way to read it. The AI visibility checker guide compares the tooling, the GEO audit checklist covers the on-site half of the work, and AI visibility data accuracy explains what any tool can and cannot know.
Run the eight steps this week and you will have a baseline nobody in your category has. If you want the shortcut version first, the free AI visibility check runs a version of this measurement on your domain and gives you a starting number to argue with.





