All Articles
    AI Visibility
    Published July 20, 20268 min read

    How to Run an AI Visibility Audit: A Step-by-Step Procedure

    An AI visibility audit measures how often AI assistants name your brand when buyers ask questions. Here is the full procedure, including the data schema, the sample size that makes a reading trustworthy, and how to read the result.

    Matiss Katanenko

    Matiss Katanenko

    Co-founder, Honeyb

    How to Run an AI Visibility Audit: A Step-by-Step Procedure

    An AI visibility audit measures how often AI assistants name your brand when buyers ask the questions that lead to a purchase. You run a fixed set of buyer prompts across several engines, repeat each prompt multiple times, record which brands get named and which domains get cited, then compute your share of answers. The whole procedure takes a marketer about half a day and costs nothing but time.

    The reason it needs a procedure rather than a few spot checks is variance. The same AI query changes its answer roughly 70% of the time (SparkToro). Anything you measure once, you have not measured.

    What the audit produces

    Three numbers, and they are the only three that matter at the end of a first audit.

    Share of answers. Of all the answers your prompt set generated, what percentage named your brand at all. This is the headline. It is the AI-era equivalent of share of voice, and it is covered in more depth in how to measure AI share of voice.

    Average position. When you are named, are you first in the list or seventh. Engines named 4.8 to 5.2 brands per answer in Honeyb's (our product) July 2026 measurement, so being named is not the same as being recommended.

    Cited domains. Which third-party pages the engine leaned on to build its answer. This is your action list, because AI visibility correlates most strongly with third-party mentions and video, not on-page work (Ahrefs).

    Step 1: Build the prompt set

    Write 20 to 30 prompts a real buyer would type. Not keywords. Full questions, in the phrasing a person actually uses.

    Cover four intent bands, roughly evenly:

    - Category discovery. "What is the best tool for X", "who are the leading providers of X" - Constrained comparison. "Best X for a team of five under $100 a month", "X for regulated industries" - Alternatives. "Alternatives to [the market leader]", "what should I use instead of X" - Problem-first. The symptom your product solves, with no category word in it at all

    Include your brand name in exactly zero of them. A prompt that names you tests whether the engine knows you exist, not whether it recommends you. If you want a starting library, 21 buyer questions AI gets asked is a usable base to adapt.

    Freeze the list before you run anything. Editing prompts mid-audit destroys comparability with your next audit.

    Step 2: Choose engines

    Four is the practical set: ChatGPT, Gemini, Perplexity and Claude. They behave differently enough that three of them will not stand in for the fourth. In Honeyb's (our product) measurement of 13 July 2026, the same prompt produced the same top brand on only 20% of ChatGPT and Perplexity pairs, and 53% of Gemini and Claude pairs.

    Citation transparency also varies. ChatGPT and Gemini often hide their citations, while Perplexity and Claude expose theirs. If your budget of effort only stretches to two engines, run one from each group so you still get a citation list.

    Use fresh sessions or logged-out windows. Personalisation and chat memory will quietly contaminate the sample otherwise.

    Step 3: Decide the run count

    This is where most manual audits fail. A single run per prompt is noise, not a reading.

    In the 13 July 2026 Honeyb (our product) measurement, 20 buyer prompts were run three times each across four engines via API, producing 240 answers. The top-ranked brand changed between identical runs on Gemini 44% of the time, Perplexity 43%, ChatGPT 35%, and Claude 28%. Brand-set overlap between runs was Claude 67%, Perplexity 61%, Gemini 54%, ChatGPT 42%.

    Top-pick change rate

    How often the top recommendation changes between identical runs

    Share of consecutive identical prompt runs where the engine's number-one recommended brand changed: Gemini 44%, Perplexity 43%, ChatGPT 35%, Claude 28%. Honeyb measurement, 13 July 2026: 20 buyer-intent prompts, 3 runs each, via API (gpt-5-mini, gemini-2.5-flash, claude-haiku-4-5, sonar).

    Read that Gemini figure literally. If you run a prompt once on Gemini and record the winner, there is close to a coin-flip chance a second run would have given you a different winner. One run tells you what happened once.

    The practical guidance:

    Runs per prompt per engineWhat it supportsWhat it cannot support
    1Nothing. A single sample from a distribution that movesAny claim at all
    3A first read on whether you appear at allRanking yourself against a close competitor
    5A share-of-answers number you can report internallyDetecting a change smaller than about 20 points
    10+Tracking movement over time on a small prompt setNothing much beyond this for a manual audit

    Three runs is the floor for a first audit: 25 prompts, 4 engines, 3 runs is 300 answers, which is a long but finishable day. If that is too much, cut prompts before you cut runs. Fifteen prompts run three times beats forty-five prompts run once, every time. The reasoning is set out in why spot-checking fails.

    Want to see this in action?

    See how every major AI model talks about your brand. Free to start.

    Free AI Check

    Step 4: Record every run in a fixed schema

    Record as you go, in one sheet, one row per run. Retrofitting structure onto screenshots afterwards does not work.

    FieldTypeExampleWhy it is there
    run_idinteger147Unique key, makes duplicates visible
    dateISO date2026-07-20Answers drift; undated rows are worthless later
    prompt_idstringP07Ties runs of the same prompt together
    prompt_textstring"best CRM for a 5 person agency"The frozen wording, verbatim
    intent_bandenumcomparisonLets you see which intents you lose
    engineenumperplexityThe four-engine split
    run_numberinteger21 to N, so variance is computable
    brands_namedlistAcme, Bolt, CorvidIn order of appearance
    brand_countinteger5Sanity check against the 4.8 to 5.2 norm
    our_positioninteger or null3null means not named, which is data
    domains_citedlistforbes.com, reddit.comBlank where the engine hides them
    answer_textstringfull textFor re-reading claims about you later

    Two fields earn their keep more than the rest. `our_position` as null rather than blank forces you to distinguish "not named" from "forgot to check". And `answer_text` is what you go back to when the audit says your share dropped and you need to know what changed, a diagnosis walked through in why AI visibility drops.

    Step 5: Compute share of answers

    Share of answers is your named-count divided by total answers, per engine and overall.

    Compute it three ways. Overall gives you the headline. Per engine tells you where the gap is, and the gaps are usually wide. Per intent band tells you what kind of buyer never hears about you, which is the most actionable cut of the three.

    Then do the same for your three closest competitors from the same rows. You already recorded every brand named, so this costs nothing extra and turns an absolute number into a ranking.

    Step 6: Log the cited domains

    Pull every domain from `domains_cited` and count frequency. This is your earned-media target list.

    Expect Reddit to dominate. Reddit is 40.1% of all AI citations and the single most-cited source (Semrush). In the Honeyb (our product) measurement, Reddit was 71 of Perplexity's 498 citations at 14%, and YouTube 40 at 8%. Citation breadth varied sharply: ChatGPT cited 445 distinct domains, Claude 194, Perplexity 142. Forbes was the only domain that appeared in all four engines' top citation lists.

    Also note that 62% of AI citations never name the brand being cited (Semrush). A page can feed the answer without your name surviving into it, which is why the cited-domain list and the brands-named list have to be tracked separately. How to get cited by AI covers what to do with the list once you have it.

    Step 7: Benchmark the result

    Share of answers has no universal pass mark, but the bands below are how to read a number in a category with five or so credible players.

    Share of answersReadingSensible next move
    0%Invisible. Engines do not associate you with the categoryThird-party coverage and category pages; nothing on-page will fix this
    1 to 15%Present but incidental. You appear when the list runs longTarget the intent bands where you scored zero
    16 to 35%Established. You are a normal member of the consideration setWork on position, not presence
    36 to 60%Strong. You are a default answer for most of your prompt setDefend it; measure monthly
    Above 60%Category-defining, or your prompt set is too narrowRe-check the prompts for wording that favours you

    The last row is not a joke. The commonest way to get a flattering audit is to write prompts that describe your product rather than the buyer's problem.

    Step 8: Set the re-measure cadence

    Monthly, on the same frozen prompt set, same engines, same run count. Quarterly is too slow for a channel that moves this fast: Reddit's ChatGPT-citation share fell from around 60% to around 10% inside a fortnight in late 2025 (Semrush).

    Only change the prompt set once a year, and when you do, run both the old and new sets once so you can bridge the series.

    When to automate instead

    The manual audit is the right first move. It is free, it teaches you how the engines actually behave, and it produces a defensible baseline.

    It stops being the right move at the point where the arithmetic turns against you. Twenty-five prompts across four engines at five runs each is 500 answers a month, by hand, forever, with the run-to-run variance above meaning a shortcut month invalidates the series. That is where a tool is worth paying for, and the entry prices are modest: Otterly at $29/mo, SE Ranking at around $55/mo, Peec at around $89/mo, AthenaHQ at around $295/mo, Profound at around $399/mo, Scrunch on custom pricing, plus Ahrefs Brand Radar on existing Ahrefs plans and Semrush AI Visibility as an add-on with a free checker.

    Honeyb (our product) sits in this set and runs the procedure above on a schedule, with a free check to start. We would rather you ran the manual audit first, because a team that has seen the variance with its own eyes reads any tool's dashboard more sceptically, and that is the correct way to read it. The AI visibility checker guide compares the tooling, the GEO audit checklist covers the on-site half of the work, and AI visibility data accuracy explains what any tool can and cannot know.

    Run the eight steps this week and you will have a baseline nobody in your category has. If you want the shortcut version first, the free AI visibility check runs a version of this measurement on your domain and gives you a starting number to argue with.

    Frequently asked questions

    How long does an AI visibility audit take?

    About half a day to a full day for a first pass. Twenty-five prompts across four engines at three runs each is 300 answers. The recording is the slow part, not the prompting, so set the schema up before you start rather than tidying screenshots afterwards.

    How many prompts do I need for the audit to be meaningful?

    Twenty to thirty, spread across category discovery, constrained comparison, alternatives and problem-first phrasing. If the workload is too high, cut the prompt count rather than the run count. Fifteen prompts run three times gives a more trustworthy reading than forty-five prompts run once.

    Why do I get a different answer every time I run the same prompt?

    Because the engines are non-deterministic. The same AI query changes its answer roughly 70% of the time (SparkToro), and in Honeyb's (our product) 13 July 2026 measurement the top-ranked brand changed between identical runs on Gemini 44% of the time and ChatGPT 35%. This is normal engine behaviour, not a fault, and it is exactly why the audit specifies repeat runs.

    Should I include my own brand name in the audit prompts?

    No. A prompt naming your brand tests whether the engine has heard of you, which is a much lower bar than whether it recommends you unprompted. Keep the prompt set brand-free so the result reflects what an actual buyer would see.

    What is a good share of answers to aim for?

    In a category with roughly five credible players, 16 to 35% means you are an established part of the consideration set, and above 36% means you are a default answer for most of your prompt set. Anything at 0% means the engines do not yet associate you with the category, which is a third-party coverage problem rather than an on-page one.

    Matiss Katanenko

    About the author

    Matiss Katanenko

    Co-founder, Honeyb

    My name is Matiss Katanenko and I co-founded Honeyb, the AI visibility platform that tracks how ChatGPT, Gemini, Claude, Perplexity and the other major AI engines talk about brands. I'm based in Riga, Latvia. Before Honeyb I spent years on the agency side running SEO and content programs for fast-growing brands across the US and Europe. That work is where I watched AI search start to compress the entire discovery channel into a four-brand short list, and decided to build the tool I wished agencies had. In my free time I'm in the sauna, on a padel court, or behind a drum kit.

    Connect on LinkedIn
    Honeyb

    Free to start

    See your brand through every major AI model.

    Run a free check in 30 seconds. The picture is usually different than you'd expect.

    ChatGPTChatGPT
    ClaudeClaude
    GeminiGemini
    PerplexityPerplexity