A prospect comparing tools for the job you do opens the Gemini app, or simply types the question into Google and reads the AI answer that now sits above the links, and Gemini returns a confident paragraph that names two or three companies and a reason to prefer each. If your company is not among them, you will never learn it from your own desk, because you did not ask the question and Gemini did not tell you it had been asked. That is the awkward shape of visibility in Google's AI: the moment that decides whether a buyer even considers you happens in a session you cannot see, phrased in words you did not choose, drawing on pages you may not have written.
So here is the case this guide makes before it reaches the how. You cannot rank-track your brand in Gemini the way you track a Google position, because Gemini has no fixed rank to track: it grounds each answer in a live Google search (grounding, in Google's own term, meaning it stitches the reply together from whatever pages rank at that moment) and then resamples, so the same question returns a different answer roughly 70% of the time (SparkToro). The only method that survives that volatility is to fix the handful of questions your buyers actually ask, run them repeatedly across days and accounts, watch both of the places Gemini can name you, and record four things every run: whether you appear, where you sit, how you are described, and which page the claim came from. Do it by hand until the sampling outgrows the half-hour it deserves, then automate it. Everything below is that method, and the reason each step is not optional.
Why a single look at Gemini is not a measurement
~70%
of repeat questions to an AI engine return a different answer
Gemini resamples too, so one reading is a sample of one. SparkToro.
<1 in 100
chance two identical questions return the same brand list
There is no fixed Gemini position to write down. SparkToro.
2.0%
of Gemini's citations come from community sources, the lowest of the big four
Gemini leans on ranking web pages, not Reddit threads. Profound.
Gemini has no rank to write down
The phrase rank tracking carries a promise inherited from the old web, where a keyword sat at position four on Tuesday and you could open a tool on Friday to see whether it had moved, because the ranking held still long enough to compare. Gemini breaks that promise in two ways at once. It resamples, so the same question put to it changes its answer roughly 70% of the time and two identical questions return the same list of brands in the same order less than one time in a hundred (SparkToro), which means the very thing you would write down, your position, is not stable enough to be a number. And it grounds, pulling a fresh set of web pages into each reply, so the answer drifts as the search results beneath it drift. The practical consequence is blunt: if you check once and see your name, you have learned what one buyer saw in one session, not where you stand, and our explainer on why the same AI question does not give everyone the same answer sets out the machinery behind the shuffle.
The two places a buyer meets your brand in Gemini
Before you can measure anything you have to be clear about where the measuring happens, because Gemini is not one surface but two, and a brand can win in one while quietly losing in the other. The first is the Gemini app and its assistant, where a user asks a question in conversation. The second is inside Google Search itself, where the same underlying model now writes the AI Overview that sits above the links and powers the fuller AI Mode, so a buyer who never opens the Gemini app still meets Gemini's judgement the instant they search. Tracking your brand in Gemini means watching both, because the answer, the sources, and even whether you appear at all can differ between them.
| Where Gemini names you | What the buyer is doing | How to check it |
|---|---|---|
| Gemini app and assistant | Asking a question in conversation, usually logged in | Run your buyer questions in a fresh, logged-out session so your own history does not flatter the answer |
| AI Overviews in Search | Typing a query and reading the AI summary above the links | Search the same buyer questions and read the Overview, noting which brands it names and links to |
| AI Mode in Search | Asking a longer, follow-up-style question inside Google | Repeat the questions in AI Mode, which can draw on a wider set of pages than the shorter Overview |
Read those three rows as different shop windows onto the same model, because a prospect who asks the app gets a conversation while a prospect who searches gets a summary, and Google has been steadily moving the second group, much the larger one, from ten blue links towards an answer it writes itself. We keep a fuller account of how the search side behaves in our piece on Google AI Mode and Gemini, and of the tooling built to watch the Overviews specifically in our guide to AI Overview trackers.
What monitoring actually has to measure
Once you accept that a single answer is noise, the job stops being do I appear and becomes four separate measurements, each of which Gemini makes awkward in its own way. Getting the four straight is what separates a real monitoring routine from a nervous habit of searching your own name.
| Signal | What it tells you | Why Gemini makes it hard to read |
|---|---|---|
| Appearance | Whether Gemini names you at all for a buyer's question | The answer changes about 70% of the time (SparkToro), so a single yes or no is a coin toss, not a finding |
| Position | Whether you are named first or buried beneath rivals | Two identical questions match the same brand list under 1 in 100 (SparkToro), so there is no fixed rank to log |
| Tone | Whether the sentence around your name helps or harms | 62% of AI citations never name the brand (Semrush), so the framing is written by pages you did not author |
| Source | Which page the claim was actually built from | Gemini often answers without showing its citations, unlike Perplexity and Claude (Semrush) |
The fourth row is the one that quietly defeats most home-made tracking, because appearance and position you can at least eyeball, and tone you can score with a rubric, as our method for tracking AI model sentiment sets out, but the source, the page that taught Gemini to call you the expensive option, is both the thing you most need and the thing Gemini is least willing to show.
What Gemini leans on, and why it changes the fix
Community citation share
Community citation share by AI engine
How you fix a bad answer depends on what the engine was reading when it wrote it, and here Gemini differs sharply from the engine most founders worry about first. Where Perplexity draws close to a fifth of its citations from community and social sources such as Reddit and forums, Gemini draws about one in fifty (Profound), the lowest of the big four, which tells you something useful and slightly counter-intuitive: winning an argument on Reddit shifts Perplexity far more than it shifts Gemini, and Gemini's answer is likelier to reflect the pages that already rank in ordinary Google search. For a brand trying to change what Gemini says, that points the effort back towards the third-party coverage and web pages Google ranks, which sits neatly with Ahrefs' finding that AI visibility correlates most with third-party mentions and video rather than with anything on your own site.
Reading a trail Gemini would rather hide
The hardest part of the job is tracing a description you dislike back to the page that caused it, and Gemini is among the least cooperative engines for the task. Gemini and ChatGPT frequently answer without showing their working, while Perplexity and Claude tend to list the sources they drew on (Semrush), which has a practical workaround: when Gemini hands you a verdict but no trail, put the same question to Perplexity and read its sources as a proxy for the kind of page Gemini most likely leaned on, then confirm the guess against what actually ranks in a plain Google search for that question, since grounding means Gemini was reading much the same web.
| Engine | Shows its sources? | What that means for tracking |
|---|---|---|
| Gemini | Often not | You see the verdict but rarely the pages behind it, so tracing tone means hunting for the source yourself |
| ChatGPT | Often not | The same limitation; the answer arrives without a reliable trail to follow |
| Perplexity | Usually yes | Citations are listed, so you can go straight to the page that set the tone |
| Claude | Usually yes | When it searches the web it exposes what it drew on |
This is why the fix for a Gemini problem is almost never a change to your own homepage, and our fuller treatment of whether AI engines cite their sources explains why the trail runs cold so often and what to do when it does.
The manual routine that survives the shuffle
A disciplined manual routine gets you a long way before you spend anything, provided you treat it as sampling rather than searching. The protocol below is the whole of it, and each step exists to cancel one of the ways a casual check quietly lies to you.
| Step | What to do | Why it matters |
|---|---|---|
| Fix the questions | Write the five to ten questions a buyer types when choosing, not your brand name | You want the answers customers see when they are deciding, not a vanity search for yourself |
| Strip the account | Run each question logged out, personalisation off, across the app, AI Overviews and AI Mode | A logged-in Google account feeds Gemini your history and flatters the result |
| Run repeats | Ask each question several times, spread across different days | One run is a sample of one against roughly 70% volatility (SparkToro) |
| Record four things | Log appearance, position, tone and any visible source for every run and every surface | A figure you cannot compare next week is not a measurement |
| Average, then judge | Treat a change as real only when it holds across runs or shows on another engine | Averaging is what separates a genuine decline from ordinary noise |
The step people skip is the dull one in the middle, running each question more than once, and it is exactly the step that makes the exercise worth doing, because a brand Gemini names in three runs out of ten and one it names in nine out of ten look identical if you only ever ask once. Spread the runs across days rather than firing them off in one sitting, and keep the wording frozen, because changing a single verb changes what Gemini grounds against, and you will no longer be comparing like with like.
When to stop tracking Gemini by hand
The manual routine has a natural expiry date, and it arrives the moment the sampling stops being a weekly half-hour and becomes a part-time job. Ten questions, run several times each, across three Gemini surfaces and spread over days, is already a few hundred readings a week before you add Perplexity to triangulate the sources, at which point a person doing it by hand will quietly cut corners, ask each question once, and drift back to the vanity check the whole method was meant to replace. That is the point to automate, and a monitoring tool earns its keep by doing the repeats you would skip, holding the wording constant, and re-scanning on a schedule so a real change surfaces in days rather than whenever you next remember. The field is priced from the vendors' own pages below, with Honeyb, our own tool, listed first because this is our guide; read its caveat with the same scepticism you would bring to any other entry.
| Tool | Entry price | Best for |
|---|---|---|
| Honeyb (our tool) | Free check, then from $29/mo | Founders who want to catch, trace and re-check Gemini answers on a budget |
| Profound | About $399/mo, demo-only | Enterprises and agencies needing the deepest analytics and white-label reporting |
| AthenaHQ | About $295/mo, free 10-minute audit | Teams wanting a fast read before they commit to a contract |
| Peec | About $89/mo | Small teams tracking a defined set of rivals |
| Otterly | $29/mo | The lowest-cost way to start watching across the main engines |
| SE Ranking | AI module about $55/mo | Teams already working in SE Ranking for their SEO |
| Semrush AI Visibility | Add-on to a Semrush plan, with a free checker | Existing Semrush users adding an AI lens |
The honest split is the one that holds across the whole category: a founder or lean team should start with a free check and, if the answers warrant it, a low-cost monitor that shows its sources and re-scans on a schedule, while the deepest analytics and the portability to carry numbers into a client report belong to Profound and are priced to match. We keep fuller, tested comparisons in our roundups of the best AI visibility tools and the tools built for monitoring brands in chat engines, both of which include Honeyb and judge it on the same criteria as everyone else.
What to do this week
So, the verdict, because a method without one is just a checklist. If you have never looked, do not start by buying anything; write down the five questions your buyers actually ask and run each one three times, logged out, across the Gemini app and the AI answer in Google Search, and you will learn more in that half-hour than in a month of typing your own name into the box. If those runs show you missing rather than mischaracterised, the problem is visibility rather than tone, and our guide on why your brand is not showing up in AI answers is the better next read; if they show you appearing but described in words you would never choose, that is a source problem living on a page you do not own, and the fix is to change the source and re-sample, not to refresh the tab and hope. The same discipline applied to the other big engine is laid out in our guide to tracking your brand in ChatGPT, and the fastest way to see today's Gemini answers across your own category, sources and all, is to run the free AI visibility check before you decide any of it is worth automating.













