You type your company name into ChatGPT, it returns a tidy paragraph that names you and says something broadly flattering, and you close the tab reassured. That reassurance is the trap this guide is written to dismantle, because the answer you just read was assembled for you, in that session, at that moment, and the next buyer who asks a similar question may see a different set of names in a different order with a different adjective attached to yours. Tracking your brand in ChatGPT is not like checking a Google ranking, where the position holds still long enough to write down and compare next week; it is closer to reading a weather system, where a single glance tells you almost nothing and only repeated readings reveal the pattern.
So here is the case this guide makes before it reaches the how. You cannot track your brand in ChatGPT with one look, because the same question returns a different answer roughly 70% of the time (SparkToro) and ChatGPT usually hides the sources behind what it says; the only method that survives that volatility is to fix a small set of the questions your buyers actually ask, strip out the personalisation that quietly flatters you, run each question several times across days and accounts, and record four separate things every run: whether you appear, where you sit, how you are described, and which page the claim came from. Do that by hand until the effort outgrows the value, then automate the sampling so a real change surfaces on its own. Everything below is that method, and the reason each step is not optional.
Why one look at ChatGPT tells you almost nothing
~70%
of repeat questions return a different answer
The reply you read once is a sample of one. SparkToro.
<1 in 100
chance two identical questions return the same brand list
There is no fixed rank in ChatGPT to write down. SparkToro.
62%
of AI citations never name the brand
The tone around your name is written by pages you did not. Semrush.
Why one look tells you almost nothing
The instinct to type your name once and trust the answer is the most common mistake here, and it fails for two reasons that compound. The first is volatility: the same question put to ChatGPT changes its answer roughly 70% of the time (SparkToro), and two identical questions return the same list of brands in the same order less than one time in a hundred, which means there is no stable position to record in the way a Google ranking gives you one. If you check once and see your name first, you have learned what one buyer saw once, not where you stand; our explainer on why ChatGPT does not give everyone the same answer sets out the machinery behind that shuffle. The second reason is closer to home, and it is that you are the worst possible person to run the check from your own laptop, because a logged-in account carries your history, your memory settings and your obvious interest in yourself into the prompt, and ChatGPT obligingly hands back a warmer answer than a cold prospect will ever see.
What tracking actually has to measure
Once you accept that a single answer is noise, the job stops being "do I appear" and becomes four separate measurements, each of which ChatGPT makes awkward in its own way. Getting the four straight is what separates a real tracking routine from a nervous habit of Googling yourself.
| Signal | What it tells you | Why it is hard to read in ChatGPT |
|---|---|---|
| Appearance | Whether ChatGPT names you at all for a buyer's question | The answer changes about 70% of the time (SparkToro), so a single yes or no is a coin toss, not a finding |
| Position | Whether you are named first or buried beneath rivals | Two identical questions match the same brand list under 1 in 100 (SparkToro), so there is no fixed rank to log |
| Tone | Whether the sentence around your name helps or harms | 62% of AI citations never name the brand (Semrush), so the tone is written by pages you did not author |
| Source | Which page the claim was actually built from | ChatGPT often hides its citations, unlike Perplexity and Claude (Semrush) |
Read the fourth row twice, because it is the one that quietly defeats most home-made tracking. Appearance and position you can at least eyeball; tone you can score with a rubric, as our method for tracking AI model sentiment sets out. But the source, the page that taught ChatGPT to call you clunky or overpriced, is the thing you most need and the thing ChatGPT is least willing to show, and without it you can see the wound without ever finding the knife.
The manual method that survives the noise
The good news is that a disciplined manual routine gets you a long way before you need to spend anything, provided you treat it as sampling rather than searching. The protocol below is the whole of it, and each step exists to cancel one of the ways a casual check lies to you.
| Step | What to do | Why it matters |
|---|---|---|
| Fix the questions | Write the five to ten questions a buyer actually types, not your brand name | You want the answers customers see when they are choosing, not a vanity search for yourself |
| Strip the memory | Run each question in a logged-out or fresh session, with personalisation and memory off | A logged-in account feeds ChatGPT your history and flatters the result |
| Run repeats | Ask each question several times, spread across different days | One run is a sample of one against roughly 70% volatility (SparkToro) |
| Record four things | Log appearance, position, tone and any visible source for every run | A figure you cannot compare next week is not a measurement |
| Average, then judge | Treat a change as real only when it holds across runs or shows on another engine | Averaging is what separates a genuine decline from ordinary noise |
The step people skip is the boring one in the middle, running each question more than once, and it is precisely the step that makes the exercise worth doing, because a brand that appears in three runs out of ten and one that appears in nine out of ten look identical if you only ever ask once. Spread the runs across days rather than firing them off in a single sitting, since answers drift as the model resamples, and keep the wording of your questions frozen, because changing a single verb changes what comes back and you will no longer be comparing like with like.
Why ChatGPT hides the ball, and where to look instead
The hardest part of the job is tracing a tone you dislike back to the page that caused it, and here the engine you are watching is the least cooperative of the lot. ChatGPT and Gemini frequently answer without showing their working, while Perplexity and Claude tend to list the sources they drew on (Semrush), which has a practical consequence for how you actually run this: when ChatGPT gives you a verdict but no trail, the fastest way to find the page behind it is to put the same question to an engine that does cite, and read Perplexity's sources as a proxy for what ChatGPT most likely leaned on.
| Engine | Shows its sources? | What that means for tracking |
|---|---|---|
| ChatGPT | Often not | You see the verdict but rarely the pages behind it, so tracing tone means hunting for the source yourself |
| Gemini | Often not | The same limitation; the answer arrives without a reliable trail to follow |
| Perplexity | Usually yes | Citations are listed, so you can go straight to the page that set the tone |
| Claude | Usually yes | When it searches the web it exposes what it drew on |
This matters because the tone you are tracking is almost never something you can edit directly. Ahrefs' analysis of what correlates with AI visibility found the strongest relationship with third-party mentions and video rather than with anything on your own site, and Semrush found that 62% of AI citations never name the brand at all, so the sentence ChatGPT writes about you is usually a paraphrase of a review, a forum thread or a comparison page you have no login to. Finding that page is the whole point of tracking the source, and our fuller treatment of whether and how ChatGPT cites its sources explains why the trail runs cold so often and what to do when it does.
When to stop tracking ChatGPT by hand
The manual routine has a natural expiry date, and it arrives the moment the sampling stops being a weekly half-hour and starts being a part-time job. Ten questions, run several times each, across days, logged and averaged, is perhaps forty readings a week for one engine; add Perplexity and Gemini to triangulate the sources and you are into the low hundreds, at which point a person doing it by hand will quietly cut corners, ask each question once, and drift back to the vanity check the whole method was meant to replace. That is the point to automate, and a monitoring tool earns its keep by doing the repeats you would skip, holding the question wording constant, and re-scanning on a schedule so a tone change surfaces in days rather than at the next time you happen to remember. We keep an honest, priced comparison of the tools built for monitoring your brand in ChatGPT, which includes Honeyb, our own tool; treat its entry with the same scepticism as any other and start from the free check rather than the paid tier.
What to do this week
So, the verdict, because a method without one is just a checklist. If you have never looked, do not start by buying anything; start by writing down the five questions your buyers actually ask and running each one three times in a logged-out session, and you will learn more from that half-hour than from a month of typing your own name into the box. If those runs show you missing rather than mischaracterised, the problem is visibility rather than tone, and our guide on why your brand is not showing up in ChatGPT is the better next read. And if the runs show you appearing but described in words you would never choose, that is a source problem living on a page you do not own, and the fix is to change the source and re-sample, not to refresh the tab and hope. The fastest way to see today's answers across the main engines, sources and all, is to run the free AI visibility check against your own category before you decide any of it is worth automating.














