A knowledge cutoff is the date after which a model stopped learning from its training data. Ask it about anything that happened later and it is working from nothing, unless it can go and look. That much is widely understood. What is less understood is that the number most people quote is the wrong one, that the newest model on the shelf is frequently not the freshest, and that for anyone trying to get a brand named in AI answers, the cutoff decides which of two entirely different mechanisms has to do the work.
This piece gives the published cutoff for every current frontier model as of July 2026, taken from the vendors' own documentation rather than from secondhand roundups, and then explains the part that actually matters commercially.
The published cutoffs, July 2026
The table below is drawn from three sources: OpenAI's model documentation, Anthropic's models overview and Google's Gemini 3 developer guide. The final column is simple arithmetic, the distance between the cutoff and the end of July 2026.
| Model | Published knowledge cutoff | Age of memory | Context window |
|---|---|---|---|
| GPT-5.6 (Sol, Terra, Luna) | 16 February 2026 | 5 months | 1.05M tokens |
| Claude Opus 5 | May 2026 | 3 months | 1M tokens |
| Claude Fable 5 | January 2026 | 6 months | 1M tokens |
| Claude Sonnet 5 | January 2026 | 6 months | 1M tokens |
| Claude Haiku 4.5 | February 2025 | 17 months | 200k tokens |
| Gemini 3 Pro | January 2025 | 18 months | 1M tokens |
| Gemini 3.5 Flash | January 2025 | 18 months | 1M tokens |
| Perplexity Sonar | None published | Not applicable | Varies by model |
Two things in that table are worth stopping on. The spread between the freshest and the stalest memory is fifteen months, across models a buyer might reasonably treat as interchangeable. And Google's Gemini 3.5 Flash, a 2026 release, carries a January 2025 cutoff, which its own documentation states plainly: "Gemini 3.5 Flash has a knowledge cutoff of January 2025." Release date and cutoff date are not the same thing and they are not even loosely correlated.
The number most people quote is the wrong one
Here is the wrinkle that almost every cutoff article misses. There are two different dates, and they are not the same.
The training data cutoff is the outer edge of the material the model was trained on. The reliable knowledge cutoff is the date through which the model's knowledge is genuinely dense and dependable. Data thins out towards the end of a training run, so the last few months before the training cutoff are represented sparsely. The model has seen a little about that period and will answer confidently about it, while actually knowing much less than it does about the year before.
Anthropic is currently the only one of the three major labs to publish both numbers side by side, and the gap it discloses is not small.
| Model | Reliable knowledge cutoff | Training data cutoff | Gap |
|---|---|---|---|
| Claude Sonnet 4.6 | August 2025 | January 2026 | 5 months |
| Claude Haiku 4.5 | February 2025 | July 2025 | 5 months |
| Claude Opus 4.6 | May 2025 | August 2025 | 3 months |
| Claude Opus 5 | May 2026 | May 2026 | None |
| Claude Fable 5 | January 2026 | January 2026 | None |
Claude Sonnet 4.6 is the instructive row. Its training data runs to January 2026, and that is the figure a spec-sheet comparison would pick up. Its reliable knowledge stops in August 2025. A buyer comparing headline numbers would rate it fresher than it behaves.
OpenAI and Google publish a single date each. That does not mean their models lack the same thinning effect at the edge of training, only that the distinction is not disclosed. The practical instruction is the same either way: treat the last few months before any published cutoff as a soft zone rather than a hard line, and do not assume a model is well informed about events that fall just inside it.

Memory versus retrieval, and why only one of them is buyable
Every one of these models can now reach past its cutoff by searching the web, though what each one retrieves rather than recalls varies more than the marketing suggests. That capability is why the cutoff has stopped being a hard ceiling on what a model can tell you, and it is also why the cutoff matters more to a brand than it used to, not less.
When a model answers a question, the information can come from one of two places. It can come from memory, meaning weights laid down during training, or it can come from retrieval, meaning documents fetched at the moment you ask. If your company existed and was written about before the cutoff, you have a chance of living in memory. If you launched afterwards, memory has nothing on you, and retrieval is the only route by which your name can appear in an answer.
That distinction is not academic. Retrieval only surfaces what it can find and rank in the moment, from a small set of pages. Our own measurement of how many sources each engine actually pulls per answer shows how narrow that window is.
Sources per answer
Average sources cited per answer, by engine
Between eight and fifteen URLs per answer, and from those the model builds a recommendation. There is no long tail. You are either in that handful or you are absent, and for a post-cutoff brand there is no fallback to memory.
The engines also differ in how much of this they show you.
| Engine | Published cutoff | Answers from retrieval | Citations exposed to the reader |
|---|---|---|---|
| ChatGPT | February 2026 | Yes, when it searches | Often hidden |
| Gemini | January 2025 | Yes, via Search Grounding | Often hidden |
| Claude | Varies by model, January to May 2026 | Yes, when it searches | Shown |
| Perplexity | None published | Yes, grounded by default | Shown |
Perplexity is the outlier worth understanding. Its documentation publishes no cutoff for the Sonar models and describes them as grounded, which is a design choice rather than an oversight. A retrieval-first engine has less need to advertise the age of its memory, because memory is not what it leans on.
The citation column matters more than it looks. Where sources are hidden, you cannot tell from the answer alone whether your brand was recalled or retrieved, which is a large part of why reading ChatGPT's sources is harder than it should be.
What a stale cutoff costs a brand
Put the two halves together and the commercial picture is clear. A company founded in 2026 is invisible to Gemini 3's memory by definition, and it is invisible to Claude Haiku 4.5's memory too. Not badly represented, simply absent. The only way that company's name reaches the answer is if the engine searches, finds a page that mentions it, and decides that page is worth citing.
This reframes what the work actually is. Getting into a training set is not a strategy, because you cannot buy it, cannot schedule it, and will not know for a year whether it worked. Getting cited is a strategy, because it turns on things you can influence: whether third parties write about you, whether those pages rank for the questions buyers ask, and whether the framing around your name is accurate.
The evidence supports that ordering. Ahrefs' analysis of AI visibility found it correlates most strongly with third-party mentions and video rather than with on-page work. Reddit alone accounts for 40.1% of all AI citations, the single most-cited source (Semrush), which is one reason why AI models cite Reddit so heavily. And being retrieved is not the same as being named: Semrush found that 62% of AI citations never mention the brand at all, so a page can be used as a source while the company behind it stays anonymous.
There is a volatility problem sitting on top of this. SparkToro found the same query changes its answer roughly 70% of the time, with two identical queries matching the same brand list less than once in a hundred runs. A model's cutoff is a fixed, published fact. Whether retrieval names you on any given Tuesday is not. That is why spot-checking a chatbot by hand will not tell you what you need to know, and why the question worth answering is a rate over time rather than a single observation.
What to do about it
For most teams the practical response is short.
- Stop optimising for the training set. You cannot influence it on any useful timescale, and the cutoffs above show how long you would be waiting.
- Assume retrieval is doing the work. Anything about your company from the last twelve months almost certainly reaches an answer through a fetched page, not through memory. Make sure such pages exist, are current, and state the facts you want repeated.
- Check what the engines currently believe. Models trained before a rebrand, a pivot or a funding round will confidently repeat the old version. A GEO audit is the structured way to find those gaps.
- Measure on a schedule, not on a hunch. Given a 70% answer-change rate, one check tells you almost nothing. A trend line tells you whether last quarter's coverage push moved anything.
Honeyb, which is our product, does the last of those: scheduled scans across the major engines, tracking how often your brand is mentioned and cited, share of voice against rivals, and the sentiment of the framing. The cutoff table above tells you which models could possibly know you from memory. Measurement tells you whether any of them actually name you.

The cutoff is a useful fact and a poor strategy. Treat it as a diagnostic, a way of knowing whether a given engine has any chance of recalling you unaided, and then put the effort into the layer you can actually move. You can see where your brand currently stands with the free AI visibility checker, which shows how often the major engines mention you today.














