Every part of a request to a language model shares one fixed budget of text. What that budget actually covers, what happens as it fills up, and why hitting it is not one single failure mode.
Key takeaways
A model’s context window is the entire amount of text it can weigh at once when producing a reply, and every part of a request draws on the same fixed budget. Anthropic’s API documentation is explicit about what counts toward it: (Anthropic, Claude API documentation — context windows) “everything in the request counts toward the context window: the system prompt, every message in messages (including tool results, images, and documents), and your tool definitions. The output… generates for the turn… counts too.” It is not a limit on your question alone. It is a shared ceiling across your question, any instructions layered underneath it, any documents attached, the whole earlier conversation, and the reply itself as it is being written.
This is why a very long document attached early in a conversation can leave noticeably less room for everything that follows — the conversation is not adding a new budget for each message, it is spending down one budget as it goes.
A context window is easy to mistake for a kind of long-term memory. It is not. The same documentation draws the distinction directly: the context window (Anthropic — context windows) “is different from the large corpus of data the language model was trained on, and instead represents a ‘working memory’ for the model.” Training happened once, before you ever opened the conversation, and shaped the model’s weights. The context window is temporary, rebuilt from scratch for every request, and holds only what is explicitly included in that request — nothing the model “remembers” from an earlier, separate conversation unless it was resupplied.
More context is not automatically better context. Anthropic’s documentation names the specific failure mode directly: (Anthropic — context windows) “as token count grows, accuracy and recall degrade, a phenomenon known as context rot”. That is precisely the phenomenon that, in Anthropic’s words, “makes curating what’s in context just as important as how much space is available” — a model reasoning over a nearly-full window is not simply “running out of room” at some hard edge; its ability to weigh everything it has been given can already be degrading well before that edge is reached.
What happens once a request actually exceeds the window differs from what happens on a length overrun in the reply. Anthropic’s stop-reason reference lists them as separate conditions: (Anthropic — stop reasons) a max_tokens result means “the response reached your max_tokens limit” — an output cap, covered on its own in why AI answers get cut off — while a distinct model_context_window_exceeded result means “the response filled the model’s context window”, and the guidance for that case is blunt: “treat the response as truncated.” The two happen for different reasons and call for different fixes, which is precisely why a product needs to check which one actually occurred rather than treating every cut-off answer the same way.
A very long conversation in a consumer chat app does not simply fail once the underlying window is full. Anthropic’s documentation notes that (Anthropic — context windows) “chat interfaces such as claude.ai can also manage the context window on a rolling ‘first in, first out’ basis” — meaning the oldest turns quietly drop out of what the model can see as new ones are added, without necessarily telling you it happened. That single mechanism is also the most common explanation for a model appearing to forget something you said earlier in a long conversation, a topic worth its own explanation — see why AI forgets what you told it.
Published window sizes vary by provider and by model, and they change often enough that a specific number is a poor thing to build a mental model around — the mechanism above is what stays true regardless of the current figure. What is worth fixing in your head is the shape of the limit: one shared budget, degrading usefulness well before the hard edge, and a silent trimming behaviour in long-running chat products that a raw API call does not have by default.
Anthropic’s own developer glossary summarizes the concept in almost the same words as its longer documentation, which is worth noting because two separately maintained pages agreeing this closely suggests the framing is stable rather than a one-off phrasing choice: (Anthropic, developer glossary) the context window “refers to the amount of text a language model can look back on and reference when generating new text… A larger context window allows the model to process and respond to more complex and lengthy prompts, while a smaller context window may limit the model’s ability to handle longer prompts or maintain coherence over extended conversations.” The glossary entry also flags that “these concepts are not unique to” any single company’s models — the mechanism, not the specific number attached to it, is the general property worth understanding.
Anthropic’s longer context-window documentation gives one concrete figure as of this writing, useful only as an illustration of scale rather than a number to design a system around: (Anthropic — context windows) several of its models currently offer “a 1M-token context window,” with others at “a 200k-token context window.” A million tokens is roughly on the order of a long novel’s worth of text, shared across everything in a single request — a genuinely large budget, and still a budget with the same edge-of-window degradation and the same shared-resource behaviour described above, not a qualitatively different kind of limit.
Related, within this hub: why AI forgets what you told it, why AI answers get cut off, and how retrieval grounds an AI answer. For connecting a model to documents that live outside the conversation, see custom AI solutions.
No. It means more text can be included in a single request. Anthropic's own documentation names a real cost to filling it — accuracy and recall degrading as token count grows, a phenomenon it calls context rot — so more included text is not automatically better reasoning.
They are separate conditions with separate causes. An output limit (max_tokens) is a cap on how long the reply itself is allowed to be. The context window filling up means the combined input and output no longer fit in the model's total budget, and the response should be treated as truncated.
Not reliably. Some chat interfaces manage long conversations on a rolling, first-in-first-out basis, quietly dropping the oldest turns as new ones are added, so material from early in a long conversation can fall out of what the model can currently see.
A short call is enough to map your data and conversations against what the window can realistically carry.