A cut-off answer looks like one problem. It is usually one of two different limits, and knowing which one happened changes what actually fixes it.
Key takeaways
A reply that stops mid-sentence looks like one problem from the outside, but it can come from two genuinely different limits, and telling them apart matters for how you fix it. Anthropic’s API documentation for its Claude models lists the distinct conditions a request can end on: (Anthropic, Claude API reference — stop reasons) a max_tokens result means “the response reached your max_tokens limit” — a length cap on the reply itself, set by whoever built the product you are using. A separate model_context_window_exceeded result means something else entirely: “the response filled the model’s context window”, and the documented guidance is to “treat the response as truncated.” The first is a ceiling on how long the answer is allowed to be. The second is the same shared budget covered in what a context window actually limits running out mid-generation, because the input — your prompt, any documents, the conversation so far — already used most of it.
A model generating one token at a time, as described in how ChatGPT generates an answer, has no built-in plan for how long its own reply should be — it discovers the end of a natural response the same way it discovers everything else, one token at a time. Left with no outer limit, that loop could in principle run for a very long time on a single request, consuming compute and, in a hosted product, cost, with nothing forcing it to stop. A max_tokens setting is the deliberate outer bound a product places on that loop — a maximum number of output tokens allowed for one reply, chosen by the people who built the product, not a property of the model’s own judgement about when it is “done.”
This is also why the same underlying model can appear to give short, clipped answers in one product and long, complete ones in another: the difference is very often not the model at all, but the length cap the surrounding product chose to set.
A response cut off by max_tokens can usually be continued or extended by simply asking, or by raising the configured limit. A response cut off because the context window itself filled up is a different situation: there may be little or no room left for a continuation in the same conversation, because everything already in the window — the conversation history, any attached documents, the reply so far — has used up the shared budget described in what a context window actually limits. The practical fix is usually to shorten what is being asked in one pass, trim earlier conversation history, or start a fresh exchange with only the material that is actually still needed — not simply to ask the model to keep going.
The two limits are not always independent of each other. Anthropic’s own documentation states that its models with the largest available context window share a 1M-token window, and that (Anthropic, Claude API documentation — context windows) “a single request to any of them can generate up to 128k output tokens (max_tokens).” Models with a smaller total window are correspondingly capped lower on output length. That is a useful, concrete illustration of the point made throughout this hub about context windows: the output cap and the input budget are not two unrelated numbers set by different parts of the system — they are drawn from, and constrained by, the same total figure.
Both reasons described above are two entries in a longer, documented list. The same reference names five other values the field can carry: end_turn for when “Claude finished its response naturally,” stop_sequence for when “Claude emitted one of your stop_sequences,” tool_use for when “Claude is calling a tool,” pause_turn for when “a server-tool loop reached its iteration limit,” and refusal for when “Claude declined to respond.” (Anthropic, Claude API reference — stop reasons) A cut-off answer is therefore one outcome among seven the field can report, not one of only two ways a response can end.
The specific figures above belong to one provider and change as models are updated, so they are worth treating as an illustration of the relationship rather than a number to design around. The relationship itself — that a length cap sits inside a larger shared budget, not outside it — is the part that holds regardless of which provider or model version is in front of you.
It is tempting to treat a cut-off answer as simply “ask again, more briefly.” That fixes the immediate symptom without addressing what a truncated answer actually is: an incomplete piece of reasoning that stopped partway through, not a shorter but still-considered one. Canada’s federal, provincial and territorial privacy commissioners make a directly relevant accountability point in their joint principles for generative AI, in language written for reliability generally but squarely applicable here: an organization relying on a generative AI tool is expected to (OPC, Principles for generative AI) “evaluate the validity and reliability of the generative AI tool for the intended purpose,” and “tools must be accurate throughout the intended lifecycle of the tool and across the variety of circumstances in which they are used.” A response that stopped mid-way through a numbered list, a calculation, or a set of instructions has not necessarily delivered a smaller but complete version of the same answer — treating it as one, rather than as an incomplete draft, is where a cut-off response quietly becomes a wrong one.
A business generating structured output — a report, an extracted data table, a long draft — with an AI model needs to know in advance which of these two limits it is closer to, because the failure looks identical in the output (a response that simply stops) but calls for a different fix. A generous max_tokens setting solves the first case. It does nothing for the second, because the second is a symptom of the input side of the same request being too large, not the output side. Distinguishing them, rather than assuming every truncated answer is simply the model running out of things to say, is usually the fastest way to actually fix it — and checking whether the answer is complete, not only whether it arrived, is worth building into the workflow either way.
Related, within this hub: what a context window actually limits and how ChatGPT generates an answer. For workflows that depend on long, complete output, see how we approach AI integration and automation.
No. It can also happen because the input already filled the shared context-window budget before the model finished writing, which is a different condition with a different fix — trimming the input, not raising an output limit.
Usually, if the cause was an output length cap. If the cause was the context window itself filling up, there may be little room left in the same conversation, and shortening what's being asked or starting fresh is usually more effective.
Whoever built the product you're using sets that limit as a parameter on each request. It is not a property of the underlying model deciding it is finished, which is why the same model can appear to give short or long replies depending on the product.
A short call is enough to work out which limit is actually being hit.