Treadstone Associates
Article · 7 min read

Why AI forgets what you told it

An assistant that seems to forget something you said a few messages ago is not being careless. It is a direct consequence of how little the underlying mechanism actually retains on its own.

Treadstone Associates · Updated 2026

Key takeaways

  • • By default, a model retains nothing between requests — each one resends the full conversation history for the model to read again.
  • • The context window that holds that history is explicitly not memory — Anthropic's own documentation calls it a temporary “working memory”.
  • • Long chat products often manage growing conversations on a rolling basis, quietly dropping the oldest turns once the window fills up.
  • • Reasoning quality can degrade before anything is actually dropped, a pattern documented as “context rot”.

There is no notebook — only what gets resent

A model does not retain anything from one exchange to the next by default. Each request is handled independently, and whatever the model appears to “know” about the current conversation is actually just whatever text was included in that specific request. Anthropic’s own documentation makes this explicit in a completely different context — explaining why sending images repeatedly is expensive — and the underlying mechanism it describes applies to everything in a conversation, not only images: (Anthropic, Claude API documentation — vision) “in multi-turn conversations and agentic workflows, each request resends the full conversation history.” There is no separate, persistent memory store the model is quietly writing to as you talk. What looks like an ongoing conversation is, mechanically, a series of separate requests, each one re-including everything said so far.

Which is also exactly what a context window is for

That resent history is what fills the model’s context window, described fully in what a context window actually limits — and a context window is explicitly not the same thing as memory. Anthropic’s documentation draws the line directly: it (Anthropic — context windows) “represents a ‘working memory’ for the model,” distinct from “the large corpus of data the language model was trained on.” A working memory, by definition, is temporary and has a limit. Once a conversation’s resent history grows large enough, something has to give, and the same documentation names the mechanism many long-running chat products use to cope with that: (Anthropic — context windows) “chat interfaces such as claude.ai can also manage the context window on a rolling ‘first in, first out’ basis” — meaning the oldest turns are quietly dropped as new ones are added, once the conversation has grown past what fits.

That is the actual mechanism behind the common complaint that an assistant “forgot” something mentioned early in a long conversation. It did not selectively decide the detail was unimportant. If it fell outside the window at the point the oldest turns were trimmed, it was no longer part of what got resent in the request that produced the later reply — it was not there to read, the same way a person cannot recall a page of a document that was never handed to them.

Two things that look like memory and are not

Some products add a genuine, separate memory feature on top of this — a system that writes specific facts to persistent storage and deliberately re-injects them into future conversations. That is a real engineering feature layered onto the stateless mechanism above, not a change to how the underlying model itself works between requests. And retrieval, covered on its own at how retrieval grounds an AI answer, can look similar from the outside — the model appears to “know” something from an earlier document — but the mechanism is the same resend-and-include pattern, just pulling from a document store instead of your own earlier messages.

The distinction is worth keeping straight for anything a business builds: if a workflow needs a fact to persist across separate sessions, that has to be engineered deliberately — stored somewhere and reintroduced into context on purpose — rather than assumed, because nothing about the base mechanism carries it forward on its own.

Why this shows up as a length problem before it shows up as a memory problem

Because everything has to be resent every time, a long conversation is not free even when nothing appears to have gone wrong — it is quietly consuming more of the shared budget described in what a context window actually limits with every added turn. Anthropic’s documentation names a real cost to that growth beyond simply running out of room: (Anthropic — context windows) “as token count grows, accuracy and recall degrade, a phenomenon known as context rot.” In other words, forgetting is not only what happens once the window is full and old turns are dropped — reasoning over an already-very-full window can already be getting worse before anything is dropped at all.

The one thing that genuinely does carry forward: fine-tuning, and it is a different mechanism entirely

There is exactly one way to make a model’s later behaviour depend on something told to it earlier, without resending the material every time — and it is not part of an ordinary conversation. Anthropic’s developer glossary defines it as a distinct process: fine-tuning is (Anthropic, developer glossary) “the process of further training a pretrained language model using additional data,” which “causes the model to start representing and mimicking the patterns and characteristics of the fine-tuning dataset.” That is genuine further training, run deliberately by whoever controls the model, on a chosen dataset — not something that happens automatically because a user typed something into a chat window enough times. Telling a model something in conversation, however many times, does not fine-tune it; it only ever populates that one conversation’s context window, which is exactly why it can vanish once that window is trimmed or the session ends.

The practical implication is blunt: if a business genuinely needs a model to reliably carry forward specific information across separate sessions and separate users — a policy, a preferred format, a standing fact about the business — that requires deliberate engineering, whether through fine-tuning, a persistent memory feature built for the purpose, or retrieval that re-supplies the same material every time it is needed. None of those happen by default, and mistaking a long conversation for a taught one is the most common way that gap goes unnoticed until it causes a problem.

Related, within this hub: what a context window actually limits, how retrieval grounds an AI answer, and why AI answers get cut off. If a workflow needs facts to persist reliably across sessions, see how we approach custom AI solutions.

Common questions

Does an AI model remember previous conversations by default?

No. Each request is handled independently, and the model only has access to whatever text is included in that specific request — typically the resent history of the current conversation, not anything from a separate, earlier one.

If a conversation gets very long, does everything in it stay available?

Not necessarily. Some products manage long conversations on a rolling basis, dropping the oldest turns as new ones are added once the conversation outgrows the available window, and reasoning quality can also degrade before anything is actually dropped.

Are memory features in some AI products a change to how the model itself works?

No. They are a separate engineering layer that stores specific information and deliberately reintroduces it into later requests. The underlying model still has no persistence of its own between requests.

Building that kind of persistent memory deliberately also imports a compliance duty that resending a conversation’s history never triggers on its own. Where what is being carried forward is personal information about a customer or user, PIPEDA’s own retention principle applies to it: “personal information shall be retained only as long as necessary for the fulfilment of those purposes,” and information “that is no longer required to fulfil the identified purposes should be destroyed, erased, or made anonymous.” (laws-lois.justice.gc.ca, PIPEDA Schedule 1, Principle 5) A memory feature built to remember indefinitely is, on that principle, built the wrong way by default — it needs its own retention and deletion rules, not just the engineering to make persistence work at all.

Need facts to persist reliably across sessions, not just within one chat?

A short call is enough to map what actually needs engineering versus what the base mechanism already handles.