Short-Term vs Long-Term Memory in LLMs, and What "Context Memory" Actually Means

AI Agent MemoryContext EngineeringGaurav Dadhich2026-09-2715 min read
Short-Term vs Long-Term Memory in LLMs, and What "Context Memory" Actually Means
On this page
  1. An LLM remembers nothing between calls
  2. Short-term memory: the context window and what fills it
  3. Long-term memory: weights and external stores
  4. Short-term vs long-term memory in LLMs, side by side
  5. What is context memory in AI systems?
  6. How short-term and long-term memory work together in a single turn
  7. How ChatGPT and Claude handle long-term memory
  8. Where each kind of memory fails
  9. When a bigger context window is enough, and when you need long-term memory
  10. How Maximem Synap handles short-term and long-term memory
  11. Frequently asked questions

Published 27 September 2026 · Definitions checked against Anthropic, Google, OpenAI and LangChain documentation and the CoALA, RAG and LongMemEval papers; product memory features and model limits read from vendor help pages on 26 September 2026.

Short-term memory in an LLM is whatever sits in the context window on the current call: the system prompt, the recent conversation, tool results and any retrieved text, all re-sent by the application on every request and gone when the request ends. Long-term memory is information that shapes the model's answers across separate calls and sessions, and it lives in one of two places: frozen into the model's weights during training, or written to an external store by the application and retrieved back into the context window when it becomes relevant. "Context memory" is the name some teams give to the first and others give to the machinery that decides what goes into it; the precise version is below.

Maximem Synap is a memory service built for the long-term side of that split. It reads a conversation as it happens, keeps discrete statements worth remembering (a fact, a preference, a dated event, a stated feeling) rather than the transcript, and hands the agent a short ranked block of them before it replies, so the context window carries the right two thousand tokens instead of the whole history. The distinction between the two kinds of memory is what makes that design necessary, so it is worth getting exactly right.

An LLM remembers nothing between calls

Every LLM API call is stateless. The model receives tokens, produces tokens, and keeps nothing. Anthropic's context window documentation describes the window as "all the text a language model can reference when generating a response", distinct from "the large corpus of data the language model was trained on", and calls it "a 'working memory' for the model". A chat feels continuous only because the application sends the earlier messages again; OpenAI's conversation state guide is explicit that even when you chain responses by ID, "all previous input tokens for responses in the chain are billed as input tokens".

That one fact organises everything else. Short-term memory is continuity you get by re-sending tokens. Long-term memory is continuity you get by storing something and deciding when to bring it back. An agent is stateless by default, and whatever it appears to remember was engineered on top by somebody.

Short-term memory: the context window and what fills it

Short-term memory is the contents of the context window for this request. Google's long-context guide uses the same comparison: "An analogy for the context window is short term memory. There is a limited amount of information that can be stored in someone's short term memory, and the same is true for generative models." In agent frameworks it has a more specific shape. LangChain's memory documentation defines short-term memory as "thread-scoped memory" that "tracks the ongoing conversation by maintaining message history within a session", persisted with a checkpointer so a thread can be resumed.

Its capacity is the model's window, which on current flagships is 1M to 1.05M tokens, while the chat apps built on them cap it far lower (ChatGPT Plus gives the Instant model 54K). We keep the full table, and what each API and app does when a conversation overflows, in LLM context windows in 2026.

Working memory is a near-synonym worth separating once. In the CoALA framework for language agents (Sumers et al.), working memory "reflects the agent's current circumstances: it stores the agent's recent perceptual input, goals, and results from intermediate, internal reasoning". For an LLM agent that is the context window plus in-run state such as the plan and scratchpad; "short-term memory" usually means the same space seen as conversation history.

A KV cache is not memory in this sense. Hugging Face's documentation describes it as storing attention calculations "so they can be reused without recomputing them", which makes generation faster within a request, and provider prompt caching reuses a prefix across requests to make it cheaper. Neither lets the model recall anything it is not being sent: Anthropic notes that cached prefixes "still occupy the context window: prompt caching changes what you pay for those tokens, not whether they count."

Because short-term memory is bounded, every serious application manages it. The options run from keeping everything until it overflows, to a sliding window that drops the oldest turns, to summarising older turns, to server-side compaction, which OpenAI and Anthropic both now offer in their APIs. Each of those keeps one conversation going longer. None of them carries anything into the next one.

Long-term memory: weights and external stores

Long-term memory is any information that influences the model beyond the current request, and it has two homes that behave nothing alike.

One home is the model's weights, which the RAG paper by Lewis et al. calls parametric memory. Everything a model learned in training persists every time the model is loaded, which is why it can answer general questions with an empty prompt. It cannot be updated per user or per fact without further training, and it has a date on it: OpenAI's GPT-4o page lists a knowledge cutoff of 1 October 2023. Nothing a user tells a deployed model changes its weights.

Outside the model sits external, or non-parametric, memory: records the application writes to a database, vector index, graph or document store, and retrieves into the context window when they are relevant. This is what "long-term memory" means in almost every product conversation, and it is the only kind an application team controls. LangChain's definition captures the scope: long-term memory "stores user-specific or application-level data across sessions and is shared across conversational threads."

External long-term memory is usually divided by what it holds. CoALA puts it this way: "Semantic memory stores facts about the world ... while episodic memory stores sequences of the agent's past behaviors," and procedural memory holds the rules that drive behaviour, which LangChain maps onto the agent's system prompt and instructions. In practice, "the user is vegetarian" is semantic, "the user's refund was delayed in June and they were frustrated" is episodic, and "always confirm the delivery address before checkout" is procedural.

Borrowing from human memory (sensory, short-term, long-term) helps up to a point. Where it breaks is that human short-term memory decays and consolidates on its own, while an LLM's short-term memory is simply text that the application chose to send, and an LLM's weights do not learn from the conversation at all. Consolidation, if it happens, is something the memory system has to do deliberately.

Short-term vs long-term memory in LLMs, side by side

Short-term memoryLong-term memory
What it isThe contents of the context window on this callInformation that persists across calls and sessions
Where it livesIn the request: system prompt, recent turns, tool output, retrieved textModel weights (parametric) or an external store (non-parametric)
LifetimeOne request; one thread if the app keeps re-sending itUntil it is updated, forgotten or deleted; weights until retrained
CapacityThe model's window (1M to 1.05M tokens on current flagships), often less in appsBounded by storage, not by the model
CostBilled as input on every call that carries itA write per new fact, plus a bounded retrieval per turn
How it changesEdit, trim, summarise or compact the promptWrite, update or delete records; retrain for weights
Typical failureOverflow, and facts lost in the middle of a long promptStale or contradictory facts, and retrieving the wrong thing
Who controls itThe application, per requestThe application (external) or the model vendor (weights)

Cost is the row teams underestimate. A conversation that re-sends its whole history grows its token bill with the square of its length; our Agentic Context Management paper shows naive accumulation "grows token cost quadratically in conversation length", and the worked numbers are in how to reduce LLM token costs in long conversations.

What is context memory in AI systems?

Context memory is the information an AI system places in front of the model for the current call, together with the rules that decide what goes in. The context window is the container, measured in tokens; context memory is what fills it and the policy that chooses it. The term has no standard definition, and pages ranking for it today use it in two opposite ways: some mean the contents of the window (short-term), and others mean a persistent layer that stores context for later (long-term). Both are describing half of the same loop.

Three terms get conflated here, and they are different things. The context window is a capacity (GPT-4o's is 128,000 tokens). Context memory is the content and selection policy for that capacity. Retrieval-augmented generation is one way to fill it, by retrieving passages from a document index. A well-built system uses long-term memory and RAG as sources, and treats context memory as the budgeted, curated result that actually reaches the model. Anthropic's documentation makes the same point in its own terms, calling "curating what's in context just as important as how much space is available."

How short-term and long-term memory work together in a single turn

Long-term memory is only useful through short-term memory, because the model can only use what is in the window. A turn in a memory-equipped agent runs as a loop: retrieve what is relevant from long-term memory, place it in the context window beside the recent conversation, let the model reason and reply, then write anything new and worth keeping back to long-term memory. CoALA names the two halves "retrieval (read from long-term memory)" and "learning (write to long-term memory)".

Writing can happen in two places. LangChain distinguishes writing "on the hot path", where the agent decides to remember something before it responds, from writing "as a background task" that does not delay the reply. And there are two philosophies for what gets written: store the transcript (or chunks of it) and search it later, or extract discrete statements as the conversation happens and store those.

An obvious objection here is that the agent framework already handles memory. It does handle one kind well. A LangGraph checkpointer persists thread state to a database so that a thread "can be resumed at any time", which is short-term memory done properly. LangGraph also ships stores for long-term memory, organised by namespace and shared across threads. What neither piece decides is which of the things a user said three weeks ago still matter, what to do when a new statement contradicts an old one, or that "my manager" and "Sarah" are the same person. Those are the long-term memory problems, and they sit on top of the framework rather than inside it; our LangGraph integration plugs into exactly that slot.

How ChatGPT and Claude handle long-term memory

ChatGPT and Claude both ship long-term memory now, built outside the model. OpenAI's Memory FAQ describes two mechanisms in ChatGPT: saved memories, which the user asks it to keep or which it saves as useful context, and reference chat history, which learns from past conversations and "can change as ChatGPT updates what is most useful to remember". Deleting a chat does not delete a saved memory created from it, and turning chat history off schedules what was learned "for deletion from OpenAI systems within 30 days".

Claude's approach, per Anthropic's help center, combines search over past chats with a memory summary: "Claude will automatically summarize your conversations and create a synthesis of key insights across your chat history", updated every 24 hours. Each project has its own separate memory space, memory is on by default for Free, Pro and Max and off by default on Team and Enterprise, and incognito chats are not saved to memory at all. Both designs confirm the split: the window stays short-term, and persistence is a separate system with its own controls. We looked at how these consumer memories behave in practice in why AI forgets.

Where each kind of memory fails

Short-term memory fails in two ways. It overflows, which produces an error, a silent truncation or an app telling you to start again. And it degrades before it overflows: Lost in the Middle found performance "significantly degrades when models must access relevant information in the middle of long contexts, even for explicitly long-context models", and Chroma's context rot study found degradation across all 18 models it tested as input length grew.

Long-term memory fails differently. It returns something stale ("lives in Chicago" after the user moved), it returns two contradictory versions and leaves the model to pick, it retrieves the wrong record, or it misses that two names refer to one person. A store of transcripts is especially prone to the contradiction problem, because it holds both the March statement and the June correction with no view about which is current. LongMemEval measures exactly these abilities (information extraction, multi-session reasoning, temporal reasoning, knowledge updates and abstention) across 500 questions, and its authors found commercial chat assistants and long-context LLMs "showing a 30% accuracy drop on memorizing information across sustained interactions".

When a bigger context window is enough, and when you need long-term memory

A larger window is enough when the job is bounded: one long document, one coding session, one conversation that will not be resumed. Everything the model needs can be sent in one request, and there is nothing to remember afterwards.

Long-term memory becomes necessary once information has to outlive the request: users who return, conversations that move between chat, email and voice, more than one agent serving the same customer, or conversations long enough that re-sending the history dominates the bill. No window size changes that, because the window is emptied at the end of every call by definition. We argue the point at length on our bigger window is not a memory page.

RAG sits between the two and is often mistaken for long-term memory. It is non-parametric memory in the Lewis et al. sense, but a document index answers "what does our documentation say", not "what did this user tell us, and is it still true". It does not decide what to keep from a conversation, resolve contradictions, track when something changed or notice that two names belong to one person. Our 50,000-document study of file search versus vector search covers the retrieval half of that question in detail.

How Maximem Synap handles short-term and long-term memory

Maximem Synap treats the two as separate problems with separate machinery. On the short-term side, Synap's short-term context is a compacted summary of the conversation plus the most recent turns kept verbatim. Agentic Compaction triggers at 3,000 tokens, 10 messages or 5 minutes idle, compresses toward a 1,500-token target while preserving the last three conversation pairs word for word, and returns a validation score and a preserved-facts count on every pass, so you can see whether the compression kept the signal.

On the long-term side, Synap reads each conversation as it arrives and stores discrete, self-contained statements (facts, preferences, episodes, emotions and dated events are the common kinds), each with a confidence score, a scope and two timestamps: when it was stored and when the thing it describes happened. Writes return immediately and extraction runs in the background. When "I moved to Berlin" arrives, it retires "I live in Chicago" rather than filing both, and the older statement is kept as history and linked to its replacement, so every memory has provenance; a poorer statement never replaces a richer one. Memory is scoped by client, customer and user, plus the conversation, and a read sees the narrowest level that applies and everything above it, never sideways into another user or tenant.

Reads are where the two meet. Before the agent replies, it asks Synap for context and gets a ranked block within a token budget (2,000 tokens by default), pre-fetched into a cache inside your own process at an asserted P75 under 15 ms in conversation. On the benchmarks built for this problem, Synap scores 92% on LongMemEval and 93.2% on LoCoMo, reproduced on our open eval harness with gpt-5-mini as both the answer model and the judge, across the full 500-question LongMemEval set and 1,540 LoCoMo questions (categories 1 to 4, adversarial category 5 excluded), as published in our July 2026 paper; these are our own runs, not an independent verification. The detail is on our evals page and in how Synap works.

Frequently asked questions

What is the difference between short-term and long-term memory in LLMs?

Short-term memory is the content of the context window on the current call, re-sent by the application on every request and lost when the request ends. Long-term memory is information that persists across calls and sessions, either in the model's weights from training or in an external store that the application writes to and retrieves from. Short-term memory is bounded by the window and billed on every call; long-term memory is bounded by storage and brought into the window only when relevant.

What is context memory in AI systems?

Context memory is the information an AI system places in the model's context window for the current call, plus the rules that decide what goes in. The context window is the capacity in tokens; context memory is the curated content that fills it, drawn from the recent conversation, long-term memory and retrieved documents. The term has no standard definition, so check whether a vendor means the window's contents or a persistent store.

Is the context window the same as short-term memory?

In practice, yes. Google's Gemini documentation describes the context window as analogous to short-term memory, and Anthropic's calls it the model's working memory. The window is the limit; short-term memory is what is in it on a given call, including the system prompt, recent turns, tool results and retrieved text.

Is RAG a form of long-term memory?

RAG is a form of non-parametric memory, because it retrieves from an external index, but it is not the same as long-term memory of a user or a conversation. A document index returns passages that match a query; it does not decide what to keep from conversations or notice that a new statement contradicts an old one.

Is the KV cache a kind of memory?

No. A KV cache stores attention calculations so a model does not recompute them while generating, and provider prompt caching reuses a repeated prefix to cut cost. Both make requests faster or cheaper, but cached tokens still occupy the context window and nothing is recalled that is not being sent.

Does ChatGPT have long-term memory?

ChatGPT does, through two mechanisms that sit outside the model: saved memories, which a user can ask it to keep, and reference chat history, which it learns from past conversations. Claude offers chat search and a memory summary updated every 24 hours, with separate memory per project. In both products the model itself remains stateless.

Do I need long-term memory if my model has a 1M-token context window?

You need it if anything has to survive the session. A 1M-token window lets one conversation run longer, but it is empty again on the next call unless the application re-sends everything, at full input cost each time, and accuracy degrades as the window fills. Long-term memory, such as Maximem Synap, stores what matters and sends a small relevant slice instead.

Sources, retrieved 26 September 2026: Anthropic, context windows; Anthropic, compaction; Google, Gemini long context; OpenAI, conversation state; OpenAI, compaction; OpenAI, GPT-4o model; LangChain, memory overview; Sumers et al., CoALA; Lewis et al., Retrieval-Augmented Generation; Hugging Face, KV cache; Liu et al., Lost in the Middle; Chroma, Context Rot; Wu et al., LongMemEval; OpenAI Help Center, Memory FAQ; Claude Help Center, chat search and memory; ChatGPT pricing; Maximem, eval harness; Maximem, Agentic Context Management.

From the team at Maximem

Stop rebuilding agent memory from scratch

Maximem Synap is the context management layer we built after hitting every problem in this post ourselves. Persistent recall across sessions, entity resolution and conscious forgetting, in Python, TypeScript and REST.

Related posts