Published 27 September 2026 · Framework behaviour, context-window figures and benchmark conditions checked against the linked sources on 26 September 2026
You give an LLM long-term memory by adding a layer outside the model that decides what is worth keeping from each conversation, stores it against the right person, updates it when a fact changes, and puts only the relevant slice back into the prompt on the next call. Maximem Synap is that layer as a managed service: your agent sends each conversation as it happens and asks for context before it replies, and on the two public long-term memory benchmarks it scores 92% on LongMemEval (500 questions) and 93.2% on LoCoMo (1,540 non-adversarial questions), results published in July 2026 and reproduced on our open evaluation harness with gpt-5-mini as both the answer model and the judge.
That answer rests on a premise worth stating exactly, because it decides the design. Every call to a language model is stateless: the model sees the tokens in that one request and nothing else. An agent appears to remember inside a session because its framework keeps the message list and sends it again on every turn, and whatever continuity it has across sessions was engineered by somebody on top of that. Long-term memory is the engineering for the second case, and most of it has nothing to do with which database you pick.
Why an LLM forgets, stated precisely
Short-term memory is whatever sits in the context window on this call: the system prompt, the recent turns, retrieved documents and tool results. It is exact, and it disappears when the session ends. Long-term memory is information stored outside the model that survives the session and is retrieved back into the context when it becomes relevant. A larger window makes short-term memory bigger without making it persistent, which is the whole difference, as our memory versus context windows page sets out.
Trouble always starts at a boundary. The user returns next week, the support thread moves from chat to email, or the agent that took the order is not the one handling the refund. At each boundary the message list is gone, and the person on the other side notices one of three things: being asked for something they already gave, being told something that stopped being true, or being treated as a stranger by a product they have used for a year. Inside one long session the related problem is growth, and compressing older turns is a legitimate short-term fix only if you can tell whether the compression kept the facts that matter.
What your agent framework already gives you
Every argument about agent memory meets the same objection first: the framework already handles it. That objection is partly right. In LangGraph, short-term memory is thread-scoped state persisted by a checkpointer, so a conversation can be paused and resumed after a restart with its history intact, and for carrying one thread through a long multi-step run it does the job it was designed for. LangGraph's documentation also offers long-term memory as stores: documents saved under custom namespaces, readable from any thread, with semantic search and content filtering.
What the framework deliberately leaves to you is every decision about what goes into that store. Its own documentation calls long-term memory "a complex challenge without a one-size-fits-all solution" and warns that with a collection of memories "the model must now delete or update existing items in the list, which can be tricky." Deciding what is worth writing, recognising that a new statement replaces an old one, resolving that "my manager" and "Sarah" are the same person, keeping one tenant's memories out of another's retrieval, and retiring what has gone stale all stay in your code. A checkpointer is the right first step and an insufficient last one for any agent that meets the same person twice.
Six ways to give an LLM long-term memory
Every approach in use today falls into one of six patterns, and each works until a specific condition arrives.
| Approach | How it works | Works well for | Breaks when |
|---|---|---|---|
| Replay the full history | Send every prior turn on every call | Short conversations | Cost grows every turn and quality falls as the prompt fills |
| Rolling summary | Compress older turns into a running summary | One long thread | Each pass loses detail, the losses compound, and nothing records where a fact came from |
| RAG over past transcripts | Chunk and embed old conversations, retrieve similar passages | Recalling what was said | It returns the March version and the June correction side by side, with no view on which is current |
| Extracted memory store | Extract discrete statements with type, time, scope and confidence; update on conflict; retrieve ranked | Production agents, built carefully | You own extraction, contradiction handling, entity resolution, consolidation and isolation |
| Model-managed memory | The model calls tools to page information between its context and an external archive, an approach described in a 2023 research paper on paging memory in and out of the context window | Autonomous, long-running agents | Memory quality rests on the model's own judgement at every step |
| Fine-tuning | Train facts or behaviour into the weights | Stable style and procedures | Per-user facts change weekly, and one user's fact cannot be deleted from weights |
A managed memory layer is the fourth row delivered as a service, so the extraction, update, scoping and consolidation work is not yours to build and operate.
Replaying everything is where almost everyone starts, and growing context windows make it tempting to stay there. Anthropic's documentation lists Claude Opus 5.5 and Claude Sonnet 5 with 1M-token windows, and the same page warns that as token count grows, accuracy and recall degrade, which it calls context rot. Cost is worse than linear: our paper on Agentic Context Management shows naive accumulation grows token cost quadratically with conversation length, while validated compaction keeps it linear. Below a few thousand tokens of history, sending everything is cheaper and more accurate than any extraction pipeline, ours included; past that point you pay more for a worse answer.
RAG vs agent memory: the difference that matters
RAG and agent memory both pull text into a prompt, which is why they get confused, but they answer different questions. RAG answers "what do our documents say?" from a corpus your team maintains, the same way for every user, and it writes nothing back. Agent memory answers "what do we know about this person, and what has happened with them?" from state the agent builds out of its own conversations, which it then has to keep current, scope to the right person, reconcile when facts conflict and eventually retire.
| RAG | Agent memory | |
|---|---|---|
| Source | Documents your team writes and maintains | The agent's own conversations and actions |
| Written by | An indexing pipeline, ahead of time | Each conversation, turn by turn |
| Scope | The same for every user | Per user, per tenant, per conversation |
| When it changes | When someone re-indexes | Continuously, as facts are superseded |
| Unit returned | A chunk of a document | A self-contained statement about someone |
| Two versions of a fact | Both retrieved side by side | The newer supersedes the older, with history kept |
| Typical failure | A retrieval miss or a stale index | A stale fact about a person, or a fact leaking across users |
| Example question | "What is the refund policy?" | "What has this customer already tried?" |
A quick test settles most cases. If a brand-new agent could answer correctly from your documents alone, it is a RAG question; if the right answer depends on what happened with this particular person, it is a memory question. Most production agents need both in the same turn, as when a support agent retrieves the current troubleshooting steps and recalls that this customer already tried a restart yesterday.
Memory systems do use retrieval internally, which is where the "memory is just RAG" argument comes from. Retrieval is one operation in a memory layer; the others are deciding what to store, resolving the same entity across sessions, choosing between contradictory facts, letting old facts decay and enforcing who may read what. A retrieval system does the first of those well and none of the rest.
Our own measurements show where the line sits. In a study of 50,000 documents and 5,000 queries, scored on MRR@10 by exact document-ID match with no model judge, vector search won the five-dataset average 0.6320 to 0.5325, and keyword search won the average of the other four once the natural-language-to-code dataset was removed, 0.5931 to 0.5614. Every one of those numbers came from documents somebody else wrote and that never change. Agent memory is neither: it grows every turn, the agent writes it, and yesterday's fact can contradict today's, so a retriever that ranks well on a static corpus will still hand the model both the old address and the new one.
Neither system should hold a value that a live system owns. The current account balance or an order's shipping status comes from a tool call to the system of record; memory holds that the customer disputed last month's invoice, and RAG holds the policy that applies. If you already run RAG, you have the retrieval half of memory and none of the write half, and pointing the same index at old transcripts looks fine in a demo because demos rarely include a user who changed their mind.
Building long-term memory: the loop and five decisions
Whatever you build on, the runtime shape is one loop: read memory before the model reasons, write memory after it acts. A turn arrives, the agent fetches the relevant memories for this person and assembles the prompt, the model reasons and calls tools, and new information from the turn is extracted and written back. Making that loop hold up in production comes down to five decisions.
What to keep. Storing every message recreates the transcript problem in a database. Extract discrete, self-contained statements instead, such as "The user's contract renews on 1 March", each with a type, a confidence score, its scope, and two timestamps: when it was stored and when the thing it describes happens, which is usually the date that matters. The common types follow the split that IBM and the LangGraph documentation both use (semantic facts, episodic events, procedural rules), and preferences deserve their own type because they change more often than facts.
What happens when a fact changes. A new statement that resembles an old one replaces it, adds to it, duplicates it, or is unrelated; "I moved to Berlin" retires "I live in Chicago". Mark the old statement historical and link it to its replacement rather than deleting it, so you can always answer why the agent believes something. And never let a poorer statement overwrite a richer one, because "takes a cholesterol medication" is not a safe replacement for "takes 20mg of Atorvastatin daily for cholesterol management".
Who a memory belongs to. Assign scope when the memory is written, not when it is read. Visibility should flow one way, from wide to narrow: a user's session can draw on what is known about their company and your product, a company-level query never reaches into one user's private memories, and one user never sees another's.
Which entity it is about. "My manager", "Sarah" and "Sarah Chen" are one person across three conversations, and without resolution the agent holds three partial records of her. Merge confident matches automatically and send ambiguous ones to review rather than guessing.
What comes back. Rank candidates on relevance, recency and confidence, then trim to a token budget, because a memory layer that returns fifty loosely related facts has recreated the long-prompt problem at a smaller size. Hybrid retrieval helps for the same reason it helps RAG: people repeat exact names, and exact names are where keyword matching is strongest.
Beneath all five sits consolidation, the background work that merges duplicates and retires what has gone stale. Redis calls selective forgetting one of the least solved parts of memory systems, and a store that only grows gets worse at retrieval as it gets bigger.
For storage, PostgreSQL with pgvector covers a large share of workloads and keeps memories transactionally consistent with the rest of your data. A graph store such as Neo4j earns its place when questions depend on relationships between entities. Markdown files in git work for one developer and one agent, and stop working the first time two sessions write facts that contradict each other.
A minimal version on Postgres
A first version needs one table and one function. Scope, time and supersession are columns from day one, because retrofitting them is where homegrown memory stores usually go wrong.
CREATE EXTENSION IF NOT EXISTS vector;
CREATE TABLE memories (
id uuid PRIMARY KEY DEFAULT gen_random_uuid(),
tenant_id text NOT NULL,
user_id text NOT NULL,
kind text NOT NULL, -- fact, preference, episode
content text NOT NULL, -- one self-contained statement
confidence real NOT NULL,
occurred_at timestamptz, -- when the described thing happens
created_at timestamptz NOT NULL DEFAULT now(),
superseded_by uuid REFERENCES memories(id),
embedding vector(384) -- match your embedding model
);
CREATE INDEX ON memories USING hnsw (embedding vector_cosine_ops);
CREATE INDEX ON memories (tenant_id, user_id) WHERE superseded_by IS NULL;
def handle_turn(tenant_id: str, user_id: str, message: str) -> str:
# Read: current memories for this person only, nearest first
memories = db.fetch(
"""SELECT content FROM memories
WHERE tenant_id = %s AND user_id = %s AND superseded_by IS NULL
ORDER BY embedding <=> %s
LIMIT 8""",
(tenant_id, user_id, embed(message)),
)
reply = llm.generate(build_prompt(memories, message))
# Write: extract statements, supersede anything they replace
for fact in extract_statements(message, reply): # an LLM call
new_id = insert_memory(tenant_id, user_id, fact)
old = find_replaced(tenant_id, user_id, fact) # same subject, new value
if old:
mark_superseded(old.id, new_id)
return reply
That is enough to watch memory work, and it is honest about what it leaves out. Extraction runs inside the turn, so the user waits for it; in production it belongs in a background job. find_replaced does the hardest work in the file in one line and needs an evaluation set of its own. Nothing resolves entities or consolidates, and the tenant boundary is a WHERE clause, which is the first thing to fail in production.
What breaks in production
Most agent memory failures in production are scoping failures, not recall failures. One shared store with a tenant identifier in the query filter is the common pattern, and when that filter is dropped or malformed, an ordinary application throws an error while a memory-backed agent returns a fluent answer built from another customer's data, with nothing in monitoring to show it happened. Filtering at read time is a bouncer checking IDs at a door the facts already walked through; scope assigned at write time is the version that survives a second tenant.
Latency comes next, because the memory read sits in the hot path of every turn. Redis notes that an external store can add a 50 to 300 ms network round trip in some architectures, which a voice agent with a few hundred milliseconds per turn can hear. Writes should never block the reply, and reads should sit as close to the agent as you can put them.
Cost has two lines, and most comparisons quote one. Extraction-based memory makes model calls on write, which our build versus buy page puts at two to three per stored memory, while a memory layer that works removes old conversation from the prompt, which is where the saving comes from. Background writes can also invalidate a prompt cache mid-conversation, so measure cache hit rates before and after; our token cost breakdown covers where caching stops helping.
Privacy is a property of the write path. Detect sensitive values before they are stored, decide per kind of data how each is kept and who can read it back, make per-person deletion reach derived embeddings as well as source text, and refuse outright to store payment card numbers, security codes, passwords and API keys. When several agents serve one customer, give them a shared layer for what is known about that customer and separate spaces for what each needs privately, with the same scope rules deciding who reads what.
Adding long-term memory with Maximem Synap
Maximem Synap takes care of the write path, scoping, entity resolution and consolidation, so your agent code keeps two calls: send the conversation, and fetch context before replying. This is the Python quickstart from the Synap documentation:
# pip install maximem-synap
# export SYNAP_API_KEY=synap_... SYNAP_INSTANCE_ID=inst_...
import asyncio
from maximem_synap import MaximemSynapSDK
async def main():
sdk = MaximemSynapSDK()
await sdk.initialize()
try:
result = await sdk.memories.create(
document="User: I prefer dark mode.\nAssistant: Noted!",
document_type="ai-chat-conversation",
user_id="user_alice",
)
# Block until ingestion finishes (for scripts and tests)
await sdk.memories.wait_for_completion(result.ingestion_id)
context = await sdk.user.context.fetch(
user_id="user_alice",
search_query=["preferences"],
)
for p in context.preferences:
print(p.content)
finally:
await sdk.shutdown()
asyncio.run(main())
A live agent skips the wait: writes return immediately and extraction runs in the background, so a memory written now is retrievable a few seconds later. Behind those two calls, Synap extracts self-contained statements with a type, confidence, scope and event time, and applies Conscious & Lossless Forgetting, so a replaced fact is retired and linked to its successor while a poorer statement never overwrites a richer one. Entity Resolution collapses the same person across conversations and channels, and scoping follows the identifiers you already pass, across your account, your tenants and individual users, with deeper hierarchies modelled in your own names. Anticipatory Retrieval pushes what the agent is likely to need next into a cache inside your own process, which is how in-conversation reads reach an asserted P75 latency under 15 ms. Custom Context Architecture generates the extraction and retrieval design for each agent from a description of what it does, and a person reviews and approves it before it applies.
Adapters for 23 agent frameworks plug into the extension point each framework already has; for LangGraph that means the checkpointer and store interfaces, walked through in our LangGraph integration guide, with a matching LangChain guide.
How to know it is working, and when you do not need it
Two public benchmarks test long-term conversational memory. LongMemEval asks 500 questions over long multi-session chat histories, probing information extraction, multi-session reasoning, temporal reasoning, knowledge updates and abstention. LoCoMo tests very long multi-session conversations; memory results conventionally exclude its adversarial category, leaving 1,540 questions. Maximem Synap scores 92% on LongMemEval (all 500 questions) and 93.2% on LoCoMo (the 1,540 non-adversarial questions), with gpt-5-mini as both the answer model and the judge, binary correct-or-wrong judging and a single run per benchmark, first published with our paper in July 2026. Those results were reproduced on our open harness, not independently verified; the harness runs against any memory system, and the category breakdown and methodology are in the results repository and on our evals page.
Benchmarks tell you whether a memory system is competent in general; your own data tells you whether it helps your agent. The shortest useful test takes fifty turns of real conversation history and runs three configurations against the same questions: the full history in the prompt, your current files or database, and the memory layer. Measure task success judged by someone who knows your domain, total cost per completed task including retries, and added latency per turn. If the memory layer does not win on at least two of those three at your current conversation length, you do not need one yet.
You also do not need a memory layer when history is short enough to send in full, or when every user should get the same answer from the same documents, which is a RAG problem. You need one the first time two sessions write facts that contradict each other, or the first time one user must not see what another wrote, and both can happen at two users. Building it yourself is entirely reasonable; the estimate is usually wrong because retrieval is the easy part, and our real cost of DIY agent memory post adds up the rest.
An LLM that remembers is an engineering outcome, not a model feature. Get the write path and the scoping right, and the agent stops asking customers for things they already said, which is the change users notice first and the one that decides whether they come back.
Frequently asked questions
Can an LLM have long-term memory on its own?
No. Every call to a language model is stateless, so the model knows only its training data and the prompt for that call. Long-term memory is built outside the model: an application stores what matters from past conversations and puts the relevant parts back into the prompt when they are needed.
What is the difference between RAG and agent memory?
RAG retrieves passages from documents your team maintains and gives every user the same answer, writing nothing back. Agent memory stores what the agent learns from its own conversations, scoped to a specific person or tenant, and keeps it current by superseding facts that change. Most production agents use both: RAG for shared knowledge, memory for what happened with this user.
Is a vector database enough for long-term memory?
No. A vector database is a retrieval index. It returns two contradictory facts about the same person side by side without comment, and it has no view on which is newer, which entity a statement refers to, whether either was deleted, or who may read it. Those decisions belong to a memory layer above the index.
Does a 1M-token context window replace long-term memory?
No. A context window is short-term memory and is gone when the session ends. Anthropic lists 1M-token windows for Claude Opus 5.5 and Claude Sonnet 5 and notes that accuracy and recall degrade as token count grows, and replaying full history makes token cost grow quadratically with conversation length.
Can fine-tuning give an LLM long-term memory?
Not for per-user facts. Fine-tuning changes model weights, which suits stable behaviour and output style, but a user's circumstances change weekly, and one user's fact cannot be deleted from weights when they ask.
What is the difference between short-term and long-term memory in an AI agent?
Short-term memory is the context window and the current thread: exact, and lost when the session ends. Long-term memory is stored outside the model, survives across sessions, and is retrieved into context when relevant. LangGraph, for example, handles short-term memory with a checkpointer and leaves long-term decisions to a separate store.
How do I stop an agent from remembering outdated facts?
Store memories as discrete statements with timestamps, and when a new statement replaces an old one, mark the old one historical and link it to its replacement instead of keeping both as current. Retrieval then returns only current statements, while the history still explains why the agent believes what it believes.
When do I not need a long-term memory layer?
When conversations are short enough to send in full, when one developer runs one agent, or when every user should get the same answer from the same documents, which is a RAG problem. A memory layer earns its cost once you have more than one user or agent and facts that change over time.
Sources, retrieved 26 September 2026: Maximem Synap documentation, Quickstart; Maximem eval harness and benchmark results repository; Dadhich, Agentic Context Management (arXiv 2607.21503); Wu et al., LongMemEval (arXiv 2410.10813); LoCoMo; LangGraph memory overview; Anthropic, Context windows; Packer et al. (arXiv 2310.08560); Redis, Long-term memory architectures for AI agents; Redis, AI agent memory vs retrieval; IBM, What is AI agent memory; pgvector; Maximem, File Search vs Vector Search for RAG; Maximem, Build vs buy agent memory; Maximem, Memory vs context windows.



