How to Add Conversational Memory to a LangChain App (2026 Guide)

Agent FrameworksAI Agent MemoryGaurav Dadhich2026-09-2716 min read
How to Add Conversational Memory to a LangChain App (2026 Guide)
On this page
  1. Which kind of memory does your app need?
  2. Step 1: thread memory with create_agent and a checkpointer
  3. Step 2: keep long threads inside the context window
  4. Step 3: remember the user across conversations with a store
  5. What about ConversationBufferMemory and RunnableWithMessageHistory?
  6. Where the built-in memory stops
  7. Step 4: add Maximem Synap with maximem-synap-langchain
  8. Failure, latency and cost
  9. Controlling what gets extracted: extraction prompts and the alternative
  10. If you are on LangGraph
  11. How to test conversational memory
  12. The short version
  13. Frequently asked questions

Published 27 September 2026 · Code checked against langchain 1.4.2, langchain-core 1.6.5, langgraph 1.2.12 and maximem-synap-langchain 0.3.0 on 26 September 2026; every Maximem Synap example was executed against the installed packages.

To add conversational memory to a LangChain app today, pass a checkpointer to create_agent and reuse one thread_id per conversation, which gives the agent the full history of that thread; to remember a person across conversations, add a LangGraph store, or plug in Maximem Synap through the maximem-synap-langchain package, which records each turn, decides what is worth keeping and returns the relevant memories to your prompt through a standard LangChain retriever or tool. The older route is closing: ConversationBufferMemory no longer ships in the langchain package, and since langchain-core 1.6.4 (21 September 2026) RunnableWithMessageHistory and BaseChatMessageHistory are both marked deprecated, with removal scheduled for 2.0.

Every LLM call is stateless. A model receives tokens, returns tokens and keeps nothing, so whatever your app "remembers" is text that your code put back into the next prompt (why AI forgets covers the mechanism). LangChain gives you two places to keep that text, and a memory layer adds a third behaviour on top: deciding which parts of a conversation are facts worth carrying forward. This guide walks through all of it in the order you will need it, with the version-checked code for each step.

Which kind of memory does your app need?

There are three different jobs hiding under the word "memory", and picking the wrong one is the most common reason a LangChain chatbot that worked in testing forgets a returning user in production.

JobWhat it answersLangChain toolScope
Thread memory"What did we say earlier in this conversation?"Checkpointer (InMemorySaver, PostgresSaver)One thread_id
User memory"What do I know about this person from last week?"LangGraph store (InMemoryStore, PostgresStore)A namespace you design, usually a user ID
Managed memory"Which of the hundreds of things this person said still matter, and which are no longer true?"A memory layer such as Maximem SynapClient, customer, user and conversation, applied at write time

LangChain's memory overview borrows a psychology split that is worth keeping in your head: semantic memory holds facts (the user is on the Pro plan), episodic memory holds experiences (the refund on 3 September went wrong), and procedural memory holds instructions (answer in bullet points). Thread memory carries all three for one conversation; the other two jobs are about which of them survive when the conversation ends.

Step 1: thread memory with create_agent and a checkpointer

LangChain 1.x treats short-term memory as part of the agent's state. You pass a checkpointer when you build the agent, and every call that carries the same thread_id resumes the same message list. This is the pattern from the current LangChain short-term memory docs, with the model string they use today:

from langchain.agents import create_agent
from langgraph.checkpoint.memory import InMemorySaver

agent = create_agent(
    model="openai:gpt-5.5",
    tools=[],
    checkpointer=InMemorySaver(),
)

config = {"configurable": {"thread_id": "1"}}
agent.invoke({"messages": [{"role": "user", "content": "Hi, my name is Bob."}]}, config)
reply = agent.invoke({"messages": [{"role": "user", "content": "What is my name?"}]}, config)
print(reply["messages"][-1].content)  # the agent answers "Bob"

InMemorySaver keeps checkpoints in a Python dictionary, so a restart wipes them. For anything deployed, the docs point you to a database-backed saver:

# pip install -U langgraph-checkpoint-postgres "psycopg[binary]"
from langchain.agents import create_agent
from langgraph.checkpoint.postgres import PostgresSaver

DB_URI = "postgresql://postgres:postgres@localhost:5432/postgres?sslmode=disable"
with PostgresSaver.from_conn_string(DB_URI) as checkpointer:
    checkpointer.setup()  # creates the tables on first run
    agent = create_agent("openai:gpt-5.5", tools=[], checkpointer=checkpointer)

Keep thread_id values under 255 characters on Postgres, because the column is bounded (LangGraph persistence docs); a UUID per conversation is the simplest convention, and checkpoints accumulate, so plan a retention job early.

What a checkpointer gives you is exact: the thread's state, resumable after a crash or a human-in-the-loop pause. What it does not give you is anything outside that thread. A returning user who opens a new conversation gets a new thread_id and a blank slate.

Step 2: keep long threads inside the context window

A thread that grows forever will eventually cost more per turn than it is worth and then overflow the model's window. LangChain 1.x handles this with middleware on the agent. The built-in SummarizationMiddleware replaces older messages with a model-written summary once a trigger is hit:

from langchain.agents import create_agent
from langchain.agents.middleware import SummarizationMiddleware
from langgraph.checkpoint.memory import InMemorySaver

agent = create_agent(
    model="openai:gpt-5.5",
    tools=[],
    middleware=[
        SummarizationMiddleware(
            model="openai:gpt-5.4-mini",
            trigger=("tokens", 4000),
            keep=("messages", 20),
        )
    ],
    checkpointer=InMemorySaver(),
)

If you would rather trim than summarise, the docs show a @before_model middleware that returns RemoveMessage(id=REMOVE_ALL_MESSAGES) followed by the messages you want to keep. Whichever you choose, keep each assistant message that requested a tool next to the tool result that answers it; most providers reject a history where one appears without the other.

Summaries are lossy by construction: a thread where the user said "ship it on Friday" and later "hold it until Monday" can be summarised with the first instruction and without the correction. For the cost side of long threads, see how to reduce LLM token costs in long conversations.

Step 3: remember the user across conversations with a store

Cross-conversation memory in LangChain is built on LangGraph stores: JSON documents saved under a namespace tuple and a key. You hand the store to create_agent, and tools reach it through ToolRuntime, following the long-term memory docs:

from dataclasses import dataclass
from typing_extensions import TypedDict

from langchain.agents import create_agent
from langchain.tools import ToolRuntime, tool
from langgraph.store.memory import InMemoryStore


@dataclass
class Context:
    user_id: str


class UserInfo(TypedDict):
    name: str


@tool
def save_user_info(user_info: UserInfo, runtime: ToolRuntime[Context]) -> str:
    """Save user info."""
    runtime.store.put(("users",), runtime.context.user_id, dict(user_info))
    return "Saved."


@tool
def get_user_info(runtime: ToolRuntime[Context]) -> str:
    """Look up user info."""
    item = runtime.store.get(("users",), runtime.context.user_id)
    return str(item.value) if item else "Unknown user"


store = InMemoryStore()  # use PostgresStore in production
agent = create_agent(
    "openai:gpt-5.5",
    tools=[save_user_info, get_user_info],
    store=store,
    context_schema=Context,
)
agent.invoke(
    {"messages": [{"role": "user", "content": "My name is John Smith"}]},
    context=Context(user_id="user_123"),
)

Identity is the distinction that matters. thread_id names a conversation and changes every session; the namespace names a person and never changes. Keep user facts in the store and conversation turns in the checkpointer, and pass both to create_agent when you need both. Configure the store with an embedding index and store.search(namespace, query=...) becomes a semantic search over what you saved.

When to write is your call. The LangChain overview describes writing "in the hot path" (a tool the agent calls before answering, so the memory is available at once but adds latency and a decision to every turn) or "in the background" (a separate job that extracts memories after the fact, with no latency cost but a delay before other threads see them).

What about ConversationBufferMemory and RunnableWithMessageHistory?

Most LangChain memory tutorials still indexed today teach ConversationBufferMemory, ConversationBufferWindowMemory, ConversationSummaryMemory and VectorStoreRetrieverMemory wired into a ConversationChain. Here is where each one stands, read from the package source rather than from a blog:

APIStatus on 26 September 2026Replacement named by LangChain
ConversationBufferMemory and the other langchain.memory classesNot in langchain 1.x; shipped in langchain-classic, deprecated since 0.3.1, removal in 2.0.0create_agent with a checkpointer or the Store API
ConversationChainlangchain-classic onlycreate_agent
RunnableWithMessageHistoryDeprecated since langchain-core 1.3.3 (5 May 2026), removal in 2.0.0"LangGraph's built-in persistence"
BaseChatMessageHistory, InMemoryChatMessageHistoryDeprecated since langchain-core 1.6.4 (21 September 2026), removal in 2.0.0The short-term memory docs

Nothing here breaks tomorrow. LangChain's release policy says deprecated features keep working through the whole 1.x series, and no 2.0 date has been announced. The practical reading is that new code should not be built on these classes, and existing code has a migration on the roadmap. The mapping is mechanical: buffer, window and summary memory become a checkpointer plus trimming or SummarizationMiddleware; entity memory and "facts about the user" become a store or a memory layer; ConversationChain becomes create_agent.

Where the built-in memory stops

LangChain's persistence layer does its job well. A Postgres checkpointer will carry a thread through restarts and interruptions, and a store will keep whatever you put in it and search it by meaning. Both are storage primitives, though, and the hard part of memory is not storage.

Four decisions are left to your code. The store saves the JSON you hand it, so something has to read the conversation and decide that "we moved to the Growth plan last month" is a fact worth saving and "thanks, that helps" is not. Something also has to settle what happens when facts collide, and LangChain's own overview is candid here: it warns that updating a single profile document gets error-prone as the profile grows, and that with a collection of memories "the model must now delete or update existing items in the list, which can be tricky"; it points to a separate package and to evaluation. Isolation is the third decision, because a namespace tuple is a naming convention, so keeping one tenant's memories out of another tenant's session is only as strong as every line of code that builds a namespace. Retrieval is the fourth: which of the hundreds of items stored for a user belong in this prompt, ranked how, inside what token budget.

Owning those four decisions means building a memory service, which is a reasonable project (the real cost of DIY agent memory lays out what it involves). The alternative is to keep LangChain for orchestration and hand the four decisions to a memory layer.

Step 4: add Maximem Synap with maximem-synap-langchain

Maximem Synap is a memory service for AI agents. Your app sends it conversation turns as they happen; Synap reads them, decides what is worth keeping and stores discrete, self-contained statements (facts, preferences, episodes and other types, each with a confidence score, a scope and a timestamp) rather than transcript chunks. Before the agent replies, you ask for context and get back a short ranked set of memories ready for the prompt. The maximem-synap-langchain package wires that into LangChain's own interfaces: a message history, a retriever, two agent tools, a callback handler, plus a runnable for the current conversation's compacted history. More on the design is in how Maximem Synap works.

Install and initialise

pip install maximem-synap-langchain langchain langchain-openai
export SYNAP_API_KEY=synap_...   # from the dashboard at synap.maximem.ai

Install it as maximem-synap-langchain and import it as synap_langchain. The SDK is async and is created once per process:

import os
from maximem_synap import MaximemSynapSDK

sdk = MaximemSynapSDK(api_key=os.environ["SYNAP_API_KEY"])
await sdk.initialize()   # once at startup
# ...
await sdk.shutdown()     # once at exit

One rule to learn before the first error: Synap conversation IDs must be UUID strings. record_message rejects a value such as "conv-123" with InvalidConversationIdError, so generate str(uuid.uuid4()) per conversation, or derive a stable UUID from your own session key with uuid.uuid5.

Recipe A: a conversational chain that remembers the user

For a chat chain (prompt, model, no tools), the cleanest setup reads memory before the model call and records the turn after it. SynapRetriever fetches memories about this user from earlier conversations, synap_st_runnable returns the current conversation's history in compacted form, and SynapCallbackHandler records the user's message and the model's reply without any change to your chain:

from operator import itemgetter

from langchain.chat_models import init_chat_model
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.runnables import RunnableLambda, RunnablePassthrough
from synap_langchain import SynapCallbackHandler, SynapRetriever, synap_st_runnable

llm = init_chat_model("openai:gpt-5.5")

prompt = ChatPromptTemplate.from_messages([
    ("system",
     "You are a support assistant.\n\n"
     "What you know about this user from earlier conversations:\n{memories}\n\n"
     "This conversation so far:\n{history}"),
    ("human", "{question}"),
])


def format_memories(docs):
    return "\n".join(f"- {d.page_content}" for d in docs) or "Nothing stored yet."


def make_chat(sdk, user_id: str, conversation_id: str):
    retriever = SynapRetriever(sdk=sdk, user_id=user_id, mode="fast", max_results=8)
    # The retriever raises SynapIntegrationError on failure; fall back to no memories.
    safe_retriever = retriever.with_fallbacks([RunnableLambda(lambda _: [])])

    chain = (
        RunnablePassthrough.assign(
            memories=itemgetter("question") | safe_retriever | format_memories,
            history=synap_st_runnable(sdk, conversation_id),
        )
        | prompt
        | llm
    )
    recorder = SynapCallbackHandler(sdk=sdk, conversation_id=conversation_id, user_id=user_id)
    return chain.with_config(callbacks=[recorder])


# per conversation
import uuid
chat = make_chat(sdk, user_id="alice", conversation_id=str(uuid.uuid4()))
reply = await chat.ainvoke({"question": "How should I handle refresh tokens?"})

Each SynapRetriever result is a LangChain Document whose metadata carries the memory type, the scope it came from and a confidence or strength score, so you can filter or format by type. mode="fast" is built for a live conversation; mode="accurate" breaks a compound question into parts and follows relationships between entities, at a higher latency, which suits a background job or an explicit lookup.

Recipe B: a create_agent agent that decides when to use memory

For a tool-calling agent, give the model the two memory tools and let it choose. SynapSearchTool appears to the model as search_memory and SynapStoreTool as store_memory; both are ordinary BaseTools, so they drop into create_agent next to your other tools, and they sit happily beside a checkpointer that keeps the thread:

import uuid

from langchain.agents import create_agent
from langgraph.checkpoint.memory import InMemorySaver
from synap_langchain import SynapSearchTool, SynapStoreTool

user_id = "alice"
conversation_id = str(uuid.uuid4())

agent = create_agent(
    "openai:gpt-5.5",
    tools=[
        SynapSearchTool(sdk=sdk, user_id=user_id, mode="fast"),
        SynapStoreTool(sdk=sdk, user_id=user_id),
    ],
    system_prompt=(
        "You are a support assistant. Call search_memory when the user refers "
        "to something from an earlier conversation. Call store_memory when the "
        "user tells you a durable fact or preference."
    ),
    checkpointer=InMemorySaver(),  # thread memory, as in Step 1
)


async def ask(text: str) -> str:
    result = await agent.ainvoke(
        {"messages": [{"role": "user", "content": text}]},
        config={"configurable": {"thread_id": conversation_id}},
    )
    answer = result["messages"][-1].text
    # Record the finished turn so Synap can extract from it.
    await sdk.conversation.record_message(
        conversation_id=conversation_id, role="user", content=text, user_id=user_id)
    await sdk.conversation.record_message(
        conversation_id=conversation_id, role="assistant", content=answer, user_id=user_id)
    return answer

Those explicit record_message calls are deliberate. SynapCallbackHandler records the latest human message every time a chat model starts, which is exactly right for a single-call chain; inside a tool loop the model is called more than once per user turn, and in our test the same user message was recorded once per model call. Record the finished turn yourself in agents, and use the callback handler for chains.

Which component does which job

ComponentLangChain interfaceUse it for
SynapCallbackHandlerAsyncCallbackHandlerRecording every turn of a chain automatically
SynapRetrieverBaseRetrieverPulling memories about the user into the prompt before the model call
synap_st_runnableLCEL RunnableThe current conversation's compacted history as a prompt variable
SynapSearchTool, SynapStoreToolBaseToolLetting an agent look things up or save something explicitly
SynapChatMessageHistoryBaseChatMessageHistoryExisting code built on RunnableWithMessageHistory

Should memory be automatic or the agent's decision? Both, split by direction. Recording should be automatic, because a model that has to remember to save things will forget to on the turns that matter, and Synap decides what to keep anyway. Retrieval before each reply should be automatic too. The tools are for the cases the automatic read does not anticipate: the user asks about something from months ago, or says "remember this".

If you already run RunnableWithMessageHistory, SynapChatMessageHistory slots in as the history factory (lambda session_id: SynapChatMessageHistory(sdk=sdk, conversation_id=session_id, user_id="alice"), with a UUID session ID). It works on langchain-core 1.6.5, and it inherits the deprecation of the two LangChain classes it builds on, so treat it as a bridge while you move to Recipe A or B.

Scoping per user and per tenant

Every component takes user_id and an optional customer_id. Maximem Synap applies isolation from those identifiers at write time across three levels (client, which is your account; customer, which is your tenant; user, which is one person) plus the conversation. A read sees the narrowest applicable scope and everything above it, never below and never sideways, so one tenant's memories do not surface in another tenant's session and you do not write namespace logic to make that true. Pass customer_id on a multi-tenant (B2B) instance; the Synap docs note that a single-tenant instance rejects it.

Failure, latency and cost

Memory sits in the path of every reply, so it matters what each piece does on a bad day. In version 0.3.0 of the package the split is: SynapCallbackHandler never raises and logs at ERROR; synap_st_runnable returns an empty string by default; SynapChatMessageHistory returns an empty history on a failed read but raises on a failed write; SynapRetriever and the two tools raise SynapIntegrationError, which is why Recipe A wraps the retriever in with_fallbacks. On the service side, Maximem Synap narrows a retrieval rather than failing it when part of the search has a problem, flags that it narrowed, and retries transient errors with backoff before raising a typed error.

Writes are asynchronous. record_message returns immediately and extraction runs in the background for several seconds, so a fact written now is not retrievable a second later. That is right for production and worth remembering in a demo.

Reads are designed for the hot path. While a conversation is going, Synap works out what the agent is likely to need next and pushes it into a cache inside your own process, so most in-conversation reads return without a network call; Maximem asserts an in-conversation P75 retrieval latency under 15 ms. Cost is per operation: a fast retrieval is 1 credit and an accurate one 3, ingestion scales with document length, and credits display at $0.00175 each. For where the data lives: managed cloud in the United States is the default, and self-hosted or air-gapped deployment is available for enterprise engagements.

Controlling what gets extracted: extraction prompts and the alternative

If you build memory on a LangGraph store, you will write an extraction prompt, because something has to turn a conversation into the items you put. A good one has a short list of allowed memory types, a rule that each memory is one self-contained sentence with no unresolved pronouns, examples that return nothing (small talk, thanks, questions the user asked), examples that return something, and structured output so parsing never fails. A minimal version for a support agent:

import uuid
from pydantic import BaseModel, Field
from langchain.chat_models import init_chat_model


class Memories(BaseModel):
    memories: list[str] = Field(default_factory=list)


EXTRACTION_PROMPT = """Extract durable facts about the user from the conversation.
Keep only: account and plan details, stated preferences, commitments with dates.
Write each memory as one standalone sentence that names the user explicitly.
Return an empty list for greetings, thanks and questions with no new information.

Example: "thanks, that fixed it" -> []
Example: "we moved to the Growth plan last month" -> ["The user moved to the Growth plan in August 2026."]

Conversation:
{conversation}"""

# transcript: the conversation as text; store: your LangGraph store; user_id: the person
extractor = init_chat_model("openai:gpt-5.4-mini").with_structured_output(Memories)
result = extractor.invoke(EXTRACTION_PROMPT.format(conversation=transcript))
for text in result.memories:
    store.put((user_id, "memories"), str(uuid.uuid4()), {"text": text})

That works, and it is where the real work starts. The prompt has no idea that "the Growth plan" retires "the Starter plan" from last month; you need an update pass for that. It writes the same fact again every time the user repeats it; you need deduplication. It was written for one agent, and your next agent (an onboarding assistant, a voice agent) needs different categories. And every change to the prompt changes what is in the store, so you need an evaluation set to know whether the change helped.

Maximem Synap does not take a hand-written extraction prompt. It generates a memory architecture for each agent from a description of what that agent does: what gets extracted, how aggressively, into which categories, how it is scoped and stored, how it is retrieved and ranked, and how long it is kept. The architecture is authored by Anthropic's Claude Opus and checked by two further frontier-model judges, it is never applied until a person reviews and approves it, and it is versioned so you can roll it back. A healthcare agent and a food-delivery agent running the same code end up extracting different things, which is the problem an extraction prompt tries to solve one agent at a time. Collisions are handled at the memory level: a new statement that replaces an old one marks the old one historical and linked to its replacement, and a poorer statement ("takes a cholesterol medication") never overwrites a richer one ("takes 20mg of Atorvastatin daily").

If you are on LangGraph

Everything above uses create_agent, which runs on LangGraph underneath. If you build graphs directly, the same split applies: compile with a checkpointer for thread state and a store for cross-thread data. For Maximem Synap, the companion maximem-synap-langgraph package provides SynapStore, a LangGraph BaseStore, so builder.compile(checkpointer=checkpointer, store=SynapStore(sdk, user_id="alice")) gives your nodes Synap-backed long-term memory through the store interface they already use; use search as the read path. The package's own documentation describes its SynapCheckpointSaver as best-effort and recommends pairing it with a real key-value saver such as PostgresSaver for checkpoint fidelity, so keep Postgres for thread state. The LangGraph integration post has the setup.

How to test conversational memory

Test the framework behaviour and the memory behaviour separately, because they fail differently. For the framework: messages accumulate within one thread_id, two thread IDs stay separate, a restart with a durable checkpointer resumes the thread, and tool calls stay paired with their results after trimming. For memory: a fact from conversation one is available in conversation two, an unrelated user never sees it, a correction ("actually, I moved to Berlin") wins over the original, and a memory is only asserted as retrievable after the write has had time to process.

Most of this runs offline: FakeMessagesListChatModel from langchain_core returns scripted replies, including tool calls, which is how we ran every Maximem Synap example in this post against the real installed packages before publishing.

The short version

For a new LangChain app in September 2026: build with create_agent, give it a Postgres checkpointer and a UUID thread_id per conversation, add SummarizationMiddleware once threads get long, and decide how the app will remember people across conversations. If that means a LangGraph store, budget for the extraction, update and evaluation work described above. If it means Maximem Synap, add SynapRetriever and SynapCallbackHandler to a chain, or the two Synap tools plus an explicit record_message to an agent, and pass a user_id everywhere. Leave ConversationBufferMemory and RunnableWithMessageHistory out of new code.

LangChain now gives you well-built, versioned primitives for thread state and a place to keep user data, and leaves the judgement about what to remember to you. Put that judgement in a memory layer and the agent gets better with every conversation instead of every sprint, and your users stop repeating themselves.

Frequently asked questions

Does LangChain have built-in conversational memory?

Yes. In LangChain 1.x, thread-level memory comes from passing a checkpointer to create_agent and reusing a thread_id, and cross-conversation memory comes from passing a LangGraph store. The pre-1.0 memory classes such as ConversationBufferMemory are not in the langchain package any more; they live in langchain-classic, marked deprecated with removal planned for 2.0.

Is RunnableWithMessageHistory deprecated?

Yes. langchain-core marks RunnableWithMessageHistory deprecated since version 1.3.3 (May 2026), and BaseChatMessageHistory since version 1.6.4 (21 September 2026), both with removal scheduled for 2.0.0. They still work throughout 1.x under LangChain's release policy, but new code should use a checkpointer with create_agent.

How do I make LangChain remember a user across different conversations?

Use a store rather than the checkpointer. A checkpointer is keyed by thread_id, so a new conversation starts empty; a LangGraph store is keyed by a namespace you control, such as the user ID, and any thread can read it. A memory layer such as Maximem Synap goes further by extracting the facts from each conversation for you and returning the relevant ones before each reply.

How do I add Maximem Synap to a LangChain app?

Install maximem-synap-langchain, create and initialise one MaximemSynapSDK with your API key, then add SynapRetriever and SynapCallbackHandler to a chain, or SynapSearchTool and SynapStoreTool to a create_agent agent. Pass a user_id to every component and a UUID conversation_id per conversation, and add customer_id on a multi-tenant instance.

Can I write a custom prompt for memory extraction in LangChain?

Yes, if you build extraction yourself: prompt a model with allowed memory types, standalone-sentence rules and empty examples, parse the output with with_structured_output, and write the results to a store. You then also own deduplication, contradiction handling and evaluation. Maximem Synap takes a different route and generates the extraction behaviour from a description of your agent, with a person approving it before it applies.

Does adding memory slow down my LangChain agent?

It adds whatever the read costs in the hot path. A database checkpointer is one read per turn. A semantic store search adds an embedding call and a vector query. Maximem Synap pre-fetches likely context into a cache inside your process and asserts an in-conversation P75 retrieval latency under 15 ms, while writes return immediately and extraction runs in the background.

Sources, retrieved 26 September 2026: LangChain short-term memory; LangChain long-term memory; LangGraph persistence; LangGraph stores; LangChain memory overview; LangChain release policy; langchain-core on PyPI (1.3.3, 1.6.4, 1.6.5 release history and source); langchain on PyPI; langchain-classic on PyPI; maximem-synap-langchain on PyPI (0.3.0); maximem-synap-langgraph on PyPI (0.4.0); Maximem Synap LangChain docs; Maximem Synap.

From the team at Maximem

Stop rebuilding agent memory from scratch

Maximem Synap is the context management layer we built after hitting every problem in this post ourselves. Persistent recall across sessions, entity resolution and conscious forgetting, in Python, TypeScript and REST.

Related posts