Sessions and file-based memory are what the Claude Agent SDK ships for persistence. A session is the conversation transcript, written to local disk; continue, resume and fork return to it, and a SessionStore adapter mirrors it to S3, Redis or a database so that another host can pick it up. CLAUDE.md files and auto memory load notes per working directory. Both are keyed to a session ID or a directory and not to an end user of your product, so nothing in the SDK carries one user's facts into their next session or revises them when they change. Maximem Synap fills that gap through the SDK's own extension points: a hook that fetches context scoped to user_id and customer_id before each prompt, and an in-process MCP server that lets the agent search and write memory itself.
Where Claude Agent SDK Excels
Anthropic packaged the agent loop behind Claude Code as a library for Python and TypeScript. Built-in tools read and edit files, run commands and search the web; hooks run your code at fixed points in the loop; subagents take focused subtasks; MCP connects external tools; and permissions decide which tool calls run without approval. If your agent has to work through a multi-step task with real tools, the SDK handles the orchestration well.
What is left open is what the agent knows about a user who comes back tomorrow in a new session.
How Claude Agent SDK Memory Works Today
Two mechanisms carry state between runs. Sessions come first. A session contains the prompt, every tool call, every tool result and every response, and the SDK writes it to ~/.claude/projects/<encoded-cwd>/ as a JSONL file automatically. That file survives a process restart. continue_conversation=True in Python (continue: true in TypeScript) picks up the most recent session in the working directory, resume takes a specific session ID, and fork_session copies a session's history into a new one.
Session files are local to the machine that created them. For serverless functions, autoscaled workers and containers, a SessionStore adapter mirrors each transcript batch to your own backend, and the TypeScript SDK repository carries reference adapters for S3, Redis and Postgres. The store is a mirror and not a replacement, and mirror writes are best-effort.
File-based memory is the second mechanism. With default options, query() reads CLAUDE.md files and rules as the Claude Code CLI does; auto memory under ~/.claude/projects/<project>/memory/ is loaded into the system prompt at session start, and the agent adds to it with the ordinary Write and Edit tools. For multi-tenant deployments, Anthropic's guidance is to give each tenant its own filesystem and turn these inputs off (settingSources: [] and CLAUDE_CODE_DISABLE_AUTO_MEMORY=1), because an SDK process otherwise picks up host-level configuration and per-directory memory.
All of this works, and all of it is addressed by session ID or working directory. You can map one long-lived session to each user and resume it, which is a reasonable design for a single ongoing thread; over time, automatic compaction replaces the early turns with a summary, and a second conversation with the same user still starts from an empty transcript. What the SDK does not have is a store of facts about a person that every session can query, that links the different names and identifiers the person appears under, and that is corrected when a fact changes. Those are memory-layer problems.
What Synap Adds
Synap is agentic context management. It does not replace sessions or CLAUDE.md. It adds a per-user store beside them, and the package (synap_claude_agent in Python, @maximem/synap-claude-agent in TypeScript) connects to the SDK at two documented plug points.
create_synap_hooks returns a hooks dictionary for ClaudeAgentOptions. Its UserPromptSubmit hook calls Synap with the prompt as the search query, scoped to the user_id and customer_id you bound at construction, and returns the result as additionalContext, which the SDK adds to what Claude sees for that turn. The same hook records the user's prompt to Synap under the SDK's session ID. Recording the assistant's reply needs an open Synap stream (sdk.instance.listen()): the TypeScript package then reports it from the Stop hook, while in Python the hooks do not report it, so you call report_assistant_turn once the run returns.
create_synap_mcp_server returns an in-process MCP server with two tools. synap_search takes a query and returns formatted context; synap_remember stores an explicit fact. Claude sees them as mcp__synap__synap_search and mcp__synap__synap_remember, and neither tool has a user argument, so the model cannot reach another user's memory by changing a parameter.
A third helper, create_synap_st_hook, injects Synap's short-term context (a compacted summary plus recent turns) for a conversation ID you pass explicitly. Retrieval has two modes: fast (vector-only, 50 to 100ms) and accurate (graph traversal + reranking, 200 to 500ms). The hooks and the MCP server default to accurate and take a mode argument.
We built this because agents kept stalling in production, and the cause was rarely tool logic; it was context that lived in a different session last Tuesday.
Production testing hit 92% LongMemEval, 93.2% on LoCoMo. Fast mode retrieves in under 100ms.
For why context management is infrastructure and not a feature, read What Is Agentic Context Management?
For build-versus-buy numbers, see The Real Cost of DIY Agent Memory
Technical Deep Dive
LongMemEval Benchmark -
LongMemEval tests whether agents recall facts across long, multi-turn conversations spanning multiple sessions. The benchmark simulates production conditions where users return days apart and expect the agent to remember prior context. Synap scores 92% on this benchmark. Baseline vector-only approaches typically score 60-70%. The gap comes from entity resolution and temporal awareness that pure vector search lacks.
Entity Resolution Mechanism
Synap tracks identity across 15 reference patterns: names, emails, phone numbers, account IDs, session IDs, API keys, and more. When an agent encounters "John" in one session and "[email protected]" in another, the resolution engine runs deterministic matching on structured fields, then probabilistic matching on unstructured references. Conflicts are resolved using temporal recency and source confidence scores. The result is a single canonical entity that accumulates context across all identifiers.
Graph Traversal in Accurate Mode
Fast mode retrieves by vector similarity alone. Accurate mode adds a graph layer that traverses relationships between entities. If you ask about "the project John mentioned," the graph finds John, traverses to projects linked to John, and returns the relevant context. This adds latency but catches connections that vector similarity misses. Reranking then scores results by recency, confidence, and query relevance.
Multi-Tenant Scoping
Synap's scope chain narrows from client to customer to user. user_id identifies the person and customer_id identifies the tenant, so a SaaS deploying agents for multiple customers uses customer_id to ensure tenant A never sees tenant B's memory. On a B2B instance customer_id is required; on a B2C instance it is rejected, because there the user is the whole identity. conversation_id is not a scope tier. It groups turns inside a scope and has to be a UUID, which the Agent SDK's session IDs already are. This scoping is enforced at the storage layer and not only in application logic.
What Synap Adds to Claude Agent SDK
Persistence
Claude Agent SDK Native. Session transcripts on local disk, resumable after a restart. A SessionStore adapter mirrors them to your own backend.
With Synap. Per-user memory survives across sessions and restarts.
Entity Resolution
Claude Agent SDK Native. Transcripts and Markdown notes. The SDK documentation describes no entity linking.
With Synap. "John" and "[email protected]" resolve to one canonical entity across every session.
Compaction
Claude Agent SDK Native. Automatic inside a session. Earlier turns are replaced by a summary.
With Synap. Server-side compaction with configurable levels (aggressive, balanced, conservative, adaptive), delivered to the agent through create_synap_st_hook.
Retrieval Latency
Claude Agent SDK Native. No retrieval step. Resume loads the stored transcript.
With Synap. 50 to 100ms fast mode. 200 to 500ms accurate mode.
Long-Term Recall
Claude Agent SDK Native. The SDK documentation gives no cross-session recall figure.
With Synap. 92% on LongMemEval.
Failure Handling
Claude Agent SDK Native. A failed SessionStore write is retried, then logged and dropped. The query continues.
With Synap. A failed fetch is logged and nothing is injected. synap_search answers that no context is available, and synap_remember returns a tool error. Your agent keeps running.
User Scoping
Claude Agent SDK Native. Session ID and working directory.
With Synap. user_id and customer_id bound at construction. conversation_id is optional and falls back to the session ID.
What Production Teams Gain
Cross-session continuity. Your user chats on Monday and returns on Wednesday in a new session. The hook fetches what Synap holds for that user before Claude sees the prompt, so Monday's facts arrive without resuming Monday's transcript.
Accuracy that ships. 92% LongMemEval, 93.2% on LoCoMo measures whether agents recall facts across long, multi-turn conversations spanning multiple sessions.
Latency that does not block. Fast retrieval: 50 to 100ms. Accurate mode with graph traversal and reranking: 200 to 500ms. Both degrade without crashing. A failure returns empty results and a log line, not a broken agent.
How to Get Started
Setup
Python
pip install maximem-synap-claude-agent
TypeScript
npm install @maximem/synap-claude-agent @maximem/synap-js-sdk @anthropic-ai/claude-agent-sdk zod
Configure your API key. Generate one from the Synap Dashboard.
.env
SYNAP_API_KEY=synap_your_key_here
ANTHROPIC_API_KEY=your-anthropic-api-key
Initialize the Synap SDK once at application startup. See SDK Initialization for the full lifecycle and configuration options.
Basic integration The smallest useful integration uses the hooks for automatic context injection and adds the MCP server so that the agent can search and store memory on its own:
Python
import asynciofrom claude_agent_sdk import ClaudeAgentOptions, ResultMessage, query from maximem_synap import MaximemSynapSDK from synap_claude_agent import create_synap_hooks, create_synap_mcp_server
async def main(): sdk = MaximemSynapSDK() # reads SYNAP_API_KEY await sdk.initialize()
options = ClaudeAgentOptions( hooks=create_synap_hooks( sdk, user_id="alice", customer_id="acme", # B2B instances only; leave out on B2C ), mcp_servers={ "synap": create_synap_mcp_server(sdk, user_id="alice", customer_id="acme"), }, allowed_tools=["mcp__synap__synap_search", "mcp__synap__synap_remember"], ) async for message in query( prompt="What did I tell you about my trial account?", options=options, ): if isinstance(message, ResultMessage) and message.subtype == "success": print(message.result) await sdk.shutdown()
asyncio.run(main())
TypeScript
import { query } from "@anthropic-ai/claude-agent-sdk"; import { SynapClient } from "@maximem/synap-js-sdk"; import { createSynapHooks, createSynapMcpServer } from "@maximem/synap-claude-agent";const sdk = new SynapClient({ apiKey: process.env.SYNAP_API_KEY }); await sdk.initialize();
for await (const message of query({ prompt: "What did I tell you about my trial account?", options: { hooks: createSynapHooks({ sdk, userId: "alice", customerId: "acme" }), mcpServers: { synap: createSynapMcpServer({ sdk, userId: "alice", customerId: "acme" }), }, allowedTools: ["mcp__synap__synap_search", "mcp__synap__synap_remember"], }, })) { if (message.type === "result" && message.subtype === "success") { console.log(message.result); } }
On each prompt, the hook wraps the fetched context in a <synap_memory> block; when Synap returns nothing, nothing is injected. Context fetch failures degrade gracefully: the error is logged and the turn proceeds without memory. Explicit writes surface their failures, because synap_remember reports an ingestion error back to the agent. The full reference is in the Claude Agent SDK integration docs.
Memory Is Infrastructure
Anthropic gave developers the agent loop behind Claude Code as a library, and its sessions hold a single conversation well. Knowing a user across conversations, resolving who they are, and retrieving the right facts for the current prompt is a different problem.
Teams either build that infrastructure themselves, or they plug in a system built for it; what the second route buys is an agent that improves with each conversation and users who stop repeating themselves.
This is why memory is infrastructure, not a feature.
Start building Claude Agents that remember across sessions → (https://synap.maximem.ai)
Synap pricing is usage-based. You pay for memory operations: storage, retrieval, compaction. No per-seat or per-framework surcharge. Starter plan: $49/month. Every new account gets $25 in free credits to test before committing. See full pricing at https://synap.maximem.ai/pricing.
Related Posts
- What Is Agentic Context Management?
- The Real Cost of DIY Agent Memory
- Skills Are the New Microservices



