Published 27 September 2026 · Model and embedding prices, PostgreSQL 19 status, benchmark definitions and Maximem Synap documentation checked on 26 September 2026. The cost arithmetic and its assumptions are in the section on what a graph costs.
For agent memory, neither wins on its own. Vector search is the right tool for finding past statements that mean something close to the current question; a knowledge graph is the right tool once the agent has to connect facts through the same person or company, or tell which of two conflicting facts is current. Start with vector and keyword search over self-contained statements, and add graph expansion when you can name a class of questions that needs it. Maximem Synap is built on that split: it stores each memory as a discrete statement rather than a transcript chunk, resolves every mention of a person to one entity, records a changed fact as a change instead of overwriting it, and expands through the entity graph in its Accurate retrieval mode, when a question spans several entities.
A temporal knowledge graph is the version of this that also records when each fact was true, and it matters for the same reason: agents are asked "what is true now" and "what was true then", and a memory that cannot tell the two apart will answer both with whichever fact ranks higher.
What each approach actually answers
Vector search turns text into embeddings (arrays of numbers that place similar meanings near each other) and returns the stored items nearest to the query. Most stores do this approximately for speed: pgvector, for instance, runs exact search by default and adds HNSW or IVFFlat indexes for approximate search, "which trades some recall for speed." The question vector search answers is "what have I stored that is similar to this?"
A knowledge graph stores entities as nodes and the relationships between them as typed, directed edges: Priya -WORKS_AT-> Acme, Acme -ON_PLAN-> Enterprise, Enterprise -RENEWS_ON-> 2027-03-01. Retrieval starts at an entity and follows edges. The question a graph answers is "what is connected to this, how, and what is true about it?"
Agent memory needs both kinds of question. A support agent asked "have we seen this error before?" wants similarity. The same agent asked "is this customer's renewal affected by the outage on their plan?" wants three facts joined through two entities, and no single stored sentence is similar to that question.
Where vector search breaks for agent memory
Vector search fails in four predictable ways once memory spans many conversations, and none of them is fixed by a better embedding model.
Similar is not current. "I live in Chicago" and "I moved to Berlin in June" embed close together, so both come back for "where does the user live?", ranked by similarity rather than by which one is still true. A retrieval layer over transcripts holds the fact as stated in March and the correction in June, and has no view about which is current; the model is left to sort it out at read time, on every call.
Chunks have no identity. "My manager", "Sarah" and "Sarah Chen" are three strings in three chunks. Nothing in a vector index knows they are one person, so a question about Sarah retrieves a third of what the agent knows about her.
Multi-hop answers are not similar to the question. "Which of this user's projects depend on the service we deprecated last month?" needs the user's projects, each project's dependencies and the deprecation date. Each fact is stored somewhere; none of them is individually close to the question, so top-k similarity misses the chain.
Exact terms match poorly by meaning. Order numbers, error codes, API names and SKUs are where embeddings blur. In our retrieval study (50,000 documents across five datasets, 5,000 queries, scored on MRR@10 by exact document match with no model judge), vector search led the five-dataset average on MRR@10 (0.6320 against 0.5325), but with the code dataset removed, plain keyword search led the average of the other four (0.5931 against 0.5614), and it indexed in 2.11 seconds against 161.6 seconds for embeddings. This one is fixed by adding keyword search alongside vectors, not by adding a graph.
What a knowledge graph costs you
A graph fixes the first three failures, and it charges for it on every write.
An extraction call per write. Someone has to turn "Sarah said Acme is moving to the annual plan next quarter" into entities and edges, and in agent memory that someone is a language model. At prices retrieved on 26 September 2026, the gap to embedding is large. Assume a conversation turn adds 400 tokens, and an extraction call sends 2,000 input tokens (the turn, the extraction instructions and the candidate entities it might match) and returns 250 tokens of entities and relationships:
| Per conversation turn | Price basis | Cost per turn | Per million turns |
|---|---|---|---|
| Embed the turn | OpenAI text-embedding-3-small, $0.02 per million tokens | $0.000008 | $8 |
| Extract a graph with Claude Haiku 4.5 | $1 input, $5 output per million tokens | $0.00325 | $3,250 |
| Extract a graph with Claude Sonnet 5 | $2 input, $10 output per million tokens | $0.0065 | $6,500 |
That is roughly 400 times the embedding cost on the cheaper model, before storage and before any follow-up call to settle an ambiguous entity match. For a customer-facing agent where a wrong fact costs a customer, it is a small price. For an internal tool processing millions of low-stakes turns, it may not pay back.
Errors become durable. A vector index that retrieves a poor chunk forgets about it by the next query. An extraction that writes Priya -WORKS_AT-> Initech from a hypothetical ("if I worked at Initech...") has created a fact that every later traversal will follow.
Entity resolution is the hard part. A graph with three nodes for Sarah is worse than no graph, because traversals from each node see a third of her relationships and report them confidently. "Acme", "Acme Corp" and "ACME Corporation" have to merge; "Sarah (my manager)" and "Sarah (my sister)" must not. Doing this well means matching on names, aliases and context, and routing the genuinely ambiguous cases to a person rather than guessing.
Empty on day one, and slower on the hot path. A new user has no graph, so early recall leans on similarity anyway. And each hop is another lookup, while expanding a neighbourhood multiplies candidates to rank; if the traversal sits inside every conversational turn, measure its latency before you ship it.
Knowledge graph vs vector search at a glance
| Vector search | Knowledge graph | |
|---|---|---|
| Stores | Embeddings of text (chunks or statements) | Entities, typed relationships, properties |
| Answers | What is similar to this? | What is connected to this, how, and what is true now? |
| Strong at | Paraphrase, fuzzy recall, "have we discussed this before?" | Multi-hop questions, one identity across many names, changing facts |
| Weak at | Current vs outdated facts, identity, multi-hop, exact identifiers | Cold start, extraction errors, cost per write, traversal latency |
| Cost per write | One embedding call | One LLM extraction call, plus entity resolution |
| Useful from | The first stored item | Once enough entities and relationships exist |
Exact identifiers belong to neither column; they are keyword search's job, which is why production memory usually runs hybrid search.
The hybrid that holds up in production
A pattern that holds up uses each index for what it is good at, in sequence. Embed the query and run similarity and keyword search to find entry points: the statements and entities most relevant to the question. From those entities, expand along relationships to pick up connected facts that were never similar to the question. Merge the candidates, rank them on relevance, recency and confidence, drop anything that has been superseded, and trim to a token budget before it reaches the prompt.
Microsoft Research's GraphRAG uses the same local-search move over documents, starting from entities and traversing their neighbourhood, and adds community summaries for corpus-wide questions, because "RAG fails on global questions directed at an entire text corpus, such as 'What are the main themes in the dataset?'" Agent memory rarely asks corpus-wide questions. It asks about specific people and their changing circumstances, which is why the expansion step matters more than the summaries.
When to add a graph
Start with vector plus keyword search over clean, self-contained statements. Add relationship expansion when at least one of these is true, and you can point at failing conversations that show it:
- The agent regularly needs two or more facts joined through an entity ("the plan of the company this user works for").
- The same person, company or project shows up under several names across conversations or channels.
- Facts change, and wrong answers come from the agent picking the outdated one.
- Someone needs to ask why the agent believes something, and "it was similar" is not an acceptable answer.
If none of these shows up in your failure logs, a graph adds cost and moving parts without changing answers. A single-session support bot, or an agent whose stored facts are independent and stable ("prefers email", "account type: enterprise"), is well served by similarity search.
The choice underneath: chunks or statements
Indexes get the attention, but the unit you store decides more. A chunk of transcript contains the fact you want and four you do not, it contains the March version and the June correction side by side, and it cannot be compared with another chunk in any meaningful way. A self-contained statement ("The user moved to Berlin in June 2026") can be compared with an existing one ("The user lives in Chicago") and a system can decide whether the new one replaces it, adds to it or duplicates it. Put chunks in a graph and you get a graph of chunks; put statements in a vector index and a lot of the "similar is not current" problem goes away, because superseded statements can be marked and filtered before ranking.
Replacement has its own trap, and it is the one systems rarely advertise. If "the user takes 20mg of atorvastatin daily for cholesterol" is replaced by a later, vaguer "the user takes a cholesterol medication", the memory has lost the dose while appearing to update. Maximem Synap will not make that trade: when a newer statement is poorer than the one it resembles, both survive. That rule sits in the update step, and no choice of index provides it.
What is a temporal knowledge graph, and does agent memory need one?
A temporal knowledge graph is a knowledge graph in which every fact carries time: when it became true, when it stopped being true, and when the system learned it. Where an ordinary knowledge graph stores Priya -WORKS_AT-> Acme as if it were permanent, a temporal one stores that edge with a validity period, so it can answer "where does Priya work now?" and "where did Priya work when she signed the contract?" and get different, correct answers. The research literature defines it the same way; a 2024 survey of temporal knowledge graph question answering writes a temporal knowledge graph as entities, relations, timestamps and facts, with facts often stored as quadruples (subject, relation, object, time). Much of that research is about forecasting which links will appear next; agent memory needs something plainer, reliable recall of what holds now and what held then.
Two clocks, not one
Database research settled this decades before AI agents. Temporal databases distinguish valid time, "the time period during or event time at which a fact is true in the real world", from transaction time, "the time at which a fact was recorded in the database." Keeping both is called bitemporal. The two clocks disagree all the time in agent memory: a user mentions in September that they changed jobs in June; a backfill imports last year's support tickets today. A memory that only knows when it stored something will treat the imported ticket as the newest truth.
Invalidate, do not delete
When a new fact contradicts an old one, a temporal graph closes the old fact's validity period instead of deleting it. The current answer is the fact whose period is still open; the old one stays queryable for "what was true then" and for audit. It also lets the system separate two cases that look identical in text: a correction ("the meeting was actually on Tuesday") fixes what was recorded, while a real change ("from next week, meetings move to Tuesday") ends one fact and starts another.
Why agents need this
Time-dependent questions are a large share of what makes memory hard. LongMemEval, a 500-question benchmark for long-term chat memory, tests five abilities: "information extraction, multi-session reasoning, temporal reasoning, knowledge updates, and abstention". Two of the five abilities, temporal reasoning and knowledge updates, are exactly the questions a memory without time gets wrong.
You may not need a graph to get it
Time and graphs are separate decisions. If your "as of" questions concern single facts ("what plan was this customer on in January?"), a table with validity columns answers them. SQL:2011 standardised application-time tables (valid time), system-versioned tables (transaction time) and bitemporal tables that combine both, and MariaDB, IBM Db2 and SQL Server implement them. A graph becomes necessary when the time-bound question crosses relationships: who managed this account when the incident happened, which team that person was on then, and which policy applied to that team at the time.
Recency weighting is not a substitute
Boosting newer memories helps ranking, but it cannot tell a new fact from an old event reported late, or a one-off exception ("for this trip, book aisle seats") from a lasting change of preference. Deciding what is current is a write-time judgement about what a new statement does to the old ones; recency is a read-time tiebreak.
What goes wrong in practice
Relative dates get resolved against the wrong anchor ("next Friday" read on the day of ingestion rather than the day it was said), imported history lands as if it were new, unknown dates are silently filled with "now", and two intervals that meet at midnight both look current. A temporal model makes these errors visible; it does not prevent them.
How Maximem Synap handles time
Maximem Synap records, for every memory, both when it was stored and when the thing it describes actually happened, which is not the same date and is often the one that matters. When something changes, nothing is destroyed: the old statement is marked historical and linked to what replaced it, along with the reason, so the history of a fact (where it came from, what it replaced, why it changed, what it used to say) can be read in the dashboard.
Keeping tenants apart when retrieval walks relationships
A similarity search can be filtered by tenant in one place. A traversal follows edges wherever they lead, so if any relationship connects two customers' data (a shared vendor, or a person who changed employers), a two-hop query can cross the boundary unless every hop is scoped. In a multi-tenant product the agent will then confidently use another tenant's fact about a shared entity.
Maximem Synap enforces its hierarchy at write time: every memory belongs to a scope (your account, your customer's organisation, one person, one conversation), and a read prefers the narrowest scope that applies and sees everything above it, never below and never sideways. One person's memory does not reach another person's session, and one tenant's memory does not reach another tenant's, whichever retrieval technique produced the candidate.
Do you need a graph database for this?
Often not. At the scale of one agent's memory, relationships can live in the database you already run. PostgreSQL's recursive queries handle bounded traversals, Apache AGE adds openCypher queries to PostgreSQL 11 through 18, and pgvector keeps embeddings in the same place. A native graph query standard for PostgreSQL (SQL/PGQ) was committed for version 19 and then rolled back; it is not in the PostgreSQL 19 release notes as of 14 September 2026. If you already run Neo4j, its vector index puts similarity search and traversal in one engine. The companion guide, what is a graph database and when should you use one, covers that decision in general.
It is also worth separating agent memory from GraphRAG. GraphRAG builds a graph once over a document collection to answer questions about the documents. Agent memory builds its graph continuously from conversations about the same recurring people, where the difficult parts are updates and identity rather than summarising a corpus.
How to test the choice on your own agent
Benchmarks tell you whether a system handles the categories; your own conversations tell you whether it handles your agent. Build a small fixture before choosing: one fact and its later replacement, a late-arriving old record, a person mentioned under three names, a two-hop question, an exact identifier, and a question whose honest answer is "I do not know". Ask each question twice, as "now" and "as of" a date, and check the evidence the retrieval returned as well as the final answer, because a model can pick the wrong fact even when the right one was retrieved.
For published numbers, check which LongMemEval abilities a result covers, which models answered and judged, and who ran it; our own runs use an open harness so the method can be re-run.
Where Maximem Synap fits
Maximem Synap is a hosted memory service that sits beside your agent: you send conversations as they happen, and before each reply you ask what is known and get back a short, ranked set of statements ready for the prompt. The work described above (deciding what to keep, resolving entities across channels, handling changed facts without losing detail, keeping scopes apart) happens inside it, so the choice between similarity and relationships becomes a per-call parameter rather than an architecture you maintain.
context = await sdk.conversation.context.fetch(
conversation_id="3f6b1a2c-4d5e-6f7a-8b9c-0d1e2f3a4b5c",
search_query=["What did Alice say about the project Bob is leading?"],
max_results=10,
mode="accurate", # "fast" for the live turn; "accurate" follows entity relationships
)
Fast mode is the default and is built for agents that are talking to someone right now; Synap also works out what the agent is likely to need next and pre-fetches it into a cache in your own process, which serves in-conversation reads at an asserted P75 under 15 ms. Accurate mode takes longer and does more: it breaks a compound question into parts, expands through the relationships between entities, and merges what it finds. How Synap works walks through the pipeline, and the real cost of DIY agent memory sets out what building the same pieces yourself involves.
Frequently asked questions
Is a knowledge graph better than vector search for AI agent memory?
Neither is better on its own. Vector search finds past statements similar in meaning to the question; a knowledge graph connects facts through the same entities and can tell which of two conflicting facts is current. Most agents should start with vector and keyword search over self-contained statements and add graph expansion when specific questions need facts joined through a person, company or project.
What is a temporal knowledge graph?
A temporal knowledge graph is a knowledge graph in which every fact carries time: when it became true, when it stopped being true, and usually when the system recorded it. That lets an agent answer both "what is true now?" and "what was true on this date?", and lets it retire an outdated fact without deleting its history.
Why use a temporal knowledge graph for agent memory?
Facts about users change, and agents are asked both current and historical questions. Without time on each fact, a memory returns whichever of two conflicting facts ranks higher. LongMemEval, a 500-question benchmark for long-term chat memory, treats temporal reasoning and knowledge updates as two of its five core abilities for this reason.
Can I use a knowledge graph and vector search together?
Using both is the usual production design: similarity and keyword search find the most relevant statements and entities, the graph expands from those entities to connected facts, and the merged results are ranked and trimmed to a token budget before they reach the prompt.
Do I need a dedicated graph database for agent memory?
Usually not at the scale of one agent's memory. Relationships can live in PostgreSQL using recursive queries or the Apache AGE extension, next to pgvector for embeddings. A dedicated graph database makes sense when traversals are deep, variable in depth and central to the product, or when you already run one.
How much more does a knowledge graph cost than vector search?
Most of the extra cost is an LLM extraction call on every write. With a turn of 400 tokens and an extraction call of 2,000 input and 250 output tokens, embedding costs about $8 per million turns with text-embedding-3-small, while extraction costs about $3,250 per million turns on Claude Haiku 4.5 and $6,500 on Claude Sonnet 5, at prices checked on 26 September 2026.
How does Maximem Synap handle knowledge graphs and changing facts?
Maximem Synap stores discrete statements rather than transcript chunks, resolves "my manager", "Sarah" and "Sarah Chen" into one entity across conversations, and expands through the entity graph in its Accurate retrieval mode. When a fact changes, the old statement is marked historical and linked to its replacement with the reason, and a newer statement that is less detailed never replaces a more detailed one.
Sources, retrieved 26 September 2026: pgvector on GitHub; Neo4j Cypher manual, vector indexes; OpenAI, text-embedding-3-small model page; Anthropic, Claude pricing; Edge et al., From Local to Global: A Graph RAG Approach, arXiv 2404.16130; Su et al., Temporal Knowledge Graph Question Answering: A Survey, arXiv 2406.14191; Wu et al., LongMemEval, arXiv 2410.10813; Wikipedia, Temporal database; PostgreSQL documentation, WITH queries; PostgreSQL 19 release notes (as of 14 September 2026); Apache AGE on GitHub; Maximem Synap documentation, context fetch; Maximem, File Search vs Vector Search for RAG.



