Published 27 September 2026 · Query-language standards, PostgreSQL 19 status and vendor documentation checked on 26 September 2026; research on graphs in agent memory checked on 27 September 2026.
A graph database stores data as nodes (the things: people, accounts, devices, services) and relationships (the typed, directed connections between them), with properties on both, and it answers questions by following those relationships instead of joining tables. Use one when your most important questions are about paths and patterns across connected data, especially when the number of hops is not known in advance, such as fraud rings, dependency maps, recommendations and knowledge graphs. Do not use one for tabular records, key lookups, append-only event logs or whole-table reporting, where a relational database is simpler and faster.
Most teams asking this question already run PostgreSQL, so the real decision is usually narrower than "graph or not": it is whether a traversal-heavy workload has outgrown recursive SQL enough to justify a second database, a new query language, a sync pipeline and another system on call.
How a graph database works
A graph database has three building blocks. Nodes carry a label such as Person or Account and any number of key-value properties. Relationships connect exactly two nodes; in Neo4j a relationship "must always have a start node, an end node, and exactly one type" and "must have a direction", and it can carry properties of its own, such as since: 2024-03-01 on a WORKS_AT edge. Properties are the attributes on either.
What makes it a graph database rather than a relational database with a diagram drawn over it is how relationships are stored. A relational database keeps a foreign key in one row and finds the matching row at query time with a join, so a question that crosses four relationships runs four joins, each matching keys across a table. A graph database persists each relationship as a record the engine can follow directly from the node it starts at; AWS puts it as relationships that "are not calculated at query times but are persisted in the database". The practical consequence, in Microsoft's wording, is that "a traversal's cost depends on the number of edges it touches (the local neighborhood), not the total size of the dataset." A friends-of-friends query on a person with forty friends touches a few thousand edges whether the database holds a million people or a billion.
That property is the whole argument for the category. Everything else (the query languages and the built-in algorithms) follows from making a hop cheap.
One question, written two ways
Take a small social graph: people, and a KNOWS relationship between them. The question is "who can Priya reach within four introductions?"
Cypher, the query language Neo4j created and that Amazon Neptune also supports as openCypher, makes the pattern the query:
MATCH (me:Person {name: 'Priya'})-[:KNOWS*1..4]->(other:Person)
WHERE other <> me
RETURN DISTINCT other.name
*1..4 means "follow between one and four KNOWS relationships". Changing the depth is a one-character edit.
PostgreSQL needs a recursive common table expression, for the same question, where "a WITH query can refer to its own output":
WITH RECURSIVE reach(person_id, depth) AS (
SELECT friend_id, 1 FROM knows WHERE person_id = 42 -- Priya
UNION
SELECT k.friend_id, r.depth + 1
FROM knows k
JOIN reach r ON k.person_id = r.person_id
WHERE r.depth < 4
)
SELECT DISTINCT person_id FROM reach WHERE person_id <> 42;
Both return the right answer. The difference shows up in two places. The SQL version re-joins the knows table at every level and stops cycles only because of the depth limit, so its cost grows with how many rows each level produces. And the SQL version is a lot harder to extend: adding "only through people who work at the same company, and return the path" is a clause in Cypher and a rewrite in SQL. If your queries look like the second one more and more often, that is the strongest signal in this article.
Property graphs, RDF and the query languages
Graph databases come in two data models. A property graph is what the example above uses: labelled nodes, typed relationships, properties on both. It is built for traversal and analytics, and it is what Neo4j, Neptune's property-graph mode, TigerGraph and most application work use.
An RDF graph follows the W3C's Resource Description Framework. Oracle's explainer describes a statement as three elements, "the subject, a predicate, and the object", with every vertex and edge identified by a URI; that shared-identifier design is why RDF suits publishing and integrating data across organisations, and why RDF stores are often what people mean by "knowledge graph" in enterprise data work. RDF is queried with SPARQL.
Query languages have been fragmented, which InfluxData's guide lists as a real cost of adoption. That is changing. ISO/IEC 39075:2024, the GQL standard, was published in April 2024, and ISO's own committee called it "the first time in more than 35 years that ISO will be releasing a new database query language". It standardises the pattern-matching style Cypher made familiar. Alongside it, SQL:2023 added SQL/PGQ (ISO/IEC 9075-16:2023), which lets a relational database define a property graph over existing tables and query it with MATCH inside SQL. Gremlin, from Apache TinkerPop, remains the step-by-step traversal language on several engines.
When a graph database is the right tool
A graph database earns its place when three things are true together: the questions are about how things connect rather than about the things themselves, the number of hops varies or is not known in advance, and the relationships carry meaning of their own (a type, a date, a weight). Microsoft's criteria say the same: choose one when "your primary questions involve paths, neighborhoods, and patterns in connected data" and "the number of hops is variable or not known in advance."
Workloads that fit are the ones where the answer is a path or a pattern. Fraud detection looks for rings, such as several accounts sharing a card, an email or an IP address across different physical locations, which AWS describes as detecting fraud "through relationship patterns". Recommendations walk from a person to what similar people bought. IT and network operations trace which services break when one host fails. Supply-chain teams trace a component back through tiers of suppliers. Identity and access management answers "who can reach this resource, through which groups and roles". Knowledge graphs connect the entities in a domain so that search and AI systems can follow them. Graph engines also ship algorithms for this shape of data, such as shortest path, centrality (who is most connected), link-based ranking and community detection (which clusters are tighter to each other than to everyone else).
A quick test before any of that: write down the five questions your application must answer fastest. If three of them contain "through", "connected to", "within N steps" or "path", you have a graph workload. If they contain "total", "average", "by month" or "where id =", you do not.
When a graph database is the wrong tool
Most data is not a graph problem, and graph vendors' own pages say so in places. AWS's guidance is that records with fixed columns and few relationships, such as an inventory of items and counts, belong in a relational database, and that "it is also important not to use graph databases simply as key-value stores", because a lookup by known key uses none of what the engine is built for.
Four workloads are a poor fit. Whole-dataset aggregation and reporting (revenue by region by month) scans everything and traverses nothing; graphs, as InfluxData notes, "struggle to process queries that span the entire database". Append-heavy time series and event logs want columnar or time-series storage. Search by meaning or by keyword wants a vector or full-text index. And a schema with two or three shallow relationships gains nothing from a graph engine that a join does not already give you, while costing you a second system.
Graph database vs relational database at a glance
| Relational database | Graph database | |
|---|---|---|
| Unit of storage | Rows in tables with a fixed schema | Nodes and relationships, each with properties; schema optional |
| How relationships work | Foreign keys, matched at query time with joins | Persisted relationships, followed directly |
| Cost of a deep query | Grows with join depth and the rows produced at each level | Grows with the edges actually touched |
| Query languages | SQL | Cypher and openCypher, Gremlin, SPARQL, GQL (ISO 2024) |
| Strongest at | Transactions on records, aggregates, reporting, constraints | Paths, patterns, variable-depth traversal, network analysis |
| Weakest at | Variable-depth and many-hop questions | Whole-dataset scans, tabular reporting, key-value lookups |
| Typical examples | PostgreSQL, MySQL, SQL Server | Neo4j, Amazon Neptune, TigerGraph, Spanner Graph |
Many production systems run both, a point Google Cloud makes directly: the relational database stays the system of record for customers, orders and transactions, and the graph holds the connections that need traversing, fed from the relational source.
Can PostgreSQL do this instead?
Often, yes, and it is worth checking before adding anything. There are three routes, and they differ in how far they go.
Recursive CTEs, shown above, are built into PostgreSQL and handle bounded traversals such as org charts, bills of materials and category trees well, provided you index the foreign keys and cap the depth.
Apache AGE is a PostgreSQL extension that adds openCypher queries over graph data stored inside Postgres. Its README lists support for PostgreSQL 11 through 18, and its latest release is v1.8.0. It keeps your data, backups and access control in one database, at the cost of running an extension that your managed Postgres provider may or may not offer.
SQL/PGQ is the standard way to put graph queries inside SQL, and it is not available in a released PostgreSQL today. A SQL/PGQ implementation was committed to the PostgreSQL 19 development branch in March 2026 and later rolled back, and the PostgreSQL 19 release notes (which still show the release date as "2026-??-??, AS OF 2026-09-14") do not list it. Treat any article that says PostgreSQL 19 ships native graph queries as unverified until those notes change.
A reasonable rule: stay in PostgreSQL while your traversals are bounded (known depth, a handful of hops) and your graph fits comfortably on one machine. Move to a dedicated graph database when variable-depth path queries become a core product feature, when you need graph algorithms at scale, or when the recursive SQL has become the part of the codebase nobody wants to touch.
What running a graph database costs a team
Transactions are not the problem some older guides suggest. InfluxData's page lists "no transactions" as a disadvantage of graph databases; Neo4j's operations manual states the opposite: "Neo4j DBMS supports transactions with full ACID properties, and it uses a write-ahead transaction log to ensure durability." Check the specific engine rather than the category.
Scaling is harder than with most databases, because a graph does not split cleanly. Shard a connected graph across machines and some hops become network calls, which is why one widely read analysis recommends sharding by something you already know about the data, such as customer ID or region, so that most traversals stay on one machine. High-degree nodes, or supernodes (a celebrity account with millions of followers, a default "Unknown" category that everything links to), turn a cheap hop into a scan of millions of edges; the usual fixes are splitting the node or indexing its relationships by type or property.
Organisational costs are larger. Your team learns a new query language and a new way of modelling. Unless the graph is the system of record, you run a pipeline that copies data from the relational source and keeps it in sync, and Microsoft notes that a standalone graph store "often introduces ETL (extract, transform, load) and governance overhead". It is one more system to secure, patch, back up and put on call. The main options are dedicated engines such as Neo4j and TigerGraph, managed services such as Amazon Neptune, relational databases that have added graph querying such as Google's Spanner Graph and Oracle, and the Postgres extension route above.
On modelling, start from the questions rather than the entities. Make something a node when you will traverse through it or attach relationships to it (a company that employees, contracts and incidents all point at); keep it as a property when you only ever read it (a person's date of birth).
Graph databases, knowledge graphs and vector databases in AI
Three terms get used interchangeably in AI writing, and they are different things. A graph database is the storage and query engine. A knowledge graph is a body of facts about a domain, modelled as entities and typed relationships, often with an agreed vocabulary; it can live in a graph database, an RDF store, or a set of Postgres tables. A vector database stores embeddings (numeric representations of meaning) and returns the items most similar to a query; pgvector, for example, does exact search by default and adds HNSW or IVFFlat indexes for approximate search, "which trades some recall for speed."
A vector database answers "what is similar to this?" A graph answers "what is connected to this, and how?" Neither replaces the other, and the boundary is not always where people expect: in our retrieval study (50,000 documents across five datasets, 5,000 queries, scored on MRR@10 by exact document match with no model judge), vector search led the five-dataset average on MRR@10 (0.6320 against 0.5325 for keyword search), but with the code dataset removed, plain keyword search led the average of the other four (0.5931 against 0.5614), and indexing took 2.11 seconds for keywords against 161.6 seconds for embeddings.
GraphRAG is the best-known use of graphs in retrieval. Microsoft Research's GraphRAG paper extracts an entity graph from a document collection and summarises its communities, because, in the authors' words, "RAG fails on global questions directed at an entire text corpus, such as 'What are the main themes in the dataset?'" The two approaches also meet inside one engine: Neo4j's vector index runs HNSW-based nearest-neighbour search on node and relationship properties, so a query can find similar entry points by meaning and then traverse from them, a pattern usually called hybrid search.
Is a graph database, or GraphRAG, enough to build agent memory?
No. A graph database is a store, and GraphRAG is a retrieval technique built on one. Both are important components of durable agent memory, and neither is a memory system on its own. A graph holds relationships once they exist; durable memory also needs a vector index for meaning and a keyword or file index for exact terms, organised together, and deciding what goes into which of them, what to keep, what to merge, what to retire and what to hand back on a given turn is the memory system's job, not the database's.
Agent memory is full of graph-shaped facts, which is why the graph is so tempting. The user works at a company, the company is on a plan, the plan renews on a date, "my manager" and "Sarah Chen" are the same person, and the user moved from Chicago to Berlin in June. Four reasons explain why storing those facts in a graph still leaves most of the work undone.
The store makes none of the decisions. In our paper on agentic context management (Dadhich, arXiv 2607.21503, July 2026) we list what a production agent platform has to decide turn by turn: which of the things just said are worth retaining at all, in what structure they should be retained, which small fraction of everything retained belongs in this turn's context, what the next turn is likely to need, and what should happen when the relevant context exceeds the budget the model can use meaningfully. As the paper puts it, "These are five different decisions made at five different moments, and a store makes none of them." A graph database answers where a relationship lives. It does not decide that the relationship was worth writing down, that three names refer to one person, that a new fact contradicts an old one, or that one customer's facts must never appear in another customer's traversal.
No single index finds everything. Retrieval over memory needs more than one signal, because the questions an agent asks are shaped differently. In our retrieval study, keyword search won outright on the two datasets whose questions turn on specific entities (HotpotQA and SciQ), vector search won by a wide margin where the words in the question and the answer differ most (natural-language-to-code, 0.91 against 0.29 MRR@10), and with the code dataset removed keyword search led the average. The paper draws the conclusion for memory directly: vector-only implementations "need to be used in conjunction with graphs so as to retain relational information in querying as well as to manage provenance", and "Pure graph traversal is expensive and brittle." The hybrid is necessary, and the paper is explicit that it is "necessary, not sufficient." Outside research reaches the same place from the retrieval side. GraphRAG-Bench (Xiang et al., 2025, revised February 2026) starts from the observation that "GraphRAG frequently underperforms vanilla RAG on many real-world tasks", and a systematic evaluation of RAG against GraphRAG (Han et al., 2025, revised March 2026) found distinct strengths for each depending on the task, with strategies that combine the two "leading to consistent performance improvements."
GraphRAG was built for a different job. GraphRAG answers global questions about a document collection that changes slowly, and it pays for that with heavy indexing: every document passes through a language model that extracts entities and summarises communities. Microsoft Research's own LazyGraphRAG announcement (November 2024) says its lighter variant's "data indexing costs are identical to vector RAG and 0.1% of the costs of full GraphRAG", which puts full GraphRAG indexing at roughly a thousand times vector indexing. Agent memory is the opposite workload: it is written on every turn, it changes constantly, and it has to be current by the next conversation. The companion piece works through what graph extraction costs per conversation turn at today's model prices.
Memory is judged on abilities no store provides. LongMemEval (Wu et al.), the most widely used benchmark for long-term memory in chat assistants, tests information extraction, multi-session reasoning, temporal reasoning, knowledge updates and abstention, and it breaks a memory system into indexing, retrieval and reading stages, each with its own design choices. Knowing that an answer changed, when it changed, and when to say "I do not know" are properties of how memory is written and read, not of the database underneath. Our paper states the dependency as a chain: answer quality is at most the minimum of extraction quality, retrieval quality and reasoning sufficiency, so a strong graph cannot compensate for weak extraction upstream of it.
| The graph database gives you | The memory system still has to decide |
|---|---|
| Nodes, relationships and properties, stored durably | Which parts of a conversation are worth keeping at all |
| Fast traversal once an entry point is known | Whether a fact belongs in the graph, the vector index, the keyword or file store, or several of them |
| Transactions on writes | That "my manager", "Sarah" and "Sarah Chen" are one entity, and what to do when it is not sure |
| A query language for paths and patterns | Which of two conflicting facts is current, and keeping the old one as history |
| Optionally, a vector index on node properties | Which user, customer and organisation each fact belongs to, enforced on every read |
| What fits in this turn's token budget, and what the next turn is likely to need | |
| What to forget deliberately, rather than lose by accident |
Maximem Synap, the memory layer we build, is designed around that split: the graph is one retrieval technique inside a memory service rather than the memory itself. Synap reads each conversation as it arrives and stores discrete, self-contained statements with a type, a confidence score, a scope and a timestamp, rather than transcript chunks. It resolves people and companies into single entities across conversations and channels, and sends genuinely ambiguous matches to a review queue rather than guessing. When a fact changes, the change is recorded, not overwritten: the old statement is marked historical and linked to the one that replaced it, with the reason. At read time, Synap's Fast mode serves agents that are talking to someone right now, and its Accurate mode breaks a compound question into parts and expands through the relationships between entities, which is the graph traversal, only when the question needs it.
Whether you build it or use a service, the order of work is the same: get extraction, entity resolution, scoping and change handling right first, then choose the stores and how they are combined. The companion piece on knowledge graph vs vector search for agent memory goes deeper on that choice, including temporal knowledge graphs, and how Synap works covers the pipeline end to end.
Frequently asked questions
What is a graph database in simple terms?
A graph database is a database that stores things as nodes and the connections between them as relationships, each with its own properties, and answers questions by following those connections. It is built for questions such as "who is connected to whom, through what, within how many steps", which relational databases answer with repeated joins.
When should I use a graph database instead of a relational database?
Use a graph database when your core queries are about paths and patterns across connected data and the number of hops is variable or unknown, as in fraud rings, dependency tracing, recommendations and knowledge graphs. Keep a relational database for records, transactions, reporting and aggregates. Many systems use both, with the relational database as the system of record and the graph fed from it.
Can PostgreSQL be used as a graph database?
PostgreSQL can answer bounded graph questions with recursive CTEs, and the Apache AGE extension adds openCypher queries over graph data stored in Postgres (PostgreSQL 11 to 18 per its README). The SQL/PGQ standard for graph queries inside SQL was committed to PostgreSQL 19 development in March 2026 and later rolled back, and it is not in the PostgreSQL 19 release notes as of 14 September 2026.
Do graph databases support ACID transactions?
Many do. Neo4j's documentation states that it "supports transactions with full ACID properties" using a write-ahead log. Some older guides list "no transactions" as a general weakness of graph databases; check the specific engine's documentation rather than relying on a category-wide claim.
What is the difference between a graph database and a knowledge graph?
A graph database is software for storing and querying connected data. A knowledge graph is the content: facts about a domain modelled as entities and typed relationships, often with a shared vocabulary. A knowledge graph can be stored in a graph database, an RDF triple store, or ordinary relational tables.
What is the difference between a graph database and a vector database?
A vector database stores embeddings and returns the items most similar in meaning to a query. A graph database stores explicit relationships and returns what is connected to a starting point and how. AI systems often use both: similarity search to find where to start, and graph traversal to follow the relationships from there.
Is GraphRAG enough to give an AI agent memory?
No. GraphRAG is a retrieval technique for answering broad questions over a document collection, and a graph database is a store. Durable agent memory also needs vector and keyword retrieval alongside the graph, plus the decisions no store makes: what to keep from each conversation, which entity a mention refers to, which of two conflicting facts is current, whose data a fact belongs to, and what fits in the current turn. Research comparing GraphRAG with plain RAG finds each wins on different tasks, and combining them performs better than either alone.
What is GQL?
GQL is the ISO standard query language for property graphs, published in April 2024 as ISO/IEC 39075:2024. It standardises the pattern-matching style that Cypher popularised, much as SQL standardised relational queries, and it sits alongside SQL/PGQ, the part of SQL:2023 that brings the same graph patterns into SQL itself.
Sources, retrieved 26 September 2026: Microsoft Learn, graph database overview; Neo4j, what is a graph database; Neo4j operations manual, database internals; Neo4j Cypher manual, vector indexes; AWS, what is a graph database; Amazon Neptune user guide; Google Cloud, what is a graph database; Oracle, what is a graph database (9 January 2026); InfluxData, a guide to graph databases; ISO/IEC 39075:2024 GQL; ISO/IEC JTC 1, article on GQL (April 2024); PostgreSQL documentation, WITH queries; PostgreSQL 19 release notes (as of 14 September 2026); depesz, SQL/PGQ commit note; Apache AGE on GitHub; pgvector on GitHub; Edge et al., From Local to Global: A Graph RAG Approach, arXiv 2404.16130; Dadhich, Agentic Context Management, arXiv 2607.21503 (retrieved 27 September 2026); Xiang et al., When to use Graphs in RAG (GraphRAG-Bench), arXiv 2506.05690 (retrieved 27 September 2026); Han et al., RAG vs. GraphRAG: A Systematic Evaluation and Key Insights, arXiv 2502.11371 (retrieved 27 September 2026); Microsoft Research, LazyGraphRAG (25 November 2024) (retrieved 27 September 2026); Wu et al., LongMemEval, arXiv 2410.10813 (retrieved 27 September 2026); DZone, Do Graph Databases Scale?; Maximem, File Search vs Vector Search for RAG.



