Mem0 vs Letta vs Zep: Agent Memory Frameworks Compared

Mem0 vs Letta vs Zep: Agent Memory Frameworks Compared

Mem0 vs Letta vs Zep: Agent Memory Frameworks Compared

Every team that ships an LLM agent eventually hits the same wall: the model forgets. Context windows are large, but stuffing the whole history into every call is slow, expensive, and noisy, and the moment a session ends the agent starts from zero. Three open-source projects have become the default answers to that problem, and they disagree sharply about what “memory” even means. This Mem0 vs Letta vs Zep comparison separates the architectures from the marketing.

The disagreement matters now because 2026 production agents run for weeks, serve the same user across channels, and must handle facts that change. A memory layer that extracts and deduplicates facts (Mem0), one that lets the agent edit its own memory (Letta), and one that models facts as a graph with time (Zep and its Graphiti engine) fail in different ways, cost different amounts, and are measured by benchmarks that vendors openly dispute.

You will leave with a mental model of each design, a clear picture of what the published benchmark numbers do and do not prove, and a decision matrix for choosing between them.

What this covers: the memory problem, each framework’s architecture, a worked comparison of extraction, retrieval, latency and deletion, the benchmark dispute, failure modes, and a recommendation checklist.

Context and Background

LLM agent memory is the set of mechanisms that let an agent carry information beyond a single context window. The simplest approach is a long context window plus a rolling summary. The next is retrieval-augmented generation (RAG) over past messages, which embeds every turn and retrieves the nearest neighbours. Both work for demos and degrade in production, because raw chat logs are redundant, contradictory over time, and full of conversational filler.

The research lineage matters because all three projects descend from it. The MemGPT paper (Packer, Wooders, Lin, Fang, Patil, Stoica and Gonzalez, arXiv:2310.08560) proposed virtual context management, borrowing the operating-system idea of paging between fast and slow memory so a model appears to have more memory than its window holds. That project was later renamed Letta. Zep published its own architecture paper in January 2025 (arXiv:2501.13956), built around a temporally aware knowledge graph engine called Graphiti. Mem0 published its paper in 2025 (arXiv:2504.19413) describing a memory-centric pipeline that extracts, consolidates and retrieves salient facts.

If you want the broader taxonomy of memory types (episodic, semantic, procedural) before comparing products, our overview of AI agent memory systems and long-term architectures covers it, and the companion piece on LLM agent memory architecture in production covers operational concerns. This post is narrower: it takes three concrete products and asks where each one’s design choices bite.

All three core projects are released under the Apache-2.0 licence, according to their GitHub repositories as read for this article. That removes licensing as a differentiator for the open-source cores, though the hosted platforms have different feature gates and pricing, covered later.

A note on method. Everything below is drawn from the vendors’ papers, repositories and documentation, plus one third-party comparison, read in October 2026. Where a number comes from a vendor it is labelled as vendor-reported. Where I could not verify a detail in this research pass, I say so rather than fill the gap.

Three Philosophies of Agent Memory

The shortest accurate summary: Mem0 treats memory as a curated list of facts, Letta treats memory as part of the agent’s own working state, and Zep treats memory as a time-aware graph. Everything else, from latency to failure modes, follows from that choice. The question to ask of each is who decides what gets remembered, and in what shape it is stored.

Mem0 vs Letta vs Zep agent memory architectures compared side by side

Figure 1: Where each framework puts the decision about what to remember, and the shape of the resulting store.

The diagram shows three pipelines converging on the same consumer, the LLM prompt. In Mem0 an extraction step sits between the conversation and the store. In Letta the agent itself calls memory tools and a hierarchy of tiers feeds the context window. In Zep an ingestion pipeline builds a graph of entities and time-stamped facts that retrieval then queries.

Mem0: extract, compare, update

Mem0 runs a two-phase pipeline. In the extraction phase an LLM reads a message pair (typically the latest user and assistant turns plus some recent context) and proposes candidate facts, such as “prefers vegetarian food” or “works at a logistics firm”. In the update phase each candidate is compared against similar existing memories and the system chooses one of four operations: ADD, UPDATE, DELETE, or NOOP. That is how it avoids storing “likes Python” fifty times and how it reconciles “lives in Pune” with a later “moved to Bangalore”.

The design is deliberately simple at the API level. The repository quickstart creates a Memory() object, calls memory.search(query=message, filters={"user_id": user_id}, top_k=3) before the model call, and memory.add(messages, user_id=user_id) afterwards. By default it assumes an OpenAI LLM and embedding model (the quickstart names gpt-5-mini and text-embedding-3-small), though other providers are configurable.

Mem0 also offers graph memory on its platform. Its paper describes a graph variant that models relationships between entities and reports it scoring roughly 2% higher overall than the base configuration on the LoCoMo benchmark, an improvement the authors treat as modest. The platform’s pricing page, as reported by a third-party comparison, gates graph memory behind its Pro tier.

Letta: the agent edits its own memory

Letta inherits MemGPT’s central idea: the agent is responsible for managing what stays in its context. A third-party comparison of the three projects describes Letta’s tiers as core memory (in-context blocks), recall memory (searchable conversation history), and archival memory (a vector-backed store the agent queries through tool calls). Core memory blocks are labelled text regions, for example a “human” block describing the user and a “persona” block describing the agent, which Letta’s documentation uses in its quickstart.

The consequence is that memory writes are tool calls made by the agent’s own reasoning loop. The agent decides to rewrite its “human” block when it learns something new, or to push older material out to archival storage. Letta also describes “sleep-time compute”, in which a second agent edits core memory while the primary agent is idle, moving expensive consolidation off the response path.

Letta’s project has also shifted shape. The current GitHub repository is organised around Letta Code, with a terminal UI, an App Server started with letta server, and desktop and web runtimes, while the retired V1 API server lives on an archive branch. Its documentation also describes a newer file-based memory system called MemFS, “git-tracked” and using agent “dreaming”. I could not verify the internals of MemFS beyond that one-line description, so treat it as an evolving area and check the current docs before building on it.

Zep and Graphiti: facts with a validity window

Zep’s core is Graphiti, which the repository describes as a framework for building and querying temporal context graphs. Unlike a static knowledge graph, it keeps track of how facts change and retains provenance back to the source data. Facts (edges between entities) carry validity windows; when new information contradicts an old fact, the old edge is invalidated rather than deleted. The repository calls this explicit bi-temporal tracking with automatic fact invalidation, and a third-party comparison names the window fields valid_at and invalid_at.

For self-hosting, Graphiti needs a graph database. The repository lists Neo4j 5.26, FalkorDB 1.1.2, and Amazon Neptune (paired with OpenSearch Serverless for full-text search); Kuzu 0.11.2 is listed as deprecated and slated for removal. Graphiti also ships a Model Context Protocol (MCP) server in its mcp_server directory, so MCP clients can read and write the graph. The older self-contained Zep Community Edition server is, per the third-party comparison, deprecated and no longer developed, so the open-source path today is Graphiti plus your own database, while the managed Zep platform adds the product features and service levels.

The point to carry forward: Zep is the only one of the three whose core data model represents time explicitly. Mem0 and Letta can store a timestamp as text, but the graph’s retrieval logic does not reason over validity intervals.

Walk-through: One Fact, Three Lifecycles

The clearest way to see the difference is to follow a single fact through each system. Suppose a support agent hears on Monday, “I’m based in Pune and I use the Pro plan,” and on Friday, “We just moved our team to Bangalore.” A good memory layer should answer “where is this customer?” correctly on the following Monday, and ideally also know where they were before.

Write path

Write path comparison for Mem0, Letta and Zep long-term memory LLM pipelines

Figure 2: The write path for the same fact in each framework, showing where an LLM call is spent and what is persisted.

In Mem0 the Friday message triggers extraction, which yields a candidate such as “team is based in Bangalore”. The update phase retrieves the similar stored memory “based in Pune”, and the LLM picks UPDATE or DELETE plus ADD. The result is one current fact. Whether the old value survives depends on configuration and platform features; the core design is to keep the store tidy, not to keep history.

In Letta the main agent sees the message in its context, reasons that the “human” block is now stale, and calls a memory-editing tool to rewrite the line. Nothing happens unless the agent decides it should, which is a feature (the agent can judge importance) and a risk (it can simply forget to write). The sleep-time agent can later reorganise the blocks when no user is waiting.

In Zep the Friday message is added as an episode. Graphiti extracts entities and relationships with an LLM, resolves them against existing nodes, and detects that “located in Bangalore” contradicts “located in Pune”. The Pune edge gets an invalidation timestamp and the Bangalore edge is added. Both remain queryable, so “where were they in September?” has an answer.

All three spend LLM calls on the write path. This is the single most underappreciated cost in agent memory: you pay inference on every ingested message to save inference later. The three frameworks differ in whether that cost is on the response path (Letta’s tool calls, if the agent writes during the turn), asynchronous (Zep’s ingestion is designed to be processed in the background, and Letta’s sleep-time agents likewise), or batched per message pair (Mem0).

Read path

At read time the shapes diverge. Mem0 runs a semantic search over its memory store, filtered by identifiers such as user_id, and returns a short list of fact strings you paste into the prompt. Its latest published algorithm description also mentions entity linking as a retrieval feature. Letta mostly does not retrieve on your behalf: core blocks are always in context, and archival or recall lookups are tool calls the agent decides to make, which adds a model round trip when it does search. Zep’s retrieval is hybrid, combining semantic similarity, full-text search and graph traversal over the knowledge graph, and returns facts together with their validity information.

Read path and retrieval latency sources for agent memory frameworks

Figure 3: Read path for each framework, with the main source of added latency marked on each.

Latency follows from this. Mem0’s read is typically one embedding call and one vector search. Letta’s read, when the agent searches archival memory, costs an extra reasoning step. Zep’s read involves a graph query plus reranking, though Zep’s paper claims latency reductions against baseline implementations.

A minimal Mem0 example

The following is the verified quickstart shape from the Mem0 repository, wrapped into a helper. It assumes an OpenAI API key is configured in the environment, since the defaults are OpenAI models.

from mem0 import Memory
from openai import OpenAI

memory = Memory()
client = OpenAI()

def chat(message: str, user_id: str) -> str:
    # 1. Retrieve relevant memories for this user
    found = memory.search(query=message, filters={"user_id": user_id}, top_k=3)
    facts = "\n".join(f"- {m['memory']}" for m in found["results"])

    # 2. Ground the model call in retrieved facts
    reply = client.chat.completions.create(
        model="gpt-5-mini",
        messages=[
            {"role": "system", "content": f"Known about the user:\n{facts}"},
            {"role": "user", "content": message},
        ],
    ).choices[0].message.content

    # 3. Persist the new exchange; Mem0 extracts and reconciles facts
    memory.add(
        [{"role": "user", "content": message},
         {"role": "assistant", "content": reply}],
        user_id=user_id,
    )
    return reply

The search and add calls and the filters/top_k/user_id arguments match the repository quickstart. The m['memory'] field name and the found["results"] shape follow that quickstart’s use of a results list, but check the response schema in the current docs before relying on the per-item key, because the SDK has been through a major algorithm revision in 2026.

For Graphiti, the equivalent entry point is an add_episode call on a Graphiti client connected to Neo4j or FalkorDB. I have not verified the current signature in this research pass, so I am not reproducing code for it here. Consult the repository README for the exact arguments, including how to pass the reference time of the message, which is what powers the temporal reasoning.

Side-by-side properties

Property Mem0 Letta Zep / Graphiti
Core idea Extract and reconcile facts Agent edits its own tiered memory Temporal knowledge graph
Who decides what to store A pipeline LLM The agent itself An ingestion pipeline LLM
Store shape Fact strings with embeddings Core blocks, recall log, archival vectors Entities, edges, episodes
Time awareness Limited, timestamps as data Limited, depends on agent behaviour First-class validity windows
Open-source licence Apache-2.0 Apache-2.0 Apache-2.0 (Graphiti)
Self-host dependency Vector store and LLM Letta server and database Graph database (Neo4j, FalkorDB, Neptune)
Main integration style Library calls Agent runtime / platform Library plus managed API, MCP server

The licence row deserves a caveat. Apache-2.0 covers the open-source code. The numbers Mem0 publishes come from its managed platform, which its own repository says includes optimizations not available in the open-source SDK, so self-hosters should expect directionally similar but not identical results.

The Benchmark Dispute, Read Carefully

Benchmarks are where the three vendors collide, and a careful reader should treat every score below as vendor-reported unless stated otherwise. The relevant evaluations are LoCoMo (long conversational memory), LongMemEval, and the Deep Memory Retrieval (DMR) benchmark introduced with MemGPT.

What each vendor published

Zep’s paper reports 94.8% on DMR versus 93.4% for MemGPT, and on LongMemEval it reports accuracy improvements of up to 18.5% together with a 90% reduction in response latency versus baseline implementations. Mem0’s paper evaluates on LoCoMo across single-hop, temporal, multi-hop and open-domain questions, and reports a 26% relative improvement in the LLM-as-a-Judge metric over OpenAI’s memory baseline, 91% lower p95 latency than a full-context approach, and more than 90% token savings.

Mem0’s repository later reports a new-algorithm score of 92.5 on LoCoMo (up from 71.4 for the previous algorithm) and 94.4 on LongMemEval (up from 67.8), plus 64.1 and 48.6 on the BEAM benchmark at 1M and 10M token scales, all measured on its managed platform with an open-sourced evaluation framework. A third-party comparison also cites an independent LongMemEval evaluation by Vectorize on March 15, 2026 that measured Mem0 at 49.0, against Mem0’s own 94.4. I have not independently reproduced either figure, and the gap is too large to explain without differences in harness, dataset variant, judge model or algorithm version.

The Zep rebuttal

In 2025 Zep published a post disputing Mem0’s comparison. Zep’s argument, which is partisan and should be read as such, is that Mem0’s evaluation misconfigured Zep in three ways. It assigned the user role to both conversation participants, which Zep’s logic would treat as a single user whose identity changed message to message. It appended timestamps to message text instead of using Zep’s created_at field, which interferes with temporal reasoning. And it ran searches sequentially rather than in parallel, inflating latency.

With those changes Zep reports a LoCoMo J score of 75.14% (plus or minus 0.17), against the 65.99% Mem0’s paper attributed to Zep. In the same post, it cites Mem0 Graph at 68.44% and Mem0 Base at 66.88%, so Zep claims roughly a 10% relative lead. Zep’s reported p95 search latency was 0.632 seconds, versus 0.657 for Mem0 Graph and 0.200 for Mem0 Base, a figure Zep says is not an apples-to-apples comparison.

Zep’s post also criticises LoCoMo itself, and these points are worth taking seriously regardless of which vendor you favour. It notes that a full-context baseline of about 73% beat Mem0’s best score of about 68% in that setup, suggesting the benchmark conversations are short enough that stuffing everything into context wins. It says LoCoMo lacks knowledge-update questions, which is exactly the capability a temporal graph is built for. And it lists category 5 as unusable for missing ground truth, along with multimodal errors, speaker misattribution and ambiguous questions.

Why published agent memory benchmark scores are not comparable across vendors

Figure 4: Sources of divergence between vendor-reported scores, from dataset variant to judge model to which algorithm version was measured.

How to read the numbers

My view, labelled as opinion: no published score here can pick your winner. The inputs to a score include the benchmark variant, the answering model, the judge model, how each competitor was configured, and which software version ran. A third-party comparison reaches the same conclusion, noting that the scores come from different harnesses and variants and are not directly comparable. Letta has a figure of 83.2 on LoCoMo in a third-party table, but I found no independent measurement behind it, and because Letta’s design puts memory decisions inside the agent loop, a retrieval-style benchmark measures only part of what it does (this is my inference, not a Letta statement).

The practical consequence is that you should build a small evaluation from your own conversations. Take 50 to 200 real multi-session transcripts, write questions whose answers changed over time, and score each framework with the same answering and judging model. That test takes a day and tells you more than any leaderboard.

Cost, Latency, Privacy and Deletion

Choosing a memory layer is also an operations decision. Four concerns decide whether a pilot survives contact with production: where inference is spent, what the hosted tiers cost, how deletion works, and how multi-tenant isolation is enforced.

Where the money goes

A memory layer adds LLM calls on ingestion. Consider a worked example with deliberately illustrative numbers: an agent handles 10,000 conversations a month, each with 20 message pairs. A per-pair extraction design makes about 200,000 extraction calls a month before any reconciliation calls. If each call averages 1,500 input tokens and 150 output tokens, that is 300 million input tokens and 30 million output tokens monthly just to write memory. The arithmetic is mine and the token counts are assumptions; plug in your own prices. The lesson is that write-path volume scales with message count, not with the number of useful facts, so batching, skipping trivial turns and using a small extraction model all matter.

Letta shifts some of that cost into the agent’s own reasoning loop and into optional sleep-time agents. Zep, on its managed platform, meters usage in credits: according to the third-party comparison, one credit per episode of up to 350 bytes plus one credit per additional 350 bytes, with a free tier of 10,000 credits a month. Hosted pricing at the time of that article put Mem0 at a free tier of 10,000 adds and 1,000 retrievals per month with a $19 Starter plan and a $249 Pro plan that adds graph memory, Letta at a free plan, a $20 Pro plan for personal use and usage-based developer pricing, and Zep’s Flex plan at $125 per month. Prices change often, so verify against each pricing page before budgeting.

Read-side cost is smaller but not zero. Mem0’s paper claims over 90% token savings versus full-context prompting, which is plausible by construction: retrieving three short facts instead of a 100,000-token history is a large reduction. That claim compares against the worst baseline, though. A well-built summary-plus-RAG system is the fairer rival.

Latency budgets

Interactive agents have a budget of roughly one to two seconds for a first token, and every memory lookup spends part of it. A single-step vector search is the cheapest read path. Graph retrieval with reranking costs more, which is why the Zep post quotes p95 search latencies in the hundreds of milliseconds (0.632 seconds corrected Zep versus 0.200 for Mem0 Base in the same post). Agent-driven retrieval in Letta adds a model round trip whenever the agent chooses to search, which can dominate everything else.

The pattern that works in all three: keep the always-needed identity facts in context (Letta’s core blocks, or a small pinned set from Mem0 or Zep), and retrieve the rest on demand. Write asynchronously after the response is sent, unless your agent must see its own write in the same turn.

Privacy, deletion and the right to be forgotten

Memory layers are personal data stores by design, so regulations such as the GDPR’s erasure right apply to what you extract, not just to raw chat logs. This is where design differences become compliance differences.

With Mem0, memories are discrete records keyed by identifiers like user_id, so erasing a user is conceptually a filtered delete. The risk is derived data: embeddings and any copies in your own logs. With Letta, memory is scattered across core blocks, recall history and archival passages, and agent-authored text can paraphrase personal data into places you did not plan for. You need a deletion routine that covers all three tiers and any sleep-time rewrites.

With Zep and Graphiti, the central design choice, invalidating rather than deleting, is the opposite of what erasure requires. Invalidated facts remain queryable by design. For a deletion request you must remove the nodes, edges and source episodes themselves, and the provenance links that make the graph auditable also make it easy to find everything about a person. I have not verified the exact deletion APIs of each product in this research pass, so confirm them in the current documentation and test them end to end before you rely on them.

For multi-tenant systems, namespace by user or organisation from day one. Mem0 uses identifier filters, and Graphiti has a group concept (its MCP server lists group management). Test that a query under one identifier can never surface another’s memories, because retrieval bugs here are data breaches. Memory is also an attack surface: a poisoned fact written once will be retrieved forever, which connects to the injection risks covered in our post on agentic AI security and prompt injection.

Trade-offs, Gotchas, and What Goes Wrong

Each design has a characteristic failure, and knowing it in advance is worth more than any feature list.

Mem0: lossy extraction. A fact the extraction prompt did not consider salient is gone, because the raw conversation is not the source of truth in the memory store. Extraction can also hallucinate, converting a hypothetical (“if I moved to Berlin…”) into a fact. Reconciliation errors compound: a wrong DELETE destroys a correct memory, and no later retrieval can resurrect it unless you keep the raw transcripts. Keep them anyway, in cheap storage, so you can re-extract when the pipeline improves.

Letta: the agent forgets to remember. Self-editing memory relies on the model’s judgement in every turn. Smaller models under-write, over-write, or corrupt a core block with a bad edit, and since core blocks are always in context, an error there affects every later response. Memory behaviour also varies by model, so a model upgrade can silently change what your agent keeps. The platform’s shift toward Letta Code and newer memory systems also means documentation and APIs have moved; pin versions.

Zep and Graphiti: graph complexity. Entity resolution is hard. If “Acme”, “Acme Corp” and “ACME Ltd.” become three nodes, the graph fragments and retrieval misses facts. Ingestion is the heaviest of the three, since each episode drives entity extraction, resolution and contradiction detection, and you now operate a graph database. The self-contained community server is deprecated, so self-hosting means running Neo4j, FalkorDB or Neptune yourself.

Shared problems. All three inherit extraction error from the underlying LLM. All three are exposed to stale-context bugs: a memory retrieved at the start of a long session may be outdated by the end. All three need an eviction or decay policy, because a memory that grows without bound degrades retrieval precision. And none removes the need for an evaluation harness. An anti-pattern worth naming: using memory as a substitute for a database. If the customer’s plan tier is authoritative in your billing system, look it up there; use memory for what users said, preferred and implied.

Practical Recommendations

Start from the failure you can least afford. If it is staleness (preferences, locations, roles and statuses that change), the temporal graph is the only design that models change explicitly, and Zep or Graphiti deserves the first pilot. If it is simplicity and speed to production, Mem0’s add-and-search API gets a working memory in an afternoon. If you are building a long-lived, autonomous agent that should curate its own knowledge and you control the runtime, Letta fits best, because memory is part of how the agent thinks rather than a sidecar.

Situation Best fit Why
Chatbot or assistant with user preferences, low ops budget Mem0 Simple API, fact store, no graph database
Facts that change and history matters (CRM, support, compliance) Zep / Graphiti Validity windows keep old and new facts
Autonomous long-running agent with its own persona and tools Letta Tiered, self-edited memory inside the agent loop
Strict erasure requirements, minimal data retained Mem0 or custom Discrete records are simplest to delete; verify APIs
Tool-using IDE or MCP clients wanting shared graph memory Graphiti MCP server Ships an MCP server with search and episode tools

Whichever you choose, follow this checklist:

  • Keep raw transcripts separately so extraction can be re-run.
  • Build a 50 to 200 conversation evaluation from your own data, including knowledge-update questions.
  • Write memory asynchronously and pin only a small set of always-needed facts in context.
  • Namespace every record by tenant and user, and test cross-tenant leakage.
  • Implement and test deletion end to end, including derived embeddings and graph edges.
  • Pin framework versions; all three changed materially in 2026.
  • Treat retrieved memory as untrusted input to the model.

Agent memory also matters beyond chat. In industrial settings, an agent that remembers equipment history, past interventions and operator corrections is closer to what we describe in our piece on agentic digital twins for AI-driven industrial analysis, where “this pump was serviced in March” must be a time-stamped fact, not an overwritten string. That is the strongest argument for temporal modelling outside consumer chat.

Frequently Asked Questions

What is the main difference between Mem0, Letta and Zep?

Mem0 extracts salient facts from conversations and reconciles them with an ADD, UPDATE, DELETE or NOOP decision. Letta, which grew out of MemGPT, lets the agent edit its own tiered memory (core, recall and archival) through tool calls. Zep builds a temporal knowledge graph with Graphiti, where facts carry validity windows and contradicted facts are invalidated rather than deleted. The key question is who decides what to remember and in what shape it is stored.

Is Letta the same as MemGPT?

Letta is the successor to MemGPT. The GitHub repository is titled “Letta (f.k.a. MemGPT)”, and the original research paper by Packer and colleagues (arXiv:2310.08560) introduced virtual context management, which moves information between a limited context window and external storage. Letta has since grown into a platform for stateful agents, now centred on Letta Code, a terminal UI and App Server, with the V1 API server retired to an archive branch.

Which agent memory framework is fastest?

No published number settles this. Mem0 Base showed 0.200 seconds p95 search latency in Zep’s own comparison post, against 0.632 for Zep, but Zep argues the comparison is not like for like. Letta’s latency depends on whether the agent chooses to search archival memory, which adds a model round trip. Measure p95 on your own workload, with your own embedding model, hosting region and concurrency, before deciding.

Can I self-host all three?

Yes, with differences. Mem0, Letta and Graphiti are all Apache-2.0 licensed. Mem0 needs an LLM, an embedding model and a vector store. Letta can run its own server, started with letta server in the current Letta Code repository. Graphiti needs a graph database such as Neo4j, FalkorDB or Amazon Neptune. Zep’s older self-contained community server is reported deprecated, so self-hosting Zep today effectively means Graphiti.

Are the LoCoMo and LongMemEval scores trustworthy?

Treat them as vendor-reported and non-comparable. Mem0 reports 92.5 on LoCoMo and 94.4 on LongMemEval for its platform’s new algorithm, while an independent Vectorize evaluation cited in a third-party comparison measured Mem0 at 49.0 on LongMemEval. Zep contests Mem0’s setup for Zep and criticises LoCoMo for short conversations and missing knowledge-update questions. Build your own evaluation from real transcripts.

Do I need a knowledge graph for agent memory?

Not usually at the start. A fact store with good extraction handles preferences and stable attributes well. A graph earns its complexity when facts change over time, when relationships between entities drive answers, or when you need an audit trail of what was believed when. Mem0’s own paper reports only a roughly 2% overall gain from its graph variant on LoCoMo, which suggests graphs matter most for temporal and relational queries.

Further Reading

By Riju — about

Comments

No comments yet. Why don’t you start the discussion?

Leave a Reply

Your email address will not be published. Required fields are marked *