/

/

/

How to Read Memory Benchmarks

Share

How to Read Memory Benchmarks

How to Read Memory Benchmarks

How to Read Memory Benchmarks

aashi dutt
Published on Oct 8, 2026

·

20 min read

How to Read Memory Benchmarks
How to Read Memory Benchmarks

On this page

The internet is filled with benchmark numbers, and everyone seeks to come out on top. However, the numbers are not just a factor in determining whose system performed the best. They also depend on the environment around the system and the components that make it up. In memory benchmarks, the same system can lead to different results simply because you used a different model as a judge. Within a single harness, the same system's score also moves with retrieval depth and with the model that extracts its memories.

Our LoCoMo scores have been reported at both 74% and 92.5%, and neither is wrong; they are just a response to a different setup. Each one answered a slightly different question like, which question categories were counted, which model wrote the answers, which model graded them, and how the system under test was configured. Change any of those and a memory score can move by twenty points.

So, let's drop the ball and understand how we can read and understand these scores.

This guide is about reading the setup behind a memory benchmark score. It covers what each benchmark measures, how an evaluation is put together, where scores go wrong, how much is noise, and finally how to test a memory system on your own workload. Basically, an ultimate guide to everything about memory benchmarks.

Looking for a quick version? We have it here: https://mem0.ai/blog/state-of-ai-agent-memory-2026

The benchmark tests

Every memory benchmark refers to at least 4 different kinds of tests, and the most misleading conclusions appear by mixing them:

  • Long context retrieval: It tests whether a model can find and use facts inside one large prompt. With no memory system involved, these benchmarks measure the context window, which is not an actual memory (or maybe a preconditioned one if you say so).

  • Cross-session recall: This actually tests a memory system like a human recalling another task or episode in between a similar task. It tests whether a system can answer questions about past conversations. This indeed is the place from which most benchmark numbers come from.

  • Conflict and update handling: This benchmark tests whether a system uses the latest version of a fact after it changes or not?

  • Agentic task completion: It’s basically what happens to humans after a test — did the agent do the right thing later using what it had learned earlier? It then grades the actions it takes rather than writing them as text.

Multi-hop QA sets such as HotPotQA and MuSiQue also appear in memory comparisons, but they test retrieval across documents, not memory over time.

Benchmark tests

Here is a quick table for an overview:

Benchmark

Year

What it tests

Scale

Grading

RULER

2024

Long-context retrieval across 13 task types

Configurable context length

Rule-based

BABILong

2024

Reasoning over facts in very long contexts

Up to millions of tokens

Exact match

NoLiMa

2025

Retrieval with minimal word overlap between question and fact

Up to 32K tokens and beyond

Accuracy

LoCoMo

2024

Single-hop, multi-hop, temporal, open-domain and adversarial QA

About 9K tokens per conversation, up to 35 sessions

F1 and BLEU-1 originally; LLM judge in most recent work

LongMemEval

2024

500 questions across extraction, multi-session reasoning, temporal reasoning, knowledge updates and abstention

S: about 115K tokens; M: about 1.5M tokens across 500 sessions

LLM judge

BEAM

2025

2,000 questions across ten memory abilities

Buckets up to 10M tokens

LLM judge

MemoryAgentBench

2025

Accurate retrieval, test-time learning, long-range understanding, selective forgetting

Information fed incrementally

Substring exact match and accuracy

MemoryArena

2026

Interdependent multi-session agent tasks: web navigation, planning, search, formal reasoning

Multi-session task loops

Task success

STATE-Bench

2026

450 tasks in customer support, travel and shopping

Five runs per task

Completion, pass^5, LLM-judged UX, cost

DolphinBench

2026

600 tool-using tasks across simulated apps

About 500K tokens of history per persona

Action-based checks, with cost and latency

How does a memory evaluation work?

Every evaluation that involves a memory system has the same moving parts: a pipeline, an answer model, a judge, and the metrics used to score the result. Note that each part can affect the score, so we’ll explore these one by one.

How des memory evaluation work
  1. Pipeline: The pipeline stores memories and retrieves them later. It has 2 phases:

    1. Ingestion: In this phase, the system reads the history and writes it into its store, say a vector index, a graph, files, or some mix.

    2. Query: In this phase, the system retrieves from that store, and an answer model responds.

    Finally, a single score blends the two phases together. A fact that was never written cannot be extracted or retrieved, no matter how good your search is. So, when the result surprises you, look through each phase to find out where that empty result came from.

  2. Extraction model: Ingestion usually runs an LLM that decides which facts to store. Mem0's open-source LongMemEval results change only that model, with the same embedder, the same vector store, and GPT-5 as answerer and judge:

Extraction model

Overall

SS-User

SS-Asst

SS-Pref

Knowledge update

Temporal

Multi-session

GPT-5

91.0%

95.7%

92.9%

93.3%

91.0%

94.7%

83.5%

GPT-OSS-120B

89.8%

95.7%

96.4%

93.3%

89.5%

80.5%

79.7%

Llama 4 Maverick

88.6%

97.1%

75.0%

93.3%

93.6%

90.2%

84.2%

Gemma 4 31B

88.6%

95.7%

83.9%

93.3%

94.9%

91.7%

78.9%

The overall scores sit within 2.4 points of each other. But the categories don’t.

Overall memory eval

So, Single-session assistant recall ranges from 75.0% to 96.4%, and temporal reasoning from 80.5% to 94.7%. The ranking also flips by category: GPT-5 leads overall and on temporal reasoning, but scores below both Llama 4 Maverick and Gemma 4 31B on knowledge updates, where Gemma leads. These ranges suggest that an overall score can hide the weakness that matters for your workload.

Note: Two cautions apply to tables like this.

The categories are small: A single-session preference has 30 questions, which is why every model scores 93.3% (28 of 30), and a 56-question category carries a margin of roughly ±7 to ±11 points by our calculation, depending on the score.

Each row is a single run: So, read the differences of a few points within a category as noise, and differences of 15 to 20 points as signal.

  1. Answer model: Most memory systems do not answer questions themselves. They pass the retrieved info to an LLM that replies to any query. This makes the score a measurement of the LLM, rather than the memory system itself. DolphinBench's results show how large the effect can be: the same memory layer scored 70.67% on the Hermes harness with GPT-5.6-Luna and 32.33% on Claude Code with Claude Sonnet 5. The harness changed along with the model, so the gap isn't the model alone, but it is far larger than most gaps between memory systems.

  2. Judge: Most current benchmarks grade text with a second LLM. So, the judge model’s prompt and its strictness can affect the overall score. We’ll cover the failure modes of this a little later.

  3. Metrics: The LoCoMo paper scored answers with F1 and BLEU-1, which measure word overlap with a reference. Some papers report overlap and judge scores side by side; most vendor results, including ours, now report only the judge score. Most vendors now report only the judge score, because the overlap metrics penalize correct answers phrased differently and the judge scores reward answers that sound right.

    Neither is wrong, but they aren't interchangeable, and a system can rank differently under each. Some benchmarks also report more than one metric. On BEAM 1M, Mem0's result is a 70.1% pass rate and a 0.641 average score; the 64.1 figure usually quoted is the average score, not the pass rate.

  4. Retrieval metrics: End-to-end accuracy doesn't tell you where a failure happened. Retrieval metrics such as recall@k (was the needed memory in the top k results?) separate a search failure from an answer failure. Some harnesses also report accuracy at several retrieval depths. Mem0's evaluation framework, for example, scores at top 10, 20, 50, and 200. A system that needs 200 memories per query to reach its score is making a different trade-off from one that gets there with 10. Deeper retrieval isn't always better, either. On LongMemEval, Mem0's platform scored 94.8% at top 50 and 94.4% at top 200, and its multi-session score fell from 93.2% to 88.0% as more memories entered the prompt.

So, that’s all the main part. But where do these scores go wrong? Let’s find out!

Where do the scores go wrong?

A benchmark score can be off for a number of reasons that have nothing to do with the system under the test. let’s discuss some common problems in order of how they enter the evaluation. These include questions like, how the questions are written, how the answers are graded, how much data there is, and how the benchmark ages.

Keyword shortcut

Most needle-in-a-haystack tests let a model succeed by matching words between the question and the buried fact. NoLiMa removed that overlap, so the model had to infer the connection. In the paper, 11 of 13 models fell to half or less of their short-context score at 32K tokens, and GPT-4o dropped from 99.3% to 69.7%.

NoLiMa tests models, not memory systems, but the same shortcut applies to retrieval as well. Embedding search and BM25 both reward queries that share vocabulary with the stored memory. If a benchmark's questions reuse the wording of the facts they target ("What is Alex's favorite restaurant?" against "Alex's favorite restaurant is..."), retrieval looks better than it will on real user queries, which rarely echo the original phrasing.

Answer-key errors

An independent audit of LoCoMo found 99 score-corrupting errors in the 1,540 scored questions. They include hallucinated facts in the key, wrong temporal reasoning and speaker attribution mistakes. A system that tracks speakers correctly can be marked wrong for disagreeing with the key. Treat a benchmark with no published audit as having an unknown error rate, not zero error rate.

Judge leniency

The same audit as above found LoCoMo's standard judge accepted 62.81% of intentionally wrong answers, while catching specific factual errors about 89% of the time. Judges tend to accept answers on the right topic even when the detail is wrong.

The wider research on LLM judges is mixed. Zheng et al. found GPT-4 as a judge agreed with human preferences more than 80% of the time, about as often as two humans agree, but also documented position, verbosity and self-enhancement bias. Position bias applies to pairwise comparisons, where a judge sees two answers; most memory benchmarks grade one answer against a reference, so it matters less here than leniency does.

Three practices make a judge stricter:

  • Give it a reference answer and ask it to compare, rather than asking whether a response seems right.

  • Break the reference into atomic facts, such as the date, the person and the place, and grade each one.

  • Run two judge models and treat answers they disagree on as uncertain.

Scale versus the context window

A memory benchmark only tests memory if the history doesn't fit comfortably in the answer model's context. LoCoMo conversations average about 9K tokens, which fits inside current context windows. LongMemEval-S, at about 115K tokens, fits in many. LongMemEval-M, at about 1.5M tokens, and BEAM's 10M-token bucket don't.

On the smaller benchmarks, a strong model with the whole transcript is a hard baseline to beat on accuracy, and the case for a memory layer rests on cost and latency instead.

Contamination

LoCoMo and LongMemEval have been public for about 2 years now, and any model trained on recent web data may have seen the conversations, the questions, or both. There's no reliable way to measure how much this inflates scores for a given answer model, which is one reason newer benchmarks with fresh or regenerated data matter

Saturation

Once top systems cluster near the ceiling, a benchmark stops separating them. Our LoCoMo and LongMemEval results show the pattern: we got 98.6% and 98.2% on LongMemEval's single-session categories, where further gains carry little information. MemoryArena makes the broader point: agents with near-saturated LoCoMo performance did poorly on its interdependent multi-session tasks. A high score on a saturated benchmark says that the system has cleared a bar, not that it is ready for long-running agent work.

Conflicting and changing facts

Updating memory when facts change is the least solved problem in this space. On MemoryAgentBench, memory agents built on GPT-4o reached about 60% on single-hop conflict resolution, and every method stayed at 7% or below on multi-hop conflict resolution. Our own numbers show the same pressure. Mem0's knowledge-update score on LongMemEval is 93.6%, below its overall 94.4%, because an add-only design keeps older facts that can surface alongside newer ones. BEAM shows the same pattern, with contradiction resolution at 0.357 on BEAM 1M and 0.325 on BEAM 10M.

So, now we know almost every possible reason your score can go wrong. Let’s understand how big of a difference that error makes in our product.

How big is a real difference?

A benchmark score is an estimate, and most published scores come without an error bar. This section shows how to estimate that margin and how to test whether a gap between two systems is real or not.

  • Sampling error: With n questions and accuracy p, the 95% margin from sampling alone is roughly 1.96 × √(p(1−p)/n). By our calculation, that's about ±1.3 points for LoCoMo's 1,540 scored questions at 92.5%, about ±2.0 points for LongMemEval's 500 questions at 94.4%, and about ±6.9 points for BEAM's 200-question 10M bucket at 48.6%.

    Sampling error
  • Judge and run variance: Sampling error is not the total error. Judges disagree with themselves across runs, and agents behave differently on repeated attempts or if the model hallucinates. Mem0's evaluation also states a ±1 point interval from judge inconsistency alone, on top of sampling error.

  • Paired tests: When two systems answer the same questions, compare them question by question rather than by overall percentage. McNemar's test uses only the questions where the systems disagree, and a paired bootstrap resamples questions to estimate the gap's interval. Both are more sensitive than comparing two independent margins, and almost no vendor comparison reports either.

  • Repeated runs: Sierra's τ-bench introduced pass^k, the share of tasks an agent completes on every one of k attempts, and STATE-Bench reports pass^5. A task that succeeds four times in five looks fine in an average and fails one user in five in production.

Any comparison would not be complete without mentioning the effect of cost, latency and tokens when models are involved in the process. Let’s look into that next.

Cost, latency and tokens

A memory layer costs something to ingest history and something to retrieve it, and it can save money by keeping prompts small. Accuracy alone hides half of that trade.

Tokens per query make results comparable on cost. Mem0's LoCoMo result, for example, uses a mean of 6,956 tokens per query at a top-200 retrieval budget. A similar score at a smaller budget would be the cheaper result, and a slightly higher score at a much larger budget may not be worth the cost.

Read cost and latency four ways:

  • Find the frontier: A system is on the Pareto frontier if nothing else is both more accurate and cheaper. Several systems are often on it at once.

  • Split the bill: Separate agent cost (the model reasoning and calling tools) from memory cost (ingestion and retrieval) to see what the memory layer itself buys.

  • Separate read and write latency: Retrieval is on the user's critical path; ingestion often runs in the background. One blended latency figure hides the one users feel.

  • Stay within one stack: In DolphinBench's results, the same 600 tests cost more than ten times as much under Claude Code with Sonnet 5 as under the Hermes harness. So, you need to compare costs within a configuration, not across them.

Case study: LoCoMo numbers

Put the public LoCoMo results side by side and the method differences become visible.

Source

System

Score

Setup notes

Zep

Temporal knowledge graph

75.14% ± 0.17

Categories 1 to 4

Letta

Filesystem agent

74.0%

Transcripts stored as files

Mem0

Mem0 platform

92.5%

Managed platform (v3 pipeline), top-200 retrieval, single pass; gpt-4o answerer and judge by harness default; ±1 judge interval

These aren't points on one line. The Zep figures differ mainly in how categories were counted. Letta's result comes from a different architecture run under its own setup. Mem0's figure is from the managed platform at a top-200 retrieval budget, with gpt-4o as answerer and judge by harness default, and Mem0's docs note that the platform includes optimizations not in the open-source SDK.
The honest reading of this table is narrow: within one setup, run by one party, some systems scored higher than others. Across rows, almost nothing can be compared directly.

A checklist for any benchmark result

Even if you skip the whole blog above, just keep this checklist in your mind before you select the benchmark for your use.

  1. What category is it? Long-context retrieval, cross-session recall, update handling or agentic tasks.

  2. Who ran it, and has anyone reproduced it? A third-party recompute or rerun carries more weight than a first-party number.

  3. Is the answer key audited? If not, its error rate is unknown.

  4. Which answer model, judge and judge prompt? Scores with different ones aren't comparable.

  5. Which metric? F1, BLEU and judge scores can rank systems differently.

  6. What retrieval budget? Top 10 and top 200 are different operating points.

  7. Were all systems configured and tuned equally? Good comparisons say so explicitly.

  8. How many questions, runs and what margin? Compare gaps to the margin, and prefer paired tests.

  9. Is there a full-context baseline? If it wins on accuracy, the claim is about cost or latency.

  10. Does the history exceed the context window? If not, the test may not be measuring memory.

  11. Are cost, latency and tokens reported? Accuracy alone is half the trade.

  12. Can you check it yourself? Published per-item outputs let you recompute a headline figure.

Running your own evaluation

No public benchmark measures a memory system on your conversations, with your model, within your cost and latency limits. Benchmarks narrow the field; your own tests decide. So, here are four steps, from least to most effort:

  1. Check a shared leaderboard: The Agent Memory Leaderboard runs one answering and grading pipeline for every entrant.

  2. Recompute a published number: When a benchmark publishes per-item results, sum the pass or fail values and compare the total with the headline. This needs no model. It confirms the arithmetic and that the outputs exist; it doesn't validate the answer key, the judge or the method.

  3. Run a smoke test: Pick 10 to 20 real tasks from your product where the agent needs something from an earlier session, and write down the correct outcome for each. With 20 tasks, the margin is roughly ±20 points by our calculation, so this test finds large failures but can't rank two good systems.

  4. Run a public benchmark on your stack: This needs your own agent, an answer model, a judge model, API keys and a memory system, and the model calls cost money. Use the official harness where one exists, start with a subset, run more than once, and report the spread rather than the best run.

A note

We built DolphinBench to test the step that QA benchmarks skip: whether an agent recognizes, without being asked, that it needs something it was told long ago. Instead of asking "Which channel does the team use?", it asks the agent to post an update and grades the tool call. Three personas each bring about 500K tokens of history, and agents take 600 tool-using tests across simulated apps. Each test is certified solvable, passing with the relevant history and failing without it, and every result reports accuracy, total cost and median latency.

At launch, results are single runs, so close margins should be read as ties, and the strongest memory layer changed with the harness and model. Per-task results are public in the repository, and submissions are open to any memory system.

Conclusion

Every score in this guide will be out of date within a year. The method won't be. Find out what a benchmark measures, how the answers were produced and graded, how large the margin is, and whether you can reproduce one number yourself, then test on your own workload.

Frequently asked questions

Q. Which memory benchmark should I trust most?

None on its own. Pick the category that matches your workload: LongMemEval-M or BEAM's larger buckets for long chat histories, MemoryAgentBench for facts that change, and an agentic benchmark such as STATE-Bench or MemoryArena if your agent takes actions. Then confirm the shortlist with a small test on your own tasks.

Q. Why do reported LoCoMo scores for the same system differ so much?

Because the setups differ. Reports vary in which question categories they count, which answer and judge models they use, how deep retrieval goes and how the system is configured. Each of these can move a score by several points. Compare scores only within one setup.

Q. Is LoCoMo still worth reporting in 2026?

As a baseline, yes. As a deciding benchmark, less so. Its conversations average about 9K tokens, which fits in current context windows; an audit found errors in 6.4% of its answer key; and top scores now sit close to its practical ceiling. Report it alongside a longer benchmark rather than on its own.

Q. Does retrieving more memories always improve accuracy?

No. In Mem0's LongMemEval results, accuracy was 94.8% at top 50 and 94.4% at top 200, and the multi-session score fell from 93.2% to 88.0% as more memories entered the prompt. Extra memories add tokens and cost, and past a point they add noise the answer model has to filter.

Q. How big does a gap need to be before it means something?

Larger than the margin. From sampling alone, that's about ±1.3 points on LoCoMo's 1,540 questions and about ±6.9 points on BEAM's 200-question 10M bucket, before judge and run variance. If two systems answered the same questions, a paired test such as McNemar's is the right check. A gap of a few points from a single run is best read as a tie.

Q. Why not put the whole history in the context window instead?

On small benchmarks, that's a strong baseline and often hard to beat on accuracy. The case for a memory layer is cost, latency and scale: it keeps prompts small, and it still works when the history outgrows the window, as in LongMemEval-M at about 1.5M tokens or BEAM's 10M bucket. A good comparison reports a full-context baseline wherever the history fits.

Further reading

These are the primary sources behind this guide, grouped by topic.

Start building with memory

Wire persistent memory into your own agent in about 15 minutes. Free tier, no credit card.

Share on:

aashi dutt

Aashi Dutt

She is a senior technical content writer at Mem0. She covers agent memory architecture and the engineering decisions behind building agents that actually remember. She experiments with new features and turns research into posts developers can put straight to use.

Start building with memory

Free tier, no card