
·
13 min read
On this page
TL;DR
When a model is handed retrieved memories, it often gives them the wrong amount of influence: irrelevant memories steer the answer, and relevant ones get ignored. MemCalib found that even the best frontier model it tested scored only 46.25 out of 100 on calibrated memory use. MemSyco-Bench found that 61 to 62% of errors in three memory systems happened after the right memory had already been retrieved.
The field has spent years making retrieval better. These papers argue that the next problem sits one step later, in how the model uses what’s retrieved.
Problem
Say, a user's memory contains three facts.
They're vegetarian.
They like mild seasoning.
Their friend likes steak.
The user asks for a dinner recommendation, and their agent suggests a lightly seasoned steak.

Figure: LLM agents over-using irrelevant memory and under-using relevant constraints(Source)
Here, Retrieval did its job because all three memories were relevant enough to come back. The model then gave the friend's preference control over the answer, used the seasoning preference correctly, and ignored the one fact that should have decided everything.
This example is from the MemCalib paper from USTC, Alibaba's Qwen Applications group, and Fudan. It describes a failure that retrieval benchmarks were never built to catch.
Reading it, we started wondering how many other groups had run into the same thing. It turned out to be at least four!
What memory benchmarks usually measure?
Most memory evaluation asks one question: did the system find the right fact? LoCoMo and LongMemEval, the two benchmarks the field reports most, test recall and question answering over long conversation histories. They're good at what they measure.

Figure: Error-cause analysis on existing memory benchmarks(Source)
But the MemSyco-Bench paper actually checked where errors actually come from on four existing benchmarks: LongMemEval, LoCoMo, STALE, and PersonaMem. On those benchmarks, 47.4 to 66.1% of all samples were cases where the evidence wasn't retrieved, and the answer was wrong. Cases where the evidence was retrieved and the answer was still wrong, made up only 5.8 to 13.7%.
That's a useful finding in two directions.
It means existing benchmarks are mostly measuring retrieval, which is fine. But it also means they barely exercise the other failure: the model had the right memory and misused it. If you only test retrieval, you'll only find retrieval problems.
Memory use as calibration
MemCalib splits each retrieved memory into atomic propositions and assigns each one a target level of influence for the current query:
Level | Meaning | Dinner example |
|---|---|---|
Ignore | No footprint on the answer | "My friend likes steak" |
Bound | Local support, doesn't decide the answer | "I like mild seasoning" |
Control | Decides or constrains the core answer | "I am vegetarian" |
Influence above the target is over-use. Influence below it is under-use and both hurt. Over-use lets irrelevant or outdated memories override current evidence. While Under-use throws away constraints that should govern the answer.
The benchmark has 15,000 examples across health, general assistance, and coding. The test set is 1,500. An LLM judge classifies how each atom was actually used, using a rubric specific to that atom.

Figure: An illustration of RPEVAL (Source)
A paper from Renmin University and Huawei, How Does Personalized Memory Shape LLM Behavior?, arrived at almost the same taxonomy. Its benchmark, RPEval, labels each stored preference Ignore, Support, or Dominate, and it frames the problem as pragmatic reasoning. A basic assistant pastes memory into the prompt. A better one infers from the query whether and how a memory should be used.
Gap measurement in 5 ways
We read through 5 papers to identify one failure:
Paper | What it tests | Headline finding |
|---|---|---|
Influence of each memory atom, three domains | Best model tested: SCS 46.25, Exact 28.40 out of 100 | |
Applying or suppressing stored preferences in formal writing | Misapplication rates up to 86.48% | |
Over-personalization with real memory systems attached | Scores 26.2% to 61.1% below the no-memory baseline | |
Classifying each preference as Ignore, Support, or Dominate | Humans: 0.86 on Ignore, models 0.06 to 0.38 | |
Memory-induced sycophancy, five tasks | 61 to 62% of errors occur after the right memory was retrieved |
Note: These papers use different taxonomies, different judges, and different data in the same direction. The magnitudes aren't comparable across rows, so don't read the table as "benchmark X is harder than benchmark Y."
What does the evidence show?
As discussed above, the memory benchmarks do not ask what happens after retrieval or how the retrieved info affects the model output, but the 5 papers show some evidence around this:
Frontier models struggle, and they lean one way: On MemCalib, GPT-5.6-SOL (Table 1) scored best with a Sample Calibration Score (SCS) of 46.25 and an Exact score of 28.40, meaning it used every atom at exactly the right level in about 28% of responses. Every model except Qwen3-8B over-used memory more than it under-used it.

Table1 : Evaluation results of representative models on the MemCalib test set(Source)
Bigger isn't reliably better: Qwen3-8B beat Qwen3.5-35B-A3B on MemCalib (SCS 31.17 vs 26.54). RPEval makes a stronger claim that more capable models are worse at ignoring irrelevant preferences, and calls it an inverse scaling effect. The table below shows Qwen2.5-7B scored 0.06 on Ignore, GPT-5 scored 0.12, and DeepSeek-V3 scored 0.38. We'd read that as "capability doesn't fix it," not "capability makes it worse."

Table 2: Performance of major LLMs on the discriminative intent matching accuracy in RPEVAL(Source)
Retrieval quality isn't the bottleneck: OP-Bench attached five memory setups (plain RAG, Mem0, MemU, MemOS) to four backbone models and scored every one below the no-memory baseline.

Figure: Comparison of over-personalization scores across different base models and memory systems. (Source)
It also measured attention by checking if models gave memory tokens more than twice the attention of the user's query, even after normalizing for length. MemSyco-Bench found that across Mem0, A-Mem, and LightMem, 61 to 62% of errors happened after the relevant memory was retrieved.
Wrong memories change factual answers: When MemSyco-Bench added a misleading memory snippet before factual questions, DeepSeek-V4-Flash's accuracy fell from 56.1% to 40.2%, and its sycophancy rate rose from 24.3% to 52.3%.
Reasoning doesn't fix it by default: BenchPreS compared reasoning and non-reasoning variants of the same models. Turning reasoning on raised the rate of correctly applied preferences, but it also raised the rate of misapplied ones. In the failure traces, the model listed the stored preferences as a checklist and executed all of them.

Figure: Performance comparison of non-reasoning and reasoning-enabled model variants in terms of Misapplication Rate (MR), Appropriate Application Rate (AAR), and IFBench score.(Source)
Updates are hard: MemSyco-Bench tested cases where a stored fact had changed. When a system retrieved both the old and updated versions together, accuracy collapsed. For Mem0 on Qwen3-8B, accuracy was 53.06% when only the updated memory came back and 26.38% when both did. The authors concluded that memory systems need temporal arbitration, not just retrieval.
What fixes exist?
Training the model can fix: MemCalib found that standard post-training methods (GRPO and on-policy self-distillation) improved one direction while worsening the other, which it calls a "calibration seesaw."

Figure: Comparison of optimization behavior: GRPO and OPSD can reduce both errors in principle but tend to prioritize one objective in practice due to coarse credit assignment, while MemCalib-RL better balances the two for better overall performance.(Source)
Proposed method: MemCalib-RL splits the reward into separate channels for each target-and-actual combination and uses counterfactual likelihoods to assign credit to specific tokens. On Qwen3-8B, SCS went from 31.17 for the base model to 79.54, and it was the only method that reduced both over-use and under-use on all three models tested. The gains also transferred to RPEval without extra training.
The catch for most builders: it needs training access and token-level likelihoods, so you can't apply it through an API.
Prompts help, but they cost: BenchPreS tested an instruction to include only preferences appropriate to the task. It cut misapplication sharply, for example from 52.99% to 14.04% on Claude Sonnet 4.5 and from 86.48% to 12.80% on Gemini 3 Pro, at a cost of 0.78 to 3.82 points of correct application. MemSyco-Bench found that a memory-caution instruction helped when memory conflicted with evidence (Full Dialog improved 31.6% on that task) but hurt personalization by 13.0 to 21.0%. Asking "Are you sure?" made things worse on average, with drops of 9.9 to 27.7% across settings, because the model doubled down on memory-shaped answers.
Structured reasoning is a middle path: RPEval's authors proposed RP-Reasoner, which treats memory use as inferring the user's intent.

Figure: RP-Reasoner(Source)
They report it resolved 80% of the bad cases they observed in a large-scale commercial assistant. That's their employer's product, and there's no public data behind the number, so treat it as a promising signal rather than a result.
Notes before you try:
We'd rather say this plainly than have a reader find it.
Data is synthetic: MemCalib, OP-Bench, and MemSyco-Bench build their examples with LLM pipelines and human review. None of it is real user traffic.
Distractors: 84.3% of its atoms in MemCalib have an Ignore target. That makes it partly a test of distractor resistance. The authors address this with per-class results: after MemCalib-RL, correct use rose in all three classes
LLM judges score everything: MemCalib's judge agreed with a human 96.7% of the time on naturally sampled cases, but only 74.0% (κ 0.610) on a stress sample, mostly at the line between Bound and Control. That's exactly the distinction the benchmark cares about. BenchPreS reports 92% agreement with a human on 100 samples.
Where Mem0 fits
Two of these papers tested Mem0, and neither result is flattering, so let's be direct.
What they measured
In OP-Bench, Mem0 on GPT-4o-mini scored 46.32 against 83.10 for no memory, a 44.3% drop. Plain RAG dropped less, and MemU and MemOS dropped more. In MemSyco-Bench, Mem0 scored well below Full Dialog on contextual scope control (13.34 vs 70.00 on Qwen3-8B), and showed the old-plus-updated failure described above. It wasn't uniformly worse, though: on Qwen3-8B's memory-evidence conflict task, Mem0 scored 21.33 against 0.67 for Full Dialog.
Which Mem0 they tested
MemSyco-Bench's appendix lists its configuration: mem0_version: "v1.1", graph memory off, top-10 retrieval, a bge-m3 embedding model, and DeepSeek-V4-Flash as the memory LLM at temperature 0.7. OP-Bench cites the 2025 Mem0 paper and uses top-5 retrieval.
What we take from it
Most of these failures happen on the model side, after retrieval, which is the point of all five papers. A memory layer decides what reaches the model. It doesn't decide how the model weighs it.
That doesn't make retrieval irrelevant.
OP-Bench notes that irrelevant memories reaching the model is part of the problem, and fewer irrelevant memories means less to misuse.
What Mem0 gives you to work with:
A relevance score on every result: On Platform, it's a combined 0 to 1 value from semantic, BM25 keyword, and entity signals. Use it to decide what reaches the model, but tune any cutoff on your own queries.
For example:
Timestamps and expiration dates: These inserted in your prompt can show the model which version of a fact is newer. This is the temporal arbitration MemSyco-Bench says is missing.

Additive write path with explicit corrections: Automatic extraction adds new facts without silently rewriting old ones, so an old fact and its update can both come back. When your app knows a fact changed, correct it with
updateordeletelike this:
What we'd tell a builder
Evaluate use, not just retrieval: Log whether the right memory was retrieved and whether the answer used it correctly. MemSyco-Bench's split between those two cases is easy to copy.
Measure both directions: A fix that cuts over-use often raises under-use. If you only track one, you'll ship the seesaw.
Be careful with blanket caution prompts: They help with conflicts and hurt personalization but test them on your own traffic.
Give the model time information: When old and new versions of a fact can both be retrieved, timestamps are the cheapest signal you have.
Conclusion
The open question these papers leave is what calibration looks like in a live system, with real users and real update patterns rather than synthetic benchmarks. None of them measure that yet. A small version is easy to start: sample a few hundred production queries, label which retrieved memories should have mattered, and score what the model actually did.
Until then, the takeaway we'd stand behind is: retrieving the right memory is necessary, and it isn't enough. How the model weighs what it's given is now the harder half of the problem.
Frequently Asked Questions
Q. If my memory system retrieves the right facts, why does my agent still get answers wrong?
Because retrieval and use are separate steps. Across Mem0, A-Mem, and LightMem, MemSyco-Bench found that 61 to 62% of errors happened after the relevant memory was already retrieved. The model had the right fact and gave it the wrong weight, or let a less relevant memory win.
Q. Can't I just tell the model to only use relevant memories?
It helps, and it has a cost. In BenchPreS, an instruction like that cut misapplied preferences sharply with only a small loss in correct application. In MemSyco-Bench, a general caution instruction helped when memory conflicted with evidence but reduced personalization by 13 to 21%. Test any instruction on both failure types before relying on it.
Q. Will a bigger model or a reasoning model fix this?
Not by default. On MemCalib, a smaller Qwen model beat a larger one. In BenchPreS, turning on reasoning raised both correct and incorrect preference use, because the model treated stored preferences as a checklist. Reasoning did help once the prompt explicitly told it to suppress inappropriate preferences.
Q. What happens to memory when a user changes a fact?
If both the old and new facts are retrieved, models often pick the wrong one. In MemSyco-Bench, Mem0's accuracy on Qwen3-8B fell from 53.06% with only the updated memory to 26.38% with both. Passing timestamps to the model, and correcting stored facts with update or delete when your app knows they changed, gives it a way to tell which is current.
Q. How do I check whether my own agent has this problem?
Take a sample of real queries and label, for each retrieved memory, whether it should have been ignored, used as support, or used to decide the answer. Then compare against what the model actually did. That's a small version of MemCalib's method, and it tells you which direction your system leans, which matters more than any single score.
References
Cao et al., MemCalib: Benchmarking and Optimizing Memory Use in LLM Agents
Yoon et al., BenchPreS: A Benchmark for Context-Aware Personalized Preference Selectivity of Persistent-Memory LLMs
Hu et al., OP-Bench: Benchmarking Over-Personalization for Memory-Augmented Personalized Conversational Agents
Xiang et al., MemSyco-Bench: Benchmarking Sycophancy in Agent Memory
More on Agent Memory

What Is Memory Staleness in AI? Causes, Risks & Solutions
Understand memory staleness in AI agents, why outdated information leads to errors, and the best techniques to maintain accurate long-term memory.
9 min read

AI Agent Memory: Build vs. Buy
Build when extraction rules, data residency, or low write volume make a vendor hard to justify. Buy when your ship date is weeks out, write volume is high, or compliance is in scope. Both paths retrieve, so token savings don't decide it.
32 min read

Memory for the Trades: Persistent Memory for Field Service AI Agents
Persistent memory gives field service AI agents what the business already knows, past repairs, equipment history, and customer preferences, so the next technician arrives prepared.
13 min read

Aashi Dutt
She is a senior technical content writer at Mem0. She covers agent memory architecture and the engineering decisions behind building agents that actually remember. She experiments with new features and turns research into posts developers can put straight to use.
Start building with memory
Free tier, no card






