Agent Memory · AI Agents

Agent Memory: External Stores vs Graphs vs Side-States vs Patches

MemGPT pages memory like an OS, MRAgent rebuilds a memory graph while reasoning, δ-mem compresses history into an 8×8 state, and EvoMem stores change patches: four answers to where agent memory should live.

Agent Memory: External Stores vs Graphs vs Side-States vs Patches

Four answers to one question

Every agent memory system is an answer to the same question: where does the past live, and who decides what gets loaded? The four systems in this comparison give four genuinely different answers, and the differences matter more than the benchmark scores, because they dictate cost, auditability, and what kinds of forgetting are even possible.

MemGPT says memory lives outside the model, in tiered external stores, and the LLM itself pages data in and out through function calls, an operating system in miniature. MRAgent also keeps memory outside the model, but as a Cue-Tag-Content graph that the agent walks and prunes while reasoning, instead of retrieving a fixed top-k in one shot. δ-mem says memory should live inside the model as a tiny 8×8 associative state updated by a delta rule, read out as low-rank corrections to attention, so there is no retrieval step at all. And EvoMem, evaluated on the EvoArena benchmark, says the failure is not where memory lives but what it records: agents need update patches that preserve how facts changed, not just the latest snapshot.

Key numbers

SystemWhere memory livesHeadline resultBenchmark and settingSame harness?
MemGPT (GPT-4 backend)External: recall + archival stores, paged by the LLM93.4% vs 35.3% for recursive summarizationDeep Memory Retrieval, 500 multi-session conversationsYes, same backend and benchmark
MRAgent (Gemini-2.5-Flash backbone)External: Cue-Tag-Content graph, actively traversed84.21 vs 68.31 for Mem0 (+23.3% relative)LoCoMo, LLM-Judge, frozen shared backboneYes, harness vs harness
MRAgent cost—118k tokens and 586s per sample vs A-Mem 632k tokens, LangMem 3,268kLongMemEvalYes, per-sample token accounting
δ-mem (frozen full-attention LLM)Internal: fixed 8×8 online state, delta-rule updates1.10× over backbone, 1.15× over strongest competing memory methodAverage across long-memory tasksYes, same backbone family
δ-mem per-benchmark—1.31× on MemoryAgentBench, 1.20× on LoCoMoStated as relative multipliersSame-harness, but ratios not absolute scores
EvoArena base agentsLatest-state memory39.6% average accuracy; chain accuracy 21.5% / 10.6% / 39.1%Terminal-Bench-Evo / SWE-Chain-Evo / PersonaMem-EvoYes, agents evaluated on all three domains
EvoMemPatches recording what changed, why, and the evidence+1.5 avg, +2.6 step, +3.7 chain accuracy; GAIA +6.1, LoCoMo +4.8EvoArena plus standard benchmarksYes, patch memory vs latest-state on same agents

Read the table for what it does not contain: no two rows share a benchmark. MemGPT’s 93.4% is a different task from MRAgent’s 84.21, and neither is directly comparable to δ-mem’s 1.15× multiplier or EvoMem’s +3.7 chain points. The honest cross-system claims are about design (cost structure, failure modes, what gets preserved), not about ranking these scores.

Paging: MemGPT and the OS analogy

MemGPT’s core move is to stop treating the context window as the memory and start treating it as RAM. The system gives the LLM a fixed system instruction, a writable working context, and tools to append, search, and evict; when the FIFO message queue approaches its token limit, the model gets a warning before overflow and decides what to flush to recall storage and what to summarize. On Deep Memory Retrieval, a benchmark of 500 conversations of 5 sessions each where a question can only be answered by recalling an earlier session, this self-managed paging reaches 93.4% with GPT-4, against 35.3% for recursive summarization, which compresses old turns into a lossy running summary. The gap is the price of summarization: the fact you need was averaged away three sessions ago.

The design’s weakness is symmetric to its strength. Every memory operation is an LLM decision, so a backend that forgets to save, or searches badly, silently loses information; the system delegates retention instead of guaranteeing it. And one user turn can trigger several searches and writes, each a paid model call, so latency and cost scale with how actively the agent manages its own memory.

Graph memory: MRAgent and active reconstruction

MRAgent keeps memory external too, but changes the retrieval contract. Instead of fetching a fixed top-k and reasoning once, the agent interleaves traversal and reasoning: each turn it selects tags from the Cue-Tag-Content graph, expands the matching branches, prunes what is irrelevant, and repeats until a checker decides the evidence answers the query. The paper proves the point formally: for any retrieval budget of at least 2, the passive retrieve-then-reason class is strictly contained in the active one. The benchmarks show it concretely: on LoCoMo, temporal questions jump from Mem0’s 61.68 to 80.37, because answering “when did the tournament happen?” requires inferring “July” as a temporal anchor from intermediate evidence, something an upfront retriever cannot do.

The striking row in the table is the cost. MRAgent does more reasoning per query than its baselines yet ends up the cheapest system tested: 118k tokens per sample on LongMemEval versus A-Mem’s 632k and LangMem’s 3,268k. The trick is deferred work: graph construction stays lightweight, and relational reasoning happens at query time, scoped to one question, with tags filtering out expensive episodic text before it is loaded. The caveat: both backbones are closed frontier models, and the gains assume an LLM reliable enough to infer the right cues mid-search.

Side-states: δ-mem and memory without retrieval

δ-mem takes the opposite bet: skip the store entirely. It bolts a fixed 8×8 associative-memory matrix onto a frozen full-attention LLM. New tokens update the state by a delta rule, writing the error between what the state predicts and what actually arrived, so stale associations get overwritten instead of accumulating; the readout injects low-rank corrections into the backbone’s attention at generation time. Because the state never grows, per-step cost is flat no matter how long the conversation runs, and there is no retrieval latency and no context re-spend.

Reported as multipliers, the gains are 1.10× over the frozen backbone on average and 1.15× over the strongest competing memory method, with 1.31× on MemoryAgentBench and 1.20× on LoCoMo. Read those carefully: they are relative ratios, so absolute scores and the identity of the baselines matter, and a 1.10× average over a memoryless model is real but modest. The ceiling question is intrinsic: an 8×8 state is a very small place to keep a very long history, and the paper does not show where it starts dropping facts the delta rule cannot preserve. This architecture also requires owning the model stack; it is not something you add to a hosted API model.

Patches: EvoMem and memory of change

EvoArena reframes the problem from capacity to validity. Its three domains (terminal tasks, software repositories, and user preferences) are packaged as versioned evolution chains, and the results are sobering: base agents average 39.6% across the evolving domains, and chain accuracy collapses far below step accuracy (21.5% on Terminal-Bench-Evo, 10.6% on SWE-Chain-Evo) because an agent that solves one snapshot still fails when the CLI, dependency, or preference changes underneath it. Latest-state memory is the culprit: when an update overwrites the old rule, the record of the change itself is gone, and later tasks that depend on the transition become unsolvable.

EvoMem’s fix is small and diagnostic: record non-additive updates as structured patches (what changed, why, how the new state differs, and which evidence triggered it) and retrieve patches alongside current memory. The gain is modest (+1.5 average on EvoArena, with +2.6 step and +3.7 chain accuracy, plus +6.1 on GAIA and +4.8 on LoCoMo), but the direction is the point: preserving update history is a better primitive than storing the latest snapshot, and the chain-level numbers say agents are far from reliable either way.

When to use which

  • You need unbounded, auditable memory on a hosted model: MemGPT-style paging. It works with off-the-shelf function-calling models, and every memory write is an explicit, inspectable action.
  • Queries are multi-hop or temporal, and one-shot top-k keeps missing: graph memory with active reconstruction, MRAgent-style. Budget for a stronger backend, since the loop lives and dies on the model’s mid-search inferences.
  • You control the model weights and want zero retrieval latency: a parametric side-state like δ-mem. Accept that you cannot audit what the 8×8 state remembers.
  • Your environment changes (APIs, repos, user preferences): patch memory on top of whatever store you already have. EvoMem’s gains are small but the failure it fixes is structural, and it is the cheapest of the four to adopt.

Limits and open questions

No controlled study puts these four systems on one benchmark with matched backbones, so design-level claims are solid and ranking-level claims are not. The dialogue-memory results (DMR, LoCoMo, LongMemEval) come from conversational settings; whether the same ordering holds for tool-use logs, code histories, or multimodal streams is untested. δ-mem’s ratios hide absolute scores, MRAgent’s best LongMemEval number (86.76) uses a mixed-backbone setup, and EvoMem’s gains, while directionally consistent, leave chain accuracy below 15% on two of three EvoArena domains. The open question that cuts across all four: every system assumes the model is the unreliable part and memory infrastructure will save it. EvoArena’s 39.6% average suggests the bottleneck may be reasoning about change, not storage.

FAQ

What are the main types of LLM agent memory?

Four architectures dominate the recent literature: external stores paged in and out by the LLM itself (MemGPT), external graphs traversed during reasoning (MRAgent), small parametric states inside the model updated by a delta rule (δ-mem), and patch histories that record how stored facts changed (EvoMem). They differ in where memory lives, who controls loading, and what kinds of forgetting are possible.

Does retrieval beat a longer context window for agent memory?

On the evidence here, yes for long-horizon recall. MemGPT reaches 93.4% on Deep Memory Retrieval where recursive summarization gets 35.3%, and δ-mem’s premise is that a bigger window does not fix utilization (the “lost in the middle” problem). But both still need the model to reason correctly over what is loaded.

Which agent memory system is cheapest?

It depends on the cost axis. MRAgent reports the lowest token bill on LongMemEval (118k tokens per sample versus A-Mem’s 632k and LangMem’s 3,268k) despite doing more reasoning. δ-mem has no retrieval cost at all but requires modifying the model. MemGPT’s cost scales with how many memory function calls each turn triggers.

Why does latest-state memory fail when facts change?

Because most memory stores keep only the latest state. EvoArena shows agents average 39.6% on evolving environments, with chain accuracy as low as 10.6% on SWE-Chain-Evo; EvoMem’s patch memory, which preserves what changed and why, recovers +3.7 chain accuracy, a real but partial fix.

Can I combine these architectures?

Yes, and the papers invite it. EvoMem is explicitly an add-on over an existing memory system rather than a replacement, and MemGPT’s tiers are agnostic to what the archival store’s index is. The genuinely exclusive choice is δ-mem, which requires access to the backbone’s attention computation.

One line: decide where memory lives before you benchmark: paging, graphs, side-states and patches optimize for different failures. Read the MemGPT, MRAgent, δ-mem and EvoArena explainers for the full breakdown of each.