NeFut Logo NeFut
中 Admin Login

[CS.AI] DyadMem: A Long-Term Memory Benchmark for User-Conditioned Agent Interaction

Published at: 2026-10-06 22:00 Last updated: 2026-10-08 01:25
#AI #Machine Learning #LLM

DyadMem is designed to benchmark how agents retain memory over long‑term interactions with users. Existing benchmarks mainly supervise user facts or preferences and ignore relationship‑specific agent memory—the way an agent should cooperate with a particular user as their shared history evolves. To address this, the authors define User‑conditioned Relational Agent Memory (URAM) and jointly annotate user‑side memory and URAM along the same multi‑session trajectories, yielding six memory categories.

The dataset comprises 3,065 episodes, 50,961 sessions, and 61,210 QA instances. Each session is equipped with gold‑standard Capture and Update annotations, query‑level Recall support, and two QA settings: Gold‑Memory (using gold memory) and Full‑Pipeline (full memory generation).

Experiments on 16 open‑weight and 4 proprietary models show strong performance on Gold‑Memory QA, but a sharp drop on Full‑Pipeline QA, highlighting the bottleneck in memory generation. Quantitative analysis further reveals low Capture recall, incomplete Recall, and unsafe‑deletion issues, even for state‑of‑the‑art LLMs.

A rigorous ablation confirms the effectiveness of URAM: all 20 models improve when URAM is incorporated. In summary, DyadMem offers a dual‑domain, full‑pipeline memory benchmark with extensive fine‑grained annotations, providing a solid foundation for advancing long‑term memory research in LLMs.

Review: DyadMem’s detailed memory taxonomy and end‑to‑end evaluation fill a critical gap in long‑term interaction benchmarks, offering valuable insights for enhancing LLM reliability in real‑world user scenarios.

Original Source: https://arxiv.org/abs/2610.03020

[h] Back to Home