Memoria-Bench: A Comprehensive Benchmark for Evaluating Memory in Long-Horizon Autonomous Agents
Qiufeng Wang, Jiaxuan Zhu, Ziteng Feng, Zhenyu Cui, Jialong Wu, Shuxia Lin, Caorui Li, Renzhao Liang
Abstract
Memory is a core capability of autonomous agents, yet existing benchmarks evaluate it primarily in constrained settings such as short dialogues or synthetic tasks, failing to reflect realistic agent deployments. We present \textbf{Memoria-Bench}, a benchmark for evaluating agent memory grounded in complete, chronologically ordered interaction trajectories that may span millions of tokens. Guided by principles of realism, domain and agent diversity, and explicit exposure of memory-centric challenges, all tasks are formulated as anti-summarization question answering, requiring fine-grained, temporally grounded memory retrieval rather than high-level abstraction. Memoria-Bench covers deep research, coding, and Science \& Development agents across seven domain categories and instantiates three task families: temporal aggregation, multi-hop memory reasoning, and long-range state tracking. Experiments on state-of-the-art long-context models and memory-augmented-based methods reveal substantial performance degradation in long, noisy trajectories, exposing a critical memory bottleneck beyond context length scaling.
BibTeX
@inproceedings{icml2026_memoriabenchacom,
title = {Memoria-Bench: A Comprehensive Benchmark for Evaluating Memory in Long-Horizon Autonomous Agents},
author = {Qiufeng Wang and Jiaxuan Zhu and Ziteng Feng and Zhenyu Cui and Jialong Wu and Shuxia Lin and Caorui Li and Renzhao Liang and Yifei Yu and Kun Wang and Qiankun Li and Guibin Zhang and Siming Huang and Xianzhen Luo and Jie Wang and Junnan Dong and Siyu An and Biao Liu and Yidong Wang and Cunxiang Wang and Yu Chen and Zhenhong Zhou and Liang Lin and Zhongxiang Sun and Deng-Bao Wang and Xu Yang and Yang Liu and Min-Ling Zhang and di yin and Xing Sun and Jiaheng Liu and Qian-Wen Zhang},
booktitle = {ICML 2026},
year = {2026}
}