NeurIPS 2025poster0 citations

MemSim: A Bayesian Simulator for Evaluating Memory of LLM-based Personal Assistants

Zeyu Zhang, Quanyu Dai, Luyu Chen, Zeren Jiang, Rui Li, Jieming Zhu, Xu Chen, Yi Xie

Abstract

LLM-based agents have been widely applied as personal assistants, capable of memorizing information from user messages and responding to personal queries. However, there still lacks an objective and automatic evaluation on their memory capability, largely due to the challenges in constructing reliable questions and answers (QAs) according to user messages. In this paper, we propose MemSim, a Bayesian simulator designed to automatically construct reliable QAs from generated user messages, simultaneously keeping their diversity and scalability. Specifically, we introduce the Bayesian Relation Network (BRNet) and a causal generation mechanism to mitigate the impact of LLM hallucinations on factual information, facilitating the automatic creation of an evaluation dataset. Based on MemSim, we generate a dataset in the daily-life scenario, named MemDaily, and conduct extensive experiments to assess the effectiveness of our approach. We also provide a benchmark for evaluating different memory mechanisms in LLM-based agents with the MemDaily dataset.

LLM-based AgentsMemoryPersonal Assistants
BibTeX
@inproceedings{
zhang2025memsim,
title={MemSim: A Bayesian Simulator for Evaluating Memory of {LLM}-based Personal Assistants},
author={Zeyu Zhang and Quanyu Dai and Luyu Chen and Zeren Jiang and Rui Li and Jieming Zhu and Xu Chen and Yi Xie and Zhenhua Dong and Ji-Rong Wen},
booktitle={The Thirty-ninth Annual Conference on Neural Information Processing Systems},
year={2025},
url={https://openreview.net/forum?id=vAT2xlaWJY}
}
MemSim: A Bayesian Simulator for Evaluating Memory of LLM-based Personal Assistants · NeurIPS 2025