2026
Beyond a Million Tokens: Benchmarking and Enhancing Long-Term Memory in LLMs
ICLR 2026poster
Evaluating the abilities of large language models (LLMs) for tasks that require long-term memory and thus long-context reasoning, for example in conversational settings, is hampered by the existing benchmarks, which often lack narrative coherence, cover narrow domains, and only test simple recall-or…