2026
AMA-Bench: Evaluating Long-Horizon Memory for Agentic Applications
ICML 2026poster
Large Language Models (LLMs) are deployed as autonomous agents in increasingly complex applications, where enabling long horizon memory is critical for achieving strong performance. However, a significant gap exists between practical applications and current evaluation standards for agent memory: ex…