2025
MMAU: A Holistic Benchmark of Agent Capabilities Across Diverse Domains
NAACL 2025findings
Recent advances in large language models (LLMs) have increased the demand for comprehensive benchmarks to evaluate their capabilities as human-like agents. Existing benchmarks, while useful, often focus on specific application scenarios, emphasizing task completion but failing to dissect the underly…