2025
GameArena: Evaluating LLM Reasoning through Live Computer Games
ICLR 2025poster
Evaluating the reasoning abilities of large language models (LLMs) is challenging. Existing benchmarks often depend on static datasets, which are vulnerable to data contamination and may get saturated over time, or on binary live human feedback that conflates reasoning with other abilities. As the m…