← Search

Anze Xie

1 accepted papers

2025

GameArena: Evaluating LLM Reasoning through Live Computer Games

ICLR 2025poster

Evaluating the reasoning abilities of large language models (LLMs) is challenging. Existing benchmarks often depend on static datasets, which are vulnerable to data contamination and may get saturated over time, or on binary live human feedback that conflates reasoning with other abilities. As the m…

Cited by 2SourcePDFScholar