2026
Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games
ICLR 2026poster
Large Language Model (LLM) agents are reshaping the game industry, by enabling more intelligent and human-preferable characters. Yet, current game benchmarks fall short of practical needs: they lack evaluations of diverse LLM capabilities across various game genres, studies of agentic modules crucia…