ICLR 2026poster0 citations

Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games

Dongmin Park, Minkyu Kim, Beongjun Choi, Junhyuck Kim, Keon Lee, Jonghyun Lee, Inkyu Park, Byeong-Uk Lee

Abstract

Large Language Model (LLM) agents are reshaping the game industry, by enabling more intelligent and human-preferable characters. Yet, current game benchmarks fall short of practical needs: they lack evaluations of diverse LLM capabilities across various game genres, studies of agentic modules crucial for complex gameplay, and fine-tuning datasets to adapt pre-trained LLMs into gaming agents. To fill these gaps, we present Orak, a benchmark for training and evaluating LLM agents across 12 popular video games spanning all major genres. Using a plug-and-play interface built on Model Context Protocol (MCP), Orak supports systematic and reproducible studies of agentic modules in varied game scenarios. We further release a fine-tuning dataset of expert LLM gameplay trajectories spanning multiple genres, turning general LLMs into effective game agents. Orak offers a comprehensive evaluation framework, including game leaderboards, LLM battle arenas, and in-depth analyses of input modality, agentic strategies, and fine-tuning effects, establishing a foundation towards versatile gaming agents. Code is available at https://anonymous.4open.science/r/Orak-5013/.

LLMAgentsBenchmarkGames
BibTeX
@inproceedings{
park2026orak,
title={Orak: A Foundational Benchmark for Training and Evaluating {LLM} Agents on Diverse Video Games},
author={Dongmin Park and Minkyu Kim and Beongjun Choi and Junhyuck Kim and Keon Lee and Jonghyun Lee and Inkyu Park and Byeong-Uk Lee and Jaeyoung Hwang and Jaewoo Ahn and Ameya Sunil Mahabaleshwarkar and Bilal Kartal and Pritam Biswas and Yoshi Suhara and Kangwook Lee and Jaewoong Cho},
booktitle={The Fourteenth International Conference on Learning Representations},
year={2026},
url={https://openreview.net/forum?id=H1ncX6O6Yh}
}
Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games · ICLR 2026