2025
LLMsPark: A Benchmark for Evaluating Large Language Models in Strategic Gaming Contexts
EMNLP 2025
As large language models (LLMs) advance across diverse tasks, the need for comprehensive evaluation beyond single metrics becomes increasingly important.To fully assess LLM intelligence, it is crucial to examine their interactive dynamics and strategic behaviors.We present LLMsPark, a game theory–ba