2026
DSGBENCH: A DIVERSE STRATEGIC GAME BENCHMARK FOR EVALUATING LLM-BASED AGENTS IN COMPLEX DECISION-MAKING ENVIRONMENTS
ICASSP 2026poster
Large language model (LLM)-based agents are increasingly applied to complex strategic environments that demand long-horizon reasoning, multi-agent interaction, and decision-making under uncertainty. However, common existing benchmarks either assess isolated skills, lack environmental diversity, or r…