ICLR 2025poster12 citations

VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents

Xiao Liu, Tianjie Zhang, Yu Gu, Iat Long Iong, Song XiXuan, Yifan Xu, Shudan Zhang, Hanyu Lai

Abstract

Large Multimodal Models (LMMs) have ushered in a new era in artificial intelligence, merging capabilities in both language and vision to form highly capable \textbf{Visual Foundation Agents} that are postulated to excel across a myriad of tasks. However, existing benchmarks fail to sufficiently challenge or showcase the full potential of LMMs as visual foundation agents in complex, real-world environments. To address this gap, we introduce VisualAgentBench (VAB), a comprehensive and unified benchmark specifically designed to train and evaluate LMMs as visual foundation agents across diverse scenarios in one standard setting, including Embodied, Graphical User Interface, and Visual Design, with tasks formulated to probe the depth of LMMs' understanding and interaction capabilities. Through rigorous testing across 9 proprietary LMM APIs and 9 open models (18 in total), we demonstrate the considerable yet still developing visual agent capabilities of these models. Additionally, VAB explores the synthesizing of visual agent trajectory data through hybrid methods including Program-based Solvers, LMM Agent Bootstrapping, and Human Demonstrations, offering insights into obstacles, solutions, and trade-offs one may meet in developing open LMM agents. Our work not only aims to benchmark existing models but also provides an instrumental playground for future development into visual foundation agents. Code, train, and test data are available at \url{https://github.com/THUDM/VisualAgentBench}.

Large Multimodal ModelsAgentsEvaluation
BibTeX
@inproceedings{
liu2025visualagentbench,
title={VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents},
author={Xiao Liu and Tianjie Zhang and Yu Gu and Iat Long Iong and Song XiXuan and Yifan Xu and Shudan Zhang and Hanyu Lai and Jiadai Sun and Xinyue Yang and Yu Yang and Zehan Qi and Shuntian Yao and Xueqiao Sun and Siyi Cheng and Qinkai Zheng and Hao Yu and Hanchen Zhang and Wenyi Hong and Ming Ding and Lihang Pan and Xiaotao Gu and Aohan Zeng and Zhengxiao Du and Chan Hee Song and Yu Su and Yuxiao Dong and Jie Tang},
booktitle={The Thirteenth International Conference on Learning Representations},
year={2025},
url={https://openreview.net/forum?id=2snKOc7TVp}
}
VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents · ICLR 2025