2025
VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents
ICLR 2025poster
Large Multimodal Models (LMMs) have ushered in a new era in artificial intelligence, merging capabilities in both language and vision to form highly capable \textbf{Visual Foundation Agents} that are postulated to excel across a myriad of tasks. However, existing benchmarks fail to sufficiently chal…