Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation
Qiyue Gao, Xinyu Pi, Kevin Liu, Junrong Chen, Ruolan Yang, Xinqi Huang, Xinyu Fang, Lu Sun
Abstract
Internal world models (WMs) enable agents to understand the world’s state and predict transitions, serving as the basis for advanced deliberative reasoning.Recent large Vision-Language Models (VLMs), such as GPT-4o and Gemini, exhibit potential as general-purpose WMs. While the latest studies have evaluated and shown limitations in specific capabilities such as visual understanding, a systematic evaluation of VLMs’ fundamental WM abilities remains absent. Drawing on comparative psychology and cognitive science, we propose a two-stage framework that assesses **perception** (visual, spatial, temporal, quantitative, and motion) and **prediction** (mechanistic simulation, transitive inference, compositional inference) to provide an atomic evaluation of VLMs as WMs. Guided by this framework, we introduce **WM-ABench**, a large-scale benchmark comprising 23 fine-grained evaluation dimensions across 6 diverse simulated environments with controlled counterfactual simulations. Through 660 experiments on 15 latest commercial and open-source VLMs, we find that these models exhibit striking limitations in basic world modeling abilities. For instance, all models perform at near-random accuracy when distinguishing motion trajectories. Additionally, they lack disentangled understanding—e.g., they tend to believe blue objects move faster than green ones. More rich results and analyses reveal significant gaps between VLMs and human-level world modeling.
BibTeX
@inproceedings{gao-etal-2025-vision,
title = "Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation",
author = "Gao, Qiyue and
Pi, Xinyu and
Liu, Kevin and
Chen, Junrong and
Yang, Ruolan and
Huang, Xinqi and
Fang, Xinyu and
Sun, Lu and
Kishore, Gautham and
Ai, Bo and
Tao, Stone and
Liu, Mengyang and
Yang, Jiaxi and
Lai, Chao-Jung and
Jin, Chuanyang and
Xiang, Jiannan and
Huang, Benhao and
Chen, Zeming and
Danks, David and
Su, Hao and
Shu, Tianmin and
Ma, Ziqiao and
Qin, Lianhui and
Hu, Zhiting",
editor = "Che, Wanxiang and
Nabende, Joyce and
Shutova, Ekaterina and
Pilehvar, Mohammad Taher",
booktitle = "Findings of the Association for Computational Linguistics: ACL 2025",
month = jul,
year = "2025",
address = "Vienna, Austria",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2025.findings-acl.1342/",
doi = "10.18653/v1/2025.findings-acl.1342",
pages = "26170--26195",
ISBN = "979-8-89176-256-5"
}