Schema-Guided World Modeling for Understanding Hierarchical Visual Dynamics
Silin Gao, Hao Zhao, Zeming Chen, Sepideh Mamooler, Antara R Bhattacharya, Qiyu Wu, Hiromi Wakaki, Yuki Mitsufuji
Abstract
Multimodal LLMs lack a systematic understanding of visual dynamics in complex human world activities, which requires the model to predict or simulate multiple levels of dynamic constituents, such as the general progression of actions and the associated changes of low-level details in the world. To address this challenge, we propose a dynamic visual schema-guided world model, DynaVieW, optimized for visual dynamic prediction and simulation. DynaVieW achieves an in-depth understanding of visual dynamics by learning interleaved state-transition sequences, where states cover broad visual scenes from video keyframes, and transitions capture comprehensive dynamic constituents within a hierarchical schema. DynaVieW jointly models transition prediction and state simulation under a mixture-of-experts architecture, with a cross-expert selective attention and a schema token re-weighted loss, to ensure effective and robust learning. DynaVieW's superior visual dynamic understanding boosts its downstream performances on both visual narrative creation and world simulation, showing improved consistency and controllability of visual generation and better instruction-following ability.
BibTeX
@inproceedings{
gao2026dynaview,
title={DynaVieW: Schema-Guided World Modeling for Understanding Hierarchical Visual Dynamics},
author={Silin Gao and Hao Zhao and Zeming Chen and Sepideh Mamooler and Antara Raaghavi Bhattacharya and Qiyu Wu and Hiromi Wakaki and Yuki Mitsufuji and Li Mi and Syrielle Montariol and Antoine Bosselut},
booktitle={Forty-third International Conference on Machine Learning},
year={2026},
url={https://openreview.net/forum?id=Gy0C1ykv2o}
}