2026
Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling
ICLR 2026poster
Videos inherently represent 2D projections of a dynamic 3D world. However, our analysis suggests that video diffusion models trained solely on raw video data often fail to capture meaningful geometric-aware structure in their learned representations. To bridge this gap between video diffusion models…