← Search

Diankun Wu

2 accepted papers

2026

Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling

ICLR 2026poster

Videos inherently represent 2D projections of a dynamic 3D world. However, our analysis suggests that video diffusion models trained solely on raw video data often fail to capture meaningful geometric-aware structure in their learned representations. To bridge this gap between video diffusion models…

Cited by 0SourcecodeScholar
2025

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence

NeurIPS 2025spotlight

Recent advancements in Multimodal Large Language Models (MLLMs) have significantly enhanced performance on 2D visual tasks. However, improving their spatial intelligence remains a challenge. Existing 3D MLLMs always rely on additional 3D or 2.5D data to incorporate spatial awareness, restricting the…

Cited by 0SourceScholar