Diff4Splat: Repurposing Video Diffusion Models for Dynamic Scene Generation
Panwang Pan, Chenguo Lin, Chenxin Li, Jingjing Zhao, Yuchen Lin, Haopeng Li, Yunlong Lin, Kairun Wen
Abstract
We introduce Diff4Splat, a feed-forward framework for dynamic scene generation from a single image. Our method synergizes the powerful generative priors of video diffusion models with geometric and motion constraints learned from a large-scale 4D dataset. Given a single image, a camera trajectory, and an optional text prompt, our model directly predicts a dynamic scene represented by a deformable 3D Gaussian field. This approach captures appearance, geometry, and motion in a single pass, eliminating the need for test-time optimization or post-hoc processing. At the core of our framework is a video latent transformer that enhances existing video diffusion models, enabling them to jointly model spatio-temporal dependencies and predict 3D Gaussian Primitives over time. Supervised by objectives targeting appearance fidelity, geometric accuracy, and motion consistency, Diff4Splat generates high-fidelity dynamic scenes within 30 seconds. We demonstrate the effectiveness of Diff4Splat across video generation, novel view synthesis, and geometry extraction, where it matches or surpasses optimization-based methods for dynamic scene synthesis while being significantly more efficient.
BibTeX
@inproceedings{cvpr2026_diff4splatrepurp,
title = {Diff4Splat: Repurposing Video Diffusion Models for Dynamic Scene Generation},
author = {Panwang Pan and Chenguo Lin and Chenxin Li and Jingjing Zhao and Yuchen Lin and Haopeng Li and Yunlong Lin and Kairun Wen and Yixuan Yuan and Yadong MU},
booktitle = {CVPR 2026},
year = {2026}
}