ICML 2026poster0 citations

Physics-Guided Motion Loss for Video Generation Model

Bowen Xue, Giuseppe Guarnera, Shuang Zhao, Zahra Montazeri

Abstract

Current video diffusion models generate visually compelling content but often violate basic laws of physics, producing subtle artifacts like rubber-sheet deformations and inconsistent object motion. We introduce a frequency-domain physics prior that improves motion plausibility without modifying model architectures. Our method decomposes common rigid motions (translation, rotation, scaling) into lightweight spectral losses computed on a low-frequency subset. Applied to Open-Sora, MVDIT, and Hunyuan, our approach improves both motion accuracy and action recognition by ~11\% on average on OpenVID-1M (relative), while maintaining visual quality. User studies show 74--83\% preference for our physics-enhanced videos. It also reduces warping error by 22--37\% (depending on the backbone) and improves temporal consistency scores. These results indicate that simple, global spectral cues are an effective drop-in regularizer for physically plausible motion in video diffusion.

DiffusionVisionRetrieval
BibTeX
@inproceedings{
xue2026physicsguided,
title={Physics-Guided Motion Loss for Video Generation Model},
author={Bowen Xue and Giuseppe Claudio Guarnera and Shuang Zhao and Zahra Montazeri},
booktitle={Forty-third International Conference on Machine Learning},
year={2026},
url={https://openreview.net/forum?id=qrwbNL3KIP}
}