Generating Long Videos of Dynamic Scenes
Tim Brooks, Janne Hellsten, Miika Aittala, Ting-chun Wang, Timo Aila, Jaakko Lehtinen, Ming-Yu Liu, Alexei A Efros
Abstract
We present a video generation model that accurately reproduces object motion, changes in camera viewpoint, and new content that arises over time. Existing video generation methods often fail to produce new content as a function of time while maintaining consistencies expected in real environments, such as plausible dynamics and object persistence. A common failure case is for content to never change due to over-reliance on inductive bias to provide temporal consistency, such as a single latent code that dictates content for the entire video. On the other extreme, without long-term consistency, generated videos may morph unrealistically between different scenes. To address these limitations, we prioritize the time axis by redesigning the temporal latent representation and learning long-term consistency from data by training on longer videos. We leverage a two-phase training strategy, where we separately train using longer videos at a low resolution and shorter videos at a high resolution. To evaluate the capabilities of our model, we introduce two new benchmark datasets with explicit focus on long-term temporal dynamics.
BibTeX
@inproceedings{
brooks2022generating,
title={Generating Long Videos of Dynamic Scenes},
author={Tim Brooks and Janne Hellsten and Miika Aittala and Ting-chun Wang and Timo Aila and Jaakko Lehtinen and Ming-Yu Liu and Alexei A Efros and Tero Karras},
booktitle={Advances in Neural Information Processing Systems},
editor={Alice H. Oh and Alekh Agarwal and Danielle Belgrave and Kyunghyun Cho},
year={2022},
url={https://openreview.net/forum?id=VnAwNNJiwDb}
}