2025
Foresight: Adaptive Layer Reuse for Accelerated and High-Quality Text-to-Video Generation
NeurIPS 2025poster
Diffusion Transformers (DiTs) achieve state-of-the-art results in text-to-image, text-to-video generation, and editing. However, their large model size and the quadratic cost of spatial-temporal attention over multiple denoising steps make video generation computationally expensive. Static caching m…