DeltaQuant: 4-bit Video Diffusion Models with Spatiotemporal Delta Smoothing
Video diffusion models have achieved remarkable generative performance, but their substantial computational and memory costs pose significant challenges for deployment, especially on consumer GPUs. As recent advances in attention optimization mitigate previous computational bottlenecks, linear layer