Mixture of Distributions Matters: Dynamic Sparse Attention for Efficient Video Diffusion Transformers
While Diffusion Transformers (DiTs) have achieved notable progress in video generation, this long-sequence generation task remains constrained by the quadratic complexity inherent to self-attention mechanisms, creating significant barriers to practical deployment. Although sparse attention methods a…