2026
Veda: Scalable Video Diffusion via Distilled Sparse Attention
ICML 2026poster
Scaling Diffusion Transformers to generate high-resolution, long videos is constrained by the quadratic cost of self-attention, and existing sparse attention methods degrade under high sparsity. We show empirically that generation quality is determined not by the sparsity ratio itself, but by how we…