2024
DiTFastAttn: Attention Compression for Diffusion Transformer Models
NeurIPS 2024poster
Diffusion Transformers (DiT) excel at image and video generation but face computational challenges due to the quadratic complexity of self-attention operators. We propose DiTFastAttn, a post-training compression method to alleviate the computational bottleneck of DiT. We identify three key redundanc…