2025
Flash Inference: Near Linear Time Inference for Long Convolution Sequence Models and Beyond
ICLR 2025poster
While transformers have been at the core of most recent advancements in sequence generative models, their computational cost remains quadratic in sequence length. Several subquadratic architectures have been proposed to address this computational issue. Some of them, including long convolution seque…