2024
Gated Slot Attention for Efficient Linear-Time Sequence Modeling
NeurIPS 2024poster
Linear attention Transformers and their gated variants, celebrated for enabling parallel training and efficient recurrent inference, still fall short in recall-intensive tasks compared to traditional Transformers and demand significant resources for training from scratch. This paper introduces Gated…