2025
Explaining Modern Gated-Linear RNNs via a Unified Implicit Attention Formulation
ICLR 2025poster
Recent advances in efficient sequence modeling have led to attention-free layers, such as Mamba, RWKV, and various gated RNNs, all featuring sub-quadratic complexity in sequence length and excellent scaling properties, enabling the construction of a new type of foundation models. In this paper, we p…