2025
Learning to (Learn at Test Time): RNNs with Expressive Hidden States
ICML 2025spotlight
Self-attention performs well in long context but has quadratic complexity. Existing RNN layers have linear complexity, but their performance in long context is limited by the expressive power of their hidden states. We present a practical framework for instantiating sequence modeling layers with lin…