Sequence Accumulation and Beyond: Infinite Context Length on Single GPU and Large Clusters
Linear sequence modeling methods, such as linear attention, state space modeling, and linear RNNs, have recently been recognized as potential alternatives to softmax attention thanks to their linear complexity and competitive performance. However, although their linear-memory advantage during traini…