Conv-Basis: A New Paradigm for Efficient Attention Inference and Gradient Computation in Transformers
The self-attention mechanism is key to the success of transformers in recent large language models (LLMs). However, the quadratic computational cost, O(n 2 ) , with respect to the input sequence length n poses a significant obstacle to further improvement and scalability in longer contexts.In this w