2024
Linear Log-Normal Attention with Unbiased Concentration
ICLR 2024poster
Transformer models have achieved remarkable results in a wide range of applications. However, their scalability is hampered by the quadratic time and memory complexity of the self-attention mechanism concerning the sequence length. This limitation poses a substantial obstacle when dealing with long…