2024
DiJiang: Efficient Large Language Models through Compact Kernelization
ICML 2024oral
In an effort to reduce the computational load of Transformers, research on linear attention has gained significant momentum. However, the improvement strategies for attention mechanisms typically necessitate extensive retraining, which is impractical for large language models with a vast array of pa…