2024
CoCA: Fusing Position Embedding with Collinear Constrained Attention in Transformers for Long Context Window Extending
ACL 2024long
Self-attention and position embedding are two crucial modules in transformer-based Large Language Models (LLMs). However, the potential relationship between them is far from well studied, especially for long context window extending. In fact, anomalous behaviors that hinder long context extrapolatio…