2023
Focused Transformer: Contrastive Training for Context Scaling
NeurIPS 2023poster
Large language models have an exceptional capability to incorporate new information in a contextual manner. However, the full potential of such an approach is often restrained due to a limitation in the effective context length. One solution to this issue is to endow an attention layer with access t…