2024
Mechanics of Next Token Prediction with Self-Attention
AISTATS 2024poster
Transformer-based language models are trained on large datasets to predict the next token given an input sequence. Despite this simple training objective, they have led to revolutionary advances in natural language processing. Underlying this success is the self-attention mechanism. In this work, we…