← Search

Muhammed E. Ildiz

2 accepted papers

2024

Mechanics of Next Token Prediction with Self-Attention

AISTATS 2024poster

Transformer-based language models are trained on large datasets to predict the next token given an input sequence. Despite this simple training objective, they have led to revolutionary advances in natural language processing. Underlying this success is the self-attention mechanism. In this work, we…

Cited by 34SourcePDFScholar
2024

Understanding Inverse Scaling and Emergence in Multitask Representation Learning

AISTATS 2024poster

Large language models exhibit strong multitasking capabilities, however, their learning dynamics as a function of task characteristics, sample size, and model complexity remain mysterious. For instance, it is known that, as the model size grows, large language models exhibit emerging abilities where…

Cited by 1SourcePDFScholar