ICLR 2025poster0 citations

Mechanism and Emergence of Stacked Attention Heads in Multi-Layer Transformers

Tiberiu Mușat

Abstract

In this paper, I introduce the retrieval problem, a simple yet common reasoning task that can be solved only by transformers with a minimum number of layers, which grows logarithmically with the input size. I empirically show that large language models can solve the task under different prompting formulations without any fine-tuning. To understand how transformers solve the retrieval problem, I train several transformers on a minimal formulation. Successful learning occurs only under the presence of an implicit curriculum. I uncover the learned mechanisms by studying the attention maps in the trained transformers. I also study the training process, uncovering that attention heads always emerge in a specific sequence guided by the implicit curriculum.

mechanistic interpretabilitylarge language modelstransformersemergent abilitiescurriculum learningreasoning
BibTeX
@inproceedings{
musat2025mechanism,
title={Mechanism and emergence of stacked attention heads in multi-layer transformers},
author={Tiberiu Mușat},
booktitle={The Thirteenth International Conference on Learning Representations},
year={2025},
url={https://openreview.net/forum?id=rUC7tHecSQ}
}
Mechanism and Emergence of Stacked Attention Heads in Multi-Layer Transformers · ICLR 2025