← Search

Andrea Agazzi

5 accepted papers

2025

A multiscale analysis of mean-field transformers in the moderate interaction regime

NeurIPS 2025oral

In this paper, we study the evolution of tokens through the depth of encoder-only transformer models at inference time by modeling them as a system of particles interacting in a mean-field way and studying the corresponding dynamics. More specifically, we consider this problem in the moderate intera…

Cited by 0SourceScholar
2025

Emergence of meta-stable clustering in mean-field transformer models

ICLR 2025oral

We model the evolution of tokens within a deep stack of Transformer layers as a continuous-time flow on the unit sphere, governed by a mean-field interacting particle system, building on the framework introduced in Geshkovski et al. (2023). Studying the corresponding mean-field Partial Differential…

Cited by 6SourcePDFScholar
2025

Quantitative convergence of trained neural networks to Gaussian processes

NeurIPS 2025poster

In this paper, we study the quantitative convergence of shallow neural networks trained via gradient descent to their associated Gaussian processes in the infinite-width limit. While previous work has established qualitative convergence under broad settings, precise, finite-width estimates rema…

Cited by 0SourceScholar
2021

Global optimality of softmax policy gradient with single hidden layer neural networks in the mean-field regime

ICLR 2021poster

We study the problem of policy optimization for infinite-horizon discounted Markov Decision Processes with softmax policy and nonlinear function approximation trained with policy gradient algorithms. We concentrate on the training dynamics in the mean-field regime, modeling e.g. the behavior of wide…

Cited by 20SourcePDFScholar