2026
Dimension-Free Minimax Rates for Learning Pairwise Interactions in Attention-Style Models
ICLR 2026poster
We study the convergence rate of learning pairwise interactions in single-layer attention-style models, where tokens interact through a weight matrix and a non-linear activation function. We prove that the minimax rate is $M^{-\frac{2\beta}{2\beta+1}}$ with $M$ being the sample size, depending only…