2026
High-Dimensional Analysis of Single-Layer Attention for Sparse-Token Classification
ICLR 2026poster
When and how can an attention mechanism learn to selectively attend to informative tokens, thereby enabling detection of weak, rare, and sparsely located features? We address these questions theoretically in a sparse-token classification model in which positive samples embed a weak signal vector in…