2026
Transformers Provably Learn Algorithmic Solutions for Graph Connectivity, But Only with the Right Data
ICML 2026poster
Transformers often fail to learn generalizable algorithms, instead relying on brittle heuristics. Using graph connectivity as a testbed, we explain this phenomenon both theoretically and empirically. We consider a simplified Transformer architecture, the Disentangled Transformer, and prove that an $…