2025
Transformers Learn Low Sensitivity Functions: Investigations and Implications
ICLR 2025poster
Transformers achieve state-of-the-art accuracy and robustness across many tasks, but an understanding of their inductive biases and how those biases differ from other neural network architectures remains elusive. In this work, we identify the sensitivity of the model to token-wise random perturbatio…