2024
A distributional simplicity bias in the learning dynamics of transformers
NeurIPS 2024poster
The remarkable capability of over-parameterised neural networks to generalise effectively has been explained by invoking a ``simplicity bias'': neural networks prevent overfitting by initially learning simple classifiers before progressing to more complex, non-linear functions. While simplicity bias…