← Search

Federica Gerace

5 accepted papers

2024

A distributional simplicity bias in the learning dynamics of transformers

NeurIPS 2024poster

The remarkable capability of over-parameterised neural networks to generalise effectively has been explained by invoking a ``simplicity bias'': neural networks prevent overfitting by initially learning simple classifiers before progressing to more complex, non-linear functions. While simplicity bias…

Cited by 7SourcePDFScholar
2024

Learning from higher-order correlations, efficiently: hypothesis tests, random features, and neural networks

NeurIPS 2024poster

Neural networks excel at discovering statistical patterns in high-dimensional data sets. In practice, higher-order cumulants, which quantify the non-Gaussian correlations between three or more variables, are particularly important for the performance of neural networks. But how efficient are neural…

Cited by 0SourcePDFScholar
2020

Critical initialisation in continuous approximations of binary neural networks

ICLR 2020poster

The training of stochastic neural network models with binary ($\pm1$) weights and activations via continuous surrogate networks is investigated. We derive new surrogates using a novel derivation based on writing the stochastic neural network as a Markov chain. This derivation also encompasses existi…

Cited by 3SourceScholar
2020

Generalisation error in learning with random features and the hidden manifold model

ICML 2020poster

We study generalised linear regression and classification for a synthetically generated dataset encompassing different problems of interest, such as learning with random features, neural networks in the lazy training regime, and the hidden manifold model. We consider the high-dimensional regime and…

Cited by 214SourcePDFScholar