ICLR 2022poster20 citations

On feature learning in neural networks with global convergence guarantees

Zhengdao Chen, Eric Vanden-Eijnden, Joan Bruna

Abstract

We study the gradient flow optimization of over-parameterized neural networks (NNs) in a setup that allows feature learning while admitting non-asymptotic global convergence guarantees. First, we prove that for wide shallow NNs under the mean-field (MF) scaling and with a general class of activation functions, when the input dimension is at least the size of the training set, the training loss converges to zero at a linear rate under gradient flow. Building upon this analysis, we study a model of wide multi-layer NNs with random and untrained weights in earlier layers, and also prove a linear-rate convergence of the training loss to zero, regardless of the input dimension. We also show empirically that, unlike in the Neural Tangent Kernel (NTK) regime, our multi-layer model exhibits feature learning and can achieve better generalization performance than its NTK counterpart.

neural networksfeature learninggradient descentglobal convergence
BibTeX
@inproceedings{
chen2022on,
title={On feature learning in neural networks with global convergence guarantees},
author={Zhengdao Chen and Eric Vanden-Eijnden and Joan Bruna},
booktitle={International Conference on Learning Representations},
year={2022},
url={https://openreview.net/forum?id=PQTW3iG4sC-}
}
On feature learning in neural networks with global convergence guarantees · ICLR 2022