← Search

Misha Belkin

10 accepted papers

2023

Aiming towards the minimizers: fast convergence of SGD for overparametrized problems

NeurIPS 2023poster

Modern machine learning paradigms, such as deep learning, occur in or close to the interpolation regime, wherein the number of model parameters is much larger than the number of data samples. In this work, we propose a regularity condition within the interpolation regime which endows the stochastic…

Cited by 17SourcePDFScholar
2023

Restricted Strong Convexity of Deep Learning Models with Smooth Activations

ICLR 2023poster

We consider the problem of optimization of deep learning models with smooth activation functions. While there exist influential results on the problem from the ``near initialization'' perspective, we shed considerable new light on the problem. In particular, we make two key technical contributions f…

Cited by 13SourcePDFScholar
2022

Benign Overfitting in Two-layer Convolutional Neural Networks

NeurIPS 2022accept

Modern neural networks often have great expressive power and can be trained to overfit the training data, while still achieving a good test performance. This phenomenon is referred to as “benign overfitting”. Recently, there emerges a line of works studying “benign overfitting” from the theoretical…

Cited by 138SourcePDFScholar
2022

Benign, Tempered, or Catastrophic: Toward a Refined Taxonomy of Overfitting

NeurIPS 2022accept

The practical success of overparameterized neural networks has motivated the recent scientific study of \emph{interpolating methods}-- learning methods which are able fit their training data perfectly. Empirically, certain interpolating methods can fit noisy training data without catastrophically ba…

Cited by 45SourcePDFScholar
2022

Transition to Linearity of General Neural Networks with Directed Acyclic Graph Architecture

NeurIPS 2022accept

In this paper we show that feedforward neural networks corresponding to arbitrary directed acyclic graphs undergo transition to linearity as their ``width'' approaches infinity. The width of these general networks is characterized by the minimum in-degree of their neurons, except for the input and f…

Cited by 6SourcePDFScholar
2022

Transition to Linearity of Wide Neural Networks is an Emerging Property of Assembling Weak Models

ICLR 2022spotlight

Wide neural networks with linear output layer have been shown to be near-linear, and to have near-constant neural tangent kernel (NTK), in a region containing the optimization path of gradient descent. These findings seem counter-intuitive since in general neural networks are highly complex models.…

Cited by 6SourcePDFScholar
2021

Risk Bounds for Over-parameterized Maximum Margin Classification on Sub-Gaussian Mixtures

NeurIPS 2021poster

Modern machine learning systems such as deep neural networks are often highly over-parameterized so that they can fit the noisy training data exactly, yet they can still achieve small test errors in practice. In this paper, we study this "benign overfitting" phenomenon of the maximum margin classifi…

Cited by 69SourcePDFScholar
2020

On the linearity of large non-linear models: when and why the tangent kernel is constant

NeurIPS 2020spotlight

The goal of this work is to shed light on the remarkable phenomenon of "transition to linearity" of certain neural networks as their width approaches infinity. We show that the "transition to linearity'' of the model and, equivalently, constancy of the (neural) tangent kernel (NTK) result from the s…

Cited by 193SourcePDFScholar