← Search

Alon Brutzkus

10 accepted papers

2024

How Uniform Random Weights Induce Non-uniform Bias: Typical Interpolating Neural Networks Generalize with Narrow Teachers

ICML 2024spotlight

A main theoretical puzzle is why over-parameterized Neural Networks (NNs) generalize well when trained to zero loss (i.e., so they interpolate the data). Usually, the NN is trained with Stochastic Gradient Descent (SGD) or one of its variants. However, recent empirical work examined the generalizati…

Cited by 5SourcePDFScholar
2022

Efficient Learning of CNNs using Patch Based Features

ICML 2022spotlight

Recent work has demonstrated the effectiveness of using patch based representations when learning from image data. Here we provide theoretical support for this observation, by showing that a simple semi-supervised algorithm that uses patch statistics can efficiently learn labels produced by a one-hi…

2021

Towards Understanding Learning in Neural Networks with Linear Teachers

ICML 2021spotlight

Can a neural network minimizing cross-entropy learn linearly separable data? Despite progress in the theory of deep learning, this question remains unsolved. Here we prove that SGD globally optimizes this learning problem for a two-layer network with Leaky ReLU activations. The learned network can i…

Cited by 29SourcePDFScholar
2019

Why do Larger Models Generalize Better? A Theoretical Perspective via the XOR Problem

ICML 2019oral

Empirical evidence suggests that neural networks with ReLU activations generalize better with over-parameterization. However, there is currently no theoretical analysis that explains this observation. In this work, we provide theoretical and empirical evidence that, in certain cases, overparameteriz…

Cited by 100SourcePDFScholar
2018

SGD Learns Over-parameterized Networks that Provably Generalize on Linearly Separable Data

ICLR 2018poster

Neural networks exhibit good generalization behavior in the over-parameterized regime, where the number of network parameters exceeds the number of observations. Nonetheless, current generalization bounds for neural networks fail to explain this phenomenon. In an attempt to bridge this gap, we study…

Cited by 305SourcePDFScholar