← Search

Libin Zhu

8 accepted papers

2025

Emergence in non-neural models: grokking modular arithmetic via average gradient outer product

ICML 2025oral

Neural networks trained to solve modular arithmetic tasks exhibit grokking, a phenomenon where the test accuracy starts improving long after the model achieves 100% training accuracy in the training process. It is often taken as an example of "emergence", where model ability manifests sharply throug…

Cited by 6SourcePDFScholar
2024

Catapults in SGD: spikes in the training loss and their impact on generalization through feature learning

ICML 2024poster

In this paper, we first present an explanation regarding the common occurrence of spikes in the training loss when neural networks are trained with stochastic gradient descent (SGD). We provide evidence that the spikes in the training loss of SGD are "catapults", an optimization phenomenon originall…

2024

Quadratic models for understanding catapult dynamics of neural networks

ICLR 2024poster

While neural networks can be approximated by linear models as their width increases, certain properties of wide neural networks cannot be captured by linear models. In this work we show that recently proposed Neural Quadratic Models can exhibit the "catapult phase" Lewkowycz et al. (2020) that arise…

2023

Neural tangent kernel at initialization: linear width suffices

UAI 2023poster

In this paper we study the problem of lower bounding the minimum eigenvalue of the neural tangent kernel (NTK) at initialization, an important quantity for the theoretical analysis of training in neural networks. We consider feedforward neural networks with smooth activation functions. Without any d…

Cited by 10SourcePDFScholar
2023

Restricted Strong Convexity of Deep Learning Models with Smooth Activations

ICLR 2023poster

We consider the problem of optimization of deep learning models with smooth activation functions. While there exist influential results on the problem from the ``near initialization'' perspective, we shed considerable new light on the problem. In particular, we make two key technical contributions f…

Cited by 13SourcePDFScholar
2022

Transition to Linearity of General Neural Networks with Directed Acyclic Graph Architecture

NeurIPS 2022accept

In this paper we show that feedforward neural networks corresponding to arbitrary directed acyclic graphs undergo transition to linearity as their ``width'' approaches infinity. The width of these general networks is characterized by the minimum in-degree of their neurons, except for the input and f…

Cited by 6SourcePDFScholar
2022

Transition to Linearity of Wide Neural Networks is an Emerging Property of Assembling Weak Models

ICLR 2022spotlight

Wide neural networks with linear output layer have been shown to be near-linear, and to have near-constant neural tangent kernel (NTK), in a region containing the optimization path of gradient descent. These findings seem counter-intuitive since in general neural networks are highly complex models.…

Cited by 6SourcePDFScholar
2020

On the linearity of large non-linear models: when and why the tangent kernel is constant

NeurIPS 2020spotlight

The goal of this work is to shed light on the remarkable phenomenon of "transition to linearity" of certain neural networks as their width approaches infinity. We show that the "transition to linearity'' of the model and, equivalently, constancy of the (neural) tangent kernel (NTK) result from the s…

Cited by 193SourcePDFScholar