← Search

Chaoyue Liu

10 accepted papers

2025

Better NTK Conditioning: A Free Lunch from (ReLU) Nonlinear Activation in Wide Neural Networks

NeurIPS 2025poster

Nonlinear activation functions are widely recognized for enhancing the expressivity of neural networks, which is the primary reason for their widespread implementation. In this work, we focus on ReLU activation and reveal a novel and intriguing property of nonlinear activations. By comparing enablin…

Cited by 0SourceScholar
2024

Catapults in SGD: spikes in the training loss and their impact on generalization through feature learning

ICML 2024poster

In this paper, we first present an explanation regarding the common occurrence of spikes in the training loss when neural networks are trained with stochastic gradient descent (SGD). We provide evidence that the spikes in the training loss of SGD are "catapults", an optimization phenomenon originall…

2024

Quadratic models for understanding catapult dynamics of neural networks

ICLR 2024poster

While neural networks can be approximated by linear models as their width increases, certain properties of wide neural networks cannot be captured by linear models. In this work we show that recently proposed Neural Quadratic Models can exhibit the "catapult phase" Lewkowycz et al. (2020) that arise…

2023

Aiming towards the minimizers: fast convergence of SGD for overparametrized problems

NeurIPS 2023poster

Modern machine learning paradigms, such as deep learning, occur in or close to the interpolation regime, wherein the number of model parameters is much larger than the number of data samples. In this work, we propose a regularity condition within the interpolation regime which endows the stochastic…

Cited by 17SourcePDFScholar
2022

Transition to Linearity of General Neural Networks with Directed Acyclic Graph Architecture

NeurIPS 2022accept

In this paper we show that feedforward neural networks corresponding to arbitrary directed acyclic graphs undergo transition to linearity as their ``width'' approaches infinity. The width of these general networks is characterized by the minimum in-degree of their neurons, except for the input and f…

Cited by 6SourcePDFScholar
2022

Transition to Linearity of Wide Neural Networks is an Emerging Property of Assembling Weak Models

ICLR 2022spotlight

Wide neural networks with linear output layer have been shown to be near-linear, and to have near-constant neural tangent kernel (NTK), in a region containing the optimization path of gradient descent. These findings seem counter-intuitive since in general neural networks are highly complex models.…

Cited by 6SourcePDFScholar
2020

On the linearity of large non-linear models: when and why the tangent kernel is constant

NeurIPS 2020spotlight

The goal of this work is to shed light on the remarkable phenomenon of "transition to linearity" of certain neural networks as their width approaches infinity. We show that the "transition to linearity'' of the model and, equivalently, constancy of the (neural) tangent kernel (NTK) result from the s…

Cited by 193SourcePDFScholar