← Search

Atsushi Yamamura

2 accepted papers

2023

Stochastic Collapse: How Gradient Noise Attracts SGD Dynamics Towards Simpler Subnetworks

NeurIPS 2023poster

In this work, we reveal a strong implicit bias of stochastic gradient descent (SGD) that drives overly expressive networks to much simpler subnetworks, thereby dramatically reducing the number of independent parameters, and improving generalization. To reveal this bias, we identify _invariant sets_,…

2023

The Asymmetric Maximum Margin Bias of Quasi-Homogeneous Neural Networks

ICLR 2023top-25%

In this work, we explore the maximum-margin bias of quasi-homogeneous neural networks trained with gradient flow on an exponential loss and past a point of separability. We introduce the class of quasi-homogeneous models, which is expressive enough to describe nearly all neural networks with homogen…

Cited by 29SourcePDFScholar