ICLR 2021poster54 citations

The inductive bias of ReLU networks on orthogonally separable data

Mary Phuong, Christoph H Lampert

Abstract

We study the inductive bias of two-layer ReLU networks trained by gradient flow. We identify a class of easy-to-learn (`orthogonally separable') datasets, and characterise the solution that ReLU networks trained on such datasets converge to. Irrespective of network width, the solution turns out to be a combination of two max-margin classifiers: one corresponding to the positive data subset and one corresponding to the negative data subset. The proof is based on the recently introduced concept of extremal sectors, for which we prove a number of properties in the context of orthogonal separability. In particular, we prove stationarity of activation patterns from some time $T$ onwards, which enables a reduction of the ReLU network to an ensemble of linear subnetworks.

inductive biasimplicit biasgradient descentReLU networksmax-marginextremal sector
BibTeX
@inproceedings{
phuong2021the,
title={The inductive bias of Re{\{}LU{\}} networks on orthogonally separable data},
author={Mary Phuong and Christoph H Lampert},
booktitle={International Conference on Learning Representations},
year={2021},
url={https://openreview.net/forum?id=krz7T0xU9Z_}
}
The inductive bias of ReLU networks on orthogonally separable data · ICLR 2021