Ridge Rider: Finding Diverse Solutions by Following Eigenvectors of the Hessian
Jack Parker-Holder, Luke Metz, Cinjon Resnick, Hengyuan Hu, Adam Lerer, Alistair Letcher, Alexander Peysakhovich, Aldo Pacchiano
Abstract
Over the last decade, a single algorithm has changed many facets of our lives - Stochastic Gradient Descent (SGD). In the era of ever decreasing loss functions, SGD and its various offspring have become the go-to optimization tool in machine learning and are a key component of the success of deep neural networks (DNNs). While SGD is guaranteed to converge to a local optimum (under loose assumptions), in some cases it may matter which local optimum is found, and this is often context-dependent. Examples frequently arise in machine learning, from shape-versus-texture-features to ensemble methods and zero-shot coordination. In these settings, there are desired solutions which SGD on
BibTeX
@inproceedings{NEURIPS2020_08425b88,
author = {Parker-Holder, Jack and Metz, Luke and Resnick, Cinjon and Hu, Hengyuan and Lerer, Adam and Letcher, Alistair and Peysakhovich, Alexander and Pacchiano, Aldo and Foerster, Jakob},
booktitle = {Advances in Neural Information Processing Systems},
editor = {H. Larochelle and M. Ranzato and R. Hadsell and M.F. Balcan and H. Lin},
pages = {753--765},
publisher = {Curran Associates, Inc.},
title = {Ridge Rider: Finding Diverse Solutions by Following Eigenvectors of the Hessian},
url = {https://proceedings.neurips.cc/paper_files/paper/2020/file/08425b881bcde94a383cd258cea331be-Paper.pdf},
volume = {33},
year = {2020}
}