Robust Bi-Tempered Logistic Loss Based on Bregman Divergences
Ehsan Amid, Manfred K. Warmuth, Rohan Anil, Tomer Koren
Abstract
We introduce a temperature into the exponential function and replace the softmax output layer of the neural networks by a high-temperature generalization. Similarly, the logarithm in the loss we use for training is replaced by a low-temperature logarithm. By tuning the two temperatures, we create loss functions that are non-convex already in the single layer case. When replacing the last layer of the neural networks by our bi-temperature generalization of the logistic loss, the training becomes more robust to noise. We visualize the effect of tuning the two temperatures in a simple setting and show the efficacy of our method on large datasets. Our methodology is based on Bregman divergences and is superior to a related two-temperature method that uses the Tsallis divergence.
BibTeX
@inproceedings{NEURIPS2019_8cd7775f,
author = {Amid, Ehsan and Warmuth, Manfred K. K and Anil, Rohan and Koren, Tomer},
booktitle = {Advances in Neural Information Processing Systems},
editor = {H. Wallach and H. Larochelle and A. Beygelzimer and F. d\textquotesingle Alch\'{e}-Buc and E. Fox and R. Garnett},
pages = {},
publisher = {Curran Associates, Inc.},
title = {Robust Bi-Tempered Logistic Loss Based on Bregman Divergences},
url = {https://proceedings.neurips.cc/paper_files/paper/2019/file/8cd7775f9129da8b5bf787a063d8426e-Paper.pdf},
volume = {32},
year = {2019}
}