← Search

Xavier Bouthillier

3 accepted papers

2018

Fast Approximate Natural Gradient Descent in a Kronecker Factored Eigenbasis

NeurIPS 2018poster

Optimization algorithms that leverage gradient covariance information, such as variants of natural gradient descent (Amari, 1998), offer the prospect of yielding more effective descent directions. For models with many parameters, the covari- ance matrix they are based on becomes gigantic, making the…

Cited by 193SourcePDFScholar
2015

Efficient Exact Gradient Update for training Deep Networks with Very Large Sparse Targets

NeurIPS 2015oral

An important class of problems involves training deep neural networks with sparse prediction targets of very high dimension D. These occur naturally in e.g. neural language models or the learning of word-embeddings, often posed as predicting the probability of next words among a vocabulary of size D…