← Search

Belhal Karimi

4 accepted papers

2020

Towards Better Generalization of Adaptive Gradient Methods

NeurIPS 2020poster

Adaptive gradient methods such as AdaGrad, RMSprop and Adam have been optimizers of choice for deep learning due to their fast training speed. However, it was recently observed that their generalization performance is often worse than that of SGD for over-parameterized neural networks. While new alg…

Cited by 25SourcePDFScholar
2019

On the Global Convergence of (Fast) Incremental Expectation Maximization Methods

NeurIPS 2019poster

The EM algorithm is one of the most popular algorithm for inference in latent data models. The original formulation of the EM algorithm does not scale to large data set, because the whole data set is required at each iteration of the algorithm. To alleviate this problem, Neal and Hinton [1998] have…

Cited by 42SourcePDFScholar