ICML 2017poster132 citations

Learning Deep Latent Gaussian Models with Markov Chain Monte Carlo

Matthew D. Hoffman

Abstract

Deep latent Gaussian models are powerful and popular probabilistic models of high-dimensional data. These models are almost always fit using variational expectation-maximization, an approximation to true maximum-marginal-likelihood estimation. In this paper, we propose a different approach: rather than use a variational approximation (which produces biased gradient signals), we use Markov chain Monte Carlo (MCMC, which allows us to trade bias for computation). We find that our MCMC-based approach has several advantages: it yields higher held-out likelihoods, produces sharper images, and does not suffer from the variational overpruning effect. MCMC’s additional computational overhead proves to be significant, but not prohibitive.

BibTeX
@InProceedings{pmlr-v70-hoffman17a,
  title = 	 {Learning Deep Latent {G}aussian Models with {M}arkov Chain {M}onte {C}arlo},
  author =       {Matthew D. Hoffman},
  booktitle = 	 {Proceedings of the 34th International Conference on Machine Learning},
  pages = 	 {1510--1519},
  year = 	 {2017},
  editor = 	 {Precup, Doina and Teh, Yee Whye},
  volume = 	 {70},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {06--11 Aug},
  publisher =    {PMLR},
  pdf = 	 {http://proceedings.mlr.press/v70/hoffman17a/hoffman17a.pdf},
  url = 	 {https://proceedings.mlr.press/v70/hoffman17a.html},
  abstract = 	 {Deep latent Gaussian models are powerful and popular probabilistic models of high-dimensional data. These models are almost always fit using variational expectation-maximization, an approximation to true maximum-marginal-likelihood estimation. In this paper, we propose a different approach: rather than use a variational approximation (which produces biased gradient signals), we use Markov chain Monte Carlo (MCMC, which allows us to trade bias for computation). We find that our MCMC-based approach has several advantages: it yields higher held-out likelihoods, produces sharper images, and does not suffer from the variational overpruning effect. MCMC’s additional computational overhead proves to be significant, but not prohibitive.}
}
Learning Deep Latent Gaussian Models with Markov Chain Monte Carlo · ICML 2017