NeurIPS 2022accept18 citations

Tight Mutual Information Estimation With Contrastive Fenchel-Legendre Optimization

Qing Guo, Junya Chen, Dong Wang, Yuewei Yang, Xinwei Deng, Jing Huang, Lawrence Carin, Fan Li

Abstract

Successful applications of InfoNCE (Information Noise-Contrastive Estimation) and its variants have popularized the use of contrastive variational mutual information (MI) estimators in machine learning . While featuring superior stability, these estimators crucially depend on costly large-batch training, and they sacrifice bound tightness for variance reduction. To overcome these limitations, we revisit the mathematics of popular variational MI bounds from the lens of unnormalized statistical modeling and convex optimization. Our investigation yields a new unified theoretical framework encompassing popular variational MI bounds, and leads to a novel, simple, and powerful contrastive MI estimator we name FLO. Theoretically, we show that the FLO estimator is tight, and it converges under stochastic gradient descent. Empirically, the proposed FLO estimator overcomes the limitations of its predecessors and learns more efficiently. The utility of FLO is verified using extensive benchmarks, and we further inspire the community with novel applications in meta-learning. Our presentation underscores the foundational importance of variational MI estimation in data-efficient learning.

mutual informationvariational inferencecontrastive learningfew-shot learningmeta learning
BibTeX
@inproceedings{
guo2022tight,
title={Tight Mutual Information Estimation With Contrastive Fenchel-Legendre Optimization},
author={Qing Guo and Junya Chen and Dong Wang and Yuewei Yang and Xinwei Deng and Jing Huang and Lawrence Carin and Fan Li and Chenyang Tao},
booktitle={Advances in Neural Information Processing Systems},
editor={Alice H. Oh and Alekh Agarwal and Danielle Belgrave and Kyunghyun Cho},
year={2022},
url={https://openreview.net/forum?id=M-seILmeISn}
}
Tight Mutual Information Estimation With Contrastive Fenchel-Legendre Optimization · NeurIPS 2022