Wasserstein Distributional Learning via Majorization-Minimization
Chengliang Tang, Nathan Lenssen, Ying Wei, Tian Zheng
Abstract
Learning function-on-scalar predictive models for conditional densities and identifying factors that influence the entire probability distribution are vital tasks in many data-driven applications. We present an efficient Majorization-Minimization optimization algorithm, Wasserstein Distributional Learning (WDL), that trains Semi-parametric Conditional Gaussian Mixture Models (SCGMM) for conditional density functions and uses the Wasserstein distance $W_2$ as a proper metric for the space of density outcomes. We further provide theoretical convergence guarantees and illustrate the algorithm using boosted machines. Experiments on the synthetic data and real-world applications demonstrate the effectiveness of the proposed WDL algorithm.
BibTeX
@InProceedings{pmlr-v206-tang23b,
title = {Wasserstein Distributional Learning via Majorization-Minimization},
author = {Tang, Chengliang and Lenssen, Nathan and Wei, Ying and Zheng, Tian},
booktitle = {Proceedings of The 26th International Conference on Artificial Intelligence and Statistics},
pages = {10703--10731},
year = {2023},
editor = {Ruiz, Francisco and Dy, Jennifer and van de Meent, Jan-Willem},
volume = {206},
series = {Proceedings of Machine Learning Research},
month = {25--27 Apr},
publisher = {PMLR},
pdf = {https://proceedings.mlr.press/v206/tang23b/tang23b.pdf},
url = {https://proceedings.mlr.press/v206/tang23b.html},
abstract = {Learning function-on-scalar predictive models for conditional densities and identifying factors that influence the entire probability distribution are vital tasks in many data-driven applications. We present an efficient Majorization-Minimization optimization algorithm, Wasserstein Distributional Learning (WDL), that trains Semi-parametric Conditional Gaussian Mixture Models (SCGMM) for conditional density functions and uses the Wasserstein distance $W_2$ as a proper metric for the space of density outcomes. We further provide theoretical convergence guarantees and illustrate the algorithm using boosted machines. Experiments on the synthetic data and real-world applications demonstrate the effectiveness of the proposed WDL algorithm.}
}