ICML 2026poster0 citations

Well-Posed KL-Regularized Control via Wasserstein and Kalman–Wasserstein KL Divergences

Viktor Stein, Adwait Datar, Nihat Ay

Abstract

Kullback-Leibler divergence (KL) regularization is widely used in reinforcement learning, but it becomes infinite under support mismatch and can degenerate in low-noise limits. Utilizing a unified information-geometric framework we introduce (Kalman)-Wasserstein-based KL analogues by replacing the Fisher–Rao geometry in the dynamical formulation of the KL with transport-based geometries, and we derive closed-form values for common distribution families. These divergences remain finite under support mismatch and yield a geometric interpretation of regularization heuristics used in Kalman ensemble methods. We demonstrate the utility of these divergences in KL-regularized optimal control. In the fully tractable setting of linear time-invariant systems with Gaussian process noise, the classical KL reduces to a quadratic control penalty that becomes singular as process noise vanishes. Our variants remove this singularity and yield well-posed problems. On a double integrator and a cart-pole example, the resulting controls outperform KL-based regularization.

RL
BibTeX
@inproceedings{
stein2026wellposed,
title={Well-Posed {KL}-Regularized Control via Wasserstein and Kalman{\textendash}Wasserstein {KL} Divergences},
author={Viktor Stein and Adwait Datar and Nihat Ay},
booktitle={Forty-third International Conference on Machine Learning},
year={2026},
url={https://openreview.net/forum?id=diF53wYIj3}
}