ICML 2025poster0 citations

Wasserstein Policy Optimization

David Pfau, Ian Davies, Diana L Borsa, João Guilherme Madeira Araújo, Brendan Daniel Tracey, Hado van Hasselt

Abstract

We introduce Wasserstein Policy Optimization (WPO), an actor-critic algorithm for reinforcement learning in continuous action spaces. WPO can be derived as an approximation to Wasserstein gradient flow over the space of all policies projected into a finite-dimensional parameter space (e.g., the weights of a neural network), leading to a simple and completely general closed-form update. The resulting algorithm combines many properties of deterministic and classic policy gradient methods. Like deterministic policy gradients, it exploits knowledge of the *gradient* of the action-value function with respect to the action. Like classic policy gradients, it can be applied to stochastic policies with arbitrary distributions over actions -- without using the reparameterization trick. We show results on the DeepMind Control Suite and a magnetic confinement fusion task which compare favorably with state-of-the-art continuous control methods.

Policy OptimizationWasserstein metricOptimal TransportGradient FlowDeep Reinforcement LearningActor-CriticContinuous Control
BibTeX
@inproceedings{
pfau2025wasserstein,
title={Wasserstein Policy Optimization},
author={David Pfau and Ian Davies and Diana L Borsa and Jo{\~a}o Guilherme Madeira Ara{\'u}jo and Brendan Daniel Tracey and Hado van Hasselt},
booktitle={Forty-second International Conference on Machine Learning},
year={2025},
url={https://openreview.net/forum?id=oAKe7MG9GM}
}
Wasserstein Policy Optimization · ICML 2025