ICLR 2026poster0 citations

Entropy-preserving reinforcement learning

Ben Lipkin, Aleksei Petrenko, Kevin Chen, Erik Wijmans, Marco Francis Cusumano-Towner, Raja Giryes, Philipp Kraehenbuehl

Abstract

Policy gradient algorithms have been a driver of much recent advancement in language model reasoning. One of their most appealing properties is the ability to learn from exploration on their own trajectories, a process crucial for discovering diverse approaches and fostering creative solutions. As we show in this paper, most policy gradient algorithms naturally reduce the entropy---and thus the diversity of explored trajectories---as part of training, yielding a policy increasingly limited in its ability to explore. However, not all algorithms exhibit this collapse in entropy equally. In this paper, we formally analyze the contributions of leading policy gradient objectives on entropy, show which mechanisms they employ to implicitly limit entropy collapse, and propose a new regularization method, REPO, that stabilizes entropy over training through the use of an adaptive controller. Models trained with REPO preserve entropy throughout training, yielding final policies that are, on average, more performant. By preserving entropy in the final policy, REPO-trained models can even be re-trained on evolving data distributions in new environments, unlike their non-entropy-preserving counterparts.

Large language modelreinforcement learningentropyGRPOPPO
BibTeX
@inproceedings{
lipkin2026entropypreserving,
title={Entropy-preserving reinforcement learning},
author={Ben Lipkin and Aleksei Petrenko and Kevin Chen and Erik Wijmans and Marco Francis Cusumano-Towner and Raja Giryes and Philipp Kraehenbuehl},
booktitle={The Fourteenth International Conference on Learning Representations},
year={2026},
url={https://openreview.net/forum?id=E8MR8jgEeZ}
}
Entropy-preserving reinforcement learning · ICLR 2026