2026
Entropy-preserving reinforcement learning
Ben Lipkin, Aleksei Petrenko, Kevin Chen, Erik Wijmans, Marco Francis Cusumano-Towner, Raja Giryes +1
ICLR 2026poster
Policy gradient algorithms have been a driver of much recent advancement in language model reasoning. One of their most appealing properties is the ability to learn from exploration on their own trajectories, a process crucial for discovering diverse approaches and fostering creative solutions. As w…