2026
Restoring Exploration after Post-Training: Latent Exploration Decoding for Large Reasoning Models
ICML 2026poster
Large Reasoning Models (LRMs) have recently achieved strong mathematical and code reasoning performance through Reinforcement Learning (RL) post-training. However, we show that modern reasoning post-training induces an unintended exploration collapse: temperature-based sampling no longer increases p…