2026
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping
ICLR 2026poster
Large Reasoning Models (LRMs) have shown remarkable capabilities in solving complex problems through reinforcement learning (RL), particularly by generating long reasoning traces. However, these extended outputs often exhibit substantial redundancy, which limits the efficiency of LRMs. In this paper…