2026
Attention as a Compass: Efficient Exploration for Process-Supervised RL in Reasoning Models
ICLR 2026poster
Reinforcement Learning (RL) has shown remarkable success in enhancing the reasoning capabilities of Large Language Models (LLMs). Process-Supervised RL (PSRL) has emerged as a more effective paradigm compared to outcome-based RL. However, existing PSRL approaches suffer from limited exploration effi…