2026
Self-Aligned Reward: Towards Effective and Efficient Reasoners
ICLR 2026poster
Reinforcement learning with verifiable rewards has significantly advanced reasoning with large language models (LLMs) in domains such as mathematics and logic. However, verifiable signals provide only coarse-grained or binary correctness feedback. This limitation results in inefficiencies like overl…