2026
Chasing the Tail: Effective Rubric-based Reward Modeling for Large Language Model Post-Training
ICLR 2026poster
Reinforcement fine-tuning (RFT) often suffers from reward over-optimization, where a policy model hacks the reward signals to achieve high scores while producing low-quality outputs. Our theoretical analysis shows that the key lies in reward misspecification at the high-reward tail: the inability to…