2026
Replay Failures as Successes: Sample-Efficient Reinforcement Learning for Instruction Following
ICML 2026poster
Reinforcement Learning (RL) has shown promise for aligning Large Language Models (LLMs) to follow instructions with various constraints. Despite the encouraging results, RL improvement inevitably relies on sampling successful, high-quality responses; however, the initial model often struggles to gen…