2026
Learning What Reinforcement Learning Can't: Interleaved Online Fine-Tuning for Hardest Questions
ICLR 2026poster
Recent advances in large language model (LLM) reasoning have shown that reasoning ability can emerge through reinforcement learning (RL). However, despite these successes, RL in its current form remains insufficient to induce capabilities that exceed the limitations of the base model, as it is prima…