2026
Breaking Barriers: Do Reinforcement Fine-tuning Gains Transfer To Unseen Domains?
ICLR 2026poster
Reinforcement post training (RPT) has recently shown promise in improving the reasoning abilities of large language models (LLMs). However, it remains unclear how well these improvements generalize to new domains, as prior work evaluates RPT models on data from the same domains used for fine-tuning.…