← Search

Caleb Biddulph

3 accepted papers

2026

Breaking Barriers: Do Reinforcement Fine-tuning Gains Transfer To Unseen Domains?

ICLR 2026poster

Reinforcement post training (RPT) has recently shown promise in improving the reasoning abilities of large language models (LLMs). However, it remains unclear how well these improvements generalize to new domains, as prior work evaluates RPT models on data from the same domains used for fine-tuning.…

Cited by 0SourceScholar
2025

MONA: Myopic Optimization with Non-myopic Approval Can Mitigate Multi-step Reward Hacking

ICML 2025poster

Future advanced AI systems may learn sophisticated strategies through reinforcement learning (RL) that humans cannot understand well enough to safely evaluate. We propose a training method which avoids agents learning undesired multi-step plans that receive high reward (multi-step "reward hacks") ev…

Cited by 1SourcePDFScholar