2026
OnePO: Direct One-stage Policy Optimization for SFT-free Domain Adaptation
ICML 2026poster
Domain adaptation transforms general-purpose LLMs into specialized experts for specific domains or tasks. This process typically follows a two-stage recipe: first, Supervised Fine-Tuning (SFT) to inject domain knowledge or induce specific behaviors (e.g., reasoning patterns), followed by Reinforceme…