2026
PSPO: Prompt-Level Prioritization and Experience-Weighted Smoothing for Efficient Policy Optimization
AAAI 2026technical
Reinforcement Fine-tuning (RFT) methods such as Group Relative Policy Optimization (GRPO) have demonstrated strong capabilities in aligning Large Language Models with human preferences. However, these approaches often suffer from limited data efficiency, necessitating extensive on-policy rollouts to