2026
OPPO: Accelerating PPO-based RLHF via Pipeline Overlap
ICLR 2026poster
Proximal Policy Optimization (PPO)-based reinforcement learning from human feedback (RLHF) is a widely adopted paradigm for aligning large language models (LLMs) with human preferences. However, its training pipeline suffers from substantial inefficiencies due to sequential multi-model dependencies…