2024
RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
ICML 2024poster
Reinforcement learning from human feedback (RLHF) has proven effective in aligning large language models (LLMs) with human preferences, but gathering high-quality preference labels is expensive. RL from AI Feedback (RLAIF), introduced in Bai et al. (2022b), offers a promising alternative that trains…