2026
Robust Preference Optimization: Aligning Language Models with Noisy Preference Feedback
ICLR 2026poster
Standard human preference-based alignment methods, such as Reinforcement Learning from Human Feedback (RLHF), are a cornerstone technology for aligning Large Language Models (LLMs) with human values. However, these methods are all underpinned by a strong assumption that the collected preference data…