2024
Provably Robust DPO: Aligning Language Models with Noisy Feedback
ICML 2024poster
Learning from preference-based feedback has recently gained traction as a promising approach to align language models with human interests. While these aligned generative models have demonstrated impressive capabilities across various tasks, their dependence on high-quality human preference data pos…