← Search

Xufei Lv

1 accepted papers

2026

The Secret Engine Behind RLHF: It's Contarstive Learning All Along

ICML 2026poster

Alignment of large language models (LLMs) with human values has recently garnered significant attention, with prominent examples including the canonical yet costly Reinforcement Learning from Human Feedback (RLHF) and the simple Direct Preference Optimization (DPO). In this work, we demonstrate that…

Cited by 0SourceScholar