2025
APLOT: Robust Reward Modeling via Adaptive Preference Learning with Optimal Transport
EMNLP 2025
The reward model (RM) plays a crucial role in aligning Large Language Models (LLMs) with human preferences through Reinforcement Learning, where the Bradley-Terry (BT) objective has been recognized as simple yet powerful, specifically for pairwise preference learning. However, BT-based RMs often str