2026
Disentangling Length Bias in Preference Learning via Response-Conditioned Modeling
ICLR 2026poster
Reinforcement Learning from Human Feedback (RLHF) has achieved considerable success in aligning large language models (LLMs) by modeling human preferences with a learnable reward model and employing a reinforcement learning algorithm to maximize the reward model's scores. However, these reward model…