2025
Gradient-Adaptive Policy Optimization: Towards Multi-Objective Alignment of Large Language Models
ACL 2025long
Reinforcement Learning from Human Feedback (RLHF) has emerged as a powerful technique for aligning large language models (LLMs) with human preferences. However, effectively aligning LLMs with diverse human preferences remains a significant challenge, particularly when they are conflict. To address t…