← Search

Yunkun Xu

2 accepted papers

2025

Gradient-Adaptive Policy Optimization: Towards Multi-Objective Alignment of Large Language Models

ACL 2025long

Reinforcement Learning from Human Feedback (RLHF) has emerged as a powerful technique for aligning large language models (LLMs) with human preferences. However, effectively aligning LLMs with diverse human preferences remains a significant challenge, particularly when they are conflict. To address t…

2021

Look Before You Leap: Safe Model-Based Reinforcement Learning with Human Intervention

CoRL 2021poster

Safety has become one of the main challenges of applying deep reinforcement learning to real world systems. Currently, the incorporation of external knowledge such as human oversight is the only means to prevent the agent from visiting the catastrophic state. In this paper, we propose MBHI, a novel…

Cited by 15SourceScholar