← Search

Xiyue Peng

5 accepted papers

2026

Keep the Best, Forget the Rest: Reliable Alignment with Order-Aware Preference Optimization

ICLR 2026poster

Direct Preference Optimization (DPO) has emerged as a powerful framework for aligning large language models (LLMs) with human preferences via pairwise comparisons. However, its performance is highly sensitive to the quality of training samples: when the reference policy is poorly aligned with human…

Cited by 0SourcecodeScholar
2026

Towards Achieving Optimal Strong Regret and Constraint Violation via Computational Efficient Model-free RL

ICML 2026poster

We study episodic constrained Markov decision processes (CMDPs) with linear function approximation, where the goal is to achieve strong regret and constraint violation guarantees without allowing error cancellations. Unlike the existing work, which focuses on either tabular CMDP or model-based reinf…

Cited by 0SourceScholar
2025

Enhancing Safety in Reinforcement Learning with Human Feedback via Rectified Policy Optimization

NeurIPS 2025poster

Balancing helpfulness and safety (harmlessness) is a critical challenge in aligning large language models (LLMs). Current approaches often decouple these two objectives, training separate preference models for helpfulness and safety, while framing safety as a constraint within a constrained Markov D…

Cited by 0SourcecodeScholar
2024

Adversarially Trained Weighted Actor-Critic for Safe Offline Reinforcement Learning

NeurIPS 2024poster

We propose WSAC (Weighted Safe Actor-Critic), a novel algorithm for Safe Offline Reinforcement Learning (RL) under functional approximation, which can robustly optimize policies to improve upon an arbitrary reference policy with limited data coverage. WSAC is designed as a two-player Stackelberg gam…

Cited by 0SourcePDFScholar
2024

Safe and Efficient: A Primal-Dual Method for Offline Convex CMDPs under Partial Data Coverage

NeurIPS 2024poster

Offline safe reinforcement learning (RL) aims to find an optimal policy using a pre-collected dataset when data collection is impractical or risky. We propose a novel linear programming (LP) based primal-dual algorithm for convex MDPs that incorporates ``uncertainty'' parameters to improve data effi…

Cited by 0SourcePDFScholar