← Search

Jiahui Zhu

4 accepted papers

2026

Keep the Best, Forget the Rest: Reliable Alignment with Order-Aware Preference Optimization

ICLR 2026poster

Direct Preference Optimization (DPO) has emerged as a powerful framework for aligning large language models (LLMs) with human preferences via pairwise comparisons. However, its performance is highly sensitive to the quality of training samples: when the reference policy is poorly aligned with human…

Cited by 0SourcecodeScholar
2025

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints

ICML 2025poster

Online safe reinforcement learning (RL) plays a key role in dynamic environments, with applications in autonomous driving, robotics, and cybersecurity. The objective is to learn optimal policies that maximize rewards while satisfying safety constraints modeled by constrained Markov decision processe…

Cited by 0SourcePDFScholar
2024

Deja vu: Contrastive Historical Modeling with Prefix-tuning for Temporal Knowledge Graph Reasoning

NAACL 2024findings

Temporal Knowledge Graph Reasoning (TKGR) is the task of inferring missing facts for incomplete TKGs in complex scenarios (e.g., transductive and inductive settings), which has been gaining increasing attention. Recently, to mitigate dependence on structured connections in TKGs, text-based methods h…

2024

Research of calibration method for fusion of LDS sensor and ToF low-cost sensor

IROS 2024poster

This paper proposes a method for calibrating the external parameters of the LDS sensor and ToF depth camera based on three cylinders. This method obtains the scanning data of the side surfaces of the three cylinders at different postures by changing the posture of the robot. For the single-line lase…

Cited by 0SourceScholar