← Search

Yiting He

3 accepted papers

2025

Model Editing as a Robust and Denoised variant of DPO: A Case Study on Toxicity

ICLR 2025poster

Recent alignment algorithms such as direct preference optimization (DPO) have been developed to improve the safety of large language models (LLMs) by training these models to match human behaviors exemplified by preference data. However, these methods are both computationally intensive and lacking…

2025

Sample Complexity of Distributionally Robust Off-Dynamics Reinforcement Learning with Online Interaction

ICML 2025poster

Off-dynamics reinforcement learning (RL), where training and deployment transition dynamics are different, can be formulated as learning in a robust Markov decision process (RMDP) where uncertainties in transition dynamics are imposed. Existing literature mostly assumes access to generative models a…

Cited by 0SourcePDFScholar