2025
Model Editing as a Robust and Denoised variant of DPO: A Case Study on Toxicity
ICLR 2025poster
Recent alignment algorithms such as direct preference optimization (DPO) have been developed to improve the safety of large language models (LLMs) by training these models to match human behaviors exemplified by preference data. However, these methods are both computationally intensive and lacking…