2025
How Does DPO Reduce Toxicity? A Mechanistic Neuron-Level Analysis
EMNLP 2025
Safety fine-tuning algorithms reduce harmful outputs in language models, yet their mechanisms remain under-explored. Direct Preference Optimization (DPO) is a popular choice of algorithm, but prior explanations—attributing its effects solely to dampened toxic neurons in the MLP layers—are incomplete