2026
Surgery: Mitigating Harmful Fine-Tuning for Large Language Models via Attention Sink
ICML 2026spotlight
Harmful fine-tuning can invalidate safety alignment of large language models, exposing significant safety risks. In this paper, we utilize the attention sink mechanism to mitigate harmful fine-tuning. Specifically, we first measure a statistic named *sink divergence* for each attention head and obse…