← Search

Xiumin Wang

1 accepted papers

2026

Surgery: Mitigating Harmful Fine-Tuning for Large Language Models via Attention Sink

ICML 2026spotlight

Harmful fine-tuning can invalidate safety alignment of large language models, exposing significant safety risks. In this paper, we utilize the attention sink mechanism to mitigate harmful fine-tuning. Specifically, we first measure a statistic named *sink divergence* for each attention head and obse…

Cited by 0SourceScholar