2026
Safety at One Shot: Patching Fine-Tuned LLMs with A Single Instance
ICLR 2026poster
Fine-tuning safety-aligned large language models (LLMs) can substantially compromise their safety. Previous approaches require many safety samples or calibration sets, which not only incur significant computational overhead during realignment but also lead to noticeable degradation in model utility.…