2025
Safeguard Fine-Tuned LLMs Through Pre- and Post-Tuning Model Merging
EMNLP 2025
Fine-tuning large language models (LLMs) for downstream tasks often leads to catastrophic forgetting, notably degrading the safety of originally aligned models. While some existing methods attempt to restore safety by incorporating additional safety data, the quality of such data typically falls sho