IJCAI 20260 citations

Towards Comprehensive Post Safety Alignment of Large Language Models via Safety Patching

Weixiang Zhao, Yulin Hu, Yang Deng, Jiahe Guo, Xingyu Sui, Yanyan Zhao, Bing Qin, Ting Liu

Abstract

Safety alignment of large language models (LLMs) has been gaining increasing attention. However, current safety-aligned LLMs suffer from the fragile and imbalanced safety mechanisms, which can still be induced to generate unsafe responses, exhibit over-safety by rejecting safe user inputs, and fail to preserve general utility after safety alignment. To this end, we propose a novel post safety alignment (PSA) method to address these inherent and emerging safety challenges, including safety enhancement, over-safety mitigation, and utility preservation. In specific, we introduce SafePatching, a novel framework for comprehensive PSA, where two distinct safety patches are developed on the harmful data to enhance safety and mitigate over-safety concerns, and then seamlessly integrated into the target LLM backbone without compromising its utility. Extensive experiments on four representative aligned LLMs, including LLaMA-2/3, Gemma and Mistral, show that SafePatching achieves a more comprehensive PSA than baseline methods, further optimizing the balance between being helpful and harmless in current aligned LLMs. Also, SafePatching demonstrates its superiority in continual PSA scenarios.

AI Ethics, Trust, Fairnes: Safety and robustnessNatural Language Processing: Language models
BibTeX
@inproceedings{ijcai2026_towardscomprehen,
  title = {Towards Comprehensive Post Safety Alignment of Large Language Models via Safety Patching},
  author = {Weixiang Zhao and Yulin Hu and Yang Deng and Jiahe Guo and Xingyu Sui and Yanyan Zhao and Bing Qin and Ting Liu},
  booktitle = {IJCAI 2026},
  year = {2026}
}