SafeInfer: Context Adaptive Decoding Time Safety Alignment for Large Language Models
Somnath Banerjee, Sayan Layek, Soham Tripathy, Shanu Kumar, Animesh Mukherjee, Rima Hazra
Abstract
Language models aligned for safety often exhibit fragile and imbalanced mechanisms, increasing the chances of producing unsafe content. In addition, editing techniques to incorporate new knowledge can further compromise safety. To tackle these issues, we propose SafeInfer, a context-adaptive, decoding-time safety alignment strategy for generating safe responses to user queries. safeInfer involves two phases: the 'safety amplification' phase, which uses safe demonstration examples to adjust the model’s hidden states and increase the likelihood of safer outputs, and the 'safety-guided decoding' phase, which influences token selection based on safety-optimized distributions to ensure the generated content adheres to ethical guidelines. Further, we introduce HarmEval, a novel benchmark for comprehensive safety evaluations, designed to address potential misuse scenarios in line with the policies of leading AI technology companies.
BibTeX
@article{Banerjee_Layek_Tripathy_Kumar_Mukherjee_Hazra_2025, title={SafeInfer: Context Adaptive Decoding Time Safety Alignment for Large Language Models}, volume={39}, url={https://ojs.aaai.org/index.php/AAAI/article/view/34927}, DOI={10.1609/aaai.v39i26.34927}, abstractNote={Language models aligned for safety often exhibit fragile and imbalanced mechanisms, increasing the chances of producing unsafe content. In addition, editing techniques to incorporate new knowledge can further compromise safety. To tackle these issues, we propose SafeInfer, a context-adaptive, decoding-time safety alignment strategy for generating safe responses to user queries.
safeInfer involves two phases: the ’safety amplification’ phase, which uses safe demonstration examples to adjust the model’s hidden states and increase the likelihood of safer outputs, and the ’safety-guided decoding’ phase, which influences token selection based on safety-optimized distributions to ensure the generated content adheres to ethical guidelines. Further, we introduce HarmEval, a novel benchmark for comprehensive safety evaluations, designed to address potential misuse scenarios in line with the policies of leading AI technology companies.}, number={26}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, author={Banerjee, Somnath and Layek, Sayan and Tripathy, Soham and Kumar, Shanu and Mukherjee, Animesh and Hazra, Rima}, year={2025}, month={Apr.}, pages={27188-27196} }