← Search

Soham Tripathy

1 accepted papers

2025

SafeInfer: Context Adaptive Decoding Time Safety Alignment for Large Language Models

AAAI 2025technical

Language models aligned for safety often exhibit fragile and imbalanced mechanisms, increasing the chances of producing unsafe content. In addition, editing techniques to incorporate new knowledge can further compromise safety. To tackle these issues, we propose SafeInfer, a context-adaptive, decodi…