2024
MoGU: A Framework for Enhancing Safety of LLMs While Preserving Their Usability
NeurIPS 2024poster
Large Language Models (LLMs) are increasingly deployed in various applications. As their usage grows, concerns regarding their safety are rising, especially in maintaining harmless responses when faced with malicious instructions. Many defense strategies have been developed to enhance the safety of…