2025
LlmFixer: Fix the Helpfulness of Defensive Large Language Models
EMNLP 2025
Defense strategies of large language models besides alignment are introduced to defend against jailbreak attacks, and they have managed to decrease the success rate of jailbreak attacks. However, these defense strategies weakened the helpfulness of large language models. In this work, we propose a u