2025
Root Defense Strategies: Ensuring Safety of LLM at the Decoding Level
ACL 2025long
Large language models (LLMs) have demonstrated immense utility across various industries. However, as LLMs advance, the risk of harmful outputs increases due to incorrect or malicious prompts. While current methods effectively address jailbreak risks, they share common limitations: 1) Judging harmfu…