2025
Gamma-Guard: Lightweight Residual Adapters for Robust Guardrails in Large Language Models
EMNLP 2025
Large language models (LLMs) are widely deployed as zero-shot evaluators for answer grading, content moderation, and document ranking. Yet studies show that guard models (Guards)—LLMs fine-tuned for safety—remain vulnerable to “jailbreak” attacks, jeopardising downstream chatbots.We confirm this wea