← Search

Yuanshu Zhao

1 accepted papers

2025

Gamma-Guard: Lightweight Residual Adapters for Robust Guardrails in Large Language Models

EMNLP 2025

Large language models (LLMs) are widely deployed as zero-shot evaluators for answer grading, content moderation, and document ranking. Yet studies show that guard models (Guards)—LLMs fine-tuned for safety—remain vulnerable to “jailbreak” attacks, jeopardising downstream chatbots.We confirm this wea