← Search

Guorui Chen

2 accepted papers

2025

LLM Jailbreak Detection for (Almost) Free!

EMNLP 2025

Large language models (LLMs) enhance security through alignment when widely used, but remain susceptible to jailbreak attacks capable of producing inappropriate content. Jailbreak detection methods show promise in mitigating jailbreak attacks through the assistance of other models or multiple model

2025

Reimagining Safety Alignment with An Image

EMNLP 2025

Large language models (LLMs) excel in diverse applications but face dual challenges: generating harmful content under jailbreak attacks and over-refusing benign queries due to rigid safety mechanisms. These issues severely affect the application of LLMs, especially in the medical and education field