← Search

Kangmin Bae

2 accepted papers

2026

Explaining Jailbreaks: Structured and Interpretable Safety Assessment for Large Language Models

IJCAI 2026

Large Language Models (LLMs) remain highly vulnerable to jailbreak attacks, yet existing evaluations rely primarily on outcome-level metrics such as Attack Success Rate (ASR), providing limited insight into how and why safety failures occur. We propose an explanation-aware safety framework that augm

Cited by 0Scholar
2026

GraphShield: Graph-Theoretic Modeling of Network-Level Dynamics for Robust Jailbreak Detection

ICLR 2026poster

Large language models (LLMs) are increasingly deployed in real-world applications but remain highly vulnerable to jailbreak prompts that bypass safety guardrails and elicit harmful outputs. We propose GraphShield, a graph-theoretic jailbreak detector that models information routing inside the LLM as…

Cited by 0SourceScholar