← Search

Sungwon Yi

3 accepted papers

2026

Dissect and Prune: Enhancing Robustness in AI-Generated Image Detection

ICML 2026poster

While existing AI-generated image detectors report high performance, we identify that this is largely driven by a critical *prediction asymmetry*: a bias toward the real class that severely limits sensitivity to generated content, especially under standard post-processing operations such as compress…

Cited by 0SourceScholar
2026

Explaining Jailbreaks: Structured and Interpretable Safety Assessment for Large Language Models

IJCAI 2026

Large Language Models (LLMs) remain highly vulnerable to jailbreak attacks, yet existing evaluations rely primarily on outcome-level metrics such as Attack Success Rate (ASR), providing limited insight into how and why safety failures occur. We propose an explanation-aware safety framework that augm

Cited by 0Scholar
2026

GraphShield: Graph-Theoretic Modeling of Network-Level Dynamics for Robust Jailbreak Detection

ICLR 2026poster

Large language models (LLMs) are increasingly deployed in real-world applications but remain highly vulnerable to jailbreak prompts that bypass safety guardrails and elicit harmful outputs. We propose GraphShield, a graph-theoretic jailbreak detector that models information routing inside the LLM as…

Cited by 0SourceScholar