2026
Furina: Fragmented Uncertainty-Driven Refusal Instability Attack
ICML 2026poster
Safety alignment in large language models (LLMs) and multimodal large language models (MLLMs) is commonly assumed to operate as a near-binary threshold mechanism. We challenge this assumption by revealing that safety behavior is governed by an \emph{instability region} where small perturbations indu…