2026
Detecting Fluent Optimization Based Adversarial Prompts via Sequential Entropy Changes
ICML 2026poster
Optimization-based adversarial suffixes can jailbreak aligned large language models (LLMs) while remaining fluent, weakening detectors based on static global or windowed perplexity statistics. We cast adversarial suffix detection as an \emph{online change-point detection} problem over the token-leve…