Detecting Fluent Optimization Based Adversarial Prompts via Sequential Entropy Changes
Optimization-based adversarial suffixes can jailbreak aligned large language models (LLMs) while remaining fluent, weakening detectors based on static global or windowed perplexity statistics. We cast adversarial suffix detection as an \emph{online change-point detection} problem over the token-leve…