2026
Response Attack: Exploiting Contextual Priming to Jailbreak Large Language Models
AAAI 2026technical
Contextual priming, where earlier stimuli covertly bias later judgments, offers an unexplored attack surface for large language models (LLMs). We uncover a contextual priming vulnerability in which the previous response in the dialogue can steer its subsequent behavior toward policy-violating conten