2026
ARGUS: Defending Against Multimodal Indirect Prompt Injection via Steering Instruction-Following Behavior
CVPR 2026
Multimodal Large Language Models (MLLMs) are increasingly vulnerable to multimodal Indirect Prompt Injection (IPI) attacks, which embed malicious instructions in images, videos, or audio to hijack model behavior. Existing defenses, designed primarily for text-only LLMs, are unsuitable for countering