Con Instruction: Universal Jailbreaking of Multimodal Large Language Models via Non-Textual Modalities
Existing attacks against multimodal language models often communicate instruction through text, either as an explicit malicious instruction or a crafted generic prompt, and accompanied by a toxic image. In contrast, here we exploit the capabilities of MLLMs in following non-textual instruction, i.e.…