2026
Transferability of Adversarial Attacks in Video-based MLLMs: A Cross-modal Image-to-Video Approach
AAAI 2026technical
Video-based multimodal large language models (V-MLLMs) have shown vulnerability to adversarial examples in video-text multimodal tasks. However, the transferability of adversarial videos to unseen models—a common and practical real-world scenario—remains unexplored. In this paper, we pioneer an in