Transferable Attacks on Open-Vocabulary Video Instance Segmentation via Dual-Objective Triggers
Minghao Shou, Kesen Wang, Tong Zhang, Han Bao, Zonghui Wang
Abstract
Open‑vocabulary video instance segmentation (OV‑VIS) couples spatial‑temporal reasoning with language grounding, yet its adversarial robustness has remained unexplored. We present the Dual-Objective Triggers (DOT), the first transferable attack on OV-VIS that simultaneously exploits the vision–language coupling and temporal coherence. DOT deploys a Dual Semantic Perturbation Module that overlays two complementary triggers: a Semantic Suppression Trigger erases the alignment between the true object and the query, while a Plausible Replacement Trigger steers the tracker toward a phantom trajectory that is visually plausible and text‑consistent. To amplify cross‑model transferability without sacrificing perceptual fidelity, we introduce Phase‑Guided Adversarial Training, which injects perturbations primarily in the phase spectrum while blending amplitudes with clean references. Extensive experiments on four state‑of‑the‑art OV‑VIS implementations demonstrate that DOT reduces mAP by up to 69.3\% and raises attack success rate by up to 98\%, outperforming the strongest baselines by a factor of 1.6$\times$ on average, while maintaining a PSNR of 52.49 dB, thus exposing critical security vulnerabilities and laying a foundation for future research on robust and trustworthy vision–language systems.
BibTeX
@inproceedings{ijcai2026_transferableatta,
title = {Transferable Attacks on Open-Vocabulary Video Instance Segmentation via Dual-Objective Triggers},
author = {Minghao Shou and Kesen Wang and Tong Zhang and Han Bao and Zonghui Wang},
booktitle = {IJCAI 2026},
year = {2026}
}