ICASSP 2025accepted0 citations

Vision-text Enhancement Network For Weakly Supervised Video Anomaly Detection

Yiheng Chen, Shuai Fu, Niantong Qin, Xinning Du, Jianping Ren, Shuhua Liu

Abstract

The recent vision-language pre-training model ImageBind has shown significant success in a wide range of visual tasks, demonstrating excellent ability of joint embedding space across different modalities in visual or textual representation. A worthwhile problem is utilizing such a strong model for weakly supervised video anomaly detection (WSVAD). Most previous works only use the single visual modality, and define anomaly detection as a simple video classification task. However, such solutions ignore the textual information in the dataset and the anomaly event localization. To address these issues, this paper proposes the vision-text enhancement network (VTENet). For text features, it adopts the frozen ImageBind model directly without any fine-tuning process. For video features, it enhances visual feature representation through the proposed temporal enhancement graph convolution module (TGC). VTENet makes full use of associations between vision and text, contains two branches to complete coarse-grained and fine-grained video anomaly detection. One branch utilizes visual features for coarse-grained binary classification, while the other compares the text features with the entire video features across the dataset to complete fine-grained video anomaly detection. Extensive experiments on XD-Violence and UCF-Crime demonstrate the superiority of the proposed approach in both coarse-grained and fine-grained tasks.

BibTeX
@inproceedings{icassp2025_visiontextenhanc,
  title = {Vision-text Enhancement Network For Weakly Supervised Video Anomaly Detection},
  author = {Yiheng Chen and Shuai Fu and Niantong Qin and Xinning Du and Jianping Ren and Shuhua Liu},
  booktitle = {ICASSP 2025},
  year = {2025}
}
Vision-text Enhancement Network For Weakly Supervised Video Anomaly Detection · ICASSP 2025