RefineVAD: Semantic-Guided Feature Recalibration for Weakly Supervised Video Anomaly Detection
Junhee Lee, ChaeBeen Bang, MyoungChul Kim, MyeongAh Cho
Abstract
Weakly-Supervised Video Anomaly Detection aims to identify anomalous events using only video-level labels, balancing annotation efficiency with practical applicability. However, existing methods often oversimplify the anomaly space by treating all abnormal events as a single category, overlooking the diverse semantic and temporal characteristics intrinsic to real-world anomalies. Inspired by how humans perceive anomalies, by jointly interpreting temporal motion patterns and semantic structures underlying different anomaly types, we propose RefineVAD, a novel framework that mimics this dual-process reasoning. Our framework integrates two core modules. The first, Motion-aware Temporal Attention and Recalibration (MoTAR), estimates motion salience and dynamically adjusts temporal focus via shift-based attention and global Transformer-based modeling. The second, Category-Oriented Refinement (CORE), injects soft anomaly category priors into the representation space by aligning segment-level features with learnable category prototypes through cross-attention. By jointly leveraging temporal dynamics and semantic structure, explicitly models both ``how
BibTeX
@inproceedings{aaai2026_refinevadsemanti,
title = {RefineVAD: Semantic-Guided Feature Recalibration for Weakly Supervised Video Anomaly Detection},
author = {Junhee Lee and ChaeBeen Bang and MyoungChul Kim and MyeongAh Cho},
booktitle = {AAAI 2026},
year = {2026}
}