ICASSP 2025accepted0 citations

Interactive and Balanced Multimodal Learning via Cross Attention and Gradient Modulation for Compressed Video Action Recognition

Xinqi Li, Shaojie Li, Ming Ma

Abstract

Compressed video action recognition is a crucial task in video processing. Compared with traditional methods, it directly processes RGB (I-frames) and motion (motion vectors and residuals) modalities, which effectively alleviates computational burdens. However, this task suffers from dynamic noise and insufficient interaction between modalities. To address these issues, we propose the Cross-Modal Fusion Modulation Network (CFM-Net), which consists of an RGB stream and a motion stream. The motion stream incorporates a Pseudo Optical Flow Generator (POFG) that reduces motion noise through adversarial learning and optical flow supervision. The interaction between streams is strengthened by the Lightweight Cross-Modal Fusion (LCF) block and the Adaptive Dynamic Gradient Modulation (ADGM) strategy. The LCF facilitates the fusion of RGB and motion modalities through cross-attention, while the ADGM enhances their interaction by balancing multimodal optimization. We evaluate the performance of CFM-Net on the UCF-101 and HMDB-51 benchmarks, demonstrating the effectiveness and accuracy of our approach.

BibTeX
@inproceedings{icassp2025_interactiveandba,
  title = {Interactive and Balanced Multimodal Learning via Cross Attention and Gradient Modulation for Compressed Video Action Recognition},
  author = {Xinqi Li and Shaojie Li and Ming Ma},
  booktitle = {ICASSP 2025},
  year = {2025}
}
Interactive and Balanced Multimodal Learning via Cross Attention and Gradient Modulation for Compressed Video Action Recognition · ICASSP 2025