ICASSP 2023accepted0 citations

CM-CS: Cross-Modal Common-Specific Feature Learning For Audio-Visual Video Parsing

Hongbo Chen, Dongchen Zhu, Guanghui Zhang, Wenjun Shi, Xiaolin Zhang, Jiamao Li

Abstract

The weakly-supervised audio-visual video parsing (AVVP) task aims to parse duration and categories of each snippet when only the video-level event labels are provided. Most methods either leverage attention mechanisms to explore cross-modal and cross-video event semantics or alleviate label noise to improve performance. However, the distributional modality discrepancy caused by the heterogeneity of signals remains a significant challenge. To this end, we propose a novel cross-modal common-specific feature learning method (cm-CS) to map the modal features into modality-common and modality-specific subspaces. The former aims to capture similar high-level scene cue across different modalities, while the later attempts to capture specific cue. The proposed method is applied among and across in-visual 2D-3D modalities, audio-visual modalities, respectively. In addition, we design a training strategy to strengthen the learning of similarity and differences across modalities. Experiments show a large improvement of our method against existing works on the Look, Listen, and Parse (LLP) dataset (e.g. from 58.9% to 62.9% in video-level visual metric).

BibTeX
@inproceedings{icassp2023_cmcscrossmodalco,
  title = {CM-CS: Cross-Modal Common-Specific Feature Learning For Audio-Visual Video Parsing},
  author = {Hongbo Chen and Dongchen Zhu and Guanghui Zhang and Wenjun Shi and Xiaolin Zhang and Jiamao Li},
  booktitle = {ICASSP 2023},
  year = {2023}
}