← Search

Yanxiang Chen

3 accepted papers

2025

Multimodal Class-aware Semantic Enhancement Network for Audio-Visual Video Parsing

AAAI 2025technical

The Audio-Visual Video Parsing task aims to recognize and temporally localize all events occurring in either the audio or visual stream, or both. Capturing accurate event semantics for each audio/visual segment is vital. Prior works directly utilize the extracted holistic audio and visual features f…

2025

Text-Infused Audio-Visual Video Parsing with Semantic-Aware Multimodal Contrastive Learning

ICASSP 2025accepted

The Audio-Visual Video Parsing task aims to recognize events occurring in video segments for each modality. Presently, the excellent performance in handling video parsing is shown by generating pseudo labels at the segment level. However, these approaches still suffer from adequate semantic learning…

Cited by 0SourceScholar
2020

Spectrogram Analysis Via Self-Attention for Realizing Cross-Model Visual-Audio Generation

ICASSP 2020accepted

Human cognition is supported by the combination of multimodal information from different sources of perception. The two most important modalities are visual and audio. Cross-modal visual-audio generation enables the synthesis of data from one modality following the acquisition of data from another.…

Cited by 0SourceScholar