← Search

Heqing Zou

9 accepted papers

2024

Cross-Modality and Within-Modality Regularization for Audio-Visual Deepfake Detection

ICASSP 2024accepted

Audio-visual deepfake detection scrutinizes manipulations in public video using complementary multimodal cues. Current methods, which train on fused multimodal data for multimodal targets face challenges due to uncertainties and inconsistencies in learned representations caused by independent modali…

Cited by 0SourceScholar
2023

Cross-Modal Global Interaction and Local Alignment for Audio-Visual Speech Recognition

IJCAI 2023poster

Audio-visual speech recognition (AVSR) research has gained a great success recently by improving the noise-robustness of audio-only automatic speech recognition (ASR) with noise-invariant visual information. However, most existing AVSR approaches simply fuse the audio and visual features by concaten…

2023

Leveraging Modality-Specific Representations for Audio-Visual Speech Recognition via Reinforcement Learning

AAAI 2023technical

Audio-visual speech recognition (AVSR) has gained remarkable success for ameliorating the noise-robustness of speech recognition. Mainstream methods focus on fusing audio and visual inputs to obtain modality-invariant representations. However, such representations are prone to over-reliance on audio…

Cited by 31SourcePDFScholar
2023

MIR-GAN: Refining Frame-Level Modality-Invariant Representations with Adversarial Network for Audio-Visual Speech Recognition

ACL 2023long

Audio-visual speech recognition (AVSR) attracts a surge of research interest recently by leveraging multimodal signals to understand human speech. Mainstream approaches addressing this task have developed sophisticated architectures and techniques for multi-modality fusion and representation learnin…

2023

UniS-MMC: Multimodal Classification via Unimodality-supervised Multimodal Contrastive Learning

ACL 2023findings

Multimodal learning aims to imitate human beings to acquire complementary information from multiple modalities for various downstream tasks. However, traditional aggregation-based multimodal fusion methods ignore the inter-modality relationship, treat each modality equally, suffer sensor noise, and…

2023

Unifying Speech Enhancement and Separation with Gradient Modulation for End-to-End Noise-Robust Speech Separation

ICASSP 2023accepted

Recent studies in neural network-based monaural speech separation (SS) have achieved a remarkable success thanks to increasing ability of long sequence modeling. However, they would degrade significantly when put under realistic noisy conditions, as the background noise could be mistaken for speaker…

Cited by 0SourceScholar
2022

Self-Critical Sequence Training for Automatic Speech Recognition

ICASSP 2022accepted

Although automatic speech recognition (ASR) task has gained remarkable success by sequence-to-sequence models, there are two main mismatches between its training and testing that might lead to performance degradation: 1) The typically used cross-entropy criterion aims to maximize log-likelihood of t…

Cited by 0SourceScholar
2022

Speech Emotion Recognition with Co-Attention Based Multi-Level Acoustic Information

ICASSP 2022accepted

Speech Emotion Recognition (SER) aims to help the machine to understand human’s subjective emotion from only audio in-formation. However, extracting and utilizing comprehensive in-depth audio information is still a challenging task. In this paper, we propose an end-to-end speech emotion recognition…

Cited by 0SourceScholar