Unbiased Multimodal Audio-to-Intent Recognition
Qian Dong, Yuezhou Dong, Ke Qin, Guiduo Duan, Tao He
Abstract
Audio-to-intent recognition is a critical task focused on identifying a speaker’s intent from spoken language. Recently, multimodal audio-to-intent approaches have emerged as the predominant strategy for enhancing audio-based intent recognition. In this study, we conduct a series of empirical experiments that reveal a significant modality bias in current multimodal audio-to-intent recognition methods. Specifically, these methods disproportionately rely on the textual modality to determine intent, often neglecting the audio data. To address this issue, we propose a context-enhanced contrastive learning framework designed to capture rich regional- and global-audio context information, thereby enabling more balanced audio-to-intent recognition. Additionally, we introduce a prototype-based intent classification strategy that encourages different intent classes and modalities to converge toward unified prototypes, leading to smoother classification boundaries as opposed to the traditionally skewed boundaries. Extensive experiments demonstrate that our approach effectively mitigates modality bias, e.g., a performance improvement of 2.12% in intent classification compared to the state-of-the-art method GZAIR on the dataset MintRec.
BibTeX
@inproceedings{icassp2025_unbiasedmultimod,
title = {Unbiased Multimodal Audio-to-Intent Recognition},
author = {Qian Dong and Yuezhou Dong and Ke Qin and Guiduo Duan and Tao He},
booktitle = {ICASSP 2025},
year = {2025}
}