ICASSP 2025accepted0 citations

Reliable Learning From LLM Features for Multimodal Emotion and Intent Joint Understanding

Xiaolin Xu, Cheng Lu, Zhaoyang Li, Yuyun Liu, Yinghao Ma, Jiahao Luo, Yuan Zong, Wenming Zheng

Abstract

This paper describes a Reliable Learning Framework (RLF) for the 1st Multimodal Emotion and Intent Joint Understanding (MEIJU) Challenge at ICASSP 2025. Our proposed RLF includes a Hierarchical Interaction Network and a Reliable Fusion Strategy. The former can excavate emotion and intent cues from the high-level semantic features of multimodal data (video, audio, and text) generated by pretrained Large Language Models (LMMs), to enhance their representations, and the latter reliably integrates multiple predictions to further improve the robustness of emotion and intent understanding. Our RLF method achieved first place on Track 2 (Mandarin) of MEIJU, with performance scores for emotion, intent, and joint recognition reaching 0.7285, 0.7456, and 0.7370.

BibTeX
@inproceedings{icassp2025_reliablelearning,
  title = {Reliable Learning From LLM Features for Multimodal Emotion and Intent Joint Understanding},
  author = {Xiaolin Xu and Cheng Lu and Zhaoyang Li and Yuyun Liu and Yinghao Ma and Jiahao Luo and Yuan Zong and Wenming Zheng},
  booktitle = {ICASSP 2025},
  year = {2025}
}