ICASSP 2025accepted0 citations

KMG-LL: Knowledge-enhanced Multimodal Graph for Dialogue Generation

Yuezhou Dong, Tao He, Qian Dong, Ke Qin

Abstract

Multimodal dialogue generation requires a comprehensive understanding of each modality and the ability to effectively integrate these modalities to produce contextually relevant and diverse responses. However, existing approaches, which primarily rely on language models, often suffer from modality bias, stemming from an over-dependence on language data in training, and inherent biases in model structure. This over-reliance results in an imbalance across modalities and leads to suboptimal performance. To address this challenge, we propose a novel dialogue generation model with Knowledge-enhanced Multimodal Graph, dubbed KMG-LL, designed to balance multimodal information and incorporate external knowledge. Specifically, we transform multimodal information into graph-structured data and integrate them with commonsense knowledge to construct knowledge-enhanced multimodal graph. To further refine the extraction of multimodal information, we propose a modal-balanced knowledge aggregation that processes modalities cross multiple levels. Extensive experiments on two multimodal dialogue datasets demonstrate that KMG-LL significantly outperforms existing baselines in multimodal dialogue generation.

BibTeX
@inproceedings{icassp2025_kmgllknowledgeen,
  title = {KMG-LL: Knowledge-enhanced Multimodal Graph for Dialogue Generation},
  author = {Yuezhou Dong and Tao He and Qian Dong and Ke Qin},
  booktitle = {ICASSP 2025},
  year = {2025}
}