AAAI 2026technical0 citations

Making Visual Dialogue More Engaging: A New Task, Method, and Metric

Guanghui Ye, Huan Zhao, Yingxue Gao, Zhixue Zhao, Kehan Wang, Xupeng Zha, Zhihua Jiang

Abstract

Large language model (LLM)-based visual dialogue (VD) systems have made response generation for image-grounded conversations more correct and coherent. However, user engagement - the extent to which a user is interested, emotionally involved, and willing to continue the conversation - remains a challenge. To fully explore engaging VD, we propose: (i) a new task named Audio-enhanced VD (AVD), which introduces additional audio dialogue contexts that can more vividly convey the speaker

BibTeX
@inproceedings{aaai2026_makingvisualdial,
  title = {Making Visual Dialogue More Engaging: A New Task, Method, and Metric},
  author = {Guanghui Ye and Huan Zhao and Yingxue Gao and Zhixue Zhao and Kehan Wang and Xupeng Zha and Zhihua Jiang},
  booktitle = {AAAI 2026},
  year = {2026}
}