ICASSP 2024accepted0 citations

Think as People: Context-Driven Multi-Image News Captioning with Adaptive Dual Attention

Qiang Yang, Xiaodong Wu, Xiuying Chen, Xin Gao, Xiangliang Zhang

Abstract

Automatic image captioning has been extensively studied, however, existing methods primarily focus on a single image. Actually, the demand for captioning multiple images and corresponding contextual information has been growing in diverse scenarios, e.g., composing news articles headlines, and electronic medical reports. In this paper, we propose a novel COntext-driven captioning approach for Multi-Image News, called COMIN, which employs a two-step attention mechanism, called adaptive dual attention, comprising global attention for grasping overall context and local attention for finer image details. It is inspired by the observation and cognitive processes of human beings where global attention and local attention are responsible for understanding the high-level features and detailing the low-level features. Experimental results on our newly contributed Star-News dataset show that our proposed model outperforms the state-of-the-art image captioning methods in multi-image captioning scenarios.

BibTeX
@inproceedings{icassp2024_thinkaspeoplecon,
  title = {Think as People: Context-Driven Multi-Image News Captioning with Adaptive Dual Attention},
  author = {Qiang Yang and Xiaodong Wu and Xiuying Chen and Xin Gao and Xiangliang Zhang},
  booktitle = {ICASSP 2024},
  year = {2024}
}
Think as People: Context-Driven Multi-Image News Captioning with Adaptive Dual Attention · ICASSP 2024