ICASSP 2025accepted0 citations

Efficient Visual Storytelling through Descriptive Words Distillation and Dynamic Decoding

Zixuan Xiong, Hai Lin, Lichen Bai, Yinghui Li, Hai-Tao Zheng, Hong-Gee Kim

Abstract

Visual storytelling, a complex task in natural language generation, aims to create coherent and engaging narratives from a sequence of images, requiring more intricate and lengthy descriptions than typical image captioning. Current methods generally employ sophisticated modal interaction modules and require substantial additional data for training. This paper introduces two innovative techniques to advance visual storytelling capabilities. First, we propose a method to distill descriptive word embeddings from model parameters as an extension to the model’s pre-trained embeddings. This approach focuses the training process on these descriptive words, allowing for efficient fine-tuning and eliminates the need for additional visual modules and reduces computational overhead. Second, we develop a dynamic decoding strategy that adapts the text generation process based on the characteristic of the image features. This ensures a improved alignment with image features and the generated stories. Our experiments on the VIST dataset demonstrate that these methods achieve an impressive performance. Ablation studies show significant improvements in training efficiency and narrative quality compared to existing approaches.

BibTeX
@inproceedings{icassp2025_efficientvisuals,
  title = {Efficient Visual Storytelling through Descriptive Words Distillation and Dynamic Decoding},
  author = {Zixuan Xiong and Hai Lin and Lichen Bai and Yinghui Li and Hai-Tao Zheng and Hong-Gee Kim},
  booktitle = {ICASSP 2025},
  year = {2025}
}