AAAI 2026technical0 citations

Tell as You Want: Customizing Image Narrative with Knowledge and Thoughts

Ziwei Yao, Qian Wang, Ruiping Wang, Xilin Chen

Abstract

With the advancement of vision-language models, image captioning has made significant progress, leading to the generation of more accurate and detailed descriptions. Current image captioning primarily focuses on describing the apparent visual characteristics, which are easily observed by most humans, but less helpful in real-world scenarios. When users seek a deeper understanding of visual content, they may be concerned with fine-grained categories, function properties, and other background knowledge, rather than merely appearances. Additionally, as users

BibTeX
@inproceedings{aaai2026_tellasyouwantcus,
  title = {Tell as You Want: Customizing Image Narrative with Knowledge and Thoughts},
  author = {Ziwei Yao and Qian Wang and Ruiping Wang and Xilin Chen},
  booktitle = {AAAI 2026},
  year = {2026}
}