AAAI 2026technical0 citations
Tell as You Want: Customizing Image Narrative with Knowledge and Thoughts
Ziwei Yao, Qian Wang, Ruiping Wang, Xilin Chen
Abstract
With the advancement of vision-language models, image captioning has made significant progress, leading to the generation of more accurate and detailed descriptions. Current image captioning primarily focuses on describing the apparent visual characteristics, which are easily observed by most humans, but less helpful in real-world scenarios. When users seek a deeper understanding of visual content, they may be concerned with fine-grained categories, function properties, and other background knowledge, rather than merely appearances. Additionally, as users
BibTeX
@inproceedings{aaai2026_tellasyouwantcus,
title = {Tell as You Want: Customizing Image Narrative with Knowledge and Thoughts},
author = {Ziwei Yao and Qian Wang and Ruiping Wang and Xilin Chen},
booktitle = {AAAI 2026},
year = {2026}
}