MncCap: Mining Neural Composition for Zero-shot Image Captioning via Text-only Training
Tongtong Liu, Chen Yang, Guoqiang Chen, Qinxu Gao, Enhua Song, Wenhui Li
Abstract
Current text-only image captioning methods leverage the shared feature space of CLIP to train zero-shot image captioning using text data only, leaving feature associations and contextual understanding not fully explored. Neurological studies have revealed that the anterior temporal lobes of the brain are responsible for binding attributes to specific individuals and the corresponding collective connections. Inspired by the above studies, we propose a novel Mining Neural Composition for zero-shot image captioning (MncCap) via text-only training to model the neural composition. During training, we combine the global and local fine-grained features provided by the text clues to achieve a stronger ability of contextual understanding. To express the relationship from discriminative information in the text, we propose a strategy of converting each candidate sentence into a text-tree. During inference, a pre-trained detector is used to obtain the ROI features in the image to improve the contextual integrity of the semantic features. Experimental results conducted on three image captioning benchmark datasets show that our framework achieves remarkable performance improvements.
BibTeX
@inproceedings{icassp2025_mnccapminingneur,
title = {MncCap: Mining Neural Composition for Zero-shot Image Captioning via Text-only Training},
author = {Tongtong Liu and Chen Yang and Guoqiang Chen and Qinxu Gao and Enhua Song and Wenhui Li},
booktitle = {ICASSP 2025},
year = {2025}
}