← Search

Enhua Song

1 accepted papers

2025

MncCap: Mining Neural Composition for Zero-shot Image Captioning via Text-only Training

ICASSP 2025accepted

Current text-only image captioning methods leverage the shared feature space of CLIP to train zero-shot image captioning using text data only, leaving feature associations and contextual understanding not fully explored. Neurological studies have revealed that the anterior temporal lobes of the brai…

Cited by 0SourceScholar