2022
Text-Only Training for Image Captioning using Noise-Injected CLIP
EMNLP 2022finding
We consider the task of image-captioning using only the CLIP model and additional text data at training time and no additional captioned images. Our approach relies on the fact that CLIP is trained to make visual and textual embeddings similar. Therefore, we only need to learn how to translate CLIP…