2023
Positive-Augmented Contrastive Learning for Image and Video Captioning Evaluation
CVPR 2023highlight
The CLIP model has been recently proven to be very effective for a variety of cross-modal tasks, including the evaluation of captions generated from vision-and-language architectures. In this paper, we propose a new recipe for a contrastive-based evaluation metric for image captioning, namely Positi…