← Search

Manuele Barraco

2 accepted papers

2023

Positive-Augmented Contrastive Learning for Image and Video Captioning Evaluation

CVPR 2023highlight

The CLIP model has been recently proven to be very effective for a variety of cross-modal tasks, including the evaluation of captions generated from vision-and-language architectures. In this paper, we propose a new recipe for a contrastive-based evaluation metric for image captioning, namely Positi…

2023

With a Little Help from Your Own Past: Prototypical Memory Networks for Image Captioning

ICCV 2023poster

Image captioning, like many tasks involving vision and language, currently relies on Transformer-based architectures for extracting the semantics in an image and translating it into linguistically coherent descriptions. Although successful, the attention operator only considers a weighted summation…

Cited by 23PDFcodeScholar