← Search

Matteo Stefanini

1 accepted papers

2020

Meshed-Memory Transformer for Image Captioning

CVPR 2020poster

Transformer-based architectures represent the state of the art in sequence modeling tasks like machine translation and language understanding. Their applicability to multi-modal contexts like image captioning, however, is still largely under-explored. With the aim of filling this gap, we present M2…

Cited by 1303PDFcodeScholar