2020
More Grounded Image Captioning by Distilling Image-Text Matching Model
CVPR 2020poster
Visual attention not only improves the performance of image captioners, but also serves as a visual interpretation to qualitatively measure the caption rationality and model transparency. Specifically, we expect that a captioner can fix its attentive gaze on the correct objects while generating the…