Show, Interpret and Tell: Entity-Aware Contextualised Image Captioning in Wikipedia
Humans exploit prior knowledge to describe images, and are able to adapt their explanation to specific contextual information given, even to the extent of inventing plausible explanations when contextual information and images do not match. In this work, we propose the novel task of captioning Wikip…