← Search

Marcal Rusinol

4 accepted papers

2019

Good News, Everyone! Context Driven Entity-Aware Captioning for News Images

CVPR 2019poster

Current image captioning systems perform at a merely descriptive level, essentially enumerating the objects in the scene and their relations. Humans, on the contrary, interpret images by integrating several sources of prior knowledge of the world. In this work, we aim to take a step closer to produc…

Cited by 191PDFcodeScholar
2019

Scene Text Visual Question Answering

ICCV 2019poster

Current visual question answering datasets do not consider the rich semantic information conveyed by text within an image. In this work, we present a new dataset, ST-VQA, that aims to highlight the importance of exploiting high-level semantic information present in images as textual cues in the Visu…

Cited by 417PDFScholar
2017

Self-Supervised Learning of Visual Features Through Embedding Images Into Text Topic Spaces

CVPR 2017poster

End-to-end training from scratch of current deep architectures for new computer vision problems would require Imagenet-scale datasets, and this is not always possible. In this paper we present a method that is able to take advantage of freely available multi-modal content to train computer vision al…

Cited by 143PDFScholar