← Search

Lluis Gomez

11 accepted papers

2023

Show, Interpret and Tell: Entity-Aware Contextualised Image Captioning in Wikipedia

AAAI 2023technical

Humans exploit prior knowledge to describe images, and are able to adapt their explanation to specific contextual information given, even to the extent of inventing plausible explanations when contextual information and images do not match. In this work, we propose the novel task of captioning Wikip…

2023

Text-DIAE: A Self-Supervised Degradation Invariant Autoencoder for Text Recognition and Document Enhancement

AAAI 2023technical

In this paper, we propose a Text-Degradation Invariant Auto Encoder (Text-DIAE), a self-supervised model designed to tackle two tasks, text recognition (handwritten or scene-text) and document image enhancement. We start by employing a transformer-based architecture that incorporates three pretext…

2020

RoadText-1K: Text Detection & Recognition Dataset for Driving Videos

ICRA 2020poster

Perceiving text is crucial to understand semantics of outdoor scenes and hence is a critical requirement to build intelligent systems for driver assistance and self-driving. Most of the existing datasets for text detection and recognition comprise still images and are mostly compiled keeping text in…

Cited by 62SourceScholar
2019

Good News, Everyone! Context Driven Entity-Aware Captioning for News Images

CVPR 2019poster

Current image captioning systems perform at a merely descriptive level, essentially enumerating the objects in the scene and their relations. Humans, on the contrary, interpret images by integrating several sources of prior knowledge of the world. In this work, we aim to take a step closer to produc…

Cited by 191PDFcodeScholar
2019

Scene Text Visual Question Answering

ICCV 2019poster

Current visual question answering datasets do not consider the rich semantic information conveyed by text within an image. In this work, we present a new dataset, ST-VQA, that aims to highlight the importance of exploiting high-level semantic information present in images as textual cues in the Visu…

Cited by 417PDFScholar
2017

Self-Supervised Learning of Visual Features Through Embedding Images Into Text Topic Spaces

CVPR 2017poster

End-to-end training from scratch of current deep architectures for new computer vision problems would require Imagenet-scale datasets, and this is not always possible. In this paper we present a method that is able to take advantage of freely available multi-modal content to train computer vision al…

Cited by 143PDFScholar