← Search

Ali Furkan Biten

5 accepted papers

2023

Show, Interpret and Tell: Entity-Aware Contextualised Image Captioning in Wikipedia

AAAI 2023technical

Humans exploit prior knowledge to describe images, and are able to adapt their explanation to specific contextual information given, even to the extent of inventing plausible explanations when contextual information and images do not match. In this work, we propose the novel task of captioning Wikip…

2023

Text-DIAE: A Self-Supervised Degradation Invariant Autoencoder for Text Recognition and Document Enhancement

AAAI 2023technical

In this paper, we propose a Text-Degradation Invariant Auto Encoder (Text-DIAE), a self-supervised model designed to tackle two tasks, text recognition (handwritten or scene-text) and document image enhancement. We start by employing a transformer-based architecture that incorporates three pretext…

2022

LaTr: Layout-Aware Transformer for Scene-Text VQA

CVPR 2022oral

We propose a novel multimodal architecture for Scene Text Visual Question Answering (STVQA), named Layout-Aware Transformer (LaTr). The task of STVQA requires models to reason over different modalities. Thus, we first investigate the impact of each modality, and reveal the importance of the language…

Cited by 109PDFcodeScholar
2019

Good News, Everyone! Context Driven Entity-Aware Captioning for News Images

CVPR 2019poster

Current image captioning systems perform at a merely descriptive level, essentially enumerating the objects in the scene and their relations. Humans, on the contrary, interpret images by integrating several sources of prior knowledge of the world. In this work, we aim to take a step closer to produc…

Cited by 191PDFcodeScholar
2019

Scene Text Visual Question Answering

ICCV 2019poster

Current visual question answering datasets do not consider the rich semantic information conveyed by text within an image. In this work, we present a new dataset, ST-VQA, that aims to highlight the importance of exploiting high-level semantic information present in images as textual cues in the Visu…

Cited by 417PDFScholar