← Search

Dimosthenis Karatzas

13 accepted papers

2026

Strategic Navigation or Stochastic Search? How Agents and Humans Reason Over Document Collections

ICML 2026oral

Multimodal agents offer a compelling path to automating complex document-intensive workflows, yet a critical question remains: do these architectures demonstrate genuine strategic reasoning, or simply conduct stochastic trial-and-error search? To address this, we introduce Agentic Document VQA, a be…

Cited by 0SourceScholar
2026

TRIM: A SELF-SUPERVISED VIDEO SUMMARIZATION FRAMEWORK MAXIMIZING TEMPORAL RELATIVE INFORMATION AND REPRESENTATIVENESS

ICASSP 2026poster

The increasing ubiquity of video content and the corresponding demand for efficient access to meaningful information have elevated video summarization and video highlights as a vital research area. However, many state-of-the-art methods depend heavily either on supervised annotations or on attention…

Cited by 0SourcePDFScholar
2025

DocMIA: Document-Level Membership Inference Attacks against DocVQA Models

ICLR 2025poster

Document Visual Question Answering (DocVQA) has introduced a new paradigm for end-to-end document understanding, and quickly became one of the standard benchmarks for multimodal LLMs. Automating document processing workflows, driven by DocVQA models, presents significant potential for many business…

2025

DocVXQA: Context-Aware Visual Explanations for Document Question Answering

ICML 2025poster

We propose **DocVXQA**, a novel framework for visually self-explainable document question answering, where the goal is not only to produce accurate answers to questions but also to learn visual heatmaps that highlight critical regions, offering interpretable justifications for the model decision. To…

2024

CoMix: A Comprehensive Benchmark for Multi-Task Comic Understanding

NeurIPS 2024poster

The comic domain is rapidly advancing with the development of single-page analysis and synthesis models. However, evaluation metrics and datasets lag behind, often limited to small-scale or single-style test sets. We introduce a novel benchmark, CoMix, designed to evaluate the multi-task capabilitie…

2023

Show, Interpret and Tell: Entity-Aware Contextualised Image Captioning in Wikipedia

AAAI 2023technical

Humans exploit prior knowledge to describe images, and are able to adapt their explanation to specific contextual information given, even to the extent of inventing plausible explanations when contextual information and images do not match. In this work, we propose the novel task of captioning Wikip…

2023

Text-DIAE: A Self-Supervised Degradation Invariant Autoencoder for Text Recognition and Document Enhancement

AAAI 2023technical

In this paper, we propose a Text-Degradation Invariant Auto Encoder (Text-DIAE), a self-supervised model designed to tackle two tasks, text recognition (handwritten or scene-text) and document image enhancement. We start by employing a transformer-based architecture that incorporates three pretext…

2020

RoadText-1K: Text Detection & Recognition Dataset for Driving Videos

ICRA 2020poster

Perceiving text is crucial to understand semantics of outdoor scenes and hence is a critical requirement to build intelligent systems for driver assistance and self-driving. Most of the existing datasets for text detection and recognition comprise still images and are mostly compiled keeping text in…

Cited by 62SourceScholar
2019

Good News, Everyone! Context Driven Entity-Aware Captioning for News Images

CVPR 2019poster

Current image captioning systems perform at a merely descriptive level, essentially enumerating the objects in the scene and their relations. Humans, on the contrary, interpret images by integrating several sources of prior knowledge of the world. In this work, we aim to take a step closer to produc…

Cited by 191PDFcodeScholar
2019

Scene Text Visual Question Answering

ICCV 2019poster

Current visual question answering datasets do not consider the rich semantic information conveyed by text within an image. In this work, we present a new dataset, ST-VQA, that aims to highlight the importance of exploiting high-level semantic information present in images as textual cues in the Visu…

Cited by 417PDFScholar
2017

Self-Supervised Learning of Visual Features Through Embedding Images Into Text Topic Spaces

CVPR 2017poster

End-to-end training from scratch of current deep architectures for new computer vision problems would require Imagenet-scale datasets, and this is not always possible. In this paper we present a method that is able to take advantage of freely available multi-modal content to train computer vision al…

Cited by 143PDFScholar