← Search

Simon Ging

3 accepted papers

2024

Open-ended VQA benchmarking of Vision-Language models by exploiting Classification datasets and their semantic hierarchy

ICLR 2024spotlight

The evaluation of text-generative vision-language models is a challenging yet crucial endeavor. By addressing the limitations of existing Visual Question Answering (VQA) benchmarks and proposing innovative evaluation methodologies, our research seeks to advance our understanding of these models’ cap…

2020

COOT: Cooperative Hierarchical Transformer for Video-Text Representation Learning

NeurIPS 2020poster

Many real-world video-text tasks involve different levels of granularity, such as frames and words, clip and sentences or videos and paragraphs, each with distinct semantics. In this paper, we propose a Cooperative hierarchical Transformer (COOT) to leverage this hierarchy information and model the…