← Search

Gregor Geigle

8 accepted papers

2025

Centurio: On Drivers of Multilingual Ability of Large Vision-Language Model

ACL 2025long

Most Large Vision-Language Models (LVLMs) to date are trained predominantly on English data, which makes them struggle to understand non-English input and fail to generate output in the desired target language. Existing efforts mitigate these issues by adding multilingual training data, but do so in…

Cited by 0SourcePDFScholar
2024

African or European Swallow? Benchmarking Large Vision-Language Models for Fine-Grained Object Classification

EMNLP 2024main

Recent Large Vision-Language Models (LVLMs) demonstrate impressive abilities on numerous image understanding and reasoning tasks. The task of fine-grained object classification (e.g., distinction between animal species), however, has been probed insufficiently, despite its downstream importance. We…

2024

Babel-ImageNet: Massively Multilingual Evaluation of Vision-and-Language Representations

ACL 2024long

Vision-and-language (VL) models with separate encoders for each modality (e.g., CLIP) have become the go-to models for zero-shot image classification and image-text retrieval. They are, however, mostly evaluated in English as multilingual benchmarks are limited in availability. We introduce Babel-Im…

2024

Does Object Grounding Really Reduce Hallucination of Large Vision-Language Models?

EMNLP 2024main

Large vision-language models (LVLMs) have recently dramatically pushed the state of the art in image captioning and many image understanding tasks (e.g., visual question answering). LVLMs, however, often hallucinate and produce captions that mention concepts that cannot be found in the image. These…

Cited by 0SourcePDFScholar
2024

InstructIR: High-Quality Image Restoration Following Human Instructions

ECCV 2024poster

"Image restoration is a fundamental problem that involves recovering a high-quality clean image from its degraded observation. All-In-One image restoration models can effectively restore images from various types and levels of degradation using degradation-specific information as prompts to guide th…

2022

FigMemes: A Dataset for Figurative Language Identification in Politically-Opinionated Memes

EMNLP 2022main

Real-world politically-opinionated memes often rely on figurative language to cloak propaganda and radical ideas to help them spread. It is not only a scientific challenge to develop machine learning models to recognize them in memes, but also sociologically beneficial to understand hidden meanings…

Cited by 22SourcePDFScholar
2022

xGQA: Cross-Lingual Visual Question Answering

ACL 2022findings

Recent advances in multimodal vision and language modeling have predominantly focused on the English language, mostly due to the lack of multilingual multimodal datasets to steer modeling efforts. In this work, we address this gap and provide xGQA, a new multilingual evaluation benchmark for the vis…

2021

AdapterDrop: On the Efficiency of Adapters in Transformers

EMNLP 2021main

Transformer models are expensive to fine-tune, slow for inference, and have large storage requirements. Recent approaches tackle these shortcomings by training smaller models, dynamically reducing the model size, and by training light-weight adapters. In this paper, we propose AdapterDrop, removing…

Cited by 268SourcePDFScholar