← Search

Georgios Pantazopoulos

5 accepted papers

2025

CROPE: Evaluating In-Context Adaptation of Vision and Language Models to Culture-Specific Concepts

NAACL 2025long

As Vision and Language models (VLMs) become accessible across the globe, it is important that they demonstrate cultural knowledge. In his paper, we introduce CROPE, a visual question answering benchmark designed to probe the knowledge of culture-specific concepts and evaluate the capacity for cultur…

2025

Evaluating Multimodal Language Models as Visual Assistants for Visually Impaired Users

ACL 2025long

This paper explores the effectiveness of Multimodal Large Language models (MLLMs) as assistive technologies for visually impaired individuals. We conduct a user survey to identify adoption patterns and key challenges users face with such technologies. Despite a high adoption rate of these models, ou…

Cited by 0SourcePDFScholar
2024

Lost in Space: Probing Fine-grained Spatial Understanding in Vision and Language Resamplers

NAACL 2024short

An effective method for combining frozen large language models (LLM) and visual encoders involves a resampler module that creates a ‘visual prompt’ which is provided to the LLM, along with the textual prompt. While this approach has enabled impressive performance across many coarse-grained tasks lik…

2024

Shaking Up VLMs: Comparing Transformers and Structured State Space Models for Vision & Language Modeling

EMNLP 2024main

This study explores replacing Transformers in Visual Language Models (VLMs) with Mamba, a recent structured state space model (SSM) that demonstrates promising performance in sequence modeling. We test models up to 3B parameters under controlled conditions, showing that Mamba-based VLMs outperforms…

2023

Multitask Multimodal Prompted Training for Interactive Embodied Task Completion

EMNLP 2023long main

Interactive and embodied tasks pose at least two fundamental challenges to existing Vision \& Language (VL) models, including 1) grounding language in trajectories of actions and observations, and 2) referential disambiguation. To tackle these challenges, we propose an Embodied MultiModal Agent (EMM…

Cited by 0SourceScholar