← Search

Alexandros Xenos

4 accepted papers

2025

Vision-Free Retrieval: Rethinking Multimodal Search with Textual Scene Descriptions

EMNLP 2025

Contrastively-trained Vision-Language Models (VLMs), such as CLIP, have become the standard approach for learning discriminative vision-language representations. However, these models often exhibit shallow language understanding, manifesting bag-of-words behaviour. These limitations are reinforced b

Cited by 0SourcePDFScholar
2025

VladVA: Discriminative Fine-tuning of LVLMs

CVPR 2025poster

Contrastively-trained Vision-Language Models (VLMs) like CLIP have become the de facto approach for discriminative vision-language representation learning. However, these models have limited language understanding, often exhibiting a "bag of words" behavior. At the same time, Large Vision-Language M…

Cited by 0SourcePDFScholar
2023

A Simple Baseline for Knowledge-Based Visual Question Answering

EMNLP 2023short main

This paper is on the problem of Knowledge-Based Visual Question Answering (KB-VQA). Recent works have emphasized the significance of incorporating both explicit (through external databases) and implicit (through LLMs) knowledge to answer questions requiring external knowledge effectively. A common l…

Cited by 0SourcecodeScholar
2022

From the Detection of Toxic Spans in Online Discussions to the Analysis of Toxic-to-Civil Transfer

ACL 2022long

We study the task of toxic spans detection, which concerns the detection of the spans that make a text toxic, when detecting such spans is possible. We introduce a dataset for this task, ToxicSpans, which we release publicly. By experimenting with several methods, we show that sequence labeling mode…