← Search

Ioanna Ntinou

2 accepted papers

2025

Vision-Free Retrieval: Rethinking Multimodal Search with Textual Scene Descriptions

EMNLP 2025

Contrastively-trained Vision-Language Models (VLMs), such as CLIP, have become the standard approach for learning discriminative vision-language representations. However, these models often exhibit shallow language understanding, manifesting bag-of-words behaviour. These limitations are reinforced b

Cited by 0SourcePDFScholar
2024

Multiscale Vision Transformers Meet Bipartite Matching for Efficient Single-stage Action Localization

CVPR 2024poster

Action Localization is a challenging problem that combines detection and recognition tasks which are often addressed separately. State-of-the-art methods rely on off-the-shelf bounding box detections pre-computed at high resolution and propose transformer models that focus on the classification task…