← Search

Ander Salaberria

2 accepted papers

2025

Vision-Language Models Struggle to Align Entities across Modalities

ACL 2025finding

Cross-modal entity linking refers to the ability to align entities and their attributes across different modalities. While cross-modal entity linking is a fundamental skill needed for real-world applications such as multimodal code generation, fake news detection, or scene understanding, it has not…

2024

BiVLC: Extending Vision-Language Compositionality Evaluation with Text-to-Image Retrieval

NeurIPS 2024poster

Existing Vision-Language Compositionality (VLC) benchmarks like SugarCrepe are formulated as image-to-text retrieval problems, where, given an image, the models need to select between the correct textual description and a synthetic hard negative text. In this work, we present the Bidirectional Visio…