← Search

Alberte Fern{\'a}ndez-Castro

1 accepted papers

2025

Concept-pedia: a Wide-coverage Semantically-annotated Multimodal Dataset

EMNLP 2025

Vision-language Models (VLMs), such as CLIP and SigLIP, have become the de facto standard for multimodal tasks, serving as essential building blocks for recent Multimodal Large Language Models, including LLaVA and PaliGemma. However, current evaluations for VLMs remain heavily anchored to ImageNet.

Cited by 0SourcePDFScholar