← Search

Michele Cafagna

2 accepted papers

2024

ViLMA: A Zero-Shot Benchmark for Linguistic and Temporal Grounding in Video-Language Models

ICLR 2024poster

With the ever-increasing popularity of pretrained Video-Language Models (VidLMs), there is a pressing need to develop robust evaluation methodologies that delve deeper into their visio-linguistic capabilities. To address this challenge, we present ViLMA (Video Language Model Assessment), a task-agno…

Cited by 13SourcePDFScholar
2022

VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic Phenomena

ACL 2022long

We propose VALSE (Vision And Language Structured Evaluation), a novel benchmark designed for testing general-purpose pretrained vision and language (V&L) models for their visio-linguistic grounding capabilities on specific linguistic phenomena. VALSE offers a suite of six tests covering various ling…