← Search

Julius Mayer

2 accepted papers

2025

Can you SPLICE it together? A Human Curated Benchmark for Probing Visual Reasoning in VLMs

EMNLP 2025

In this work, we introduce SPLICE, a human-curated benchmark derived from the COIN instructional video dataset, designed to probe event-based reasoning across multiple dimensions: temporal, causal, spatial, contextual, and general knowledge. SPLICE includes 3,381 human-filtered videos spanning 12 ca

Cited by 0SourcePDFScholar
2025

iVISPAR — An Interactive Visual-Spatial Reasoning Benchmark for VLMs

EMNLP 2025

Vision-Language Models (VLMs) are known to struggle with spatial reasoning and visual alignment. To help overcome these limitations, we introduce iVISPAR, an interactive multimodal benchmark designed to evaluate the spatial reasoning capabilities of VLMs acting as agents. iVISPAR is based on a varia