← Search

Imanol Miranda

2 accepted papers

2026

TABLET: A Large-Scale Dataset for Robust Visual Table Understanding

ICLR 2026poster

While table understanding increasingly relies on pixel-only settings, current benchmarks predominantly use synthetic renderings that lack the complexity and visual diversity of real-world tables. Additionally, existing visual table understanding (VTU) datasets offer fixed examples with single visual…

Cited by 0SourcecodeScholar
2024

BiVLC: Extending Vision-Language Compositionality Evaluation with Text-to-Image Retrieval

NeurIPS 2024poster

Existing Vision-Language Compositionality (VLC) benchmarks like SugarCrepe are formulated as image-to-text retrieval problems, where, given an image, the models need to select between the correct textual description and a synthetic hard negative text. In this work, we present the Bidirectional Visio…