← Search

Iñigo Alonso

8 accepted papers

2026

TABLET: A Large-Scale Dataset for Robust Visual Table Understanding

ICLR 2026poster

While table understanding increasingly relies on pixel-only settings, current benchmarks predominantly use synthetic renderings that lack the complexity and visual diversity of real-world tables. Additionally, existing visual table understanding (VTU) datasets offer fixed examples with single visual…

Cited by 0SourcecodeScholar
2025

Vision-Language Models Struggle to Align Entities across Modalities

ACL 2025finding

Cross-modal entity linking refers to the ability to align entities and their attributes across different modalities. While cross-modal entity linking is a fundamental skill needed for real-world applications such as multimodal code generation, fake news detection, or scene understanding, it has not…

2021

Semi-Supervised Semantic Segmentation With Pixel-Level Contrastive Learning From a Class-Wise Memory Bank

ICCV 2021poster

This work presents a novel approach for semi-supervised semantic segmentation. The key element of this approach is our contrastive learning module that enforces the segmentation network to yield similar pixel-level feature representations for same-class samples across the whole dataset. To achieve t…

Cited by 288PDFcodeScholar
2020

3D-MiniNet: Learning a 2D Representation From Point Clouds for Fast and Efficient 3D LIDAR Semantic Segmentation

RA-L 2020

LIDAR semantic segmentation is an essential task that provides 3D semantic information about the environment to robots. Fast and efficient semantic segmentation methods are needed to match the strong computational and temporal restrictions of many real-world robotic applications. This work presents

Cited by 162SourcecodeScholar
2019

Enhancing V-SLAM Keyframe Selection with an Efficient ConvNet for Semantic Analysis

ICRA 2019poster

Selecting relevant visual information from a video is a challenging task on its own and even more in robotics, due to strong computational restrictions. This work proposes a novel keyframe selection strategy based on image quality and semantic information, which boosts strategies currently used in V…

Cited by 19SourceScholar