← Search

Pablo Arbelaez

9 accepted papers

2026

Benchmarking Open-ended Segmentation

ICLR 2026poster

Open-ended segmentation requires models capable of generating free-form descriptions of previously unseen concepts and regions. Despite advancements in model development, current evaluation protocols for open-ended segmentation tasks fail to capture the true semantic accuracy of the generated descri…

Cited by 0SourceScholar
2025

A Standardized Benchmark for Multilabel Antimicrobial Peptide Classification

NeurIPS 2025poster

Antimicrobial peptides have emerged as promising molecules to combat antimicrobial resistance. However, fragmented datasets, inconsistent annotations, and the lack of standardized benchmarks hinder computational approaches and slow down the discovery of new candidates. To address these challenges, w…

Cited by 0SourceScholar
2024

Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

CVPR 2024poster

We present Ego-Exo4D a diverse large-scale multimodal multiview video dataset and benchmark challenge. Ego-Exo4D centers around simultaneously-captured egocentric and exocentric video of skilled human activities (e.g. sports music dance bike repair). 740 participants from 13 cities worldwide perform…

2020

Active Speakers in Context

CVPR 2020poster

Current methods for active speaker detection focus on modeling audiovisual information from a single speaker. This strategy can be adequate for addressing single-speaker scenarios, but it prevents accurate detection when the task is to identify who of many candidate speakers are talking. This paper…

Cited by 104PDFcodeScholar
2018

Dynamic Multimodal Instance Segmentation Guided by Natural Language Queries

ECCV 2018poster

We address the problem of segmenting an object given a natural language expression that describes it. Current techniques tackle this task by either ( extit{i}) directly or recursively merging linguistic and visual information in the channel dimension and then performing convolutions; or by ( extit{i…

2015

Aligning 3D Models to RGB-D Images of Cluttered Scenes

CVPR 2015poster

The goal of this work is to represent objects in an RGB-D scene with corresponding 3D models from a library. We approach this problem by first detecting and segmenting object instances in the scene and then using a convolutional neural network (CNN) to predict the pose of the object. This CNN is tra…

Cited by 320SourcePDFScholar
2015

Hypercolumns for Object Segmentation and Fine-Grained Localization

CVPR 2015poster

Recognition algorithms based on convolutional networks (CNNs) typically use the output of the last layer as feature representation. However, the information in this layer may be too coarse to allow precise localization. On the contrary, earlier layers may be precise in localization but will not capt…

Cited by 2008SourcePDFScholar