← Search

Albert Clapés

4 accepted papers

2026

AdaSpot: Spend Resolution Where It Matters for Precise Event Spotting

CVPR 2026

Precise Event Spotting aims to localize fast-paced actions or events in videos with high temporal precision, a key task for applications in sports analytics, robotics, and autonomous systems. Existing methods typically process all frames uniformly, overlooking the inherent spatio-temporal redundancy

Cited by 0SourcecodeScholar
2026

Beyond Caption-Based Queries in Video Moment Retrieval

CVPR 2026

Current Video Moment Retrieval (VMR) models are trained on videos paired with captions, which are written by annotators after watching the videos. These captions are used as textual queries---which we term caption-based queries. This annotation process induces a visual bias, leading to overly descri

Cited by 0SourceScholar
2025

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding

ICCV 2025poster

Video Temporal Grounding (VTG) involves Moment Retrieval (MR) and Highlight Detection (HD) based on textual queries. For this, most methods rely solely on final-layer features of frozen large pre-trained backbones, limiting their adaptability to new domains. While full fine-tuning is often impractic…

2023

Gloss-Free Sign Language Translation: Improving from Visual-Language Pretraining

ICCV 2023poster

Sign Language Translation (SLT) is a challenging task due to its cross-domain nature, involving the translation of visual-gestural language to text. Many previous methods employ an intermediate representation,i.e., gloss sequences, to facilitate SLT, thus transforming it into a two-stage task of sig…

Cited by 62PDFcodeScholar