← Search

David Pujol-Perich

2 accepted papers

2026

Beyond Caption-Based Queries in Video Moment Retrieval

CVPR 2026

Current Video Moment Retrieval (VMR) models are trained on videos paired with captions, which are written by annotators after watching the videos. These captions are used as textual queries---which we term caption-based queries. This annotation process induces a visual bias, leading to overly descri

Cited by 0SourceScholar
2025

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding

ICCV 2025poster

Video Temporal Grounding (VTG) involves Moment Retrieval (MR) and Highlight Detection (HD) based on textual queries. For this, most methods rely solely on final-layer features of frozen large pre-trained backbones, limiting their adaptability to new domains. While full fine-tuning is often impractic…