← Search

Fabian Caba

10 accepted papers

2024

Sync from the Sea: Retrieving Alignable Videos from Large-Scale Datasets

ECCV 2024oral

"Temporal video alignment aims to synchronize the key events like object interactions or action phase transitions in two videos. Such methods could benefit various video editing, processing, and understanding tasks. However, existing approaches operate under the restrictive assumption that a suitabl…

Cited by 1SourcePDFScholar
2023

Efficient Adaptive Human-Object Interaction Detection with Concept-guided Memory

ICCV 2023poster

Human Object Interaction (HOI) detection aims to localize and infer the relationships between a human and an object. Arguably, training supervised models for this task from scratch presents challenges due to the performance drop over rare classes and the high computational cost and time required to…

Cited by 24PDFcodeScholar
2022

MAD: A Scalable Dataset for Language Grounding in Videos From Movie Audio Descriptions

CVPR 2022poster

The recent and increasing interest in video-language research has driven the development of large-scale datasets that enable data-intensive machine learning techniques. In comparison, limited effort has been made at assessing the fitness of these datasets for the video-language grounding task. Recen…

Cited by 123PDFcodeScholar
2022

MovieCuts: A New Dataset and Benchmark for Cut Type Recognition

ECCV 2022poster

"Understanding movies and their structural patterns is a crucial task in decoding the craft of video editing. While previous works have developed tools for general analysis, such as detecting characters or recognizing cinematography properties at the shot level, less effort has been devoted to under…

2022

The Anatomy of Video Editing: A Dataset and Benchmark Suite for AI-Assisted Video Editing

ECCV 2022poster

"Machine learning is transforming the video editing industry. Recent advances in computer vision have leveled-up video editing tasks such as intelligent reframing, rotoscoping, color grading, or applying digital makeups. However, most of the solutions have focused on video manipulation and VFX. This…

2022

vCLIMB: A Novel Video Class Incremental Learning Benchmark

CVPR 2022oral

Continual learning (CL) is under-explored in the video domain. The few existing works contain splits with imbalanced class distributions over the tasks, or study the problem in unsuitable datasets. We introduce vCLIMB, a novel video continual learning benchmark. vCLIMB is a standardized test-bed to…

Cited by 46PDFScholar
2021

Learning To Cut by Watching Movies

ICCV 2021poster

Video content creation keeps growing at an incredible pace; yet, creating engaging stories remains challenging and requires non-trivial video editing expertise. Many video editing components are astonishingly hard to automate primarily due to the lack of raw video materials. This paper focuses on a…

Cited by 26PDFcodeScholar
2021

MAAS: Multi-Modal Assignation for Active Speaker Detection

ICCV 2021poster

Active speaker detection requires a solid integration of multi-modal cues. While individual modalities can approximate a solution, accurate predictions can only be achieved by explicitly fusing the audio and visual features and modeling their temporal progression. Despite its inherent muti-modal nat…

Cited by 64PDFcodeScholar
2020

Active Speakers in Context

CVPR 2020poster

Current methods for active speaker detection focus on modeling audiovisual information from a single speaker. This strategy can be adequate for addressing single-speaker scenarios, but it prevents accurate detection when the task is to identify who of many candidate speakers are talking. This paper…

Cited by 104PDFcodeScholar
2020

Temporally Distributed Networks for Fast Video Semantic Segmentation

CVPR 2020poster

We present TDNet, a temporally distributed network designed for fast and accurate video semantic segmentation. We observe that features extracted from a certain high-level layer of a deep CNN can be approximated by composing features extracted from several shallower sub-networks. Leveraging the inhe…

Cited by 250PDFScholar