← Search

Fabian Caba Heilbron

18 accepted papers

2025

Discovering Divergent Representations between Text-to-Image Models

ICCV 2025poster

In this paper, we investigate when and how visual representations learned by two different generative models diverge from each other. Specifically, given two text-to-image models, our goal is to discover visual attributes that appear in images generated by one model but not the other, along with the…

Cited by 0SourcePDFScholar
2025

Improving Personalized Search with Regularized Low-Rank Parameter Updates

CVPR 2025highlight

Personalized vision-language retrieval seeks to recognize new concepts (e.g. "my dog Fido") from only a few examples. This task is challenging because it requires not only learning a new concept from a few images, but also integrating the personal and general knowledge together to recognize the conc…

2025

ResidualViT for Efficient Temporally Dense Video Encoding

ICCV 2025poster

Several video understanding tasks, such as natural language temporal video grounding, temporal activity localization, and audio description generation, require "temporally dense" reasoning over frames sampled at high temporal resolution. However, computing frame-level features for these tasks is com…

Cited by 0SourcePDFScholar
2024

Concept Weaver: Enabling Multi-Concept Fusion in Text-to-Image Models

CVPR 2024poster

While there has been significant progress in customizing text-to-image generation models generating images that combine multiple personalized concepts remains challenging. In this work we introduce Concept Weaver a method for composing customized text-to-image diffusion models at inference time. Spe…

Cited by 12SourcePDFScholar
2024

Scaling Up Video Summarization Pretraining with Large Language Models

CVPR 2024poster

Long-form video content constitutes a significant portion of internet traffic making automated video summarization an essential research problem. However existing video summarization datasets are notably limited in their size constraining the effectiveness of state-of-the-art methods for generalizat…

Cited by 13SourcePDFScholar
2024

Towards Automated Movie Trailer Generation

CVPR 2024poster

Movie trailers are an essential tool for promoting films and attracting audiences. However the process of creating trailers can be time-consuming and expensive. To streamline this process we propose an automatic trailer generation framework that generates plausible trailers from a full movie by auto…

Cited by 3SourcePDFScholar
2023

Localizing Moments in Long Video Via Multimodal Guidance

ICCV 2023poster

The recent introduction of the large-scale, long-form MAD and Ego4D datasets has enabled researchers to investigate the performance of current state-of-the-art methods for video grounding in the long-form setup, with interesting findings: current grounding methods alone fail at tackling this challen…

Cited by 26PDFcodeScholar
2023

Long-range Multimodal Pretraining for Movie Understanding

ICCV 2023poster

Learning computer vision models from (and for) movies has a long-standing history. While great progress has been attained, there is still a need for a pretrained multimodal model that can perform well in the ever-growing set of movie understanding tasks the community has been establishing. In this w…

Cited by 16PDFScholar
2023

Meta-Personalizing Vision-Language Models To Find Named Instances in Video

CVPR 2023poster

Large-scale vision-language models (VLM) have shown impressive results for language-guided search applications. While these models allow category-level queries, they currently struggle with personalized searches for moments in a video where a specific object instance such as "My dog Biscuit" appears…

2023

PIVOT: Prompting for Video Continual Learning

CVPR 2023poster

Modern machine learning pipelines are limited due to data availability, storage quotas, privacy regulations, and expensive annotation processes. These constraints make it difficult or impossible to train and update large-scale models on such dynamic annotated sets. Continual learning directly approa…

Cited by 60SourcePDFScholar
2021

Real-Time Semantic Segmentation With Fast Attention

RA-L 2021

In deep CNN based models for semantic segmentation, high accuracy relies on rich spatial context (large receptive fields) and fine spatial details (high resolution), both of which incur high computational costs. In this letter, we propose a novel architecture that addresses both challenges and achie

Cited by 143SourcecodeScholar
2018

Action Search: Spotting Actions in Videos and Its Application to Temporal Action Localization

ECCV 2018poster

State-of-the-art temporal action detectors inefficiently search the entire video for specific actions. Despite the encouraging progress these methods achieve, it is crucial to design automated approaches that only explore parts of the video which are the most relevant to the actions being searched f…

2018

Diagnosing Error in Temporal Action Detectors

ECCV 2018poster

Despite the recent progress in video understanding and the continuous rate of improvement in temporal action localization throughout the years, it is still unclear how far (or close?) we are to solving the problem. To this end, we introduce a new diagnostic tool to analyze the performance of tempora…

2018

What do I Annotate Next? An Empirical Study of Active Learning for Action Localization

ECCV 2018poster

Despite tremendous progress achieved in temporal action localization, state-of-the-art methods still struggle to train accurate models when annotated data is scarce. In this paper, we introduce a novel active learning framework for temporal localization that aims to mitigate this data dependency iss…

Cited by 51SourcePDFScholar
2017

SCC: Semantic Context Cascade for Efficient Action Detection

CVPR 2017poster

Despite the recent advances in large-scale video analysis, action detection remains as one of the most challenging unsolved problems in computer vision. This snag is in part due to the large volume of data that needs to be analyzed to detect actions in videos. Existing approaches have mitigated the…

Cited by 111PDFScholar
2016

Fast Temporal Activity Proposals for Efficient Detection of Human Actions in Untrimmed Videos

CVPR 2016poster

In many large-scale video analysis scenarios, one is interested in localizing and recognizing human activities that occur in short temporal intervals within long untrimmed videos. Current approaches for activity detection still struggle to handle large-scale video collections and the task remains re…

Cited by 355PDFScholar
2015

ActivityNet: A Large-Scale Video Benchmark for Human Activity Understanding

CVPR 2015poster

In spite of many dataset efforts for human action recognition, current computer vision algorithms are still severely limited in terms of the variability and complexity of the actions that they can recognize. This is in part due to the simplicity of current benchmarks, which mostly focus on simple ac…

Cited by 3269SourcePDFScholar
2015

Robust Manhattan Frame Estimation From a Single RGB-D Image

CVPR 2015poster

This paper proposes a new framework for estimating the Manhattan Frame (MF) of an indoor scene from a single RGB-D image. Our technique formulates this problem as the estimation of a rotation matrix that best aligns the normals of the captured scene to a canonical world axes. By introducing sparsity…

Cited by 44SourcePDFScholar