← Search

Vittorio Ferrari

50 accepted papers

2026

MaskInversion: Localized Embeddings via Optimization of Explainability Maps

ICLR 2026poster

Vision-language foundation models such as CLIP have achieved tremendous results in global vision-language alignment, but still show some limitations in creating representations for specific image regions. To address this problem, we propose MaskInversion, a method that leverages the feature represe…

Cited by 0SourcecodeScholar
2024

Grounding Everything: Emerging Localization Properties in Vision-Language Transformers

CVPR 2024poster

Vision-language foundation models have shown remarkable performance in various zero-shot settings such as image retrieval classification or captioning. But so far those models seem to fall behind when it comes to zero-shot localization of referential expressions and objects in images. As a result th…

2023

Agile Modeling: From Concept to Classifier in Minutes

ICCV 2023poster

The application of computer vision methods to nuanced, subjective concepts is growing. While crowdsourcing has served the vision community well for most objective tasks (such as labeling a "zebra"), it now falters on tasks where there is substantial subjectivity in the concept (such as identifying "…

Cited by 14PDFScholar
2023

CAD-Estate: Large-scale CAD Model Annotation in RGB Videos

ICCV 2023poster

We propose a method for annotating videos of complex multi-object scenes with a globally-consistent 3D representation of the objects. We annotate each object with a CAD model from a database, and place it in the 3D coordinate frame of the scene with a 9-DoF pose transformation. Our method is semi-au…

Cited by 7PDFcodeScholar
2023

Connecting Vision and Language With Video Localized Narratives

CVPR 2023highlight

We propose Video Localized Narratives, a new form of multimodal video annotations connecting vision and language. In the original Localized Narratives, annotators speak and move their mouse simultaneously on an image, thus grounding each word with a mouse trace segment. However, this is challenging…

2023

Encyclopedic VQA: Visual Questions About Detailed Properties of Fine-Grained Categories

ICCV 2023poster

We propose Encyclopedic-VQA, a large scale visual question answering (VQA) dataset featuring visual questions about detailed properties of fine-grained categories and instances. It contains 221k unique question+answer pairs each matched with (up to) 5 images, resulting in a total of 1M VQA samples.…

Cited by 38PDFcodeScholar
2023

Estimating Generic 3D Room Structures from 2D Annotations

NeurIPS 2023poster

Indoor rooms are among the most common use cases in 3D scene understanding. Current state-of-the-art methods for this task are driven by large annotated datasets. Room layouts are especially important, consisting of structural elements in 3D, such as wall, floor, and ceiling. However, they are diffi…

2023

NAVI: Category-Agnostic Image Collections with High-Quality 3D Shape and Pose Annotations

NeurIPS 2023poster

Recent advances in neural reconstruction enable high-quality 3D object reconstruction from casually captured image collections. Current techniques mostly analyze their progress on relatively simple image collections where SfM techniques can provide ground-truth (GT) camera poses. We note that SfM te…

2023

StoryBench: A Multifaceted Benchmark for Continuous Story Visualization

NeurIPS 2023poster

Generating video stories from text prompts is a complex task. In addition to having high visual quality, videos need to realistically adhere to a sequence of text prompts whilst being consistent throughout the frames. Creating a benchmark for video generation requires data annotated over time, which…

2023

Tracking by 3D Model Estimation of Unknown Objects in Videos

ICCV 2023poster

Most model-free visual object tracking methods formulate the tracking task as object location estimation given by a 2D segmentation or a bounding box in each video frame. We argue that this representation is limited and instead propose to guide and improve 2D tracking with an explicit object represe…

Cited by 7PDFScholar
2022

How Stable Are Transferability Metrics Evaluations?

ECCV 2022poster

"Transferability metrics is a maturing field with increasing interest, which aims at providing heuristics for selecting the most suitable source models to transfer to a given target dataset, without fine-tuning them all. However, existing works rely on custom experimental setups which differ across…

2022

Motion-From-Blur: 3D Shape and Motion Estimation of Motion-Blurred Objects in Videos

CVPR 2022poster

We propose a method for jointly estimating the 3D motion, 3D shape, and appearance of highly motion-blurred objects from a video. To this end, we model the blurred appearance of a fast moving object in a generative fashion by parametrizing its 3D position, rotation, velocity, acceleration, bounces,…

Cited by 10PDFcodeScholar
2022

RayTran: 3D Pose Estimation and Shape Reconstruction of Multiple Objects from Videos with Ray-Traced Transformers

ECCV 2022poster

"We propose a transformer-based neural network architecture for multi-object 3D reconstruction from RGB videos. It relies on two alternative ways to represent its knowledge: as a global 3D grid of features and an array of view-specific 2D grids. We progressively exchange information between the two…

2022

The Missing Link: Finding Label Relations across Datasets

ECCV 2022poster

"Computer Vision is driven by the many datasets which can be used for training or evaluating novel methods. Each of these dataset, however, has its own design principles resulting in a different set of labels,different appearance domains and different annotation instructions. In this paper we explor…

2022

Transferability Estimation Using Bhattacharyya Class Separability

CVPR 2022poster

Transfer learning has become a popular method for leveraging pre-trained models in computer vision. However, without performing computationally expensive fine-tuning, it is difficult to quantify which pre-trained source models are suitable for a specific target task, or, conversely, to which tasks a…

Cited by 80PDFcodeScholar
2022

Transferability Metrics for Selecting Source Model Ensembles

CVPR 2022oral

We address the problem of ensemble selection in transfer learning: Given a large pool of source models we want to select an ensemble of models which, after fine-tuning on the target training set, yields the best performance on the target test set. Since fine-tuning all possible ensembles is computat…

Cited by 31PDFScholar
2022

Uncertainty-Aware Deep Multi-View Photometric Stereo

CVPR 2022poster

This paper presents a simple and effective solution to the longstanding classical multi-view photometric stereo (MVPS) problem. It is well-known that photometric stereo (PS) is excellent at recovering high-frequency surface details, whereas multi-view stereo (MVS) can help remove the low-frequency d…

Cited by 45PDFScholar
2022

Urban Radiance Fields

CVPR 2022poster

The goal of this work is to perform 3D reconstruction and novel view synthesis from data captured by scanning platforms commonly deployed for world mapping in urban outdoor environments (e.g., Street View). Given a sequence of posed RGB images and lidar sweeps acquired by cameras and scanners moving…

Cited by 343PDFScholar
2021

DeFMO: Deblurring and Shape Recovery of Fast Moving Objects

CVPR 2021poster

Objects moving at high speed appear significantly blurred when captured with cameras. The blurry appearance is especially ambiguous when the object has complex shape or texture. In such cases, classical methods, or even humans, are unable to recover the object's appearance and motion. We propose a m…

Cited by 50PDFcodeScholar
2021

Shape from Blur: Recovering Textured 3D Shape and Motion of Fast Moving Objects

NeurIPS 2021poster

We address the novel task of jointly reconstructing the 3D shape, texture, and motion of an object from a single motion-blurred image. While previous approaches address the deblurring problem only in the 2D image domain, our proposed rigorous modeling of all object properties in the 3D domain enable…

2021

Sharf: Shape-conditioned Radiance Fields from a Single View

ICML 2021spotlight

We present a method for estimating neural scenes representations of objects given only a single image. The core of our method is the estimation of a geometric scaffold for the object and its use as a guide for the reconstruction of the underlying radiance field. Our formulation is based on a generat…

Cited by 124SourcePDFScholar
2021

Telling the What While Pointing to the Where: Multimodal Queries for Image Retrieval

ICCV 2021poster

Most existing image retrieval systems use text queries as a way for the user to express what they are looking for. However, fine-grained image retrieval often requires the ability to also express where in the image the content they are looking for is. The text modality can only cumbersomely express…

Cited by 26PDFScholar
2021

Towards Reusable Network Components by Learning Compatible Representations

AAAI 2021technical

This paper proposes to make a first step towards compatible and hence reusable network components. Rather than training networks for different tasks independently, we adapt the training process to produce network components that are compatible across tasks. In particular, we split a network into two…

Cited by 14SourcePDFScholar
2021

Uncalibrated Neural Inverse Rendering for Photometric Stereo of General Surfaces

CVPR 2021poster

This paper presents an uncalibrated deep neural network framework for the photometric stereo problem. For training models to solve the problem, existing neural network-based methods either require exact light directions or ground-truth surface normals of the object or both. However, in practice, it…

Cited by 64PDFScholar
2020

C-Flow: Conditional Generative Flow Models for Images and 3D Point Clouds

CVPR 2020poster

Flow-based generative models have highly desirable properties like exact log-likelihood evaluation and exact latent-variable inference, however they are still in their infancy and have not received as much attention as alternative generative models. In this paper, we introduce C-Flow, a novel condit…

Cited by 113PDFScholar
2020

CoReNet: Coherent 3D Scene Reconstruction from a Single RGB Image

ECCV 2020poster

Advances in deep learning techniques have allowed recent work to reconstruct the shape of a single object given only one RBG image as input. Building on common encoder-decoder architectures for this task, we propose three extensions: (1) ray-traced skip connections that propagate local 2D informatio…

2020

Connecting Vision and Language with Localized Narratives

ECCV 2020poster

We propose Localized Narratives, a new form of multimodal image annotations connecting vision and language. We ask annotators to describe an image with their voice while simultaneously hovering their mouse over the region they are describing. Since the voice and the mouse pointer are synchronized, w…

2020

Continuous Adaptation for Interactive Object Segmentation by Learning from Corrections

ECCV 2020poster

In interactive object segmentation a user collaborates with a computer vision model to segment an object. Recent works employ convolutional neural networks for this task: Given an image and a set of corrections made by the user as input, they output a segmentation mask. These approaches achieve stro…

Cited by 65SourcePDFScholar
2019

Interactive Full Image Segmentation by Considering All Regions Jointly

CVPR 2019poster

We address interactive full image annotation, where the goal is to accurately segment all object and stuff regions in an image. We propose an interactive, scribble-based annotation framework which operates on the whole image to produce segmentations for all regions. This enables sharing scribble cor…

Cited by 97PDFScholar
2018

Learning Intelligent Dialogs for Bounding Box Annotation

CVPR 2018poster

We introduce Intelligent Annotation Dialogs for bounding box annotation. We train an agent to automatically choose a sequence of actions for a human annotator to produce a bounding box in a minimal amount of time. Specifically, we consider two actions: box verification, where the annotator verifies…

2018

Revisiting Knowledge Transfer for Training Object Class Detectors

CVPR 2018poster

We propose to revisit knowledge transfer for training object detectors on target classes from weakly supervised training images, helped by a set of source classes with bounding-box annotations. We present a unified knowledge transfer framework based on training a single neural network multi-class ob…

Cited by 89SourcePDFScholar
2017

Action Tubelet Detector for Spatio-Temporal Action Localization

ICCV 2017poster

Current state-of-the-art approaches for spatio-temporal action localization rely on detections at the frame level that are then linked or tracked across time. In this paper, we leverage the temporal continuity of videos instead of operating at the frame level. We propose the ACtion Tubelet detector…

Cited by 437PDFcodeScholar
2017

Extreme Clicking for Efficient Object Annotation

ICCV 2017poster

Manually annotating object bounding boxes is central to building computer vision datasets, and it is very time consuming (annotating ILSVRC [53] took 35s for one high-quality box [62]). It involves clicking on imaginary corners of a tight box around the object. This is difficult as these corners are…

Cited by 330PDFScholar
2017

Training Object Class Detectors With Click Supervision

CVPR 2017spotlight

Training object class detectors typically requires a large set of images with objects annotated by bounding boxes. However, manually drawing bounding boxes is very time consuming. In this paper we greatly reduce annotation time by proposing center-click annotations: we ask annotators to click on the…

Cited by 157PDFScholar
2016

Discovering the Physical Parts of an Articulated Object Class From Multiple Videos

CVPR 2016poster

We propose a motion-based method to discover the physical parts of an articulated object class (e.g. head/torso/leg of a horse) from multiple videos. The key is to find object regions that exhibit consistent motion relative to the rest of the object, across multiple videos. We can then learn a locat…

Cited by 14PDFScholar
2016

How Hard Can It Be? Estimating the Difficulty of Visual Search in an Image

CVPR 2016poster

We address the problem of estimating image difficulty defined as the human response time for solving a visual search task. We collect human annotations of image difficulty for the PASCAL VOC 2012 data set through a crowd-sourcing platform. We then analyze what human interpretable image properties ca…

Cited by 164PDFScholar
2016

We Don't Need No Bounding-Boxes: Training Object Class Detectors Using Only Human Verification

CVPR 2016spotlight

Training object class detectors typically requires a large set of images in which objects are annotated by bounding-boxes. However, manually drawing bounding-boxes is very time consuming. We propose a new scheme for training object detectors which only requires annotators to verify bounding-boxes pr…

Cited by 179PDFScholar
2015

An Active Search Strategy for Efficient Object Class Detection

CVPR 2015poster

Object class detectors typically apply a window classifier to all the windows in a large set, either in a sliding window manner or using object proposals. In this paper, we develop an active search strategy that sequentially chooses the next window to evaluate based on all the information gathered b…

Cited by 84SourcePDFScholar
2015

Articulated Motion Discovery Using Pairs of Trajectories

CVPR 2015poster

We propose an unsupervised approach for discovering characteristic motion patterns in videos of highly articulated objects performing natural, unscripted behaviors, such as tigers in the wild. We discover consistent patterns in a bottom-up manner by analyzing the relative displacements of large numb…

Cited by 50SourcePDFScholar
2015

Joint Calibration of Ensemble of Exemplar SVMs

CVPR 2015poster

We present a method for calibrating the Ensemble of Exemplar SVMs model. Unlike the standard approach, which calibrates each SVM independently, our method optimizes their joint performance as an ensemble. We formulate joint calibration as a constrained optimization problem and devise an efficient op…

Cited by 14SourcePDFScholar