← Search

Ignacio Rocco

16 accepted papers

2026

Efficiently Reconstructing Dynamic Scenes One D4RT at a Time

CVPR 2026

Understanding and reconstructing the complex geometry and motion of dynamic 4D scenes from video remains a formidable challenge in computer vision. This paper introduces D4RT, a simple yet powerful feedforward network designed to efficiently solve this task. D4RT utilizes a unified transformer archi

Cited by 0SourceScholar
2025

Direct Motion Models for Assessing Generated Videos

ICML 2025poster

A current limitation of video generative video models is that they generate plausible looking frames, but poor motion --- an issue that is not well captured by FVD and other popular methods for evaluating generated videos. Here we go beyond FVD by developing a metric which better measures plausible…

2025

TAPNext: Tracking Any Point (TAP) as Next Token Prediction

ICCV 2025poster

Tracking Any Point (TAP) in a video is a challenging computer vision problem with many demonstrated applications in robotics, video editing, and 3D reconstruction. Existing methods for TAP rely heavily on complex tracking-specific inductive biases and heuristics, limiting their generality and potent…

2024

CoTracker: It is Better to Track Together

ECCV 2024poster

"We introduce , a transformer-based model that tracks a large number of 2D points in long video sequences. Differently from most existing approaches that track points independently, tracks them jointly, accounting for their dependencies. We show that joint tracking significantly improves tracking ac…

2024

TAPVid-3D: A Benchmark for Tracking Any Point in 3D

NeurIPS 2024poster

We introduce a new benchmark, TAPVid-3D, for evaluating the task of long-range Tracking Any Point in 3D (TAP-3D). While point tracking in two dimensions (TAP-2D) has many benchmarks measuring performance on real-world videos, such as TAPVid-DAVIS, three-dimensional point tracking has none. To this e…

2023

Common Pets in 3D: Dynamic New-View Synthesis of Real-Life Deformable Categories

CVPR 2023highlight

Obtaining photorealistic reconstructions of objects from sparse views is inherently ambiguous and can only be achieved by learning suitable reconstruction priors. Earlier works on sparse rigid object reconstruction successfully learned such priors from large datasets such as CO3D. In this paper, we…

2023

DynamicStereo: Consistent Dynamic Depth From Stereo Videos

CVPR 2023poster

We consider the problem of reconstructing a dynamic scene observed from a stereo camera. Most existing methods for depth from stereo treat different stereo frames independently, leading to temporally inconsistent depth predictions. Temporal consistency is especially important for immersive AR or VR…

2023

Replay: Multi-modal Multi-view Acted Videos for Casual Holography

ICCV 2023poster

We introduce Replay, a collection of multi-view, multi-modal videos of humans interacting socially. Each scene is filmed in high production quality, from different viewpoints with several static cameras, as well as wearable action cameras, and recorded with a large array of microphones at different…

Cited by 7PDFcodeScholar
2022

BodyMap: Learning Full-Body Dense Correspondence Map

CVPR 2022poster

Dense correspondence between humans carries powerful semantic information that can be utilized to solve fundamental problems for full-body understanding such as in-the-wild surface matching, tracking and reconstruction. In this paper we present BodyMap, a new framework for obtaining high-definition…

Cited by 22PDFScholar
2022

KeyTr: Keypoint Transporter for 3D Reconstruction of Deformable Objects in Videos

CVPR 2022oral

We consider the problem of reconstructing the depth of dynamic objects from videos. Recent progress in dynamic video depth prediction has focused on improving the output of monocular depth estimators by means of multi-view constraints while imposing little to no restrictions on the deformation of th…

Cited by 13PDFScholar
2020

Efficient Neighbourhood Consensus Networks via Submanifold Sparse Convolutions

ECCV 2020poster

In this work we target the problem of estimating accurately localised correspondences between a pair of images. We adopt the recent Neighbourhood Consensus Networks that have demonstrated promising performance for difficult correspondence problems and propose modifications to overcome their main lim…

Cited by 208SourcePDFScholar
2019

D2-Net: A Trainable CNN for Joint Description and Detection of Local Features

CVPR 2019poster

In this work we address the problem of finding reliable pixel-level correspondences under difficult imaging conditions. We propose an approach where a single convolutional neural network plays a dual role: It is simultaneously a dense feature descriptor and a feature detector. By postponing the dete…

Cited by 909PDFcodeScholar
2019

Is This the Right Place? Geometric-Semantic Pose Verification for Indoor Visual Localization

ICCV 2019poster

Visual localization in large and complex indoor scenes, dominated by weakly textured rooms and repeating geometric patterns, is a challenging problem with high practical relevance for applications such as Augmented Reality and robotics. To handle the ambiguities arising in this scenario, a common st…

Cited by 59PDFScholar
2018

Neighbourhood Consensus Networks

NeurIPS 2018spotlight

We address the problem of finding reliable dense correspondences between a pair of images. This is a challenging task due to strong appearance differences between the corresponding scene elements and ambiguities generated by repetitive patterns. The contributions of this work are threefold. First, i…

Cited by 519SourcePDFScholar