← Search

Tomas Jakab

10 accepted papers

2026

EgoEdit: Dataset, Real-Time Streaming Model, and Benchmark for Egocentric Video Editing

CVPR 2026

We study instruction-guided editing of egocentric videos for interactive AR applications. While recent AI video editors perform well on third-person footage, egocentric views present unique challenges -- including rapid egomotion, and frequent hand-object interactions -- that create a significant do

Cited by 0SourcecodeScholar
2025

DualPM: Dual Posed-Canonical Point Maps for 3D Shape and Pose Reconstruction

CVPR 2025highlight

The choice of data representation is a key factor in the success of deep learning in geometric tasks. For instance, DUSt3R has recently introduced the concept of viewpoint- invariant point maps, generalizing depth prediction, and showing that one can reduce all the key problems in the 3D reconstruct…

Cited by 1SourcePDFScholar
2025

VMem: Consistent Interactive Video Scene Generation with Surfel-Indexed View Memory

ICCV 2025accepted

We propose a novel memory module for building video generators capable of interactively exploring environments. Previous approaches have achieved similar results either by out-painting 2D views of a scene while incrementally reconstructing its 3D geometry--which quickly accumulates errors--or by usi…

Cited by 0SourcePDFScholar
2024

Instant Uncertainty Calibration of NeRFs Using a Meta-Calibrator

ECCV 2024poster

"Neural Radiance Fields (NeRFs) have markedly improved novel view synthesis, but accurate uncertainty quantification in their image predictions remains an open problem. The prevailing methods for estimating uncertainty, including the state-of-the-art Density-aware NeRF Ensembles (DANE) [?], quantify…

2024

Learning the 3D Fauna of the Web

CVPR 2024poster

Learning 3D models of all animals in nature requires massively scaling up existing solutions. With this ultimate goal in mind we develop 3D-Fauna an approach that learns a pan-category deformable 3D animal model for more than 100 animal species jointly. One crucial bottleneck of modeling animals is…

Cited by 19SourcePDFScholar
2024

Scene-Conditional 3D Object Stylization and Composition

ECCV 2024poster

"Recently, 3D generative models have made impressive progress, enabling the generation of almost arbitrary 3D assets from text or image inputs. However, these approaches generate objects in isolation without any consideration for the scene where they will eventually be placed. In this paper, we prop…

Cited by 1SourcePDFScholar
2023

MagicPony: Learning Articulated 3D Animals in the Wild

CVPR 2023poster

We consider the problem of predicting the 3D shape, articulation, viewpoint, texture, and lighting of an articulated animal like a horse given a single test image as input. We present a new method, dubbed MagicPony, that learns this predictor purely from in-the-wild single-view images of the object…

2021

KeypointDeformer: Unsupervised 3D Keypoint Discovery for Shape Control

CVPR 2021poster

We introduce KeypointDeformer, a novel unsupervised method for shape control through automatically discovered 3D keypoints. We cast this as the problem of aligning a source 3D object to a target 3D object from the same object category. Our method analyzes the difference between the shapes of the two…

Cited by 70PDFcodeScholar
2020

Self-Supervised Learning of Interpretable Keypoints From Unlabelled Videos

CVPR 2020oral

We propose a new method for recognizing the pose of objects from a single image that for learning uses only unlabelled videos and a weak empirical prior on the object poses. Video frames differ primarily in the pose of the objects they contain, so our method distils the pose information by analyzing…

Cited by 100PDFScholar
2018

Unsupervised Learning of Object Landmarks through Conditional Image Generation

NeurIPS 2018poster

We propose a method for learning landmark detectors for visual objects (such as the eyes and the nose in a face) without any manual supervision. We cast this as the problem of generating images that combine the appearance of the object as seen in a first example image with the geometry of the object…

Cited by 285SourcePDFScholar