← Search

David Novotny

36 accepted papers

2026

ActionMesh: Animated 3D Mesh Generation with Temporal 3D Diffusion

CVPR 2026

Generating animated 3D objects is at the heart of many applications, yet most advanced works are typically difficult to apply in practice because of their limited setup, their long runtime, or their limited quality. We introduce ActionMesh, a generative model that predicts production-ready 3D meshes

Cited by 0SourcecodeScholar
2026

SceneScribe-1M: A Large-Scale Video Dataset with Comprehensive Geometric and Semantic Annotations

CVPR 2026

The convergence of 3D geometric perception and video synthesis has created an unprecedented demand for large-scale video data that is rich in both semantic and spatio-temporal information. While existing datasets have advanced either 3D understanding or video generation, a significant gap remains in

Cited by 0SourceScholar
2025

PartGen: Part-level 3D Generation and Reconstruction with Multi-view Diffusion Models

CVPR 2025highlight

Text- or image-to-3D generators and 3D scanners can now produce 3D assets with high-quality shapes and textures, but as single, fused entities lacking meaningful structure. In contrast, most applications and creative workflows require 3D assets to be composed of distinct, meaningful parts that can b…

Cited by 5SourcePDFScholar
2025

Twinner: Shining Light on Digital Twins in a Few Snaps

CVPR 2025poster

We present the first large reconstruction model, Twinner, capable of recovering a scene's illumination as well as an object's geometry and material properties from only a few posed images. Twinner is based on the Large Reconstruction Model and innovates in three key ways:1) We introduce a memory-eff…

Cited by 0SourcePDFScholar
2025

UnCommon Objects in 3D

CVPR 2025poster

We introduce Uncommon Objects in 3D (uCO3D), a new object-centric dataset for 3D deep learning and 3D generative AI. uCO3D is the largest publicly-available collection of high-resolution videos of objects with 3D annotations that ensures full-360 degree coverage. uCO3D is significantly more diverse…

2025

VGGT: Visual Geometry Grounded Transformer

CVPR 2025award

We present VGGT, a feed-forward neural network that directly infers all key 3D attributes of a scene, including camera parameters, point maps, depth maps, and 3D point tracks, from one, a few, or hundreds of its views. This approach is a step forward in 3D computer vision, where models have typicall…

2025

WildCAT3D: Appearance-Aware Multi-View Diffusion in the Wild

NeurIPS 2025poster

Despite recent advances in sparse novel view synthesis (NVS) applied to object-centric scenes, scene-level NVS remains a challenge. A central issue is the lack of available clean multi-view training data, beyond manually curated datasets with limited diversity, camera variation, or licensing issues.…

Cited by 0SourceScholar
2024

Animal Avatars: Reconstructing Animatable 3D Animals from Casual Videos

ECCV 2024oral

"We present a method to build animatable dog avatars from monocular videos. This is challenging as animals display a range of (unpredictable) non-rigid movements and have a variety of appearance details (e.g., fur, spots, tails). We develop an approach that links the video frames via a 4D solution t…

2024

Meta 3D AssetGen: Text-to-Mesh Generation with High-Quality Geometry, Texture, and PBR Materials

NeurIPS 2024poster

We present Meta 3D AssetGen (AssetGen), a significant advancement in text-to-3D generation which produces faithful, high-quality meshes with texture and material control. Compared to works that bake shading in the 3D object’s appearance, AssetGen outputs physically-based rendering (PBR) materials, s…

2024

VGGSfM: Visual Geometry Grounded Deep Structure From Motion

CVPR 2024highlight

Structure-from-motion (SfM) is a long-standing problem in the computer vision community which aims to reconstruct the camera poses and 3D structure of a scene from a set of unconstrained 2D images. Classical frameworks solve this problem in an incremental manner by detecting and matching keypoints r…

2024

ViewDiff: 3D-Consistent Image Generation with Text-to-Image Models

CVPR 2024poster

3D asset generation is getting massive amounts of attention inspired by the recent success on text-guided 2D content creation. Existing text-to-3D methods use pretrained text-to-image diffusion models in an optimization problem or fine-tune them on synthetic data which often results in non-photoreal…

2023

Common Pets in 3D: Dynamic New-View Synthesis of Real-Life Deformable Categories

CVPR 2023highlight

Obtaining photorealistic reconstructions of objects from sparse views is inherently ambiguous and can only be achieved by learning suitable reconstruction priors. Earlier works on sparse rigid object reconstruction successfully learned such priors from large datasets such as CO3D. In this paper, we…

2023

HOLODIFFUSION: Training a 3D Diffusion Model Using 2D Images

CVPR 2023poster

Diffusion models have emerged as the best approach for generative modeling of 2D images. Part of their success is due to the possibility of training them on millions if not billions of images with a stable learning objective. However, extending these models to 3D remains difficult for two reasons. F…

Cited by 121SourcePDFScholar
2023

HoloFusion: Towards Photo-realistic 3D Generative Modeling

ICCV 2023poster

Diffusion-based image generators can now produce high-quality and diverse samples, but their success has yet to fully translate to 3D generation: existing diffusion methods can either generate low-resolution but 3D consistent outputs, or detailed 2D views of 3D objects with potential structural defe…

Cited by 40PDFcodeScholar
2023

PoseDiffusion: Solving Pose Estimation via Diffusion-aided Bundle Adjustment

ICCV 2023poster

Camera pose estimation is a long-standing computer vision problem that to date often relies on classical methods, such as handcrafted keypoint matching, RANSAC and bundle adjustment. In this paper, we propose to formulate the Structure from Motion (SfM) problem inside a probabilistic diffusion frame…

Cited by 83PDFcodeScholar
2023

Replay: Multi-modal Multi-view Acted Videos for Casual Holography

ICCV 2023poster

We introduce Replay, a collection of multi-view, multi-modal videos of humans interacting socially. Each scene is filmed in high production quality, from different viewpoints with several static cameras, as well as wearable action cameras, and recorded with a large array of microphones at different…

Cited by 7PDFcodeScholar
2022

KeyTr: Keypoint Transporter for 3D Reconstruction of Deformable Objects in Videos

CVPR 2022oral

We consider the problem of reconstructing the depth of dynamic objects from videos. Recent progress in dynamic video depth prediction has focused on improving the output of monocular depth estimators by means of multi-view constraints while imposing little to no restrictions on the deformation of th…

Cited by 13PDFScholar
2022

iSDF: Real-Time Neural Signed Distance Fields for Robot Perception

RSS 2022poster

We present iSDF, a continual learning system for real-time signed distance field (SDF) reconstruction. Given a stream of posed depth images from a moving camera, it trains a randomly initialised neural network to map input 3D coordinate to approximate signed distance. The model is self-supervised by…

2021

Common Objects in 3D: Large-Scale Learning and Evaluation of Real-Life 3D Category Reconstruction

ICCV 2021poster

Traditional approaches for learning 3D object categories have been predominantly trained and evaluated on synthetic datasets due to the unavailability of real 3D-annotated category-centric data. Our main goal is to facilitate advances in this field by collecting real-world data in a magnitude simila…

Cited by 491PDFcodeScholar
2021

DensePose 3D: Lifting Canonical Surface Maps of Articulated Objects to the Third Dimension

ICCV 2021poster

We tackle the problem of monocular 3D reconstruction of articulated objects like humans and animals. Our key contribution is DensePose 3D, a novel parametric model of an articulated mesh, which can be learned in a self-supervised fashion from 2D image annotations only. This is in stark contrast with…

Cited by 5PDFScholar
2021

Discovering Relationships Between Object Categories via Universal Canonical Maps

CVPR 2021poster

We tackle the problem of learning the geometry of multiple categories of deformable objects jointly. Recent work has shown that it is possible to learn a unified dense pose predictor for several categories of related objects. However, training such models requires to initialize inter-category corres…

Cited by 24PDFScholar
2021

NeuroMorph: Unsupervised Shape Interpolation and Correspondence in One Go

CVPR 2021poster

We present NeuroMorph, a new neural network architecture that takes as input two 3D shapes and produces in one go, i.e. in a single feed forward pass, a smooth interpolation and point-to-point correspondences between them. The interpolation, expressed as a deformation field, changes the pose of the…

Cited by 81PDFScholar
2021

Unsupervised Learning of 3D Object Categories From Videos in the Wild

CVPR 2021poster

Recently, numerous works have attempted to learn 3D reconstructors of textured 3D models of visual categories given a training set of annotated static images of objects. In this paper, we seek to decrease the amount of needed supervision by leveraging a collection of object-centric videos captured i…

Cited by 81PDFScholar
2020

3D Multi-bodies: Fitting Sets of Plausible 3D Human Models to Ambiguous Image Data

NeurIPS 2020spotlight

We consider the problem of obtaining dense 3D reconstructions of deformable objects from single and partially occluded views. In such cases, the visual evidence is usually insufficient to identify a 3D reconstruction uniquely, so we aim at recovering several plausible reconstructions compatible with…

Cited by 94SourcePDFScholar
2020

Canonical 3D Deformer Maps: Unifying parametric and non-parametric methods for dense weakly-supervised category reconstruction

NeurIPS 2020poster

We propose the Canonical 3D Deformer Map, a new representation of the 3D shape of common object categories that can be learned from a collection of 2D images of independent objects. Our method builds in a novel way on concepts from parametric deformation models, non-parametric 3D reconstruction, and…

2020

Continuous Surface Embeddings

NeurIPS 2020poster

In this work, we focus on the task of learning and representing dense correspondences in deformable object categories. While this problem has been considered before, solutions so far have been rather ad-hoc for specific object types (i.e., humans), often with significant manual work involved. Howeve…

2019

C3DPO: Canonical 3D Pose Networks for Non-Rigid Structure From Motion

ICCV 2019oral

We propose C3DPO, a method for extracting 3D models of deformable objects from 2D keypoint annotations in unconstrained images. We do so by learning a deep network that reconstructs a 3D object from a single view at a time, accounting for partial occlusions, and explicitly factoring the effects of v…

Cited by 137PDFcodeScholar
2019

Correlated Uncertainty for Learning Dense Correspondences from Noisy Labels

NeurIPS 2019poster

Many machine learning methods depend on human supervision to achieve optimal performance. However, in tasks such as DensePose, where the goal is to establish dense visual correspondences between images, the quality of manual annotations is intrinsically limited. We address this issue by augmenting n…

Cited by 33SourcePDFScholar
2019

PerspectiveNet: A Scene-consistent Image Generator for New View Synthesis in Real Indoor Environments

NeurIPS 2019poster

Given a set of a reference RGBD views of an indoor environment, and a new viewpoint, our goal is to predict the view from that location. Prior work on new-view generation has predominantly focused on significantly constrained scenarios, typically involving artificially rendered views of isolated CAD…

Cited by 24SourcePDFScholar
2018

Self-Supervised Learning of Geometrically Stable Features Through Probabilistic Introspection

CVPR 2018poster

Self-supervision can dramatically cut back the amount of manually-labelled data required to train deep neural networks. While self-supervision has usually been considered for tasks such as image classification, in this paper we aim at extending it to geometry-oriented tasks such as semantic matching…

Cited by 87SourcePDFScholar
2018

Semi-convolutional Operators for Instance Segmentation

ECCV 2018poster

Object detection and instance segmentation are dominated by region-based methods such as Mask RCNN. However, there is a growing interest in reducing these problems to pixel labeling tasks, as the latter could be more efficient, could be integrated seamlessly in image-to-image network architectures a…

Cited by 110SourcePDFScholar
2017

AnchorNet: A Weakly Supervised Network to Learn Geometry-Sensitive Features for Semantic Matching

CVPR 2017poster

Despite significant progress of deep learning in recent years, state-of-the-art semantic matching methods still rely on legacy features such as SIFT or HoG. We argue that the strong invariance properties that are key to the success of recent deep architectures on the classification task make them un…

Cited by 71PDFScholar