← Search

Andrea Tagliasacchi

65 accepted papers

2026

Lyra: Generative 3D Scene Reconstruction via Video Diffusion Model Self-Distillation

ICLR 2026poster

The ability to generate virtual environments is crucial for applications ranging from gaming to physical AI domains such as robotics, autonomous driving, and industrial AI. Current learning-based 3D reconstruction methods rely on the availability of captured real-world multi-view data, which is not…

Cited by 0SourcecodeScholar
2026

MeshSplatting: Differentiable Rendering with Opaque Meshes

CVPR 2026

Primitive-based splatting methods like 3D Gaussian Splatting (3DGS) have revolutionized novel view synthesis with real-time rendering. However, their point-based representations remain incompatible with mesh-based pipelines that power AR/VR and game engines. We present Mesh Splatting, a mesh-based r

Cited by 0SourcecodeScholar
2026

ORBIT: Benchmarking SfM in the Wild with 360deg Video

CVPR 2026

Structure-from-Motion (SfM) is a cornerstone of 3D perception, yet current methods often fail when applied to complex videos involving challenging camera motions or dynamic scenes.Compounding the problem, the field lacks reliable ground-truth benchmarks for such difficult scenarios, making it hard t

Cited by 0SourceScholar
2026

Semantic Foam: Unifying Spatial and Semantic Scene Decomposition

CVPR 2026

Modern scene reconstruction methods, such as 3D Gaussian Splatting, deliver photo-realistic novel view synthesis at real-time speeds, yet their adoption in interactive graphics applications has been limited. A major bottleneck is the difficulty of interacting with these representations compared to t

Cited by 0SourceScholar
2026

Spherical Voronoi: Directional Appearance as a Differentiable Partition of the Sphere

CVPR 2026

Radiance field methods (e.g. 3D Gaussian Splatting) have emerged as a powerful paradigm for novel view synthesis, yet their appearance modeling often relies on Spherical Harmonics (SH), which impose fundamental limitations. SH struggle with high-frequency signals, exhibit Gibbs ringing artifacts, an

Cited by 0SourcecodeScholar
2025

3D Gaussian Flats: Hybrid 2D/3D Photometric Scene Reconstruction

NeurIPS 2025poster

Recent advances in radiance fields and novel view synthesis enable creation of realistic digital twins from photographs. However, current methods struggle with flat, texture-less surfaces, creating uneven and semi-transparent reconstructions, due to an ill-conditioned photometric reconstruction obje…

Cited by 0SourceScholar
2025

AC3D: Analyzing and Improving 3D Camera Control in Video Diffusion Transformers

CVPR 2025poster

Numerous works have recently integrated 3D camera control into foundational text-to-video models, but the resulting camera control is often imprecise, and video generation quality suffers. In this work, we analyze camera motion from a first principles perspective, uncovering insights that enable pre…

Cited by 10SourcePDFScholar
2025

Controlling Space and Time with Diffusion Models

ICLR 2025poster

We present 4DiM, a cascaded diffusion model for 4D novel view synthesis (NVS), supporting generation with arbitrary camera trajectories and timestamps, in natural scenes, conditioned on one or more images. With a novel architecture and sampling procedure, we enable training on a mixture of 3D (with…

2025

PIAD: Pose and Illumination agnostic Anomaly Detection

CVPR 2025poster

We introduce the Pose and Illumination agnostic Anomaly Detection (PIAD) problem, a generalization of pose-agnostic anomaly detection (PAD). Being illumination agnostic is critical, as it relaxes the assumption that training data for an object has to be acquired in the same light configuration of th…

2025

Radiant Foam: Real-Time Differentiable Ray Tracing

ICCV 2025poster

Research on differentiable scene representations is consistently moving towards more efficient, real-time models. Recently, this has led to the popularization of splatting methods, which eschew the traditional ray-based rendering of radiance fields in favor of rasterization. This has yielded a signi…

Cited by 0SourcePDFScholar
2025

RoMo: Robust Motion Segmentation Improves Structure from Motion

ICCV 2025poster

There has been extensive progress in the reconstruction and generation of 4D scenes from monocular casually-captured video. Estimating accurate camera poses from videos through structure-from-motion (SfM) relies on robustly separating static and dynamic parts of a video. We propose a novel approach…

Cited by 0SourcePDFScholar
2025

SMITE: Segment Me In TimE

ICLR 2025poster

Segmenting an object in a video presents significant challenges. Each pixel must be accurately labeled, and these labels must remain consistent across frames. The difficulty increases when the segmentation is with arbitrary granularity, meaning the number of segments can vary arbitrarily, and masks…

2025

StochasticSplats: Stochastic Rasterization for Sorting-Free 3D Gaussian Splatting

ICCV 2025poster

3D Gaussian splatting (3DGS) is a popular radiance field method, with many application-specific extensions. Most variants rely on the same core algorithm: depth-sorting of Gaussian splats then rasterizing in primitive order. This ensures correct alpha compositing, but can cause rendering artifacts d…

Cited by 0SourcePDFScholar
2025

VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control

ICLR 2025poster

Modern text-to-video synthesis models demonstrate coherent, photorealistic generation of complex videos from a text description. However, most existing models lack fine-grained control over camera movement, which is critical for downstream applications related to content creation, visual effects, an…

Cited by 38SourcePDFScholar
2024

"PointNeRF++: A multi-scale, point-based Neural Radiance Field"

ECCV 2024poster

"Point clouds offer an attractive source of information to complement images in neural scene representations, especially when few images are available. Neural rendering methods based on point clouds do exist, but they do not perform well when the point cloud is sparse or incomplete, which is often t…

2024

3D Gaussian Splatting as Markov Chain Monte Carlo

NeurIPS 2024spotlight

While 3D Gaussian Splatting has recently become popular for neural rendering, current methods rely on carefully engineered cloning and splitting strategies for placing Gaussians, which does not always generalize and may lead to poor-quality renderings. For many real-world scenes this leads to their…

2024

4D-fy: Text-to-4D Generation Using Hybrid Score Distillation Sampling

CVPR 2024poster

Recent breakthroughs in text-to-4D generation rely on pre-trained text-to-image and text-to-video models to generate dynamic 3D scenes. However current text-to-4D methods face a three-way tradeoff between the quality of scene appearance 3D structure and motion. For example text-to-image models and t…

2024

Accelerating Neural Field Training via Soft Mining

CVPR 2024poster

We present an approach to accelerate Neural Field training by efficiently selecting sampling locations. While Neural Fields have recently become popular it is often trained by uniformly sampling the training domain or through handcrafted heuristics. We show that improved convergence and final traini…

2024

BANF: Band-Limited Neural Fields for Levels of Detail Reconstruction

CVPR 2024poster

Largely due to their implicit nature neural fields lack a direct mechanism for filtering as Fourier analysis from discrete signal processing is not directly applicable to these representations. Effective filtering of neural fields is critical to enable level-of-detail processing in downstream applic…

Cited by 3SourcePDFScholar
2024

Bayes' Rays: Uncertainty Quantification for Neural Radiance Fields

CVPR 2024highlight

Neural Radiance Fields (NeRFs) have shown promise in applications like view synthesis and depth estimation but learning from multiview images faces inherent uncertainties. Current methods to quantify them are either heuristic or computationally demanding. We introduce BayesRays a post-hoc framework…

2024

Lagrangian Hashing for Compressed Neural Field Representations

ECCV 2024poster

"We present Lagrangian Hashing, a representation for neural fields combining the characteristics of fast training NeRF methods that rely on Eulerian grids (i.e. InstantNGP), with those that employ points equipped with features as a way to represent information (e.g. 3D Gaussian Splatting or PointNeR…

Cited by 1SourcePDFScholar
2024

Neural Fields as Distributions: Signal Processing Beyond Euclidean Space

CVPR 2024poster

Neural fields have emerged as a powerful and broadly applicable method for representing signals. However in contrast to classical discrete digital signal processing the portfolio of tools to process such representations is still severely limited and restricted to Euclidean domains. In this paper we…

Cited by 1SourcePDFScholar
2024

TC4D: Trajectory-Conditioned Text-to-4D Generation

ECCV 2024poster

"Recent techniques for text-to-4D generation synthesize dynamic 3D scenes using supervision from pre-trained text-to-video models. However, existing representations, such as deformation models or time-dependent neural representations, are limited in the amount of motion they can generate—they cannot…

Cited by 37SourcePDFScholar
2024

Unsupervised Keypoints from Pretrained Diffusion Models

CVPR 2024highlight

Unsupervised learning of keypoints and landmarks has seen significant progress with the help of modern neural network architectures but performance is yet to match the supervised counterpart making their practicability questionable. We leverage the emergent knowledge within text-to-image diffusion m…

2024

pixelSplat: 3D Gaussian Splats from Image Pairs for Scalable Generalizable 3D Reconstruction

CVPR 2024poster

We introduce pixelSplat a feed-forward model that learns to reconstruct 3D radiance fields parameterized by 3D Gaussian primitives from pairs of images. Our model features real-time and memory-efficient rendering for scalable training as well as fast 3D reconstruction at inference time. To overcome…

2023

BlendFields: Few-Shot Example-Driven Facial Modeling

CVPR 2023poster

Generating faithful visualizations of human faces requires capturing both coarse and fine-level details of the face geometry and appearance. Existing methods are either data-driven, requiring an extensive corpus of data not publicly accessible to the research community, or fail to capture fine detai…

Cited by 8SourcePDFScholar
2023

CC3D: Layout-Conditioned Generation of Compositional 3D Scenes

ICCV 2023poster

In this work, we introduce CC3D, a conditional generative model that synthesizes complex 3D scenes conditioned on 2D semantic scene layouts, trained using single-view images. Different from most existing 3D GANs that limit their applicability to aligned single objects, we focus on generating complex…

Cited by 44PDFScholar
2023

CUF: Continuous Upsampling Filters

CVPR 2023poster

Neural fields have rapidly been adopted for representing 3D signals, but their application to more classical 2D image-processing has been relatively limited. In this paper, we consider one of the most important operations in image processing: upsampling. In deep learning, learnable upsampling layers…

Cited by 11SourcePDFScholar
2023

MobileNeRF: Exploiting the Polygon Rasterization Pipeline for Efficient Neural Field Rendering on Mobile Architectures

CVPR 2023poster

Neural Radiance Fields (NeRFs) have demonstrated amazing ability to synthesize images of 3D scenes from novel views. However, they rely upon specialized volumetric rendering algorithms based on ray marching that are mismatched to the capabilities of widely deployed graphics hardware. This paper intr…

2023

NeuMap: Neural Coordinate Mapping by Auto-Transdecoder for Camera Localization

CVPR 2023poster

This paper presents an end-to-end neural mapping method for camera localization, dubbed NeuMap, encoding a whole scene into a grid of latent codes, with which a Transformer-based auto-decoder regresses 3D coordinates of query pixels. State-of-the-art feature matching methods require each scene to be…

2023

Neural Fields with Hard Constraints of Arbitrary Differential Order

NeurIPS 2023poster

While deep learning techniques have become extremely popular for solving a broad range of optimization problems, methods to enforce hard constraints during optimization, particularly on deep neural networks, remain underdeveloped. Inspired by the rich literature on meshless interpolation and its ext…

Cited by 8SourcePDFScholar
2023

Novel View Synthesis with Diffusion Models

ICLR 2023poster

We present 3DiM (pronounced "three-dim"), a diffusion model for 3D novel view synthesis from as few as a single image. The core of 3DiM is an image-to-image diffusion model -- 3DiM takes a single reference view and their poses as inputs, and generates a novel view via diffusion. 3DiM can then genera…

Cited by 277SourcePDFScholar
2023

OpenScene: 3D Scene Understanding With Open Vocabularies

CVPR 2023poster

Traditional 3D scene understanding approaches rely on labeled 3D datasets to train a model for a single task with supervision. We propose OpenScene, an alternative approach where a model predicts dense features for 3D scene points that are co-embedded with text and image pixels in CLIP feature space…

Cited by 341SourcePDFScholar
2023

RobustNeRF: Ignoring Distractors With Robust Losses

CVPR 2023highlight

Neural radiance fields (NeRF) excel at synthesizing new views given multi-view, calibrated images of a static scene. When scenes include distractors, which are not persistent during image capture (moving objects, lighting variations, shadows), artifacts appear as view-dependent effects or 'floaters'…

2023

SparsePose: Sparse-View Camera Pose Regression and Refinement

CVPR 2023poster

Camera pose estimation is a key step in standard 3D reconstruction pipelines that operates on a dense set of images of a single object or scene. However, methods for pose estimation often fail when there are only a few images available because they rely on the ability to robustly identify and match…

Cited by 45SourcePDFScholar
2023

Unsupervised Semantic Correspondence Using Stable Diffusion

NeurIPS 2023poster

Text-to-image diffusion models are now capable of generating images that are often indistinguishable from real images. To generate such images, these models must understand the semantics of the objects they are asked to generate. In this work we show that, without any training, one can leverage this…

2023

nerf2nerf: Pairwise Registration of Neural Radiance Fields

ICRA 2023poster

We introduce a technique for pairwise registration of neural fields that extends classical optimization-based local registration (i.e. ICP) to operate on Neural Radiance Fields (NeRF)-neural 3D scene representations trained from collections of calibrated images. NeRF does not decompose illumination…

Cited by 33SourcecodeScholar
2022

CoNeRF: Controllable Neural Radiance Fields

CVPR 2022poster

We extend neural 3D representations to allow for intuitive and interpretable user control beyond novel view rendering (i.e. camera control). We allow the user to annotate which part of the scene one wishes to control with just a small number of mask annotations in the training images. Our key idea i…

Cited by 109PDFcodeScholar
2022

D^2NeRF: Self-Supervised Decoupling of Dynamic and Static Objects from a Monocular Video

NeurIPS 2022accept

Given a monocular video, segmenting and decoupling dynamic objects while recovering the static environment is a widely studied problem in machine intelligence. Existing solutions usually approach this problem in the image domain, limiting their performance and understanding of the environment. We in…

2022

Kubric: A Scalable Dataset Generator

CVPR 2022poster

Data is the driving force of machine learning, with the amount and quality of training data often being more important for the performance of a system than architecture and training details. But collecting, processing and annotating real data at scale is difficult, expensive, and frequently raises a…

Cited by 249PDFcodeScholar
2022

MetaPose: Fast 3D Pose From Multiple Views Without 3D Supervision

CVPR 2022poster

In the era of deep learning, human pose estimation from multiple cameras with unknown calibration has received little attention to date. We show how to train a neural model to perform this task with high precision and minimal latency overhead. The proposed model takes into account joint location unc…

Cited by 34PDFcodeScholar
2022

Neural Descriptor Fields: SE(3)-Equivariant Object Representations for Manipulation

ICRA 2022poster

We present Neural Descriptor Fields (NDFs), an object representation that encodes both points and relative poses between an object and a target (such as a robot gripper or a rack used for hanging) via category-level descriptors. We employ this representation for object manipulation, where given a ta…

Cited by 184SourcecodeScholar
2022

Panoptic Neural Fields: A Semantic Object-Aware Neural Scene Representation

CVPR 2022poster

We present PanopticNeRF, an object-aware neural scene representation that decomposes a scene into a set of objects (things) and background (stuff). Each object is represented by a separate MLP that takes a position, direction, and time and outputs density and radiance. The background is represented…

Cited by 293PDFScholar
2022

Scene Representation Transformer: Geometry-Free Novel View Synthesis Through Set-Latent Scene Representations

CVPR 2022poster

A classical problem in computer vision is to infer a 3D scene representation from few images that can be used to render novel views at interactive rates. Previous work focuses on reconstructing pre-defined 3D representations, e.g. textured meshes, or implicit representations, e.g. radiance fields, a…

Cited by 208PDFScholar
2022

Urban Radiance Fields

CVPR 2022poster

The goal of this work is to perform 3D reconstruction and novel view synthesis from data captured by scanning platforms commonly deployed for world mapping in urban outdoor environments (e.g., Street View). Given a sequence of posed RGB images and lidar sweeps acquired by cameras and scanners moving…

Cited by 343PDFScholar
2021

COTR: Correspondence Transformer for Matching Across Images

ICCV 2021poster

We propose a novel framework for finding correspondences in images based on a deep neural network that, given two images and a query point in one of them, finds its correspondence in the other. By doing so, one has the option to query only the points of interest and retrieve sparse correspondences,…

Cited by 320PDFcodeScholar
2021

Canonical Capsules: Self-Supervised Capsules in Canonical Pose

NeurIPS 2021poster

We propose a self-supervised capsule architecture for 3D point clouds. We compute capsule decompositions of objects through permutation-equivariant attention, and self-supervise the process by training with pairs of randomly rotated objects. Our key idea is to aggregate the attention masks into sema…

2021

MIST: Multiple Instance Spatial Transformer

CVPR 2021poster

We propose a deep network that can be trained to tackle image reconstruction and classification problems that involve detection of multiple object instances, without any supervision regarding their whereabouts. The network learns to extract the most significant top-K patches, and feeds these patches…

Cited by 15PDFcodeScholar
2021

Unsupervised Part Representation by Flow Capsules

ICML 2021spotlight

Capsule networks aim to parse images into a hierarchy of objects, parts and relations. While promising, they remain limited by an inability to learn effective low level part descriptions. To address this issue we propose a way to learn primary capsule encoders that detect atomic parts from a single…

Cited by 49SourcePDFScholar
2021

Vector Neurons: A General Framework for SO(3)-Equivariant Networks

ICCV 2021poster

Invariance and equivariance to the rotation group have been widely discussed in the 3D deep learning community for pointclouds. Yet most proposed methods either use complex mathematical tools that may limit their accessibility, or are tied to specific input data types and network architectures. In t…

Cited by 345PDFcodeScholar
2020

ACNe: Attentive Context Normalization for Robust Permutation-Equivariant Learning

CVPR 2020poster

Many problems in computer vision require dealing with sparse, unordered data in the form of point clouds. Permutation-equivariant networks have become a popular solution - they operate on individual data points with simple perceptrons and extract contextual information with global pooling. This can…

Cited by 196PDFcodeScholar
2020

CoSE: Compositional Stroke Embeddings

NeurIPS 2020poster

We present a generative model for stroke-based drawing tasks which is able to model complex free-form structures. While previous approaches rely on sequence-based models for drawings of basic objects or handwritten text, we propose a model that treats drawings as a collection of strokes that can be…

2020

CvxNet: Learnable Convex Decomposition

CVPR 2020oral

Any solid object can be decomposed into a collection of convex polytopes (in short, convexes). When a small number of convexes are used, such a decomposition can be thought of as a piece-wise approximation of the geometry. This decomposition is fundamental in computer graphics, where it provides one…

Cited by 292PDFScholar
2020

Deep Implicit Volume Compression

CVPR 2020oral

We describe a novel approach for compressing truncated signed distance fields (TSDF) stored in 3D voxel grids, and their corresponding textures. To compress the TSDF, our method relies on a block-based neural network architecture trained end-to-end, achieving state-of-the-art rate-distortion trade-o…

Cited by 52PDFcodeScholar
2020

NASA Neural Articulated Shape Approximation

ECCV 2020poster

Efficient representation of articulated objects such as human bodies is an important problem in computer vision and graphics. To efficiently simulate deformation, existing approaches represent 3D objects using polygonal meshes and deform them using skinning techniques. This paper introduces neural a…

Cited by 263SourcePDFScholar
2020

PIE-NET: Parametric Inference of Point Cloud Edges

NeurIPS 2020poster

We introduce an end-to-end learnable technique to robustly identify feature edges in 3D point cloud data. We represent these edges as a collection of parametric curves (i.e.,~lines, circles, and B-splines). Accordingly, our deep neural network, coined PIE-NET, is trained for parametric inference of…

Cited by 125SourcePDFScholar
2020

ShapeFlow: Learnable Deformation Flows Among 3D Shapes

NeurIPS 2020spotlight

We present ShapeFlow, a flow-based model for learning a deformation space for entire classes of 3D shapes with large intra-class variations. ShapeFlow allows learning a multi-template deformation space that is agnostic to shape topology, yet preserves fine geometric details. Different from a generat…

Cited by 101SourcePDFScholar
2019

Linearized Multi-Sampling for Differentiable Image Transformation

ICCV 2019oral

We propose a novel image sampling method for differentiable image transformation in deep neural networks. The sampling schemes currently used in deep learning, such as Spatial Transformer Networks, rely on bilinear interpolation, which performs poorly under severe scale changes, and more importantly…

Cited by 27PDFcodeScholar
2019

Volumetric Capture of Humans With a Single RGBD Camera via Semi-Parametric Learning

CVPR 2019poster

Volumetric (4D) performance capture is fundamental for AR/VR content generation. Whereas previous work in 4D performance capture has shown impressive results in studio settings, the technology is still far from being accessible to a typical consumer who, at best, might own a single RGBD sensor. Thus…

Cited by 47PDFScholar
2018

Espresso: Efficient Forward Propagation for Binary Deep Neural Networks

ICLR 2018poster

There are many applications scenarios for which the computational performance and memory footprint of the prediction phase of Deep Neural Networks (DNNs) need to be optimized. Binary Deep Neural Networks (BDNNs) have been shown to be an effective way of achieving this objective. In this pape…

2017

Low-Dimensionality Calibration Through Local Anisotropic Scaling for Robust Hand Model Personalization

ICCV 2017poster

We present a robust algorithm for personalizing a sphere-mesh tracking model to a user from a collection of depth measurements. Our core contribution is to demonstrate how simple geometric reasoning can be exploited to build a shape-space, and how its performance is comparable to shape-spaces constr…

Cited by 61PDFcodeScholar