← Search

Angela Dai

68 accepted papers

2026

Animating the Uncaptured: Humanoid Mesh Animation with Video Diffusion Models

ICLR 2026poster

Animation of humanoid characters is essential in various graphics applications, but require significant time and cost to create realistic animations. We propose an approach to synthesize 4D animated sequences of input static 3D humanoid meshes, leveraging strong generalized motion priors from genera…

Cited by 0SourceScholar
2025

DiffuMatch: Category-Agnostic Spectral Diffusion Priors for Robust Non-rigid Shape Matching

ICCV 2025poster

Deep functional maps have recently emerged as a powerful tool for solving non-rigid shape correspondence tasks. Methods that use this approach combine the power and flexibility of the functional map framework, with data-driven learning for improved accuracy and generality. However, most existing met…

2025

ExCap3D: Expressive 3D Scene Understanding via Object Captioning with Varying Detail

ICCV 2025poster

Generating text descriptions of objects in 3D indoor scenes is an important building block of embodied understanding. Existing methods do this by describing objects at a single level of detail, which often does not capture fine-grained details such as varying textures, materials, and shapes of the p…

Cited by 0SourcePDFScholar
2025

GaussianSpeech: Audio-Driven Personalized 3D Gaussian Avatars

ICCV 2025poster

We introduce GaussianSpeech, a novel approach that synthesizes high-fidelity animation sequences of photorealistic and personalized multi-view consistent 3D human head avatars from spoken audio at real-time rendering rates. To capture the expressive and detailed nature of human heads, including skin…

2025

MeshArt: Generating Articulated Meshes with Structure-Guided Transformers

CVPR 2025poster

Articulated 3D object generation is fundamental for creating realistic, functional, and interactable virtual assets which are not simply static. We introduce MeshArt, a hierarchical transformer-based approach to generate articulated 3D meshes with clean, compact geometry, reminiscent of human-crafte…

Cited by 1SourcePDFScholar
2025

MeshPad: Interactive Sketch-Conditioned Artist-Reminiscent Mesh Generation and Editing

ICCV 2025poster

We introduce MeshPad, a generative approach that creates 3D meshes from sketch inputs. Building on recent advances in artist-reminiscent triangle mesh generation, our approach addresses the need for interactive mesh creation. To this end, we focus on enabling consistent edits by decomposing editing…

Cited by 0SourcePDFScholar
2025

Non-Gaited Legged Locomotion With Monte-Carlo Tree Search and Supervised Learning

RA-L 2025

Legged robots are able to navigate complex terrains by continuously interacting with the environment through careful selection of contact sequences and timings. However, the combinatorial nature behind contact planning hinders the applicability of such optimization problems on hardware. In this work

Cited by 8SourceScholar
2025

Open-YOLO 3D: Towards Fast and Accurate Open-Vocabulary 3D Instance Segmentation

ICLR 2025oral

Recent works on open-vocabulary 3D instance segmentation show strong promise but at the cost of slow inference speed and high computation requirements. This high computation cost is typically due to their heavy reliance on aggregated clip features from multi-view, which require computationally expen…

2025

PrEditor3D: Fast and Precise 3D Shape Editing

CVPR 2025poster

We propose a training-free approach to 3D editing that enables the editing of a single shape and the reconstruction of a mesh within a few minutes. Leveraging 4-view images, user-guided text prompts, and rough 2D masks, our method produces an edited 3D mesh that aligns with the prompt. For this, our…

Cited by 3SourcePDFScholar
2025

QuickSplat: Fast 3D Surface Reconstruction via Learned Gaussian Initialization

ICCV 2025poster

Surface reconstruction is fundamental to computer vision and graphics, enabling applications in 3D modeling, mixed reality, robotics, and more. Existing approaches based on volumetric rendering obtain promising results, but optimize on a per-scene basis, resulting in a slow optimization that can str…

Cited by 0SourcePDFScholar
2025

SceneFactor: Factored Latent 3D Diffusion for Controllable 3D Scene Generation

CVPR 2025poster

We present SceneFactor, a diffusion-based approach for large-scale 3D scene generation that enables controllable generation and effortless editing. SceneFactor enables text-guided 3D scene synthesis through our factored diffusion formulation, leveraging latent semantic and geometric manifolds for ge…

Cited by 6SourcePDFScholar
2024

Coherent 3D Scene Diffusion From a Single RGB Image

NeurIPS 2024poster

We present a novel diffusion-based approach for coherent 3D scene reconstruction from a single RGB image. Our method utilizes an image-conditioned 3D scene diffusion model to simultaneously denoise the 3D poses and geometries of all objects within the scene. Motivated by the ill-posed nature of th…

Cited by 1SourcePDFScholar
2024

DPHMs: Diffusion Parametric Head Models for Depth-based Tracking

CVPR 2024poster

We introduce Diffusion Parametric Head Models (DPHMs) a generative model that enables robust volumetric head reconstruction and tracking from monocular depth sequences. While recent volumetric head models such as NPHMs can now excel in representing high-fidelity head geometries tracking and reconstr…

Cited by 6SourcePDFScholar
2024

DiffuScene: Denoising Diffusion Models for Generative Indoor Scene Synthesis

CVPR 2024poster

We present DiffuScene for indoor 3D scene synthesis based on a novel scene configuration denoising diffusion model. It generates 3D instance properties stored in an unordered object set and retrieves the most similar geometry for each object configuration which is characterized as a concatenation of…

Cited by 99SourcePDFScholar
2024

DrivAerNet++: A Large-Scale Multimodal Car Dataset with Computational Fluid Dynamics Simulations and Deep Learning Benchmarks

NeurIPS 2024poster

We present DrivAerNet++, the largest and most comprehensive multimodal dataset for aerodynamic car design. DrivAerNet++ comprises 8,000 diverse car designs modeled with high-fidelity computational fluid dynamics (CFD) simulations. The dataset includes diverse car configurations such as fastback, not…

2024

FaceTalk: Audio-Driven Motion Diffusion for Neural Parametric Head Models

CVPR 2024poster

We introduce FaceTalk a novel generative approach designed for synthesizing high-fidelity 3D motion sequences of talking human heads from input audio signal. To capture the expressive detailed nature of human heads including hair ears and finer-scale eye movements we propose to couple speech signal…

2024

FutureHuman3D: Forecasting Complex Long-Term 3D Human Behavior from Video Observations

CVPR 2024poster

We present a generative approach to forecast long-term future human behavior in 3D requiring only weak supervision from readily available 2D human action data. This is a fundamental task enabling many downstream applications. The required ground-truth data is hard to capture in 3D (mocap suits expen…

Cited by 3SourcePDFScholar
2024

MeshGPT: Generating Triangle Meshes with Decoder-Only Transformers

CVPR 2024highlight

We introduce MeshGPT a new approach for generating triangle meshes that reflects the compactness typical of artist-created meshes in contrast to dense triangle meshes extracted by iso-surfacing methods from neural fields. Inspired by recent advances in powerful large language models we adopt a seque…

Cited by 124SourcePDFScholar
2024

PaSCo: Urban 3D Panoptic Scene Completion with Uncertainty Awareness

CVPR 2024poster

We propose the task of Panoptic Scene Completion (PSC) which extends the recently popular Semantic Scene Completion (SSC) task with instance-level information to produce a richer understanding of the 3D scene. Our PSC proposal utilizes a hybrid mask-based technique on the nonempty voxels from sparse…

2023

HyperDiffusion: Generating Implicit Neural Fields with Weight-Space Diffusion

ICCV 2023poster

Implicit neural fields, typically encoded by a multilayer perceptron (MLP) that maps from coordinates (e.g., xyz) to signals (e.g., signed distances), have shown remarkable promise as a high-fidelity and compact representation. However, the lack of a regular and explicit grid structure also makes it…

Cited by 128PDFcodeScholar
2023

Mask3D: Pre-Training 2D Vision Transformers by Learning Masked 3D Priors

CVPR 2023poster

Current popular backbones in computer vision, such as Vision Transformers (ViT) and ResNets are trained to perceive the world from 2D images. However, to more effectively understand 3D structural priors in 2D backbones, we propose Mask3D to leverage existing large-scale RGB-D data in a self-supervis…

Cited by 16SourcePDFScholar
2023

Neural Part Priors: Learning To Optimize Part-Based Object Completion in RGB-D Scans

CVPR 2023highlight

3D scene understanding has seen significant advances in recent years, but has largely focused on object understanding in 3D scenes with independent per-object predictions. We thus propose to learn Neural Part Priors (NPPs), parametric spaces of objects and their parts, that enable optimizing to fit…

Cited by 8SourcePDFScholar
2023

ObjectMatch: Robust Registration Using Canonical Object Correspondences

CVPR 2023poster

We present ObjectMatch, a semantic and object-centric camera pose estimator for RGB-D SLAM pipelines. Modern camera pose estimators rely on direct correspondences of overlapping regions between frames; however, they cannot align camera frames with little or no overlap. In this work, we propose to le…

Cited by 10SourcePDFScholar
2023

Panoptic Lifting for 3D Scene Understanding With Neural Fields

CVPR 2023highlight

We propose Panoptic Lifting, a novel approach for learning panoptic 3D volumetric representations from images of in-the-wild scenes. Once trained, our model can render color images together with 3D-consistent panoptic segmentation from novel viewpoints. Unlike existing approaches which use 3D input…

Cited by 134SourcePDFScholar
2022

4DContrast: Contrastive Learning with Dynamic Correspondences for 3D Scene Understanding

ECCV 2022poster

"We present a new approach to instill 4D dynamic object priors into learned 3D representations by unsupervised pre-training. We observe that dynamic movement of an object through an environment provides important cues about its objectness, and thus propose to imbue learned 3D representations with su…

Cited by 67SourcePDFScholar
2022

Language-Grounded Indoor 3D Semantic Segmentation in the Wild

ECCV 2022poster

"Recent advances in 3D semantic segmentation with deep neural networks have shown remarkable success, with rapid performance increase on available datasets. However, current 3D semantic segmentation benchmarks contain only a small number of categories -- less than 30 for ScanNet and SemanticKITTI, f…

2022

PatchComplete: Learning Multi-Resolution Patch Priors for 3D Shape Completion on Unseen Categories

NeurIPS 2022accept

While 3D shape representations enable powerful reasoning in many visual and perception applications, learning 3D shape priors tends to be constrained to the specific categories trained on, leading to an inefficient learning process, particularly for general applications with unseen categories. Thus…

2022

Pose2Room: Understanding 3D Scenes from Human Activities

ECCV 2022poster

"With wearable IMU sensors, one can estimate human poses from wearable devices without requiring visual input. In this work, we pose the question: Can we reason about object structure in real-world environments solely from human trajectory information? Crucially, we observe that human motion and int…

Cited by 17SourcePDFScholar
2022

Texturify: Generating Textures on 3D Shape Surfaces

ECCV 2022poster

"Texture cues on 3D objects are key to compelling visual representations, with the possibility to create high visual fidelity with inherent spatial consistency across different views. Since the availability of textured 3D shapes remains very limited, learning a 3D-supervised data-driven method that…

Cited by 73SourcePDFScholar
2021

NPMs: Neural Parametric Models for 3D Deformable Shapes

ICCV 2021poster

Parametric 3D models have enabled a wide variety of tasks in computer graphics and vision, such as modeling human bodies, faces, and hands. However, the construction of these parametric models is often tedious, as it requires heavy manual tweaking, and they struggle to represent additional complexit…

Cited by 118PDFcodeScholar
2021

Neural Deformation Graphs for Globally-Consistent Non-Rigid Reconstruction

CVPR 2021poster

We introduce Neural Deformation Graphs for globally-consistent deformation tracking and 3D reconstruction of non-rigid objects. Specifically, we implicitly model a deformation graph via a deep neural network. This neural deformation graph does not rely on any object-specific structure and, thus, can…

Cited by 83PDFcodeScholar
2021

Panoptic 3D Scene Reconstruction From a Single RGB Image

NeurIPS 2021poster

Richly segmented 3D scene reconstructions are an integral basis for many high-level scene understanding tasks, such as for robotics, motion planning, or augmented reality. Existing works in 3D perception from a single RGB image tend to focus on geometric reconstruction only, or geometric reconstruc…

2021

Patch2CAD: Patchwise Embedding Learning for In-the-Wild Shape Retrieval From a Single Image

ICCV 2021poster

3D perception of object shapes from RGB image input is fundamental towards semantic scene understanding, grounding image-based perception in our spatially 3-dimensional real-world environments. To achieve a mapping between image views of objects and 3D shapes, we leverage CAD model priors from exist…

Cited by 36PDFScholar
2021

Pri3D: Can 3D Priors Help 2D Representation Learning?

ICCV 2021poster

Recent advances in 3D perception have shown impressive progress in understanding geometric structures of 3D shapes and even scenes. Inspired by these advances in geometric understanding, we aim to imbue image-based perception with representations learned under geometric constraints. We introduce an…

Cited by 88PDFcodeScholar
2021

RetrievalFuse: Neural 3D Scene Reconstruction With a Database

ICCV 2021poster

3D reconstruction of large scenes is a challenging problem due to the high-complexity nature of the solution space, in particular for generative neural networks. In contrast to traditional generative learned models which encode the full generative process into a neural network and can struggle with…

Cited by 38PDFcodeScholar
2021

SPSG: Self-Supervised Photometric Scene Generation From RGB-D Scans

CVPR 2021poster

We present SPSG, a novel approach to generate high-quality, colored 3D models of scenes from RGB-D scan observations by learning to infer unobserved scene geometry and color in a self-supervised fashion. Our self-supervised approach learns to jointly inpaint geometry and color by correlating an inco…

Cited by 42PDFcodeScholar
2021

Seeing Behind Objects for 3D Multi-Object Tracking in RGB-D Sequences

CVPR 2021poster

Multi-object tracking from RGB-D video sequences is a challenging problem due to the combination of changing viewpoints, motion, and occlusions over time. We observe that having the complete geometry of objects aids in their tracking, and thus propose to jointly infer the complete geometry of object…

Cited by 28PDFScholar
2021

Towards Part-Based Understanding of RGB-D Scans

CVPR 2021poster

Recent advances in 3D semantic scene understanding have shown impressive progress in 3D instance segmentation, enabling object-level reasoning about 3D scenes; however, a finer-grained understanding is required to enable interactions with objects and their functional understanding. Thus, we propose…

Cited by 12PDFScholar
2021

TransformerFusion: Monocular RGB Scene Reconstruction using Transformers

NeurIPS 2021poster

We introduce TransformerFusion, a transformer-based 3D scene reconstruction approach. From an input monocular RGB video, the video frames are processed by a transformer network that fuses the observations into a volumetric feature grid representing the scene; this feature grid is then decoded into a…

Cited by 158SourcePDFScholar
2020

Adversarial Texture Optimization From RGB-D Scans

CVPR 2020poster

Realistic color texture generation is an important step in RGB-D surface reconstruction, but remains challenging in practice due to inaccuracies in reconstructed geometry, misaligned camera poses, and view-dependent imaging artifacts. In this work, we present a novel approach for color texture gener…

Cited by 59PDFcodeScholar
2020

Mask2CAD: 3D Shape Prediction by Learning to Segment and Retrieve

ECCV 2020poster

Object recognition has seen significant progress in the image domain, with focus primarily on 2D perception. We propose to leverage existing large-scale datasets of 3D models to understand the underlying 3D structure of objects seen in an image by constructing a CAD-based representation of the objec…

Cited by 95SourcePDFScholar
2020

Neural Non-Rigid Tracking

NeurIPS 2020poster

We introduce a novel, end-to-end learnable, differentiable non-rigid tracker that enables state-of-the-art non-rigid reconstruction by a learned robust optimization. Given two input RGB-D frames of a non-rigidly moving object, we employ a convolutional neural network to predict dense correspondences…

2020

SG-NN: Sparse Generative Neural Networks for Self-Supervised Scene Completion of RGB-D Scans

CVPR 2020poster

We present a novel approach that converts partial and noisy RGB-D scans into high-quality 3D scene reconstructions by inferring unobserved scene geometry. Our approach is fully self-supervised and can hence be trained solely on incomplete, real-world scans. To achieve, self-supervision, we remove fr…

Cited by 176PDFcodeScholar
2020

SceneCAD: Predicting Object Alignments and Layouts in RGB-D Scans

ECCV 2020poster

We present a novel approach to reconstructing lightweight, CAD-based representations of scanned 3D environments from commodity RGB-D sensors. Our key idea is to jointly optimize for both CAD model alignments as well as layout estimations of the scanned scene, explicitly modeling inter-relationships…

Cited by 73SourcePDFScholar
2019

Scan2CAD: Learning CAD Model Alignment in RGB-D Scans

CVPR 2019oral

We present Scan2CAD, a novel data-driven method that learns to align clean 3D CAD models from a shape database to the noisy and incomplete geometry of a commodity RGB-D scan. For a 3D reconstruction of an indoor scene, our method takes as input a set of CAD models, and predicts a 9DoF pose that alig…

Cited by 294PDFScholar
2018

ScanComplete: Large-Scale Scene Completion and Semantic Segmentation for 3D Scans

CVPR 2018poster

We introduce ScanComplete, a novel data-driven approach for taking an incomplete 3D scan of a scene as input and predicting a complete 3D model along with per-voxel semantic labels. The key contribution of our method is its ability to handle large scenes with varying spatial extent, managing the cub…

Cited by 375SourcePDFScholar
2017

ScanNet: Richly-Annotated 3D Reconstructions of Indoor Scenes

CVPR 2017spotlight

A key requirement for leveraging supervised deep learning methods is the availability of large, labeled datasets. Unfortunately, in the context of RGB-D scene understanding, very little data is available -- current datasets cover a small range of scene views and have limited semantic annotations.…

Cited by 5003PDFScholar
2017

Shape Completion Using 3D-Encoder-Predictor CNNs and Shape Synthesis

CVPR 2017spotlight

We introduce a data-driven approach to complete partial 3D shapes through a combination of volumetric deep neural networks and 3D shape synthesis. From a partially-scanned input shape, our method first infers a low-resolution -- but complete -- output. To this end, we introduce a 3D-Encoder-Predicto…

Cited by 804PDFScholar
2016

Volumetric and Multi-View CNNs for Object Classification on 3D Data

CVPR 2016spotlight

3D shape models are becoming widely available and easier to capture, making available 3D information crucial for progress in object classification. Current state-of-the-art methods rely on CNNs to address this problem. Recently, we witness two types of CNNs being developed: CNNs based upon volumetri…

Cited by 2061PDFScholar