← Search

Or Litany

42 accepted papers

2026

DiffusionHarmonizer: Bridging Neural Reconstruction and Photorealistic Simulation with Online Diffusion Enhancer

CVPR 2026

Simulation is essential to the development and evaluation of autonomous robots such as self-driving vehicles. Neural reconstruction is emerging as a promising solution as it enables simulating a wide variety of scenarios from real-world data alone in an automated and scalable way. However, while met

Cited by 0SourcecodeScholar
2026

Realiz3D: 3D Generation Made Photorealistic via Domain-Aware Learning

CVPR 2026

We often aim to generate images that are both photorealistic and 3D-consistent, adhering to precise geometry, material, and viewpoint controls. Typically, this is achieved by fine-tuning an image generator, pre-trained on billions of real images, using renders of synthetic 3D assets, where annotatio

Cited by 0SourceScholar
2026

SpaceControl: Introducing Test-Time Spatial Control to 3D Generative Modeling

ICLR 2026poster

Generative methods for 3D assets have recently achieved remarkable progress, yet providing intuitive and precise control over the object geometry remains a key challenge. Existing approaches predominantly rely on text or image prompts, which often fall short in geometric specificity: language can be…

Cited by 0SourceScholar
2026

Time-to-Move: Training-Free Motion-Controlled Video Generation via Dual-Clock Denoising

ICLR 2026poster

Diffusion-based video generation can create realistic videos, yet existing image and text-based conditioning fails to offer precise motion control. Prior methods for motion control typically rely on displacement-based conditioning and require model-specific fine-tuning, which is computationally expe…

Cited by 0SourcecodeScholar
2025

A Lesson in Splats: Teacher-Guided Diffusion for 3D Gaussian Splats Generation with 2D Supervision

ICCV 2025poster

We present a novel framework for training 3D image-conditioned diffusion models using only 2D supervision. Recovering 3D structure from 2D images is inherently ill-posed due to the ambiguity of possible reconstructions, making generative models a natural choice. However, most existing 3D generative…

Cited by 0SourcePDFScholar
2025

MonSTeR: a Unified Model for Motion, Scene, Text Retrieval

ICCV 2025poster

Intention drives human movement in complex environments, but such movement can only happen if the surrounding context supports it. Despite the intuitive nature of this mechanism, existing research has not yet provided tools to evaluate the alignment between skeletal movement (motion), intention (tex…

2025

OmniRe: Omni Urban Scene Reconstruction

ICLR 2025spotlight

We introduce OmniRe, a comprehensive system for efficiently creating high-fidelity digital twins of dynamic real-world scenes from on-device logs. Recent methods using neural fields or Gaussian Splatting primarily focus on vehicles, hindering a holistic framework for all dynamic foregrounds demanded…

2024

3DiffTection: 3D Object Detection with Geometry-Aware Diffusion Features

CVPR 2024poster

3DiffTection introduces a novel method for 3D object detection from single images utilizing a 3D-aware diffusion model for feature extraction. Addressing the resource-intensive nature of annotating large-scale 3D image data our approach leverages pretrained diffusion models traditionally used for 2D…

Cited by 13SourcePDFScholar
2024

Dynamic LiDAR Re-simulation using Compositional Neural Fields

CVPR 2024highlight

We introduce DyNFL a novel neural field-based approach for high-fidelity re-simulation of LiDAR scans in dynamic driving scenes. DyNFL processes LiDAR measurements from dynamic environments accompanied by bounding boxes of moving objects to construct an editable neural field. This field comprising s…

2024

EmerNeRF: Emergent Spatial-Temporal Scene Decomposition via Self-Supervision

ICLR 2024poster

We present EmerNeRF, a simple yet powerful approach for learning spatial-temporal representations of dynamic driving scenes. Grounded in neural fields, EmerNeRF simultaneously captures scene geometry, appearance, motion, and semantics via self-bootstrapping. EmerNeRF hinges upon two core components:…

2024

The Nerfect Match: Exploring NeRF Features for Visual Localization

ECCV 2024poster

"In this work, we propose the use of Neural Radiance Fields () as a scene representation for visual localization. Recently, has been employed to enhance pose regression and scene coordinate regression models by augmenting the training database, providing auxiliary supervision through rendered images…

Cited by 9SourcePDFScholar
2024

Zero-to-Hero: Enhancing Zero-Shot Novel View Synthesis via Attention Map Filtering

NeurIPS 2024poster

Generating realistic images from arbitrary views based on a single source image remains a significant challenge in computer vision, with broad applications ranging from e-commerce to immersive virtual experiences. Recent advancements in diffusion models, particularly the Zero-1-to-3 model, have been…

Cited by 2SourcePDFScholar
2023

Fast Monocular Scene Reconstruction With Global-Sparse Local-Dense Grids

CVPR 2023poster

Indoor scene reconstruction from monocular images has long been sought after by augmented reality and robotics developers. Recent advances in neural field representations and monocular priors have led to remarkable results in scene-level surface reconstructions. The reliance on Multilayer Perceptron…

Cited by 9SourcePDFScholar
2023

Mask3D: Mask Transformer for 3D Semantic Instance Segmentation

ICRA 2023poster

Modern 3D semantic instance segmentation approaches predominantly rely on specialized voting mechanisms followed by carefully designed geometric clustering techniques. Building on the successes of recent Transformer-based methods for object detection and image segmentation, we propose the first Tran…

Cited by 272SourcecodeScholar
2023

Neural Kernel Surface Reconstruction

CVPR 2023highlight

We present a novel method for reconstructing a 3D implicit surface from a large-scale, sparse, and noisy point cloud. Our approach builds upon the recently introduced Neural Kernel Fields (NKF) representation. It enjoys similar generalization capabilities to NKF, while simultaneously addressing its…

Cited by 87SourcePDFScholar
2023

Neural LiDAR Fields for Novel View Synthesis

ICCV 2023poster

We present Neural Fields for LiDAR (NFL), a method to optimise a neural field scene representation from LiDAR measurements, with the goal of synthesizing realistic LiDAR scans from novel viewpoints. NFL combines the rendering power of neural fields with a detailed, physically motivated model of the…

Cited by 61PDFScholar
2023

Towards Viewpoint Robustness in Bird's Eye View Segmentation

ICCV 2023poster

Autonomous vehicles (AV) require that neural networks used for perception be robust to different viewpoints if they are to be deployed across many types of vehicles without the repeated cost of data collection and labeling for each. AV companies typically focus on collecting data from diverse scenar…

Cited by 15PDFScholar
2023

Trace and Pace: Controllable Pedestrian Animation via Guided Trajectory Diffusion

CVPR 2023poster

We introduce a method for generating realistic pedestrian trajectories and full-body animations that can be controlled to meet user-defined goals. We draw on recent advances in guided diffusion modeling to achieve test-time controllability of trajectories, which is normally only associated with rule…

Cited by 118SourcePDFScholar
2022

GET3D: A Generative Model of High Quality 3D Textured Shapes Learned from Images

NeurIPS 2022accept

As several industries are moving towards modeling massive 3D virtual worlds, the need for content creation tools that can scale in terms of the quantity, quality, and diversity of 3D content is becoming evident. In our work, we aim to train performant 3D generative models that synthesize textured me…

2022

Generating Useful Accident-Prone Driving Scenarios via a Learned Traffic Prior

CVPR 2022poster

Evaluating and improving planning for autonomous vehicles requires scalable generation of long-tail traffic scenarios. To be useful, these scenarios must be realistic and challenging, but not impossible to drive through safely. In this work, we introduce STRIVE, a method to automatically generate ch…

Cited by 157PDFScholar
2022

LION: Latent Point Diffusion Models for 3D Shape Generation

NeurIPS 2022accept

Denoising diffusion models (DDMs) have shown promising results in 3D point cloud synthesis. To advance 3D DDMs and make them useful for digital artists, we require (i) high generation quality, (ii) flexibility for manipulation and applications such as conditional synthesis and shape interpolation, a…

2022

Language-Grounded Indoor 3D Semantic Segmentation in the Wild

ECCV 2022poster

"Recent advances in 3D semantic segmentation with deep neural networks have shown remarkable success, with rapid performance increase on available datasets. However, current 3D semantic segmentation benchmarks contain only a small number of categories -- less than 30 for ScanNet and SemanticKITTI, f…

2022

MvDeCor: Multi-View Dense Correspondence Learning for Fine-Grained 3D Segmentation

ECCV 2022poster

"We propose to utilize self-supervised techniques in the 2D domain for fine-grained 3D shape segmentation tasks. This is inspired by the observation that view-based surface representations are more effective at modeling high-resolution surface details and texture than their 3D counterparts based on…

Cited by 13SourcePDFScholar
2022

Neural Fields As Learnable Kernels for 3D Reconstruction

CVPR 2022poster

We present Neural Kernel Fields: a novel method for reconstructing implicit 3D shapes based on a learned kernel ridge regression. Our technique achieves state-of-the-art results when reconstructing 3D objects and large scenes from sparse oriented points, and can reconstruct shape categories outside…

Cited by 82PDFScholar
2021

3DIoUMatch: Leveraging IoU Prediction for Semi-Supervised 3D Object Detection

CVPR 2021poster

3D object detection is an important yet demanding task that heavily relies on difficult to obtain 3D annotations. To reduce the required amount of supervision, we propose 3DIoUMatch, a novel semi-supervised method for 3D object detection applicable to both indoor and outdoor scenes. We leverage a te…

Cited by 155PDFcodeScholar
2021

DIB-R++: Learning to Predict Lighting and Material with a Hybrid Differentiable Renderer

NeurIPS 2021poster

We consider the challenging problem of predicting intrinsic object properties from a single image by exploiting differentiable renderers. Many previous learning-based approaches for inverse graphics adopt rasterization-based renderers and assume naive lighting and material models, which often fail t…

Cited by 68SourcePDFScholar
2021

ReLMoGen: Integrating Motion Generation in Reinforcement Learning for Mobile Manipulation

ICRA 2021poster

Many Reinforcement Learning (RL) approaches use joint control signals (positions, velocities, torques) as action space for continuous control tasks. We propose to lift the action space to a higher level in the form of subgoals for a motion generator (a combination of motion planner and trajectory ex…

Cited by 81SourceScholar
2021

Vector Neurons: A General Framework for SO(3)-Equivariant Networks

ICCV 2021poster

Invariance and equivariance to the rotation group have been widely discussed in the 3D deep learning community for pointclouds. Yet most proposed methods either use complex mathematical tools that may limit their accessibility, or are tied to specific input data types and network architectures. In t…

Cited by 345PDFcodeScholar
2021

Weakly Supervised Learning of Rigid 3D Scene Flow

CVPR 2021poster

We propose a data-driven scene flow estimation algorithm exploiting the observation that many 3D scenes can be explained by a collection of agents moving as rigid bodies. At the core of our method lies a deep architecture able to reason at the object-level by considering 3D scene flow in conjunction…

Cited by 114PDFcodeScholar
2020

ImVoteNet: Boosting 3D Object Detection in Point Clouds With Image Votes

CVPR 2020poster

3D object detection has seen quick progress thanks to advances in deep learning on point clouds. A few recent works have even shown state-of-the-art performance with just point clouds input (e.g. VoteNet). However, point cloud data have inherent limitations. They are sparse, lack color information a…

Cited by 348PDFcodeScholar
2020

PointContrast: Unsupervised Pre-training for 3D Point Cloud Understanding

ECCV 2020poster

Arguably one of the top success stories of deep learning is transfer learning. The finding that pre-training a network on a rich source set (g, ImageNet) can help boost performance once fine-tuned on a usually much smaller target set, has been instrumental to many applications in language and vision…

2020

Towards Precise Completion of Deformable Shapes

ECCV 2020poster

According to Aristotle, a philosopher in Ancient Greece, {\it ``the whole is greater than the sum of its parts''}. This statement was adopted to explain human perception by the Gestalt psychology school of thought in the twentieth century. Here, we claim that observing part of an object which was pr…

2019

Deep Hough Voting for 3D Object Detection in Point Clouds

ICCV 2019oral

Current 3D object detection methods are heavily influenced by 2D detectors. In order to leverage architectures in 2D detectors, they often convert 3D point clouds to regular grids (i.e., to voxel grids or to bird's eye view images), or rely on detection in 2D images to propose 3D boxes. Few works ha…

Cited by 1587PDFcodeScholar
2018

Deformable Shape Completion With Graph Convolutional Autoencoders

CVPR 2018poster

The availability of affordable and portable depth sensors has made scanning objects and people simpler than ever. However, dealing with occlusions and missing parts is still a significant challenge. The problem of reconstructing a (possibly non-rigidly moving) 3D object from a single or multiple par…

Cited by 290SourcePDFScholar
2017

Deep Functional Maps: Structured Prediction for Dense Shape Correspondence

ICCV 2017poster

We introduce a new framework for learning dense correspondence between deformable 3D shapes. Existing learning based approaches model shape correspondence as a labelling problem, where each point of a query shape receives a label identifying a point on some reference domain; the correspondence is th…

Cited by 345PDFcodeScholar