← Search

Shenlong Wang

70 accepted papers

2026

Generalizable Sparse-View 3D Reconstruction from Unconstrained Images

CVPR 2026

Reconstructing 3D scenes from sparse, unposed images remains challenging under real-world conditions with varying illumination and transient occlusions. Existing methods rely on scene-specific optimization with appearance embeddings or dynamic masks, requiring extensive per-scene training and failin

Cited by 0SourcecodeScholar
2026

MUSE: Multimodal Uncertainty Quantification of State Estimation

ICRA 2026poster

Accurate visual state estimation has been a central topic in robotics with a wide range of applications in robot navigation, autonomous driving, and autonomous flight. Recent advances in robot perception have led to significant improvements in the accuracy and robustness of state estimation, yet a f…

2026

SAGE: Scalable Agentic 3D Scene Generation for Embodied AI

CVPR 2026

Real-world data collection for embodied agents remains costly and unsafe, calling for scalable, realistic, and simulator-ready 3D environments. However, existing scene-generation systems often rely on rule-based or task-specific pipelines, yielding artifacts and physically invalid scenes. We present

Cited by 0SourcecodeScholar
2025

AD-GS: Object-Aware B-Spline Gaussian Splatting for Self-Supervised Autonomous Driving

ICCV 2025poster

Modeling and rendering dynamic urban driving scenes is crucial for self-driving simulation. Current high-quality methods typically rely on costly manual object tracklet annotations, while self-supervised approaches fail to capture dynamic object motions accurately and decompose scenes properly, resu…

Cited by 0SourcePDFScholar
2025

Controllable Weather Synthesis and Removal with Video Diffusion Models

ICCV 2025poster

Generating realistic and controllable weather effects in videos is valuable for many applications. Physics-based weather simulation requires precise reconstructions that are hard to scale to in-the-wild videos, while current video editing often lacks realism and control.In this work, we introduce We…

Cited by 0SourcePDFScholar
2025

DRAWER: Digital Reconstruction and Articulation With Environment Realism

CVPR 2025poster

Creating virtual digital replicas from real-world data unlocks significant potential across domains like gaming and robotics. In this paper, we present DRAWER, a novel framework that converts a video of a static indoor scene into a photorealistic and interactive digital environment. Our approach cen…

2025

Demeter: A Parametric Model of Crop Plant Morphology from the Real World

ICCV 2025poster

Learning 3D parametric shape models of objects has gained popularity in vision and graphics and has showed broad utility in 3D reconstruction, generation, understanding, and simulation. While powerful models exist for humans and animals, equally expressive approaches for modeling plants are lacking.…

Cited by 0SourcePDFScholar
2025

HoloScene: Simulation‑Ready Interactive 3D Worlds from a Single Video

NeurIPS 2025poster

Digitizing the physical world into accurate simulation‑ready virtual environments offers significant opportunities in a variety of fields such as augmented and virtual reality, gaming, and robotics. However, current 3D reconstruction and scene-understanding methods commonly fall short in one or more…

Cited by 0SourceScholar
2025

Human-like Navigation in a World Built for Humans

CoRL 2025poster

When navigating in a man-made environment they haven’t visited before—like an office building—humans employ behaviors such as reading signs and asking others for directions. These behaviors help humans reach their destinations efficiently by reducing the need to search through large areas. Existing…

Cited by 0SourceScholar
2025

IRIS: Inverse Rendering of Indoor Scenes from Low Dynamic Range Images

CVPR 2025poster

Inverse rendering seeks to recover 3D geometry, surface material, and lighting from captured images, enabling advanced applications such as novel-view synthesis, relighting, and virtual object insertion. However, most existing techniques rely on high dynamic range (HDR) images as input, limiting acc…

Cited by 4SourcePDFScholar
2025

InvRGB+L: Inverse Rendering of Complex Scenes with Unified Color and LiDAR Reflectance Modeling

ICCV 2025poster

We present InvRGB+L, a novel inverse rendering model that reconstructs large, relightable, and dynamic scenes from a single RGB+LiDAR sequence. Conventional inverse graphics methods rely primarily on RGB observations and use LiDAR mainly for geometric information, often resulting in suboptimal mater…

Cited by 0SourcePDFScholar
2025

LIFe-GoM: Generalizable Human Rendering with Learned Iterative Feedback Over Multi-Resolution Gaussians-on-Mesh

ICLR 2025poster

Generalizable rendering of an animatable human avatar from sparse inputs relies on data priors and inductive biases extracted from training on large data to avoid scene-specific optimization and to enable fast reconstruction. This raises two main challenges: First, unlike iterative gradient-based ad…

Cited by 0SourcePDFScholar
2025

Map Space Belief Prediction for Manipulation-Enhanced Mapping

RSS 2025poster

Searching for objects in cluttered environments requires selecting efficient viewpoints and manipulation actions to remove occlusions and reduce uncertainty in object locations, shapes, and categories. In this work, we address the problem of manipulation-enhanced semantic mapping, where a robot has…

Cited by 2PDFScholar
2025

NoPo-Avatar: Generalizable and Animatable Avatars from Sparse Inputs without Human Poses

NeurIPS 2025poster

We tackle the task of recovering an animatable 3D human avatar from a single or a sparse set of images. For this task, beyond a set of images, many prior state-of-the-art methods use accurate “ground-truth” camera poses and human poses as input to guide reconstruction at test-time. We show that pose…

Cited by 0SourceScholar
2025

PhysGen3D: Crafting a Miniature Interactive World from a Single Image

CVPR 2025poster

Envisioning physically plausible outcomes from a single image requires a deep understanding of the world's dynamics. To address this, we introduce MiniTwin, a novel framework that transforms a single image into an amodal, camera-centric, interactive 3D scene. By combining advanced image-based geomet…

Cited by 3SourcePDFScholar
2025

PhysTwin: Physics-Informed Reconstruction and Simulation of Deformable Objects from Videos

ICCV 2025poster

Creating a physical digital twin of a real-world object has immense potential in robotics, content creation, and XR. In this paper, we present PhysTwin, a novel framework that uses sparse videos of dynamic objects in interaction to produce a photo- and physically realistic, real-time interactive vir…

2025

Uni4D: Unifying Visual Foundation Models for 4D Modeling from a Single Video

CVPR 2025highlight

This paper presents a unified approach to understanding dynamic scenes from casual videos. Large pretrained vision foundation models, such as vision-language, video depth prediction, motion tracking, and segmentation models, offer promising capabilities. However, training a single model for comprehe…

2025

Visual Sync: Multi‑Camera Synchronization via Cross‑View Object Motion

NeurIPS 2025poster

Today, people can easily record memorable moments, ranging from concerts, sports events, lectures, family gatherings, and birthday parties with multiple consumer cameras. However, synchronizing these cross‑camera streams remains challenging. Existing methods assume controlled settings, specific targ…

Cited by 0SourceScholar
2025

VoxelSplat: Dynamic Gaussian Splatting as an Effective Loss for Occupancy and Flow Prediction

CVPR 2025poster

Recent advancements in camera-based occupancy prediction have focused on the simultaneous prediction of 3D semantics and scene flow, a task that presents significant challenges due to specific difficulties, e.g., occlusions and unbalanced dynamic environments. In this paper, we analyze these challen…

2024

DiffTune-MPC: Closed-Loop Learning for Model Predictive Control

RA-L 2024

Model predictive control (MPC) has been applied to many platforms in robotics and autonomous systems for its capability to predict a system's future behavior while incorporating constraints that a system may have. To enhance the performance of a system with an MPC controller, one can manually tune t

Cited by 25SourceScholar
2024

GoMAvatar: Efficient Animatable Human Modeling from Monocular Video Using Gaussians-on-Mesh

CVPR 2024poster

We introduce GoMAvatar a novel approach for real-time memory-efficient high-quality animatable human modeling. GoMAvatar takes as input a single monocular video to create a digital avatar capable of re-articulation in new poses and real-time rendering from novel viewpoints while seamlessly integrati…

Cited by 34SourcePDFScholar
2024

On the Overconfidence Problem in Semantic 3D Mapping

ICRA 2024poster

Semantic 3D mapping, the process of fusing depth and image segmentation information between multiple views to build 3D maps annotated with object classes in real-time, is a recent topic of interest. This paper highlights the fusion overconfidence problem, in which conventional mapping methods assign…

Cited by 6SourcecodeScholar
2024

Physical Property Understanding from Language-Embedded Feature Fields

CVPR 2024poster

Can computers perceive the physical properties of objects solely through vision? Research in cognitive science and vision science has shown that humans excel at identifying materials and estimating their physical properties based purely on visual appearance. In this paper we present a novel approach…

Cited by 12SourcePDFScholar
2024

RoboEXP: Action-Conditioned Scene Graph via Interactive Exploration for Robotic Manipulation

CoRL 2024poster

We introduce the novel task of interactive scene exploration, wherein robots autonomously explore environments and produce an action-conditioned scene graph (ACSG) that captures the structure of the underlying environment. The ACSG accounts for both low-level information (geometry and semantics) and…

Cited by 21SourcecodeScholar
2024

SuperGaussian: Repurposing Video Models for 3D Super Resolution

ECCV 2024poster

"We present a simple, modular, and generic method that upsamples coarse 3D models by adding geometric and appearance details. While generative 3D models now exist, they do not yet match the quality of their counterparts in image and video domains. We demonstrate that it is possible to directly repur…

2024

Video2Game: Real-time Interactive Realistic and Browser-Compatible Environment from a Single Video

CVPR 2024poster

Creating high-quality and interactive virtual environments such as games and simulators often involves complex and costly manual modeling processes. In this paper we present Video2Game a novel approach that automatically converts videos of real-world scenes into realistic and interactive game enviro…

Cited by 12SourcePDFScholar
2023

Building Rearticulable Models for Arbitrary 3D Objects From 4D Point Clouds

CVPR 2023poster

We build rearticulable models for arbitrary everyday man-made objects containing an arbitrary number of parts that are connected together in arbitrary ways via 1-degree-of-freedom joints. Given point cloud videos of such everyday objects, our method identifies the distinct object parts, what parts a…

2023

ClimateNeRF: Extreme Weather Synthesis in Neural Radiance Field

ICCV 2023poster

Physical simulations produce excellent predictions of weather effects. Neural radiance fields produce SOTA scene models. We describe a novel NeRF-editing procedure that can fuse physical simulations with NeRF models of scenes, producing realistic movies of physical phenomena in those scenes. Our app…

Cited by 32PDFScholar
2023

ContactGen: Generative Contact Modeling for Grasp Generation

ICCV 2023poster

This paper presents a novel object-centric contact representation ContactGen for hand-object interaction. The ContactGen comprises 3 components: a contact map indicates the contact location, a part map represents the contact hand part, and a direction map tells the contact direction within each part…

Cited by 30PDFcodeScholar
2023

MapPrior: Bird's-Eye View Map Layout Estimation with Generative Models

ICCV 2023poster

Despite tremendous advancements in bird's-eye view (BEV) perception, existing models fall short in generating realistic and coherent semantic map layouts, and they fail to account for uncertainties arising from partial sensor information (such as occlusion or limited coverage). In this work, we intr…

Cited by 14PDFScholar
2023

Sim-on-Wheels: Physical World in the Loop Simulation for Self-Driving

RA-L 2023

We present Sim-on-Wheels, a safe, realistic, and vehicle-in-loop framework to test autonomous vehicles' performance in the real world under safety-critical scenarios. Sim-on-wheels runs on a self-driving vehicle operating in the physical world. It creates virtual traffic participants with risky beha

Cited by 11SourcecodeScholar
2023

Structure from Duplicates: Neural Inverse Graphics from a Pile of Objects

NeurIPS 2023poster

Abstract Our world is full of identical objects (\emph{e.g.}, cans of coke, cars of same model). These duplicates, when seen together, provide additional and strong cues for us to effectively reason about 3D. Inspired by this observation, we introduce Structure from Duplicates (SfD), a novel inverse…

2022

CASA: Category-agnostic Skeletal Animal Reconstruction

NeurIPS 2022accept

Recovering a skeletal shape from a monocular video is a longstanding challenge. Prevailing nonrigid animal reconstruction methods often adopt a control-point driven animation model and optimize bone transforms individually without considering skeletal topology, yielding unsatisfactory shape and arti…

Cited by 33SourcePDFScholar
2022

NeurMiPs: Neural Mixture of Planar Experts for View Synthesis

CVPR 2022poster

We present Neural Mixtures of Planar Experts (NeurMiPs), a novel planar-based scene representation for modeling geometry and appearance. NeurMiPs leverages a collection of local planar experts in 3D space as the scene representation. Each planar expert consists of the parameters of the local rectang…

Cited by 30PDFcodeScholar
2022

SGAM: Building a Virtual 3D World through Simultaneous Generation and Mapping

NeurIPS 2022accept

We present simultaneous generation and mapping (SGAM), a novel 3D scene generation algorithm. Our goal is to produce a realistic, globally consistent 3D world on a large scale. Achieving this goal is challenging and goes beyond the capacities of existing 3D generation or video generation approaches,…

2022

Virtual Correspondence: Humans as a Cue for Extreme-View Geometry

CVPR 2022poster

Recovering the spatial layout of the cameras and the geometry of the scene from extreme-view images is a longstanding challenge in computer vision. Prevailing 3D reconstruction algorithms often adopt the image matching paradigm and presume that a portion of the scene is co-visible across images, yie…

Cited by 27PDFScholar
2021

GeoSim: Realistic Video Simulation via Geometry-Aware Composition for Self-Driving

CVPR 2021poster

Scalable sensor simulation is an important yet challenging open problem for safety-critical domains such as self-driving. Current works in image simulation either fail to be photorealistic or do not model the 3D environment and the dynamic objects within, losing high-level control and physical reali…

Cited by 106PDFScholar
2021

S3: Neural Shape, Skeleton, and Skinning Fields for 3D Human Modeling

CVPR 2021poster

Constructing and animating humans is an important component for building virtual worlds in a wide variety of applications such as virtual reality or robotics testing in simulation. As there are exponentially many variations of humans with different shape, pose and clothing, it is critical to develop…

Cited by 85PDFScholar
2021

SceneGen: Learning To Generate Realistic Traffic Scenes

CVPR 2021poster

We consider the problem of generating realistic traffic scenes automatically. Existing methods typically insert actors into the scene according to a set of hand-crafted heuristics and are limited in their ability to model the true complexity and diversity of real traffic scenes, thus inducing a cont…

Cited by 118PDFScholar
2020

Conditional Entropy Coding for Efficient Video Compression

ECCV 2020poster

We propose a very simple and efficient video compression framework that only focuses on modeling the conditional entropy between frames. Unlike prior learning-based approaches, we reduce complexity by not performing any form of explicit transformations between frames and assume each frame is encoded…

Cited by 73SourcePDFScholar
2020

DSDNet: Deep Structured self-Driving Network

ECCV 2020poster

In this paper, we propose the Deep Structured self-Driving Network (DSDNet), which performs object detection, motion prediction, and motion planning with a single neural network. Towards this goal, we develop a deep structured energy based model which considers the interactions between actors and pr…

Cited by 118SourcePDFScholar
2020

Deep Feedback Inverse Problem Solver

ECCV 2020poster

We present an efficient, effective, and generic approach towards solving inverse problems. The key idea is to leverage the feedback signal provided by the forward process and learn an iterative update model. Specifically, in each iteration, the neural network takes the feedback as input and outputs…

2020

LiDARsim: Realistic LiDAR Simulation by Leveraging the Real World

CVPR 2020oral

We tackle the problem of producing realistic simulations of LiDAR point clouds, the sensor of preference for most self-driving vehicles. We argue that, by leveraging real data, we can simulate the complex world more realistically compared to employing virtual worlds built from CAD/procedural models.…

Cited by 265PDFScholar
2020

MuSCLE: Multi Sweep Compression of LiDAR using Deep Entropy Models

NeurIPS 2020poster

We present a novel compression algorithm for reducing the storage of LiDAR sensory data streams. Our model exploits spatio-temporal relationships across multiple LIDAR sweeps to reduce the bitrate of both geometry and intensity values. Towards this goal, we propose a novel conditional entropy model…

2020

OctSqueeze: Octree-Structured Entropy Model for LiDAR Compression

CVPR 2020oral

We present a novel deep compression algorithm to reduce the memory footprint of LiDAR point clouds. Our method exploits the sparsity and structural redundancy between points to reduce the bitrate. Towards this goal, we first encode the point cloud into an octree, a data-efficient structure suitable…

Cited by 216PDFScholar
2020

Pit30M: A Benchmark for Global Localization in the Age of Self-Driving Cars

IROS 2020poster

We are interested in understanding whether retrieval-based localization approaches are good enough in the context of self-driving vehicles. Towards this goal, we introduce Pit30M, a new image and LiDAR dataset with over 30 million frames, which is 10 to 100 times larger than those used in previous w…

Cited by 15SourcecodeScholar
2019

Convolutional Recurrent Network for Road Boundary Extraction

CVPR 2019poster

Creating high definition maps that contain precise information of static elements of the scene is of utmost importance for enabling self driving cars to drive safely. In this paper, we tackle the problem of drivable road boundary extraction from LiDAR and camera imagery. Towards this goal, we design…

Cited by 86PDFScholar
2019

DeepPruner: Learning Efficient Stereo Matching via Differentiable PatchMatch

ICCV 2019poster

Our goal is to significantly speed up the runtime of current state-of-the-art stereo algorithms to enable real-time inference. Towards this goal, we developed a differentiable PatchMatch module that allows us to discard most disparities without requiring full cost volume evaluation. We then exploit…

Cited by 342PDFScholar
2019

Efficient Graph Generation with Graph Recurrent Attention Networks

NeurIPS 2019poster

We propose a new family of efficient and expressive deep generative models of graphs, called Graph Recurrent Attention Networks (GRANs). Our model generates graphs one block of nodes and associated edges at a time. The block size and sampling stride allow us to trade off sample quality for efficienc…

2019

Exploiting Sparse Semantic HD Maps for Self-Driving Vehicle Localization

IROS 2019poster

In this paper we propose a novel semantic localization algorithm that exploits multiple sensors and has precision on the order of a few centimeters. Our approach does not require detailed knowledge about the appearance of the world, and our maps require orders of magnitude less storage than maps uti…

Cited by 147SourceScholar
2019

Learning to Localize Through Compressed Binary Maps

CVPR 2019poster

One of the main difficulties of scaling current localization systems to large environments is the on-board storage required for the maps. In this paper we propose to learn to compress the map representation such that it is optimal for the localization task. As a consequence, higher compression rates…

Cited by 39PDFScholar
2018

Deep Continuous Fusion for Multi-Sensor 3D Object Detection

ECCV 2018poster

In this paper, we propose a novel 3D object detector that can exploit both LIDAR as well as cameras to perform very accurate localization. Towards this goal, we design an end-to-end learnable architecture that exploits continuous convolutions to fuse image and LIDAR feature maps at different levels…

Cited by 1168SourcePDFScholar
2018

Deep Multi-Sensor Lane Detection

IROS 2018poster

Reliable and accurate lane detection has been a long-standing problem in the field of autonomous driving. In recent years, many approaches have been developed that use images (or videos) as input and reason in image space. In this paper we argue that accurate image estimates do not translate to prec…

Cited by 108SourceScholar
2018

Deep Parametric Continuous Convolutional Neural Networks

CVPR 2018poster

Standard convolutional neural networks assume a grid structured input is available and exploit discrete convolutions as their fundamental building blocks. This limits their applicability to many real-world applications. In this paper we propose Parametric Continuous Convolution, a new learnable oper…

Cited by 561SourcePDFScholar
2017

Find your way by observing the sun and other semantic cues

ICRA 2017poster

In this paper we present a robust, efficient and affordable approach to self-localization which requires neither GPS nor knowledge about the appearance of the world. Towards this goal, we utilize freely available cartographic maps and derive a probabilistic model that exploits semantic cues in the f…

Cited by 60SourceScholar
2017

TorontoCity: Seeing the World With a Million Eyes

ICCV 2017spotlight

In this paper we introduce the TorontoCity benchmark, which covers the full greater Toronto area (GTA) with 712.5km2 of land, 8439km of road and around 400, 000 buildings. Our benchmark provides different perspectives of the world captured from airplanes, drones and cars driving around the city. Man…

Cited by 217PDFScholar
2016

HD Maps: Fine-Grained Road Segmentation by Parsing Ground and Aerial Images

CVPR 2016poster

In this paper we present an approach to enhance existing maps with fine grained segmentation categories such as parking spots and sidewalk, as well as the number and location of road lanes. Towards this goal, we propose an efficient approach that is able to estimate these fine grained categories by…

Cited by 181PDFScholar