← Search

Wei-Chiu Ma

48 accepted papers

2026

PolaRiS: Scalable Real-to-Sim Evaluations for Generalist Robot Policies

RSS 2026poster

A significant challenge for robot learning research is our ability to accurately measure and compare the performance of robot policies. Benchmarking in robotics is historically challenging due to the stochasticity, reproducibility, and time-consuming nature of real-world rollouts. This challenge is …

Cited by 26SourceScholar
2026

SAGE: Scalable Agentic 3D Scene Generation for Embodied AI

CVPR 2026

Real-world data collection for embodied agents remains costly and unsafe, calling for scalable, realistic, and simulator-ready 3D environments. However, existing scene-generation systems often rely on rule-based or task-specific pipelines, yielding artifacts and physically invalid scenes. We present

Cited by 0SourcecodeScholar
2026

X-Diffusion: Training Diffusion Policies on Cross-Embodiment Human Demonstrations

ICRA 2026poster

Human videos are a scalable source of training data for robot learning. However, humans and robots significantly differ in embodiment, making many human actions infeasible for direct execution on a robot. Still, these demonstrations convey rich object-interaction cues and task intent. Our goal is to…

2025

Beyond the Frame: Generating 360deg Panoramic Videos from Perspective Videos

ICCV 2025poster

360deg videos have emerged as a promising medium to represent our dynamic visual world. Compared to the "tunnel vision" of standard cameras, their borderless field of view offers a more complete perspective of our surroundings. While existing video models excel at producing standard videos, their ab…

Cited by 0SourcePDFScholar
2025

Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Model

CVPR 2025poster

Multimodal language models (MLLMs) are increasingly being applied in real-world environments, necessitating their ability to interpret 3D spaces and comprehend temporal dynamics. Current methods often rely on specialized architectural designs or task-specific fine-tuning to achieve this. We introduc…

Cited by 1SourcePDFScholar
2025

DRAWER: Digital Reconstruction and Articulation With Environment Realism

CVPR 2025poster

Creating virtual digital replicas from real-world data unlocks significant potential across domains like gaming and robotics. In this paper, we present DRAWER, a novel framework that converts a video of a static indoor scene into a photorealistic and interactive digital environment. Our approach cen…

2025

Eval3D: Interpretable and Fine-grained Evaluation for 3D Generation

CVPR 2025poster

Despite the unprecedented progress in the field of 3D generation, current systems still often fail to produce high-quality 3D assets that are visually appealing and geometrically and semantically consistent across multiple viewpoints. To effectively assess the quality of the generated 3D data, there…

Cited by 1SourcePDFScholar
2025

Prompting with the Future: Open-World Model Predictive Control with Interactive Digital Twins

RSS 2025poster

Recent advancements in open-world robot manipulation have been largely driven by vision-language models (VLMs). While these models exhibit strong generalization ability in high-level planning, they struggle to predict low-level robot controls due to limited physical-world understanding. To address t…

Cited by 0PDFScholar
2025

Sparse Voxels Rasterization: Real-time High-fidelity Radiance Field Rendering

CVPR 2025poster

We propose an efficient radiance field rendering algorithm that incorporates a rasterization process on adaptive sparse voxels without neural networks or 3D Gaussians. There are two key contributions coupled with the proposed system. The first is to adaptively and explicitly allocate sparse voxels t…

2025

X-Sim: Cross-Embodiment Learning via Real-to-Sim-to-Real

CoRL 2025oral

Human videos offer a scalable way to train robot manipulation policies, but lack the action labels needed by standard imitation learning algorithms. Existing cross-embodiment approaches try to map human motion to robot actions, but often fail when the embodiments differ significantly. We propose X-S…

Cited by 0SourceScholar
2024

BLINK: Multimodal Large Language Models Can See but Not Perceive

ECCV 2024poster

"We introduce , a new benchmark for multimodal language models (LLMs) that focuses on core visual perception abilities not found in other evaluations. Most of the tasks can be solved by humans “within a blink” (, relative depth estimation, visual correspondence, forensics detection, and multi-view r…

2024

ExtraNeRF: Visibility-Aware View Extrapolation of Neural Radiance Fields with Diffusion Models

CVPR 2024poster

We propose ExtraNeRF a novel method for extrapolating the range of views handled by a Neural Radiance Field (NeRF). Our main idea is to leverage NeRFs to model scene-specific fine-grained details while capitalizing on diffusion models to extrapolate beyond our observed data. A key ingredient is to t…

Cited by 3SourcePDFScholar
2024

From an Image to a Scene: Learning to Imagine the World from a Million 360° Videos

NeurIPS 2024poster

Three-dimensional (3D) understanding of objects and scenes play a key role in humans' ability to interact with the world and has been an active area of research in computer vision, graphics, and robotics. Large scale synthetic and object-centric 3D datasets have shown to be effective in training mod…

2024

Multilingual Diversity Improves Vision-Language Representations

NeurIPS 2024spotlight

Massive web-crawled image-text datasets lay the foundation for recent progress in multimodal learning. These datasets are designed with the goal of training a model to do well on standard computer vision benchmarks, many of which, however, have been shown to be English-centric (e.g., ImageNet). Cons…

Cited by 8SourcePDFScholar
2024

Video2Game: Real-time Interactive Realistic and Browser-Compatible Environment from a Single Video

CVPR 2024poster

Creating high-quality and interactive virtual environments such as games and simulators often involves complex and costly manual modeling processes. In this paper we present Video2Game a novel approach that automatically converts videos of real-world scenes into realistic and interactive game enviro…

Cited by 12SourcePDFScholar
2023

Learning Compact Representations for LiDAR Completion and Generation

CVPR 2023poster

LiDAR provides accurate geometric measurements of the 3D world. Unfortunately, dense LiDARs are very expensive and the point clouds captured by low-beam LiDAR are often sparse. To address these issues, we present UltraLiDAR, a data-driven framework for scene-level LiDAR completion, LiDAR generation,…

Cited by 44SourcePDFScholar
2023

Neural Lighting Simulation for Urban Scenes

NeurIPS 2023poster

Different outdoor illumination conditions drastically alter the appearance of urban scenes, and they can harm the performance of image-based robot perception systems if not seen during training. Camera simulation provides a cost-effective solution to create a large dataset of images captured under d…

Cited by 11SourcePDFScholar
2023

Structure from Duplicates: Neural Inverse Graphics from a Pile of Objects

NeurIPS 2023poster

Abstract Our world is full of identical objects (\emph{e.g.}, cans of coke, cars of same model). These duplicates, when seen together, provide additional and strong cues for us to effectively reason about 3D. Inspired by this observation, we introduce Structure from Duplicates (SfD), a novel inverse…

2023

UniSim: A Neural Closed-Loop Sensor Simulator

CVPR 2023highlight

Rigorously testing autonomy systems is essential for making safe self-driving vehicles (SDV) a reality. It requires one to generate safety critical scenarios beyond what can be collected safely in the world, as many scenarios happen rarely on our roads. To accurately evaluate performance, we need to…

Cited by 201SourcePDFScholar
2023

What Happened 3 Seconds Ago? Inferring the Past With Thermal Imaging

CVPR 2023poster

Inferring past human motion from RGB images is challenging due to the inherent uncertainty of the prediction problem. Thermal images, on the other hand, encode traces of past human-object interactions left in the environment via thermal radiation measurement. Based on this observation, we collect th…

2022

CADSim: Robust and Scalable in-the-wild 3D Reconstruction for Controllable Sensor Simulation

CoRL 2022poster

Realistic simulation is key to enabling safe and scalable development of self-driving vehicles. A core component is simulating the sensors so that the entire autonomy system can be tested in simulation. Sensor simulation involves modeling traffic participants, such as vehicles, with high-quality app…

Cited by 27SourceScholar
2022

MIRA: Mental Imagery for Robotic Affordances

CoRL 2022poster

Humans form mental images of 3D scenes to support counterfactual imagination, planning, and motor control. Our abilities to predict the appearance and affordance of the scene from previously unobserved viewpoints aid us in performing manipulation tasks (e.g., 6-DoF kitting) with a level of ease that…

Cited by 33SourceScholar
2022

NeurMiPs: Neural Mixture of Planar Experts for View Synthesis

CVPR 2022poster

We present Neural Mixtures of Planar Experts (NeurMiPs), a novel planar-based scene representation for modeling geometry and appearance. NeurMiPs leverages a collection of local planar experts in 3D space as the scene representation. Each planar expert consists of the parameters of the local rectang…

Cited by 30PDFcodeScholar
2022

SGAM: Building a Virtual 3D World through Simultaneous Generation and Mapping

NeurIPS 2022accept

We present simultaneous generation and mapping (SGAM), a novel 3D scene generation algorithm. Our goal is to produce a realistic, globally consistent 3D world on a large scale. Achieving this goal is challenging and goes beyond the capacities of existing 3D generation or video generation approaches,…

2022

Virtual Correspondence: Humans as a Cue for Extreme-View Geometry

CVPR 2022poster

Recovering the spatial layout of the cameras and the geometry of the scene from extreme-view images is a longstanding challenge in computer vision. Prevailing 3D reconstruction algorithms often adopt the image matching paradigm and presume that a portion of the scene is co-visible across images, yie…

Cited by 27PDFScholar
2021

S3: Neural Shape, Skeleton, and Skinning Fields for 3D Human Modeling

CVPR 2021poster

Constructing and animating humans is an important component for building virtual worlds in a wide variety of applications such as virtual reality or robotics testing in simulation. As there are exponentially many variations of humans with different shape, pose and clothing, it is critical to develop…

Cited by 85PDFScholar
2020

Conditional Entropy Coding for Efficient Video Compression

ECCV 2020poster

We propose a very simple and efficient video compression framework that only focuses on modeling the conditional entropy between frames. Unlike prior learning-based approaches, we reduce complexity by not performing any form of explicit transformations between frames and assume each frame is encoded…

Cited by 73SourcePDFScholar
2020

Deep Feedback Inverse Problem Solver

ECCV 2020poster

We present an efficient, effective, and generic approach towards solving inverse problems. The key idea is to leverage the feedback signal provided by the forward process and learn an iterative update model. Specifically, in each iteration, the neural network takes the feedback as input and outputs…

2020

LiDARsim: Realistic LiDAR Simulation by Leveraging the Real World

CVPR 2020oral

We tackle the problem of producing realistic simulations of LiDAR point clouds, the sensor of preference for most self-driving vehicles. We argue that, by leveraging real data, we can simulate the complex world more realistically compared to employing virtual worlds built from CAD/procedural models.…

Cited by 265PDFScholar
2020

PolyTransform: Deep Polygon Transformer for Instance Segmentation

CVPR 2020poster

In this paper, we propose PolyTransform, a novel instance segmentation algorithm that produces precise, geometry-preserving masks by combining the strengths of prevailing segmentation approaches and modern polygon-based methods. In particular, we first exploit a segmentation network to generate inst…

Cited by 218PDFScholar
2020

Recovering and Simulating Pedestrians in the Wild

CoRL 2020

Sensor simulation is a key component for testing the performance of self-driving vehicles and for data augmentation to better train perception systems. Typical approaches rely on artists to create both 3D assets and their animations to generate a new scenario. This, however, does not scale. In contr

Cited by 0SourcePDFScholar
2020

Weakly-supervised 3D Shape Completion in the Wild

ECCV 2020poster

3D shape completion for real data is important but challenging, since partial point clouds acquired by real-world sensors are usually sparse, noisy and unaligned. Different from previous methods, we address the problem of learning 3D complete shape from unaligned and real-world partial point clouds.…

Cited by 65SourcePDFScholar
2019

Convolutional Recurrent Network for Road Boundary Extraction

CVPR 2019poster

Creating high definition maps that contain precise information of static elements of the scene is of utmost importance for enabling self driving cars to drive safely. In this paper, we tackle the problem of drivable road boundary extraction from LiDAR and camera imagery. Towards this goal, we design…

Cited by 86PDFScholar
2019

DAGMapper: Learning to Map by Discovering Lane Topology

ICCV 2019poster

One of the fundamental challenges to scale self-driving is being able to create accurate high definition maps (HD maps) with low cost. Current attempts to automate this pro- cess typically focus on simple scenarios, estimate independent maps per frame or do not have the level of precision required b…

Cited by 133PDFScholar
2019

DeepPruner: Learning Efficient Stereo Matching via Differentiable PatchMatch

ICCV 2019poster

Our goal is to significantly speed up the runtime of current state-of-the-art stereo algorithms to enable real-time inference. Towards this goal, we developed a differentiable PatchMatch module that allows us to discard most disparities without requiring full cost volume evaluation. We then exploit…

Cited by 310PDFScholar
2019

Exploiting Sparse Semantic HD Maps for Self-Driving Vehicle Localization

IROS 2019poster

In this paper we propose a novel semantic localization algorithm that exploits multiple sensors and has precision on the order of a few centimeters. Our approach does not require detailed knowledge about the appearance of the world, and our maps require orders of magnitude less storage than maps uti…

Cited by 147SourceScholar
2018

Deep Parametric Continuous Convolutional Neural Networks

CVPR 2018poster

Standard convolutional neural networks assume a grid structured input is available and exploit discrete convolutions as their fundamental building blocks. This limits their applicability to many real-world applications. In this paper we propose Parametric Continuous Convolution, a new learnable oper…

Cited by 561SourcePDFScholar
2018

Hierarchical Recurrent Attention Networks for Structured Online Maps

CVPR 2018poster

In this paper, we tackle the problem of online road network extraction from sparse 3D point clouds. Our method is inspired by how an annotator builds a lane graph, by first identifying how many lanes there are and then drawing each one in turn. We develop a hierarchical recurrent network that atten…

Cited by 77SourcePDFScholar
2018

Single Image Intrinsic Decomposition without a Single Intrinsic Image

ECCV 2018poster

Intrinsic image decomposition---decomposing a natural image into a set of images corresponding to different physical causes---is one of the key and fundamental problems of computer vision. Previous intrinsic decomposition approaches either address the problem in a fully supervised manner, or require…

Cited by 81SourcePDFScholar
2018

SurfConv: Bridging 3D and 2D Convolution for RGBD Images

CVPR 2018poster

The last few years have seen approaches trying to combine the increasing popularity of depth sensors and the success of the convolutional neural networks. Using depth as additional channel alongside the RGB input has the scale variance problem present in image convolution based approaches. On the ot…

2017

Find your way by observing the sun and other semantic cues

ICRA 2017poster

In this paper we present a robust, efficient and affordable approach to self-localization which requires neither GPS nor knowledge about the appearance of the world. Towards this goal, we utilize freely available cartographic maps and derive a probabilistic model that exploits semantic cues in the f…

Cited by 60SourceScholar
2017

Forecasting Interactive Dynamics of Pedestrians With Fictitious Play

CVPR 2017poster

We develop predictive models of pedestrian dynamics by encoding the coupled nature of multi-pedestrian interaction using game theory and deep learning-based visual analysis to estimate person-specific behavior parameters. We focus on predictive models since they are important for developing interact…

Cited by 215PDFScholar
2015

How Do We Use Our Hands? Discovering a Diverse Set of Common Grasps

CVPR 2015poster

Our aim is to show how state-of-the-art computer vision techniques can be used to advance prehensile analysis (i.e., understanding the functionality of human hands). Prehensile analysis is a broad field of multi-disciplinary interest, where researchers painstakingly manually analyze hours of hand-ob…

Cited by 76SourcePDFScholar