← Search

Gengshan Yang

18 accepted papers

2025

Agent-to-Sim: Learning Interactive Behavior Models from Casual Longitudinal Videos

ICLR 2025poster

We present Agent-to-Sim (ATS), a framework for learning interactive behavior models of 3D agents from casual longitudinal video collections. Different from prior works that rely on marker-based tracking and multiview cameras, ATS learns natural behaviors of animal agents non-invasively through video…

2024

SplaTAM: Splat Track & Map 3D Gaussians for Dense RGB-D SLAM

CVPR 2024poster

Dense simultaneous localization and mapping (SLAM) is crucial for robotics and augmented reality applications. However current methods are often hampered by the non-volumetric or implicit way they represent a scene. This work introduces SplaTAM an approach that for the first time leverages explicit…

2024

Tactile DreamFusion: Exploiting Tactile Sensing for 3D Generation

NeurIPS 2024poster

3D generation methods have shown visually compelling results powered by diffusion image priors. However, they often fail to produce realistic geometric details, resulting in overly smooth surfaces or geometric details inaccurately baked in albedo maps. To address this, we introduce a new method that…

2023

PPR: Physically Plausible Reconstruction from Monocular Videos

ICCV 2023oral

Given monocular videos, we build 3D models of articulated objects and environments whose 3D configurations satisfy dynamics and contact constraints. At its core, our method leverages differentiable physics simulation to aid visual reconstructions. We couple differentiable physics simulation with dif…

Cited by 31PDFcodeScholar
2023

Reconstructing Animatable Categories From Videos

CVPR 2023poster

Building animatable 3D models is challenging due to the need for 3D scans, laborious registration, and manual rigging. Recently, differentiable rendering provides a pathway to obtain high-quality 3D models from monocular videos, but these are limited to rigid categories or single instances. We prese…

2023

SLoMo: A General System for Legged Robot Motion Imitation From Casual Videos

RA-L 2023

We present SLoMo: a first-of-its-kind framework for transferring skilled motions from casually captured “in-the-wild” video footage of humans and animals to legged robots. SLoMo works in three stages: 1) synthesize a physically plausible reconstructed key-point trajectory from monocular videos; 2) o

Cited by 29SourcecodeScholar
2023

Total-Recon: Deformable Scene Reconstruction for Embodied View Synthesis

ICCV 2023poster

We explore the task of embodied view synthesis from monocular videos of deformable scenes. Given a minute-long RGBD video of people interacting with their pets, we render the scene from novel camera trajectories derived from the in-scene motion of actors: (1) egocentric cameras that simulate the poi…

Cited by 20PDFcodeScholar
2022

BANMo: Building Animatable 3D Neural Models From Many Casual Videos

CVPR 2022oral

Prior work for articulated 3D shape reconstruction often relies on specialized multi-view and depth sensors or pre-built deformable 3D models. Such methods do not scale to diverse sets of objects in the wild. We present a method that requires neither of them. It builds high-fidelity, articulated 3D…

Cited by 202PDFcodeScholar
2021

LASR: Learning Articulated Shape Reconstruction From a Monocular Video

CVPR 2021poster

Remarkable progress has been made in 3D reconstruction of rigid structures from a video or a collection of images. However, it is still challenging to reconstruct nonrigid structures from RGB inputs, due to the under-constrained nature of this problem. While template-based approaches, such as parame…

Cited by 129PDFcodeScholar
2021

NeRS: Neural Reflectance Surfaces for Sparse-view 3D Reconstruction in the Wild

NeurIPS 2021poster

Recent history has seen a tremendous growth of work exploring implicit representations of geometry and radiance, popularized through Neural Radiance Fields (NeRF). Such works are fundamentally based on a (implicit) {\em volumetric} representation of occupancy, allowing them to model diverse scene s…

2021

ViSER: Video-Specific Surface Embeddings for Articulated 3D Shape Reconstruction

NeurIPS 2021spotlight

We introduce ViSER, a method for recovering articulated 3D shapes and dense3D trajectories from monocular videos. Previous work on high-quality reconstruction of dynamic 3D shapes typically relies on multiple camera views, strong category-specific priors, or 2D keypoint supervision. We show that no…