← Search

Bokui Shen

13 accepted papers

2025

Make a Donut: Hierarchical EMD-Space Planning for Zero-Shot Deformable Manipulation With Tools

RA-L 2025

Deformable object manipulation stands as one of the most captivating yet formidable challenges in robotics. While previous techniques have predominantly relied on learning latent dynamics through demonstrations, typically represented as either particles or images, there exists a pertinent limitation

Cited by 4SourceScholar
2024

CAD: Photorealistic 3D Generation via Adversarial Distillation

CVPR 2024poster

The increased demand for 3D data in AR/VR robotics and gaming applications gave rise to powerful generative pipelines capable of synthesizing high-quality 3D objects. Most of these models rely on the Score Distillation Sampling (SDS) algorithm to optimize a 3D representation such that the rendered i…

Cited by 13SourcePDFScholar
2024

MultiPhys: Multi-Person Physics-aware 3D Motion Estimation

CVPR 2024poster

We introduce MultiPhys a method designed for recovering multi-person motion from monocular videos. Our focus lies in capturing coherent spatial placement between pairs of individuals across varying degrees of engagement. MultiPhys being physically aware exhibits robustness to jittering and occlusion…

Cited by 5SourcePDFScholar
2024

SAGE: Bridging Semantic and Actionable Parts for GEneralizable Articulated-Object Manipulation under Language Instructions

RSS 2024poster

To interact with daily-life articulated objects of diverse structures and functionalities, understanding the object parts plays a central role in both user instruction comprehension and task execution. However, the possible discordance between the semantic meaning and physics functionalities of the…

Cited by 0SourcePDFScholar
2023

Banana: Banach Fixed-Point Network for Pointcloud Segmentation with Inter-Part Equivariance

NeurIPS 2023spotlight

Equivariance has gained strong interest as a desirable network property that inherently ensures robust generalization. However, when dealing with complex systems such as articulated objects or multi-object scenes, effectively capturing inter-part transformations poses a challenge, as it becomes enta…

Cited by 15SourcePDFScholar
2023

COPILOT: Human-Environment Collision Prediction and Localization from Egocentric Videos

ICCV 2023poster

The ability to forecast human-environment collisions from egocentric observations is vital to enable collision avoidance in applications such as VR, AR, and wearable assistive robotics. In this work, we introduce the challenging problem of predicting collisions in diverse environments from multi-vie…

Cited by 3PDFcodeScholar
2023

GINA-3D: Learning To Generate Implicit Neural Assets in the Wild

CVPR 2023poster

Modeling the 3D world from sensor data for simulation is a scalable way of developing testing and validation environments for robotic learning problems such as autonomous driving. However, manually creating or re-creating real-world-like environments is difficult, expensive, and not scalable. Recent…

Cited by 21SourcePDFScholar
2023

NAP: Neural 3D Articulated Object Prior

NeurIPS 2023poster

We propose Neural 3D Articulated object Prior (NAP), the first 3D deep generative model to synthesize 3D articulated object models. Despite the extensive research on generating 3D static objects, compositions, or scenes, there are hardly any approaches on capturing the distribution of articulated ob…

Cited by 16SourcePDFScholar
2023

PointOdyssey: A Large-Scale Synthetic Dataset for Long-Term Point Tracking

ICCV 2023oral

We introduce PointOdyssey, a large-scale synthetic dataset, and data generation framework, for the training and evaluation of long-term fine-grained tracking algorithms. Our goal is to advance the state-of-the-art by placing emphasis on long videos with naturalistic motion. Toward the goal of natura…

Cited by 143PDFcodeScholar
2022

ACID: Action-Conditional Implicit Visual Dynamics for Deformable Object Manipulation

RSS 2022poster

Manipulating volumetric deformable objects in the real world, like plush toys and pizza dough, bring substantial challenges due to infinite shape variations, non-rigid motions, and partial observability. We introduce ACID, an action-conditional visual dynamics model for volumetric deformable objects…

Cited by 41SourcePDFScholar
2022

ADeLA: Automatic Dense Labeling With Attention for Viewpoint Shift in Semantic Segmentation

CVPR 2022oral

We describe a method to deal with performance drop in semantic segmentation caused by viewpoint changes within multi-camera systems, where temporally paired images are readily available, but the annotations may only be abundant for a few typical views. Existing methods alleviate performance drop via…

Cited by 6PDFScholar
2021

iGibson 1.0: A Simulation Environment for Interactive Tasks in Large Realistic Scenes

IROS 2021poster

We present iGibson 1.0, a novel simulation environment to develop robotic solutions for interactive tasks in large-scale realistic scenes. Our environment contains 15 fully interactive home-sized scenes with 108 rooms populated with rigid and articulated objects. The scenes are replicas of real-worl…

Cited by 193SourceScholar
2021

iGibson 2.0: Object-Centric Simulation for Robot Learning of Everyday Household Tasks

CoRL 2021poster

Recent research in embodied AI has been boosted by the use of simulation environments to develop and train robot learning approaches. However, the use of simulation has skewed the attention to tasks that only require what robotics simulators can simulate: motion and physical contact. We present iGib…

Cited by 268SourceScholar