← Search

Kaichun Mo

30 accepted papers

2026

DreamDojo: A Real-Time Robot World Model from Large-Scale Human Videos

ICML 2026spotlight

Being able to simulate the outcomes of actions in varied environments will revolutionize the development of generalist agents at scale. However, modeling these world dynamics, especially for dexterous robotics tasks, poses significant challenges due to limited data coverage and scarce action labels.…

Cited by 81SourceScholar
2026

PointWorld: Scaling 3D World Models for In-The-Wild Robotic Manipulation

CVPR 2026

Humans anticipate, from a glance and a contemplated action of their bodies, how the 3D world will respond, a capability that is equally vital for robotic manipulation. We introduce PointWorld, a large pre-trained 3D world model that unifies state and action in a shared 3D space as 3D point flows: gi

Cited by 0SourcecodeScholar
2025

3D-MVP: 3D Multiview Pretraining for Manipulation

CVPR 2025poster

Recent works have shown that visual pretraining on egocentric datasets using masked autoencoders (MAE) can improve generalization for downstream robotics tasks. However, these approaches pretrain only on 2D images, while many robotics applications require 3D scene understanding. In this work, we pro…

Cited by 0SourcePDFScholar
2025

MatchMaker: Automated Asset Generation for Robotic Assembly

ICRA 2025

Robotic assembly remains a significant challenge due to complexities in visual perception, functional grasping, contact-rich manipulation, and performing high-precision tasks. Simulation-based learning and sim-to-real transfer have led to recent success in solving assembly tasks in the presence of o

Cited by 3SourcecodeScholar
2024

Category-Level Multi-Part Multi-Joint 3D Shape Assembly

CVPR 2024poster

Shape assembly composes complex shapes geometries by arranging simple part geometries and has wide applications in autonomous robotic assembly and CAD modeling. Existing works focus on geometry reasoning and neglect the actual physical assembly process of matching and fitting joints which are the co…

Cited by 15SourcePDFScholar
2024

Summarize the Past to Predict the Future: Natural Language Descriptions of Context Boost Multimodal Object Interaction Anticipation

CVPR 2024poster

We study object interaction anticipation in egocentric videos. This task requires an understanding of the spatio-temporal context formed by past actions on objects coined "action context". We propose TransFusion a multimodal transformer-based architecture for short-term object interaction anticipati…

Cited by 5SourcePDFScholar
2024

URDFormer: A Pipeline for Constructing Articulated Simulation Environments from Real-World Images

RSS 2024poster

Constructing accurate and targeted simulation scenes that are both visually and physically realistic is a problem of significant practical interest in domains ranging from robotics to computer vision. This problem has become even more relevant as researchers wielding large data-hungry learning metho…

Cited by 21SourcePDFScholar
2023

COPILOT: Human-Environment Collision Prediction and Localization from Egocentric Videos

ICCV 2023poster

The ability to forecast human-environment collisions from egocentric observations is vital to enable collision avoidance in applications such as VR, AR, and wearable assistive robotics. In this work, we introduce the challenging problem of predicting collisions in diverse environments from multi-vie…

Cited by 3PDFcodeScholar
2023

DualAfford: Learning Collaborative Visual Affordance for Dual-gripper Manipulation

ICLR 2023poster

It is essential yet challenging for future home-assistant robots to understand and manipulate diverse 3D objects in daily human environments. Towards building scalable systems that can perform diverse manipulation tasks over various 3D shapes, recent works have advocated and demonstrated promising r…

Cited by 16SourcePDFScholar
2023

JacobiNeRF: NeRF Shaping With Mutual Information Gradients

CVPR 2023poster

We propose a method that trains a neural radiance field (NeRF) to encode not only the appearance of the scene but also semantic correlations between scene points, regions, or entities -- aiming to capture their mutual co-variation patterns. In contrast to the traditional first-order photometric reco…

2023

STOW: Discrete-Frame Segmentation and Tracking of Unseen Objects for Warehouse Picking Robots

CoRL 2023poster

Segmentation and tracking of unseen object instances in discrete frames pose a significant challenge in dynamic industrial robotic contexts, such as distribution warehouses. Here, robots must handle object rearrangements, including shifting, removal, and partial occlusion by new items, and track the…

Cited by 6SourceScholar
2023

Towards Learning Geometric Eigen-Lengths Crucial for Fitting Tasks

ICML 2023poster

Some extremely low-dimensional yet crucial geometric eigen-lengths often determine the success of some geometric tasks. For example, the *height* of an object is important to measure to check if it can fit between the shelves of a cabinet, while the *width* of a couch is crucial when trying to move…

Cited by 5SourcePDFScholar
2023

Where2Explore: Few-shot Affordance Learning for Unseen Novel Categories of Articulated Objects

NeurIPS 2023poster

Articulated object manipulation is a fundamental yet challenging task in robotics. Due to significant geometric and semantic variations across object categories, previous manipulation models struggle to generalize to novel categories. Few-shot learning is a promising solution for alleviating this is…

Cited by 39SourcePDFScholar
2022

AdaAfford: Learning to Adapt Manipulation Affordance for 3D Articulated Objects via Few-Shot Interactions

ECCV 2022poster

"Perceiving and interacting with 3D articulated objects, such as cabinets, doors, and faucets, pose particular challenges for future home-assistant robots performing daily tasks in human environments. Besides parsing the articulated parts and joint parameters, researchers recently advocate learning…

Cited by 68SourcePDFScholar
2022

Fixing Malfunctional Objects With Learned Physical Simulation and Functional Prediction

CVPR 2022poster

This paper studies the problem of fixing malfunctional 3D objects. While previous works focus on building passive perception models to learn the functionality from static 3D objects, we argue that functionality is reckoned with respect to the physical interactions between the object and the user. Gi…

Cited by 6PDFScholar
2022

GIMO: Gaze-Informed Human Motion Prediction in Context

ECCV 2022poster

"Predicting human motion is critical for assistive robots and AR/VR applications, where the interaction with humans needs to be safe and comfortable. Meanwhile, an accurate prediction depends on understanding both the scene context and human intentions. Even though many works study scene-aware human…

2022

IFR-Explore: Learning Inter-object Functional Relationships in 3D Indoor Scenes

ICLR 2022poster

Building embodied intelligent agents that can interact with 3D indoor environments has received increasing research attention in recent years. While most works focus on single-object or agent-object visual functionality and affordances, our work proposes to study a novel, underexplored, kind of visu…

Cited by 7SourcePDFScholar
2022

Object Pursuit: Building a Space of Objects via Discriminative Weight Generation

ICLR 2022poster

We propose a framework to continuously learn object-centric representations for visual learning and understanding. Existing object-centric representations either rely on supervisions that individualize objects in the scene, or perform unsupervised disentanglement that can hardly deal with complex sc…

2022

VAT-Mart: Learning Visual Action Trajectory Proposals for Manipulating 3D ARTiculated Objects

ICLR 2022poster

Perceiving and manipulating 3D articulated objects (e.g., cabinets, doors) in human environments is an important yet challenging task for future home-assistant robots. The space of 3D articulated objects is exceptionally rich in their myriad semantic categories, diverse shape geometry, and complicat…

Cited by 104SourcePDFScholar
2021

O2O-Afford: Annotation-Free Large-Scale Object-Object Affordance Learning

CoRL 2021poster

Contrary to the vast literature in modeling, perceiving, and understanding agent-object (e.g., human-object, hand-object, robot-object) interaction in computer vision and robotics, very few past works have studied the task of object-object interaction, which also plays an important role in robotic m…

Cited by 73SourceScholar
2021

Where2Act: From Pixels to Actions for Articulated 3D Objects

ICCV 2021poster

One of the fundamental goals of visual perception is to allow agents to meaningfully interact with their environment. In this paper, we take a step towards that long-term goal -- we extract highly localized actionable information related to elementary actions such as pushing or pulling for articulat…

Cited by 202PDFcodeScholar
2020

Generative 3D Part Assembly via Dynamic Graph Learning

NeurIPS 2020poster

Autonomous part assembly is a challenging yet crucial task in 3D computer vision and robotics. Analogous to buying an IKEA furniture, given a set of 3D parts that can assemble a single shape, an intelligent agent needs to perceive the 3D part geometry, reason to propose pose estimations for the inpu…

Cited by 100SourcePDFScholar
2020

Learning 3D Part Assembly from a Single Image

ECCV 2020poster

Autonomous assembly is a crucial capability for robots in many applications. For this task, several problems such as obstacle avoidance, motion planning, and actuator control have been extensively studied in robotics. However, when it comes to task specification, the space of possibilities remains u…

2020

Learning to Group: A Bottom-Up Framework for 3D Part Discovery in Unseen Categories

ICLR 2020poster

We address the problem of learning to discover 3D parts for objects in unseen categories. Being able to learn the geometry prior of parts and transfer this prior to unseen categories pose fundamental challenges on data-driven shape segmentation approaches. Formulated as a contextual bandit problem,…

Cited by 45SourcecodeScholar
2020

PT2PC: Learning to Generate 3D Point Cloud Shapes from Part Tree Conditions

ECCV 2020poster

Generative 3D shape modeling is a fundamental research area in computer vision and interactive computer graphics, with many real-world applications. This paper investigates the novel problem of generating a 3D point cloud geometry for a shape from a symbolic part tree representation. In order to lea…

Cited by 49SourcePDFScholar
2020

SAPIEN: A SimulAted Part-Based Interactive ENvironment

CVPR 2020oral

Building home assistant robots has long been a goal for vision and robotics researchers. To achieve this task, a simulated environment with physically realistic simulation, sufficient articulated objects, and transferability to the real robot is indispensable. Existing environments achieve these req…

Cited by 560PDFcodeScholar
2019

PartNet: A Large-Scale Benchmark for Fine-Grained and Hierarchical Part-Level 3D Object Understanding

CVPR 2019poster

We present PartNet: a consistent, large-scale dataset of 3D objects annotated with fine-grained, instance-level, and hierarchical 3D part information. Our dataset consists of 573,585 part instances over 26,671 3D models covering 24 object categories. This dataset enables and serves as a catalyst for…

Cited by 849PDFScholar
2017

PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation

CVPR 2017oral

Point cloud is an important type of geometric data structure. Due to its irregular format, most researchers transform such data to regular 3D voxel grids or collections of images. This, however, renders data unnecessarily voluminous and causes issues. In this paper, we design a novel type of neural…

Cited by 19941PDFScholar