← Search

Danfei Xu

67 accepted papers

2026

Compositional Diffusion with Guided search for Long-Horizon Planning

ICLR 2026oral

Generative models have emerged as powerful tools for planning, with compositional approaches offering particular promise for modeling long-horizon task distributions by composing together local, modular generative models. This compositional paradigm spans diverse domains, from multi-step manipulatio…

Cited by 0SourcecodeScholar
2026

Compositional Visual Planning via Inference-Time Diffusion Scaling

ICLR 2026poster

Diffusion models excel at short-horizon robot planning, yet scaling them to long-horizon tasks remains challenging due to computational constraints and limited training data. Existing compositional approaches stitch together short segments by separately denoising each component and averaging overla…

Cited by 0SourcecodeScholar
2026

Counterfactual VLA: Self-Reflective Vision-Language-Action Model with Adaptive Reasoning

CVPR 2026

Recent reasoning-augmented Vision-Language-Action (VLA) models have improved the interpretability of end-to-end autonomous driving by generating intermediate reasoning traces. Yet these models primarily describe what they perceive and intend to do, rarely questioning whether their planned actions ar

Cited by 0SourceScholar
2026

EMMA: Scaling Mobile Manipulation Via Egocentric Human Data

ICRA 2026poster

Scaling mobile manipulation imitation learning is bottlenecked by expensive mobile robot teleoperation. We present Egocentric Mobile MAnipulation (EMMA), an end-to-end framework training mobile manipulation policies from human mobile manipulation data with static robot data, sidestepping mobile tele…

2026

EMMA: Scaling Mobile Manipulation via Egocentric Human Data

RA-L 2026

Scaling mobile manipulation imitation learning is bottlenecked by expensive mobile robot teleoperation. We present Egocentric Mobile MAnipulation (EMMA), an end-to-end framework training mobile manipulation policies from human mobile manipulation data with static robot data, sidestepping mobile tele

Cited by 31SourcecodeScholar
2026

EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World

RSS 2026poster

Robot learning increasingly depends on large and diverse data, yet robot data collection remains expensive and difficult to scale. Egocentric human data offer a promising alternative by capturing rich manipulation behavior across everyday environments. However, existing human datasets are often limi…

Cited by 0SourceScholar
2026

Emergence of Human to Robot Transfer in Vision-Language-Action Models

RSS 2026poster

Vision-language-action (VLA) models can enable broad open world generalization, but require large and diverse datasets. It is appealing to consider whether some of this data can come from human videos, which cover diverse real-world situations and are easy to obtain. However, it is difficult to trai…

Cited by 0SourceScholar
2026

KinDER: A Physical Reasoning Benchmark for Robot Learning and Planning

RSS 2026poster

Robotic systems that interact with the physical world must reason about kinematic and dynamic constraints imposed by their own embodiment, their environment, and the task at hand. We introduce KinDER, a benchmark for Kinematic and Dynamic Embodied Reasoning that targets physical reasoning challenges…

Cited by 0SourceScholar
2026

Opt2Skill: Imitating Dynamically-Feasible Whole-Body Trajectories for Versatile Humanoid Loco-Manipulation

ICRA 2026poster

Humanoid robots are designed to perform diverse loco-manipulation tasks. However, they face challenges due to their high-dimensional and unstable dynamics, as well as the complex contact-rich nature of the tasks. Model-based optimal control methods offer flexibility to define precise motion but are …

2026

ReSteer: Quantifying and Refining the Steerability of Multitask Robot Policies

RSS 2026poster

Despite strong multi-task pretraining, existing policies often exhibit poor task steerability. For example, a robot may fail to respond to a new instruction “put the bowl in the sink” when moving towards the oven, executing “close the oven”, even though it can complete both tasks when executed separ…

Cited by 0SourceScholar
2026

Uncertainty-driven 3D Gaussian Splatting Active Mapping via Anisotropic Visibility Field

CVPR 2026

We present Gaussian Splatting Anisotropic Visibility Field (GAVIS), a novel framework for uncertainty quantification and active mapping in 3DGS. Our key insight is that regions unseen from the training views yield unreliable predictions from the 3DGS. To address this, we introduce a principled and e

Cited by 0SourcecodeScholar
2025

DreamDrive: Generative 4D Scene Modeling from Street View Images

ICRA 2025

Synthesizing photo-realistic visual observations from an ego vehicle's driving trajectory is a critical step towards scalable training of self-driving models. Reconstruction-based methods create 3D scenes from driving logs and synthesize geometry-consistent driving videos through neural rendering, b

Cited by 24SourceScholar
2025

EgoBridge: Domain Adaptation for Generalizable Imitation from Egocentric Human Data

NeurIPS 2025poster

Egocentric human experience data presents a vast resource for scaling up end-to-end imitation learning for robotic manipulation. However, significant domain gaps in visual appearance, sensor modalities, and kinematics between human and robot impede knowledge transfer. This paper presents EgoBridge,…

Cited by 0SourcecodeScholar
2025

EgoMimic: Scaling Imitation Learning via Egocentric Video

ICRA 2025

The scale and diversity of demonstration data required for imitation learning is a significant challenge. We present EgoMimic, a full-stack framework which scales manipulation via human embodiment data, specifically egocentric human videos paired with 3D hand tracking. EgoMimic achieves this through

Cited by 136SourcecodeScholar
2025

Generalizable Domain Adaptation for Sim-and-Real Policy Co-Training

NeurIPS 2025poster

Behavior cloning has shown promise for robot manipulation, but real-world demonstrations are costly to acquire at scale. While simulated data offers a scalable alternative, particularly with advances in automated demonstration generation, transferring policies to the real world is hampered by variou…

Cited by 0SourceScholar
2025

Generative Trajectory Stitching through Diffusion Composition

NeurIPS 2025spotlight

Effective trajectory stitching for long-horizon planning is a significant challenge in robotic decision-making. While diffusion models have shown promise in planning, they are limited to solving tasks similar to those seen in their training data. We propose CompDiffuser, a novel generative approach…

Cited by 0SourceScholar
2025

ImMimic: Cross-Domain Imitation from Human Videos via Mapping and Interpolation

CoRL 2025oral

Learning robot manipulation from abundant human videos offers a scalable alternative to costly robot-specific data collection. However, domain gaps across visual, morphological, and physical aspects hinder direct imitation. To effectively bridge the domain gap, we propose ImMimic, an embodiment-agno…

Cited by 0SourceScholar
2025

Joint Model-based Model-free Diffusion for Planning with Constraints

CoRL 2025poster

Model-free diffusion planners have shown great promise for robot motion planning, but practical robotic systems often require combining them with model-based optimization modules to enforce constraints, such as safety. Na\"ively integrating these modules presents compatibility challenges when diffus…

Cited by 9SourceScholar
2025

LoRA3D: Low-Rank Self-Calibration of 3D Geometric Foundation models

ICLR 2025spotlight

Emerging 3D geometric foundation models, such as DUSt3R, offer a promising approach for in-the-wild 3D vision tasks. However, due to the high-dimensional nature of the problem space and scarcity of high-quality 3D data, these pre-trained models still struggle to generalize to many challenging circum…

2025

Opt2Skill: Imitating Dynamically-Feasible Whole-Body Trajectories for Versatile Humanoid Loco-Manipulation

RA-L 2025

Humanoid robots are designed to perform diverse loco-manipulation tasks. However, they face challenges due to their high-dimensional and unstable dynamics, as well as the complex contact-rich nature of the tasks. Model-based optimal control methods offer flexibility to define precise motion but are

Cited by 50SourcecodeScholar
2025

RAIL: Reachability-Aided Imitation Learning for Safe Policy Execution

ICRA 2025

Imitation learning (IL) has shown great success in learning complex robot manipulation tasks. However, there remains a need for practical safety methods to justify widespread deployment. In particular, it is important to certify that a system obeys hard constraints on unsafe behavior in settings whe

Cited by 3SourcecodeScholar
2025

SAIL: Faster-than-Demonstration Execution of Imitation Learning Policies

CoRL 2025oral

Offline Imitation Learning (IL) methods such as Behavior Cloning are effective at acquiring complex robotic manipulation skills. However, existing IL-trained policies are confined to execute the task at the same speed as shown in demonstration data. This limits the task throughput of a robotic…

Cited by 0SourcecodeScholar
2025

STORM: Spatio-TempOral Reconstruction Model For Large-Scale Outdoor Scenes

ICLR 2025poster

We present STORM, a spatio-temporal reconstruction model designed for reconstructing dynamic outdoor scenes from sparse observations. Existing dynamic reconstruction methods often rely on per-scene optimization, dense observations across space and time, and strong motion supervision, resulting in le…

2025

What Matters in Learning from Large-Scale Datasets for Robot Manipulation

ICLR 2025poster

Imitation learning from large multi-task demonstration datasets has emerged as a promising path for building generally-capable robots. As a result, 1000s of hours have been spent on building such large-scale datasets around the globe. Despite the continuous growth of such efforts, we still lack a sy…

Cited by 3SourcePDFScholar
2024

EmerNeRF: Emergent Spatial-Temporal Scene Decomposition via Self-Supervision

ICLR 2024poster

We present EmerNeRF, a simple yet powerful approach for learning spatial-temporal representations of dynamic driving scenes. Grounded in neural fields, EmerNeRF simultaneously captures scene geometry, appearance, motion, and semantics via self-bootstrapping. EmerNeRF hinges upon two core components:…

2024

Generative Factor Chaining: Coordinated Manipulation with Diffusion-based Factor Graph

CoRL 2024poster

Learning to plan for multi-step, multi-manipulator tasks is notoriously difficult because of the large search space and the complex constraint satisfaction problems. We present Generative Factor Chaining (GFC), a composable generative model for planning. GFC represents a planning problem as a spatia…

Cited by 3SourceScholar
2024

Large Spatial Model: End-to-end Unposed Images to Semantic 3D

NeurIPS 2024poster

Reconstructing and understanding 3D structures from a limited number of images is a classical problem in computer vision. Traditional approaches typically decompose this task into multiple subtasks, involving several stages of complex mappings between different data representations. For example, den…

2024

MimicTouch: Leveraging Multi-modal Human Tactile Demonstrations for Contact-rich Manipulation

CoRL 2024poster

Tactile sensing is critical to fine-grained, contact-rich manipulation tasks, such as insertion and assembly. Prior research has shown the possibility of learning tactile-guided policy from teleoperated demonstration data. However, to provide the demonstration, human users often rely on visual feedb…

Cited by 16SourceScholar
2024

NOD-TAMP: Generalizable Long-Horizon Planning with Neural Object Descriptors

CoRL 2024poster

Solving complex manipulation tasks in household and factory settings remains challenging due to long-horizon reasoning, fine-grained interactions, and broad object and scene diversity. Learning skills from demonstrations can be an effective strategy, but such methods often have limited generalizabil…

Cited by 0SourcecodeScholar
2024

Neural Visibility Field for Uncertainty-Driven Active Mapping

CVPR 2024poster

This paper presents Neural Visibility Field (NVF) a novel uncertainty quantification method for Neural Radiance Fields (NeRF) applied to active mapping. Our key insight is that regions not visible in the training views lead to inherently unreliable color predictions by NeRF at this region resulting…

Cited by 4SourcePDFScholar
2024

Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment Collaboration

ICRA 2024

Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for man

Cited by 910SourcecodeScholar
2024

Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment Collaboration0

ICRA 2024poster

Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for man…

Cited by 259SourcecodeScholar
2023

Generative Skill Chaining: Long-Horizon Skill Planning with Diffusion Models

CoRL 2023poster

Long-horizon tasks, usually characterized by complex subtask dependencies, present a significant challenge in manipulation planning. Skill chaining is a practical approach to solving unseen tasks by combining learned skill priors. However, such methods are myopic if sequenced greedily and face scala…

Cited by 79SourcecodeScholar
2023

Guided Conditional Diffusion for Controllable Traffic Simulation

ICRA 2023poster

Controllable and realistic traffic simulation is critical for developing and verifying autonomous vehicles. Typical heuristic-based traffic models offer flexible control to make vehicles follow specific trajectories and traffic rules. On the other hand, data-driven approaches generate realistic and…

Cited by 167SourcecodeScholar
2023

Human-in-the-Loop Task and Motion Planning for Imitation Learning

CoRL 2023poster

Imitation learning from human demonstrations can teach robots complex manipulation skills, but is time-consuming and labor intensive. In contrast, Task and Motion Planning (TAMP) systems are automated and excel at solving long-horizon tasks, but they are difficult to apply to contact-rich tasks. In…

Cited by 21SourcecodeScholar
2023

Language-Guided Traffic Simulation via Scene-Level Diffusion

CoRL 2023oral

Realistic and controllable traffic simulation is a core capability that is necessary to accelerate autonomous vehicle (AV) development. However, current approaches for controlling learning-based traffic models require significant domain expertise and are difficult for practitioners to use. To remedy…

Cited by 94SourceScholar
2023

Learning to Discern: Imitating Heterogeneous Human Demonstrations with Preference and Representation Learning

CoRL 2023poster

Practical Imitation Learning (IL) systems rely on large human demonstration datasets for successful policy learning. However, challenges lie in maintaining the quality of collected data and addressing the suboptimal nature of some demonstrations, which can compromise the overall dataset quality and…

Cited by 9SourceScholar
2023

MimicPlay: Long-Horizon Imitation Learning by Watching Human Play

CoRL 2023oral

Imitation learning from human demonstrations is a promising paradigm for teaching robots manipulation skills in the real world. However, learning complex long-horizon tasks often requires an unattainable amount of demonstrations. To reduce the high data requirement, we resort to human play data - vi…

Cited by 187SourcecodeScholar
2023

ProgPrompt: Generating Situated Robot Task Plans using Large Language Models

ICRA 2023poster

Task planning can require defining myriad domain knowledge about the world in which a robot needs to act. To ameliorate that effort, large language models (LLMs) can be used to score potential next actions during task planning, and even generate action sequences directly, given an instruction in nat…

Cited by 893SourcecodeScholar
2022

AdvDO: Realistic Adversarial Attacks for Trajectory Prediction

ECCV 2022poster

"Trajectory prediction is essential for autonomous vehicles (AVs) to plan correct and safe driving behaviors. While many prior works aim to achieve higher prediction accuracy, few studies the adversarial robustness of their methods. To bridge this gap, we propose to study the adversarial robustness…

Cited by 89SourcePDFScholar
2022

Robust Trajectory Prediction against Adversarial Attacks

CoRL 2022oral

Trajectory prediction using deep neural networks (DNNs) is an essential component of autonomous driving (AD) systems. However, these methods are vulnerable to adversarial attacks, leading to serious consequences such as collisions. In this work, we identify two key ingredients to defend trajectory…

Cited by 47SourceScholar
2021

Co-GAIL: Learning Diverse Strategies for Human-Robot Collaboration

CoRL 2021poster

We present a method for learning human-robot collaboration policy from human-human collaboration demonstrations. An effective robot assistant must learn to handle diverse human behaviors shown in the demonstrations and be robust when the humans adjust their strategies during online task execution. O…

Cited by 49SourceScholar
2021

Deep Affordance Foresight: Planning Through What Can Be Done in the Future

ICRA 2021poster

Planning in realistic environments requires searching in large planning spaces. Affordances are a powerful concept to simplify this search, because they model what actions can be successful in a given situation. However, the classical notion of affordance is not suitable for long horizon planning be…

Cited by 94SourcecodeScholar
2021

Generalization Through Hand-Eye Coordination: An Action Space for Learning Spatially-Invariant Visuomotor Control

IROS 2021poster

Imitation Learning (IL) is an effective framework to learn visuomotor skills from offline demonstration data. However, IL methods often fail to generalize to new scene configurations not covered by training data. On the other hand, humans can manipulate objects in varying conditions. Key to such cap…

Cited by 35SourceScholar
2021

What Matters in Learning from Offline Human Demonstrations for Robot Manipulation

CoRL 2021oral

Imitating human demonstrations is a promising approach to endow robots with various manipulation capabilities. While recent advances have been made in imitation learning and batch (offline) reinforcement learning, a lack of open-source human datasets and reproducible learning methods make assessing…

Cited by 523SourcecodeScholar
2020

6-PACK: Category-level 6D Pose Tracker with Anchor-Based Keypoints

ICRA 2020poster

We present 6-PACK, a deep learning approach to category-level 6D object pose tracking on RGB-D data. Our method tracks in real time novel object instances of known object categories such as bowls, laptops, and mugs. 6-PACK learns to compactly represent an object by a handful of 3D keypoints, based o…

Cited by 190SourcecodeScholar
2020

GTI: Learning to Generalize across Long-Horizon Tasks from Human Demonstrations

RSS 2020poster

Imitation learning is an effective and safe technique to train robot policies in the real world because it does not depend on an expensive random exploration process. However, due to the lack of exploration, learning policies that generalize beyond the demonstrated behaviors is still an open challen…

Cited by 174SourcePDFScholar
2020

Procedure Planning in Instructional Videos

ECCV 2020poster

In this paper, we study the problem of procedure planning in instructional videos, which can be seen as the first step towards enabling autonomous agents to plan for complex tasks in everyday settings such as cooking. Given the current visual observation of the world and a visual goal, we ask the qu…

Cited by 119SourcePDFScholar
2019

Continuous Relaxation of Symbolic Planner for One-Shot Imitation Learning

IROS 2019poster

We address one-shot imitation learning, where the goal is to execute a previously unseen task based on a single demonstration. While there has been exciting progress in this direction, most of the approaches still require a few hundred tasks for meta-training, which limits the scalability of the app…

Cited by 48SourceScholar
2019

DenseFusion: 6D Object Pose Estimation by Iterative Dense Fusion

CVPR 2019poster

A key technical challenge in performing 6D object pose estimation from RGB-D image is to fully leverage the two complementary data sources. Prior works either extract information from the RGB image and depth separately or use costly post-processing steps, limiting their performances in highly clutte…

Cited by 1292PDFScholar
2019

Neural Task Graphs: Generalizing to Unseen Tasks From a Single Video Demonstration

CVPR 2019oral

Our goal is to generate a policy to complete an unseen task given just a single video demonstration of the task in a given domain. We hypothesize that to successfully generalize to unseen complex tasks from a single video demonstration, it is necessary to explicitly incorporate the compositional str…

Cited by 173PDFScholar
2019

Regression Planning Networks

NeurIPS 2019poster

Recent learning-to-plan methods have shown promising results on planning directly from observation space. Yet, their ability to plan for long-horizon tasks is limited by the accuracy of the prediction model. On the other hand, classical symbolic planners show remarkable capabilities in solving long-…

2019

Situational Fusion of Visual Representation for Visual Navigation

ICCV 2019poster

A complex visual navigation task puts an agent in different situations which call for a diverse range of visual perception abilities. For example, to "go to the nearest chair", the agent might need to identify a chair in a living room using semantics, follow along a hallway using vanishing point cue…

Cited by 77PDFScholar
2018

Neural Task Programming: Learning to Generalize Across Hierarchical Tasks

ICRA 2018poster

In this work, we propose a novel robot learning framework called Neural Task Programming (NTP), which bridges the idea of few-shot learning from demonstration and neural program induction. NTP takes as input a task specification (e.g., video demonstration of a task) and recursively decomposes it int…

Cited by 257SourcecodeScholar
2015

Folding deformable objects using predictive simulation and trajectory optimization

IROS 2015poster

Robotic manipulation of deformable objects remains a challenging task. One such task is folding a garment autonomously. Given start and end folding positions, what is an optimal trajectory to move the robotic arm to fold a garment? Certain trajectories will cause the garment to move, creating wrinkl…

Cited by 158SourceScholar
2015

Regrasping and unfolding of garments using predictive thin shell modeling

ICRA 2015poster

Deformable objects such as garments are highly unstructured, making them difficult to recognize and manipulate. In this paper, we propose a novel method to teach a two-arm robot to efficiently track the states of a garment from an unknown state to a known state by iterative regrasping. The problem i…

Cited by 95SourceScholar