← Search

Benjamin Burchfiel

20 accepted papers

2026

Geometry-aware 4D Video Generation for Robot Manipulation

ICLR 2026poster

Understanding and predicting dynamics of the physical world can enhance a robot's ability to plan and interact effectively in complex environments. While recent video generation models have shown strong potential in modeling dynamic scenes, generating videos that are both temporally coherent and geo…

Cited by 0SourcecodeScholar
2026

Using Non-Expert Data to Robustify Imitation Learning Via Offline Reinforcement Learning

ICRA 2026poster

Imitation learning has proven effective for training robots to perform complex tasks from expert human demonstrations. However, it remains limited by its reliance on high-quality, task-specific data, restricting adaptability to the diverse range of real-world object configurations and scenarios. In …

2025

Adaptive Compliance Policy: Learning Approximate Compliance for Diffusion Guided Control

ICRA 2025

Compliance plays a crucial role in manipulation, as it balances between the concurrent control of position and force under uncertainties. Yet compliance is often overlooked by today's visuomotor policies that solely focus on position control. This paper introduces Adaptive Compliance Policy (ACP), a

Cited by 57SourcecodeScholar
2025

Diffusion Policy Policy Optimization

ICLR 2025poster

We introduce Diffusion Policy Policy Optimization, DPPO, an algorithmic framework including best practices for fine-tuning diffusion-based policies (e.g. Diffusion Policy) in continuous control and robot learning tasks using the policy gradient (PG) method from reinforcement learning (RL). PG method…

Cited by 270SourcePDFScholar
2025

GHIL-Glue: Hierarchical Control with Filtered Subgoal Images

ICRA 2025

Image and video generative models that are pretrained on Internet-scale data can greatly increase the generalization capacity of robot learning systems. These models can function as high-level planners, generating intermediate sub-goals for low-level goal-conditioned policies to reach. However, the

Cited by 9SourcecodeScholar
2025

One Demo is Worth a Thousand Trajectories: Action-View Augmentation for Visuomotor Policies

CoRL 2025poster

Visuomotor policies for manipulation have demonstrated remarkable potential in modeling complex robotic behaviors, yet minor alterations in the robot’s initial configuration and unseen obstacles easily lead to out-of-distribution observations. Without extensive data collection effort, these result i…

Cited by 0SourceScholar
2025

PolyTouch: A Robust Multi-Modal Tactile Sensor for Contact-Rich Manipulation Using Tactile-Diffusion Policies

ICRA 2025

Achieving robust dexterous manipulation in un-structured domestic environments remains a significant challenge in robotics. Even with state-of-the-art robot learning methods, haptic-oblivious control strategies (i.e. those relying only on external vision and/or proprioception) often fall short due t

Cited by 23SourceScholar
2025

Should VLMs be Pre-trained with Image Data?

ICLR 2025poster

Pre-trained LLMs that are further trained with image data perform well on vision-language tasks. While adding images during a second training phase effectively unlocks this capability, it is unclear how much of a gain or loss this two-step pipeline gives over VLMs which integrate images earlier int…

Cited by 0SourcePDFScholar
2025

Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets

RSS 2025poster

Imitation learning has emerged as a promising approach towards building generalist robots. However, the reliance on high-quality expert demonstrations poses a challenge in scaling imitation learning for large-scale robot foundation models. On the other hand, large amounts of video data depicting a w…

Cited by 2PDFScholar
2024

ManiWAV: Learning Robot Manipulation from In-the-Wild Audio-Visual Data

CoRL 2024poster

Audio signals provide rich information for the robot interaction and object properties through contact. These information can surprisingly ease the learning of contact-rich robot manipulation skills, especially when the visual information alone is ambiguous or incomplete. However, the usage of audio…

Cited by 24SourceScholar
2024

OpenVLA: An Open-Source Vision-Language-Action Model

CoRL 2024poster

Large policies pretrained on a combination of Internet-scale vision-language data and diverse robot demonstrations have the potential to change how we teach robots new skills: rather than training new behaviors from scratch, we can fine-tune such vision-language-action (VLA) models to obtain robust,…

Cited by 437SourceScholar
2024

Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots

RSS 2024poster

We present Universal Manipulation Interface (UMI) -- a data collection and policy learning framework that allows direct skill transfer from in-the-wild human demonstrations to deployable robot policies. UMI employs hand-held grippers coupled with careful interface design to enable portable, low-cost…

Cited by 235SourcePDFScholar
2023

AdaptSim: Task-Driven Simulation Adaptation for Sim-to-Real Transfer

CoRL 2023poster

Simulation parameter settings such as contact models and object geometry approximations are critical to training robust manipulation policies capable of transferring from simulation to real-world deployment. There is often an irreducible gap between simulation and reality: attempting to match the dy…

Cited by 16SourceScholar
2023

Bag All You Need: Learning a Generalizable Bagging Strategy for Heterogeneous Objects

IROS 2023poster

We introduce a practical robotics solution for the task of heterogeneous bagging, requiring the placement of multiple rigid and deformable objects into a deformable bag. This is a difficult task as it features complex interactions between multiple highly deformable objects under limited observabilit…

Cited by 19SourceScholar
2023

Cloth Funnels: Canonicalized-Alignment for Multi-Purpose Garment Manipulation

ICRA 2023poster

Automating garment manipulation is challenging due to extremely high variability in object configurations. To reduce this intrinsic variation, we introduce the task of “canonicalized-alignment” that simplifies downstream applications by reducing the possible garment configurations. This task can be…

Cited by 46SourceScholar
2022

DextAIRity: Deformable Manipulation Can be a Breeze

RSS 2022poster

This paper introduces DextAIRity, an approach to manipulate deformable objects using active airflow. In contrast to conventional contact-based quasi-static manipulations, DextAIRity allows the system to apply dense forces on out-of-contact surfaces, expands the system's reach range, and provides saf…

Cited by 59SourcePDFScholar
2022

Iterative Residual Policy for Goal-Conditioned Dynamic Manipulation of Deformable Objects

RSS 2022poster

This paper tackles the task of goal-conditioned dynamic manipulation of deformable objects. This task is highly challenging due to its complex dynamics (introduced by object deformation and high-speed action) and strict task requirements (defined by a precise goal specification). To address these ch…

Cited by 90SourcePDFScholar
2019

Grounding Language Attributes to Objects using Bayesian Eigenobjects

IROS 2019poster

We develop a system to disambiguate object instances within the same class based on simple physical descriptions. The system takes as input a natural language phrase and a depth image containing a segmented object and predicts how similar the observed object is to the object described by the phrase.…

Cited by 23SourceScholar
2018

Hybrid Bayesian Eigenobjects: Combining Linear Subspace and Deep Network Methods for 3D Robot Vision

IROS 2018poster

We introduce Hybrid Bayesian Eigenobjects (HBEOs), a novel representation for 3D objects designed to allow a robot to jointly estimate the pose, class, and full 3D geometry of a novel object observed from a single viewpoint in a single practical framework. By combining both linear subspace methods a…

Cited by 6SourceScholar