← Search

Zipeng Fu

12 accepted papers

2025

CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models

CVPR 2025poster

Vision-language-action models (VLAs) have shown potential in leveraging pretrained vision-language models and diverse robot demonstrations for learning generalizable sensorimotor control. While this paradigm effectively utilizes large-scale data from both robotic and non-robotic sources, current VLA…

2025

Helpful DoggyBot: Open-World Object Fetching using Legged Robots and Vision-Language Models

IROS 2025

Learning-Based methods have achieved strong performance for quadrupedal locomotion. However, several challenges prevent quadrupeds from learning helpful indoor skills that require interaction with environments and humans: lack of end-effectors for manipulation, limited semantic under-standing using

Cited by 15SourceScholar
2024

HumanPlus: Humanoid Shadowing and Imitation from Humans

CoRL 2024poster

One of the key arguments for building robots that have similar form factors to human beings is that we can leverage the massive human data for training.Yet, doing so has remained challenging in practice due to the complexities in humanoid perception and control, lingering physical gaps between human…

Cited by 109SourceScholar
2024

Mobile ALOHA: Learning Bimanual Mobile Manipulation using Low-Cost Whole-Body Teleoperation

CoRL 2024poster

Imitation learning from human demonstrations has shown impressive performance in robotics. However, most results focus on table-top manipulation, lacking the mobility and dexterity necessary for generally useful tasks. In this work, we develop a system for imitating mobile manipulation tasks that ar…

Cited by 15SourceScholar
2024

Mobility VLA: Multimodal Instruction Navigation with Long-Context VLMs and Topological Graphs

CoRL 2024poster

An elusive goal in navigation research is to build an intelligent agent that can understand multimodal instructions including natural language and image, and perform useful navigation. To achieve this, we study a widely useful category of navigation tasks we call Multimodal Instruction Navigation wi…

Cited by 20SourceScholar
2024

UMI-on-Legs: Making Manipulation Policies Mobile with Manipulation-Centric Whole-body Controllers

CoRL 2024poster

We introduce UMI-on-Legs, a new framework that combines real-world and simulation data for quadruped manipulation systems. We scale task-centric data collection in the real world using a handheld gripper (UMI), providing a cheap way to demonstrate task-relevant manipulation skills without a robot.…

Cited by 45SourceScholar
2023

Robot Parkour Learning

CoRL 2023oral

Parkour is a grand challenge for legged locomotion that requires robots to overcome various obstacles rapidly in complex environments. Existing methods can generate either diverse but blind locomotion skills or vision-based but specialized skills by using reference animal data or complex rewards. Ho…

Cited by 195SourcecodeScholar
2022

Coupling Vision and Proprioception for Navigation of Legged Robots

CVPR 2022poster

We exploit the complementary strengths of vision and proprioception to develop a point-goal navigation system for legged robots, called VP-Nav. Legged systems are capable of traversing more complex terrain than wheeled robots, but to fully utilize this capability, we need a high-level path planner i…

Cited by 73PDFcodeScholar
2022

Deep Whole-Body Control: Learning a Unified Policy for Manipulation and Locomotion

CoRL 2022oral

An attached arm can significantly increase the applicability of legged robots to several mobile manipulation tasks that are not possible for the wheeled or tracked counterparts. The standard modular control pipeline for such legged manipulators is to decouple the controller into that of manipulation…

Cited by 165SourcecodeScholar
2021

Minimizing Energy Consumption Leads to the Emergence of Gaits in Legged Robots

CoRL 2021poster

Legged locomotion is commonly studied and expressed as a discrete set of gait patterns, like walk, trot, gallop, which are usually treated as given and pre-programmed in legged robots for efficient locomotion at different speeds. However, fixing a set of pre-programmed gaits limits the generality of…

Cited by 138SourcecodeScholar