← Search

Yifeng Zhu

27 accepted papers

2026

Differentiable Laplacian Matrix Guided Superpixel Segmentation

CVPR 2026

Superpixels partition an image into perceptually coherent regions, reducing the cost of downstream vision tasks. Modern deep learning methods excel at superpixel generation but often yield irregular boundaries and isolated pixels, necessitating non-differentiable post-processing to enforce connectiv

Cited by 0SourcecodeScholar
2026

IMPASTO: Integrating Model-Based Planning with Learned Dynamics Models for Robotic Oil Painting Reproduction

ICRA 2026poster

Robotic reproduction of oil paintings using soft brushes and pigments requires force-sensitive control of deformable tools, prediction of brushstroke effects, and multi-step stroke planning, often without human step-by-step demonstrations or faithful simulators. Given only a sequence of target oil p…

2025

BUMBLE: Unifying Reasoning and Acting with Vision-Language Models for Building-wide Mobile Manipulation

ICRA 2025

To operate at a building scale, service robots must perform long-horizon mobile manipulation tasks by navigating to different rooms, accessing multiple floors, and interacting with a wide and unseen range of everyday objects. We refer to these tasks as Building-wide Mobile Manipulation. To tackle th

Cited by 31SourceScholar
2025

ComposableNav: Instruction-Following Navigation in Dynamic Environments via Composable Diffusion

CoRL 2025poster

This paper considers the problem of enabling robots to navigate dynamic environments while following instructions. The challenge lies in the combinatorial nature of instruction specifications: each instruction can include multiple specifications, and the number of possible specification combination…

Cited by 0SourceScholar
2025

LodeStar: Long-horizon Dexterity via Synthetic Data Augmentation from Human Demonstrations

CoRL 2025poster

Developing robotic systems capable of robustly executing long-horizon manipulation tasks with human-level dexterity is challenging, as such tasks require both physical dexterity and seamless sequencing of manipulation skills while robustly handling environment variations. While imitation learning of…

Cited by 0SourcecodeScholar
2025

SocialNav-SUB: Benchmarking VLMs for Scene Understanding in Social Robot Navigation

CoRL 2025poster

Robot navigation in dynamic, human-centered environments requires socially-compliant decisions grounded in robust scene understanding, including spatiotemporal awareness and the ability to interpret human intentions. Recent Vision-Language Models (VLMs) show exhibit promising capabilities such as ob…

Cited by 0SourceScholar
2024

Harmon: Whole-Body Motion Generation of Humanoid Robots from Language Descriptions

CoRL 2024poster

Humanoid robots, with their human-like embodiment, have the potential to integrate seamlessly into human environments. Critical to their coexistence and cooperation with humans is the ability to understand natural language communications and exhibit human-like behaviors. This work focuses on generat…

Cited by 9SourcecodeScholar
2024

INTERPRET: Interactive Predicate Learning from Language Feedback for Generalizable Task Planning

RSS 2024poster

Learning abstract state representations and knowledge is crucial for long-horizon robot planning. We present InterPreT, an LLM-powered framework for robots to learn symbolic predicates from language feedback of human non-experts during embodied interaction. The learned predicates provide relational…

2024

LOTUS: Continual Imitation Learning for Robot Manipulation Through Unsupervised Skill Discovery

ICRA 2024poster

We introduce LOTUS, a continual imitation learning algorithm that empowers a physical robot to continuously and efficiently learn to solve new manipulation tasks throughout its lifespan. The core idea behind LOTUS is constructing an ever-growing skill library from a sequence of new tasks with a smal…

Cited by 26SourcecodeScholar
2024

OKAMI: Teaching Humanoid Robots Manipulation Skills through Single Video Imitation

CoRL 2024poster

We study the problem of teaching humanoid robots manipulation skills by imitating from single video demonstrations. We introduce OKAMI, a method that generates a manipulation plan from a single RGB-D video and derives a policy for execution. At the heart of our approach is object-aware retargeting,…

Cited by 33SourcecodeScholar
2024

Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment Collaboration

ICRA 2024

Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for man

Cited by 910SourcecodeScholar
2024

Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment Collaboration0

ICRA 2024poster

Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for man…

Cited by 259SourcecodeScholar
2023

A Joint Modeling of Vision-Language-Action for Target-oriented Grasping in Clutter

ICRA 2023poster

We focus on the task of language-conditioned grasping in clutter, in which a robot is supposed to grasp the target object based on a language instruction. Previous works separately conduct visual grounding to localize the target object, and generate a grasp for that object. However, these works requ…

Cited by 49SourcecodeScholar
2023

LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning

NeurIPS 2023poster

Lifelong learning offers a promising paradigm of building a generalist agent that learns and adapts over its lifespan. Unlike traditional lifelong learning problems in image and text domains, which primarily involve the transfer of declarative knowledge of entities and concepts, lifelong learning i…

Cited by 108SourcePDFScholar
2023

Learning Generalizable Manipulation Policies with Object-Centric 3D Representations

CoRL 2023poster

We introduce GROOT, an imitation learning method for learning robust policies with object-centric and 3D priors. GROOT builds policies that generalize beyond their initial training conditions for vision-based manipulation. It constructs object-centric 3D representations that are robust toward backgr…

Cited by 48SourcecodeScholar
2023

Learning to Walk by Steering: Perceptive Quadrupedal Locomotion in Dynamic Environments

ICRA 2023poster

We tackle the problem of perceptive locomotion in dynamic environments. In this problem, a quadrupedal robot must exhibit robust and agile walking behaviors in response to environmental clutter and moving obstacles. We present a hierarchical learning framework, named PRELUDE, which decomposes the pr…

Cited by 10SourcecodeScholar
2023

Symbolic State Space Optimization for Long Horizon Mobile Manipulation Planning

IROS 2023poster

In existing task and motion planning (TAMP) research, it is a common assumption that experts manually specify the state space for task-level planning. A well-developed state space enables the desirable distribution of limited computational resources between task planning and motion planning. However…

Cited by 6SourceScholar
2022

Bottom-Up Skill Discovery From Unsegmented Demonstrations for Long-Horizon Robot Manipulation

RA-L 2022

We tackle real-world long-horizon robot manipulation tasks through skill discovery. We present a bottom-up approach to learning a library of reusable skills from unsegmented demonstrations and use these skills to synthesize prolonged robot behaviors. Our method starts with constructing a hierarchica

Cited by 108SourceScholar
2022

VIOLA: Object-Centric Imitation Learning for Vision-Based Robot Manipulation

CoRL 2022poster

We introduce VIOLA, an object-centric imitation learning approach to learning closed-loop visuomotor policies for robot manipulation. Our approach constructs object-centric representations based on general object proposals from a pre-trained vision model. VIOLA uses a transformer-based policy to rea…

Cited by 19SourcecodeScholar
2022

Visually Grounded Task and Motion Planning for Mobile Manipulation

ICRA 2022poster

Task and motion planning (TAMP) algorithms aim to help robots achieve task-level goals, while maintaining motion-level feasibility. This paper focuses on TAMP domains that involve robot behaviors that take extended periods of time (e.g., long-distance navigation). In this paper, we develop a visual…

Cited by 32SourceScholar
2021

Fast Uncertainty Quantification for Deep Object Pose Estimation

ICRA 2021poster

Deep learning-based object pose estimators are often unreliable and overconfident especially when the input image is outside the training domain, for instance, with sim2real transfer. Efficient and robust uncertainty quantification (UQ) in pose estimators is critically needed in many robotic tasks.…

Cited by 35SourceScholar
2021

Hierarchical Planning for Long-Horizon Manipulation with Geometric and Symbolic Scene Graphs

ICRA 2021poster

We present a visually grounded hierarchical planning algorithm for long-horizon manipulation tasks. Our algorithm offers a joint framework of neuro-symbolic task planning and low-level motion generation conditioned on the specified goal. At the core of our approach is a two-level scene graph represe…

Cited by 138SourceScholar
2021

Machine versus Human Attention in Deep Reinforcement Learning Tasks

NeurIPS 2021poster

Deep reinforcement learning (RL) algorithms are powerful tools for solving visuomotor decision tasks. However, the trained models are often difficult to interpret, because they are represented as end-to-end deep neural networks. In this paper, we shed light on the inner workings of such trained mod…

Cited by 28SourcePDFScholar
2021

Synergies Between Affordance and Geometry: 6-DoF Grasp Detection via Implicit Representations

RSS 2021poster

Grasp detection in clutter requires the robot to reason about the 3D scene from incomplete and noisy perception. In this work; we draw insight that 3D reconstruction and grasp learning are two intimately connected tasks; both of which require a fine-grained understanding of local geometry details. W…

Cited by 169SourcePDFScholar
2020

Human Gaze Assisted Artificial Intelligence: A Review

IJCAI 2020poster

Human gaze reveals a wealth of information about internal cognitive state. Thus, gaze-related research has significantly increased in computer vision, natural language processing, decision learning, and robotics in recent years. We provide a high-level overview of the research efforts in these field…

Cited by 0SourcePDFScholar