← Search

Zheyu Zhuang

11 accepted papers

2026

PALM: Enhanced Generalizability for Local Visuomotor Policies via Perception Alignment

RA-L 2026

Generalizing beyond the training domain in image-based behavior cloning remains challenging. Existing methods address individual axes of generalization, workspace shifts, viewpoint changes, and cross-embodiment transfer, yet they are typically developed in isolation and often rely on complex pipelin

Cited by 1SourceScholar
2025

Feature Extractor or Decision Maker: Rethinking the Role of Visual Encoders in Visuomotor Policies

ICRA 2025

An end-to-end (E2E) visuomotor policy is typically treated as a unified whole, but recent approaches using out-of-domain (OOD) data to pretrain the visual encoder have cleanly separated the visual encoder from the network, with the remainder referred to as the policy. We propose Visual Alignment Tes

Cited by 1SourceScholar
2025

MirrorDuo: Reflection-Consistent Visuomotor Learning from Mirrored Demonstration Pairs

CoRL 2025poster

Image-based behaviour cloning leverages demonstrations captured from ubiquitous RGB cameras, enabling impressive visuomotor performance. However, it remains constrained by the cost of collecting sufficiently diverse demonstrations, especially for generalizing across workspace variations. We propose…

Cited by 0SourceScholar
2024

Enhancing Visual Domain Robustness in Behaviour Cloning via Saliency-Guided Augmentation

CoRL 2024poster

In vision-based behaviour cloning (BC), traditional image-level augmentation methods such as pixel shifting enhance in-domain performance but often struggle with visual domain shifts, including distractors, occlusion, and changes in lighting and backgrounds. Conversely, superimposition-based augment…

Cited by 2SourceScholar
2024

Raising Body Ownership in End-to-End Visuomotor Policy Learning via Robot-Centric Pooling

IROS 2024poster

We present Robot-centric Pooling (RcP), a novel pooling method designed to enhance end-to-end visuomo-tor policies by enabling differentiation between the robots and similar entities or their surroundings. Given an image-proprioception pair, RcP guides the aggregation of image features by highlighti…

Cited by 0SourcecodeScholar
2022

GoferBot: A Visual Guided Human-Robot Collaborative Assembly System

IROS 2022poster

The current transformation towards smart manufacturing has led to a growing demand for human-robot collaboration (HRC) in the manufacturing process. Perceiving and understanding the human co-worker's behaviour introduces challenges for collaborative robots to efficiently and effectively perform task…

Cited by 10SourceScholar
2021

Stereo Hybrid Event-Frame (SHEF) Cameras for 3D Perception

IROS 2021poster

Stereo camera systems play an important role in robotics applications to perceive the 3D world. However, conventional cameras have drawbacks such as low dynamic range, motion blur and latency due to the underlying frame- based mechanism. Event cameras address these limitations as they report the bri…

Cited by 29SourcecodeScholar
2020

LyRN (Lyapunov Reaching Network): A Real-Time Closed Loop approach from Monocular Vision

ICRA 2020poster

We propose a closed-loop, multi-instance control algorithm for visually guided reaching based on novel learning principles. A control Lyapunov function methodology is used to design a reaching action for a complex multi-instance task in the case where full state information (poses of all potential r…

Cited by 8SourceScholar
2019

Learning Real-time Closed Loop Robotic Reaching from Monocular Vision by Exploiting A Control Lyapunov Function Structure

IROS 2019poster

Visual reaching and grasping is a fundamental problem in robotics research. This paper proposes a novel approach based on deep learning a control Lyapunov function and its derivatives by encouraging a differential constraint in addition to vanilla regression that directly regresses independent joint…

Cited by 3SourceScholar