← Search

Zhongxiang Zhou

15 accepted papers

2026

Geometry-Guided Modeling of Foundation Features Enables Generalizable Object Shape Deformation Learning

ICML 2026poster

Monocular 3D shape recovery is fundamental to geometric understanding, yet achieving robust generalization across arbitrary viewpoints and unseen object categories remains a significant challenge. In this paper, we present a generalizable deformation learning framework that reconstructs 3D objects b…

Cited by 0SourceScholar
2026

Toward Embodiment Equivariant Vision-Language-Action Policy

ICRA 2026poster

Vision-language-action policies learn manipulation skills across tasks, environments and embodiments through large-scale pre-training. However, their ability to generalize to novel robot configurations remains limited. Most approaches emphasize model size, dataset scale and diversity while paying le…

2025

Adaptive Neural Uncalibrated Visual Servo with Zero-shot Transfer of Extrinsics and Scenes

IROS 2025

Deploying visual servo controller to novel scenes with uncertain parameters requires additional manual effort for calibration. Traditional methods tackle this problem by online estimating the Jacobian matrix. However, they struggle in challenging scenes due to intrinsic limitations. For instance, im

Cited by 0SourceScholar
2025

CNSv2: Probabilistic Correspondence Encoded Neural Image Servo

ICRA 2025

Visual servo based on traditional image matching methods often requires accurate keypoint correspondence for high precision control. However, keypoint detection or matching tends to fail in challenging scenarios with inconsistent illuminations or textureless objects, resulting significant performanc

Cited by 2SourceScholar
2025

ColaDex: Contact-guided Optimization and VLM-assisted Selection for Task-oriented Dexterous Grasp Generation

IROS 2025

Task-oriented dexterous grasp generation aims to generate stable and functional grasps that enable a robotic hand to effectively interact with objects to accomplish specific tasks. However, generating high-dimensional hand configurations that seamlessly adapt to diverse task requirements and object

Cited by 0SourceScholar
2025

Disambiguate Gripper State in Grasp-Based Tasks: Pseudo-Tactile as Feedback Enables Pure Simulation Learning

IROS 2025

Grasp-based manipulation tasks are fundamental to robots interacting with their environments, yet gripper state ambiguity significantly reduces the robustness of imitation learning policies for these tasks. Data-driven solutions face the challenge of high real-world data costs, while simulation data

Cited by 2SourceScholar
2024

Adapting for Calibration Disturbances: A Neural Uncalibrated Visual Servoing Policy

ICRA 2024poster

Visual servoing (VS) is a widely used technique in industries where there are hundreds of robots, but it requires accurate camera calibration including camera intrinsic and extrinsic parameters. However, it is labour-intensive to calibrate robots one-by-one in practical use. In this paper, we propos…

Cited by 1SourceScholar
2023

A Hyper-Network Based End-to-End Visual Servoing With Arbitrary Desired Poses

RA-L 2023

Recently, several works achieve end-to-end visual servoing (VS) for robotic manipulation by replacing traditional controller with differentiable neural networks, but lose the ability to servo arbitrary desired poses. This letter proposes a differentiable architecture for arbitrary pose servoing: a h

Cited by 8SourceScholar
2023

A Joint Modeling of Vision-Language-Action for Target-oriented Grasping in Clutter

ICRA 2023poster

We focus on the task of language-conditioned grasping in clutter, in which a robot is supposed to grasp the target object based on a language instruction. Previous works separately conduct visual grounding to localize the target object, and generate a grasp for that object. However, these works requ…

Cited by 49SourcecodeScholar
2023

Open-Set Object Detection Using Classification-Free Object Proposal and Instance-Level Contrastive Learning

RA-L 2023

Detecting both known and unknown objects is a fundamental skill for robot manipulation in unstructured environments. Open-set object detection (OSOD) is a promising direction to handle the problem consisting of two subtasks: objects and background separation, and open-set object classification. In t

Cited by 21SourceScholar
2022

Learning to Fill the Seam by Vision: Sub-millimeter Peg-in-hole on Unseen Shapes in Real World

ICRA 2022poster

In the peg insertion task, human pays attention to the seam between the peg and the hole and tries to fill it continuously with visual feedback. By imitating the human's behavior, we design architectures with position and orientation estimators based on the seam representation for pose alignment, wh…

Cited by 18SourcecodeScholar
2021

Assembly Sequence Generation for New Objects via Experience Learned from Similar Object

IROS 2021poster

Assembly orders of components have direct influence on feasibility and efficiency of assembly process in manufacturing and are usually defined by experienced operators. To automate the assembly sequence generation process, we present a method using the idea of case-based reasoning, which can take ad…

Cited by 2SourceScholar
2021

Learn to Differ: Sim2Real Small Defection Segmentation Network

IROS 2021poster

Recent studies on deep-learning-based small defection segmentation approaches are trained in specific settings and tend to be limited by fixed context. Throughout the training, the network inevitably learns the representation of the background of the training data before figuring out the defection.…

Cited by 0SourcecodeScholar
2021

REDE: End-to-End Object 6D Pose Robust Estimation Using Differentiable Outliers Elimination

RA-L 2021

Object 6D pose estimation is a fundamental task in many applications. Conventional methods solve the task by detecting and matching the keypoints, then estimating the pose. Recent efforts bringing deep learning into the problem mainly overcome the vulnerability of conventional methods to environment

Cited by 41SourcecodeScholar