← Search

Hanbo Zhang

21 accepted papers

2026

Hypothesis-driven Model Expansion under Uncertainty for Open-World Robot Planning

RSS 2026poster

We consider an open-world planning setting in which service robots must operate in unknown environments with incomplete knowledge of objects and actions. Traditional closed-world approaches with pre-programmed knowledge bases fail when robots encounter unexpected situations and tasks, posing a funda…

Cited by 0SourceScholar
2026

Rectifying Gradient Trajectories: A Hierarchical Geometric Framework with Structural Constraints for Few-Shot EEG Adaptation

ICML 2026poster

Few-shot EEG domain adaptation faces severe data heterogeneity and optimization instability. While prevalent "symmetric alignment" methods typically seek a compromised shared subspace, they often falter when domain discrepancies are vast, leading to mutual interference and negative transfer. To over…

Cited by 0SourceScholar
2025

Chain-of-Action: Trajectory Autoregressive Modeling for Robotic Manipulation

NeurIPS 2025poster

We present Chain-of-Action (CoA), a novel visuomotor policy paradigm built upon Trajectory Autoregressive Modeling. Unlike conventional approaches that predict next step action(s) forward, CoA generates an entire trajectory by explicit backward reasoning with task-specific goals through an action-le…

Cited by 0SourceScholar
2025

MimicFunc: Imitating Tool Manipulation from a Single Human Video via Functional Correspondence

CoRL 2025poster

Imitating tool manipulation from human videos offers an intuitive approach to teaching robots, while also providing a promising and scalable alternative to labor-intensive teleoperation data collection for visuomotor policy learning. While humans can mimic tool manipulation behavior by observing oth…

Cited by 0SourceScholar
2025

ProcWorld: Benchmarking Large Model Planning in Reachability-Constrained Environments

EMNLP 2025

We introduce ProcWorld, a large-scale benchmark for partially observable embodied spatial reasoning and long-term planning with large language models (LLM) and vision language models (VLM). ProcWorld features a wide range of challenging embodied navigation and object manipulation tasks, covering 16

Cited by 0SourcePDFScholar
2025

RTV-Bench: Benchmarking MLLM Continuous Perception, Understanding and Reasoning through Real-Time Video

NeurIPS 2025poster

Multimodal Large Language Models (MLLMs) increasingly excel at perception,understanding, and reasoning. However, current benchmarks inadequately evaluate their ability to perform these tasks continuously in dynamic, real-world environments. To bridge this gap, we introduce RT V-Bench, a fine-grained…

Cited by 0SourcecodeScholar
2025

Towards Extrinsic Dexterity Grasping in Unrestricted Environments

IROS 2025

Grasping large and flat objects (e.g., a book or a pan) is often regarded as an ungraspable task, which poses significant challenges due to the unreachable grasping poses. Prior research has exploited environmental interactions through Extrinsic Dexterity, utilizing external structures such as walls

Cited by 0SourcecodeScholar
2024

Towards Unified Interactive Visual Grounding in The Wild

ICRA 2024poster

Interactive visual grounding in Human-Robot Interaction (HRI) is challenging yet practical due to the inevitable ambiguity in natural languages. It requires robots to disambiguate the user’s input by active information gathering. Previous approaches often rely on predefined templates to ask disambig…

Cited by 3SourcecodeScholar
2024

Vision-Language Foundation Models as Effective Robot Imitators

ICLR 2024spotlight

Recent progress in vision language foundation models has shown their ability to understand multimodal data and resolve complicated vision language tasks, including robotics manipulation. We seek a straightforward way of making use of existing vision-language models (VLMs) with simple fine-tuning on…

Cited by 133SourcePDFScholar
2024

What Matters in Training a GPT4-Style Language Model with Multimodal Inputs?

NAACL 2024long

Recent advancements in GPT-4V have displayed remarkable multi-modal capabilities in processing image inputs and following open-ended instructions. Despite these advancements, there is considerable scope for enhancing open-source multi-modal LLMs, especially in terms of multi-modal understanding accu…

2022

REGRAD: A Large-Scale Relational Grasp Dataset for Safe and Object-Specific Robotic Grasping in Clutter

RA-L 2022

Despite the impressive progress achieved in robotic grasping, robots are not skilled in sophisticated tasks (e.g. search and grasp a specified target in clutter). Such tasks involve not only grasping but the comprehensive perception of the world (e.g. the object relationships). Recently, encouraging

Cited by 51SourcecodeScholar
2021

REGNet: REgion-based Grasp Network for End-to-end Grasp Detection in Point Clouds

ICRA 2021poster

Reliable robotic grasping in unstructured environments is a crucial but challenging task. The main problem is to generate the optimal grasp of novel objects from partial noisy observations. This paper presents an end-to-end grasp detection network taking one single-view point cloud as input to tackl…

Cited by 104SourcecodeScholar
2019

A Multi-task Convolutional Neural Network for Autonomous Robotic Grasping in Object Stacking Scenes

IROS 2019poster

Autonomous robotic grasping plays an important role in intelligent robotics. However, how to help the robot grasp specific objects in object stacking scenes is still an open problem, because there are two main challenges for autonomous robots: (1) it is a comprehensive task to know what and how to g…

Cited by 87SourceScholar
2019

ROI-based Robotic Grasp Detection for Object Overlapping Scenes

IROS 2019poster

Grasp detection considering the affiliations between grasps and their owner in object overlapping scenes is a necessary and challenging task for the practical use of the robotic grasping approach. In this paper, a robotic grasp detection algorithm named ROI-GD is proposed to provide a feasible solut…

Cited by 213SourceScholar
2019

Task-oriented Grasping in Object Stacking Scenes with CRF-based Semantic Model

IROS 2019poster

In task-oriented grasping, the robot is supposed to manipulate the objects in a task-compatible manner, which is more important but more challenging than just stably grasping. However, most of existing works perform task-oriented grasping only in single object scenes. This greatly limits their pract…

Cited by 24SourceScholar
2018

Fully Convolutional Grasp Detection Network with Oriented Anchor Box

IROS 2018poster

In this paper, we present a real-time approach to predict multiple grasping poses for a parallel-plate robotic gripper using RGB images. A model with oriented anchor box mechanism is proposed and a new matching strategy is used during the training process. An end-to-end fully convolutional neural ne…

Cited by 246SourceScholar