← Search

Bin He

46 accepted papers

2026

Autonomous UAV–Quadruped Docking in Complex Terrains Via Active Posture Alignment and Constraint-Aware Control

ICRA 2026poster

Autonomous docking between Unmanned Aerial Vehicles (UAVs) and ground robots is essential for heterogeneous systems, yet most existing approaches target wheeled platforms whose limited mobility constrains exploration in complex terrains. Quadruped robots offer superior adaptability but undergo frequ…

2026

DLM: Unified Decision Language Models for Offline Multi-Agent Sequential Decision Making

ICML 2026spotlight

Building scalable and reusable multi-agent decision policies from offline datasets remains a challenge in offline multi-agent reinforcement learning (MARL), as existing methods often rely on fixed observation formats and action spaces that limit generalization. In contrast, large language models (LL…

Cited by 0SourceScholar
2026

Kaiwu: A Multimodal Manipulation Dataset and Framework for Robot Learning and Human-Robot Interaction

ICRA 2026poster

Cutting-edge robot learning techniques including foundation models and imitation learning from humans all pose huge demands on large-scale and high-quality datasets which constitute one of the bottleneck in the general intelligent robot fields. This paper presents the Kaiwu multimodal dataset to add…

2026

Morphogenetic Assembly and Adaptive Control for Heterogeneous Modular Robots

ICRA 2026poster

This paper presents a closed-loop automation framework for heterogeneous modular robots, encompassing the entire pipeline from morphological construction to adaptive control. Within this framework, a mobile manipulator manipulates heterogeneous functional modules—including structural, joint, and whe…

2026

Offline-Trained GAN-Augmented Highly Adaptive Control with Multi-DoF Fusion for Pneumatic Soft Surgical Robots (I)

ICRA 2026poster

Pneumatic soft robots are well-suited for minimally invasive surgery owing to their compliance and safe interaction with tissues. However, achieving highly adaptive control is difficult owing to modeling inaccuracies, inter-chamber coupling, and disturbances from surgical instruments. Non-learning a…

Cited by 0Scholar
2026

PaiP: An Operational Aware Interactive Planner for Unknown Cabinet Environments

ICRA 2026poster

Box/cabinet scenarios pose with stacked objects significant challenges for robotic motion due to visual occlusions and constrained free space. Traditional collision-free trajectory planning methods often fail when no collision-free paths exist, and may even lead to catastrophic collisions caused by …

2026

Predicting Tactile Sensory Outcome of Physical Human-Robot Interaction Through Embodied Learning Strategy

RA-L 2026

Estimation of robotic dynamic states in physical human-robot interaction (pHRI) is crucial for robots to handle real-world uncertainties. Since wearable tactile sensor arrays have been increasingly applied in robots, the tactile sensory outcomes (TSOs) during pHRIs would provide practical and ideal

Cited by 0SourceScholar
2026

Self-Organised Sequential Multi-Agent Reinforcement Learning for Closely Cooperation Tasks

ICRA 2026poster

Cooperative tasks are common in multi-agent systems, with closely cooperative tasks being a special case of this, where a change in the state of the environment requires multiple agents to perform a specific operation at the same time. Take a box-pushing task as an example, the box is heavy and requ…

Cited by 0SourceScholar
2026

VO-DP: Semantic-Geometric Adaptive Diffusion Policy for Vision-Only Robotic Manipulation

ICRA 2026poster

In the context of imitation learning, visuomotor-based diffusion policy learning is one of the main directions in robotic manipulation. Most of these approaches rely on point clouds as observation inputs and construct scene representations through point clouds feature learning, which enables them to…

2026

When Birds Meet Fish: Vision-Force Fusion for Autonomous Underwater Docking in Cross-Domain Avian-Aquatic Collaboration

ICRA 2026poster

Unmanned aerial–aquatic vehicles (UAAVs) provide cross-domain adaptability and broad visions, while autonomous underwater vehicles (AUVs) support long-duration operations. This work integrates the two by developing a rapid underwater docking and releasing system. An autonomous clamping mechanism is …

Cited by 0Scholar
2025

Align-A-Video: Deterministic Reward Tuning of Image Diffusion Models for Consistent Video Editing

CVPR 2025poster

Due to control limitations in the denoising process and the lack of training, zero-shot video editing methods often struggle to meet user instructions, resulting in generated videos that are visually unappealing and fail to fully satisfy expectations. To address this problem, we propose Align-A-Vide…

Cited by 0SourcePDFScholar
2025

Bridging Training and Execution via Dynamic Directed Graph-Based Communication in Cooperative Multi-Agent Systems

AAAI 2025technical

Multi-agent systems must learn to communicate and understand interactions between agents to achieve cooperative goals in partially observed tasks. However, existing approaches lack a dynamic directed communication mechanism and rely on global states, thus diminishing the role of communication in cen…

2025

KINND: A Keyframe Insertion Framework via Neural Network Decision-Making for VSLAM

RA-L 2025

Keyframe insertion is critical for the performance and robustness of SLAM systems. However, traditional heuristic-based methods often lead to suboptimal keyframe selection, compromising the accuracy of localization and mapping. To address this, we propose KINND, a lightweight neural network-based fr

Cited by 4SourceScholar
2025

Learning Efficient Robotic Garment Manipulation with Standardization

ICML 2025poster

Garment manipulation is a significant challenge for robots due to the complex dynamics and potential self-occlusion of garments. Most existing methods of efficient garment unfolding overlook the crucial role of standardization of flattened garments, which could significantly simplify downstream task…

2025

NeuTRL: Neural Trust-Guided Reinforcement Learning for Human-Robot Collaboration

RA-L 2025

Reinforcement Learning from Human Feedback (RLHF) enables robots to learn cooperative strategies aligned with human expectations by incorporating feedback into the learning process. However, existing RLHF methods rely on explicit query-based feedback, which is limited for complex, long-horizon tasks

Cited by 6SourceScholar
2025

Rotation Invariant Spatial Networks for Single-View Point Cloud Classification

IJCAI 2025

Point cloud classification is critical for three-dimensional scene understanding. However, in real-world scenarios, depth cameras often capture partial, single-view point clouds of objects with different poses, making their accurate classification a challenge. In this paper, we propose a novel point

2025

STC-Tracker: Spatiotemporal-Consistent Multi-Robot Collaboration Framework for Long-Term Dynamic Object Tracking

IROS 2025

Multi-robot cooperative tracking, as a vital sub-field of multi-robot collaboration, exhibits significant potential in areas such as military reconnaissance and emergency rescue. Conventional dynamic object tracking methods often face issues of incomplete target detection and even loss in complex sc

Cited by 0SourceScholar
2025

Self-Organised Sequential Multi-Agent Reinforcement Learning for Closely Cooperation Tasks

RA-L 2025

Cooperative tasks are common in multi-agent systems, with closely cooperative tasks being a special case of this, where a change in the state of the environment requires multiple agents to perform a specific operation at the same time. Take a box-pushing task as an example, the box is heavy and requ

Cited by 0SourceScholar
2025

Sensing Differently: Unifying Vision, Language, Posture and Tactile in Robotic Perception

IROS 2025

Multi-modal fusion perception enhances robotic performance in complex tasks by providing more comprehensive information than single modality. While tactile and proprioceptive sensing are effective for direct contact tasks like grasping, current research mainly focuses on vision-language fusion, negl

Cited by 0SourceScholar
2025

Uni-Zipper: A Multi-modal Perception Framework of Deformable Objects with Unpaired Data

IROS 2025

Multi-modal perception plays a crucial role in preventing deformation and damage during the robotic manipulation of deformable objects. However, integrating new heterogeneous modalities into existing robotic perception frameworks remains a significant challenge, primarily due to the need for massive

Cited by 0SourceScholar
2024

AnimatableDreamer: Text-Guided Non-rigid 3D Model Generation and Reconstruction with Canonical Score Distillation

ECCV 2024poster

"Advances in 3D generation have facilitated sequential 3D model generation (a.k.a 4D generation), yet its application for animatable objects with large motion remains scarce. Our work proposes AnimatableDreamer, a text-to-4D generation framework capable of generating diverse categories of non-rigid…

2024

Closely Cooperative Multi-Agent Reinforcement Learning Based on Intention Sharing and Credit Assignment

RA-L 2024

Collaborative tasks are important in multi-agent systems. Multi-agent reinforcement learning is a commonly used technique for solving multi-agent cooperative policy learning. The closely collaborative task is a special but common case within cooperative tasks, where the change in the environmental s

Cited by 2SourceScholar
2024

Learning Cross Dimension Scene Representation for Interactive Navigation Agents in Obstacle-Cluttered Environments

RA-L 2024

Embodied visual navigation has witnessed significant advancements. However, most studies commonly assume that environments are static and contain at least one collision-free path. In human environments, agents frequently encounter challenges when navigating through scenes with disarranged objects. I

Cited by 2SourceScholar
2024

Trust Recognition in Human-Robot Cooperation Using EEG

ICRA 2024poster

Collaboration between humans and robots is becoming increasingly crucial in our daily life. In order to accomplish efficient cooperation, trust recognition is vital, empowering robots to predict human behaviors and make trust-aware decisions. Consequently, there is an urgent need for a generalized a…

Cited by 3SourcecodeScholar
2024

Ultrafast capturing in-flight objects with reprogrammable working speed ranges

ICRA 2024poster

In-flight high-speed object capturing is crucial in nature to improve survival and adaptation to the environment, such as the predation of frogs, leopards, and eagles. Despite its ubiquitousness in nature, capturing fast-moving objects is extremely challenging in engineering implementations. In this…

Cited by 0SourceScholar
2024

X-Tacformer : Spatio-tempral Attention Model for Tactile Recognition

ICRA 2024poster

Recently, tactile sensing has attracted great interests in robotics, especially for exploring unstructured objects. Sensor arrays play an important role in the exploration, which generates rich spatio-temporal information. In this work, we propose an efficient tactile recognition model, X-Tacformer.…

Cited by 0SourceScholar
2023

Continuous-Time LiDAR-Inertial-Vehicle Odometry Method with Lateral Acceleration Constraint

ICRA 2023poster

In this paper, we propose a continuous-time-based LiDAR-inertial-vehicle odometry method, which can tightly fuse the data from Light Detection And Ranging (LiDAR), inertial measurement units (IMU), and vehicle measurements. The lateral acceleration constraint is further added to trajectory estimatio…

Cited by 3SourceScholar
2023

GAN-Based Editable Movement Primitive From High-Variance Demonstrations

RA-L 2023

Movement Primitive (MP) is a promising Learning from Demonstration (LfD) framework, which is commonly used to learn movements from human demonstrations and adapt the learned movements to new task scenes. A major goal of MP research is to improve the adaptability of MP to various target positions and

Cited by 4SourceScholar
2023

Goal-Conditioned Reinforcement Learning With Disentanglement-Based Reachability Planning

RA-L 2023

Goal-Conditioned Reinforcement Learning (GCRL) can enable agents to spontaneously set diverse goals to learn a set of skills. Despite the excellent works proposed in various fields, reaching distant goals in temporally extended tasks remains a challenge for GCRL. Current works tackled this problem b

Cited by 6SourceScholar
2023

R-LIOM: Reflectivity-Aware LiDAR-Inertial Odometry and Mapping

RA-L 2023

With the advent of solid-state LiDAR, a series of related studies have boosted the development of Simultaneous Localization and Mapping (SLAM). However, existing methods cannot work well in indoor environments. In the letter, the reflectivity measurement of the solid-state LiDAR is exploited to impr

Cited by 9SourceScholar
2023

s-Adaptive Decoupled Prototype for Few-Shot Object Detection

ICCV 2023poster

Meta-learning-based few-shot detectors use one K-average-pooled prototype (averaging along K-shot dimension) in both Region Proposal Network (RPN) and Detection head (DH) for query detection. Such plain operation would harm the FSOD performance in two aspects: 1) the poor quality of the prototype, a…

Cited by 14PDFScholar
2022

A Novel Neural Multi-Store Memory Network for Autonomous Visual Navigation in Unknown Environment

RA-L 2022

Learning to achieve a user-specified objective from a random position in unseen environments is challenging for image-guided navigation agents. The abilities of long-horizon reasoning and semantic understanding are still lacking. Inspired by the human memory mechanism, we introduce a neural multi-st

Cited by 24SourceScholar
2022

KQA Pro: A Dataset with Explicit Compositional Programs for Complex Question Answering over Knowledge Base

ACL 2022long

Complex question answering over knowledge base (Complex KBQA) is challenging because it requires various compositional reasoning capabilities, such as multi-hop inference, attribute comparison, set operation, etc. Existing benchmarks have some shortcomings that limit the development of Complex KBQA:…

2022

Weakly Supervised Disentangled Representation for Goal-Conditioned Reinforcement Learning

RA-L 2022

Goal-conditioned reinforcement learning is a crucial yet challenging algorithm which enables agents to achieve multiple user-specified goals when learning a set of skills in a dynamic environment. However, it typically requires millions of the environmental interactions explored by agents, which is

Cited by 7SourceScholar
2021

A Global Past-Future Early Exit Method for Accelerating Inference of Pre-trained Language Models

NAACL 2021long

Early exit mechanism aims to accelerate the inference speed of large-scale pre-trained language models. The essential idea is to exit early without passing through all the inference layers at the inference stage. To make accurate predictions for downstream tasks, the hierarchical linguistic informat…

2021

Be Careful about Poisoned Word Embeddings: Exploring the Vulnerability of the Embedding Layers in NLP Models

NAACL 2021long

Recent studies have revealed a security threat to natural language processing (NLP) models, called the Backdoor Attack. Victim models can maintain competitive performance on clean samples while behaving abnormally on samples with a specific trigger word inserted. Previous backdoor attacking methods…

2021

Benchmarking Commonsense Knowledge Base Population with an Effective Evaluation Dataset

EMNLP 2021main

Reasoning over commonsense knowledge bases (CSKB) whose elements are in the form of free-text is an important yet hard task in NLP. While CSKB completion only fills the missing links within the domain of the CSKB, CSKB population is alternatively proposed with the goal of reasoning unseen assertions…

2021

Neural Network Surgery: Injecting Data Patterns into Pre-trained Models with Minimal Instance-wise Side Effects

NAACL 2021long

Side effects during neural network tuning are typically measured by overall accuracy changes. However, we find that even with similar overall accuracy, existing tuning methods result in non-negligible instance-wise side effects. Motivated by neuroscientific evidence and theoretical results, we demon…

Cited by 13SourcePDFScholar