← Search

Huimin LU

48 accepted papers

2026

DeRVOS: Decoupling Consistent Trajectory Generation and Multimodal Understanding for Referring Video Object Segmentation

CVPR 2026

Referring video object segmentation (RVOS) aims to segment objects within a video according to natural language expressions. Unlike earlier works focusing on static single-object scenarios, recent studies address more complex motion scenes. Previous methods typically adopt a query-based, logically m

Cited by 0SourceScholar
2026

Efficient Hierarchical Reinforcement Learning with Dynamic Kolmogorov–Arnold Network for Long-Horizon Robotic Manipulation

ICRA 2026poster

Long-horizon robotic manipulation remains a critical challenge in robotics. Hierarchical reinforcement learning offers a promising solution, but often suffers from an imbalance dilemma: simplifying skill learning increases the complexity of planning, thereby expanding the solution space and computat…

Cited by 0Scholar
2026

Grasp Like Humans: Learning Generalizable Multi-Fingered Grasping from Human Proprioceptive Sensorimotor Integration

ICRA 2026poster

Tactile and kinesthetic perceptions are crucial for human dexterous manipulation, enabling reliable grasping of objects via proprioceptive sensorimotor integration. For robotic hands, even though acquiring such tactile and kinesthetic feedback is feasible, establishing a direct mapping from this sen…

2026

HumanoidExo: Scalable Whole-Body Humanoid Manipulation Via Wearable Exoskeleton

ICRA 2026poster

A significant bottleneck in humanoid policy learning is the acquisition of large-scale, diverse datasets, as collecting reliable real-world data remains both difficult and cost-prohibitive. To address this limitation, we introduce HumanoidExo, a novel system that transfers human motion to whole-body…

2026

PILaN: Generating Task-Individual Independent Customized Assistive Control on a Hip-Knee Powered Exoskeleton

RA-L 2026

Generating task-individual independent customized assistive control is a big challenge of wearable powered exoskeletons, which usually relies on accurate and robust motion intention perceptions (MIP) of the wearer's limb. Traditional physics-based models and deep learning models, as two commonly use

Cited by 0SourceScholar
2026

SurfAAV: Design and Implementation of a Novel Multimodal Surfing Aquatic-Aerial Vehicle

ICRA 2026poster

Despite significant advancements in the research of aquatic-aerial robots, existing configurations struggle to efficiently perform underwater, surface, and aerial movement. In this paper, we propose a novel multimodal surfing aquatic-aerial vehicle, SurfAAV, which efficiently integrates underwater n…

2026

TiCoSS: Tightening the Coupling between Semantic Segmentation and Stereo Matching within a Joint Learning Framework (I)

ICRA 2026poster

Semantic segmentation and stereo matching, respectively analogous to the ventral and dorsal streams in our human brain, are two key components of autonomous driving perception systems. Addressing these two tasks with separate networks is no longer the mainstream direction in developing computer visi…

Cited by 0Scholar
2025

A Novel Decomposed Feature-Oriented Framework for Open-Set Semantic Segmentation on LiDAR Data

ICRA 2025

Semantic segmentation is a key technique that enables mobile robots to understand and navigate surrounding environments autonomously. However, most existing works focus on segmenting known objects, overlooking the identification of unknown classes, which is common in real-world applications. In this

Cited by 1SourcecodeScholar
2025

BEVDiffLoc: End-to-End LiDAR Global Localization in BEV View based on Diffusion Model

IROS 2025

Localization is one of the core parts of modern robotics. Classic localization methods typically follow the retrieve-then-register paradigm, achieving remarkable success. Recently, the emergence of end-to-end localization approaches has offered distinct advantages, including a streamlined system arc

Cited by 1SourcecodeScholar
2025

C-TRAC: Terrain-Adaptive Control for Articulated Tracked Robots via Contact-Aware Reinforcement Learning

IROS 2025

Articulated tracked robots face significant challenges in maintaining stable locomotion over uneven terrain due to unknown contact points between tracks and ground, which are critical for dynamic control. Unlike legged robots, where contact locations can be predicted, tracked systems require real-ti

Cited by 0SourceScholar
2025

DVRP-MHSI: Dynamic Visualization Research Platform for Multimodal Human-Swarm Interaction

RA-L 2025

In recent years, there has been a significant amount of research on algorithms and control methods for distributed collaborative robots. However, the emergence of collective behavior in a swarm is still difficult to predict and control. Nevertheless, human interaction with the swarm helps render the

Cited by 4SourcecodeScholar
2025

Dual-Arm Hierarchical Planning for Laboratory Automation: Vibratory Sieve Shaker Operations

IROS 2025

This paper addresses the challenges of automating vibratory sieve shaker operations in a materials laboratory, focusing on three critical tasks: 1) dual-arm lid manipulation in 3 cm clearance spaces, 2) bimanual handover in overlapping workspaces, and 3) obstructed powder sample container delivery w

Cited by 0SourceScholar
2025

Efficient Instance Motion-Aware Point Cloud Scene Prediction

IROS 2025

Point cloud prediction (PCP) aims to forecast future 3D point clouds of scenes by leveraging sequential historical LiDAR scans, offering a promising avenue to enhance the perceptual capabilities of autonomous systems. However, existing methods mostly adopt an end-to-end approach without explicitly m

Cited by 0SourcecodeScholar
2025

Efficient Multimodal 3D Object Detector via Instance-Level Contrastive Distillation

IROS 2025

Multimodal 3D object detectors leverage the strengths of both geometry-aware LiDAR point clouds and semantically rich RGB images to enhance detection performance. However, the inherent heterogeneity between these modalities, including unbalanced convergence and modal misalignment, poses significant

Cited by 1SourcecodeScholar
2025

Generalizable Zero-Shot Object Pose Estimation for Bin-Picking

ICRA 2025

Unordered grasping in industrial robotic manipulation requires precise six-degree-of-freedom (6D) pose estimation. However, existing methods often struggle with unknown objects and require retraining, limiting their practicality. Traditional 3D point-pair feature methods, while training-free, perfor

Cited by 0SourceScholar
2025

Image-Goal Navigation Using Refined Feature Guidance and Scene Graph Enhancement

IROS 2025

In this paper, we introduce a novel image-goal navigation approach, named RFSG. Our focus lies in leveraging the fine-grained connections between goals, observations, and the environment within limited image data, all the while keeping the navigation architecture simple and lightweight. To this end,

Cited by 3SourcecodeScholar
2025

InsCMPR: Efficient Cross-Modal Place Recognition via Instance-Aware Hybrid Mamba-Transformer

ICRA 2025

Place recognition is an important technique for autonomous mobile robotic applications. While single-modal sensor-based approaches have shown satisfactory performance, cross-modal place recognition remains underexplored due to the challenge of bridging the cross-modal heterogeneity gap. In this work

Cited by 1SourcecodeScholar
2025

Leveraging Semantic Graphs for Efficient and Robust LiDAR SLAM

IROS 2025

Accurate and robust simultaneous localization and mapping (SLAM) is crucial for autonomous mobile systems, typically achieved by leveraging the geometric features of the environment. Incorporating semantics provides a richer scene representation that not only enhances localization accuracy in SLAM b

Cited by 3SourcecodeScholar
2025

LuSeg: Efficient Negative and Positive Obstacles Segmentation via Contrast-Driven Multi-Modal Feature Fusion on the Lunar

IROS 2025

As lunar exploration missions grow increasingly complex, ensuring safe and autonomous rover-based surface exploration has become one of the key challenges in lunar exploration tasks. In this work, we have developed a lunar surface simulation system called the Lunar Exploration Simulator System (LESS

Cited by 3SourcecodeScholar
2025

Multiple Rotation Averaging with Constrained Reweighting Deep Matrix Factorization

ICRA 2025

Multiple rotation averaging plays a crucial role in computer vision and robotics domains. The conventional optimization-based methods optimize a nonlinear cost function based on certain noise assumptions, while most previous learning-based methods require ground truth labels in the supervised traini

Cited by 0SourceScholar
2025

NeuroVE: Brain-Inspired Linear-Angular Velocity Estimation With Spiking Neural Networks

RA-L 2025

Vision-based ego-velocity estimation is a fundamental problem in robot state estimation. However, the constraints of frame-based cameras, including motion blur and insufficient frame rates in dynamic settings, readily lead to the failure of conventional velocity estimation techniques. Mammals exhibi

Cited by 5SourceScholar
2025

NuExo: A Wearable Exoskeleton Covering all Upper Limb ROM for Outdoor Data Collection and Teleoperation of Humanoid Robots

IROS 2025

The evolution from motion capture and teleoperation to robot skill learning has emerged as a hotspot and critical pathway for advancing embodied intelligence. However, existing systems still face a persistent gap in simultaneously achieving four objectives: accurate tracking of full upper limb movem

Cited by 7SourcecodeScholar
2025

RLCNet: A Novel Deep Feature-Matching-Based Method for Online Target-Free Radar-LiDAR Calibration

ICRA 2025

While millimeter-wave radars are widely used in robotics and autonomous driving, extrinsic calibration with other sensors remains challenging due to the sparsity and uncertainty of radar point clouds. In this paper, we propose a novel deep feature-matching-based online extrinsic calibration approach

Cited by 0SourcecodeScholar
2025

ResLPR: A LiDAR Data Restoration Network and Benchmark for Robust Place Recognition Against Weather Corruptions

IROS 2025

LiDAR-based place recognition (LPR) is a key component for autonomous driving, and its resilience to environmental corruption is critical for safety in high-stakes applications. While state-of-the-art (SOTA) LPR methods perform well in clean weather, they still struggle with weather-induced corrupti

Cited by 7SourcecodeScholar
2025

Self-Supervised Diffusion-Based Scene Flow Estimation and Motion Segmentation With 4D Radar

RA-L 2025

Scene flow estimation (SFE) and motion segmentation (MOS) using 4D radar are emerging yet challenging tasks in robotics and autonomous driving applications. Existing LiDAR- or RGB-D-based point cloud processing methods often deliver suboptimal performance on radar data due to radar signals' highly s

Cited by 1SourcecodeScholar
2025

SurfAAV: Design and Implementation of a Novel Multimodal Surfing Aquatic-Aerial Vehicle

RA-L 2025

Despite significant advancements in the research of aquatic-aerial robots, existing configurations struggle to efficiently perform underwater, surface, and aerial movement. In this paper, we propose a novel multimodal surfing aquaticaerial vehicle, SurfAAV, which efficiently integrates underwater na

Cited by 1SourceScholar
2025

UGNA-VPR: A Novel Training Paradigm for Visual Place Recognition Based on Uncertainty-Guided NeRF Augmentation

RA-L 2025

Visual place recognition (VPR) is crucial for robots to identify previously visited locations, playing an important role in autonomous navigation in both indoor and outdoor environments. However, most existing VPR datasets are limited to single-viewpoint scenarios, leading to reduced recognition acc

Cited by 1SourcecodeScholar
2025

UniDetox: Universal Detoxification of Large Language Models via Dataset Distillation

ICLR 2025poster

We present UniDetox, a universally applicable method designed to mitigate toxicity across various large language models (LLMs). Previous detoxification methods are typically model-specific, addressing only individual models or model families, and require careful hyperparameter tuning due to the trad…

2024

Diffusion-Based Point Cloud Super-Resolution for mmWave Radar Data

ICRA 2024poster

The millimeter-wave radar sensor maintains stable performance under adverse environmental conditions, making it a promising solution for all-weather perception tasks, such as outdoor mobile robotics. However, the radar point clouds are relatively sparse and contain massive ghost points, which greatl…

Cited by 7SourceScholar
2024

RadarMOSEVE: A Spatial-Temporal Transformer Network for Radar-Only Moving Object Segmentation and Ego-Velocity Estimation

AAAI 2024technical

Moving object segmentation (MOS) and Ego velocity estimation (EVE) are vital capabilities for mobile systems to achieve full autonomy. Several approaches have attempted to achieve MOSEVE using a LiDAR sensor. However, LiDAR sensors are typically expensive and susceptible to adverse weather condition…

2024

SGLC: Semantic Graph-Guided Coarse-Fine-Refine Full Loop Closing for LiDAR SLAM

RA-L 2024

Loop closing is a crucial component in SLAM that helps eliminate accumulated errors through two main steps: loop detection and loop pose correction. The first step determines whether loop closing should be performed, while the second estimates the 6-DoF pose to correct odometry drift. Current method

Cited by 12SourcecodeScholar
2024

Spatio-Temporal Calibration for Omni-Directional Vehicle-Mounted Event Cameras

RA-L 2024

We present a solution to the problem of spatio-temporal calibration for event cameras mounted on an onmi-directional vehicle. Different from traditional methods that typically determine the camera's pose with respect to the vehicle's body frame using alignment of trajectories, our approach leverages

Cited by 7SourcecodeScholar
2024

SuperFusion: Multilevel LiDAR-Camera Fusion for Long-Range HD Map Generation

ICRA 2024poster

High-definition (HD) semantic map generation of the environment is an essential component of autonomous driving. Existing methods have achieved good performance in this task by fusing different sensor modalities, such as LiDAR and camera. However, current works are based on raw data or network featu…

Cited by 54SourcecodeScholar
2024

TSCM: A Teacher-Student Model for Vision Place Recognition Using Cross-Metric Knowledge Distillation

ICRA 2024poster

Visual place recognition (VPR) plays a pivotal role in autonomous exploration and navigation of mobile robots within complex outdoor environments. While cost-effective and easily deployed, camera sensors are sensitive to lighting and weather changes, and even slight image alterations can greatly aff…

Cited by 1SourcecodeScholar
2023

ElC-OIS: Ellipsoidal Clustering for Open-World Instance Segmentation on LiDAR Data

IROS 2023poster

Open-world Instance Segmentation (OIS) is a challenging task that aims to accurately segment every object instance appearing in the current observation, regardless of whether these instances have been labeled in the training set. This is important for safety-critical applications such as robust auto…

Cited by 4SourcecodeScholar
2023

Hybrid Map-Based Path Planning for Robot Navigation in Unstructured Environments

IROS 2023poster

Fast and accurate path planning is important for ground robots to achieve safe and efficient autonomous navigation in unstructured outdoor environments. However, most existing methods exploiting either 2D or 2.5D maps struggle to balance the efficiency and safety for ground robots navigating in such…

Cited by 14SourcecodeScholar
2023

InsMOS: Instance-Aware Moving Object Segmentation in LiDAR Data

IROS 2023poster

Identifying moving objects is a crucial capability for autonomous navigation, consistent map generation, and future trajectory prediction of objects. In this paper, we propose a novel network that addresses the challenge of segmenting moving objects in 3D LiDAR scans. Our approach not only predicts…

Cited by 32SourcecodeScholar
2022

Feature Distillation Interaction Weighting Network for Lightweight Image Super-resolution

AAAI 2022technical

Convolutional neural networks based single-image superresolution (SISR) has made great progress in recent years. However, it is difficult to apply these methods to real-world scenarios due to the computational and memory cost. Meanwhile, how to take full advantage of the intermediate features under…

2021

A Shared Control Framework for Human-Multirobot Foraging With Brain-Computer Interface

RA-L 2021

With the rapid development of multi-robot systems (MRSs), they can be widely used to perform various tasks in typical environments. However, the inevitable disadvantages of onboard sensor errors, communication delays, and underspecified environmental factors seriously affect the operation of MRSs. T

Cited by 7SourceScholar
2021

Enhancing Audio-Visual Association with Self-Supervised Curriculum Learning

AAAI 2021technical

The recent success of audio-visual representations learning can be largely attributed to their pervasive concurrency property, which can be used as a self-supervision signal and extract correlation information. While most recent works focus on capturing the shared associations between the audio and…

Cited by 26SourcePDFScholar
2021

Keypoint Matching for Point Cloud Registration Using Multiplex Dynamic Graph Attention Networks

RA-L 2021

The registration of point clouds is a key ingredient of LiDAR-based SLAM systems and mapping approaches. A challenging task in this context is finding the right data association between 3D points. This paper proposes a novel and flexible graph network architecture to tackle the keypoint matching pro

Cited by 54SourceScholar
2021

Partial Feature Selection and Alignment for Multi-Source Domain Adaptation

CVPR 2021poster

Multi-Source Domain Adaptation (MSDA), which dedicates to transfer the knowledge learned from multiple source domains to an unlabeled target domain, has drawn increasing attention in the research community. By assuming that the source and target domains share consistent key feature representations a…

Cited by 41PDFScholar
2021

Robust Motion Averaging under Maximum Correntropy Criterion

ICRA 2021poster

Recently, the motion averaging method has been introduced as an effective means to solve the multi-view registration problem. This method aims to recover global motions from a set of relative motions, where the original method is sensitive to outliers due to using the Frobenius norm error in the opt…

Cited by 9SourceScholar
2021

Shared Control Based on a Brain-Computer Interface for Human-Multirobot Cooperation

RA-L 2021

Currently, distributed multi-robot systems (MRSs) can meet the requirements of various tasks in complex environments. Nevertheless, the inevitable disadvantages of robot sensor errors, communication delays, and obstructive environmental factors hinder the operation of MRSs. Therefore, a shared contr

Cited by 14SourceScholar
2019

CAN: Contextual Aggregating Network for Semantic Segmentation

ICASSP 2019accepted

Fully convolutional neural networks (FCNs) have shown great success in dense estimation tasks. One key pillar of such progress is mining multi-scale context cues from features in different convolutional layers. This paper introduces contextual aggregating network(CAN), a generic convolutional featur…

Cited by 0SourceScholar
2015

Long range traversable region detection based on superpixels clustering for mobile robots

IROS 2015poster

Traversable region detection is important for autonomous visual navigation of mobile robots. Only short range traversable regions can be detected using traditional methods based on stereo vision because of the limited image resolution and baseline of stereo vision. In this paper, we propose a novel…

Cited by 17SourceScholar