← Search

Hesheng Wang

96 accepted papers

2026

ARFlow: Auto-regressive Optical Flow Estimation for Arbitrary-Length Videos via Progressive Next-Frame Forecasting

ICLR 2026poster

Optical flow estimation is a fundamental computer vision task that predicts per-pixel displacements from consecutive images. Recent works attempt to exploit temporal cues to improve the estimation performance. However, their temporal modeling is restricted to short video sequences due to the unaffor…

Cited by 0SourceScholar
2026

ARMFlow: AutoRegressive MeanFlow for Online 3D Human Reaction Generation

CVPR 2026

3D human reaction generation faces three main challenges: (1) high motion fidelity, (2) real-time inference, and (3) autoregressive adaptability for online scenarios. Existing methods fail to meet all three simultaneously. We propose ARMFlow, a MeanFlow-based autoregressive framework that models tem

Cited by 0SourcecodeScholar
2026

DIAL-GS: Dynamic Instance Aware Reconstruction for Label-Free Street Scenes with 4D Gaussian Splatting

ICRA 2026poster

Urban scene reconstruction is critical for autonomous driving, enabling structured 3D representations for data synthesis and closed-loop testing. Supervised approaches rely on costly human annotations and lack scalability, while current self-supervised methods often confuse static and dynamic elemen…

2026

Delving into Non-Exchangeability for Conformal Prediction in Graph-Structured Multivariate Time Series

ICML 2026poster

Point forecasting for graph-structured multivariate time series is a fundamental problem, but rigorous uncertainty quantification for such predictions is still underexplored. Conformal prediction (CP) offers uncertainty estimation with a solid coverage guarantee under the exchangeability assumption,…

Cited by 0SourceScholar
2026

Diffusion Stabilizer Policy for Automated Surgical Robot Manipulations

ICRA 2026poster

Intelligent surgical robots have the potential to revolutionize clinical practice by enabling more precise and automated surgical procedures. However, the automation of such robot for surgical tasks remains under-explored compared to recent advancements in solving household manipulation tasks. These…

2026

GaussianDWM: 3D Gaussian Driving World Model for Unified Scene Understanding and Multi-Modal Generation

CVPR 2026

Driving World Models (DWMs) have been developing rapidly with the advances of generative models. However, existing DWMs lack 3D scene understanding capabilities and can only generate content conditioned on input data, without the ability to interpret or reason about the driving environment. Moreover

Cited by 0SourcecodeScholar
2026

Leveraging Geometric Prior Uncertainty and Complementary Constraints for High-Fidelity Neural Indoor Surface Reconstruction

ICRA 2026poster

Neural implicit surface reconstruction with signed distance function has made significant progress, but recovering fine details such as thin structures and complex geometries remains challenging due to unreliable or noisy geometric priors. Existing approaches rely on implicit uncertainty that arises…

2026

MRASfM: Multi-Camera Reconstruction and Aggregation through Structure-From-Motion in Driving Scenes

ICRA 2026poster

Structure from Motion (SfM) estimates camera poses and reconstructs point clouds, forming a foundation for various tasks. However, applying SfM to driving scenes captured by multi-camera systems presents significant difficulties, including unreliable pose estimation, excessive outliers in road surfa…

2026

S$^2$-MLLM: Boosting Spatial Reasoning Capability of MLLMs for 3D Visual Grounding with Structural Guidance

CVPR 2026

3D Visual Grounding (3DVG) focuses on locating objects in 3D scenes based on natural language descriptions, serving as a fundamental task for embodied AI and robotics. Recent advances in Multi-modal Large Language Models (MLLMs) have motivated research into extending them to 3DVG. However, MLLMs pri

Cited by 0SourcecodeScholar
2026

StreamVLO: Streaming Visual-LiDAR Odometry with Cumulative Drift Compensation

CVPR 2026

We propose StreamVLO, a streaming visual-LiDAR odometry framework that performs unified spatio-temporal correlation with Mamba models and tackles the long-standing cumulative drift problem via an online Cumulative Drift Compensation scheme for localization in 4D dynamic environments. Specifically, S

Cited by 0SourceScholar
2026

UTTG: A Universal Teleoperation Framework Via Online Trajectory Generation

ICRA 2026poster

Teleoperation is crucial for hazardous environment operations and serves as a key tool for collecting expert demonstrations in robot learning. However, existing methods face robotic hardware dependency and control frequency mismatches between teleoperation devices and robotic platforms. Our approach…

Cited by 0codeScholar
2026

WeatherCity: Urban Scene Reconstruction with Controllable Multi-Weather Transformation

CVPR 2026

Editable high-fidelity 4D scenes are crucial for autonomous driving, as they can be applied to end-to-end training and closed-loop simulation. However, existing reconstruction methods are primarily limited to replicating observed scenes and lack the capability for diverse weather simulation. While i

Cited by 0SourcecodeScholar
2025

ASCENT: Autonomous Skill Learning Toward Complex Embodied Tasks With Foundation Models

ICRA 2025

Collecting data from simulated scenarios for training robotic skills provides a safer and more controllable alternative to real-world environments. However, it demands considerable effort, including the manual construction of simulation environments, the careful design of tasks, and the challenge of

Cited by 0SourceScholar
2025

CrossBEV-PR: Cross-modal Visual-LiDAR Place Recognition via BEV Feature Distillation

IROS 2025

Utilizing 2D images for place recognition within 3D point cloud maps presents significant challenges in autonomous driving applications, primarily due to the inherent cross-modal disparity between visual and LiDAR data. In this study, we propose a novel cross-modal visual-LiDAR place recognition met

Cited by 0SourcecodeScholar
2025

DVN-SLAM: Dynamic Visual Neural Slam Based on Local-Global Encoding

ICRA 2025

Recent research on Simultaneous Localization and Mapping (SLAM) based on implicit representation has shown promising results in indoor environments. However, some challenges remain: the limited scene representation capability of implicit encoding, the uncertainty in the rendering process from implic

Cited by 12SourceScholar
2025

Deformable Gaussian Splatting for Efficient and High-Fidelity Reconstruction of Surgical Scenes

ICRA 2025

Efficient and high-fidelity reconstruction of deformable surgical scenes is a critical yet challenging task. Building on recent advancements in 3D Gaussian splatting, current methods have seen significant improvements in both reconstruction quality and rendering speed. However, two major limitations

Cited by 5SourceScholar
2025

Diff-IP2D: Diffusion-Based Hand-Object Interaction Prediction on Egocentric Videos

IROS 2025

Understanding how humans would behave during hand-object interaction (HOI) is vital for applications in service robot manipulation and extended reality. To achieve this, some recent works simultaneously forecast hand trajectories and object affordances on human egocentric videos. The joint predictio

Cited by 22SourcecodeScholar
2025

FLAME: Learning to Navigate with Multimodal LLM in Urban Environments

AAAI 2025technical

Large Language Models (LLMs) have demonstrated potential in Vision-and-Language Navigation (VLN) tasks, yet current applications face challenges. While LLMs excel in general conversation scenarios, they struggle with specialized navigation tasks, yielding suboptimal performance compared to specializ…

2025

Foresee and Act Ahead: Task Prediction and Pre-Scheduling Enabled Efficient Robotic Warehousing

ICRA 2025

In warehousing systems, to enhance efficiency amid surging demand volumes, much attention has been placed on how to reasonably allocate tasks of delivery to robots. However, the labor of robots is still inevitably wasted to some extent. In this paper, we propose a pre-scheduling enhanced warehousing

Cited by 1SourceScholar
2025

FreeDriveRF: Monocular RGB Dynamic NeRF Without Poses for Autonomous Driving via Point-Level Dynamic-Static Decoupling

ICRA 2025

Dynamic scene reconstruction for autonomous driving enables vehicles to perceive and interpret complex scene changes more precisely. Dynamic Neural Radiance Fields (NeRFs) have recently shown promising capability in scene modeling. However, many existing methods rely heavily on accurate poses inputs

Cited by 4SourcecodeScholar
2025

Fusion Scene Context: Robust and Efficient LiDAR Place Recognition Across Season

IROS 2025

Place recognition is an important component for autonomous robot navigation. Many existing LiDAR-based place recognition methods encode the structural information of 3D LiDAR data into 2D image representations. However, most of these intermediates only exploit the projection in a single view, ignori

Cited by 0SourceScholar
2025

Improved 2D Hand Trajectory Prediction with Multi-View Consistency

IROS 2025

Forecasting how human hands would move around target objects on egocentric videos can provide prior knowledge to enhance the path planning capabilities of service robots and assistive wearable devices. During the hand-object interaction process, head movements always occur concurrently to provide ob

Cited by 0SourcecodeScholar
2025

MNE-SLAM: Multi-Agent Neural SLAM for Mobile Robots

CVPR 2025poster

Neural implicit scene representations have recently shown promising results in dense visual SLAM. However, existing implicit SLAM algorithms are constrained to single-agent scenarios, and fall difficulty in large indoor scenes and long sequences. Existing multi-agent SLAM frameworks cannot meet the…

2025

Novel Diffusion Models for Multimodal 3D Hand Trajectory Prediction

IROS 2025

Predicting hand motion is critical for understanding human intentions and bridging the action space between human movements and robot manipulations. Existing hand trajectory prediction (HTP) methods forecast the future hand waypoints in 3D space conditioned on past egocentric observations. However,

Cited by 5SourcecodeScholar
2025

Planning from Imagination: Episodic Simulation and Episodic Memory for Vision-and-Language Navigation

AAAI 2025technical

Humans navigate unfamiliar environments using episodic simulation and episodic memory, which facilitate a deeper understanding of the complex relationships between environments and objects. Developing an imaginative memory system inspired by human mechanisms can enhance the navigation performance of…

Cited by 0SourcePDFScholar
2025

RL-GSBridge: 3D Gaussian Splatting Based Real2Sim2Real Method for Robotic Manipulation Learning

ICRA 2025

Sim-to-Real refers to the process of transferring policies learned in simulation to the real world, which is crucial for achieving practical robotics applications. However, recent Sim2real methods either rely on a large amount of augmented data or large learning models, which is inefficient for spec

Cited by 2SourceScholar
2025

SGLoc: Semantic Localization System for Camera Pose Estimation from 3D Gaussian Splatting Representation

IROS 2025

We propose SGLoc, a novel localization system that directly regresses camera poses from 3D Gaussian Splatting (3DGS) representation by leveraging semantic information. Our method utilizes the semantic relationship between 2D image and 3D scene representation to estimate the 6DoF pose without prior p

Cited by 1SourcecodeScholar
2025

Seeing through Uncertainty: Robust Task-Oriented Optimization in Visual Navigation

NeurIPS 2025poster

Visual navigation is a fundamental problem in embodied AI, yet practical deployments demand long-horizon planning capabilities to address multi-objective tasks. A major bottleneck is data scarcity: policies learned from limited data often overfit and fail to generalize OOD. Existing neural network-b…

Cited by 0SourceScholar
2025

SemAlign3D: Semantic Correspondence between RGB-Images through Aligning 3D Object-Class Representations

CVPR 2025poster

Semantic correspondence made tremendous progress through the recent advancements of large vision models (LVM). While these LVMs have been shown to reliably capture local semantics, the same can currently not be said for capturing global geometric relationships between semantic object regions. This p…

Cited by 0SourcePDFScholar
2025

TopoLiDM: Topology-Aware LiDAR Diffusion Models for Interpretable and Realistic LiDAR Point Cloud Generation

IROS 2025

LiDAR scene generation is critical for mitigating real-world LiDAR data collection costs and enhancing the robustness of downstream perception tasks in autonomous driving. However, existing methods commonly struggle to capture geometric realism and global topological consistency. Recent LiDAR Diffus

Cited by 5SourcecodeScholar
2025

Towards Autonomous Indoor Parking: A Globally Consistent Semantic SLAM System and A Semantic Localization Subsystem

IROS 2025

We propose a globally consistent semantic SLAM system (GCSLAM) and a semantic-fusion localization subsystem (SF-Loc), which achieves accurate semantic mapping and robust localization in complex parking lots. Visual cameras (front-view and surround-view), IMU, and wheel encoder form the input sensor

Cited by 3SourceScholar
2025

Wonder Wins Ways: Curiosity-Driven Exploration through Multi-Agent Contextual Calibration

NeurIPS 2025poster

Autonomous exploration in complex multi-agent reinforcement learning (MARL) with sparse rewards critically depends on providing agents with effective intrinsic motivation. While artificial curiosity offers a powerful self-supervised signal, it often confuses environmental stochasticity with meaningf…

Cited by 0SourceScholar
2024

3DSFLabelling: Boosting 3D Scene Flow Estimation by Pseudo Auto-labelling

CVPR 2024poster

Learning 3D scene flow from LiDAR point clouds presents significant difficulties including poor generalization from synthetic datasets to real scenes scarcity of real-world 3D labels and poor performance on real sparse LiDAR point clouds. We present a novel approach from the perspective of auto-labe…

2024

An Origami-Inspired Pneumatic Continuum Module with Active Variable Stiffness

IROS 2024poster

This paper presents a novel pneumatic continuum module featuring high contraction-ratio, bidirectional actuation, and active stiffness regulation. The module comprises four linear pneumatic actuators integrating rigid polygon origami frame into soft bellow. This integration not only helps to regulat…

Cited by 0SourceScholar
2024

Cam4DOcc: Benchmark for Camera-Only 4D Occupancy Forecasting in Autonomous Driving Applications

CVPR 2024poster

Understanding how the surrounding environment changes is crucial for performing downstream tasks safely and reliably in autonomous driving applications. Recent occupancy estimation techniques using only camera images as input can provide dense occupancy representations of large-scale scenes based on…

2024

Cooperative Path Planning for Four-Way Shuttle Vehicles in Storage and Retrieval Systems: A Hierarchically Dynamic Graph-Based Approach

IROS 2024poster

Recently, Shuttle-based Storage and Retrieval Systems (SBS/RSs) have garnered significant attention from both academia and industry, owing to their high spatial utilization and rapid response speed. However, the weak connectivity of roadmaps in densely stored environments increases the likelihood of…

Cited by 0SourceScholar
2024

DDS-SLAM: Dense Semantic Neural SLAM for Deformable Endoscopic Scenes

IROS 2024poster

Estimating camera motion and continuously reconstructing dense scenes in deformable environments presents a complex and open challenge. Many existing approaches tend to rely on assumptions about the scene’s topology or the nature of deformable motion. However, these assumptions do not hold true in m…

Cited by 2SourcecodeScholar
2024

DSLO: Deep Sequence LiDAR Odometry Based on Inconsistent Spatio-temporal Propagation

IROS 2024poster

This paper introduces a 3D point cloud sequence learning model based on inconsistent spatio-temporal propagation for LiDAR odometry, termed DSLO. It consists of a pyramid structure with a spatial information reuse strategy, a sequential pose initialization module, a gated hierarchical pose refinemen…

Cited by 0SourcecodeScholar
2024

Decentralized Multi-Robot Navigation Coupled with Spatial-Temporal RetNet Based on Deep Reinforcement Learning

IROS 2024poster

Navigating robots through dynamic multi-robot environments, avoiding collisions with both other robots and obstacles, has emerged as a central challenge in robotics. The existing approaches fall short in allowing the policy network to effectively capture spatial-temporal reciprocal collision avoidan…

Cited by 0SourceScholar
2024

Decentralized Trajectory Planning for Formation Flight in Unknown and Dense Environments

IROS 2024poster

For aerial swarms, formation flight has been applied in various scenes. However, most existing works do not consider balancing the conflicting requirements among keeping formation, keeping the smoothness of trajectories, and obstacle avoidance within the limited time. To address this issue, we propo…

Cited by 0SourceScholar
2024

DifFlow3D: Toward Robust Uncertainty-Aware Scene Flow Estimation with Iterative Diffusion-Based Refinement

CVPR 2024poster

Scene flow estimation which aims to predict per-point 3D displacements of dynamic scenes is a fundamental task in the computer vision field. However previous works commonly suffer from unreliable correlation caused by locally constrained searching ranges and struggle with accumulated inaccuracy aris…

2024

Enhancing Exploratory Capability of Visual Navigation Using Uncertainty of Implicit Scene Representation

IROS 2024

In the context of visual navigation in unknown scenes, both “exploration” and “exploitation” are equally crucial. Robots must first establish environmental cognition through exploration and then utilize the cognitive information to accomplish target searches. However, most existing methods for image

Cited by 2SourcecodeScholar
2024

LHMap-loc: Cross-Modal Monocular Localization Using LiDAR Point Cloud Heat Map

ICRA 2024poster

Localization using a monocular camera in the pre-built LiDAR point cloud map has drawn increasing attention in the field of autonomous driving and mobile robotics. However, there are still many challenges (e.g. difficulties of map storage, poor localization robustness in large scenes) in accurately…

Cited by 1SourcecodeScholar
2024

Multi-Agent Teamwise Cooperative Path Finding and Traffic Intersection Coordination

IROS 2024poster

When coordinating the motion of connected autonomous vehicles at a signal-free intersection, the vehicles from each direction naturally forms a team and each team seeks to minimize their own traversal time through the intersection, without concerning the traversal times of other teams. Since the int…

Cited by 1SourceScholar
2024

SNI-SLAM: Semantic Neural Implicit SLAM

CVPR 2024poster

We propose SNI-SLAM a semantic SLAM system utilizing neural implicit representation that simultaneously performs accurate semantic mapping high-quality surface reconstruction and robust camera tracking. In this system we introduce hierarchical semantic representation to allow multi-level semantic co…

2024

SoftNeRF: A Self-Modeling Soft Robot Plugin for Various Tasks

IROS 2024poster

Building a self-model for robots, enabling them to simulate their physical selves and predict future states without direct interaction with the physical world, is crucial for robot motion planning and control. Existing self-modeling methods primarily focus on rigid robots and typically require signi…

Cited by 0SourcecodeScholar
2024

Spherical Frustum Sparse Convolution Network for LiDAR Point Cloud Semantic Segmentation

NeurIPS 2024poster

LiDAR point cloud semantic segmentation enables the robots to obtain fine-grained semantic information of the surrounding environment. Recently, many works project the point cloud onto the 2D image and adopt the 2D Convolutional Neural Networks (CNNs) or vision transformer for LiDAR point cloud sema…

2024

Toward Universal and Scalable Road Graph Partitioning for Efficient Multi-Robot Path Planning

IROS 2024

To date, multi-robot path planning has primarily been addressed by centralized solvers, typically aiming to maintain optimality. However, given its NP-hard nature, directly applying existing solvers in large and complex scenarios proves inefficient. A promising alternative lies in adopting a divide-

Cited by 1SourceScholar
2023

DELFlow: Dense Efficient Learning of Scene Flow for Large-Scale Point Clouds

ICCV 2023poster

Point clouds are naturally sparse, while image pixels are dense. The inconsistency limits feature fusion from both modalities for point-wise scene flow estimation. Previous methods rarely predict scene flow from the entire point clouds of the scene with one-time inference due to the memory inefficie…

Cited by 11PDFcodeScholar
2023

Lyapunov Constrained Safe Reinforcement Learning for Multicopter Visual Servoing

IROS 2023poster

Traditional methods based on Lyapunov analysis and learning-based approaches such as reinforcement learning (RL) are two powerful tools in visual servo tasks. Traditional methods are interpretable and their stability can be guar-anteed by Lyapunov analysis. However, they tend to have a high dependen…

Cited by 0SourceScholar
2023

RLSAC: Reinforcement Learning Enhanced Sample Consensus for End-to-End Robust Estimation

ICCV 2023poster

Robust estimation is a crucial and still challenging task, which involves estimating model parameters in noisy environments. Although conventional sampling consensus-based algorithms sample several times to achieve robustness, these algorithms cannot use data features and historical information effe…

Cited by 6PDFcodeScholar
2023

RegFormer: An Efficient Projection-Aware Transformer Network for Large-Scale Point Cloud Registration

ICCV 2023poster

Although point cloud registration has achieved remarkable advances in object-level and indoor scenes, large-scale registration methods are rarely explored. Challenges mainly arise from the huge point number, complex distribution, and outliers of outdoor LiDAR scans. In addition, most existing regist…

Cited by 58PDFcodeScholar
2023

SeasonDepth: Cross-Season Monocular Depth Prediction Dataset and Benchmark Under Multiple Environments

IROS 2023poster

Different environments pose a great challenge to the outdoor robust visual perception for long-term autonomous driving, and the generalization of learning-based algorithms on different environments is still an open problem. Although monocular depth prediction has been well studied recently, few work…

Cited by 20SourcecodeScholar
2023

Self-supervised Multi-frame Monocular Depth Estimation with Pseudo-LiDAR Pose Enhancement

ICRA 2023poster

Depth estimation is one of the most important tasks in scene understanding. In the existing joint self-supervised learning approaches of depth-pose estimation, depth estimation and pose estimation networks are independent of each other. They only use the adjacent image frames for pose estimation and…

Cited by 5SourceScholar
2023

TransLO: A Window-Based Masked Point Transformer Framework for Large-Scale LiDAR Odometry

AAAI 2023technical

Recently, transformer architecture has gained great success in the computer vision community, such as image classification, object detection, etc. Nonetheless, its application for 3D vision remains to be explored, given that point cloud is inherently sparse, irregular, and unordered. Furthermore, ex…

2023

VDBblox: Accurate and Efficient Distance Fields for Path Planning and Mesh Reconstruction

IROS 2023poster

Highly accurate and efficient map in unknown and complex environments is essential for robotics navigation. Traditionally, mobile robot platforms are often computationally constrained when using multiple sensors to process large amounts of input data. In previous works, some of them have been deploy…

Cited by 3SourcecodeScholar
2022

FusionNet: Coarse-to-Fine Extrinsic Calibration Network of LiDAR and Camera with Hierarchical Point-pixel Fusion

ICRA 2022poster

In this paper, we propose a novel network, Fusion-Net, which can estimate the extrinsic calibration matrix between LiDAR and a monocular RGB camera with high accuracy and robustness. FusionNet is a coarse-to-fine method, providing an online and end-to-end solution that can automatically detect and c…

Cited by 18SourceScholar
2022

Keypoint-Based Planar Bimanual Shaping of Deformable Linear Objects Under Environmental Constraints With Hierarchical Action Framework

RA-L 2022

This letter addresses the problem of contact-based manipulation of deformable linear objects (DLOs) towards desired shapes with a dual-arm robotic system. To alleviate the burden of high-dimensional continuous state-action spaces, we model DLOs as kinematic multibody systems via our proposed keypoin

Cited by 42SourceScholar
2022

LNC Assisted Localization and Mapping in Pipe Environment

IROS 2022poster

Regular maintenance of pipelines is an important task to ensure oil transportation and other operation (sewers, nature gas). Precise localization of pipeline damage can greatly improve the efficiency of maintenance work. Since the texture similarity and illumination change of pipe, traditional local…

Cited by 3SourceScholar
2022

Vision-Based Contact Point Selection for the Fully Non-Fixed Contact Manipulation of Deformable Objects

RA-L 2022

Most of the research on deformable objects manipulation (DOM) is under the assumption of fixed contact (FC). Although there are some attempts to break this assumption, research on DOM with the fully non-fixed contact (FNFC), which refers that the object is not fixedly connected to both the end-effec

Cited by 7SourceScholar
2022

What Matters for 3D Scene Flow Network

ECCV 2022poster

"3D scene flow estimation from point clouds is a low-level 3D motion perception task in computer vision. Flow embedding is a commonly used technique in scene flow estimation, and it encodes the point motion between two consecutive frames. Thus, it is critical for the flow embeddings to capture the c…

2021

A Registration-aided Domain Adaptation Network for 3D Point Cloud Based Place Recognition

IROS 2021poster

In the field of large-scale SLAM for autonomous driving and mobile robotics, 3D point cloud based place recognition has aroused significant research interest due to its robustness to changing environments with drastic daytime and weather variance. However, it is time-consuming and effort-costly to o…

Cited by 11SourceScholar
2021

Distributed Rendezvous Control of Networked Uncertain Robotic Systems with Bearing Measurements

ICRA 2021poster

In this paper, the distributed rendezvous control problem of networked uncertain robotic systems with bearing measurements is investigated. The network topology of the multi-robot systems is described by an undirected graph. The dynamics of robots is modeled by Euler-Lagrange equation with unknown i…

Cited by 0SourceScholar
2021

Hybrid Vision/Force Control for Interaction with the Bottle-like Object

ICRA 2021poster

This study proposes a hybrid vision/force control scheme for interaction with the inner surface of the bottle-like object. Based on the geometry of the object, a new generalized constraint called the bottleneck (BN) constraint is proposed, which ensures the tool passes through a fixed 3-D region and…

Cited by 1SourceScholar
2021

PWCLO-Net: Deep LiDAR Odometry in 3D Point Clouds Using Hierarchical Embedding Mask Optimization

CVPR 2021poster

A novel 3D point cloud learning model for deep LiDAR odometry, named PWCLO-Net, using hierarchical embedding mask optimization is proposed in this paper. In this model, the Pyramid, Warping, and Cost volume (PWC) structure for the LiDAR odometry task is built to refine the estimated pose in a coarse…

Cited by 81PDFcodeScholar
2021

Soft Manipulator Fault Detection and Identification Using ANC-based LSTM

IROS 2021poster

Timely fault detection and identification (FDI) of soft manipulators are critical in the design of surgical systems to improve reliability. However, due to the intrinsic compliance of soft manipulators, their end effectors vibrate during the dynamic control process, which introduces noise into the m…

Cited by 7SourcecodeScholar
2021

Toward State-Unsaturation Guaranteed Fault Detection Method in Visual Servoing of Soft Robot Manipulators

IROS 2021poster

This paper puts forward a novel sensor-less fault detection method with only task errors feedback and applies it to visual servoing tasks of soft robot manipulators. The method is developed by introducing a suitably designed endogenous accessory signal (EAS). On the one hand, EAS transforms the chan…

Cited by 4SourceScholar
2021

Towards Collision Detection, Localization and Force Estimation for a Soft Cable-driven Robot Manipulator

ICRA 2021poster

Soft robots have been applied widely to various constrained scenarios due to the advantages over traditional rigid manipulators such as softness, deformability and adaptability to constrained surroundings. To make full use of this merit, this paper proposes a method that integrates collision detecti…

Cited by 1SourceScholar
2021

Unsupervised Learning of 3D Scene Flow from Monocular Camera

ICRA 2021poster

Scene flow represents the motion of points in the 3D space, which is the counterpart of the optical flow that represents the motion of pixels in the 2D image. However, it is difficult to obtain the ground truth of scene flow in the real scenes, and recent studies are based on synthetic data for trai…

Cited by 19SourcecodeScholar
2020

A Synchronization Approach for Achieving Cooperative Adaptive Cruise Control Based Non-Stop Intersection Passing

ICRA 2020poster

Cooperative adaptive cruise control (CACC) of intelligent vehicles contributes to improving cruise control performance, reducing traffic congestion, saving energy and increasing traffic flow capacity. In this paper, we resolve the CACC problem from the viewpoint of synchronization control, our main…

Cited by 7SourceScholar
2020

End-to-End 3D Point Cloud Learning for Registration Task Using Virtual Correspondences

IROS 2020poster

3D Point cloud registration is still a very challenging topic due to the difficulty in finding the rigid transformation between two point clouds with partial correspondences, and it's even harder in the absence of any initial estimation information. In this paper, we present an end-to-end deep-learn…

Cited by 26SourcecodeScholar
2020

Hierarchical Quadtree Feature Optical Flow Tracking Based Sparse Pose-Graph Visual-Inertial SLAM

ICRA 2020poster

Accurate, robust and real-time localization under constrained-resources is a critical problem to be solved. In this paper, we present a new sparse pose-graph visual-inertial SLAM (SPVIS). Unlike the existing methods that are costly to deal with a large number of redundant features and 3D map points,…

Cited by 9SourceScholar
2020

Robust Dynamic State Estimation for Lateral Control of an Industrial Tractor Towing Multiple Passive Trailers

IROS 2020poster

In this paper, we propose a dynamic state estimation framework for lateral control of a heavy tractor-trailers system using only mass-produced low-cost sensors. This issue is challenging since the lateral velocity of the lead tractor is difficult to measure directly. The performance of existing dyna…

Cited by 0SourceScholar
2020

Robust Path Following of the Tractor-Trailers System in GPS-Denied Environments

RA-L 2020

This letter reports a general path following framework for the tractor-trailers system in Global Positioning System (GPS)-denied environments. Compared to existing methods, this approach prioritizes a robust, cost-optimized, and easy-to-implement solution. First, to achieve accurate path following,

Cited by 27SourceScholar
2019

A Hierarchical Framework for Coordinating Large-Scale Robot Networks

ICRA 2019poster

In this paper, we study the cooperative path planning and motion coordination problems of the multi-robot system with large number of robots, aiming for practical applications in robotic warehouses and automated transportation systems. Particularly, we solve the life-long planning problem and guaran…

Cited by 15SourceScholar
2019

LPD-Net: 3D Point Cloud Learning for Large-Scale Place Recognition and Environment Analysis

ICCV 2019poster

Point cloud based place recognition is still an open issue due to the difficulty in extracting local features from the raw 3D point cloud and generating the global descriptor, and it's even harder in the large-scale dynamic environments. In this paper, we develop a novel deep neural network, named L…

Cited by 345PDFScholar
2019

Learning Actions from Human Demonstration Video for Robotic Manipulation

IROS 2019poster

Learning actions from human demonstration is an emerging trend for designing intelligent robotic systems, which can be referred as video to command. The performance of such approach highly relies on the quality of video captioning. However, the general video captioning methods focus more on the unde…

Cited by 32SourceScholar
2019

Local Pose optimization with an Attention-based Neural Network

IROS 2019poster

In this paper, we propose a novel pose optimizer which can be inserted into either supervised or unsupervised end-to-end visual odometry for the purpose of local pose optimization. The pose optimizer is an analogue of the pose graph optimization used in traditional VSLAM algorithms. Local pose optim…

Cited by 2SourceScholar
2019

Long-Term Visual Inertial SLAM based on Time Series Map Prediction

IROS 2019poster

With the advance in the field of mobile robots, autonomous robots are required for long-term deployment in dynamic and complex environments. However, the performance of Visual Inertial SLAM systems in long-term operation is not satisfactory, and most long-term SLAM systems assumes periodic changes i…

Cited by 16SourceScholar
2019

Retrieval-based Localization Based on Domain-invariant Feature Learning under Changing Environments

IROS 2019poster

Visual localization is a crucial problem in mobile robotics and autonomous driving. One solution is to retrieve images with known pose from a database for the localization of query images. However, in environments with drastically varying conditions (e.g. illumination changes, seasons, occlusion, dy…

Cited by 30SourcecodeScholar
2019

SeqLPD: Sequence Matching Enhanced Loop-Closure Detection Based on Large-Scale Point Cloud Description for Self-Driving Vehicles

IROS 2019poster

Place recognition and loop-closure detection are main challenges in the localization, mapping and navigation tasks of self-driving vehicles. In this paper, we solve the loop-closure detection problem by incorporating the deep-learning based point cloud description method and the coarse-to-fine seque…

Cited by 70SourceScholar
2019

Unsupervised Learning of Monocular Depth and Ego-Motion Using Multiple Masks

ICRA 2019poster

A new unsupervised learning method of depth and ego-motion using multiple masks from monocular video is proposed in this paper. The depth estimation network and the ego-motion estimation network are trained according to the constraints of depth and ego-motion without truth values. The main contribut…

Cited by 38SourcecodeScholar
2019

Vision-Based Dynamic Control of Car-Like Mobile Robots

ICRA 2019poster

Most existing controllers for Car-Like Mobile Robots (CLMR) are designed to handle dynamic effects by decoupling speed and steering controls, also assume that full states are accessible, which are unrealistic for real-world applications. This paper presents a combined speed and steering control syst…

Cited by 9SourceScholar
2018

A Failure-Tolerant Approach to Synchronous Formation Control of Mobile Robots Under Communication Delays

ICRA 2018poster

Robot malfunction is inevitable in practical applications of the robot formation control due to uncontrolled crashing, system malfunction or communication loss. In this paper, we study the synchronous formation control problem in the presence of robot malfunctions. Our main idea is to improve the ne…

Cited by 7SourceScholar
2018

Stabilize an Unsupervised Feature Learning for LiDAR-based Place Recognition

IROS 2018poster

Place recognition is one of the major challenges for the LiDAR-based effective localization and mapping task. Traditional methods are usually relying on geometry matching to achieve place recognition, where a global geometry map need to be restored. In this paper, we accomplish the place recognition…

Cited by 21SourceScholar
2018

Vision-Based State Estimation and Trajectory Tracking Control of Car-Like Mobile Robots with Wheel Skidding and Slipping

IROS 2018poster

Most existing trajectory tracking controllers are based on non-skidding and non-slipping assumptions, also assume that full states are accessible, which is unrealistic for real-world applications due to tire-road interaction. This paper presents a novel vision-based approach to achieve high performa…

Cited by 12SourceScholar
2017

A unified leader-follower scheme for mobile robots with uncalibrated on-board camera

ICRA 2017poster

This paper studies the problem of image-based leader-follower formation control for mobile robots, where the controller is designed independently of the leader's motion. An adaptive control scheme, which is suitable for both omnidirectional and perspective cameras, is proposed. The proposed approach…

Cited by 14SourceScholar
2017

Design, modeling and experimental validation of a scissor mechanisms enabled compliant modular earthworm-like robot

IROS 2017poster

Inspired by natural earthworm locomotion behavior and segmental muscle motion mechanism, this paper presents our recently developed compliant modular earthwormlike robot with the novel segmental muscle-mimetic design unit that is capable of efficiently mimicking earthworms' segmental muscle contract…

Cited by 26SourceScholar
2016

Adaptive 3D pose computation of suturing needle using constraints from static monocular image feedback

IROS 2016poster

In this paper, we address the problem of the image-based 3D pose computation of a semi-circle suturing needle using monocular image feedback for laparoscopy. We propose a constrained two-degree-of-freedom (2-DOF) geometry-based modelling method to parametrise the needle's 6-DOF pose, including depth…

Cited by 14SourceScholar
2016

Development of a robotic system for orthodontic archwire bending

ICRA 2016

Customized archwires are demanded in the lingual orthodontic treatment for patients suffering from malocclusion. Traditionally, these archwires could only be bent by experienced orthodontists manually. This pattern requires a specialized skills training and occupies long charside time, but still can

Cited by 16SourceScholar
2016

Robust image-based computation of the 3D position of RCM instruments and its application to image-guided manipulation

ICRA 2016

In this paper, we address the 3D position control of RCM-constrained instruments with monocular cameras. To compute the instrument's position from a single 2D image, we develop an innovative gradient descent algorithm which rotates and translates a line segment (over the plane spanned by the imaged

Cited by 2SourceScholar
2015

A gradient-based self-healing algorithm for mobile robot formation

IROS 2015poster

In this paper, we investigate the self-healing problem of mobile robot formation after some robots have been damaged, and present a gradient-based algorithm which enables mobile robots to restore the topology of the formation through local interactions among neighboring robots. Firstly, in order to…

Cited by 10SourceScholar