← Search

Rong Xiong

120 accepted papers

2026

$\pi$-BA: Probabilistic Neural Bundle Adjustment With Iterative Cycle Optimization for Driving Scene Reconstruction

RA-L 2026

Urban scene reconstruction under noisy camera poses remains a critical challenge for autonomous driving. While recent neural dense Bundle Adjustment (BA) methods have shown promising results in specific settings, their performance often degrades in real-world urban scenarios due to noisy corresponde

Cited by 0SourceScholar
2026

AnyAmber: A Generalist for Versatile Anonymous Bearing and Range Based Position Tracking

RSS 2026poster

Position tracking based on bearing measurements and Ultra-wideband (UWB) ranging is widely used in robotic navigation tasks. However, due to variations in the number of robots, anchor configurations, UWB tag layouts, and the presence or absence of anonymous visual observations, existing methods typi…

Cited by 0SourceScholar
2026

Efficient Alignment of Unconditioned Action Prior for Language-Conditioned Pick and Place in Clutter (I)

ICRA 2026poster

We study the task of language-conditioned pick and place in clutter, where a robot should grasp a target object in open clutter and move it to a specified place. Some approaches learn end-to-end policies with features from vision foundation models, requiring large datasets. Others combine foundation…

Cited by 0codeScholar
2026

Geometry-Guided Modeling of Foundation Features Enables Generalizable Object Shape Deformation Learning

ICML 2026poster

Monocular 3D shape recovery is fundamental to geometric understanding, yet achieving robust generalization across arbitrary viewpoints and unseen object categories remains a significant challenge. In this paper, we present a generalizable deformation learning framework that reconstructs 3D objects b…

Cited by 0SourceScholar
2026

Predictive Local Planning with Multi-Step Reward and Q-Value Forecasting

ICRA 2026poster

Planning in dynamic environments often relies on explicit future observation prediction or value-based estimation, both of which can be brittle or hard to generalize in uncertain settings. We propose a novel model-based reinforcement learning framework that performs trajectory rollout and optimizati…

Cited by 0Scholar
2026

Shape Sensing and Tip Tracking Via Reciprocating Magnet in the Soft Continuum Robot

ICRA 2026poster

Soft continuum robots, attributable to inherently compliant trunks and shape manipulability, have been widely deployed in complex scenarios requiring safe human-robot interaction. However, their nonlinear deformations and hyperredundant degrees of freedom pose substantial challenges for full-body sh…

Cited by 0Scholar
2026

Sym-Servo: Disambiguate Symmetric Object Pose by End-To-End Optimal Visual Servo

ICRA 2026poster

Controlling symmetric objects is an indispensable but challenging task in robotic manipulation. Mainstream perception-action frameworks rely on accurate 6D pose estimation to guide the controller. However, the majority of existing 6D pose estimation methods for symmetric objects are designed to outp…

Cited by 0Scholar
2026

Toward Embodiment Equivariant Vision-Language-Action Policy

ICRA 2026poster

Vision-language-action policies learn manipulation skills across tasks, environments and embodiments through large-scale pre-training. However, their ability to generalize to novel robot configurations remains limited. Most approaches emphasize model size, dataset scale and diversity while paying le…

2026

UnIRe: Unsupervised Instance Decomposition for Dynamic Urban Scene Reconstruction

ICRA 2026poster

Reconstructing and decomposing dynamic urban scenes is crucial for autonomous driving, urban planning, and scene editing. However, existing methods fail to perform instance-aware decomposition without manual annotations, which is crucial for instance-level scene editing. We propose UnIRe, a 3D Gauss…

2025

Adaptive Neural Uncalibrated Visual Servo with Zero-shot Transfer of Extrinsics and Scenes

IROS 2025

Deploying visual servo controller to novel scenes with uncertain parameters requires additional manual effort for calibration. Traditional methods tackle this problem by online estimating the Jacobian matrix. However, they struggle in challenging scenes due to intrinsic limitations. For instance, im

Cited by 0SourceScholar
2025

Adaptive Wavelet-Positional Encoding for High-Frequency Information Learning in Implicit Neural Representation

AAAI 2025technical

Implicit Neural Representation (INR) has shown great potential in constructing the complex nature signal as a continuous implicit function. However, the representation results are incomplete since different components of the signal correspond to different frequencies and neural network inherently te…

Cited by 0SourcePDFScholar
2025

BEV-DWPVO: BEV-Based Differentiable Weighted Procrustes for Low Scale-Drift Monocular Visual Odometry on Ground

RA-L 2025

Monocular Visual Odometry (MVO) provides a cost-effective, real-time positioning solution for autonomous vehicles. However, MVO systems face the common issue of lacking inherent scale information from monocular cameras. Traditional methods have good interpretability but can only obtain relative scal

Cited by 2SourceScholar
2025

CNSv2: Probabilistic Correspondence Encoded Neural Image Servo

ICRA 2025

Visual servo based on traditional image matching methods often requires accurate keypoint correspondence for high precision control. However, keypoint detection or matching tends to fail in challenging scenarios with inconsistent illuminations or textureless objects, resulting significant performanc

Cited by 2SourceScholar
2025

CarPlanner: Consistent Auto-regressive Trajectory Planning for Large-Scale Reinforcement Learning in Autonomous Driving

CVPR 2025poster

Trajectory planning is vital for autonomous driving, ensuring safe and efficient navigation in complex environments. While recent learning-based methods, particularly reinforcement learning (RL), have shown promise in specific scenarios, RL planners struggle with training inefficiencies and managing…

2025

ColaDex: Contact-guided Optimization and VLM-assisted Selection for Task-oriented Dexterous Grasp Generation

IROS 2025

Task-oriented dexterous grasp generation aims to generate stable and functional grasps that enable a robotic hand to effectively interact with objects to accomplish specific tasks. However, generating high-dimensional hand configurations that seamlessly adapt to diverse task requirements and object

Cited by 0SourceScholar
2025

DORec: Decomposed Object Reconstruction and Segmentation Utilizing 2D Self-Supervised Features

RA-L 2025

Recovering 3D geometry and textures of individual objects is crucial for many robotics applications, such as manipulation, pose estimation, and autonomous driving. However, decomposing a target object from a complex background is challenging. Most existing approaches rely on costly manual labels to

Cited by 1SourceScholar
2025

Deformation Configuration Estimation for Soft Continuum Robot Utilizing Seq2Seq Learning

RA-L 2025

Inspired by biological tentacles, soft continuum robots exhibit the potential for navigating through narrow spaces and operating in complex environments, offering extensive application possibilities. However, owing to their inherent compliance, soft continuum robots may undergo unpredictable deforma

Cited by 1SourceScholar
2025

Disambiguate Gripper State in Grasp-Based Tasks: Pseudo-Tactile as Feedback Enables Pure Simulation Learning

IROS 2025

Grasp-based manipulation tasks are fundamental to robots interacting with their environments, yet gripper state ambiguity significantly reduces the robustness of imitation learning policies for these tasks. Data-driven solutions face the challenge of high real-world data costs, while simulation data

Cited by 2SourceScholar
2025

Grounding 3D Object Affordance with Language Instructions, Visual Observations and Interactions

CVPR 2025poster

Grounding 3D object affordance is a task that locates objects in 3D space where they can be manipulated, which links perception and action for embodied intelligence. For example, for an intelligent robot, it is necessary to accurately ground the affordance of an object and grasp it according to huma…

Cited by 0SourcePDFScholar
2025

High-Precision and High-Efficiency Trajectory Tracking for Excavators Based on Closed-Loop Dynamics

IROS 2025

The complex nonlinear dynamics of hydraulic excavators, such as time delays and control coupling, pose significant challenges to achieving high-precision trajectory tracking. Traditional control methods often fall short in such applications due to their inability to effectively handle these nonlinea

Cited by 0SourcecodeScholar
2025

Human-guided robotic-assistance handheld continuum medical robot system

IROS 2025

Nowadays, laparoscopic surgery procedures face a trade-off between expensive, complex robotic systems and manual instruments with limited functionality. Fully robotic solutions offer precision but lack portability and intuitive control, while manual tools rely solely on the surgeon’s dexterity, limi

Cited by 0SourceScholar
2025

LI-GS: Gaussian Splatting With LiDAR Incorporated for Accurate Large-Scale Reconstruction

RA-L 2025

Large-scale 3D reconstruction is critical in the field of robotics, and the potential of 3D Gaussian Splatting (3DGS) for achieving accurate object-level reconstruction has been demonstrated. However, ensuring geometric accuracy in outdoor and unbounded scenes remains a significant challenge. This s

Cited by 32SourceScholar
2025

Mr. Virgil: Learning Multi-robot Visual-range Relative Localization

IROS 2025

Ultra-wideband (UWB)-vision fusion localization has achieved extensive applications in the domain of multiagent relative localization. The challenging matching problem between robots and visual detection renders existing methods highly dependent on identity-encoded hardware or delicate tuning algori

Cited by 0SourcecodeScholar
2025

Ms. NAMI: Multimodal Semantic Navigation on Relative Metric Intention Graph

ICRA 2025

Embodied navigation in unknown environments presents the significant challenge of integrating tasks with multimodal goals into a unified framework. In this paper, we propose the Multimodal Semantic Navigation on Relative Metric Intention Graph (Ms. NAMI), a framework that integrates various navigati

Cited by 0SourceScholar
2025

Natural Humanoid Robot Locomotion with Generative Motion Prior

IROS 2025

Natural and lifelike locomotion remains a fundamental challenge for humanoid robots to interact with human society. However, previous methods either neglect motion naturalness or rely on unstable and ambiguous style rewards. In this paper, we propose a novel Generative Motion Prior (GMP) that provid

Cited by 10SourceScholar
2025

PanopticSplatting: End-to-End Panoptic Gaussian Splatting

IROS 2025

Open-vocabulary panoptic reconstruction is a challenging task for simultaneous scene reconstruction and understanding. Recently, methods have been proposed for 3D scene understanding based on Gaussian splatting. However, these methods are multi-staged, suffering from the accumulated errors and the d

Cited by 2SourceScholar
2025

RISED: Accurate and Efficient RGB-Colorized Mapping Using Image Selection and Point Cloud Densification

ICRA 2025

Recent advances in robotics have underscored the critical role of colorized point clouds in enhancing environmental perception accuracy. However, conventional multisensor fusion Simultaneous Localization and Mapping (SLAM) systems typically employ all available images indiscriminately for point clou

Cited by 1SourceScholar
2025

Reinforcement Learning for Adaptive Planner Parameter Tuning: A Perspective on Hierarchical Architecture

ICRA 2025

Automatic parameter tuning methods for planning algorithms, which integrate pipeline approaches with learning-based techniques, are regarded as promising due to their stability and capability to handle highly constrained environments. While existing parameter tuning methods have demonstrated conside

Cited by 2SourceScholar
2025

Sparse Hierarchical LiDAR Bundle Adjustment for Online Collaborative Localization and Mapping

RA-L 2025

This letter presents a sparse hierarchical LiDAR bundle adjustment method for online multi-robot collaborative simultaneous localization and mapping (C-SLAM). The motivation behind this work is that the pose graph cannot directly reflect map inconsistencies. As a result, the map divergence across mu

Cited by 1SourceScholar
2025

Temporal-Spatial Representation Fusion for Dexterous Manipulation Learning with Unpaired Visual-Action Data

IROS 2025

Supervised behavioral cloning using robot visual-action data has been widely investigated in robot manipulation. However, these methods typically require simultaneous acquisition of visual and action data, which makes them difficult to utilize unpaired visual-action datasets: e.g. videos on Internet

Cited by 0SourceScholar
2024

A Fast Motion and Foothold Planning Framework for Legged Robots on Discrete Terrain

IROS 2024poster

Legged robot proved their capability to cross complex terrain in recent research, yet the autonomy of robots on discrete terrain still needs to be enhanced since it requires a full stack framework. This paper introduces a real-time motion and foothold planning framework tailored for legged robots na…

Cited by 0SourceScholar
2024

Adapting for Calibration Disturbances: A Neural Uncalibrated Visual Servoing Policy

ICRA 2024poster

Visual servoing (VS) is a widely used technique in industries where there are hundreds of robots, but it requires accurate camera calibration including camera intrinsic and extrinsic parameters. However, it is labour-intensive to calibrate robots one-by-one in practical use. In this paper, we propos…

Cited by 1SourceScholar
2024

Advancing Virtual Reality Interaction: A Ring-Shaped Controller and Pose Tracking

ICRA 2024poster

Ensuring robust tracking of controllers’ movement is critical for human-robot interaction in virtual reality (VR) scenarios. This paper proposes a robust tracking algorithm based on a novel wearable ring-shaped controller equipped with an inertial measurement unit (IMU) and a light-emitting diode (L…

Cited by 0SourceScholar
2024

BEV-ODOM: Reducing Scale Drift in Monocular Visual Odometry with BEV Representation

IROS 2024poster

Monocular visual odometry (MVO) is vital in autonomous navigation and robotics, providing a cost-effective and flexible motion tracking solution, but the inherent scale ambiguity in monocular setups often leads to cumulative errors over time. In this paper, we present BEV-ODOM, a novel MVO framework…

Cited by 1SourceScholar
2024

Demonstration Data-Driven Parameter Adjustment for Trajectory Planning in Highly Constrained Environments

RA-L 2024

Trajectory planning in highly constrained environments is crucial for robotic navigation. Classical algorithms are widely used for their interpretability, generalization, and system robustness. However, these algorithms often require parameter retuning when adapting to new scenarios. To address this

Cited by 2SourceScholar
2024

EDA: Evolving and Distinct Anchors for Multimodal Motion Prediction

AAAI 2024technical

Motion prediction is a crucial task in autonomous driving, and one of its major challenges lands in the multimodality of future behaviors. Many successful works have utilized mixture models which require identification of positive mixture components, and correspondingly fall into two main lines: pre…

2024

Efficient Global Trajectory Planning for Multi-robot System with Affinely Deformable Formation

IROS 2024

Global trajectory planning is crucial for long-range formation navigation tasks of multi-robot systems in efficiency improvement and energy saving, whose main challenges are the joint space constraints of the whole team and the long-range deployment. To overcome the above difficulties, we reformulat

Cited by 1SourceScholar
2024

Enhancing Closed-Loop Performance in Learning-Based Vehicle Motion Planning by Integrating Rule-Based Insights

RA-L 2024

This letter introduces an innovative vehicle motion planning method that leverages the integration of rule-based insights to significantly improve closed-loop performance within a learning-based framework. We first employ rule-based methods to heuristically search and generate a diverse set of traje

Cited by 2SourceScholar
2024

Learning Hierarchical Graph-Based Policy for Goal-Reaching in Unknown Environments

RA-L 2024

Goal-reaching in unknown environments is one of the essential tasks in robot applications. Large-scale perception and long-horizon decision-making are the keys to solving this task as the operation scope expands or complexity rises. Existing navigation methods may suffer from degraded performance in

Cited by 6SourceScholar
2024

Learning the Inverse Kinematics of Magnetic Continuum Robot for Teleoperated Navigation

IROS 2024

Magnetic continuum robots are subject to external magnetic fields and deformed remotely, simplifying the robot’s transmission mechanism and providing it with significant potential for miniaturization and operational flexibility. However, modeling magnetic field distribution generated by permanent ma

Cited by 2SourceScholar
2024

Let Occ Flow: Self-Supervised 3D Occupancy Flow Prediction

CoRL 2024poster

Accurate perception of the dynamic environment is a fundamental task for autonomous driving and robot systems. This paper introduces Let Occ Flow, the first self-supervised work for joint 3D occupancy and occupancy flow prediction using only camera inputs, eliminating the need for 3D annotations. Ut…

Cited by 10SourceScholar
2024

NGEL-SLAM: Neural Implicit Representation-based Global Consistent Low-Latency SLAM System

ICRA 2024poster

Neural implicit representations have emerged as a promising solution for providing dense geometry in Simultaneous Localization and Mapping (SLAM). However, existing methods in this direction fall short in terms of global consistency and low latency. This paper presents NGEL-SLAM to tackle the above…

Cited by 29SourceScholar
2024

OTVIC: A Dataset with Online Transmission for Vehicle-to-Infrastructure Cooperative 3D Object Detection

IROS 2024poster

Vehicle-to-infrastructure cooperative 3D object detection (VIC3D) is a task that leverages both vehicle and roadside sensors to jointly perceive the surrounding environment. However, considering the high speed of vehicles, the real-time requirements, and the limitations of communication bandwidth, r…

Cited by 1SourceScholar
2024

Online Trajectory Deformation and Tracking for Self-entanglement-free Differential-Driven Robots

ICRA 2024poster

This paper introduces an optimisation-based trajectory deformation and tracking algorithm for tethered differential-driven mobile robots. The motivation of this work is to generate self-entanglement-free (SEF) commands for a tethered differential-driven robot to track a path. Whilst existing path pl…

Cited by 0SourceScholar
2024

Optimal Non-Redundant Manipulator Surface Coverage with Rank-Deficient Manipulability Constraints

RSS 2024poster

A generalised solver for the manipulator non-revisiting coverage path planning (NCPP) problem is proposed in this paper. Nonlinear manipulator kinematics and the imposition of task-specific constraints dictate that applying conventional coverage path planning (CPP) solutions based on 2D template mat…

Cited by 0SourcePDFScholar
2024

PEP: Policy-Embedded Trajectory Planning for Autonomous Driving

RA-L 2024

Autonomous driving demands proficient trajectory planning to ensure safety and comfort. This letter introduces Policy-Embedded Planner (PEP), a novel framework that enhances closed-loop performance of imitation learning (IL) based planners by embedding a neural policy for sequential ego pose generat

Cited by 8SourceScholar
2024

PanopticRecon: Leverage Open-vocabulary Instance Segmentation for Zero-shot Panoptic Reconstruction

IROS 2024

Panoptic reconstruction is a challenging task in 3D scene understanding. However, most existing methods heavily rely on pre-trained semantic segmentation models and known 3D object bounding boxes for 3D panoptic segmentation, which is not available for in-the-wild scenes. In this paper, we propose a

Cited by 8SourceScholar
2024

RGBD-based Image Goal Navigation with Pose Drift: A Topo-metric Graph based Approach

ICRA 2024poster

Image-goal navigation in unknown environments with sensor error is of considerable difficulty for autonomous robots. In this paper, we propose a drift-resisting topo-metric graph to map the environment and localize the robot using only relative poses. The error-sharing mechanism under this represent…

Cited by 1SourceScholar
2024

Scale Disparity of Instances in Interactive Point Cloud Segmentation

IROS 2024poster

Interactive point cloud segmentation has become a pivotal task for understanding 3D scenes, enabling users to guide segmentation models with simple interactions such as clicks, therefore significantly reducing the effort required to tailor models to diverse scenarios and new categories. However, in…

Cited by 2SourceScholar
2024

Semantics-aware Motion Retargeting with Vision-Language Models

CVPR 2024poster

Capturing and preserving motion semantics is essential to motion retargeting between animation characters. However most of the previous works neglect the semantic information or rely on human-designed joint-level representations. Here we present a novel Semantics-aware Motion reTargeting (SMT) metho…

Cited by 5SourcePDFScholar
2024

Soft Hybrid Actuated Hierarchical Bronchoscope Robot for Deep Lung Examination

RA-L 2024

Lungdiseases are becoming one of the world's most serious health issues. Soft bronchoscope robots can achieve safe and controllable lung navigation, which will be crucial for the future early examination of lung diseases. However, due to the single driving method, the large size, and insufficient fl

Cited by 10SourceScholar
2024

Tree-based Representation of Locally Shortest Paths for 2D k-Shortest Non-homotopic Path Planning

ICRA 2024poster

A novel algorithm to solve the 2D k-shortest non-homotopic path planning (k-SNPP) task is proposed in this paper. The task is of practical significance as a sub-module for higherlevel planning and scheduling tasks, and is gaining increasing attention and focus in recent years. There have existed alg…

Cited by 2SourceScholar
2024

VIVO: A Visual-Inertial-Velocity Odometry with Online Calibration in Challenging Condition

IROS 2024poster

State estimation is a central component of autonomous navigation. To date, many methods presented have a disruptive potential for application, such as visual-inertial odometry (VIO), wheel and leg odometry (for short, body odometry). However, most of them are prone to fail in some challenging condit…

Cited by 0SourceScholar
2024

Vertebrae-based Global X-ray to CT Registration for Thoracic Surgeries

IROS 2024poster

X-ray to CT registration is an essential technique to provide on-site guidance for clinicians and medical robots by aligning preoperative information with intraoperative images. Current methods focus on local registration with small capture ranges and necessitate a manual initial alignment before pr…

Cited by 0SourcecodeScholar
2024

ν-DBA: Neural Implicit Dense Bundle Adjustment Enables Image-Only Driving Scene Reconstruction

IROS 2024poster

The joint optimization of the sensor trajectory and 3D map is a crucial characteristic of bundle adjustment (BA), essential for autonomous driving. This paper presents ν-DBA, a novel framework implementing geometric dense bundle adjustment (DBA) using 3D neural implicit surfaces for map parametrizat…

Cited by 0SourceScholar
2023

A Hyper-Network Based End-to-End Visual Servoing With Arbitrary Desired Poses

RA-L 2023

Recently, several works achieve end-to-end visual servoing (VS) for robotic manipulation by replacing traditional controller with differentiable neural networks, but lose the ability to servo arbitrary desired poses. This letter proposes a differentiable architecture for arbitrary pose servoing: a h

Cited by 8SourceScholar
2023

A Joint Modeling of Vision-Language-Action for Target-oriented Grasping in Clutter

ICRA 2023poster

We focus on the task of language-conditioned grasping in clutter, in which a robot is supposed to grasp the target object based on a language instruction. Previous works separately conduct visual grounding to localize the target object, and generate a grasp for that object. However, these works requ…

Cited by 49SourcecodeScholar
2023

A Two-Stage Based Social Preference Recognition in Multi-Agent Autonomous Driving System

IROS 2023poster

Multi-Agent Reinforcement Learning (MARL) has become a promising solution for constructing a multi-agent autonomous driving system (MADS) in complex and dense scenarios. But most methods consider agents acting selfishly, which leads to conflict behaviors. Some existing works incorporate the concept…

Cited by 2SourceScholar
2023

An Efficient Multi-solution Solver for the Inverse Kinematics of 3-Section Constant-Curvature Robots

RSS 2023poster

Piecewise constant curvature is a popular kinematics framework for continuum robots. Computing the model parameters from the desired end pose, known as the inverse kinematics problem, is fundamental in manipulation, tracking and planning tasks. In this paper, we propose an efficient multi-solution s…

Cited by 7SourcePDFScholar
2023

C2: Co-design of Robots via Concurrent-Network Coupling Online and Offline Reinforcement Learning

IROS 2023poster

With the increasing computing power, using data-driven approaches to co-design a robot's morphology and controller has become a promising way. However, most existing data-driven methods require training the controller for each morphology to calculate fitness, which is time-consuming. In contrast, th…

Cited by 3SourceScholar
2023

DAMS-LIO: A Degeneration-Aware and Modular Sensor-Fusion LiDAR-inertial Odometry

ICRA 2023poster

With robots being deployed in increasingly complex environments like underground mines and planetary surfaces, the multi-sensor fusion method has gained more and more attention which is a promising solution to state estimation in the such scene. The fusion scheme is a central component of these meth…

Cited by 13SourceScholar
2023

DeepRING: Learning Roto-translation Invariant Representation for LiDAR based Place Recognition

ICRA 2023poster

LiDAR based place recognition is popular for loop closure detection and re-localization. In recent years, deep learning brings improvements to place recognition by learnable feature extraction. However, these methods degenerate when the robot re-visits previous places with a large perspective differ…

Cited by 12SourceScholar
2023

Distributed Initialization for Visual-Inertial-Ranging Odometry with Position-Unknown UWB Network

ICRA 2023poster

In recent years, the visual-inertial-ranging (VIR) state estimator with a position-unknown UWB network has become popular. However, most existing VIR methods leverage centralized algorithms to initialize the UWB anchors, which are challenging to be applied to massive UWB networks. In this paper, we…

Cited by 4SourceScholar
2023

Exploiting Point-Wise Attention in 6D Object Pose Estimation Based on Bidirectional Prediction

RA-L 2023

Traditional geometric registration based estimation methods only exploit the CAD model implicitly, which leads to their dependence on observation quality and deficiency to occlusion.To address the problem,the letter proposes a bidirectional correspondence prediction network with a point-wise attenti

Cited by 1SourceScholar
2023

Failure-aware Policy Learning for Self-assessable Robotics Tasks

ICRA 2023poster

Self-assessment rules play an essential role in safe and effective real-world robotic applications, which verify the feasibility of the selected action before actual execution. But how to utilize the self-assessment results to re-choose actions remains a challenge. Previous methods eliminate the sel…

Cited by 2SourceScholar
2023

Map-Based Visual-Inertial Localization: Consistency and Complexity

RA-L 2023

Drift-free localization is essential for autonomous vehicles. In this letter, we address the problem by proposing a filter-based framework, which integrates the visual-inertial odometry and the measurements from the pre-built map. In this framework, the transformation between the odometry frame and

Cited by 6SourceScholar
2023

NF-Atlas: Multi-Volume Neural Feature Fields for Large Scale LiDAR Mapping

RA-L 2023

LiDAR Mapping has been a long-standing problem in robotics. Recent progress in neural implicit representation has brought new opportunities to robotic mapping. In this letter, we propose the multi-volume neural feature fields, called NF-Atlas, which bridge the neural feature volumes with pose graph

Cited by 21SourceScholar
2023

Open-Set Object Detection Using Classification-Free Object Proposal and Instance-Level Contrastive Learning

RA-L 2023

Detecting both known and unknown objects is a fundamental skill for robot manipulation in unstructured environments. Open-set object detection (OSOD) is a promising direction to handle the problem consisting of two subtasks: objects and background separation, and open-set object classification. In t

Cited by 21SourceScholar
2023

Robust Real-Time Motion Retargeting via Neural Latent Prediction

IROS 2023poster

Human-robot motion retargeting is a crucial approach for fast learning motion skills. Achieving real-time retargeting demands high levels of synchronization and accuracy. Even though existing retargeting methods have swift calculation, they still cause time-delay effect on the synchronous retargetin…

Cited by 1SourceScholar
2023

Self-Entanglement-Free Tethered Path Planning for Non-Particle Differential-Driven Robot

ICRA 2023poster

A novel mechanism to derive self-entanglement-free path for tethered differential-driven robots is proposed in this work. The problem is tailored to the applications of tethered robots without an omni-directional tether re-tractor which is often encountered when an omni-directional tether retracting…

Cited by 4SourceScholar
2023

UrbanGIRAFFE: Representing Urban Scenes as Compositional Generative Neural Feature Fields

ICCV 2023poster

Generating photorealistic images with controllable camera pose and scene contents is essential for many applications including AR/VR and simulation. Despite the fact that rapid progress has been made in 3D-aware generative models, most existing methods focus on object-centric images and are not appl…

Cited by 17PDFScholar
2022

A Visual Navigation Perspective for Category-Level Object Pose Estimation

ECCV 2022poster

"This paper studies category-level object pose estimation based on a single monocular image. Recent advances in pose-aware generative models have paved the way for addressing this challenging task using analysis-by-synthesis. The idea is to sequentially update a set of latent variables,e.g., pose, s…

2022

DXQ-Net: Differentiable LiDAR-Camera Extrinsic Calibration Using Quality-aware Flow

IROS 2022poster

Accurate LiDAR-camera extrinsic calibration is a precondition for many multi-sensor systems in mobile robots. Most calibration methods rely on laborious manual operations and calibration targets. While working online, the calibration methods should be able to extract information from the environment…

Cited by 36SourceScholar
2022

Domain Generalization for Vision-based Driving Trajectory Generation

ICRA 2022poster

One of the challenges in vision-based driving trajectory generation is dealing with out-of-distribution scenarios. In this paper, we propose a domain generalization method for vision-based driving trajectory generation for autonomous vehicles in urban environments, which can be seen as a solution to…

Cited by 5SourceScholar
2022

Efficient Object Manipulation to an Arbitrary Goal Pose: Learning-Based Anytime Prioritized Planning

ICRA 2022poster

We focus on the task of object manipulation to an arbitrary goal pose, in which a robot is supposed to pick an assigned object to place at the goal position with a specific orientation. However, limited by the execution space of the manipulator with gripper, one-step picking, moving and releasing mi…

Cited by 12SourceScholar
2022

FEJ-VIRO: A Consistent First-Estimate Jacobian Visual-Inertial-Ranging Odometry

IROS 2022poster

In recent years, Visual-Inertial Odometry (VIO) has achieved many significant progresses. However, VIO meth-ods suffer from localization drift over long trajectories. In this paper, we propose a First-Estimates Jacobian Visual-Inertial-Ranging Odometry (FEJ-VIRO) to reduce the localization drifts of…

Cited by 16SourceScholar
2022

Fusing Priori and Posteriori Metrics for Automatic Dataset Annotation of Planar Grasping

CoRL 2022poster

Grasp detection based on deep learning has been a research hot spot in recent years. The performance of grasping detection models relies on high-quality, large-scale grasp datasets. Taking comprehensive consideration of quality, extendability, and annotation cost, metric-based simulation methodolog…

Cited by 0SourceScholar
2022

Kinematic Motion Retargeting via Neural Latent Optimization for Learning Sign Language

RA-L 2022

Motion retargeting from a human demonstration to a robot is an effective way to reduce the professional requirements and workload of robot programming, but faces the challenges resulting from the differences between humans and robots. Traditional optimization-based methods are time-consuming and rel

Cited by 31SourceScholar
2022

Learning Interpretable BEV Based VIO without Deep Neural Networks

CoRL 2022poster

Monocular visual-inertial odometry (VIO) is a critical problem in robotics and autonomous driving. Traditional methods solve this problem based on filtering or optimization. While being fully interpretable, they rely on manual interference and empirical parameter tuning. On the other hand, learning-…

Cited by 3SourceScholar
2022

Learning Observation-Based Certifiable Safe Policy for Decentralized Multi-Robot Navigation

ICRA 2022poster

Safety is of great importance in multi-robot navigation problems. In this paper, we propose a control barrier function (CBF) based optimizer that ensures robot safety with both high probability and flexibility, using only sensor measurement. The optimizer takes action commands from the policy networ…

Cited by 12SourcecodeScholar
2022

Learning to Fill the Seam by Vision: Sub-millimeter Peg-in-hole on Unseen Shapes in Real World

ICRA 2022poster

In the peg insertion task, human pays attention to the seam between the peg and the hole and tries to fill it continuously with visual feedback. By imitating the human's behavior, we design architectures with position and orientation estimators based on the seam representation for pose alignment, wh…

Cited by 18SourcecodeScholar
2022

Leveraging Local Planar Motion Property for Robust Visual Matching and Localization

RA-L 2022

One primary difficulty preventing the visual localization for service robots is the robustness against changes, including environmental changes and perspective changes. In recent years, learning-based feature matching methods have been widely studied and effectively verified in practical application

Cited by 4SourceScholar
2022

One RING to Rule Them All: Radon Sinogram for Place Recognition, Orientation and Translation Estimation

IROS 2022poster

LiDAR-based global localization is a fundamental problem for mobile robots. It consists of two stages, place recognition and pose estimation, which yields the current orientation and translation, using only the current scan as query and a database of map scans. Inspired by the definition of a recogn…

Cited by 26SourceScholar
2022

SO-PFH: Semantic Object-based Point Feature Histogram for Global Localization in Parking Lot

IROS 2022poster

Global localization is essential for autonomous mobile systems, especially indoor applications where the GPS signal is denied. Although the appearance-based methods have been successfully applied in various localization tasks, they face various challenges such as light variation, viewpoint changing,…

Cited by 3SourceScholar
2022

Towards Two-view 6D Object Pose Estimation: A Comparative Study on Fusion Strategy

IROS 2022poster

Current RGB-based 6D object pose estimation methods have achieved noticeable performance on datasets and real world applications. However, predicting 6D pose from single 2D image features is susceptible to disturbance from changing of environment and textureless or resemblant object surfaces. Hence,…

Cited by 3SourceScholar
2022

Translation Invariant Global Estimation of Heading Angle Using Sinogram of LiDAR Point Cloud

ICRA 2022poster

Global point cloud registration is an essential module for localization, of which the main difficulty exists in estimating the rotation globally without initial value. With the aid of gravity alignment, the degree of freedom in point cloud registration could be reduced to 4DoF, in which only the hea…

Cited by 10SourceScholar
2021

Assembly Sequence Generation for New Objects via Experience Learned from Similar Object

IROS 2021poster

Assembly orders of components have direct influence on feasibility and efficiency of assembly process in manufacturing and are usually defined by experienced operators. To automate the assembly sequence generation process, we present a method using the idea of case-based reasoning, which can take ad…

Cited by 2SourceScholar
2021

CORAL: Colored structural representation for bi-modal place recognition

IROS 2021poster

Place recognition is indispensable for a drift-free localization system. Due to the variations of the environment, place recognition using single-modality has limitations. In this paper, we propose a bi-modal place recognition method, which can extract a compound global descriptor from the two modal…

Cited by 36SourceScholar
2021

Deep Samplable Observation Model for Global Localization and Kidnapping

RA-L 2021

Global localization and kidnapping are two challenging problems in robot localization. The popular method, Monte Carlo Localization (MCL) addresses the problem by iteratively updating a set of particles with a “sampling-weighting” loop. Sampling is decisive to the performance of MCL [1]. However, tr

Cited by 19SourcecodeScholar
2021

Dynamic Movement Primitive based Motion Retargeting for Dual-Arm Sign Language Motions

ICRA 2021poster

We aim to develop an efficient programming method for equipping service robots with the skill of performing sign language motions. This paper addresses the problem of transferring complex dual-arm sign language motions characterized by the coordination among arms and hands from human to robot, which…

Cited by 26SourceScholar
2021

Efficient Learning of Goal-Oriented Push-Grasping Synergy in Clutter

RA-L 2021

We focus on the task of goal-oriented grasping, in which a robot is supposed to grasp a pre-assigned goal object in clutter and needs some pre-grasp actions such as pushes to enable stable grasps. However, in this task, the robot gets positive rewards from environment only when successfully grasping

Cited by 91SourcecodeScholar
2021

Imitation Learning of Hierarchical Driving Model: From Continuous Intention to Continuous Trajectory

RA-L 2021

One of the challenges to reduce the gap between the machine and the human level driving is how to endow the system with the learning capacity to deal with the coupled complexity of environments, intentions, and dynamics. In this letter, we propose a hierarchical driving model with explicit models of

Cited by 19SourcecodeScholar
2021

Learn to Differ: Sim2Real Small Defection Segmentation Network

IROS 2021poster

Recent studies on deep-learning-based small defection segmentation approaches are trained in specific settings and tend to be limited by fixed context. Throughout the training, the network inevitably learns the representation of the background of the training data before figuring out the defection.…

Cited by 0SourcecodeScholar
2021

Learning World Transition Model for Socially Aware Robot Navigation

ICRA 2021poster

Moving in dynamic pedestrian environments is one of the important requirements for autonomous mobile robots. We present a model-based reinforcement learning approach for robots to navigate through crowded environments. The navigation policy is trained with both real interaction data from multi-agent…

Cited by 31SourcecodeScholar
2021

Neural Motion Prediction for In-flight Uneven Object Catching

IROS 2021poster

In-flight objects capture is extremely challenging. The robot is required to complete trajectory prediction, interception position calculation and motion planning within tens of milliseconds. As in-flight uneven objects are affected by various kinds of forces, which leads to the time-varying acceler…

Cited by 13SourceScholar
2021

Optimal Object Placement for Minimum Discontinuity Non-revisiting Coverage Task

ICRA 2021poster

This work considers the optimal non-revisiting coverage tasks with a single non-redundant manipulator for the case when the object can be positioned at a predefined set of locations within the workcell. The scenario is often encountered in typical industrial settings, for instance when the object pr…

Cited by 2SourceScholar
2021

PREGAN: Pose Randomization and Estimation for Weakly Paired Image Style Translation

RA-L 2021

Utilizing the trained model under different conditions without data annotation is attractive for robot applications. Towards this goal, one class of methods is to translate the image style from another environment to the one on which models are trained. In this letter, we propose a weakly-paired set

Cited by 1SourcecodeScholar
2021

REDE: End-to-End Object 6D Pose Robust Estimation Using Differentiable Outliers Elimination

RA-L 2021

Object 6D pose estimation is a fundamental task in many applications. Conventional methods solve the task by detecting and matching the keypoints, then estimating the pose. Recent efforts bringing deep learning into the problem mainly overcome the vulnerability of conventional methods to environment

Cited by 41SourcecodeScholar
2021

Robust localization for planar moving robot in changing environment: A perspective on density of correspondence and depth

ICRA 2021poster

Visual localization for planar moving robot is important to various indoor service robotic applications. To handle the textureless areas and frequent human activities in indoor environments, a novel robust visual localization algorithm which leverages dense correspondence and sparse depth for planar…

Cited by 7SourcecodeScholar
2020

Adversarial Feature Disentanglement for Place Recognition Across Changing Appearance

ICRA 2020poster

When robots move autonomously for long-term, varied appearance such as the transition from day to night and seasonal variation brings challenges to visual place recognition. Defining an appearance condition (e.g. a season, a kind of weather) as a domain, we consider that the desired representation f…

Cited by 27SourceScholar
2020

Deep Phase Correlation for End-to-End Heterogeneous Sensor Measurements Matching

CoRL 2020

The crucial step for localization is to match the current observation to the map. When the two sensor modalities are significantly different, matching becomes challenging. In this paper, we present an end-to-end deep phase correlation network (DPCN) to match heterogeneous sensor measurements. In DPC

2020

Efficient two step optimization for large embedded deformation graph based SLAM

ICRA 2020poster

Embedded deformation graph is a widely used technique in deformable geometry and graphical problems. Although the technique has been transmitted to stereo (or RGB-D) camera based SLAM applications, it remains challenging to compromise the computational cost as the model grows. In practice, the proce…

Cited by 3SourceScholar
2020

Globally optimal consensus maximization for robust visual inertial localization in point and line map

IROS 2020poster

Map based visual inertial localization is a crucial step to reduce the drift in state estimation of mobile robots. The underlying problem for localization is to estimate the pose from a set of 3D-2D feature correspondences, of which the main challenge is the presence of outliers, especially in chang…

Cited by 6SourceScholar
2020

Learning hierarchical behavior and motion planning for autonomous driving

IROS 2020poster

Learning-based driving solution, a new branch for autonomous driving, is expected to simplify the modeling of driving by learning the underlying mechanisms from data. To improve the tactical decision-making for learning-based driving solution, we introduce hierarchical behavior and motion planning (…

Cited by 48SourceScholar
2020

Learning-based Optimization Algorithms Combining Force Control Strategies for Peg-in-Hole Assembly

IROS 2020poster

In this paper, an approach for automatic peg-in-hole assembly is proposed. The task is divided into two main steps: searching phase and inserting phase. First, a multilayer perceptron network is designed to address the hole search problem and a hybrid force position controller is introduced to ensur…

Cited by 34SourceScholar
2020

Non-revisiting Coverage Task with Minimal Discontinuities for Non-redundant Manipulators

RSS 2020poster

A theoretically complete solution to the optimal Non-revisiting Coverage Path Planning (NCPP) problem of any arbitrarily-shaped object with a non-redundant manipulator is proposed in this work. Given topological graphs of surface cells corresponding to feasible and continuous manipulator configurati…

Cited by 8SourcePDFScholar
2019

2-Entity RANSAC for robust visual localization in changing environment

IROS 2019poster

Visual localization has attracted considerable attention due to its low-cost and stable sensor, which is desired in many applications, such as autonomous driving, inspection robots and unmanned aerial vehicles. However, current visual localization methods still struggle with environmental changes ac…

Cited by 11SourceScholar
2019

Communication constrained cloud-based long-term visual localization in real time

IROS 2019poster

Visual localization is one of the primary capabilities for mobile robots. Long-term visual localization in real time is particularly challenging, in which the robot is required to efficiently localize itself using visual data where appearance may change significantly over time. In this paper, we pro…

Cited by 8SourceScholar
2018

Laser Map Aided Visual Inertial Localization in Changing Environment

IROS 2018poster

Long-term visual localization in outdoor environment is a challenging problem, especially faced with the cross-seasonal, bi-directional tasks and changing environment. In this paper we propose a novel visual inertial localization framework that localizes against the LiDAR-built map. Based on the geo…

Cited by 37SourceScholar
2018

Predicting Objective Function Change in Pose-Graph Optimization

IROS 2018poster

Robust online incremental SLAM applications require metrics to evaluate the impact of current measurements. Despite its prevalence in graph pruning, information-theoretic metrics solely are insufficient to detect outliers. The optimal value of the objective function is a better choice to detect outl…

Cited by 5SourceScholar
2017

Planar scan matching using incident angle

IROS 2017poster

The main contribution of this paper is a planar scan matching algorithm that makes use of the incident angle of a scan point as a feature to enhance the robustness to large relative transformations, particularly in orientation. A new definition of the incident angle is introduced and its consistency…

Cited by 3SourceScholar
2015

Active control of under-actuated foot tilting for humanoid push recovery

IROS 2015poster

We propose a novel control framework to demonstrate a unique foot tilting maneuver based on ankle torque control for humanoid balance recovery. The framework consists of the variable impedance regulation at the center of mass of the robot based on the ankle torque control, the virtual stoppers to pr…

Cited by 11SourceScholar
2015

Probabilistic graph based spatial assembly relation inference for programming of assembly task by demonstration

IROS 2015poster

In robot programming by demonstration (PBD) for assembly tasks, one of the important topics is to inference the poses and spatial relations of parts during the demonstration. In this paper, we propose a world model called assembly graph (AG) to achieve this task. The model is able to represent the p…

Cited by 9SourceScholar