← Search

Jianmin Ji

50 accepted papers

2026

A Soft-Rigid Hybrid Robot-Assisted Feeding System with a Tendon-Driven Continuum Robot

ICRA 2026poster

Active delivery of food to a human mouth in a controlled and safe manner remains a key challenge for robot‑assisted feeding systems (RAFSs). Existing RAFS designs struggle to simultaneously achieve efficiency and safety: rigid manipulators offer fast and accurate motion but risk hazardous contact, w…

Cited by 0SourceScholar
2026

Bridging Language and Physics: Automated Design of Continuum Robots with Large Language Models

RSS 2026poster

Large language models (LLMs) have recently emerged as a promising tool for automating robot design from high-level specifications, yet they remain ineffective for robots operating under complex physical interactions. This limitation stems from the gap between language-based reasoning and the physica…

Cited by 0SourceScholar
2026

EcoVLA: Environment-Aware Adaptive Pruning with Interleaved Inference Orchestration for Vision-Language-Action Models

ICML 2026spotlight

While Vision-Language-Action (VLA) models hold promise in embodied intelligence, their large parameter counts lead to substantial inference latency that hinders real-time manipulation, motivating parameter sparsification. However, as the environment evolves during VLA execution, the optimal sparsity…

Cited by 0SourceScholar
2026

Learning Surgical Robotic Manipulation with 3D Spatial Priors

CVPR 2026

Achieving 3D spatial awareness is crucial for surgical robotic manipulation, where precise and delicate operations are required. Existing methods either explicitly reconstruct the surgical scene prior to manipulation, or enhance multi-view features by adding wrist-mounted cameras to supplement the d

Cited by 0SourceScholar
2026

SoPE: Spherical Coordinate-Based Positional Embedding for Enhancing Spatial Perception of 3D LVLMs

CVPR 2026

3D Large Vision-Language Models (3D LVLMs) built upon Large Language Models (LLMs) have achieved remarkable progress across various multimodal tasks. However, their inherited position-dependent modeling mechanism, Rotary Position Embedding (RoPE), remains suboptimal for 3D multimodal understanding.

Cited by 0SourceScholar
2026

Vulnerability-Aware Robust Multimodal Adversarial Training

AAAI 2026technical

Multimodal learning has shown significant superiority on various tasks by integrating multiple modalities. However, the interdependencies among modalities increase the susceptibility of multimodal models to adversarial attacks. Existing methods mainly focus on attacks on specific modalities or indis

Cited by 0SourcePDFScholar
2025

AAOPL: Automated Articulated Object Parameter Learning for Open-World Robotics

IROS 2025

Articulated objects are ubiquitous in daily environments, and effective manipulation of these objects is essential for advancing open-world robotics. Existing approaches, which rely heavily on large-scale data collection or simulation, often face limitations in real-world applications, including iss

Cited by 0SourceScholar
2025

CAFE-AD: Cross-Scenario Adaptive Feature Enhancement for Trajectory Planning in Autonomous Driving

ICRA 2025

Imitation learning based planning tasks on the nuPlan dataset have gained great interest due to their potential to generate human-like driving behaviors. However, open-loop training on the nuPlan dataset tends to cause causal confusion during closed-loop testing, and the dataset also presents a long

Cited by 2SourcecodeScholar
2025

CELLmap: Enhancing LiDAR SLAM Through Elastic and Lightweight Spherical Map Representation

ICRA 2025

SLAM is a fundamental capability of unmanned systems, with LiDAR-based SLAM gaining widespread adoption due to its high precision. Current SLAM systems can achieve centimeter-level accuracy within a short period. However, there are still several challenges when dealing with largescale mapping tasks

Cited by 2SourceScholar
2025

GARD: A Geometry-Informed and Uncertainty-Aware Baseline Method for Zero-Shot Roadside Monocular Object Detection

RA-L 2025

Roadside camera-based perception methods are in high demand for developing efficient vehicle-infrastructure collaborative perception systems. By focusing on object-level depth prediction, we explore the potential benefits of integrating environmental priors into such systems and propose a geometry-b

Cited by 1SourceScholar
2025

GraspCoT: Integrating Physical Property Reasoning for 6-DoF Grasping under Flexible Language Instructions

ICCV 2025poster

Flexible instruction-guided 6-DoF grasping is a significant yet challenging task for real-world robotic systems. Existing methods utilize the contextual understanding capabilities of the large language models (LLMs) to establish mappings between expressions and targets, allowing robots to comprehend…

2025

Hierarchical Framework for Constrained Dual-Arm Cooperative Manipulation with Whole-Body Collision Avoidance

IROS 2025

Dual-arm robotic systems hold great potential for complex bimanual tasks that require intricate and coordinated manipulation, such as holding and transporting a tray with a cup of coffee while navigating through cluttered environments. However, these tasks pose significant challenges due to the inhe

Cited by 0SourceScholar
2025

Improving Efficiency of Answer Set Planning with Rough Solutions from Large Language Models for Robotic Task Planning

IJCAI 2025

Answer Set Programming (ASP) planning can be used to refine the rough solutions generated by Large Language Models (LLMs) to handle specific restrictions of actions, i.e., reconstruct the rough solutions to be executable, for robotic task planning. However, it is still challenging to efficiently sol

2025

MT-PCR: Leveraging Modality Transformation for Large-Scale Point Cloud Registration with Limited Overlap

ICRA 2025

Large-scale scene point cloud registration with limited overlap is a challenging task due to computational load and constrained data acquisition. To tackle these issues, we propose a point cloud registration method, MT-PCR, based on Modality Transformation. MT-PCR leverages a Bird's Eye View (BEV) c

Cited by 0SourceScholar
2025

NaviDiffuser: Tackling Multi-Objective Robot Navigation by Weight Range Guided Diffusion Model

IROS 2025

The data-driven paradigm has shown great potential in solving many decision-making tasks. In the robot navigation realm, it also sparked a new trend. People believe powerful data-driven methods can learn efficient and general navigation policies from a vast offline dataset. However, robot navigation

Cited by 0SourceScholar
2025

OG-Gaussian: Occupancy Based Street Gaussians for Autonomous Driving

ICRA 2025

Accurate and realistic 3D scene reconstruction enables the lifelike creation of autonomous driving simulation environments. With advancements in 3D Gaussian Splatting (3DGS), previous studies have applied it to reconstruct complex dynamic driving scenes. These methods typically require expensive LiD

Cited by 5SourceScholar
2025

Perception Helps Planning: Facilitating Multi-Stage Lane-Level Integration via Double-Edge Structures

RA-L 2025

When planning for autonomous driving, it is crucial to consider essential traffic elements such as lanes, intersections, traffic regulations, and dynamic agents. However, they are often overlooked by the traditional end-to-end planning methods, likely leading to inefficiencies and non-compliance wit

Cited by 1SourceScholar
2025

STDArm: Transfer Visuomotor Policy From Static Data Training to Dynamic Robot Manipulation

RSS 2025poster

Learning visuomotor policy from human demonstrations serves as an effective method for robots to acquire complex tasks. However, data collection on mobile platforms such as drones is extremely challenging, resulting in most research being conducted with robots in stationary conditions for data colle…

Cited by 0PDFScholar
2025

SpatialSplat: Efficient Semantic 3D from Sparse Unposed Images

ICCV 2025poster

A major breakthrough in 3D reconstruction is the feedforward paradigm to generate pixel-wise 3D points or Gaussian primitives from sparse, unposed images. To further incorporate semantics while avoiding the significant memory and storage costs of high-dimensional semantic features, existing methods…

Cited by 0SourcePDFScholar
2025

Traffic Scenario Logic: A Spatial-Temporal Logic for Modeling and Reasoning of Urban Traffic Scenarios

AAAI 2025technical

Formal representations of traffic scenarios can be used to generate test cases for the safety verification of autonomous driving. However, most existing methods are limited to highway or highly simplified intersection scenarios due to the intricacy and diversity of traffic scenarios. In response, we…

2024

BEVoxSeg: BEV-Voxel Representation for Fast and Accurate Camera-Based 3D Segmentation

ICASSP 2024accepted

Recent research has demonstrated the advantages of Bird’s-eye-view (BEV) representation in the field of 3D perception. However, due to the lack of height information, BEV representation alone is insufficient to accurately reconstruct the complete surrounding 3D scene. On the other hand, voxel repres…

Cited by 0SourceScholar
2024

CRPlace: Camera-Radar Fusion with BEV Representation for Place Recognition

IROS 2024poster

The integration of complementary characteristics from camera and radar data has emerged as an effective approach in 3D object detection. However, such fusion-based methods remain unexplored for place recognition, an equally important task for autonomous systems. Given that place recognition relies o…

Cited by 4SourceScholar
2024

CalibFormer: A Transformer-based Automatic LiDAR-Camera Calibration Network

ICRA 2024poster

The fusion of LiDARs and cameras has been increasingly adopted in autonomous driving for perception tasks. The performance of such fusion-based algorithms largely depends on the accuracy of sensor calibration, which is challenging due to the difficulty of identifying common features across different…

Cited by 14SourceScholar
2024

EdgeCalib: Multi-Frame Weighted Edge Features for Automatic Targetless LiDAR-Camera Calibration

RA-L 2024

In multimodal perception systems, achieving precise extrinsic calibration between LiDAR and camera is of critical importance. However, the pre-calibrated extrinsic parameters may gradually drift during operation, leading to a decrease in the accuracy of the perception system. It is challenging to ad

Cited by 20SourceScholar
2024

FARFusion: A Practical Roadside Radar-Camera Fusion System for Far-Range Perception

RA-L 2024

Far-range perception through roadside sensors is crucial to the effectiveness of intelligent transportation systems. The main challenge of far-range perception is due to the difficulty of performing accurate object detection and tracking under far distances <italic xmlns:mml="http://www.w3.org/1998/

Cited by 24SourceScholar
2024

LDP: A Local Diffusion Planner for Efficient Robot Navigation and Collision Avoidance

IROS 2024poster

The conditional diffusion model has been demonstrated as an efficient tool for learning robot policies, owing to its advancement to accurately model the conditional distribution of policies. The intricate nature of real-world scenarios, characterized by dynamic obstacles and maze-like structures, un…

Cited by 13SourceScholar
2024

MM-Gaussian: 3D Gaussian-based Multi-modal Fusion for Localization and Reconstruction in Unbounded Scenes

IROS 2024poster

Localization and mapping are critical tasks for various applications such as autonomous vehicles and robotics. The challenges posed by outdoor environments present particular complexities due to their unbounded characteristics. In this work, we present MM-Gaussian, a LiDAR-camera multimodal fusion s…

Cited by 12SourceScholar
2024

NaviFormer: A Data-Driven Robot Navigation Approach via Sequence Modeling and Path Planning with Safety Verification

ICRA 2024poster

Reinforcement learning has shown great potential in improving the performance of robot navigation. In response to the increasing deployments of mobile robots within various scenarios, a data-driven paradigm of navigation approach with safety verification is preferred where one can train RL algorithm…

Cited by 1SourceScholar
2024

OCC-VO: Dense Mapping via 3D Occupancy-Based Visual Odometry for Autonomous Driving

ICRA 2024poster

Visual Odometry (VO) plays a pivotal role in autonomous systems, with a principal challenge being the lack of depth information in camera images. This paper introduces OCC-VO, a novel framework that capitalizes on recent advances in deep learning to transform 2D camera images into 3D semantic occupa…

Cited by 8SourcecodeScholar
2024

PathRL: An End-to-End Path Generation Method for Collision Avoidance via Deep Reinforcement Learning

ICRA 2024poster

Robot navigation using deep reinforcement learning (DRL) has shown great potential in improving the performance of mobile robots. Nevertheless, most existing DRL-based navigation methods primarily focus on training a policy that directly commands the robot with low-level controls, like linear and an…

Cited by 8SourceScholar
2024

Profiling Power Consumption in Low-Speed Autonomous Guided Vehicles

RA-L 2024

The increasing demand for automation has led to a rise in the use of low-speed Autonomous guided vehicles (AGVs). However, AGVs rely on batteries for their power source, which limits their operational time and affects their overall performance. To optimize their energy usage and enhance their batter

Cited by 7SourceScholar
2024

SDAC: A Multimodal Synthetic Dataset for Anomaly and Corner Case Detection in Autonomous Driving

AAAI 2024technical

Nowadays, closed-set perception methods for autonomous driving perform well on datasets containing normal scenes. However, they still struggle to handle anomalies in the real world, such as unknown objects that have never been seen while training. The lack of public datasets to evaluate the model pe…

Cited by 3SourcePDFScholar
2023

Bi-LRFusion: Bi-Directional LiDAR-Radar Fusion for 3D Dynamic Object Detection

CVPR 2023poster

LiDAR and Radar are two complementary sensing approaches in that LiDAR specializes in capturing an object's 3D shape while Radar provides longer detection ranges as well as velocity hints. Though seemingly natural, how to efficiently combine them for improved feature representation is still unclear.…

2023

CluB: Cluster Meets BEV for LiDAR-Based 3D Object Detection

NeurIPS 2023poster

Currently, LiDAR-based 3D detectors are broadly categorized into two groups, namely, BEV-based detectors and cluster-based detectors. BEV-based detectors capture the contextual information from the Bird's Eye View (BEV) and fill their center voxels via feature diffusion with a stack of convolution l…

Cited by 6SourcePDFScholar
2023

Reinforcement Learning for Robot Navigation with Adaptive Forward Simulation Time (AFST) in a Semi-Markov Model

IROS 2023poster

Deep reinforcement learning (DRL) algorithms have proven effective in robot navigation, especially in unknown environments, by directly mapping perception inputs into robot control commands. However, most existing methods ignore the local minimum problem in navigation and thereby cannot handle compl…

Cited by 0SourcecodeScholar
2023

Training a Non-Cooperator to Identify Vulnerabilities and Improve Robustness for Robot Navigation

RA-L 2023

Autonomous mobile robots have become popular in various applications coexisting with humans, which requires robots to navigate efficiently and safely in crowd environments with diverse pedestrians. Pedestrians may cooperate with the robot by avoiding it actively or ignoring the robot during their wa

Cited by 2SourceScholar
2022

${\mathsf{EZFusion}}$: A Close Look at the Integration of LiDAR, Millimeter-Wave Radar, and Camera for Accurate 3D Object Detection and Tracking

RA-L 2022

A recent trend is to combine multiple sensors ( <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">i.e.</i> , cameras, LiDARs and millimeter-wave Radars) to achieve robust multi-modal perception for autonomous systems such as self-driving vehicles. Alth

Cited by 11SourceScholar
2022

A Reinforcement Learning Method for Motion Control With Constraints on an HPN Arm

RA-L 2022

Soft robotic arms have shown great potential toward applications to human daily lives, which is mainly due to their infinite passive degrees of freedom and intrinsic safety. There are tasks in lives that require the motion of the robot to meet some certain pose constraints that have not been impleme

Cited by 6SourceScholar
2022

Learning to Socially Navigate in Pedestrian-rich Environments with Interaction Capacity

ICRA 2022poster

Existing navigation policies for autonomous robots tend to focus on collision avoidance while ignoring human-robot interactions in social life. For instance, robots can pass along the corridor safer and easier if pedestrians notice them. Sounds have been considered as an efficient way to attract the…

Cited by 18SourceScholar
2022

Manipulation Planning From Demonstration Via Goal-Conditioned Prior Action Primitive Decomposition and Alignment

RA-L 2022

Manipulation plays a vital role in robotics but is left unsolved. Recent work attempts to leverage the hierarchical structure of tasks via using action primitives. However, due to trajectory distribution shift, prior action primitives could hardly be adapted to new tasks. In this letter, we propose

Cited by 15SourceScholar
2022

PFilter: Building Persistent Maps through Feature Filtering for Fast and Accurate LiDAR-based SLAM

IROS 2022poster

Simultaneous localization and mapping (SLAM) based on laser sensors has been widely adopted by mobile robots and autonomous vehicles. These SLAM systems are required to support accurate localization with limited computational resources. In particular, point cloud registration, i.e., the process of m…

Cited by 21SourceScholar
2022

Transferring Knowledge from Structure-aware Self-attention Language Model to Sequence-to-Sequence Semantic Parsing

COLING 2022main

Semantic parsing considers the task of mapping a natural language sentence into a target formal representation, where various sophisticated sequence-to-sequence (seq2seq) models have been applied with promising results. Generally, these target representations follow a syntax formalism that limits pe…

Cited by 2SourcePDFScholar
2021

3D Segmentation Learning From Sparse Annotations and Hierarchical Descriptors

RA-L 2021

One of the main obstacles to 3D semantic segmentation is the significant amount of endeavor required to generate expensive point-wise annotations for fully supervised training. To alleviate manual efforts, we propose GIDSeg, a novel approach that can simultaneously learn segmentation from sparse ann

Cited by 3SourceScholar
2021

Crowd-Aware Robot Navigation for Pedestrians with Multiple Collision Avoidance Strategies via Map-based Deep Reinforcement Learning

IROS 2021poster

It is challenging for a mobile robot to navigate through human crowds. Existing approaches usually assume that pedestrians follow a predefined collision avoidance strategy, like social force model (SFM) or optimal reciprocal collision avoidance (ORCA). However, their performances commonly need to be…

Cited by 41SourceScholar
2021

DRQN-based 3D Obstacle Avoidance with a Limited Field of View

IROS 2021poster

In this paper, we propose a map-based end-to-end DRL approach for three-dimensional (3D) obstacle avoidance in a partially observed environment, which is applied to achieve autonomous navigation for an indoor mobile robot using a depth camera with a narrow field of view. We first train a neural netw…

Cited by 10SourceScholar
2021

Towards an Online RRT-based Path Planning Algorithm for Ackermann-steering Vehicles

ICRA 2021poster

It is challenging to develop an online path planning algorithm for Ackermann-steering vehicles to find collision-free and kinematically-feasible paths, that is efficient for dense environments, adaptable to various environments, and suitable for environments with narrow passages. In this paper, we p…

Cited by 10SourcecodeScholar
2019

A Multi-Domain Feature Learning Method for Visual Place Recognition

ICRA 2019poster

Visual Place Recognition (VPR) is an important component in both computer vision and robotics applications, thanks to its ability to determine whether a place has been visited and where specifically. A major challenge in VPR is to handle changes of environmental conditions including weather, season…

Cited by 37SourceScholar
2019

MRS-VPR: a multi-resolution sampling based global visual place recognition method

ICRA 2019poster

Place recognition and loop closure detection are challenging for long-term visual navigation tasks. SeqSLAM is considered to be one of the most successful approaches to achieve long-term localization under varying environmental conditions and changing viewpoints. SeqSLAM uses a brute-force sequentia…

Cited by 20SourceScholar