← Search

Jiming Chen

52 accepted papers

2026

A Super-Resolution and Multi-Axis Tactile Sensor with Soft Artificial Skin

RSS 2026poster

To achieve human-like skin tactile perception with super-resolution, the method of introducing a soft layer on sensing array has attracted increasing attention. Due to the limitations of sensing units principle, most existing tactile sensors can only sense normal force. However, multi-dimensional fo…

Cited by 0SourceScholar
2026

DiffPBR: Point-Based Rendering via Spatial-Aware Residual Diffusion

ICLR 2026poster

Neural radiance fields and 3D Gaussian splatting (3DGS) have significantly advanced 3D reconstruction and novel view synthesis (NVS). Yet, achieving high-fidelity and view-consistent renderings directly from point clouds---without costly per-scene optimization---remains a core challenge. In this wor…

Cited by 0SourceScholar
2026

MIMIC: Mask-Injected Manipulation Video Generation with Interaction Control

ICLR 2026poster

Embodied intelligence faces a fundamental bottleneck from limited large-scale interaction data. Video generation offers a scalable alternative, but manipulation videos remain particularly challenging, as they require capturing subtle, contact-rich dynamics. Despite recent advances, video diffusion m…

Cited by 0SourceScholar
2026

The Developments and Challenges towards Dexterous and Embodied Robotic Manipulation: A Survey

ICRA 2026poster

Achieving human-like dexterous robotic manipulation remains a central goal and a pivotal challenge in robotics. The development of Artificial Intelligence (AI) has allowed rapid progress in robotic manipulation. This survey summarizes the evolution of robotic manipulation from mechanical programming…

2026

VELR: Efficient Video Reward Feedback via Ensemble Latent Reward Models

ICML 2026poster

Reward feedback learning (ReFL) is effective for both text-to-image (T2I) and text-to-video (T2V) generation with image reward models (RMs). However, image RMs are misaligned with temporal objectives of T2V, motivating ReFL with video reward models. Nevertheless, directly deploying video RMs is impr…

Cited by 0SourceScholar
2025

A Hybrid Mapping Method: Balancing Efficiency and Intuitiveness in Lateral Teleoperation

IROS 2025

Mobile manipulators integrate the locomotion flexibility of quadruped robots with the operational capabilities of robotic manipulators. This integrated system is particularly effective for teleoperating explosive ordnance disposal (EOD) tasks in hazardous environments, enabling the safe handling of

Cited by 0SourceScholar
2025

AF-RLIO: Adaptive Fusion of Radar-LiDAR-Inertial Information for Robust Odometry in Challenging Environments

ICRA 2025

In robotic navigation, maintaining precise pose estimation and navigation in complex and dynamic environments is crucial. However, environmental challenges such as smoke, tunnels, and adverse weather can significantly degrade the performance of single-sensor systems like LiDAR or GPS, compromising t

Cited by 3SourcecodeScholar
2025

Can't Slow Me Down: Learning Robust and Hardware-Adaptive Object Detectors against Latency Attacks for Edge Devices

CVPR 2025poster

Object detection is a fundamental enabler for many real-time downstream applications such as autonomous driving, augmented reality and supply chain management. However, the algorithmic backbone of neural networks is brittle to imperceptible perturbations in the system inputs, which were generally kn…

2025

DHC-ME: A Decentralized Hybrid Cooperative Approach for Multi-Robot Autonomous Exploration

IROS 2025

Multi-robot exploration in unknown environments is a fundamental task for multi-robot systems, which requires the coordination of the robots to avoid collisions and conflicts while performing task allocation. Existing exploration strategies improve the efficiency of multi-robot exploration by modeli

Cited by 0SourcecodeScholar
2025

Dashing for the Golden Snitch: Multi-Drone Time-Optimal Motion Planning with Multi-Agent Reinforcement Learning

ICRA 2025

Recent innovations in autonomous drones have facilitated time-optimal flight in single-drone configurations, and enhanced maneuverability in multi-drone systems by applying optimal control and learning-based methods. However, few studies have achieved time-optimal motion planning for multi-drone sys

Cited by 9SourcecodeScholar
2025

Debiasing Trace Guidance: Top-down Trace Distillation and Bottom-up Velocity Alignment for Unsupervised Anomaly Detection

ICCV 2025poster

The leak of anomalous information from input condition poses a great challenge to reconstruction-based anomaly detection. Recent diffusion-based methods respond to this issue by suppressing anomaly information for condition injection or in-sampling inversion. However, since they treat conditions as…

Cited by 0SourcePDFScholar
2025

Fed-DFA: Federated Distillation for Heterogeneous Model Fusion Through the Adversarial Lens

AAAI 2025technical

Most of the federated learning techniques are limited to homogeneous model fusion. With the rapid growth of smart applications on resource-constrained edge devices, it becomes a barrier to accommodate their heterogeneous computing power and memory in the real world. Federated Distillation is a promi…

Cited by 1SourcePDFScholar
2025

Gate-Aware Online Planning for Two-Player Autonomous Drone Racing

ICRA 2025

The flying speed of autonomous quadrotors has increased significantly in the field of autonomous drone racing. However, most research primarily focuses on the aggressive flight of a single quadrotor, simplifying the racing gate traversal problem to a waypoint passing problem that neglects the orient

Cited by 5SourceScholar
2025

Hand-held Object Reconstruction from RGB Video with Dynamic Interaction

CVPR 2025poster

This work aims to reconstruct the 3D geometry of a rigid object manipulated by one or both hands using monocular RGB video. Previous methods rely on Structure-from-Motion or hand priors to estimate relative motion between the object and camera, which typically assume textured objects or single-hand…

2025

Multi-Robot Autonomous 3D Reconstruction Using Gaussian Splatting With Semantic Guidance

RA-L 2025

Implicit neural representations and 3D Gaussian splatting (3DGS) have shown great potential for scene reconstruction. Recent studies have expanded their applications in autonomous reconstruction through task assignment methods. However, these methods are mainly limited to a single robot, and rapid r

Cited by 4SourceScholar
2025

Online Motion Planning for Quadrotor Multi-Point Navigation Using Efficient Imitation Learning-Based Strategy

IROS 2025

Over the past decade, there has been a remarkable surge in utilizing quadrotors for various purposes due to their simple structure and aggressive maneuverability. One of the key challenges is online time-optimal trajectory generation and control technique. This paper proposes an imitation learning-b

Cited by 1SourceScholar
2025

RFMPose: Generative Category-level Object Pose Estimation via Riemannian Flow Matching

NeurIPS 2025poster

We introduce RFMPose, a novel generative framework for category-level 6D object pose estimation that learns deterministic pose trajectories through Riemannian Flow Matching (RFM). Existing discriminative approaches struggle with multi-hypothesis predictions (e.g., symmetry ambiguities) and often req…

Cited by 0SourceScholar
2025

Safety-Critical Online Quadrotor Trajectory Planner for Agile Flights in Unknown Environments

ICRA 2025

Autonomous high-speed flight in unknown, clut-tered environments is essential for a variety of quadrotor applications, such as inspection, search, and rescue. In this study, we propose a novel trajectory planner designed to achieve efficient, high-speed, collision-free flights in such environments.

Cited by 3SourceScholar
2025

TaskExp: Enhancing Generalization of Multi-Robot Exploration with Multi-Task Pre-Training

ICRA 2025

We aim to develop a general multi-agent reinforcement learning (MARL) policy that enables a group of robots to efficiently explore large-scale, unknown environments with random pose initialization. Existing MARL-based multi-robot exploration methods face challenges in reliably mapping observations t

Cited by 1SourceScholar
2025

Temporal-Spatial Representation Fusion for Dexterous Manipulation Learning with Unpaired Visual-Action Data

IROS 2025

Supervised behavioral cloning using robot visual-action data has been widely investigated in robot manipulation. However, these methods typically require simultaneous acquisition of visual and action data, which makes them difficult to utilize unpaired visual-action datasets: e.g. videos on Internet

Cited by 0SourceScholar
2025

UpViTaL: Unpaired Visual-Tactile Self-Supervised Representation Learning for Dexterous Robotic Manipulation

ICRA 2025

Visual and tactile pretraining have been extensively studied in dexterous robot manipulation tasks. However, existing methods typically require the simultaneous acquisition of visual and tactile data, making it difficult to utilize low-cost, unpaired visual-tactile datasets. Moreover, these methods

Cited by 1SourceScholar
2025

VTAO-BiManip: Masked Visual-Tactile-Action Pre-training with Object Understanding for Bimanual Dexterous Manipulation

IROS 2025

Bimanual dexterous manipulation remains a significant challenge in robotics due to the high DoFs of each hand and their coordination. Existing single-hand manipulation techniques often leverage human demonstrations to guide RL methods but fail to generalize to complex bimanual tasks involving multip

Cited by 4SourceScholar
2025

VTDexManip: A Dataset and Benchmark for Visual-tactile Pretraining and Dexterous Manipulation with Reinforcement Learning

ICLR 2025poster

Vision and touch are the most commonly used senses in human manipulation. While leveraging human manipulation videos for robotic task pretraining has shown promise in prior works, it is limited to image and language modalities and deployment to simple parallel grippers. In this paper, aiming to addr…

2024

A Consistency-Aware Spot-Guided Transformer for Versatile and Hierarchical Point Cloud Registration

NeurIPS 2024poster

Deep learning-based feature matching has shown great superiority for point cloud registration in the absence of pose priors. Although coarse-to-fine matching approaches are prevalent, the coarse matching of existing methods is typically sparse and loose without consideration of geometric consistency…

2024

AnomalyCLIP: Object-agnostic Prompt Learning for Zero-shot Anomaly Detection

ICLR 2024poster

Zero-shot anomaly detection (ZSAD) requires detection models trained using auxiliary data to detect anomalies without any training sample in a target dataset. It is a crucial task when training data is not accessible due to various concerns, e.g., data privacy, yet it is challenging since the models…

2024

Autonomous Implicit Indoor Scene Reconstruction with Frontier Exploration

ICRA 2024poster

Implicit neural representations have demonstrated significant promise for 3D scene reconstruction. Recent works have extended their applications to autonomous implicit reconstruction through the Next Best View (NBV) based method. However, the NBV method cannot guarantee complete scene coverage and o…

Cited by 3SourceScholar
2024

CAMInterHand: Cooperative Attention for Multi-View Interactive Hand Pose and Mesh Reconstruction

ICRA 2024poster

Interactive hand mesh reconstruction from singleview images poses a significant challenge with the severe occlusion and depth ambiguity inherent in interactive hand gestures. Recent approaches that employ probabilistic models and tokenpruned techniques have shown decent results in multi-view human b…

Cited by 0SourceScholar
2024

Differentially Private No-regret Exploration in Adversarial Markov Decision Processes

UAI 2024poster

We study learning adversarial Markov decision process (MDP) in the episodic setting under the constraint of differential privacy (DP). This is motivated by the widespread applications of reinforcement learning (RL) in non-stationary and even adversarial scenarios, where protecting users’ sensitive i…

Cited by 1SourcePDFScholar
2024

In-Hand 3D Object Reconstruction from a Monocular RGB Video

AAAI 2024technical

Our work aims to reconstruct a 3D object that is held and rotated by a hand in front of a static RGB camera. Previous methods that use implicit neural representations to recover the geometry of a generic hand-held object from multi-view images achieved compelling results in the visible part of the o…

2024

InterRep: A Visual Interaction Representation for Robotic Grasping

ICRA 2024poster

Recently, pre-trained vision models have gained significant attention in motor control, showcasing impressive performance across diverse robotic learning tasks. While previous works predominantly concentrate on the significance of the pre-training phase, the equally important task of extracting more…

Cited by 1SourceScholar
2024

KDD-LOAM: Jointly Learned Keypoint Detector and Descriptors Assisted LiDAR Odometry and Mapping

ICRA 2024poster

Sparse keypoint matching based on distinct 3D feature representations can improve the efficiency and robustness of point cloud registration. Existing learning-based 3D descriptors and keypoint detectors are either independent or loosely coupled, so they cannot fully adapt to each other. In this work…

Cited by 5SourceScholar
2024

LESS-Map: Lightweight and Evolving Semantic Map in Parking Lots for Long-term Self-Localization

ICRA 2024poster

Precise and long-term stable localization is essential in parking lots for tasks like autonomous driving or autonomous valet parking, etc. Existing methods rely on a fixed and memory-inefficient map, which lacks robust data association approaches. And it is not suitable for precise localization or l…

Cited by 2SourceScholar
2024

MAexp: A Generic Platform for RL-based Multi-Agent Exploration

ICRA 2024poster

The sim-to-real gap poses a significant challenge in RL-based multi-agent exploration due to scene quantization and action discretization. Existing platforms suffer from the inefficiency in sampling and the lack of diversity in Multi-Agent Reinforcement Learning (MARL) algorithms across different sc…

Cited by 4SourcecodeScholar
2024

Masked Visual-Tactile Pre-training for Robot Manipulation

ICRA 2024poster

Recent works on the pretraining for robot manipulation have demonstrated that representations learning from large human manipulation data can generalize well to new manipulation tasks and environments. However, these approaches mainly focus on human vision or natural language, neglecting tactile fee…

Cited by 6SourceScholar
2024

PointAD: Comprehending 3D Anomalies from Points and Pixels for Zero-shot 3D Anomaly Detection

NeurIPS 2024poster

Zero-shot (ZS) 3D anomaly detection is a crucial yet unexplored field that addresses scenarios where target 3D training samples are unavailable due to practical concerns like privacy protection. This paper introduces PointAD, a novel approach that transfers the strong generalization capabilities of…

2024

TPGP: Temporal-Parametric Optimization with Deep Grasp Prior for Dexterous Motion Planning

ICRA 2024poster

Grasping motion planning aims to find a feasible grasping trajectory in the configuration space given an input target grasp. While optimizing grasp motion with two or three-fingered grippers has been well studied, the study on natural grasp motion planning with a dexterous hand remains a very challe…

Cited by 2SourceScholar
2024

The Joint-Space Reconstruction of Human Fingers by using a Highly Under-Actuated Exoskeleton

ICRA 2024poster

Hand motion tracking is essential in many fields, e.g., immersive virtual reality, teleoperation of robotic hand, and hand rehabilitation of stroke patient, as human hand plays a crucial role in our daily life. The highly under-actuated hand exoskeleton, which can track the 6-DoF motions of each fin…

Cited by 4SourceScholar
2024

iBoW3D: Place Recognition Based on Incremental and General Bag of Words in 3D Scans

ICRA 2024poster

Existing methods for place recognition in 3D point clouds either ignore partial structure information by converting 3D scans to 2D images or construct constrained bag-of-words (BoW) representations reliant on specific feature extraction algorithms. In this paper, we propose a novel method based on i…

Cited by 0SourceScholar
2023

Aggressive Trajectory Generation for a Swarm of Autonomous Racing Drones

IROS 2023poster

Autonomous drone racing is becoming an excellent platform to challenge quadrotors' autonomy techniques including planning, navigation and control technologies. However, most research on this topic mainly focuses on single drone scenarios. In this paper, we describe a novel time-optimal trajectory ge…

Cited by 6SourceScholar
2023

Contact2Grasp: 3D Grasp Synthesis via Hand-Object Contact Constraint

IJCAI 2023poster

3D grasp synthesis generates grasping poses given an input object. Existing works tackle the problem by learning a direct mapping from objects to the distributions of grasping poses. However, because the physical contact is sensitive to small changes in pose, the high-nonlinear mapping between 3D ob…

Cited by 15SourcePDFScholar
2023

Detecting Multivariate Time Series Anomalies with Zero Known Label

AAAI 2023technical

Multivariate time series anomaly detection has been extensively studied under the one-class classification setting, where a training dataset with all normal instances is required. However, preparing such a dataset is very laborious since each single data instance should be fully guaranteed to be nor…

2023

DexRepNet: Learning Dexterous Robotic Grasping Network with Geometric and Spatial Hand-Object Representations

IROS 2023poster

Robotic dexterous grasping is a challenging problem due to the high degree of freedom (DoF) and complex contacts of multi-fingered robotic hands. Existing deep re-inforcement learning (DRL) based methods leverage human demonstrations to reduce sample complexity due to the high dimensional action spa…

Cited by 18SourceScholar
2023

Efficient View Path Planning for Autonomous Implicit Reconstruction

ICRA 2023poster

Implicit neural representations have shown promising potential for 3D scene reconstruction. Recent work applies it to autonomous 3D reconstruction by learning information gain for view path planning. Effective as it is, the computation of the information gain is expensive, and compared with that usi…

Cited by 20SourceScholar
2023

F&F Attack: Adversarial Attack against Multiple Object Trackers by Inducing False Negatives and False Positives

ICCV 2023poster

Multi-object tracking (MOT) aims to build moving trajectories for number-agnostic objects. Modern multi-object trackers commonly follow the tracking-by-detection strategy. Therefore, fooling detectors can be an effective solution but it usually requires attacks in multiple successive frames, resulti…

Cited by 10PDFScholar
2023

ImmFusion: Robust mmWave-RGB Fusion for 3D Human Body Reconstruction in All Weather Conditions

ICRA 2023poster

3D human reconstruction from RGB images achieves decent results in good weather conditions but degrades dramatically in rough weather. Complementary, mmWave radars have been employed to reconstruct 3D human joints and meshes in rough weather. However, combining RGB and mmWave signals for robust all-…

Cited by 24SourceScholar
2023

InterTracker: Discovering and Tracking General Objects Interacting with Hands in the Wild

IROS 2023poster

Understanding human interaction with objects is an important research topic for embodied Artificial Intelligence and identifying the objects that humans are interacting with is a primary problem for interaction understanding. Existing methods rely on frame-based detectors to locate interacting objec…

Cited by 1SourceScholar
2023

NeurAR: Neural Uncertainty for Autonomous 3D Reconstruction With Implicit Neural Representations

RA-L 2023

Implicit neural representations have shown compelling results in offline 3D reconstruction and also recently demonstrated the potential for online SLAM systems. However, applying them to autonomous 3D reconstruction, where a robot is required to explore a scene and plan a view path for the reconstru

Cited by 91SourceScholar
2023

Seal-3D: Interactive Pixel-Level Editing for Neural Radiance Fields

ICCV 2023poster

With the popularity of implicit neural representations, or neural radiance fields (NeRF), there is a pressing need for editing methods to interact with the implicit 3D models for tasks like post-processing reconstructed scenes and 3D content creation. While previous works have explored NeRF editing…

Cited by 18PDFcodeScholar