← Search

Jia Pan

121 accepted papers

2026

Binary-Gaussian: Compact and Progressive Representation for 3D Gaussian Segmentation

AAAI 2026technical

3D Gaussian Splatting (3D-GS) has emerged as an efficient 3D representation and a promising foundation for semantic tasks like segmentation. However, existing 3D-GS-based segmentation methods typically rely on high-dimensional category features, which introduce substantial memory overhead. Moreover,

Cited by 0SourcePDFScholar
2026

NeuPAN: Direct Point Robot Navigation with End-to-End Model-Based Learning (Abstract Reprint)

AAAI 2026technical

Navigating a nonholonomic robot in a cluttered, unknown environment requires accurate perception and precise motion control for real-time collision avoidance. This article presents neural proximal alternating-minimization network (NeuPAN): a real-time, highly accurate, map-free, easy-to-deploy, and

Cited by 0SourcePDFScholar
2026

Semantically Structured Mixture-of-Experts for Compositional Robotic Manipulation

RSS 2026poster

Diffusion-based policies have established a new standard for precise robotic manipulation but face a critical scalability bottleneck: high-performance models are computationally expensive, while lightweight alternatives often fail to generalize across diverse multi-task environments. Mixture-of-Expe…

Cited by 0SourceScholar
2026

Shared Autonomy Assisted by Impedance-Driven Anisotropic Guidance Field

RA-L 2026

Shared autonomy (SA) enables robots to infer human intent and assist in its achievement. While most research focuses on improving intent inference, it overlooks whether humans can understand the robot's intent in return. Without such mutual understanding, collaboration becomes less effective, degrad

Cited by 0SourceScholar
2025

A Joint Learning of Force Feedback of Robotic Manipulation and Textual Cues for Granular Materials Classification

RA-L 2025

Granular materials (GMs) are formed by a collection of particles. Even if their visual representation is straightforward, it can be seriously affected in the visually constrained environment. Based on frequency features observed in force signals, this paper proposes a non-visual classifier, <bold xm

Cited by 22SourceScholar
2025

Col-OLHTR: A Novel Framework for Multimodal Online Handwritten Text Recognition

ICASSP 2025accepted

Online Handwritten Text Recognition (OLHTR) has gained considerable attention for its diverse range of applications. Current approaches usually treat OLHTR as a sequence recognition task, employing either a single trajectory or image encoder, or multi-stream encoders, combined with a CTC or attentio…

Cited by 0SourceScholar
2025

DAWN: Dynamic Frame Avatar with Non-autoregressive Diffusion Framework for Talking head Video Generation

ICLR 2025poster

Talking head generation intends to produce vivid and realistic talking head videos from a single portrait and speech audio clip. Although significant progress has been made in diffusion-based talking head generation, almost all methods rely on autoregressive strategies, which suffer from limited con…

2025

EventSync: Joint Recovery of Temporal Offsets and Relative Orientations for Wide-Baseline Event Cameras

IROS 2025

Event-Based cameras offer significant advantages due to their high temporal resolution and low power consumption. However, when deploying multiple such cameras, a critical challenge emerges: each camera operates on an independent time system, resulting in temporal misalignment that severely degrades

Cited by 0SourcecodeScholar
2025

GauSS-MI: Gaussian Splatting Shannon Mutual Information for Active 3D Reconstruction

RSS 2025poster

This research tackles the challenge of real-time active view selection and uncertainty quantification on visual quality for active 3D reconstruction. Visual quality is a critical aspect of 3D reconstruction. Recent advancements such as Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) h…

Cited by 0PDFcodeScholar
2025

Incorporating Spatial Cues in Modular Speaker Diarization for Multi-channel Multi-party Meetings

ICASSP 2025accepted

Although fully end-to-end speaker diarization systems have made significant progress in recent years, modular systems often achieve superior results in real-world scenarios due to their greater adaptability and robustness. Historically, modular speaker diarization methods have seldom discussed how t…

Cited by 0SourceScholar
2025

Lattice Boltzmann Model for Learning Real-World Pixel Dynamicity

NeurIPS 2025poster

This work proposes the Lattice Boltzmann Model (LBM) to learn real-world pixel dynamicity for visual tracking. LBM decomposes visual representations into dynamic pixel lattices and solves pixel motion states through collision-streaming processes. Specifically, the high-dimensional distribution of t…

Cited by 0SourceScholar
2025

Magnetometer-Calibrated Hybrid Transformer for Robust Inertial Tracking in Robotics

ICRA 2025

Inertial tracking is vital for autonomous robots and has gained popularity with the ubiquity of low-cost Inertial Measurement Units (IMUs) and deep learning-powered tracking algorithms. Existing works, however, have not fully utilized IMU measurements, particularly magnetometers, nor maximized the p

Cited by 1SourcecodeScholar
2025

Optimizing Efficiency of Mixed Traffic Through Reinforcement Learning: A Topology-Independent Approach and Benchmark

ICRA 2025

This paper presents a mixed traffic control policy designed to optimize traffic efficiency across diverse road topologies, addressing issues of congestion prevalent in urban environments. A model-free reinforcement learning (RL) approach is developed to manage large-scale traffic flow, using data co

Cited by 0SourceScholar
2025

Understanding Particles From Video: Property Estimation of Granular Materials via Visuo-Haptic Learning

RA-L 2025

Granular materials (GMs) are ubiquitous in daily life. Understanding their properties is also important, especially in agriculture and industry. However, existing works require dedicated measurement equipment and also need large human efforts to handle a large number of particles. In this paper, we

Cited by 3SourceScholar
2024

A Spatial Long-Term Iterative Mask Estimation Approach for Multi-Channel Speaker Diarization and Speech Recognition

ICASSP 2024accepted

Deep learning (DL)-based speaker diarization methods have proven powerful performance comparing to traditional clustering-based methods for multi-talker speech diarization and recognition in farfield scenes. However, most DL-based approaches cannot utilize the spatial information well due to the poo…

Cited by 0SourceScholar
2024

DaDiff: Domain-aware Diffusion Model for Nighttime UAV Tracking

IROS 2024poster

Domain adaptation is an inspiring solution to the misalignment issue of day/night image features for nighttime UAV tracking. However, the one-step adaptation paradigm is inadequate in addressing the prevalent difficulties posed by low-resolution (LR) objects when viewed from the UAVs at night, owing…

Cited by 1SourcecodeScholar
2024

Efficient Planar Fabric Repositioning: Deformation-Aware RRT* for Non-Prehensile Fabric Manipulation

RA-L 2024

Fabrics present significant challenges to robotic manipulation due to their complex dynamics and infinite degrees of freedom. This letter proposes a non-prehensile approach to aligning a fabric cut piece to a specified target pose, which is a common step for many garment manufacturing tasks. Compare

Cited by 4SourceScholar
2024

HHD-GP: Incorporating Helmholtz-Hodge Decomposition into Gaussian Processes for Learning Dynamical Systems

NeurIPS 2024poster

Machine learning models provide alternatives for efficiently recognizing complex patterns from data, but the main concern in applying them to modeling physical systems stems from their physics-agnostic design, leading to learning methods that lack interpretability, robustness, and data efficiency. T…

Cited by 0SourcePDFScholar
2024

Implicit Enhancement of Target Speaker in Speaker-Adaptive ASR through Efficient Joint Optimization

ICASSP 2024accepted

In multi-speaker scenarios, automatic speech recognition (ASR) models rely on pre-processed audio after speaker separation. However, when the target speaker is not accurately separated, ASR models face limitations in reaching their peak performance. To address this issue, we propose a speaker-adapti…

Cited by 0SourceScholar
2024

LASIL: Learner-Aware Supervised Imitation Learning For Long-term Microscopic Traffic Simulation

CVPR 2024poster

Microscopic traffic simulation plays a crucial role in transportation engineering by providing insights into individual vehicle behavior and overall traffic flow. However creating a realistic simulator that accurately replicates human driving behaviors in various traffic conditions presents signific…

2024

Language-Augmented Symbolic Planner for Open-World Task Planning

RSS 2024poster

Enabling robotic agents to perform complex long-horizon tasks has been a long-standing goal in robotics and artificial intelligence (AI). Despite the potential shown by large language models (LLMs), their planning capabilities remain limited to short-horizon tasks and they are unable to replace the…

2024

Memory-Constrained Semantic Segmentation for Ultra-High Resolution UAV Imagery

RA-L 2024

Ultra-high resolution image segmentation poses a formidable challenge for UAVs with limited computation resources. Moreover, with multiple deployed tasks (e.g., mapping, localization, and decision making), the demand for a memory efficient model becomes more urgent. This letter delves into the intri

Cited by 13SourceScholar
2024

NAMER: Non-Autoregressive Modeling for Handwritten Mathematical Expression Recognition

ECCV 2024poster

"Recently, Handwritten Mathematical Expression Recognition (HMER) has gained considerable attention in pattern recognition for its diverse applications in document understanding. Current methods typically approach HMER as an image-to-sequence generation task within an autoregressive (AR) encoder-dec…

Cited by 2SourcePDFScholar
2024

NBV/NBC Planning Considering Confidence Obtained From Shape Completion Learning

RA-L 2024

In this letter, we present a novel approach for planning an object's Next Best Views (NBV) so that a depth camera can collect the object's surface point cloud and reconstruct its 3D model with a small number of consequent views. Our focus is especially on thin and curved metal plates, and we use a r

Cited by 3SourceScholar
2024

NetTrack: Tracking Highly Dynamic Objects with a Net

CVPR 2024poster

The complex dynamicity of open-world objects presents non-negligible challenges for multi-object tracking (MOT) often manifested as severe deformations fast motion and occlusions. Most methods that solely depend on coarse-grained object cues such as boxes and the overall appearance of the object are…

Cited by 15SourcePDFScholar
2024

Progressive Representation Learning for Real-Time UAV Tracking

IROS 2024poster

Visual object tracking has significantly promoted autonomous applications for unmanned aerial vehicles (UAVs). However, learning robust object representations for UAV tracking is especially challenging in complex dynamic environments, when confronted with aspect ratio change and occlusion. These cha…

Cited by 5SourcecodeScholar
2024

Prompt-Driven Temporal Domain Adaptation for Nighttime UAV Tracking

IROS 2024poster

Nighttime UAV tracking under low-illuminated scenarios has achieved great progress by domain adaptation (DA). However, previous DA training-based works are deficient in narrowing the discrepancy of temporal contexts for UAV trackers. To address the issue, this work proposes a prompt-driven temporal…

Cited by 3SourcecodeScholar
2024

Semantics-Aware Receding Horizon Planner for Object-Centric Active Mapping

RA-L 2024

The escalating demands for real-time scene comprehension in modern industries underscore the growing significance of semantic information in the daily tasks of robots, particularly in areas like autonomous inspection and target searching. This letter introduces a semantics-aware receding horizon pla

Cited by 14SourceScholar
2024

The Multimodal Information Based Speech Processing (MISP) 2023 Challenge: Audio-Visual Target Speaker Extraction

ICASSP 2024accepted

Previous Multimodal Information based Speech Processing (MISP) challenges mainly focused on audio-visual speech recognition (AVSR) with commendable success. However, the most advanced back-end recognition systems often hit performance limits due to the complex acoustic environments. This has prompte…

Cited by 0SourceScholar
2023

An Experimental Study on Sound Event Localization and Detection Under Realistic Testing Conditions

ICASSP 2023accepted

We study four data augmentation (DA) techniques and two model architectures on realistic data for sound event localization and detection (SELD). First, based on ResNet-Conformer (RC), we compare the four DA approaches on the realistic DCASE 2022 SELD test set which is often not easy to handle due to…

Cited by 0SourceScholar
2023

Cascaded Denoising Transformer for UAV Nighttime Tracking

RA-L 2023

The automation of unmanned aerial vehicles (UAVs) has been greatly promoted by visual object tracking methods with onboard cameras. However, the random and complicated real noise produced by the cameras seriously hinders the performance of state-of-the-art (SOTA) UAV trackers, especially in low-illu

Cited by 11SourceScholar
2023

Fast Event-based Double Integral for Real-time Robotics

ICRA 2023poster

Motion deblurring is a critical ill-posed problem that is important in many vision-based robotics applications. The recently proposed event-based double integral (EDI) provides a theoretical framework for solving the deblurring prob-lem with the event camera and generating clear images at high frame…

Cited by 6SourcecodeScholar
2023

Hierarchical Temporal Transformer for 3D Hand Pose Estimation and Action Recognition From Egocentric RGB Videos

CVPR 2023poster

Understanding dynamic hand motions and actions from egocentric RGB videos is a fundamental yet challenging task due to self-occlusion and ambiguity. To address occlusion and ambiguity, we develop a transformer-based framework to exploit temporal information for robust estimation. Noticing the differ…

2023

Incorporating Visual Information Reconstruction into Progressive Learning for Optimizing audio-visual Speech Enhancement

ICASSP 2023accepted

Video information has been widely introduced to speech enhancement as its contribution at low signal-to-noise ratios (SNRs). Conventional audio-visual speech enhancement networks take noisy speech and video as input and learn features of clean speech directly. To reduce the large SNR gap between the…

Cited by 0SourceScholar
2023

Loss Function Design for DNN-Based Sound Event Localization and Detection on Low-Resource Realistic Data

ICASSP 2023accepted

This study focuses on the design of a loss function for a deep neural network (DNN)-based model with two branches, which is used to solve sound event localization and detection (SELD) on low-resource realistic data. To this end, we employ a secondary network for audio classification, which provides…

Cited by 0SourceScholar
2023

POMDP-Guided Active Force-Based Search for Robotic Insertion

IROS 2023poster

In robotic insertion tasks where the uncertainty exceeds the allowable tolerance, a good search strategy is essential for successful insertion and significantly influences efficiency. The commonly used blind search method is time-consuming and does not exploit the rich contact information. In this p…

Cited by 3SourceScholar
2023

RDA: An Accelerated Collision Free Motion Planner for Autonomous Navigation in Cluttered Environments

RA-L 2023

Autonomous motion planning is challenging in multi-obstacle environments due to nonconvex collision avoidance constraints. Directly applying numerical solvers to these nonconvex formulations fails to exploit the constraint structures, resulting in excessive computation time. In this letter, we prese

Cited by 49SourcecodeScholar
2023

Reducing the GAP Between Streaming and Non-Streaming Transducer-Based ASR by Adaptive Two-Stage Knowledge Distillation

ICASSP 2023accepted

Transducer is one of the mainstream frameworks for streaming speech recognition. There is a performance gap between the streaming and non-streaming transducer models due to limited context. To reduce this gap, an effective way is to ensure that their hidden and output distributions are consistent, w…

Cited by 0SourceScholar
2023

Self-Supervised Audio-Visual Speech Representations Learning by Multimodal Self-Distillation

ICASSP 2023accepted

In this work, we present a novel method, named AV2vec, for learning audio-visual speech representations by multimodal self-distillation. AV2vec has a student and a teacher module, in which the student performs a masked latent feature regression task using the multimodal target features generated onl…

Cited by 0SourceScholar
2023

Summary on the Multimodal Information Based Speech Processing (MISP) 2022 Challenge

ICASSP 2023accepted

The Multimodal Information based Speech Processing (MISP) 2022 challenge aimed to enhance speech processing performance in harsh acoustic environments by leveraging additional modalities such as video or text. The challenge included two tracks: audio-visual speaker diarization (AVSD) and audio-visua…

Cited by 0SourceScholar
2023

TacGNN: Learning Tactile-Based In-Hand Manipulation With a Blind Robot Using Hierarchical Graph Neural Network

RA-L 2023

In this letter, we propose a novel framework for tactile-based dexterous manipulation learning with a blind anthropomorphic robotic hand, i.e. without visual sensing. First, object-related states were extracted from the raw tactile signals by a graph-based perception model - TacGNN. The resulting ta

Cited by 36SourceScholar
2023

The Multimodal Information Based Speech Processing (Misp) 2022 Challenge: Audio-Visual Diarization And Recognition

ICASSP 2023accepted

The Multi-modal Information based Speech Processing (MISP) challenge aims to extend the application of signal processing technology in specific scenarios by promoting the research into wake-up words, speaker diarization, speech recognition, and other technologies. The MISP2022 challenge has two trac…

Cited by 0SourceScholar
2023

Tight Collision Probability for UAV Motion Planning in Uncertain Environment

IROS 2023poster

Operating unmanned aerial vehicles (UAVs) in complex environments that feature dynamic obstacles and external disturbances poses significant challenges, primarily due to the inherent uncertainty in such scenarios. Additionally, inaccurate robot localization and modeling errors further exacerbate the…

Cited by 10SourcecodeScholar
2023

Vision-based Six-Dimensional Peg-in-Hole for Practical Connector Insertion

ICRA 2023poster

We study six-dimensional (6D) perceptive peg-in-hole problem for practical connector insertion task in this paper. To enable the manipulator system to handle different types of pegs in complex environment, we develop a perceptive robotic assembly system that utilizes an in-hand RGB-D camera for peg-…

Cited by 11SourceScholar
2023

mCLIP: Multilingual CLIP via Cross-lingual Transfer

ACL 2023long

Large-scale vision-language pretrained (VLP) models like CLIP have shown remarkable performance on various downstream cross-modal tasks. However, they are usually biased towards English due to the lack of sufficient non-English image-text pairs. Existing multilingual VLP methods often learn retrieva…

2022

A Generalized Continuous Collision Detection Framework of Polynomial Trajectory for Mobile Robots in Cluttered Environments

RA-L 2022

In this letter, we introduce a generalized continuous collision detection (CCD) framework for the mobile robot along the polynomial trajectory in cluttered environments including various static obstacle models. Specifically, we find that the collision conditions between robots and obstacles could be

Cited by 16SourceScholar
2022

An Efficient Centralized Planner for Multiple Automated Guided Vehicles at the Crossroad of Polynomial Curves

RA-L 2022

In this letter, we introduce acentralized planner with low computational cost to schedule the motions of multiple Automated Guided Vehicles (AGVs) at the intersection of pre-defined polynomial curves. In particular, we find that the collision conditions between two AGVs along polynomial paths can be

Cited by 20SourceScholar
2022

Deep Reinforcement Learning for Robot Collision Avoidance With Self-State-Attention and Sensor Fusion

RA-L 2022

3D LiDAR sensors can provide 3D point clouds of the environment, and are widely used in automobile navigation; while 2D LiDAR sensors can only provide point cloud in a 2D sweeping plane, and then are only used for navigating robots of small height, e.g., floor mopping robots. In this letter, we prop

Cited by 56SourceScholar
2022

DiffSRL: Learning Dynamical State Representation for Deformable Object Manipulation With Differentiable Simulation

RA-L 2022

Dynamic state representation learning is essential for robot learning. Good latent space that can accurately describe dynamic transition and constraints can significantly accelerate reinforcement learning training as well as reduce motion planning complexity. However, deformable object have very com

Cited by 16SourceScholar
2022

DynamicFilter: an Online Dynamic Objects Removal Framework for Highly Dynamic Environments

ICRA 2022poster

Emergence of massive dynamic objects will diversify spatial structures when robots navigate in urban environments. Therefore, the online removal of dynamic objects is critical. In this paper, we introduce a novel online removal framework for highly dynamic urban environments. The framework consists…

Cited by 38SourceScholar
2022

Faithful Extreme Rescaling via Generative Prior Reciprocated Invertible Representations

CVPR 2022oral

This paper presents a Generative prior ReciprocAted Invertible rescaling Network (GRAIN) for generating faithful high-resolution (HR) images from low-resolution (LR) invertible images with an extreme upscaling factor (64x). Previous researches have leveraged the prior knowledge of a pretrained GAN m…

Cited by 15PDFcodeScholar
2022

High-Resolution Face Swapping via Latent Semantics Disentanglement

CVPR 2022poster

We present a novel high-resolution face swapping method using the inherent prior knowledge of a pre-trained GAN model. Although previous research can leverage generative priors to produce high-resolution results, their quality can suffer from the entangled semantics of the latent space. We explicitl…

Cited by 96PDFcodeScholar
2022

Learn to Predict How Humans Manipulate Large-Sized Objects From Interactive Motions

RA-L 2022

Understanding human intentions during interactions has been a long-lasting theme, that has applications in human-robot interaction, virtual reality and surveillance. In this study, we focus on full-body human interactions with large-sized daily objects and aim to predict the future states of objects

Cited by 35SourceScholar
2022

ModLaNets: Learning Generalisable Dynamics via Modularity and Physical Inductive Bias

ICML 2022oral

Deep learning models are able to approximate one specific dynamical system but struggle at learning generalisable dynamics, where dynamical systems obey the same laws of physics but contain different numbers of elements (e.g., double- and triple-pendulum systems). To relieve this issue, we proposed…

2022

Reinforcement Learned Distributed Multi-Robot Navigation With Reciprocal Velocity Obstacle Shaped Rewards

RA-L 2022

The challenges to solving the collision avoidance problem lie in adaptively choosing optimal robot velocities in complex scenarios full of interactive obstacles. In this letter, we propose a distributed approach for multi-robot navigation which combines the concept of reciprocal velocity obstacle (R

Cited by 147SourcecodeScholar
2022

Siamese Object Tracking for Vision-Based UAM Approaching with Pairwise Scale-Channel Attention

IROS 2022poster

Although the manipulating of the unmanned aerial manipulator (UAM) has been widely studied, vision-based UAM approaching, which is crucial to the subsequent manipulating, generally lacks effective design. The key to the visual UAM approaching lies in object tracking, while current UAM tracking typic…

Cited by 11SourcecodeScholar
2022

The First Multimodal Information Based Speech Processing (Misp) Challenge: Data, Tasks, Baselines And Results

ICASSP 2022accepted

In this paper we discuss the rational of the Multi-model Information based Speech Processing (MISP) Challenge, and provide a detailed description of the data recorded, the two evaluation tasks and the corresponding baselines, followed by a summary of submitted systems and evaluation results. The MIS…

Cited by 0SourceScholar
2022

Towards Making the Most of Cross-Lingual Transfer for Zero-Shot Neural Machine Translation

ACL 2022long

This paper demonstrates that multilingual pretraining and multilingual fine-tuning are both critical for facilitating cross-lingual transfer in zero-shot translation, where the neural machine translation (NMT) model is tested on source languages unseen during supervised training. Following this idea…

2022

Visual-tactile Sensing for Real-time Liquid Volume Estimation in Grasping

IROS 2022poster

We propose a deep visuo-tactile model for real-time estimation of the liquid inside a deformable container in a proprioceptive way. We fuse two sensory modalities, i.e., the raw visual inputs from the RGB camera and the tactile cues from our specific tactile sensor without any extra sensor calibrati…

Cited by 16SourceScholar
2021

A Computational Framework for Robot Hand Design via Reinforcement Learning

IROS 2021poster

Robot hand is essential for a fully functional robot and designing a good robot hand is a sophisticated job that challenges the designer’s knowledge and experience. This paper presents a computational framework for automatic optimal robot hand design based on reinforcement learning (RL), which consi…

Cited by 8SourceScholar
2021

An Efficient and Responsive Robot Motion Controller for Safe Human-Robot Collaboration

RA-L 2021

Safety and efficiency are two crucial factors for human-robot collaboration. It is challenging to ensure human safety while not sacrificing the task efficiency. In this letter, we present a reinforcement learning (RL) based method with a hazard estimator to balance these two factors. Our method has

Cited by 16SourceScholar
2021

An Overconstrained Robotic Leg with Coaxial Quasi-direct Drives for Omni-directional Ground Mobility

ICRA 2021poster

Planar mechanisms dominate modern designs of legged robots with remote actuator placement for robust agility in ground mobility. This paper presents a novel design of robotic leg modules using the Bennett linkage, driven by two coaxially arranged quasi-direct actuators capable of omnidirectional gro…

Cited by 7SourceScholar
2021

Efficient SE(3) Reachability Map Generation via Interplanar Integration of Intra-planar Convolutions

ICRA 2021poster

Convolution has been used for fast computation of reachability maps, but it has high computational costs when performing SE(3) convolution operations for general joint arrangements in industrial robots and 3D workspace. Its application is also limited to planar robots, 2D workspace, or robots with s…

Cited by 7SourceScholar
2021

Encirclement Guaranteed Cooperative Pursuit with Robust Model Predictive Control

IROS 2021poster

This paper studies a novel encirclement guaranteed cooperative pursuit problem involving N pursuers and a single evader in an unbounded two-dimensional game domain. Throughout the game, the pursuers are required to maintain encirclement of the evader, i.e., the evader should always stay inside the c…

Cited by 12SourceScholar
2021

Learning-Based Optoelectronically Innervated Tactile Finger for Rigid-Soft Interactive Grasping

RA-L 2021

This letter presents a novel design of a soft tactile finger with omni-directional adaptation using multi-channel optical fibers for rigid-soft interactive grasping. Machine learning methods are used to train a model for real-time prediction of force, torque, and contact using the tactile data colle

Cited by 22SourceScholar
2021

Projecting Your View Attentively: Monocular Road Scene Layout Estimation via Cross-View Transformation

CVPR 2021poster

HD map reconstruction is crucial for autonomous driving. LiDAR-based methods are limited due to the deployed expensive sensors and time-consuming computation. Camera-based methods usually need to separately perform road segmentation and view transformation, which often causes distortion and the abse…

Cited by 115PDFcodeScholar
2021

Zero-Shot Cross-Lingual Transfer of Neural Machine Translation with Multilingual Pretrained Encoders

EMNLP 2021main

Previous work mainly focuses on improving cross-lingual transfer for NLU tasks with a multilingual pretrained encoder (MPE), or improving the performance on supervised machine translation with BERT. However, it is under-explored that whether the MPE can help to facilitate the cross-lingual transfera…

2020

A Flexible Dual-Core Optical Waveguide Sensor for Simultaneous and Continuous Measurement of Contact Force and Position

IROS 2020poster

Having the merits of chemical inertness and immunity to electromagnetic interference, light weight, small size, and softness, optical waveguides have attracted much attention in making tactile sensors recently. This paper presents a new design of waveguide using two layers of cores, one of which has…

Cited by 6SourceScholar
2020

A Two-Stage Reinforcement Learning Approach for Multi-UAV Collision Avoidance Under Imperfect Sensing

RA-L 2020

Unlike autonomous ground vehicles (AGVs), unmanned aerial vehicles (UAVs) have a higher dimensional configuration space, which makes the motion planning of multi-UAVs a challenging task. In addition, uncertainties and noises are more significant in UAV scenarios, which increases the difficulty of au

Cited by 106SourceScholar
2020

An Actor-Critic Approach for Legible Robot Motion Planner

ICRA 2020poster

In human-robot collaboration, it is crucial for the robot to make its intentions clear and predictable to the human partners. Inspired by the mutual learning and adaptation of human partners, we suggest an actor-critic approach for a legible robot motion planner. This approach includes two neural ne…

Cited by 24SourceScholar
2020

Augmented Memory for Correlation Filters in Real-Time UAV Tracking

IROS 2020poster

The outstanding computational efficiency of discriminative correlation filter (DCF) fades away with various complicated improvements. Previous appearances are also gradually forgotten due to the exponential decay of historical views in traditional appearance updating scheme of DCF framework, reducin…

Cited by 44SourcecodeScholar
2020

Configuration Space Decomposition for Learning-based Collision Checking in High-DOF Robots

IROS 2020poster

Motion planning for robots of high degrees-of-freedom (DOFs) is an important problem in robotics with sampling-based methods in configuration space \mathcal{C}\mathcal{C} as one popular solution. Recently, machine learning methods have been introduced into sampling-based motion planning methods, whi…

Cited by 8SourceScholar
2020

DeepMNavigate: Deep Reinforced Multi-Robot Navigation Unifying Local & Global Collision Avoidance

IROS 2020poster

We present a novel algorithm (DeepMNavigate) for global multi-agent navigation in dense scenarios using deep reinforcement learning (DRL). Our approach uses local and global information for each robot from motion information maps. We use a three-layer CNN that takes these maps as input to generate a…

Cited by 28SourceScholar
2020

High-Resolution Attention Network with Acoustic Segment Model for Acoustic Scene Classification

ICASSP 2020accepted

The spectral information of acoustic scenes is diverse and complex, which poses challenges for acoustic scene tasks. To improve the classification performance, a variety of convolutional neural networks (CNNs) are proposed to extract richer semantic information of scene utterances. However, the diff…

Cited by 0SourceScholar
2020

Learning Resilient Behaviors for Navigation Under Uncertainty

ICRA 2020poster

Deep reinforcement learning has great potential to acquire complex, adaptive behaviors for autonomous agents automatically. However, the underlying neural network polices have not been widely deployed in real-world applications, especially in these safety-critical tasks (e.g., autonomous driving). O…

Cited by 28SourceScholar
2019

A Two-stage Single-channel Speaker-dependent Speech Separation Approach for Chime-5 Challenge

ICASSP 2019accepted

In this paper, we design a two-stage single-channel speaker-dependent speech separation approach for the CHiME-5 Challenge, targeting the problem of far-field and multi-talker conversational speech recognition in dinner party scenarios involving background noises, reverberations and overlapping spee…

Cited by 0SourceScholar
2019

Compact Reachability Map for Excavator Motion Planning

IROS 2019poster

In this paper, we propose a novel compact reachability map representation for excavator motion planning. The constructed reachability map can concisely encode the bucket’s reachable pose and the translation capability limited by excavator’s kinematic structure. By explicitly exploiting the property…

Cited by 24SourceScholar
2019

Context-Aware Spatio-Recurrent Curvilinear Structure Segmentation

CVPR 2019poster

Curvilinear structures are frequently observed in various images in different forms, such as blood vessels or neuronal boundaries in biomedical images. In this paper, we propose a novel curvilinear structure segmentation approach using context-aware spatio-recurrent networks. Instead of directly seg…

Cited by 26PDFScholar
2019

Getting Robots Unfrozen and Unlost in Dense Pedestrian Crowds

RA-L 2019

Our goal is to navigate a mobile robot to navigate through environments with dense crowds, e.g., shopping malls, canteens, train stations, or airport terminals. In these challenging environments, existing approaches suffer from two common problems: the robot may get frozen and cannot make any progre

Cited by 67SourceScholar
2019

Plant Phenotyping by Deep-Learning-Based Planner for Multi-Robots

RA-L 2019

Manual plant phenotyping is slow, error prone, and labor intensive. In this letter, we present an automated robotic system for fast, precise, and noninvasive measurements using a new deep-learning-based next-best view planning pipeline. Specifically, we first use a deep neural network to estimate a

Cited by 69SourceScholar
2019

Visualizing the Invisible: Occluded Vehicle Segmentation and Recovery

ICCV 2019poster

In this paper, we propose a novel iterative multi-task framework to complete the segmentation mask of an occluded vehicle and recover the appearance of its invisible parts. In particular, firstly, to improve the quality of the segmentation completion, we present two coupled discriminators that intro…

Cited by 45PDFScholar
2018

Intervention Aided Reinforcement Learning for Safe and Practical Policy Optimization in Navigation

CoRL 2018

Combining deep neural networks with reinforcement learning has shown great potential in the next-generation intelligent control. However, there are challenges in terms of safety and cost in practical applications. In this pa- per, we propose the Intervention Aided Reinforcement Learning (IARL) frame

2018

Towards Optimally Decentralized Multi-Robot Collision Avoidance via Deep Reinforcement Learning

ICRA 2018poster

Developing a safe and efficient collision avoidance policy for multiple robots is challenging in the decentralized scenarios where each robot generates its paths without observing other robots' states and intents. While other distributed multi-robot collision avoidance systems exist, they often requ…

Cited by 652SourceScholar
2016

An empirical comparison among the effect of different supports in sequential robotic manipulation

IROS 2016poster

Pick-and-place regrasp extends the manipulation capability of a robot by using a sequence of regrasps to accomplish tasks that are not possible using a single grasp due to constraints such as kinematics or collisions between the robot and the environment. Previous work on pick-and-place only leverag…

Cited by 5SourceScholar
2016

Efficient Penetration Depth Computation Between Rigid Models Using Contact Space Propagation Sampling

RA-L 2016

We present a novel method to compute the approximate global penetration depth (PD) between two nonconvex geometric models. Our approach consists of two phases: offline precomputation and run-time queries. In the first phase, our formulation uses a novel sampling algorithm to precompute an approximat

Cited by 8SourceScholar
2015

Leveraging appearance priors in non-rigid registration, with application to manipulation of deformable objects

IROS 2015poster

Manipulation of deformable objects is a widely applicable but challenging task in robotics. One promising nonparametric approach for this problem is trajectory transfer, in which a non-rigid registration is computed between the starting scene of the demonstration and the scene at test time. This reg…

Cited by 56SourceScholar