← Search

Rui Huang

89 accepted papers

2026

A Multisensory Neurofeedback–Based Immersive BCI Paradigm for Emotion Regulation

ICRA 2026poster

Enhancing brain activation efficiency is crucial in developing brain computer interface (BCI) paradigm for cognitive rehabilitation. However, the existing BCI paradigms mostly achieved limited sensory-activation without sufficient feedback of mind and body, significantly limiting the user engagement…

Cited by 0Scholar
2026

A Novel DNN-Based Semi-Parametric Calibration Method for Parallel Robots Considering Non-Kinematic Parameters

RA-L 2026

The pointing accuracy of pose adjusting parallel robots (PAPRs) is critical for the imaging quality of Cherenkov telescopes, making kinematic calibration crucial for improvement. However, non-geometric error sources like elastic deformation and joint clearances create an inevitable difference betwee

Cited by 0SourceScholar
2026

A Spatiotemporal Brain Activity Visualization and Assessment Framework for Human-Robot Cognitive Interaction Training

ICRA 2026poster

Accurately assessing brain activity to modulate training parameters online is crucial for improving the human-robot cognitive interaction (HRCI) performance in closed-loop brain training. The major challenge for this technique lies in how to accurately model and characterize the intrinsic behavior o…

Cited by 0Scholar
2026

ADGaussian: Generalizable Gaussian Splatting for Autonomous Driving Via Multi-Modal Joint Learning

ICRA 2026poster

We present a novel approach, termed ADGaussian, for generalizable street scene reconstruction. The proposed method enables high-quality rendering from merely single-view input. Unlike prior Gaussian Splatting methods that primarily focus on geometry refinement, we emphasize the importance of joint o…

2026

AERO-MPPI: Anchor-Guided Ensemble Trajectory Optimization for Agile Mapless Drone Navigation

ICRA 2026poster

Agile mapless navigation in cluttered 3D environments poses significant challenges for autonomous drones. Conventional mapping–planning–control pipelines incur high computational cost and propagate estimation errors. We present AERO-MPPI, a fully GPU-accelerated framework that unifies perception and…

2026

Accelerating Diffusion Model Training under Minimal Budgets: A Condensation-Based Perspective

CVPR 2026

Diffusion models have achieved remarkable performance on a wide range of generative tasks, yet training them from scratch is notoriously resource-intensive, typically requiring millions of training images and many GPU days. Motivated by a data-centric view of this bottleneck, we adopt a condensation

Cited by 0SourcecodeScholar
2026

Agile Trajectory Planning and Large Obstacle Avoidance for Formation Flight Using a Virtual Core

ICRA 2026poster

Current methods for formation flight primarily focus on maintaining formations, often neglecting the swarm's agility. Furthermore, most of these approaches fail to leverage global information from the swarm for obstacle avoidance, making them incapable of generating efficient and safe trajectories i…

Cited by 0SourceScholar
2026

Beyond Fixed Formulas: Data-Driven Linear Predictor for Efficient Diffusion Models

CVPR 2026

Diffusion Transformers (DiTs) have achieved state-of-the-art image and video generation performance, but sampling remains expensive due to repeated transformer forward passes over many timesteps. Feature caching offers a training-free way to accelerate inference by reusing or forecasting hidden repr

Cited by 0SourcecodeScholar
2026

CaReBench: A Fine-grained Benchmark for Video Captioning and Retrieval

ICLR 2026poster

Video understanding, including video captioning and retrieval, is still a great challenge for video-language models (VLMs). The existing video retrieval and caption benchmarks only include short descriptions, limits their ability of detailed video understanding evaluation. To address this problem, w…

Cited by 0SourcecodeScholar
2026

SIGN: Safety-Aware Image-Goal Navigation for Autonomous Drones Via Reinforcement Learning

ICRA 2026poster

Image-goal navigation (ImageNav) tasks a robot with autonomously exploring an unknown environment and reaching a location that visually matches a given target image. While prior works primarily study ImageNav for ground robots, enabling this capability for autonomous drones is substantially more cha…

2026

SIGN: Safety-Aware Image-Goal Navigation for Autonomous Drones via Reinforcement Learning

RA-L 2026

Image-goal navigation (ImageNav) tasks a robot with autonomously exploring an unknown environment and reaching a location that visually matches a given target image. While prior works primarily study ImageNav for ground robots, enabling this capability for autonomous drones is substantially more cha

Cited by 1SourcecodeScholar
2026

Sparse Poisson Gamma Belief Networks for High-Dimensional Sparse Count Data

AAAI 2026technical

Bayesian networks play a crucial role in various domains for unsupervised feature extraction and data interpretation. The Poisson gamma belief networks (PGBNs), as a type of Bayesian networks, have shown promise in analyzing high-dimensional count data. However, PGBNs encounter significant challenge

Cited by 0SourcePDFScholar
2026

TransforMARS: Fault-Tolerant Self-Reconfiguration for Arbitrary-Shaped Modular Aerial Robot Systems

ICRA 2026poster

Modular Aerial Robot Systems (MARS) consist of multiple drone modules that are physically bound together to form a single structure for flight. Exploiting structural redundancy, MARS can be reconfigured into different formations to mitigate unit or rotor failures and maintain stable flight. Prior wo…

Cited by 0codeScholar
2025

3D-MOOD: Lifting 2D to 3D for Monocular Open-Set Object Detection

ICCV 2025poster

Monocular 3D object detection is valuable for various applications such as robotics and AR/VR. Existing methods are confined to closed-set settings, where the training and testing sets consist of the same scenes and/or object categories. However, real-world applications often introduce new environme…

2025

Agile Trajectory Planning and Large Obstacle Avoidance for Formation Flight Using a Virtual Core

RA-L 2025

Current methods for formation flight primarily focus on maintaining formations, often neglecting the swarm's agility. Furthermore, most of these approaches fail to leverage global information from the swarm for obstacle avoidance, making them incapable of generating efficient and safe trajectories i

Cited by 0SourceScholar
2025

DEFOM-Stereo: Depth Foundation Model Based Stereo Matching

CVPR 2025poster

Stereo matching is a key technique for metric depth estimation in computer vision and robotics. Real-world challenges like occlusion and non-texture hinder accurate disparity estimation from binocular matching cues. Recently, monocular relative depth estimation has shown remarkable generalization us…

2025

DONIS: Importance Sampling for Training Physics-Informed DeepONet

IJCAI 2025

Deep Operator Network (DeepONet) effectively learns complex operator mappings, especially for systems governed by differential equations. Physics-informed DeepONet (PI-DeepONet) extends these capabilities by integrating physical constraints, enabling robust performance with limited or no labeled dat

2025

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-centric 3D Visual Grounding

ICLR 2025poster

Enabling intelligent agents to comprehend and interact with 3D environments through natural language is crucial for advancing robotics and human-computer interaction. A fundamental task in this field is ego-centric 3D visual grounding, where agents locate target objects in real-world 3D spaces based…

Cited by 0SourcePDFScholar
2025

Engaging Mind and Body: An Immersive BCI Paradigm with Motion-Panoramic Virtual Reality

IROS 2025

Brain-computer interface (BCI) is an important technology in developing the closed-loop brain training system for cognitive functional rehabilitation. Most of existing BCI paradigms have not ensured desired immersiveness of mind and body, thereby limiting participants’ engagement in training tasks.

Cited by 0SourceScholar
2025

MARS-FTCP: Robust Fault-Tolerant Control and Agile Trajectory Planning for Modular Aerial Robot Systems

IROS 2025

Modular Aerial Robot Systems (MARS) consist of multiple drone units that can self-reconfigure to adapt to various mission requirements and fault conditions. However, existing fault-tolerant control methods exhibit significant oscillations during docking and separation, impacting system stability. To

Cited by 4SourcecodeScholar
2025

Plug-and-Play Multi-Domain Fusion Adaptation for Cross-Subject EEG-Based Motor Imagery Classification

ICRA 2025

Motor imagery (MI) classification in rehabilitation brain-computer interfaces (RBCIs) faces significant challenges due to the variability of electroencephalography (EEG) signals across subjects. Existing methods typically require extensive EEG data collection from each new subject, which is time-con

Cited by 1SourceScholar
2025

RecFlow: An Industrial Full Flow Recommendation Dataset

ICLR 2025poster

Industrial recommendation systems (RS) rely on the multi-stage pipeline to balance effectiveness and efficiency when delivering items from a vast corpus to users. Existing RS benchmark datasets primarily focus on the exposure space, where novel RS algorithms are trained and evaluated. However, when…

2025

Robust Self-Reconfiguration for Fault-Tolerant Control of Modular Aerial Robot Systems

ICRA 2025

Modular Aerial Robotic Systems (MARS) consist of multiple drone units assembled into a single, integrated rigid flying platform. With inherent redundancy, MARS can self-reconfigure into different configurations to mitigate rotor or unit failures and maintain stable flight. However, existing works on

Cited by 9SourcecodeScholar
2025

SFFCE-CD: Spatial And Frequency Feature Cross Enhancement For Change Detection

ICASSP 2025accepted

Most existing change detection (CD) methods focus on spatial domain modeling, while ignore the rich information of frequency domain. In this paper, we enhance the feature representative ability with spatial and frequency feature cross enhancement. Specifically, we propose a Change Feature Extract Mo…

Cited by 0SourceScholar
2025

Self-Supervised Enhancement for Depth from a Lightweight ToF Sensor with Monocular Images

IROS 2025

Depth map enhancement using paired high-resolution RGB images offers a cost-effective solution for improving low-resolution depth data from lightweight ToF sensors. Nevertheless, naively adopting a depth estimation pipeline to fuse the two modalities requires groundtruth depth maps for supervision.

Cited by 0SourcecodeScholar
2025

Uncertain Pushing Adaptive Coordinated Control for the Human-Exoskeleton-Walker System

RA-L 2025

Lower Limb Exoskeletons are potential in the gait training for patients with gait disorders. For patients in the early rehabilitation stages with weak upper limb strength, it is challenge to keep balance by themselves only. A mobile robotic walker is helpful to maintain the walking balance, with the

Cited by 0SourceScholar
2025

Video Perception Models for 3D Scene Synthesis

NeurIPS 2025poster

Automating the expert-dependent and labor-intensive task of 3D scene synthesis would significantly benefit fields such as architectural design, robotics simulation, and virtual reality. Recent approaches to 3D scene synthesis often rely on the commonsense reasoning of large language models (LLMs) or…

Cited by 0SourceScholar
2025

Wavelet-Assisted Multi-Frequency Attention Network for Pansharpening

AAAI 2025technical

Pansharpening aims to combine a high-resolution panchromatic (PAN) image with a low-resolution multispectral (LRMS) image to produce a high-resolution multispectral (HRMS) image. Although pansharpening in the frequency domain offers clear advantages, most existing methods either continue to operate…

2024

A Saliency Enhanced Feature Fusion Based Multiscale RGB-D Salient Object Detection Network

ICASSP 2024accepted

Multiscale convolutional neural network (CNN) has demonstrated remarkable capabilities in solving various vision problems. However, fusing features of different scales always results in large model sizes, impeding the application of multiscale CNNs in RGB-D saliency detection. In this paper, we prop…

Cited by 0SourceScholar
2024

A Sim-to-Real Instance Segmentation Framework for Densely Stacked Cartons

RA-L 2024

Robotic picking systems in automated logistics require accurate segmentation and localization of densely stacked cartons. However, the lack of comprehensive and diverse datasets for this task poses a significant challenge. Furthermore, existing instance segmentation methods struggle to meet the accu

Cited by 1SourceScholar
2024

Cross Branch Feature Fusion Decoder for Consistency Regularization-Based Semi-Supervised Change Detection

ICASSP 2024accepted

Semi-supervised change detection (SSCD) utilizes partially labeled data and a large amount of unlabeled data to detect changes. However, the transformer-based SSCD network does not perform as well as the convolution-based SSCD network due to the lack of labeled data. To overcome this limitation, we…

Cited by 0SourceScholar
2024

Incremental 3D Reconstruction through a Hybrid Explicit-and-Implicit Representation

ICRA 2024poster

3D reconstruction is an important task in computer vision and is widely used in robotics and autonomous driving. When building large-scale scenes, limitations in computing resources and the difficulty of accessing the entire dataset in a single task are inevitable. Therefore, an incremental reconstr…

Cited by 0SourceScholar
2024

Joint-Loss Enhanced Self-Supervised Learning for Refinement-Coupled Object 6D Pose Estimation

ICRA 2024poster

6D object pose estimation plays a crucial role in robot grasping and manipulation. However, the prevalent methods for 6D object pose estimation heavily rely on 6D annotated data to train deep neural networks, which poses challenges due to the difficulty in obtaining sufficient pose annotations. To a…

Cited by 0SourceScholar
2024

Mol-Instructions: A Large-Scale Biomolecular Instruction Dataset for Large Language Models

ICLR 2024poster

Large Language Models (LLMs), with their remarkable task-handling capabilities and innovative outputs, have catalyzed significant advancements across a spectrum of fields. However, their proficiency within specialized domains such as biomolecular studies remains limited. To address this challenge, w…

2024

MorAL: Learning Morphologically Adaptive Locomotion Controller for Quadrupedal Robots on Challenging Terrains

RA-L 2024

Due to the rapid development of the quadruped robot industry in the past decade, various commercial quadruped robots have emerged with distinct physical attributes. Different from the previous work in which the designed controller is robot-specific, this article proposes a learning-based control fra

Cited by 37SourceScholar
2024

Negative-Binomial Randomized Gamma Dynamical Systems for Heterogeneous Overdispersed Count Time Sequences

IJCAI 2024poster

Modeling count-valued time sequences has been receiving growing interests because count time sequences naturally arise in physical and social domains. Poisson gamma dynamical systems (PGDSs) are newly-developed methods, which can well capture the expressive latent transition structure and bursty dyn…

Cited by 0SourcePDFScholar
2024

Predicting Bird's-Eye-View Semantic Representations Using Correlated Context Learning

RA-L 2024

We redefine the concept of bird's-eye-view (BEV) imaging for machine cognition tasks, emphasizing its power as an image interpretation tool. Humans intuitively translate two-dimensional (2D) images into BEV representations by discerning and integrating spatial information, such as position and morph

Cited by 6SourceScholar
2024

SPY-Watermark: Robust Invisible Watermarking for Backdoor Attack

ICASSP 2024accepted

Backdoor attack aims to deceive a victim model when facing backdoor instances while maintaining its performance on benign data. Current methods use manual patterns or special perturbations as triggers, while they often overlook the robustness against data corruption, making backdoor attacks easy to…

Cited by 0SourceScholar
2024

Segment3D: Learning Fine-Grained Class-Agnostic 3D Segmentation without Manual Labels

ECCV 2024poster

"Current 3D scene segmentation methods are heavily dependent on manually annotated 3D training datasets. Such manual annotations are labor-intensive, and often lack fine-grained details. Furthermore, models trained on this data typically struggle to recognize object classes beyond the annotated trai…

Cited by 34SourcePDFScholar
2024

SmartEdit: Exploring Complex Instruction-based Image Editing with Multimodal Large Language Models

CVPR 2024highlight

Current instruction-based image editing methods such as InstructPix2Pix often fail to produce satisfactory results in complex scenarios due to their dependence on the simple CLIP text encoder in diffusion models. To rectify this this paper introduces SmartEdit a novel approach of instruction-based i…

2024

Towards Cross-View-Consistent Self-Supervised Surround Depth Estimation

IROS 2024poster

Depth estimation is a cornerstone for autonomous driving, yet acquiring per-pixel depth ground truth for supervised learning is challenging. Self-Supervised Surround Depth Estimation (SSSDE) from consecutive images offers an economical alternative. While previous SSSDE methods have proposed differen…

Cited by 0SourcecodeScholar
2024

Training an Open-Vocabulary Monocular 3D Detection Model without 3D Data

NeurIPS 2024poster

Open-vocabulary 3D object detection has recently attracted considerable attention due to its broad applications in autonomous driving and robotics, which aims to effectively recognize novel classes in previously unseen domains. However, existing point cloud-based open-vocabulary 3D detection models…

Cited by 4SourcePDFScholar
2023

Background-Mixed Augmentation for Weakly Supervised Change Detection

AAAI 2023technical

Change detection (CD) is to decouple object changes (i.e., object missing or appearing) from background changes (i.e., environment variations) like light and season variations in two images captured in the same scene over a long time span, presenting critical applications in disaster management, urb…

2023

Learning To Generate Language-Supervised and Open-Vocabulary Scene Graph Using Pre-Trained Visual-Semantic Space

CVPR 2023poster

Scene graph generation (SGG) aims to abstract an image into a graph structure, by representing objects as graph nodes and their relations as labeled edges. However, two knotty obstacles limit the practicability of current SGG methods in real-world scenarios: 1) training SGG models requires time-cons…

2023

ScaleMix: Intra- And Inter-Layer Multiscale Feature Combination for Change Detection

ICASSP 2023accepted

Change detection (CD) aims at finding change objects from bi-temporal images, which has wide applications in different vision tasks. Previous CD methods focus more on fusing inter-layer multiscale features while ignoring the intra-layer multiscale characteristics, which hurts the integrity of change…

Cited by 0SourceScholar
2023

Weak6D: Weakly Supervised 6D Pose Estimation With Iterative Annotation Resolver

RA-L 2023

6D object pose estimation is an essential task in vision-based robotic grasping and manipulation. Prior works always train models with a large number of pose annotated images, limiting the efficiency of model transfer between different scenarios. This letter presents an end-to-end model named <itali

Cited by 8SourceScholar
2022

A Novel Multimodal Human-Exoskeleton Interface Based on EEG and sEMG Activity for Rehabilitation Training

ICRA 2022poster

Despite the advances in the field of human-robot interface (HRI) based on biological neural signal, the use of the sole electroencephalography (EEG) signal to help robotic exoskeleton predict the limb movement is currently no mature in rehabilitation training, due to its unreliability. Multimodal HR…

Cited by 8SourceScholar
2022

Attention-Based Deep Driving Model for Autonomous Vehicles with Surround-View Cameras

IROS 2022poster

Experienced human drivers always make safe driving decisions by selectively observing the front, rear and side- view mirrors. Several end - to-end methods have been pro-posed to learn driving models with multi-view visual infor-mation. However, these benchmark methods lack semantic understanding of…

Cited by 0SourceScholar
2022

Efficient Knowledge Distillation from Model Checkpoints

NeurIPS 2022accept

Knowledge distillation is an effective approach to learn compact models (students) with the supervision of large and strong models (teachers). As empirically there exists a strong correlation between the performance of teacher and student models, it is commonly believed that a high performing teache…

2022

Exploring Structure-Aware Transformer Over Interaction Proposals for Human-Object Interaction Detection

CVPR 2022poster

Recent high-performing Human-Object Interaction (HOI) detection techniques have been highly influenced by Transformer-based object detector (i.e., DETR). Nevertheless, most of them directly map parametric interaction queries into a set of HOI predictions through vanilla Transformer in a one-stage ma…

Cited by 93PDFcodeScholar
2022

Fully Attentional Network for Semantic Segmentation

AAAI 2022technical

Recent non-local self-attention methods have proven to be effective in capturing long-range dependencies for semantic segmentation. These methods usually form a similarity map of R^(CxC) (by compressing spatial dimensions) or R^(HWxHW) (by compressing channels) to describe the feature relations alon…

2022

Human-exoskeleton Cooperative Balance Strategy for a Human-powered Augmentation Lower Exoskeleton

IROS 2022poster

Lower Limb Exoskeletons (LLE) have received considerable interest in strength augmentation, rehabilitation, and walking assistance scenarios. For strength augmentation, LLE is expected to have the capability of reducing metabolic energy. However, the energy for adjusting Center of Gravity (CoG) is a…

Cited by 2SourceScholar
2022

Salient-to-Broad Transition for Video Person Re-Identification

CVPR 2022poster

Due to the limited utilization of temporal relations in video re-id, the frame-level attention regions of mainstream methods are partial and highly similar. To address this problem, we propose a Salient-to-Broad Module (SBM) to enlarge the attention regions gradually. Specifically, in SBM, while the…

Cited by 69PDFcodeScholar
2021

BiCnet-TKS: Learning Efficient Spatial-Temporal Representation for Video Person Re-Identification

CVPR 2021poster

In this paper, we present an efficient spatial-temporal representation for video person re-identification (reID). Firstly, we propose a Bilateral Complementary Network (BiCnet) for spatial complementarity modeling. Specifically, BiCnet contains two branches. Detail Branch processes frames at origina…

Cited by 126PDFcodeScholar
2021

Coupled Segmentation and Edge Learning via Dynamic Graph Propagation

NeurIPS 2021poster

Image segmentation and edge detection are both central problems in perceptual grouping. It is therefore interesting to study how these two tasks can be coupled to benefit each other. Indeed, segmentation can be easily transformed into contour edges to guide edge learning. However, the converse is no…

Cited by 14SourcePDFScholar
2021

Estimating the Center of Mass of Human-Exoskeleton Systems with Physically Coupled Serial Chain

IROS 2021poster

Estimating the center of mass (CoM) is essential for both gait planning and controlling of lower limb exoskeletons. Different from CoM estimation in human and humanoid robots, a critical issue in human-exoskeleton systems pis how to describe the effect of physical human-exoskeleton interactions in e…

Cited by 1SourceScholar
2021

IMENet: Joint 3D Semantic Scene Completion and 2D Semantic Segmentation through Iterative Mutual Enhancement

IJCAI 2021poster

3D semantic scene completion and 2D semantic segmentation are two tightly correlated tasks that are both essential for indoor scene understanding, because they predict the same semantic classes, using positively correlated high-level features. Current methods use 2D features extracted from early-fus…

Cited by 16SourcePDFScholar
2021

Kld Minimization-Based Constrained Measurement Filtering For Two-Step TDOA Indoor Tracking

ICASSP 2021accepted

This paper presents an enhanced two-step method for tracking an indoor point target using the time difference of arrival (TDOA) measurements from an ultra wideband (UWB) positioning system. Again, the algorithm preprocesses the raw TDOAs and then feeds the results to a recursively bounded grid-based…

Cited by 0SourceScholar
2021

Not All Images are Worth 16x16 Words: Dynamic Transformers for Efficient Image Recognition

NeurIPS 2021poster

Vision Transformers (ViT) have achieved remarkable success in large-scale image recognition. They split every 2D image into a fixed number of patches, each of which is treated as a token. Generally, representing an image with more tokens would lead to higher prediction accuracy, while it also result…

2021

On the Importance of Gradients for Detecting Distributional Shifts in the Wild

NeurIPS 2021poster

Detecting out-of-distribution (OOD) data has become a critical component in ensuring the safe deployment of machine learning models in the real world. Existing OOD detection approaches primarily rely on the output or feature space for deriving OOD scores, while largely overlooking information from t…

2021

Sparse Single Sweep LiDAR Point Cloud Segmentation via Learning Contextual Shape Priors from Scene Completion

AAAI 2021technical

LiDAR point cloud analysis is a core task for 3D computer vision, especially for autonomous driving. However, due to the severe sparsity and noise interference in the single sweep LiDAR point cloud, the accurate semantic segmentation is non-trivial to achieve. In this paper, we propose a novel spars…

2021

Synergetic Gait Prediction for Stroke Rehabilitation with Varying Walking Speeds

IROS 2021poster

Lower Limb Exoskeletons (LLEs) are promising in gait rehabilitation for stroke survivors. In gait training of post-stroke patients with LLEs, one of the main challenges is how to generate appropriate gait patterns from the sound leg to the paretic leg for different patients with varying walking spee…

Cited by 8SourceScholar
2021

TemporalFusion: Temporal Motion Reasoning with Multi-Frame Fusion for 6D Object Pose Estimation

IROS 2021poster

6D object pose estimation is an essential task in vision-based robotic grasping and manipulation. Prior works extract spatial features by fusing the RGB image and depth without considering the temporal motion information, limiting their performance in heavy occlusion robotic grasping scenarios. In t…

Cited by 4SourcecodeScholar
2021

Towards Fully Autonomous Ultrasound Scanning Robot With Imitation Learning Based on Clinical Protocols

RA-L 2021

Ultrasound scanning plays an important role in modern clinical examinations. Thanks to its small footprint, low cost, and popularity, it has been widely used in annual physical examinations and many other diagnosis and intervention procedures. However, the scanning results depend heavily on the clin

Cited by 64SourceScholar
2021

Uncertainty-Guided Transformer Reasoning for Camouflaged Object Detection

ICCV 2021poster

Spotting objects that are visually adapted to their surroundings is challenging for both humans and AI. Conventional generic / salient object detection techniques are suboptimal for this task because they tend to only discover easy and clear objects, while overlooking the difficult-to-detect ones wi…

Cited by 296PDFcodeScholar
2020

An LSTM Approach to Temporal 3D Object Detection in LiDAR Point Clouds

ECCV 2020poster

Detecting objects in 3D LiDAR data is a core technology for autonomous driving and other robotics applications. Although LiDAR data is acquired over time, most of the 3D object detection algorithms propose object bounding boxes independently for each frame and neglect the useful information availabl…

Cited by 136SourcePDFScholar
2020

Data-Driven Reinforcement Learning for Walking Assistance Control of a Lower Limb Exoskeleton with Hemiplegic Patients

ICRA 2020poster

Lower limb exoskeleton (LLE) has received considerable interests in strength augmentation, rehabilitation and walking assistance scenarios. For walking assistance, the LLE is expected to have the capability of controlling the affected leg to track the unaffected leg’s motion naturally. An important…

Cited by 35SourceScholar
2020

DiPE: Deeper into Photometric Errors for Unsupervised Learning of Depth and Ego-motion from Monocular Videos

IROS 2020poster

Unsupervised learning of depth and ego-motion from unlabelled monocular videos has recently drawn great attention, which avoids the use of expensive ground truth in the supervised one. It achieves this by using the photometric errors between the target view and the synthesized views from its adjacen…

Cited by 24SourcecodeScholar
2020

Glance and Focus: a Dynamic Approach to Reducing Spatial Redundancy in Image Classification

NeurIPS 2020poster

The accuracy of deep convolutional neural networks (CNNs) generally improves when fueled with high resolution images. However, this often comes at a high computational cost and high memory footprint. Inspired by the fact that not all regions in an image are task-relevant, we propose a novel framewor…

2020

Learning Memory Augmented Cascading Network for Compressed Sensing of Images

ECCV 2020poster

In this paper, we propose a cascading network for compressed sensing of images with progressive reconstruction. Specifically, we decompose the complex reconstruction mapping into the cascade of incremental detail reconstruction (IDR) modules and measurement residual updating (MRU) modules. The IDR m…

2020

Robust Tdoa Indoor Tracking Using Constrained Measurement Filtering and Grid-Based Filtering

ICASSP 2020accepted

This paper considers exploiting the time difference of arrival (TDOA) measurements from a ultra wideband (UWB) indoor positioning system to locate a moving point target. In indoor environments, measured TDOAs are subject to large errors due to multipath and/or non-line-of-sight (NLOS) propagation. B…

Cited by 0SourceScholar
2019

Adaptive Gait Planning for Walking Assistance Lower Limb Exoskeletons in Slope Scenarios

ICRA 2019poster

Lower-limb exoskeleton has gained considerable interests in walking assistance applications for paraplegic patients. In walking assistance of paraplegic patients, the exoskeleton should have the ability to help patients to walk over different terrains in the daily life, such as slope terrains. One c…

Cited by 11SourceScholar
2019

End-to-End Driving Model for Steering Control of Autonomous Vehicles with Future Spatiotemporal Features

IROS 2019poster

End-to-end deep learning has gained considerable interests in autonomous driving vehicles in both academic and industrial fields, especially in decision making process. One critical issue in decision making process of autonomous driving vehicles is steering control. Researchers has already trained d…

Cited by 43SourceScholar
2018

Learning-based Walking Assistance Control Strategy for a Lower Limb Exoskeleton with Hemiplegia Patients

IROS 2018poster

Lower exoskeleton has gained considerable interests in walking assistance applications for both paraplegia and hemiplegia patients. In walking assistance of hemiplegia patients, the exoskeleton should have the ability to control the affected leg to follow the unaffected leg's motion naturally. One c…

Cited by 27SourceScholar
2017

Beyond Face Rotation: Global and Local Perception GAN for Photorealistic and Identity Preserving Frontal View Synthesis

ICCV 2017poster

Photorealistic frontal view synthesis from a single face image has a wide range of applications in the field of face recognition. Although data-driven deep learning methods have been proposed to address this problem by seeking solutions from ample face data, this problem is still challenging because…

Cited by 849PDFScholar
2017

Learning Dynamic Siamese Network for Visual Object Tracking

ICCV 2017poster

How to effectively learn temporal variation of target appearance, to exclude the interference of cluttered background, while maintaining real-time response, is an essential problem of visual object tracking. Recently, Siamese networks have shown great potentials of matching based trackers in achievi…

Cited by 1046PDFScholar
2016

Hierarchical Interactive Learning for a HUman-Powered Augmentation Lower EXoskeleton

ICRA 2016poster

Learning by demonstration methods have gained considerable interest in human-coupled robot control. It aims at modeling the goal motion trajectories through human demonstration. However, in lower exoskeleton control, the physical human-robot interaction is changing from pilot to pilot or even for on…

Cited by 70SourceScholar
2016

Learning Cooperative Primitives with physical Human-Robot Interaction for a HUman-powered Lower EXoskeleton

IROS 2016poster

Human-powered lower exoskeletons have gained considerable interests from both academia and industry over the past few decades, and thus have seen increasing applications in areas of human locomotion assistance and strength augmentation. One of the most important aspects in those applications is to a…

Cited by 20SourceScholar
2015

A Spatio-Temporal Appearance Representation for Viceo-Based Pedestrian Re-Identification

ICCV 2015poster

Pedestrian re-identification is a difficult problem due to the large variations in a person's appearance caused by different poses and viewpoints, illumination changes, and occlusions. Spatial alignment is commonly used to address these issues by treating the appearance of different body parts indep…

Cited by 287PDFScholar
2015

Interactive learning for sensitivity factors of a human-powered augmentation lower exoskeleton

IROS 2015poster

Sensitivity Amplification Control (SAC) algorithm was first proposed in the augmentation applications of Berkeley Lower Extremity Exoskeleton (BLEEX). The SAC algorithm is widely used in human augmentation applications since it just need the information from the exoskeleton robot, so that the comple…

Cited by 50SourceScholar