← Search

Wei Gao

99 accepted papers

2026

$AutoDrive\text{-}P^3$: Unified Chain of Perception–Prediction–Planning Thought via Reinforcement Fine-Tuning

ICLR 2026poster

Vision-language models (VLMs) are increasingly being adopted for end-to-end autonomous driving systems due to their exceptional performance in handling long-tail scenarios. However, current VLM-based approaches suffer from two major limitations: 1) Some VLMs directly output planning results without…

Cited by 0SourceScholar
2026

Dual-Path Condition Alignment for Diffusion Transformers

ICLR 2026poster

Denoising-based generative models have been significantly advanced by representation-alignment (REPA) loss, which leverages pre-trained visual encoders to guide intermediate network features. However, REPA's reliance on external visual encoders introduces two critical challenges: potential \textit{d…

Cited by 0SourcecodeScholar
2026

InfiniBench: Infinite Benchmarking for Visual Spatial Reasoning with Customizable Scene Complexity

CVPR 2026

Modern vision-language models (VLMs) are expected to have abilities of spatial reasoning with diverse scene complexities, but evaluating such abilities is difficult due to the lack of benchmarks that are not only diverse and scalable but also fully customizable. Existing benchmarks offer limited cus

Cited by 0SourcecodeScholar
2026

LangEditor: Natural Language-Driven 4D Editing for Improved Controllability of Dynamic Driving Scenes

ICRA 2026poster

Diverse and realistic data are essential for developing reliable autonomous driving (AD) systems, yet collecting and annotating large-scale real-world driving datasets is costly and time-consuming. Recent advances in synthetic scene generation and editing have enabled the creation of diverse driving…

Cited by 0Scholar
2026

MMBERT: Scaled Mixture-of-Experts Multimodal BERT for Robust Chinese Hate Speech Detection Under Cloaking Perturbations

AAAI 2026technical

Hate speech detection on Chinese social networks presents distinct challenges, particularly due to the widespread use of cloaking techniques designed to evade conventional text-based detection systems. Although large language models (LLMs) have recently improved hate speech detection capabilities, t

Cited by 0SourcePDFScholar
2026

Probabilistic Concept Graph Reasoning for Multimodal Misinformation Detection

CVPR 2026

Multimodal misinformation poses an escalating challenge that often evades traditional detectors, which are opaque black boxes and fragile against new manipulation tactics. We present Probabilistic Concept Graph Reasoning (PCGR), an interpretable, modular, and evolvable framework that reframes multim

Cited by 0SourcecodeScholar
2026

SRIF: A Safer Ranking Inference Framework for Diffusion Policy Models Without Retraining

RA-L 2026

Recently, diffusion policy models have been applied in the field of robotics. Most existing methods use all observations as condition inputs to the diffusion model. With the denoising of the diffusion model, Gaussian noise gradually becomes an action sequence. However, since these constraints are im

Cited by 0SourceScholar
2025

AdaDPCC: Adaptive Rate Control and Rate-Distortion-Complexity Optimization for Dynamic Point Cloud Compression

AAAI 2025technical

Dynamic point cloud compression (DPCC) is crucial in applications like autonomous driving and AR/VR. Current compression methods face challenges with complexity management and rate control. This paper introduces a novel dynamic coding framework that supports variable bitrate and computational comple…

Cited by 0SourcePDFScholar
2025

Causal Enhanced Autoregressive Model for Monocular Image-Goal Navigation in Unknown Map Environment

RA-L 2025

Monocular image-goal navigation in an outdoor environment is a challenging task. Robots have to face monocular scale uncertainty and complex environments. Recently, implementations based on imitation learning have made significant progress. However, robots tend to focus too much on the current state

Cited by 0SourceScholar
2025

CausalAbstain: Enhancing Multilingual LLMs with Causal Reasoning for Trustworthy Abstention

ACL 2025finding

Large Language Models (LLMs) often exhibit knowledge disparities across languages. Encouraging LLMs to abstain when faced with knowledge gaps is a promising strategy to reduce hallucinations in multilingual settings. Current abstention strategies for multilingual scenarios primarily rely on generati…

2025

DGTalker: Disentangled Generative Latent Space Learning for Audio-Driven Gaussian Talking Heads

ICCV 2025poster

In this work, we investigate the generation of high-fidelity, audio-driven 3D Gaussian talking heads from monocular videos. We present DGTalker, an innovative framework designed for real-time, high-fidelity, and 3D-aware talking head synthesis. By leveraging Gaussian generative priors and treating t…

Cited by 0SourcePDFScholar
2025

Dual-Process Watermarked Diffusion: Integrating Watermarking With Denoising in Point Clouds

ICASSP 2025accepted

The integration of depth sensing and laser scanning technologies has propelled point cloud data to the forefront of 3D graphical modeling. This paper addresses a critical gap in the literature: the protection of intellectual property in generating point clouds using Diffusion Models (DMs). We introd…

Cited by 0SourceScholar
2025

Flow4Agent: Long-form Video Understanding via Motion Prior from Optical Flow

ICCV 2025poster

Long-form video understanding has always been a challenging problem due to the significant redundancy in both temporal and spatial contents. This challenge is further exacerbated by the limited context length of Multimodal Large Language Models (MLLMs). To address this issue, many previous works hav…

Cited by 0SourcePDFScholar
2025

High-Precision Transformer-Based Visual Servoing for Humanoid Robots in Aligning Tiny Objects

IROS 2025

High-precision tiny object alignment remains a common and critical challenge for humanoid robots in real world. To address this problem, this paper proposes a vision-based framework for precisely estimating and controlling the relative position between a handheld tool and a target object for humanoi

Cited by 1SourceScholar
2025

Implanting Robust Watermarks in Latent Diffusion Models for Video Generation

ICASSP 2025accepted

In the dynamic realm of digital media, latent diffusion models (LDM) have revolutionized the generation of videos, surpassing the capabilities of traditional generative models. This paper presents Stable Video Signature, a pioneering watermarking framework for LDM in video generation. Addressing the…

Cited by 0SourceScholar
2025

Lightweight Self-Supervised Monocular Depth Estimation for All-Day Scenes Using Generative Adversarial Network

ICASSP 2025accepted

Self-supervised monocular depth estimation (MDE) has achieved performance levels comparable to supervised methods in well-lit environments. However, current methods struggle particularly with challenging nighttime scenes. Existing all-day self-supervised MDE methods often rely on specialized nightti…

Cited by 0SourceScholar
2025

One-Pass Feature Evolvable Learning with Theoretical Guarantees

ICML 2025poster

Feature evolvable learning studies the scenario where old features will vanish and new features will emerge when learning with data streams, and various methods have been developed by utilizing some useful relationships from old features to new features, rather than re-training from scratch. In this…

Cited by 0SourcePDFScholar
2025

PhyT2V: LLM-Guided Iterative Self-Refinement for Physics-Grounded Text-to-Video Generation

CVPR 2025poster

Text-to-video (T2V) generation has been recently enabled by transformer-based diffusion models, but current T2V models lack capabilities in adhering to the real-world common knowledge and physical rules, due to their limited understanding of physical realism and deficiency in temporal modeling. Exis…

2025

Point Cloud Semantic Segmentation with Sparse and Inhomogeneous Annotations

AAAI 2025technical

Utilizing uniformly distributed sparse annotations, weakly supervised learning alleviates the heavy reliance on fine-grained annotations in point cloud semantic segmentation tasks. However, few works discuss the inhomogeneity of sparse annotations, albeit it is common in real-world scenarios. Theref…

2025

ProGait: A Multi-Purpose Video Dataset and Benchmark for Transfemoral Prosthesis Users

ICCV 2025poster

Prosthetic legs play a pivotal role in clinical rehabilitation, allowing individuals with lower-limb amputations the ability to regain mobility and improve their quality of life. Gait analysis is fundamental for optimizing prosthesis design and alignment, directly impacting the mobility and life qua…

2025

ProMind-LLM: Proactive Mental Health Care via Causal Reasoning with Sensor Data

ACL 2025finding

Mental health risk is a critical global public health challenge, necessitating innovative and reliable assessment methods. With the development of large language models (LLMs), they stand out to be a promising tool for explainable mental health care applications. Nevertheless, existing approaches pr…

Cited by 0SourcePDFScholar
2025

Sparse-to-Dense: A Free Lunch for Lossless Acceleration of Video Understanding in LLMs

ACL 2025short

Due to the auto-regressive nature of current video large language models (Video-LLMs), the inference latency increases as the input sequence length grows, posing challenges for the efficient processing of video sequences that are usually very long. We observe that during decoding, the attention scor…

Cited by 0SourcePDFScholar
2025

Stochasticity-aware No-Reference Point Cloud Quality Assessment

IJCAI 2025

The evolution of point cloud processing algorithms necessitates an accurate assessment for their quality. Previous works consistently regard point cloud quality assessment (PCQA) as a MOS regression problem and devise a deterministic mapping, ignoring the stochasticity in generating MOS from subject

Cited by 0SourcePDFScholar
2025

Tackling Intertwined Data and Device Heterogeneities in Federated Learning with Unlimited Staleness

AAAI 2025technical

Federated Learning (FL) can be affected by data and device heterogeneities, caused by clients' different local data distributions and latencies in uploading model updates (i.e., staleness). Traditional schemes consider these heterogeneities as two separate and independent aspects, but this assumptio…

2025

Text-Driven Fashion Image Editing with Compositional Concept Learning and Counterfactual Abduction

CVPR 2025poster

Fashion image editing is a valuable tool for designers to convey their creative ideas by visualizing design concepts. With the recent advances in text editing methods, significant progress has been made in fashion image editing. However, they face two key challenges: spurious correlations in trainin…

Cited by 0SourcePDFScholar
2025

UniPCGC: Towards Practical Point Cloud Geometry Compression via an Efficient Unified Approach

AAAI 2025technical

Learning-based point cloud compression methods have made significant progress in terms of performance. However, these methods still encounter challenges including high complexity, limited compression modes, and a lack of support for variable rate, which restrict the practical application of these me…

2025

VE-Bench: Subjective-Aligned Benchmark Suite for Text-Driven Video Editing Quality Assessment

AAAI 2025technical

Text-driven video editing has recently experienced rapid development. Despite this, evaluating edited videos remains a considerable challenge. Current metrics tend to fail to align with human perceptions, and effective quantitative metrics for video editing are still notably absent. To address this,…

2025

Zero-shot Quantization for Large-kernels via Shape-based Distribution and Diversity Self-distillation

ICASSP 2025accepted

Zero-shot quantization (ZSQ) has emerged as an effective method to reduce model complexity and memory footprint without using original training data, thereby mitigating data privacy and security concerns during model deployment. Recently, Large-Kernel Convolutional Neural Networks (LKCNNs) have achi…

Cited by 0SourceScholar
2024

Active Loop Closure for OSM-guided Robotic Mapping in Large-Scale Urban Environments

IROS 2024poster

The autonomous mapping of large-scale urban scenes presents significant challenges for autonomous robots. To mitigate the challenges, global planning, such as utilizing prior GPS trajectories from OpenStreetMap (OSM), is often used to guide the autonomous navigation of robots for mapping. However, d…

Cited by 2SourceScholar
2024

Chain of Preference Optimization: Improving Chain-of-Thought Reasoning in LLMs

NeurIPS 2024poster

The recent development of chain-of-thought (CoT) decoding has enabled large language models (LLMs) to generate explicit logical reasoning paths for complex problem-solving. However, research indicates that these paths are not always deliberate and optimal. The tree-of-thought (ToT) method employs tr…

2024

Distribution Guidance Network for Weakly Supervised Point Cloud Semantic Segmentation

NeurIPS 2024poster

Despite alleviating the dependence on dense annotations inherent to fully supervised methods, weakly supervised point cloud semantic segmentation suffers from inadequate supervision signals. In response to this challenge, we introduce a novel perspective that imparts auxiliary constraints by regulat…

Cited by 2SourcePDFScholar
2024

Efficient Point Cloud Attribute Compression Framework using Attribute-Guided Graph Fourier Transform

ICASSP 2024accepted

The Graph Fourier Transform (GFT) has achieved remarkable success in point cloud attribute compression due to its adaptability in handling irregular signals. However, the conventional graph-based attribute compression method mostly relies on geometry information to construct the Laplace matrix. In t…

Cited by 0SourceScholar
2024

Fast Inter-frame Motion Prediction for Compressed Dynamic Point Cloud Attribute Enhancement

AAAI 2024technical

Recent years have witnessed the success of deep learning methods in quality enhancement of compressed point cloud. However, existing methods focus on geometry and attribute enhancement of single-frame point cloud. This paper proposes a novel compressed quality enhancement method for dynamic point cl…

Cited by 14SourcePDFScholar
2024

Hierarchical Trajectory Deformation Algorithm With Hybrid Controller for Active Lower Limb Rehabilitation

RA-L 2024

Robot-aided active rehabilitation has shown to be an effective treatment approach for hemiplegic patients. This paper presents an active control framework for lower limb rehabilitation, combining an interaction layer with a hierarchical trajectory deformation algorithm (HTDA), and an assist-as-neede

Cited by 5SourceScholar
2024

Less Is More: Label Recommendation for Weakly Supervised Point Cloud Semantic Segmentation

AAAI 2024technical

Weak supervision has proven to be an effective strategy for reducing the burden of annotating semantic segmentation tasks in 3D space. However, unconstrained or heuristic weakly supervised annotation forms may lead to suboptimal label efficiency. To address this issue, we propose a novel label recom…

Cited by 20SourcePDFScholar
2024

Lightweight Structured Line Map Based Visual Localization

RA-L 2024

Visual localization, also known as camera pose estimation, is a crucial component of many applications, such as robotics, autonomous driving, and augmented reality. Traditional visual localization algorithms typically run on point cloud maps generated by algorithms such as Structure-from-Motion (SfM

Cited by 10SourcecodeScholar
2024

MA-Stereo: Real-Time Stereo Matching via Multi-Scale Attention Fusion and Spatial Error-Aware Refinement

RA-L 2024

Stereo matching is a fundamental task in computer vision. Real-time stereo matching has recently shown great potential in robotics and autonomous driving applications. However, the existing cost aggregation in real-time stereo matching suffers from accuracy limitations in ill-posed regions. Furtherm

Cited by 4SourceScholar
2024

Multimodal Misinformation Detection by Learning from Synthetic Data with Multimodal LLMs

EMNLP 2024finding

Detecting multimodal misinformation, especially in the form of image-text pairs, is crucial. Obtaining large-scale, high-quality real-world fact-checking datasets for training detectors is costly, leading researchers to use synthetic datasets generated by AI technologies. However, the generalizabili…

2024

Reinforcement Retrieval Leveraging Fine-grained Feedback for Fact Checking News Claims with Black-Box LLM

COLING 2024main

Retrieval-augmented language models have exhibited promising performance across various areas of natural language processing (NLP), including fact-critical tasks. However, due to the black-box nature of advanced large language models (LLMs) and the non-retrieval-oriented supervision signal of specif…

2024

Reinforcement Tuning for Detecting Stances and Debunking Rumors Jointly with Large Language Models

ACL 2024findings

Learning multi-task models for jointly detecting stance and verifying rumors poses challenges due to the need for training data of stance at post level and rumor veracity at claim level, which are difficult to obtain. To address this issue, we leverage large language models (LLMs) as the foundation…

2024

Spined Torso Renders Advanced Mobility for Quadrupedal Locomotion

ICRA 2024poster

Animals possessing spinal columns often exhibit exceptional agility for highly dynamic locomotion. The spine grants the trunk with increased degrees of freedom, thereby endowing diverse postures. This paper presents the development of a robot STRAY for quadrupedal locomotion, featuring a four-degree…

Cited by 2SourceScholar
2024

StreamFlow: Streamlined Multi-Frame Optical Flow Estimation for Video Sequences

NeurIPS 2024poster

Prior multi-frame optical flow methods typically estimate flow repeatedly in a pair-wise manner, leading to significant computational redundancy. To mitigate this, we implement a Streamlined In-batch Multi-frame (SIM) pipeline, specifically tailored to video inputs to minimize redundant calculations…

2024

Torque Ripple Reduction in Quasi-Direct Drive Motors Through Angle-Based Repetitive Learning Observer and Model Predictive Torque Controller

IROS 2024poster

Torque ripple reduction in quasi-direct drive (QDD) motors is crucial in their robotic applications for dynamic locomotion and dexterous manipulation. In this paper, we present a novel approach for reducing torque ripples of QDD motors, which integrates an angle-based repetitive learning observer (A…

Cited by 0SourceScholar
2024

Towards Green AI in Fine-tuning Large Language Models via Adaptive Backpropagation

ICLR 2024poster

Fine-tuning is essential to adapting pre-trained large language models to downstream applications. With the increasing popularity of LLM-enabled applications, fine-tuning has been performed intensively worldwide, incurring a tremendous amount of computing costs that correspond to big carbon footprin…

2023

Edge Devices Friendly Self-Supervised Monocular Depth Estimation via Knowledge Distillation

RA-L 2023

Self-supervised monocular depth estimation (MDE) has great potential for deployment in a wide range of applications, including virtual reality, autonomous driving, and robotics. Nevertheless, most previous studies focused on complex architectures to pursue better performance in MDE. In this letter,

Cited by 11SourceScholar
2023

Improving Graph Representation for Point Cloud Segmentation via Attentive Filtering

CVPR 2023poster

Recently, self-attention networks achieve impressive performance in point cloud segmentation due to their superiority in modeling long-range dependencies. However, compared to self-attention mechanism, we find graph convolutions show a stronger ability in capturing local geometry information with le…

Cited by 43SourcePDFScholar
2023

On the Exploration of Local Significant Differences For Two-Sample Test

NeurIPS 2023poster

Recent years have witnessed increasing attentions on two-sample test with diverse real applications, while this work takes one more step on the exploration of local significant differences for two-sample test. We propose the ME$_\text{MaBiD}$, an effective test for two-sample testing, and the basic…

Cited by 1SourcePDFScholar
2023

On the Gini-impurity Preservation For Privacy Random Forests

NeurIPS 2023spotlight

Random forests have been one successful ensemble algorithms in machine learning. Various techniques have been utilized to preserve the privacy of random forests from anonymization, differential privacy, homomorphic encryption, etc., whereas it rarely takes into account some crucial ingredients of le…

Cited by 17SourcePDFScholar
2023

Prompt to be Consistent is Better than Self-Consistent? Few-Shot and Zero-Shot Fact Verification with Pre-trained Language Models

ACL 2023findings

Few-shot or zero-shot fact verification only relies on a few or no labeled training examples. In this paper, we propose a novel method called ProToCo, to Prompt pre-trained language models (PLMs) To be Consistent, for improving the factuality assessment capability of PLMs in the few-shot and zero-sh…

2023

WSDMS: Debunk Fake News via Weakly Supervised Detection of Misinforming Sentences with Contextualized Social Wisdom

EMNLP 2023long main

Fake news debunking primarily focuses on determining the truthfulness of news articles, which oversimplifies the issue as fake news often combines elements of both truth and falsehood. Thus, it becomes crucial to identify specific instances of misinformation within the articles. In this research, we…

Cited by 0SourcecodeScholar
2022

DialMed: A Dataset for Dialogue-based Medication Recommendation

COLING 2022main

Medication recommendation is a crucial task for intelligent healthcare systems. Previous studies mainly recommend medications with electronic health records (EHRs). However, some details of interactions between doctors and patients may be ignored or omitted in EHRs, which are essential for automatic…

2022

DialogConv: A Lightweight Fully Convolutional Network for Multi-view Response Selection

EMNLP 2022main

Current end-to-end retrieval-based dialogue systems are mainly based on Recurrent Neural Networks or Transformers with attention mechanisms. Although promising results have been achieved, these models often suffer from slow inference or huge number of parameters. In this paper, we propose a novel li…

2022

HIRL: Hybrid Image Restoration Based on Hierarchical Deep Reinforcement Learning via Two-Step Analysis

ICASSP 2022accepted

The restoration of hybrid distorted images in real-world scenarios is still a difficult problem due to the fact that the degrading types and degrees are always unknown. Previous studies typically utilize multiple recovery tools to restore images. However, each tool adopted inevitably introduces addi…

Cited by 0SourceScholar
2022

JE2NET: Joint Exploitation and Exploration in Reinforcement Learning Based Image Restoration

ICASSP 2022accepted

Previous reinforcement learning (RL) based image restoration studies typically train RL agents to search for recovery tools from a constructed toolset and iteratively recover images. However, we argue that these agents rely on pre-trained RL models with fixed-length paths for restoration, which perf…

Cited by 0SourceScholar
2022

OctAttention: Octree-Based Large-Scale Contexts Model for Point Cloud Compression

AAAI 2022technical

In point cloud compression, sufficient contexts are significant for modeling the point cloud distribution. However, the contexts gathered by the previous voxel-based methods decrease when handling sparse point clouds. To address this problem, we propose a multiple-contexts deep learning framework ca…

2022

Optimization of Compressive Light Field Display in Dual-Guided Learning

ICASSP 2022accepted

Glass-free compressive light field (CLF) display gains much attention due to their compatibility in holographic-like and three-dimensional (3D) demonstration. Opposite to other analogous devices, CLF display can provide binocular and motion parallaxes by stacking multiple liquid crystal screens with…

Cited by 0SourceScholar
2022

PU-Refiner: A Geometry Refiner with Adversarial Learning for Point Cloud Upsampling

ICASSP 2022accepted

We present PU-Refiner, a generative adversarial network for point cloud upsampling. The generator of our network includes a coarse feature expansion module to create coarse upsampled features, a geometry generation module to regress a coarse point cloud from the coarse upsampled features, and a prog…

Cited by 0SourceScholar
2022

Self-Supervised Arbitrary-Scale Point Clouds Upsampling via Implicit Neural Representation

CVPR 2022poster

Point clouds upsampling is a challenging issue to generate dense and uniform point clouds from the given sparse input. Most existing methods either take the end-to-end supervised learning based manner, where large amounts of pairs of sparse input and dense ground-truth are exploited as supervision i…

Cited by 63PDFcodeScholar
2022

Transient Analysis of Clustered Multitask Diffusion RLS Algorithm

ICASSP 2022accepted

In this paper, we propose a novel clustered multitask diffusion RLS (MT-DRLS) algorithm over network to further improve the performance of its counterpart, the multitask diffusion LMS (MT-DLMS) algorithm. Its transient behavior is investigated, in the mean and mean-square error sense. Simulation res…

Cited by 0SourceScholar
2021

BSP-MonoLoc: Basic Semantic Primitives based Monocular Localization on Roads

IROS 2021poster

Robust visual localization in traffic scenes is a fundamental problem for self-driving vehicles. However, it is still challenging to achieve accurate localization performance because of drastic viewpoint and illumination changes. To address the issues, we design a novel monocular localization framew…

Cited by 3SourceScholar
2021

Improving Transferability of Adversarial Patches on Face Recognition With Generative Models

CVPR 2021poster

Face recognition is greatly improved by deep convolutional neural networks (CNNs). Recently, these face recognition models have been used for identity authentication in security sensitive applications. However, deep CNNs are vulnerable to adversarial patches, which are physically realizable and stea…

Cited by 128PDFScholar
2021

Privacy-Preserving Collaborative Learning With Automatic Transformation Search

CVPR 2021poster

Collaborative learning has gained great popularity due to its benefit of data privacy protection: participants can jointly train a Deep Learning model without sharing their training sets. However, recent works discovered that an adversary can fully recover the sensitive training samples from the sha…

Cited by 61PDFScholar
2021

TS-CAM: Token Semantic Coupled Attention Map for Weakly Supervised Object Localization

ICCV 2021poster

Weakly supervised object localization (WSOL) is a challenging problem when given image category labels but requires to learn object localization models. Optimizing a convolutional neural network (CNN) for classification tends to activate local discriminative regions while ignoring complete object ex…

Cited by 254PDFcodeScholar
2020

Acoustic Scene Classification Using Deep Residual Networks with Late Fusion of Separated High and Low Frequency Paths

ICASSP 2020accepted

We investigate the problem of acoustic scene classification, using a deep residual network applied to log-mel spectrograms complemented by log-mel deltas and delta-deltas. We design the network to take into account that the temporal and frequency axes in spectrograms represent fundamentally differen…

Cited by 0SourceScholar
2020

Fast, Versatile, and Open-loop Stable Running Behaviors with Proprioceptive-only Sensing using Model-based Optimization

ICRA 2020poster

As we build our legged robots smaller and cheaper, stable and agile control without expensive inertial sensors becomes increasingly important. We seek to enable versatile dynamic behaviors on robots with limited modes of state feedback, specifically proprioceptive-only sensing. This work uses model-…

Cited by 10SourceScholar
2020

Risk-constrained Motion Planning for Robot Locomotion: Formulation and Running Robot Demonstration

IROS 2020poster

Robots encounter many risks that threaten the success of practical locomotion tasks. Legs break, electrical components overheat, and feet can unexpectedly slip. When all risks cannot be completely avoided, how does a robot decide its best action? We present a method for planning robot motions by rea…

Cited by 10SourceScholar
2019

FilterReg: Robust and Efficient Probabilistic Point-Set Registration Using Gaussian Filter and Twist Parameterization

CVPR 2019oral

Probabilistic point-set registration methods have been gaining more attention for their robustness to noise, outliers and occlusions. However, these methods tend to be much slower than the popular iterative closest point (ICP) algorithms, which severely limits their usability. In this paper, we cont…

Cited by 162PDFScholar
2019

Visual-Inertial Odometry Tightly Coupled with Wheel Encoder Adopting Robust Initialization and Online Extrinsic Calibration

IROS 2019poster

Combining camera, IMU and wheel encoder is a wise choice for car positioning because of the low cost and complementarity of the sensors. We propose a novel extended visual-inertial odometry algorithm tightly fusing data from the above three sensors. Firstly we propose an IMU-odometer pre-integration…

Cited by 67SourceScholar
2018

Dynamic Actuator Selection and Robust State-Feedback Control of Networked Soft Actuators

ICRA 2018poster

The design of robots that are light, soft, powerful is a grand challenge. Since they can easily adapt to dynamic environments, soft robotic systems have the potential of changing the status-quo of bulky robotics. A crucial component of soft robotics is a soft actuator that is activated by external s…

Cited by 22SourceScholar
2017

Intention-Net: Integrating Planning and Deep Learning for Goal-Directed Autonomous Navigation

CoRL 2017

How can a delivery robot navigate reliably to a destination in a new office building, with minimal prior information? To tackle this challenge, this paper introduces a two-level hierarchical approach, which integrates model-free deep learning and model-based path planning. At the low level, a neural

Cited by 0SourcePDFScholar
2015

Convergence analysis of the augmented complex klms algorithm with pre-tuned dictionary

ICASSP 2015accepted

Complex kernel-based adaptive algorithms have been recently introduced for complex-valued nonlinear system identification. These algorithms are built upon the same framework as complex linear adaptive filtering techniques and Wirtinger's calculus in complex reproducing kernel Hilbert spaces. In this…

Cited by 0SourceScholar