← Search

Gang Wang

136 accepted papers

2026

A Type of Actuator with Large Deformation and Load Capacity: Design and Modeling

ICRA 2026poster

Flexible actuators have garnered extensive attention due to their flexibility and versatility. However, they still exhibit significant limitations in load capacity and structural stiffness. We have developed a multifunctional rigid-flexible coupled actuator with large deformation and high load capac…

Cited by 0Scholar
2026

CAPTCHA Solving for Native GUI Agents: Automated Reasoning-Action Data Generation and Self-Corrective Training

ICML 2026poster

GUI agents are rapidly shifting from multi-module pipelines to end-to-end, native vision-language models (VLMs) that perceive raw screenshots and directly interact with digital devices. Despite rapid progress on general GUI tasks, CAPTCHA solving remains a major challenge. On the other hand, althoug…

Cited by 0SourceScholar
2026

Decoupled Heuristic Multi-Vehicle Emergency Trajectory Planning for Sudden Obstacles

RA-L 2026

The emergence of sudden obstacles can significantly reduce the feasible space and may induce locally non-convex or fragmented space, especially in densely clustered scenarios, making vehicle trajectory planning remarkably challenging. Current methods face computational bottlenecks when generating em

Cited by 0SourceScholar
2026

Demystifying and Enhancing the Efficiency of Large Language Model Based Search Agents

ICLR 2026poster

Large Language Model (LLM)-based search agents have shown remarkable capabilities in solving complex tasks by dynamically decomposing problems and addressing them through interleaved reasoning and retrieval. However, this interleaved paradigm introduces substantial efficiency bottlenecks. First, we…

Cited by 0SourcecodeScholar
2026

Hierarchical Value-Decomposed Offline Reinforcement Learning for Whole-Body Control

ICLR 2026poster

Scaling imitation learning to high-DoF whole-body robots is fundamentally limited by the \textbf{curse of dimensionality} and the prohibitive cost of collecting expert demonstrations. We argue that the core bottleneck is paradigmatic: real-world supervision for whole-body control is inherently imper…

Cited by 0SourceScholar
2026

Instance-wise Adaptive Scheduling via Derivative-Free Meta-Learning

ICLR 2026poster

Deep Reinforcement Learning has achieved remarkable progress in solving NP-hard scheduling problems. However, existing methods primarily focus on optimizing average performance over training instances, overlooking the core objective of solving each individual instance with high quality. While severa…

Cited by 0SourceScholar
2026

Learned Image Compression via Sparse Attention and Adaptive Frequency

CVPR 2026

Learned image compression (LIC) methods surpass traditional algorithms in rate-distortion (RD) performance, but still struggle to optimally balance effectiveness and efficiency. Moreover, although recent studies have demonstrated the effectiveness of utilizing frequency-domain information, they typi

Cited by 0SourceScholar
2026

Localized Coverage Planning for a Heat Transfer Tube Inspection Robot

ICRA 2026poster

The heat transfer tubes of the steam generator are critical components of the nuclear power system and require regular inspection to ensure safety. The SG-Climbot, a quadruped heat transfer tube inspection robot, is equipped with a guiding device capable of simultaneously aligning with and inspectin…

Cited by 0SourceScholar
2026

Mixture-of-World Models: Scaling Multi-Task Reinforcement Learning with Modular Latent Dynamics

ICLR 2026poster

A fundamental challenge in multi-task reinforcement learning (MTRL) is achieving sample efficiency in visual domains where tasks exhibit significant heterogeneity in both observations and dynamics. Model-based RL (MBRL) offers a promising path to sample efficiency through world models, but standard…

Cited by 0SourceScholar
2026

Object-Centric World Models from Few-Shot Annotations for Sample-Efficient Reinforcement Learning

ICLR 2026poster

While deep reinforcement learning (DRL) from pixels has achieved remarkable success, its sample inefficiency remains a critical limitation for real-world applications. Model-based RL (MBRL) addresses this by learning a world model to generate simulated experience, but standard approaches that rely o…

Cited by 0SourceScholar
2026

Rethinking Intermediate Representation for VLM-based Robot Manipulation

CVPR 2026

Vision-Language Model (VLM) is now an important component to enable robust robot manipulation. Yet, using it to translate human instructions into an action-resolvable intermediate representation often needs a tradeoff between VLM-comprehensibility and generalizability. Inspired by context-free gramm

Cited by 0SourceScholar
2025

A Chaotic Dynamics Framework Inspired by Dorsal Stream for Event Signal Processing

ICML 2025poster

Event cameras are bio-inspired vision sensors that encode visual information with high dynamic range, high temporal resolution, and low latency. Current state-of-the-art event stream processing methods rely on end-to-end deep learning techniques. However, these models are heavily dependent on data s…

Cited by 0SourcePDFScholar
2025

A Near-Field 3D Parameter Estimation Method Based on a Symmetric Enhanced Nested Array

ICASSP 2025accepted

In this paper, a high-precision three-dimensional (3-D) near-field (NF) localization method is proposed under an underdetermined case based on a symmetric enhanced nested array (SENA). Firstly, the symmetry of the array and the fourth-order cumulant (FOC) are utilized to construct the equivalent vir…

Cited by 0SourceScholar
2025

A Novel Effective Loop Gait and Stabilizing Morphology Parameterization in Snake Robots

IROS 2025

Improving motion speed and efficiency remains a critical challenge in snake robots gait control. This paper introduces the Loop gait, a novel locomotion gait designed to enhance both speed and energy efficiency of snake robots without passive wheels. Compared to Crawler gait and S-pedal gait, which

Cited by 0SourceScholar
2025

Adaptive Lossless Compression for Genomics Data by Multiple (s, k)-mer Encoding and XLSTM

ICASSP 2025accepted

Learning-based lossless compressors have been validated to have competitive advantages in genomics data (GD) compression. However, learning-based GD-dedicated compressors typically need to be pre-trained on multi-source data and then are directly used to compress another target data, we denote them…

Cited by 0SourceScholar
2025

Bio-inspired Shape Self-Assembly in Large-Scale Swarm Robots Under Information Asymmetry *

IROS 2025

This study investigates the problem of large-scale swarm robots shape self-assembly problem under conditions of information asymmetry. Existing methods assume complete sharing of global information; however, this assumption has significant limitations in terms of resource consumption and swarm emerg

Cited by 0SourceScholar
2025

CART: A Generative Cross-Modal Retrieval Framework With Coarse-To-Fine Semantic Modeling

ACL 2025long

Cross-modal retrieval aims to search for instances, which are semantically related to the query through the interaction of different modal data. Traditional solutions utilize a single-tower or dual-tower framework to explicitly compute the score between queries and candidates, which is challenged by…

Cited by 0SourcePDFScholar
2025

Distributed Flow Shop Scheduling for Heterogeneous Serial Lines With Dueling Double DQN Improved Discrete Particle Swarm Optimization

RA-L 2025

With the development of technology and the changing market demands, the influence of distributed small-batch flexible production modes in the manufacturing industry has been expanding. Meanwhile, production lines with random failures and finite buffer capacities are receiving academic attention due

Cited by 0SourceScholar
2025

DyMoDreamer: World Modeling with Dynamic Modulation

NeurIPS 2025poster

A critical bottleneck in deep reinforcement learning (DRL) is sample inefficiency, as training high-performance agents often demands extensive environmental interactions. Model-based reinforcement learning (MBRL) mitigates this by building world models that simulate environmental dynamics and genera…

Cited by 0SourceScholar
2025

Enabling In-Flight Metamorphosis in Multirotors with a Center-Driven Scissor Extendable Airframe for Adaptive Navigation

ICRA 2025

To address complex mission tasks, multirotors benefit from in-flight reconfiguration that enhances their morphological adaptability. This paper presents the Center-Driven Scissor Extendable Airframe (CDSEA), a novel one-degree-of-freedom (DOF) morphing airframe designed to replace traditional fixed-

Cited by 0SourceScholar
2025

EvdCLIP: Improving Vision-Language Retrieval with Entity Visual Descriptions from Large Language Models

AAAI 2025technical

Vision-language retrieval (VLR) has attracted significant attention in both academia and industry, which involves using text (or images) as queries to retrieve corresponding images (or text). However, existing methods often neglect the rich visual semantics knowledge of entities, thus leading to inc…

Cited by 0SourcePDFScholar
2025

Fourth-Order Cumulant Based 3-D Near-Field Underdetermined Parameter Estimation With Exact Spatial Propagation Model

ICASSP 2025accepted

Based on the exact spherical wavefront model, an under-determined estimation method for three-dimensional (3-D) parameters of near-field (NF) sources using L-shaped nested arrays is proposed, referred to as the cumulant algorithm. This algorithm leverages the temporal-spatial domain cumulants of NF…

Cited by 0SourceScholar
2025

Genomics Data Lossless Compression with (S, K)-Mer Encoding and Deep Neural Networks

AAAI 2025technical

Learning-based compression shows competitive compression ratios for genomics data. It often includes three types of compressors: static, adaptive and semi-adaptive. However, these existing compressors suffer from inferior compression ratios or throughput, and adaptive compressors also faces model c…

2025

HFedPFS: Heterogeneous Federated Learning with Personalized Data Feature Sharing

ICASSP 2025accepted

Federated learning (FL) is a distributed machine learning technique enabling multiple clients to jointly train a global model while preserving the privacy of their non-IID (non-independent and identically) data. However, traditional FL approaches require clients to use the same model structure as th…

Cited by 0SourceScholar
2025

JTD-UAV: MLLM-Enhanced Joint Tracking and Description Framework for Anti-UAV Systems

CVPR 2025poster

Unmanned Aerial Vehicles (UAVs) are widely adopted across various fields, yet they raise significant privacy and safety concerns, demanding robust monitoring solutions. Existing anti-UAV methods primarily focus on position tracking but fail to capture UAV behavior and intent. To address this, we int…

Cited by 0SourcePDFScholar
2025

Localized Coverage Planning for a Heat Transfer Tube Inspection Robot

RA-L 2025

The heat transfer tubes of the steam generator are critical components of the nuclear power system and require regular inspection to ensure safety. The SG-Climbot, a quadruped heat transfer tube inspection robot, is equipped with a guiding device capable of simultaneously aligning with and inspectin

Cited by 1SourceScholar
2025

Multi-Vehicle Cooperative Persistent Coverage for Random Target Search

RA-L 2025

This letter investigates the target search problem for a network of autonomous vehicles, aiming to maximize the detection of randomly appearing targets within a given area. Considering no prior knowledge of the targets is available, we propose a multi-vehicle cooperative persistent coverage scheme u

Cited by 4SourceScholar
2025

Multi-source Data Lossless Compression via Parallel Expansion Mapping and xLSTM

ICASSP 2025accepted

Explosive growth of multi-source data (MSD) poses challenges in data transmitting and storing. Neural Network (NN)-based lossless compressors are an important type of compression approaches to alleviate these problems. However, existing NN-based lossless compressors suffer from poor compression rati…

Cited by 0SourceScholar
2025

Optical Flow Estimation for Tiny Objects: New Problem, Specialized Benchmark, and Bioinspired Scheme

IJCAI 2025

Optical flow is pivotal in video-based tasks, yet existing methods mostly focus on medium-/large-size objects, while underperforming when characterizing the motion of tiny objects. To bridge this gap, we introduce the On-off Time-delay with Hassenstein-Reichardt correlator (OTHR), a computationally

2025

Perplexity Trap: PLM-Based Retrievers Overrate Low Perplexity Documents

ICLR 2025poster

Previous studies have found that PLM-based retrieval models exhibit a preference for LLM-generated content, assigning higher relevance scores to these documents even when their semantic quality is comparable to human-written ones. This phenomenon, known as source bias, threatens the sustainable deve…

2025

PurpCode: Reasoning for Safer Code Generation

NeurIPS 2025poster

We introduce PurpCode, the first post-training recipe for training safe code reasoning models towards generating secure code and defending against malicious cyberactivities. PurpCode trains a reasoning model in two stages: (i) Rule Learning, which explicitly teaches the model to reference cybersafet…

Cited by 0SourceScholar
2025

Q-PRM: Adaptive Query Rewriting for Retrieval-Augmented Generation via Step-level Process Supervision

EMNLP 2025

Query rewriting plays a pivotal role in Retrieval-Augmented Generation (RAG) by refining real-world queries of varying complexity. Existing approaches typically rely on outcome-supervised training or heuristic rules to guide the rewriting process. However, these paradigms often struggle to handle qu

Cited by 0SourcePDFScholar
2025

Robust Offline Imitation Learning Through State-level Trajectory Stitching

IROS 2025

Imitation learning (IL) has proven effective for enabling robots to acquire visuomotor skills through expert demonstrations. However, traditional IL methods are limited by their reliance on high-quality, often scarce, expert data, and suffer from covariate shift. To address these challenges, recent

Cited by 0SourcecodeScholar
2025

Steering Away from Harm: An Adaptive Approach to Defending Vision Language Model Against Jailbreaks

CVPR 2025poster

Vision Language Models (VLMs) can produce unintended and harmful content when exposed to adversarial attacks, particularly because their vision capabilities create new vulnerabilities. Existing defenses, such as input preprocessing, adversarial training, and response evaluation-based methods, are of…

2025

The Emperor's New Clothes in Benchmarking? A Rigorous Examination of Mitigation Strategies for LLM Benchmark Data Contamination

ICML 2025poster

Benchmark Data Contamination (BDC)—the inclusion of benchmark testing samples in the training set—has raised increasing concerns in Large Language Model (LLM) evaluation, leading to falsely inflated performance estimates and undermining evaluation reliability. To address this, researchers have propo…

2025

Tracking Tiny Drones against Clutter: Large-Scale Infrared Benchmark with Motion-Centric Adaptive Algorithm

ICCV 2025poster

Tracking flying drones in infrared videos is a crucial yet challenging task. Existing drone trackers and datasets have limitations in dealing with and characterizing tiny targets (<=20x20 pixels) against highly complex backgrounds. To tackle this issue, we have developed a large-scale benchmark for…

2025

pFedES: Generalized Proxy Feature Extractor Sharing for Model Heterogeneous Personalized Federated Learning

AAAI 2025technical

Federated learning (FL), as a privacy-preserving collaborative machine learning paradigm, has attracted significant interest from industry and academia. To allow each data owner (FL client) to train a heterogeneous and personalized local model based on its local data distribution, system resources a…

2024

3-D Near-Field Localization by Jointly Exploiting Spatial and Temporal Information Based on a Nonuniform Cross Array

ICASSP 2024accepted

In this paper, an underdetermined three-dimensional (3-D) near-field source localization method is proposed, based on a two-dimensional (2-D) symmetric nonuniform cross array. Firstly, the fourth-order cumulant of the near-field observations with multiple delay lags is exploited to construct virtual…

Cited by 0SourceScholar
2024

A Framework for Real-time Generation of Multi-directional Traversability Maps in Unstructured Environments

ICRA 2024poster

In complex unstructured environments, accurate terrain traversability analysis is a fundamental requirement for the successful execution of any movements of ground robots, especially given that terrain traversability often exhibits anisotropy. However, the difficulty in obtaining multi-directional t…

Cited by 1SourceScholar
2024

A New Fourth-Order Sparse Array Generator Based on Sum-Difference Co-Array Analysis

ICASSP 2024accepted

In this paper, based on sum-difference co-array analysis, a new fourth-order sparse array called sum-difference-FODC (SD-FODC) is proposed, allowing the construction of a fourth-order DCA with long consecutive lags using the continuous segments in the second-order DCA and SCA of the original array.…

Cited by 0SourceScholar
2024

BE-SLAM: BEV-Enhanced Dynamic Semantic SLAM with Static Object Reconstruction

IROS 2024poster

The quality of a robot’s environmental perception determines whether it can achieve more intelligent applications, such as semantic interaction with humans. SLAM, on the other hand, is one of the crucial capabilities for a robot to perceive its environment. However, when only a monocular image is pr…

Cited by 0SourceScholar
2024

Cocktail: A Comprehensive Information Retrieval Benchmark with LLM-Generated Documents Integration

ACL 2024findings

The proliferation of Large Language Models (LLMs) has led to an influx of AI-generated content (AIGC) on the internet, transforming the corpus of Information Retrieval (IR) systems from solely human-written to a coexistence with LLM-generated content. The impact of this surge in AIGC on IR systems r…

2024

DGA-GNN: Dynamic Grouping Aggregation GNN for Fraud Detection

AAAI 2024technical

Fraud detection has increasingly become a prominent research field due to the dramatically increased incidents of fraud. The complex connections involving thousands, or even millions of nodes, present challenges for fraud detection tasks. Many researchers have developed various graph-based methods t…

2024

Enhancing Joint Dynamics Modeling for Underwater Robotics Through Stochastic Extension

RA-L 2024

Accurate joint dynamics models are essential for the compliance and robustness of robot control, especially for robots operating in complex underwater environments. To improve the precision of joint dynamics models, much research focuses on refining specific parameters or incorporating previously ov

Cited by 1SourceScholar
2024

FedSSA: Semantic Similarity-based Aggregation for Efficient Model-Heterogeneous Personalized Federated Learning

IJCAI 2024poster

Federated learning (FL) is a privacy-preserving collaboratively machine learning paradigm. Traditional FL requires all data owners (a.k.a. FL clients) to train the same local model. This design is not well-suited for scenarios involving data and/or system heterogeneity. Model-Heterogeneous Personali…

2024

Federated Model Heterogeneous Matryoshka Representation Learning

NeurIPS 2024poster

Model heterogeneous federated learning (MHeteroFL) enables FL clients to collaboratively train models with heterogeneous structures in a distributed fashion. However, existing MHeteroFL methods rely on training loss to transfer knowledge between the client model and the server model, resulting in li…

Cited by 6SourcePDFScholar
2024

Increasing the Absolute Position Accuracy of Industrial Robots by Means of a Deep Continual Evidential Regression Model *

ICRA 2024poster

The use of industrial robots represents a key technology for increasing productivity and efficiency in manufacturing. However, their low absolute position accuracy still denies the broad substitution of machine tools by industrial robots. In this paper, a data-driven method for accuracy enhancement…

Cited by 0SourceScholar
2024

Learning Locomotion for Quadruped Robots via Distributional Ensemble Actor-Critic

RA-L 2024

Domain randomization introduces perturbations in the simulation to make controllers less susceptible to the reality gap, which enables remarkable sim-to-real transfer on real quadruped robots. However, aleatoric uncertainty originating from perturbations could often lead to suboptimal controllers. I

Cited by 10SourceScholar
2024

Multi-Constellation-Inspired Single-Shot Global LiDAR Localization

AAAI 2024technical

Global localization is a challenging task for intelligent robots, as its accuracy directly contributes to the performance of downstream navigation and planning tasks. However, existing literature focus more on the place retrieval and the success rate of localization, with limited attention given to…

2024

RICE: Breaking Through the Training Bottlenecks of Reinforcement Learning with Explanation

ICML 2024spotlight

Deep reinforcement learning (DRL) is playing an increasingly important role in real-world applications. However, obtaining an optimally performing DRL agent for complex tasks, especially with sparse rewards, remains a significant challenge. The training of a DRL agent can be often trapped in a bottl…

2024

Three-Dimensional Spatial-Temporal Near-Field Passive Localization Based on an Exact Spatial Propagation Model

ICASSP 2024accepted

Based on the exact source-sensor spatial geometry, a three-dimensional (3-D) spatial-temporal localization algorithm for multiple near-field (NF) sources is proposed without adopting the Fresnel approximation, which simplifies the spatial phase difference by Taylors polynomial. In addition, consider…

Cited by 1SourceScholar
2024

Unified Single-Stage Transformer Network for Efficient RGB-T Tracking

IJCAI 2024poster

Most existing RGB-T tracking networks extract modality features in a separate manner, which lacks interaction and mutual guidance between modalities. This limits the network's ability to adapt to the diverse dual-modality appearances of targets and the dynamic relationships between the modalities. A…

2024

Weak-eval-Strong: Evaluating and Eliciting Lateral Thinking of LLMs with Situation Puzzles

NeurIPS 2024poster

While advancements in NLP have significantly improved the performance of Large Language Models (LLMs) on tasks requiring vertical thinking, their lateral thinking capabilities remain under-explored and challenging to measure due to the complexity of assessing creative thought processes and the scarc…

2023

An Effective Deployment of Contrastive Learning in Multi-label Text Classification

ACL 2023findings

The effectiveness of contrastive learning technology in natural language processing tasks is yet to be explored and analyzed. How to construct positive and negative samples correctly and reasonably is the core challenge of contrastive learning. It is even harder to discover contrastive objects in mu…

Cited by 29SourcePDFScholar
2023

Bias Reduced Semidefinite Relaxation Method for Multistatic Localization in the Absence of Transmitter Position And Its Synchronization

ICASSP 2023accepted

This paper addresses the challenging problem of multistatic localization of a stationary object with a set of synchronized receivers, when the transmitter position is unknown and the synchronization with the transmitter is unavailable. Using the time delay measurements from the direct and indirect p…

Cited by 0SourceScholar
2023

Efficient and Robust Time-Optimal Trajectory Planning and Control for Agile Quadrotor Flight

RA-L 2023

Agile quadrotor flight relies on rapidly planning and accurately tracking time-optimal trajectories, a technology critical to their application in the wild. However, the computational burden of computing time-optimal trajectories based on the full quadrotor dynamics (typically on the order of minute

Cited by 30SourcecodeScholar
2023

JEPOO: Highly Accurate Joint Estimation of Pitch, Onset and Offset for Music Information Retrieval

IJCAI 2023poster

Melody extraction is a core task in music information retrieval, and the estimation of pitch, onset and offset are key sub-tasks in melody extraction. Existing methods have limited accuracy, and work for only one type of data, either single-pitch or multi-pitch. In this paper, we propose a highly ac…

Cited by 7SourcePDFScholar
2023

LeanSpeech: The Microsoft Lightweight Speech Synthesis System for Limmits Challenge 2023

ICASSP 2023accepted

This paper describes the Microsoft Text-to-Speech (TTS) system: LeanSpeech for LIMMITS (Lightweight, Multi-speaker, Multi-lingual Indic TTS) Challenge 2023<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup>, which is part of ICASSP2023 to encourage…

Cited by 0SourceScholar
2023

Out-of-Distribution Detection based on In-Distribution Data Patterns Memorization with Modern Hopfield Energy

ICLR 2023poster

Out-of-Distribution (OOD) detection is essential for safety-critical applications of deep neural networks. OOD detection is challenging since DNN models may produce very high logits value even for OOD samples. Hence, it is of great difficulty to discriminate OOD data by directly adopting Softmax on…

2023

Real-Time Whole-Body Collision Avoidance and Path Following of a Snake Robot Through MPC-based Optimization Strategies

IROS 2023poster

The work in this paper delves into the challenge of whole elongated body's obstacle avoidance during path following for a class of bionic snake robots. Currently, most studies focus solely on preventing the robot's head from colliding with obstacles through designed controllers. However, due to the…

Cited by 5SourceScholar
2023

STORM: Efficient Stochastic Transformer based World Models for Reinforcement Learning

NeurIPS 2023poster

Recently, model-based reinforcement learning algorithms have demonstrated remarkable efficacy in visual input environments. These approaches begin by constructing a parameterized simulation world model of the real environment through self-supervised learning. By leveraging the imagination of the wo…

2023

Self-Supervised Learning with Explorative Knowledge Distillation

ICASSP 2023accepted

Previous paradigms have combined self-supervised learning (SSL) with knowledge distillation to compress a self-supervised teacher model into a smaller student. In this work, we devise a self-supervised explorative distillation (SSED) algorithm to improve the representation quality of the lightweight…

Cited by 0SourceScholar
2023

ZoomTrack: Target-aware Non-uniform Resizing for Efficient Visual Tracking

NeurIPS 2023spotlight

Recently, the transformer has enabled the speed-oriented trackers to approach state-of-the-art (SOTA) performance with high-speed thanks to the smaller input size or the lighter feature extraction backbone, though they still substantially lag behind their corresponding performance-oriented versions.…

2022

AdaPID: An Adaptive PID Optimizer for Training Deep Neural Networks

ICASSP 2022accepted

Deep neural networks (DNNs) have well-documented merits in learning nonlinear functions in high-dimensional spaces. Stochastic gradient descent (SGD)-type optimization algorithms are the ‘workhorse’ for training DNNs. Nonetheless, such algorithms often suffer from slow convergence, sizable fluctuati…

Cited by 0SourceScholar
2022

Conjugate Augmented Spatial-Temporal Near-Field Sources Localization with Cross Array

ICASSP 2022accepted

A new near-field source localization method is proposed for two-dimensional (2-D) direction-of-arrival (DOA) and range estimation based on a symmetrical cross array. It first employs the conjugate symmetry property of the signal auto-correlation at different time delays to construct a conjugate augm…

Cited by 0SourceScholar
2022

Event-Triggered Tracking Control Scheme for Quadrotors with External Disturbances: Theory and Validations

ICRA 2022poster

This article studies the tracking control of a quadrotor unmanned aerial vehicle (UAV) under time-varying external disturbances. An event-triggered sliding mode control (SMC) strategy is proposed by introducing a new triggering condition form of desired trajectory, quadrotor position, and velocity.…

Cited by 6SourceScholar
2022

Knowledge Distillation via the Target-Aware Transformer

CVPR 2022oral

Knowledge distillation becomes a de facto standard to improve the performance of small neural networks. Most of the previous works propose to regress the representational features from the teacher to the student in a one-to-one spatial matching fashion. However, people tend to overlook the fact that…

Cited by 150PDFcodeScholar
2022

Learning from Students: Online Contrastive Distillation Network for General Continual Learning

IJCAI 2022poster

The goal of General Continual Learning (GCL) is to preserve learned knowledge and learn new knowledge with constant memory from an infinite data stream where task boundaries are blurry. Distilling the model's response of reserved samples between the old and the new models is an effective way to achi…

2022

Semidefinite Relaxation Method for Moving Object Localization Using a Stationary Transmitter at Unknown Position

ICASSP 2022accepted

This paper addresses the multistatic localization of a moving object in position and velocity using time delay (TD) and Doppler frequency shift (DFS) measurements, where the position of the transmitter is unknown and has not yet been synchronized with the receivers. Based on the TD and DFS measureme…

Cited by 0SourceScholar
2022

Task Autonomous Medical Robot for Both Incision Stapling and Staples Removal

RA-L 2022

Surgical incision is a pervasive procedure in medical environments. Stapling is an incision closure method that is comparable to stitching. While incision closure is performed immediately after a surgery, staples are removed several weeks after the closure. The workload of surgeons and the probabili

Cited by 20SourceScholar
2021

Collaborative Uncertainty in Multi-Agent Trajectory Forecasting

NeurIPS 2021poster

Uncertainty modeling is critical in trajectory-forecasting systems for both interpretation and safety reasons. To better predict the future trajectories of multiple agents, recent works have introduced interaction modules to capture interactions among agents. This approach leads to correlations amon…

Cited by 24SourcePDFScholar
2021

Comparative Validation Study on Bioinspired Morphology-Adaptation Flight Performance of a Morphing Quad-Rotor

RA-L 2021

Inspired by the advantages from natural birds' shape-morphing adapted flight performance, this work investigates how the bioinspired adaptative morph induced inertia variation affects the flight performance of a morphing quad-rotor. Extensive numerical and experimental evaluations on a custom-built

Cited by 20SourceScholar
2021

MDANet: Multi-Modal Deep Aggregation Network for Depth Completion

ICRA 2021poster

Depth completion aims to recover the dense depth map from sparse depth data and RGB image respectively. However, due to the huge difference between the multi-modal signal input, vanilla convolutional neural network and simple fusion strategy cannot extract features from sparse data and aggregate mul…

Cited by 17SourcecodeScholar
2021

Model-Based Trajectory Prediction and Hitting Velocity Control for a New Table Tennis Robot

IROS 2021poster

Currently, most table tennis robots concentrate on the canonical position control problem while ignoring the actual velocity control requirements. In this paper, we consider these requirements and propose a new table tennis robot framework. First, a tailor-made mechanical structure is designed such…

Cited by 20SourceScholar
2021

Morphologically Adapatative Quad-Rotor Towards Acquiring High-Performance Flight: A Comparative Study and Validation

ICRA 2021poster

This paper presents our comparative study on how the flight performances of an in-flight morphing quad-rotor are affected by the morph induced inertia variation. A custom-built in-flight morphing quad-rotor was employed in numerical and experimental tests for the study and analysis. In these tests,…

Cited by 0SourceScholar
2021

Route Coverage Testing for Autonomous Vehicles via Map Modeling

ICRA 2021poster

Autonomous vehicles (AVs) play an important role in transforming our transportation systems and relieving traffic congestion. To guarantee their safety, AVs must be sufficiently tested before they are deployed to public roads. Existing testing often focuses on AVs’ collision avoidance on a given rou…

Cited by 34SourceScholar
2020

DA4AD: End-to-End Deep Attention-based Visual Localization for Autonomous Driving

ECCV 2020poster

We present a visual localization framework based on novel deep attention aware features for autonomous driving that achieves centimeter level localization accuracy. Conventional approaches to the visual localization problem rely on handcrafted features or human-made objects on the road. They are kno…

Cited by 57SourcePDFScholar
2020

Decentralized TD Tracking with Linear Function Approximation and its Finite-Time Analysis

NeurIPS 2020poster

The present contribution deals with decentralized policy evaluation in multi-agent Markov decision processes using temporal-difference (TD) methods with linear function approximation for scalability. The agents cooperate to estimate the value function of such a process by observing continual state t…

Cited by 40SourcePDFScholar
2020

Distributed Consensus Control of Multiple UAVs in a Constrained Environment

ICRA 2020poster

In this paper, we investigate the consensus problem of multiple unmanned aerial vehicles (UAVs) in the presence of environmental constraints under a general communication topology containing a directed spanning tree. First, based on a position transformation function, we propose a novel dynamic refe…

Cited by 15SourceScholar
2020

Finite-Time Analysis of Decentralized Temporal-Difference Learning with Linear Function Approximation

AISTATS 2020poster

Motivated by the emerging use of multi-agent reinforcement learning (MARL) in engineering applications such as networked robotics, swarming drones, and sensor networks, we investigate the policy evaluation problem in a fully decentralized setting, using temporal-difference (TD) learning with linear…

Cited by 63SourcePDFScholar
2020

Finite-Time Error Bounds for Biased Stochastic Approximation with Applications to Q-Learning

AISTATS 2020poster

Inspired by the widespread use of Q-learning algorithms in reinforcement learning (RL), this present paper studies a class of biased stochastic approximation (SA) procedures under an ‘ergodic-like’ assumption on the underlying stochastic noise sequence. Leveraging a \emph{multistep Lyapunov functio…

Cited by 10SourcePDFScholar
2020

Learning connectivity and higher-order interactions in radial distribution grids

ICASSP 2020accepted

To perform any meaningful optimization task, distribution grid operators need to know the topology of their grids. Although power grid topology identification and verification has been recently studied, discovering instantaneous interplay among subsets of buses, also known as higher-order interactio…

Cited by 0SourceScholar
2019

An Approximation-Free Simple Control Scheme for Uncertain Quadrotor Systems: Theory and Validations

IROS 2019poster

In this paper, a simple tracking control scheme is proposed for quadrotor systems with uncertain dynamics. It precludes the necessity for prohibitive analytic computation of the derivatives of the desired (virtual) attitude that is typically employed in controlling quadrotor systems. Moreover, this…

Cited by 3SourceScholar
2019

Boundary-Aware Feature Propagation for Scene Segmentation

ICCV 2019poster

In this work, we address the challenging issue of scene segmentation. To increase the feature similarity of the same object while keeping the feature discrimination of different objects, we explore to propagate information throughout the image under the control of objects' boundaries. To this end, w…

Cited by 286PDFcodeScholar
2019

Semantic Correlation Promoted Shape-Variant Context for Segmentation

CVPR 2019oral

Context is essential for semantic segmentation. Due to the diverse shapes of objects and their complex layout in various scene images, the spatial scales and shapes of contexts for different objects have very large variation. It is thus ineffective or inefficient to aggregate various context informa…

Cited by 215PDFcodeScholar
2019

Unpaired Image Captioning via Scene Graph Alignments

ICCV 2019poster

Most of current image captioning models heavily rely on paired image-caption datasets. However, getting large scale image-caption paired data is labor-intensive and time-consuming. In this paper, we present a scene graph-based approach for unpaired image captioning. Our framework comprises an image…

Cited by 209PDFScholar
2018

A Bi-Directional Message Passing Model for Salient Object Detection

CVPR 2018poster

Recent progress on salient object detection is beneficial from Fully Convolutional Neural Network (FCN). The saliency cues contained in multi-level convolutional features are complementary for detecting salient objects. How to integrate multi-level features becomes an open problem in saliency detect…

Cited by 579SourcePDFScholar
2018

Adaptive Path Following of Snake Robot on Ground with Unknown and Varied Friction Coefficients

IROS 2018poster

This paper investigates the straight path following problem for a class of underactuated bio-inspired snake robots on ground with unknown and varied friction coefficients. Existing works usually design control input requiring the exact values of these friction coefficients, which however rely on the…

Cited by 11SourceScholar
2018

Adaptive Path Following of Underactuated Snake Robot on Unknown and Varied Frictions Ground: Theory and Validations

RA-L 2018

This letter investigates the straight path following problem for a class of underactuated bioinspired snake robots on ground with unknown and varied friction coefficients. Existing works usually design control input requiring the exact values of these friction coefficients, which however rely on the

Cited by 37SourceScholar
2018

Context Contrasted Feature and Gated Multi-Scale Aggregation for Scene Segmentation

CVPR 2018poster

Scene segmentation is a challenging task as it need label every pixel in the image. It is crucial to exploit discriminative context and aggregate multi-scale features to achieve better segmentation. In this paper, we first propose a novel context contrasted local feature that not only leverages the…

Cited by 450SourcePDFScholar
2018

Dpca: Dimensionality Reduction for Discriminative Analytics of Multiple Large-Scale Datasets

ICASSP 2018accepted

Principal component analysis (PCA) has well-documented merits for data extraction and dimensionality reduction. PCA deals with a single dataset at a time, and it is challenged when it comes to analyzing multiple datasets. Yet in certain setups, one wishes to extract the most significant information…

Cited by 0SourceScholar
2018

Dual Attention Matching Network for Context-Aware Feature Sequence Based Person Re-Identification

CVPR 2018poster

Typical person re-identification (ReID) methods usually describe each pedestrian with a single feature vector and match them in a task-specific metric space. However, the methods based on a single feature vector are not sufficient enough to overcome visual ambiguity, which frequently occurs in real…

Cited by 502SourcePDFScholar
2018

Look, Imagine and Match: Improving Textual-Visual Cross-Modal Retrieval With Generative Models

CVPR 2018poster

Textual-visual cross-modal retrieval has been a hot research topic in both computer vision and natural language processing communities. Learning appropriate representations for multi-modal data is crucial for the cross-modal retrieval performance. Unlike existing image-text retrieval approaches that…

Cited by 476SourcePDFScholar
2018

Motion-Guided Cascaded Refinement Network for Video Object Segmentation

CVPR 2018poster

Deep CNNs have achieved superior performance in many tasks of computer vision and image understanding. However, it is still difficult to effectively apply deep CNNs to video object segmentation(VOS) since treating video frames as separate and static will lose the information hidden in motion. To tac…

Cited by 131SourcePDFScholar
2018

Progressive Attention Guided Recurrent Network for Salient Object Detection

CVPR 2018poster

Effective convolutional features play an important role in saliency estimation but how to learn powerful features for saliency is still a challenging task. FCN-based methods directly apply multi-level convolutional features without distinction, which leads to sub-optimal results due to the distracti…

Cited by 756SourcePDFScholar
2018

SSNet: Scale Selection Network for Online 3D Action Prediction

CVPR 2018poster

In action prediction (early action recognition), the goal is to predict the class label of an ongoing action using its observed part so far. In this paper, we focus on online action prediction in streaming 3D skeleton sequences. A dilated convolutional network is introduced to model the motion dynam…

Cited by 68SourcePDFScholar
2018

Stochastic Downsampling for Cost-Adjustable Inference and Improved Regularization in Convolutional Networks

CVPR 2018poster

It is desirable to train convolutional networks (CNNs) to run more efficiently during inference. In many cases however, the computational budget that the system has for inference cannot be known beforehand during training, or the inference budget is dependent on the changing real-time resource avail…

Cited by 19SourcePDFScholar
2017

Blob Reconstruction Using Unilateral Second Order Gaussian Kernels With Application to High-ISO Long-Exposure Image Denoising

ICCV 2017poster

Blob detection and image denoising are fundamental, and sometimes related, tasks in computer vision. In this paper, we propose a blob reconstruction method using scale-invariant normalized unilateral second order Gaussian kernels. Unlike other blob detection methods, our method suppresses non-blob s…

Cited by 17PDFScholar
2017

Episodic CAMN: Contextual Attention-Based Memory Networks With Iterative Feedback for Scene Labeling

CVPR 2017poster

Scene labeling can be seen as a sequence-sequence prediction task (pixels-labels), and it is quite important to leverage relevant context to enhance the performance of pixel classification. In this paper, we introduce an episodic attention-based memory network to achieve the goal. We present a unifi…

Cited by 17PDFScholar
2017

Global Context-Aware Attention LSTM Networks for 3D Action Recognition

CVPR 2017poster

Long Short-Term Memory (LSTM) networks have shown superior performance in 3D human action recognition due to their power in modeling the dynamics and dependencies in sequential data. Since not all joints are informative for action analysis and the irrelevant joints often bring a lot of noise, we nee…

Cited by 822PDFScholar
2017

SPARTA: Sparse phase retrieval via Truncated Amplitude flow

ICASSP 2017accepted

A linear-time algorithm termed SPARse Truncated Amplitude flow (SPARTA) is developed for the phase retrieval (PR) of sparse signals. Upon formulating the sparse PR as a non-convex empirical loss minimization task, SPARTA emerges as an iterative solver consisting of two components: s1) a sparse ortho…

Cited by 0SourceScholar
2016

Context-Aware Gaussian Fields for Non-Rigid Point Set Registration

CVPR 2016poster

Point set registration (PSR) is a fundamental problem in computer vision and pattern recognition, and it has been successfully applied to many applications. Although widely used, existing PSR methods cannot align point sets robustly under degradations, such as deformation, noise, occlusion, outlier,…

Cited by 34PDFScholar
2016

NTU RGB+D: A Large Scale Dataset for 3D Human Activity Analysis

CVPR 2016poster

Recent approaches in depth-based human activity analysis achieved outstanding performance and proved the effectiveness of 3D representation for classification of action classes. Currently available depth-based and RGB+D-based action recognition benchmarks have a number of limitations, including the…

Cited by 3655PDFcodeScholar
2016

Solving Random Systems of Quadratic Equations via Truncated Generalized Gradient Flow

NeurIPS 2016poster

This paper puts forth a novel algorithm, termed \emph{truncated generalized gradient flow} (TGGF), to solve for $\bm{x}\in\mathbb{R}^n/\mathbb{C}^n$ a system of $m$ quadratic equations $y_i=|\langle\bm{a}_i,\bm{x}\rangle|^2$, $i=1,2,\ldots,m$, which even for $\left\{\bm{a}_i\in\mathbb{R}^n/\mathbb{C…

Cited by 54SourcePDFScholar
2015

Adaptive censoring for large-scale regressions

ICASSP 2015accepted

Albeit being in the big data era, a significant percentage of data accrued can be overlooked while maintaining reasonable quality of statistical inference at affordable complexity. By capitalizing on data redundancy, interval censoring is leveraged here to cope with the scarcity of resources needed…

Cited by 0SourceScholar
2015

Integrating Parametric and Non-Parametric Models For Scene Labeling

CVPR 2015poster

We adopt Convolutional Neural Networks (CNN) as our parametric model to learn discriminative features and classifiers for local patch classification. As visually similar pixels are indistinguishable from local context, we alleviate such ambiguity by putting a global scene constraint. We estimate the…

Cited by 56SourcePDFScholar
2015

Multi-Manifold Deep Metric Learning for Image Set Classification

CVPR 2015poster

In this paper, we propose a multi-manifold deep metric learning (MMDML) method for image set classification, which aims to recognize an object of interest from a set of image instances captured from varying viewpoints or under varying illuminations. Motivated by the fact that manifold can be effecti…

Cited by 242SourcePDFScholar