← Search

Lin Zhao

48 accepted papers

2026

AERO-MPPI: Anchor-Guided Ensemble Trajectory Optimization for Agile Mapless Drone Navigation

ICRA 2026poster

Agile mapless navigation in cluttered 3D environments poses significant challenges for autonomous drones. Conventional mapping–planning–control pipelines incur high computational cost and propagate estimation errors. We present AERO-MPPI, a fully GPU-accelerated framework that unifies perception and…

2026

Finite-Time Analysis of Actor-Critic Methods with Deep Neural Network Approximation

ICLR 2026poster

Actor–critic (AC) algorithms underpin many of today’s most successful reinforcement learning (RL) applications, yet their finite-time convergence in realistic settings remains largely underexplored. Existing analyses often rely on oversimplified formulations and are largely confined to linear functi…

Cited by 0SourceScholar
2026

HierAmp: Coarse-to-Fine Autoregressive Amplification for Generative Dataset Distillation

CVPR 2026

Dataset distillation often prioritizes global semantic proximity when creating small surrogate datasets for original large-scale ones. However, object semantics are inherently hierarchical. For example, the position and appearance of a bird's eyes are constrained by the outline of its head. Global p

Cited by 0SourcecodeScholar
2026

MacroNav: Multi-Task Context Representation Learning Enables Efficient Navigation in Unknown Environments

RA-L 2026

Autonomous navigation in unknown environments requires multi-scale spatial understanding that captures geometric details, topological connectivity, and global structure to support high-level decision making under partial observability. Existing approaches struggle to efficiently capture such multi-s

Cited by 1SourceScholar
2026

SIGN: Safety-Aware Image-Goal Navigation for Autonomous Drones Via Reinforcement Learning

ICRA 2026poster

Image-goal navigation (ImageNav) tasks a robot with autonomously exploring an unknown environment and reaching a location that visually matches a given target image. While prior works primarily study ImageNav for ground robots, enabling this capability for autonomous drones is substantially more cha…

2026

SIGN: Safety-Aware Image-Goal Navigation for Autonomous Drones via Reinforcement Learning

RA-L 2026

Image-goal navigation (ImageNav) tasks a robot with autonomously exploring an unknown environment and reaching a location that visually matches a given target image. While prior works primarily study ImageNav for ground robots, enabling this capability for autonomous drones is substantially more cha

Cited by 1SourcecodeScholar
2026

TransforMARS: Fault-Tolerant Self-Reconfiguration for Arbitrary-Shaped Modular Aerial Robot Systems

ICRA 2026poster

Modular Aerial Robot Systems (MARS) consist of multiple drone modules that are physically bound together to form a single structure for flight. Exploiting structural redundancy, MARS can be reconfigured into different formations to mitigate unit or rotor failures and maintain stable flight. Prior wo…

Cited by 0codeScholar
2026

VIPO: Value Function Inconsistency Penalized Offline Reinforcement Learning

ICML 2026poster

Offline reinforcement learning (RL) learns effective policies from pre-collected datasets, offering a practical solution for applications where online interactions are risky or costly. Model-based approaches are particularly advantageous for offline RL, owing to their data efficiency and generalizab…

Cited by 0SourceScholar
2025

Efficient Anchor Graph Clustering Through Enhanced Within-Cluster Homogeneity

ICASSP 2025accepted

Anchor-based clustering methods have gained attention for their efficiency in subspace, multi-view, and ensemble clustering tasks. Most existing methods focus on using anchors to reduce computational complexity in the original data space. However, clustering directly on anchors, followed by label pr…

Cited by 0SourceScholar
2025

FD2-Net: Frequency-Driven Feature Decomposition Network for Infrared-Visible Object Detection

AAAI 2025technical

Infrared-visible object detection (IVOD) seeks to harness the complementary information in infrared and visible images, thereby enhancing the performance of detectors in complex environments. However, existing methods often neglect the frequency characteristics of complementary information, such as…

Cited by 2SourcePDFScholar
2025

FED-PsyAU: Privacy-Preserving Micro-Expression Recognition via Psychological AU Coordination and Dynamic Facial Motion Modeling

ICCV 2025poster

Micro-expressions (MEs) are brief, low-intensity, often localized facial expressions. They could reveal genuine emotions individuals may attempt to conceal, valuable in contexts like criminal interrogation and psychological counseling. However, ME recognition (MER) faces challenges, such as small sa…

2025

Incomplete Multi-view Clustering via Hierarchical Semantic Alignment and Cooperative Completion

NeurIPS 2025poster

Incomplete multi-view data, where certain views are entirely missing for some samples, poses significant challenges for traditional multi-view clustering methods. Existing deep incomplete multi-view clustering approaches often rely on static fusion strategies or two-stage pipelines, leading to subop…

Cited by 0SourcecodeScholar
2025

Label-Efficient Data Augmentation with Video Diffusion Models for Guidewire Segmentation in Cardiac Fluoroscopy

AAAI 2025technical

The accurate segmentation of guidewires in interventional cardiac fluoroscopy videos is crucial for computer-aided navigation tasks. Although deep learning methods have demonstrated high accuracy and robustness in wire segmentation, they require substantial annotated datasets for generalizability, u…

Cited by 0SourcePDFScholar
2025

Learning Attribute-Aware Hash Codes for Fine-Grained Image Retrieval via Query Optimization

ICML 2025poster

Fine-grained hashing has become a powerful solution for rapid and efficient image retrieval, particularly in scenarios requiring high discrimination between visually similar categories. To enable each hash bit to correspond to specific visual attributes, we propose a novel method that harnesses lear…

Cited by 0SourcePDFScholar
2025

MARS-FTCP: Robust Fault-Tolerant Control and Agile Trajectory Planning for Modular Aerial Robot Systems

IROS 2025

Modular Aerial Robot Systems (MARS) consist of multiple drone units that can self-reconfigure to adapt to various mission requirements and fault conditions. However, existing fault-tolerant control methods exhibit significant oscillations during docking and separation, impacting system stability. To

Cited by 4SourcecodeScholar
2025

Pre-training a Density-Aware Pose Transformer for Robust LiDAR-based 3D Human Pose Estimation

AAAI 2025technical

With the rapid development of autonomous driving, LiDAR-based 3D Human Pose Estimation (3D HPE) is becoming a research focus. However, due to the noise and sparsity of LiDAR-captured point clouds, robust human pose estimation remains challenging. Most of the existing methods use temporal information…

2025

Prototype-based Contrastive Learning with Stage-wise Progressive Augmentation for Self-Supervised Fine-Grained Learning

ICCV 2025poster

In this paper, we mitigate the problem of Self-Supervised Learning (SSL) for fine-grained representation learning, aimed at distinguishing subtle differences within highly similar subordinate categories. Our preliminary analysis shows that SSL, especially the multi-stage alignment strategy, performs…

2025

Provable Discriminative Hyperspherical Embedding for Out-of-Distribution Detection

AAAI 2025technical

Out-of-distribution (OOD) detection aims to identify the test examples that do not belong to the distribution of training data. The distance-based methods, which identify OOD examples based on their distances from the centroids of in-distribution (ID) examples, have demonstrated promising OOD detect…

2025

Robust Self-Reconfiguration for Fault-Tolerant Control of Modular Aerial Robot Systems

ICRA 2025

Modular Aerial Robotic Systems (MARS) consist of multiple drone units assembled into a single, integrated rigid flying platform. With inherent redundancy, MARS can self-reconfigure into different configurations to mitigate rotor or unit failures and maintain stable flight. However, existing works on

Cited by 9SourcecodeScholar
2025

Taming Diffusion for Dataset Distillation with High Representativeness

ICML 2025poster

Recent deep learning models demand larger datasets, driving the need for dataset distillation to create compact, cost-efficient datasets while maintaining performance. Due to the powerful image generation capability of diffusion, it has been introduced to this field for generating distilled images.…

2024

A Comprehensive Framework for Occluded Human Pose Estimation

ICASSP 2024accepted

Occlusion presents a significant challenge in human pose estimation. The challenges posed by occlusion can be attributed to the following factors: 1) Data: The collection and annotation of occluded human pose samples are relatively challenging. 2) Feature: Occlusion can cause feature confusion due t…

Cited by 0SourceScholar
2024

An Asymmetric Augmented Self-Supervised Learning Method for Unsupervised Fine-Grained Image Hashing

CVPR 2024poster

Unsupervised fine-grained image hashing aims to learn compact binary hash codes in unsupervised settings addressing challenges posed by large-scale datasets and dependence on supervision. In this paper we first identify a granularity gap between generic and fine-grained datasets for unsupervised has…

Cited by 3SourcePDFScholar
2024

FlashEval: Towards Fast and Accurate Evaluation of Text-to-image Diffusion Generative Models

CVPR 2024poster

In recent years there has been significant progress in the development of text-to-image generative models. Evaluating the quality of the generative models is one essential step in the development process. Unfortunately the evaluation process could consume a significant amount of computational resour…

2024

Global Optimality of Single-Timescale Actor-Critic under Continuous State-Action Space: A Study on Linear Quadratic Regulator

IJCAI 2024poster

Actor-critic methods have achieved state-of-the-art performance in various challenging tasks. However, theoretical understandings of their performance remain elusive and challenging. Existing studies mostly focus on practically uncommon variants such as double-loop or two-timescale stepsize actor-cr…

Cited by 1SourcePDFScholar
2024

Long-tailed Object Detection Pretraining: Dynamic Rebalancing Contrastive Learning with Dual Reconstruction

NeurIPS 2024poster

Pre-training plays a vital role in various vision tasks, such as object recognition and detection. Commonly used pre-training methods, which typically rely on randomized approaches like uniform or Gaussian distributions to initialize model parameters, often fall short when confronted with long-taile…

Cited by 1SourcePDFScholar
2024

SHaRPose: Sparse High-Resolution Representation for Human Pose Estimation

AAAI 2024technical

High-resolution representation is essential for achieving good performance in human pose estimation models. To obtain such features, existing works utilize high-resolution input images or fine-grained image tokens. However, this dense high-resolution representation brings a significant computational…

2024

TMFN: A Target-oriented Multi-grained Fusion Network for End-to-end Aspect-based Multimodal Sentiment Analysis

COLING 2024main

End-to-end multimodal aspect-based sentiment analysis (MABSA) combines multimodal aspect terms extraction (MATE) with multimodal aspect sentiment classification (MASC), aiming to simultaneously extract aspect words and classify the sentiment polarity of each aspect. However, existing MABSA methods h…

Cited by 3SourcePDFScholar
2024

Towards Learning a Generalist Model for Embodied Navigation

CVPR 2024highlight

Building a generalist agent that can interact with the world is an ultimate goal for humans thus spurring the research for embodied navigation where an agent is required to navigate according to instructions or respond to queries. Despite the major progress attained previous works primarily focus on…

Cited by 47SourcePDFScholar
2023

Coupling Artificial Neurons in BERT and Biological Neurons in the Human Brain

AAAI 2023technical

Linking computational natural language processing (NLP) models and neural responses to language in the human brain on the one hand facilitates the effort towards disentangling the neural representations underpinning language perception, on the other hand provides neurolinguistics evidence to evaluat…

2023

Deep Dive Into Gradients: Better Optimization for 3D Object Detection With Gradient-Corrected IoU Supervision

CVPR 2023poster

Intersection-over-Union (IoU) is the most popular metric to evaluate regression performance in 3D object detection. Recently, there are also some methods applying IoU to the optimization of 3D bounding box regression. However, we demonstrate through experiments and mathematical proof that the 3D IoU…

2023

Fine-grained Artificial Neurons in Audio-transformers for Disentangling Neural Auditory Encoding

ACL 2023findings

The Wav2Vec and its variants have achieved unprecedented success in computational auditory and speech processing. Meanwhile, neural encoding studies that integrate the superb representation capability of Wav2Vec and link those representations to brain activities have provided novel insights into a f…

2023

Global Convergence of Two-Timescale Actor-Critic for Solving Linear Quadratic Regulator

AAAI 2023technical

The actor-critic (AC) reinforcement learning algorithms have been the powerhouse behind many challenging applications. Nevertheless, its convergence is fragile in general. To study its instability, existing works mostly consider the uncommon double-loop variant or basic models with finite state and…

Cited by 12SourcePDFScholar
2023

Learning Agile Flight Maneuvers: Deep SE(3) Motion Planning and Control for Quadrotors

ICRA 2023poster

Agile flights of autonomous quadrotors in clut-tered environments require constrained motion planning and control subject to translational and rotational dynamics. Tra-ditional model-based methods typically demand complicated design and heavy computation. In this paper, we develop a novel deep reinf…

Cited by 5SourceScholar
2023

Tightly-Coupled Visual- DVL- Inertial Odometry for Robot-Based Ice-Water Boundary Exploration

IROS 2023poster

Underwater robots, like Autonomous Underwater Vehicles (AUVs) and Remotely Operated Vehicles (ROVs), are promising tools for the exploration and study of the under-ice environment and the ecosystems that thrive there. However, state estimation is a well-known problem for robotic systems, especially,…

Cited by 12SourcecodeScholar
2022

Deterministic policy gradient: Convergence analysis

UAI 2022poster

The deterministic policy gradient (DPG) method proposed in Silver et al. [2014] has been demonstrated to exhibit superior performance particularly for applications with multi-dimensional and continuous action spaces. However, it remains unclear whether DPG converges, and if so, how fast it converges…

Cited by 25SourcePDFScholar
2022

Sem-Aug: Improving Camera-LiDAR Feature Fusion With Semantic Augmentation for 3D Vehicle Detection

RA-L 2022

Camera-LiDAR fusion provides precise distance measurements and fine-grained textures, making it a promising option for 3D vehicle detection in autonomous driving scenarios. Previous camera-LiDAR based 3D vehicle detection approaches mainly focused on employing image-based pre-trained models to fetch

Cited by 19SourceScholar
2021

Deep Symmetric Network for Underexposed Image Enhancement With Recurrent Attentional Learning

ICCV 2021poster

Underexposed image enhancement is of importance in many research domains. In this paper, we take this problem as image feature transformation between the underexposed image and its paired enhanced version, and we propose a deep symmetric network for the issue. Our symmetric network adapts invertible…

Cited by 70PDFScholar
2020

Dynamic Object Tracking for Self-Driving Cars Using Monocular Camera and LIDAR

IROS 2020poster

The detection and tracking of dynamic traffic participants (e.g., pedestrians, cars, and bicyclists) plays an important role in reliable decision-making and intelligent navigation for autonomous vehicles. However, due to the rapid movement of the target, most current vision-based tracking methods, w…

Cited by 16SourceScholar
2019

A Neural Network Based Ranking Framework to Improve ASR with NLU Related Knowledge Deployed

ICASSP 2019accepted

This work proposes a new neural network framework to simultaneously rank multiple hypotheses generated by one or more automatic speech recognition (ASR) engines for a speech utterance. Features fed in the framework not only include those calculated from the ASR information, but also involve natural…

Cited by 0SourceScholar