← Search

Yong Zhao

42 accepted papers

2026

$A_2$DEPT: Large Language Model–Driven Automated Algorithm Design via Evolutionary Program Trees

ICML 2026poster

Designing heuristics for combinatorial optimization problems (COPs) is a fundamental yet challenging task that traditionally requires extensive domain expertise. Recently, Large Language Model (LLM)-based Automated Heuristic Design (AHD) has shown promise in autonomously generating heuristic compone…

Cited by 0SourceScholar
2026

Beyond the Mean: Gaussian Distributional Successor Features for Zero-Shot Non-Linear Reward Adaptation

IJCAI 2026

Zero-shot transfer for offline reinforcement learning involves generalizing to a wide range of tasks that are often risk-sensitive, without new interaction. We show here that although Successor Features (SFs) provide a principled framework for transfer through abstracting dynamics from rewards, thei

Cited by 0Scholar
2026

Monocular Localization With Vector HD Maps Using Geometric-Context Data Association

RA-L 2026

Due to their low cost and wide field of view, monocular cameras hold great promise for visual localization. Nevertheless, significant challenges persist in associating detected semantic features with map landmarks, arising from sensing noise, depth cues, and dynamic driving scenes. Most existing app

Cited by 0SourceScholar
2026

Towards Autonomous UAV Visual Object Search in City Space: Benchmark and Agentic Methodology

AAAI 2026technical

Aerial Visual Object Search (AVOS) tasks in urban environments require Unmanned Aerial Vehicles (UAVs) to autonomously search for and identify target objects based on visual inputs without external guidance. Existing approaches struggle in complex urban environments due to redundant semantic process

Cited by 0SourcePDFScholar
2025

Aligning Large Language Models for Faithful Integrity Against Opposing Argument

AAAI 2025technical

Large Language Models (LLMs) have demonstrated impressive capabilities in complex reasoning tasks. However, they can be easily misled by unfaithful arguments during conversations, even when their original statements are correct. To this end, we investigate the problem of maintaining faithful integri…

2025

CityEQA: A Hierarchical LLM Agent on Embodied Question Answering Benchmark in City Space

EMNLP 2025

Embodied Question Answering (EQA) has primarily focused on indoor environments, leaving the complexities of urban settings—spanning environment, action, and perception—largely unexplored. To bridge this gap, we introduce CityEQA, a new task where an embodied agent answers open-vocabulary questions t

2025

Knowledge Boundary of Large Language Models: A Survey

ACL 2025long

Although large language models (LLMs) store vast amount of knowledge in their parameters, they still have limitations in the memorization and utilization of certain knowledge, leading to undesired behaviors such as generating untruthful and inaccurate responses. This highlights the critical need to…

2025

RetinaStereo: Dynamic-Volume Stereo Matching Network

ICASSP 2025accepted

Existing stereo matching techniques often struggle with detailing subtle objects on depth edges. To alleviate this problem, we introduced the Dynamic-Range Disparity Initialization module, which integrates three complementary branches: the dynamic dense volume for localized disparity sampling, the s…

Cited by 0SourceScholar
2025

VIP: Vision Instructed Pre-training for Robotic Manipulation

ICML 2025poster

The effectiveness of scaling up training data in robotic manipulation is still limited. A primary challenge in manipulation is the tasks are diverse, and the trained policy would be confused if the task targets are not specified clearly. Existing works primarily rely on text instruction to describe…

Cited by 0SourcePDFScholar
2024

Don’t Just Say “I don’t know”! Self-aligning Large Language Models for Responding to Unknown Questions with Explanations

EMNLP 2024main

Despite the remarkable abilities of Large Language Models (LLMs) to answer questions, they often display a considerable level of overconfidence even when the question does not have a definitive answer. To avoid providing hallucinated answers to these unknown questions, existing studies typically inv…

2024

Efficient Fusion of Depth Information for Defocus Deblurring

ICASSP 2024accepted

Defocus deblurring is a classic problem in image restoration tasks. The formation of its defocus blur is related to depth. Recently, the use of dual-pixel sensor designed according to depth-disparity characteristics has brought great improvements to the defocus deblurring task. However, the difficul…

Cited by 0SourceScholar
2024

Robust Visual Localization System With HD Map Based on Joint Probabilistic Data Association

RA-L 2024

Localization based on a high-definition (HD) map is a pivotal technology for autonomous driving. Nonetheless, establishing precise data association (DA) between detected landmarks and map landmarks presents a formidable challenge when leveraging prior information on maps. Traditional DA algorithms r

Cited by 8SourceScholar
2023

High-Frequency Stereo Matching Network

CVPR 2023highlight

In the field of binocular stereo matching, remarkable progress has been made by iterative methods like RAFT-Stereo and CREStereo. However, most of these methods lose information during the iterative process, making it difficult to generate more detailed difference maps that take full advantage of hi…

Cited by 86SourcePDFScholar
2023

Improving Transformer-Based Networks with Locality for Automatic Speaker Verification

ICASSP 2023accepted

Recently, Transformer-based architectures have been explored for speaker embedding extraction. Although the Transformer employs the self-attention mechanism to efficiently model the global interaction between token embeddings, it is inadequate for capturing short-range local context, which is essent…

Cited by 17SourceScholar
2023

Learnable Blur Kernel for Single-Image Defocus Deblurring in the Wild

AAAI 2023technical

Recent research showed that the dual-pixel sensor has made great progress in defocus map estimation and image defocus deblurring. However, extracting real-time dual-pixel views is troublesome and complex in algorithm deployment. Moreover, the deblurred image generated by the defocus deblurring netwo…

Cited by 6SourcePDFScholar
2022

CRMRL: Collaborative Relationship Meta Reinforcement Learning for Effectively Adapting to Type Changes in Multi-Robotic System

RA-L 2022

Multi-agent reinforcement learning methods have been widely used for multi-robotic systems, and meta-learning methods are also applied to help robots reuse prior experiences to guide new tasks learning. But in some multi-robotic tasks, the robot types cannot be determined in advance or may dynamical

Cited by 7SourceScholar
2021

A Unified Multi-Scenario Attacking Network for Visual Object Tracking

AAAI 2021technical

Existing methods of adversarial attacks successfully generate adversarial examples to confuse Deep Neural Networks (DNNs) of image classification and object detection, resulting in wrong predictions. However, these methods are difficult to attack models of video object tracking, because the tracking…

Cited by 19SourcePDFScholar
2021

Decentralized Multi-Robot Collision Avoidance in Complex Scenarios With Selective Communication

RA-L 2021

Deep reinforcement learning has been demonstrated to be an effective solution to the multi-robot collision avoidance problem. However, with existing methods, robots typically generate actions only based on local observations, sometimes augmented with global communication. Their performance deteriora

Cited by 28SourceScholar
2021

Hierarchical Context Guided Aggregation Network for Stereo Matching

ICASSP 2021accepted

Nowadays, CNN-based stereo matching methods achieved remarkable performance, and how to efficiently exploit contextual information in cost aggregation stage is the key to improve performance. In this paper, we propose a simple yet efficient network named Hierarchical Context Guided Aggregation Netwo…

Cited by 0SourceScholar
2021

Microsoft Speaker Diarization System for the Voxceleb Speaker Recognition Challenge 2020

ICASSP 2021accepted

This paper describes the Microsoft speaker diarization system for monaural multi-talker recordings in the wild, evaluated at the diarization track of the VoxCeleb Speaker Recognition Challenge (VoxSRC) 2020. We will first explain our system design to address issues in handling real multi-talker reco…

Cited by 0SourceScholar
2020

DenseFusion: Large-Scale Online Dense Pointcloud and DSM Mapping for UAVs

IROS 2020poster

With the rapidly developing unmanned aerial vehicles, the requirements of generating maps efficiently and quickly are increasing. To realize online mapping, we develop a real-time dense mapping framework named DenseFusion which can incrementally generates dense geo-referenced 3D point cloud, digital…

Cited by 11SourceScholar
2020

Hijacking Tracker: A Powerful Adversarial Attack on Visual Tracking

ICASSP 2020accepted

Visual object tracking has made important breakthroughs with the assistance of deep learning models. Unfortunately, recent research has clearly proved that deep learning models are vulnerable to malicious adversarial attacks, which mislead the models making wrong decisions by perturbing the input im…

Cited by 0SourceScholar
2020

Improving Deep CNN Networks with Long Temporal Context for Text-Independent Speaker Verification

ICASSP 2020accepted

Deep CNN networks have shown great success in various tasks for text-independent speaker recognition. In this paper, we explore two approaches for modeling long temporal contexts to improve the performance of the ResNet networks. The first approach is simply integrating the utterance-level mean and…

Cited by 0SourceScholar
2020

Non-Local Nested Residual Attention Network for Stereo Image Super-Resolution

ICASSP 2020accepted

Nowadays CNN-based stereo image super-resolution(SR) methods have obtained remarkable performance. However, most of existing methods only superficially portrayed the low layer features without considering the uneven distribution of information, which is insufficient because stereo image warping and…

Cited by 0SourceScholar
2020

One-Shot Adversarial Attacks on Visual Tracking With Dual Attention

CVPR 2020poster

Almost all adversarial attacks in computer vision are aimed at pre-known object categories, which could be offline trained for generating perturbations. But as for visual object tracking, the tracked target categories are normally unknown in advance. However, the tracking algorithms also have potent…

Cited by 100PDFScholar
2020

Salience-Guided Cascaded Suppression Network for Person Re-Identification

CVPR 2020poster

Employing attention mechanisms to model both global and local features as a final pedestrian representation has become a trend for person re-identification (Re-ID) algorithms. A potential limitation of these methods is that they focus on the most salient features, but the re-identification of a pers…

Cited by 307PDFScholar
2019

Discriminative Features Reconstruction Network for Semantic Segmentation

ICASSP 2019accepted

Thanks to the development of convolutional neural networks (CNNs), researchers have proposed lots of effective semantic segmentation models. However, there are still two problems disturbing researchers, one of which is objects misidentification on the image level and another one is poor performance…

Cited by 0SourceScholar
2019

Multi-attention Network for Thoracic Disease Classification and Localization

ICASSP 2019accepted

The chest X-ray is one of the most commonly available radiological examinations for diagnosing lung diseases. This task remains a major challenge due to 1) the shortage of accurate annotations for chest X-ray examinations, 2) the diversity of lesion areas on X-rays from different thoracic disease an…

Cited by 0SourceScholar
2019

Non-Local Recurrent Neural Memory for Supervised Sequence Modeling

ICCV 2019oral

Typical methods for supervised sequence modeling are built upon the recurrent neural networks to capture temporal dependencies. One potential limitation of these methods is that they only model explicitly information interactions between adjacent time steps in a sequence, hence the high-order intera…

Cited by 13PDFcodeScholar
2019

Scanet: Spatial-channel Attention Network for 3D Object Detection

ICASSP 2019accepted

This paper aims to achieve high-accuracy 3D object detection, in which we propose a novel Spatial-Channel Attention Network (SCANet), a two-stage detector that takes both LIDAR point clouds and RGB images as input to generate 3D object estimates. The first stage is a 3D region proposal network (RPN)…

Cited by 0SourceScholar
2019

TerrainFusion: Real-time Digital Surface Model Reconstruction based on Monocular SLAM

IROS 2019poster

This paper presents an algorithm which can generate live digtial surface model (DSM) during the flight based on simultaneous localization and mapping (SLAM). We process the keyframe which is output by a monocular SLAM system to generate a local DSM, and fuse the local DSM to the global tiled DSM inc…

Cited by 19SourceScholar
2018

Depth Super-Resolution Using Joint Adaptive Weighted Least Squares And Patching Gradient

ICASSP 2018accepted

This paper presents a flexible framework for the challenging task of color-guided depth upsampling. Some state-of-the-art approaches apply an aligned RGB image for depth recovery. Unfortunately, these kinds of methods may result in texture copying artifacts and edge blurring artifacts. To address th…

Cited by 0SourceScholar
2018

Domain and Speaker Adaptation for Cortana Speech Recognition

ICASSP 2018accepted

Voice assistant represents one of the most popular and important scenarios for speech recognition. In this paper, we propose two adaptation approaches to customize a multi-style well-trained acoustic model towards its subsidiary domain of Cortana assistant. First, we present anchor-based speaker ada…

Cited by 0SourceScholar
2018

Exploring Sequential Characteristics in Speaker Bottleneck Feature for Text-Dependent Speaker Verification

ICASSP 2018accepted

In this paper, given the speaker bottleneck feature vectors extracted with speaker discriminant neural networks, we focus on using the sequential speaker characteristics for text-dependent speaker verification. In each evaluation trial, speaker supervectors are used as the representations of the seq…

Cited by 0SourceScholar
2018

Hard Shadows Removal Using an Approximate Illumination Invariant

ICASSP 2018accepted

Hard shadows detection and removal from foreground masks is a challenging step in change detection. This paper gives a simple and effective method to address hard shadows. There are inside portion and boundary portion in hard shadows. Pixel-wise neighborhood ratio is calculated to remove the most of…

Cited by 0SourceScholar
2017

Extended low-rank plus diagonal adaptation for deep and recurrent neural networks

ICASSP 2017accepted

Recently, the low-rank plus diagonal (LRPD) adaptation was proposed for speaker adaptation of deep neural network (DNN) models. The LRPD restructures the adaptation matrix as a superposition of a diagonal matrix and a product of two low-rank matrices. In this paper, we extend the LRPD adaptation int…

Cited by 0SourceScholar
2016

Map2DFusion: Real-time incremental UAV image mosaicing based on monocular SLAM

IROS 2016poster

In this paper we present a real-time approach to stitch large-scale aerial images incrementally. A monocular SLAM system is used to estimate camera position and attitude, and meanwhile 3D point cloud map is generated. When GPS information is available, the estimated trajectory is transformed to WGS8…

Cited by 96SourceScholar
2015

Investigating online low-footprint speaker adaptation using generalized linear regression and click-through data

ICASSP 2015accepted

To develop speaker adaptation algorithms for deep neural network (DNN) that are suitable for large-scale online deployment, it is desirable that the adaptation model be represented in a compact form and learned in an unsupervised fashion. In this paper, we propose a novel low-footprint adaptation te…

Cited by 0SourceScholar