← Search

Zheng-Hua Tan

32 accepted papers

2026

A TEXT-TO-TEXT ALIGNMENT ALGORITHM FOR BETTER EVALUATION OF MODERN SPEECH RECOGNITION SYSTEMS

ICASSP 2026poster

Modern neural networks have greatly improved performance across speech recognition benchmarks. However, gains are often driven by frequent words with limited semantic weight, which can obscure meaningful differences in word error rate, the primary evaluation metric. Errors in rare terms, named entit…

Cited by 0SourcePDFScholar
2026

BiSSL: Enhancing the Alignment Between Self-Supervised Pretraining and Downstream Fine-Tuning via Bilevel Optimization

ICML 2026poster

Models initialized from self-supervised pretraining may suffer from poor alignment with downstream tasks, limiting the extent to which subsequent fine-tuning can adapt relevant representations acquired during the pretraining phase. To mitigate this, we introduce BiSSL, a novel bilevel training frame…

Cited by 2SourceScholar
2026

QUANTIZATION-BASED SCORE CALIBRATION FOR FEW-SHOT KEYWORD SPOTTING WITH DYNAMIC TIME WARPING IN NOISY ENVIRONMENTS

ICASSP 2026poster

Detecting occurrences of keywords with keyword spotting (KWS) systems requires thresholding continuous detection scores. Selecting appropriate thresholds is a non-trivial task, typically relying on optimizing performance on a validation dataset. However, such greedy threshold selection often leads t…

Cited by 0SourcePDFScholar
2025

Deep Feedback Cancellation for Hearing Aids with Improved System Stability and Sound Quality

ICASSP 2025accepted

Acoustic feedback cancellation is an important task in audio processing systems, aiming to mitigate the effects of feedback loops on system stability and sound quality. State-of-the-art methods rely on adaptive filtering algorithms and face challenges in balancing between rapid convergence and low s…

Cited by 0SourceScholar
2025

Detecting and Defending Against Adversarial Attacks on Automatic Speech Recognition via Diffusion Models

ICASSP 2025accepted

Automatic speech recognition (ASR) systems are known to be vulnerable to adversarial attacks. This paper addresses detection and defence against targeted white-box attacks on speech signals for ASR systems. While existing work has utilised diffusion models (DMs) to purify adversarial examples, achie…

Cited by 0SourceScholar
2025

Optimal Sensor Scheduling and Selection for Continuous-Discrete Kalman Filtering with Auxiliary Dynamics

ICML 2025poster

We study the Continuous-Discrete Kalman Filter (CD-KF) for State-Space Models (SSMs) where continuous-time dynamics are observed via multiple sensors with discrete, irregularly timed measurements. Our focus extends to scenarios in which the measurement process is coupled with the states of an auxili…

2024

Diffusion-Based Speech Enhancement in Matched and Mismatched Conditions Using a Heun-Based Sampler

ICASSP 2024accepted

Diffusion models are a new class of generative models that have recently been applied to speech enhancement successfully. Previous works have demonstrated their superior performance in mismatched conditions compared to state-of-the art discriminative models. However, this was investigated with a sin…

Cited by 0SourceScholar
2024

Masked Autoencoders with Multi-Window Local-Global Attention Are Better Audio Learners

ICLR 2024poster

In this work, we propose a Multi-Window Masked Autoencoder (MW-MAE) fitted with a novel Multi-Window Multi-Head Attention (MW-MHA) module that facilitates the modelling of local-global interactions in every decoder transformer block through attention heads of several distinct local and global window…

Cited by 4SourcePDFScholar
2024

PAC-Bayes Generalisation Bounds for Dynamical Systems including Stable RNNs

AAAI 2024technical

In this paper, we derive a PAC-Bayes bound on the generalisation gap, in a supervised time-series setting for a special class of discrete-time non-linear dynamical systems. This class includes stable recurrent neural networks (RNN), and the motivation for this work was its application to RNNs. In or…

2024

PAC-Bayesian Error Bound, via Rényi Divergence, for a Class of Linear Time-Invariant State-Space Models

ICML 2024poster

In this paper we derive a PAC-Bayesian error bound for a class of stochastic dynamical systems with inputs, namely, for linear time-invariant stochastic state-space models (stochastic LTI systems for short). This class of systems is widely used in control engineering and econometrics, in particular,…

Cited by 1SourcePDFScholar
2024

Self-Supervised Pretraining for Robust Personalized Voice Activity Detection in Adverse Conditions

ICASSP 2024accepted

In this paper, we propose the use of self-supervised pretraining on a large unlabelled data set to improve the performance of a personalized voice activity detection (VAD) model in adverse conditions. We pretrain a long short-term memory (LSTM)-encoder using the autoregressive predictive coding (APC…

Cited by 0SourceScholar
2023

Filterbank Learning for Noise-Robust Small-Footprint Keyword Spotting

ICASSP 2023accepted

In the context of keyword spotting (KWS), the replacement of handcrafted speech features by learnable features has not yielded superior KWS performance. In this study, we demonstrate that filterbank learning outperforms handcrafted speech features for KWS whenever the number of filterbank channels i…

Cited by 0SourceScholar
2023

Radio Sensing with Large Intelligent Surface for 6G

ICASSP 2023accepted

This paper leverages the potential of Large Intelligent Surfaces (LIS) for radio sensing in 6G wireless networks. By taking advantage of arbitrary communication signals occurring in the scenario, we apply direct processing to the output signal from the LIS to obtain a radio map that describes the ph…

Cited by 0SourceScholar
2022

Joint Far- and Near-End Speech Intelligibility Enhancement Based on the Approximated Speech Intelligibility Index

ICASSP 2022accepted

This paper considers speech enhancement of signals picked up in one noisy environment which must be presented to a listener in another noisy environment. Recently, it has been shown that an optimal solution to this problem requires the consideration of the noise sources in both environments jointly.…

Cited by 0SourceScholar
2022

Summary on the ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Grand Challenge

ICASSP 2022accepted

The ICASSP 2022 Multi-channel Multi-party Meeting Transcription Grand Challenge (M2MeT) focuses on one of the most valuable and the most challenging scenarios of speech technologies. The M2MeT challenge has particularly set up two tracks, speaker diarization (track 1) and multi-speaker automatic spe…

Cited by 0SourceScholar
2021

Audio-Visual Speech Inpainting with Deep Learning

ICASSP 2021accepted

In this paper, we present a deep-learning-based framework for audio-visual speech inpainting, i.e., the task of restoring the missing parts of an acoustic speech signal from reliable audio context and uncorrupted visual information. Recent work focuses solely on audio-only methods and generally aims…

Cited by 31SourceScholar
2021

Joint Maximum Likelihood Estimation of Power Spectral Densities and Relative Acoustic Transfer Functions for Acoustic Beamforming

ICASSP 2021accepted

Acoustic beamforming is crucial for many applications where ex-traction of a target signal from a noisy environment is required. In order to implement practical beamformers, e.g. the multichannel Wiener filter (MWF), estimation of the target and noise power spectral densities (PSDs), and the relativ…

Cited by 8SourceScholar
2020

Adversarial Example Detection by Classification for Deep Speech Recognition

ICASSP 2020accepted

Machine Learning systems are vulnerable to adversarial attacks and will highly likely produce incorrect outputs under these attacks. There are white-box and black-box attacks regarding to adversary's access level to the victim learning algorithm. To defend the learning systems from these attacks, ex…

Cited by 0SourceScholar
2020

Maximum Likelihood Estimation of the Interference-Plus-Noise Cross Power Spectral Density Matrix for Own Voice Retrieval

ICASSP 2020accepted

In headset and hearing aid applications, it is of interest to retrieve the user's own voice in a noisy environment, e.g. for telephony applications. To do so, the cross-power spectral density (CPSD) of the interference-plus-noise is required. In this paper, a novel maximum likelihood (ML) estimator…

Cited by 0SourceScholar
2019

Effects of Lombard Reflex on the Performance of Deep-learning-based Audio-visual Speech Enhancement Systems

ICASSP 2019accepted

Humans tend to change their way of speaking when they are immersed in a noisy environment, a reflex known as Lombard effect. Current speech enhancement systems based on deep learning do not usually take into account this change in the speaking style, because they are trained with neutral (non-Lombar…

Cited by 0SourceScholar
2019

On Training Targets and Objective Functions for Deep-learning-based Audio-visual Speech Enhancement

ICASSP 2019accepted

Audio-visual speech enhancement (AV-SE) is the task of improving speech quality and intelligibility in a noisy environment using audio and visual information from a talker. Recently, deep learning techniques have been adopted to solve the AV-SE task in a supervised manner. In this context, the choic…

Cited by 0SourceScholar
2018

Monaural Speech Enhancement Using Deep Neural Networks by Maximizing a Short-Time Objective Intelligibility Measure

ICASSP 2018accepted

In this paper we propose a Deep Neural Network (D NN) based Speech Enhancement (SE) system that is designed to maximize an approximation of the Short-Time Objective Intelligibility (STOI) measure. We formalize an approximate-STOI cost function and derive analytical expressions for the gradients requ…

Cited by 0SourceScholar
2017

A non-intrusive Short-Time Objective Intelligibility measure

ICASSP 2017accepted

We propose a non-intrusive intelligibility measure for noisy and non-linearly processed speech, i.e. a measure which can predict intelligibility from a degraded speech signal without requiring a clean reference signal. The proposed measure is based on the Short-Time Objective Intelligibility (STOI)…

Cited by 0SourceScholar
2017

Permutation invariant training of deep models for speaker-independent multi-talker speech separation

ICASSP 2017accepted

We propose a novel deep learning training criterion, named permutation invariant training (PIT), for speaker independent multi-talker speech separation, commonly known as the cocktail-party problem. Different from the multi-class regression technique and the deep clustering (DPCL) technique, our nov…

Cited by 0SourceScholar
2017

RedDots replayed: A new replay spoofing attack corpus for text-dependent speaker verification research

ICASSP 2017accepted

This paper describes a new database for the assessment of automatic speaker verification (ASV) vulnerabilities to spoofing attacks. In contrast to other recent data collection efforts, the new database has been designed to support the development of replay spoofing countermeasures tailored towards t…

Cited by 123SourceScholar
2016

A method for predicting the intelligibility of noisy and non-linearly enhanced binaural speech

ICASSP 2016accepted

We propose and evaluate a binaural speech intelligibility measure. The measure is a binaural extension of the Short-Time Objective Intelligibility (STOI) measure and focuses on predicting the intelligibility of noisy speech which has been enhanced by a speech processing algorithm (e.g. in a hearing…

Cited by 0SourceScholar
2016

Informed Direction of Arrival estimation using a spherical-head model for Hearing Aid applications

ICASSP 2016accepted

In this paper, we propose a Direction of Arrival (DoA) estimator for a Hearing Aid System (HAS) which can connect to a wireless microphone worn by a target talker. The wireless microphone "informs" the HAS about the almost noise-free content of the target sound, and the proposed DoA estimator uses t…

Cited by 7SourceScholar
2015

A heuristic approach for a social robot to navigate to a person based on audio and range information

IROS 2015poster

The use of social robots for elderly care is becoming ever more relevant, thus introducing new challenges which need to be solved to achieve acceptable performance. One fundamental task for a social robot is to move to the person of interest in order to start interacting or perform a service. In thi…

Cited by 8SourceScholar
2015

Maximum likelihood approach to "informed" Sound Source Localization for Hearing Aid applications

ICASSP 2015accepted

Most state-of-the-art Sound Source Localization (SSL) algorithms have been proposed for applications which are “uninformed” about the target sound content; however, utilizing a wireless microphone worn by a target talker, enables recent Hearing Aid Systems (HASs) to access to an almost noise-free so…

Cited by 0SourceScholar
2015

On the influence of microphone array geometry on HRTF-based Sound Source Localization

ICASSP 2015accepted

The direction dependence of Head Related Transfer Functions (HRTFs) forms the basis for HRTF-based Sound Source Localization (SSL) algorithms. In this paper, we show how spectral similarities of the HRTFs of different directions in the horizontal plane influence performance of HRTF-based SSL algorit…

Cited by 0SourceScholar
2015

Source-specific informative prior for i-vector extraction

ICASSP 2015accepted

An i-vector is a low-dimensional fixed-length representation of a variable-length speech utterance, and is defined as the posterior mean of a latent variable conditioned on the observed feature sequence of an utterance. The assumption is that the prior for the latent variable is non-informative, sin…

Cited by 0SourceScholar