← Search

Ke Tan

13 accepted papers

2025

Reexamining the Efficacy of MetricGAN for Speech Enhancement

ICASSP 2025accepted

MetricGAN, a notable generative approach, provides an effective framework to train speech enhancement models to produce high metric scores. However, we identify two key limitations of current MetricGAN-family models, i.e. neglecting certain mainstream metrics during evaluation and conducting evaluat…

Cited by 0SourceScholar
2024

A Closer Look at Wav2vec2 Embeddings for On-Device Single-Channel Speech Enhancement

ICASSP 2024accepted

Self-supervised learned models have been found to be very effective for tasks such as automatic speech recognition, speaker identification, and others. However, their utility in speech enhancement systems is yet to be firmly established, and perhaps slightly misunderstood. In this paper, we investig…

Cited by 0SourceScholar
2024

Audiovisual Speaker Separation with Full- and Sub-Band Modeling in the Time-Frequency Domain

ICASSP 2024accepted

We introduce a new deep learning model for talker-independent audiovisual speaker separation in noisy conditions in the time-frequency domain. The inputs to the model include noisy multi-talker mixtures and the corresponding cropped face images. Our approach incorporates cross-attention audiovisual…

Cited by 0SourceScholar
2023

Leveraging Heteroscedastic Uncertainty in Learning Complex Spectral Mapping for Single-Channel Speech Enhancement

ICASSP 2023accepted

Most speech enhancement (SE) models learn a point estimate and do not make use of uncertainty estimation in the learning process. In this paper, we show that modeling heteroscedastic uncertainty by minimizing a multivariate Gaussian negative log-likelihood (NLL) improves SE performance at no extra c…

Cited by 2SourceScholar
2023

Torchaudio-Squim: Reference-Less Speech Quality and Intelligibility Measures in Torchaudio

ICASSP 2023accepted

Measuring quality and intelligibility of a speech signal is usually a critical step in development of speech processing systems. To enable this, a variety of metrics to measure quality and intelligibility under different assumptions have been developed. Through this paper, we introduce tools and a s…

Cited by 124SourceScholar
2022

Location-Based Training for Multi-Channel Talker-Independent Speaker Separation

ICASSP 2022accepted

Permutation-invariant training (PIT) is a dominant approach for addressing the permutation ambiguity problem in talker-independent speaker separation. Leveraging spatial information afforded by microphone arrays, we propose a new training approach to resolving permutation ambiguities for multi-chann…

Cited by 0SourceScholar
2021

Real-Time Speech Enhancement for Mobile Communication Based on Dual-Channel Complex Spectral Mapping

ICASSP 2021accepted

Speech quality and intelligibility can be severely degraded by back-ground noise in mobile communication. In order to attenuate back-ground noise, speech enhancement systems have been integrated into mobile phones, and a microphone array is typically deployed to improve the enhancement performance.…

Cited by 4SourceScholar
2020

Improving Robustness of Deep Learning Based Monaural Speech Enhancement Against Processing Artifacts

ICASSP 2020accepted

In voice telecommunication, the intelligibility and quality of speech signals can be severely degraded by background noise if the speaker at the transmitting end talks in a noisy environment. Therefore, a speech enhancement system is typically integrated into the transmitter device or the receiver d…

Cited by 12SourceScholar
2019

Complex Spectral Mapping with a Convolutional Recurrent Network for Monaural Speech Enhancement

ICASSP 2019accepted

Phase is important for perceptual quality in speech enhancement. However, it seems intractable to directly estimate phase spectrogram through supervised learning due to lack of clear structure in phase spectrogram. Complex spectral mapping aims to estimate the real and imaginary spectrograms of clea…

Cited by 0SourceScholar
2019

Deep Learning Based Phase Reconstruction for Speaker Separation: A Trigonometric Perspective

ICASSP 2019accepted

This study investigates phase reconstruction for deep learning based monaural talker-independent speaker separation in the short-time Fourier transform (STFT) domain. The key observation is that, for a mixture oftwo sources, with their magnitudes accurately estimated and under a geometric constraint…

Cited by 0SourceScholar
2019

Real-time Speech Enhancement Using an Efficient Convolutional Recurrent Network for Dual-microphone Mobile Phones in Close-talk Scenarios

ICASSP 2019accepted

In mobile speech communication, the quality and intelligibility of the received speech can be severely degraded by background noise if the far-end talker is in an adverse acoustic environment. Therefore, speech enhancement algorithms are typically integrated into mobile phones to remove background n…

Cited by 0SourceScholar
2018

Gated Residual Networks with Dilated Convolutions for Supervised Speech Separation

ICASSP 2018accepted

In supervised speech separation, deep neural networks (DNNs) are typically employed to predict an ideal time-frequency (T-F) mask in order to remove background interference. However, the performance of DNNs is frequently degraded for untrained noises and speakers. Inspired by recent research on dila…

Cited by 0SourceScholar