← Search

Nana Hou

7 accepted papers

2026

ALIGNING GENERATIVE SPEECH ENHANCEMENT WITH PERCEPTUAL FEEDBACK

ICASSP 2026oral

Language Model (LM)-based speech enhancement (SE) has recently emerged as a promising direction, but existing approaches predominantly rely on token-level likelihood objectives that weakly reflect human perception. This mismatch limits progress, as optimizing signal accuracy does not always improve…

Cited by 0SourcePDFScholar
2022

Interactive Feature Fusion for End-to-End Noise-Robust Speech Recognition

ICASSP 2022accepted

Speech enhancement (SE) aims to suppress the additive noise from noisy speech signals to improve the speech’s perceptual quality and intelligibility. However, the over-suppression phenomenon in the enhanced speech might degrade the performance of downstream automatic speech recognition (ASR) task du…

Cited by 0SourceScholar
2022

Noise-Robust Speech Recognition With 10 Minutes Unparalleled In-Domain Data

ICASSP 2022accepted

Noise-robust speech recognition systems require large amounts of training data including noisy speech data and corresponding transcripts to achieve state-of-the-art performances in face of various practical environments. However, such plenty of in-domain data is not always available in the real-life…

Cited by 0SourceScholar
2022

Self-Critical Sequence Training for Automatic Speech Recognition

ICASSP 2022accepted

Although automatic speech recognition (ASR) task has gained remarkable success by sequence-to-sequence models, there are two main mismatches between its training and testing that might lead to performance degradation: 1) The typically used cross-entropy criterion aims to maximize log-likelihood of t…

Cited by 0SourceScholar
2021

Learning Disentangled Feature Representations for Speech Enhancement Via Adversarial Training

ICASSP 2021accepted

Neural speech enhancement degrades significantly in face of unseen noise. To address such mismatch, we propose to learn noise-agnostic feature representations by disentanglement learning, which removes the unspecified noise factor, while keeping the specified factors of variation associated with the…

Cited by 0SourceScholar
2020

Time-Domain Neural Network Approach for Speech Bandwidth Extension

ICASSP 2020accepted

In this paper, we study the time-domain neural network approach for speech bandwidth extension. We propose a network architecture, named multi-scale fusion neural network (MfNet), that gradually restores the low-frequency signal and predicts the high-frequency signal through the exchange of informat…

Cited by 0SourceScholar