← Search

Yanhui Tu

7 accepted papers

2024

A Spatial Long-Term Iterative Mask Estimation Approach for Multi-Channel Speaker Diarization and Speech Recognition

ICASSP 2024accepted

Deep learning (DL)-based speaker diarization methods have proven powerful performance comparing to traditional clustering-based methods for multi-talker speech diarization and recognition in farfield scenes. However, most DL-based approaches cannot utilize the spatial information well due to the poo…

Cited by 0SourceScholar
2020

2D-to-2D Mask Estimation for Speech Enhancement Based on Fully Convolutional Neural Network

ICASSP 2020accepted

In recent years, the deep learning-based approaches are popular in the field of singe-channel speech enhancement. Convolutional neural networks (CNNs) are a standard component of many current speech enhancement system. In this study, we design a new Fully CNN (FCNN)-based regression model, which can…

Cited by 6SourceScholar
2020

High-Resolution Attention Network with Acoustic Segment Model for Acoustic Scene Classification

ICASSP 2020accepted

The spectral information of acoustic scenes is diverse and complex, which poses challenges for acoustic scene tasks. To improve the classification performance, a variety of convolutional neural networks (CNNs) are proposed to extract richer semantic information of scene utterances. However, the diff…

Cited by 0SourceScholar
2019

DNN Training Based on Classic Gain Function for Single-channel Speech Enhancement and Recognition

ICASSP 2019accepted

For conventional single-channel speech enhancement based on noise power spectrum, the speech gain function, which suppresses background noise at each time-frequency bin, is calculated by prior signal-to-noise-ratio (SNR). Hence, accurate prior SNR estimation is paramount for successful noise suppres…

Cited by 0SourceScholar
2018

A Hybrid Approach to Combining Conventional and Deep Learning Techniques for Single-Channel Speech Enhancement and Recognition

ICASSP 2018accepted

Conventional speech-enhancement techniques employ statistical signal-processing algorithms. They are computationally efficient and improve speech quality even under unknown noise conditions. For these reasons, they are preferred for deployment in unpredictable environments. One limitation of these a…

Cited by 0SourceScholar
2015

Speech Separation based on signal-noise-dependent deep neural networks for robust speech recognition

ICASSP 2015accepted

In this paper, we propose a new signal-noise-dependent (SND) deep neural network (DNN) framework to further improve the separation and recognition performance of the recently developed technique for general DNN-based speech separation. We adopt a divide and conquer strategy to design the proposed SN…

Cited by 0SourceScholar