← Search

Hu Hu

8 accepted papers

2022

A Variational Bayesian Approach to Learning Latent Variables for Acoustic Knowledge Transfer

ICASSP 2022accepted

We propose a variational Bayesian (VB) approach to learning distributions of latent variables in deep neural network (DNN) models for cross-domain knowledge transfer, to address acoustic mismatches between training and testing conditions. Instead of carrying out point estimation in conventional maxi…

Cited by 0SourceScholar
2021

A Two-Stage Approach to Device-Robust Acoustic Scene Classification

ICASSP 2021accepted

To improve device robustness, a highly desirable key feature of a competitive data-driven acoustic scene classification (ASC) system, a novel two-stage system based on fully convolutional neural networks (CNNs) is proposed. Our two-stage system leverages on an ad-hoc score combination based on two C…

Cited by 0SourceScholar
2021

REDAT: Accent-Invariant Representation for End-To-End ASR by Domain Adversarial Training with Relabeling

ICASSP 2021accepted

Accents mismatching is a critical problem for end-to-end ASR. This paper aims to address this problem by building an accent-robust RNN-T system with domain adversarial training (DAT). We unveil the magic behind DAT and provide, for the first time, a theoretical guarantee that DAT learns accent-invar…

Cited by 0SourceScholar
2020

Exploring Pre-Training with Alignments for RNN Transducer Based End-to-End Speech Recognition

ICASSP 2020accepted

Recently, the recurrent neural network transducer (RNN-T) architecture has become an emerging trend in end-to-end automatic speech recognition research due to its advantages of being capable for online streaming speech recognition. However, RNN-T training is made difficult by the huge memory require…

Cited by 0SourceScholar
2020

L-Vector: Neural Label Embedding for Domain Adaptation

ICASSP 2020accepted

We propose a novel neural label embedding (NLE) scheme for the domain adaptation of a deep neural network (DNN) acoustic model with unpaired data samples from source and target domains. With NLE method, we distill the knowledge from a powerful source-domain DNN into a dictionary of label embeddings,…

Cited by 0SourceScholar
2020

Tensor-To-Vector Regression for Multi-Channel Speech Enhancement Based on Tensor-Train Network

ICASSP 2020accepted

We propose a tensor-to-vector regression approach to multi-channel speech enhancement in order to address the issue of input size explosion and hidden-layer size expansion. The key idea is to cast the conventional deep neural network (DNN) based vector-to-vector regression formulation under a tensor…

Cited by 0SourceScholar
2018

Generative Adversarial Networks Based Data Augmentation for Noise Robust Speech Recognition

ICASSP 2018accepted

Data augmentation is an effective method to increase the size of training data and reduce the mismatch between training and testing for noise robust speech recognition. Different from the traditional approaches by directly adding noise to the original waveform, in this work we utilize generative adv…

Cited by 0SourceScholar