← Search

Qiushi Zhu

6 accepted papers

2024

An Experimental Comparison of Noise-Robust Text-To-Speech Synthesis Systems Based On Self-Supervised Representation

ICASSP 2024accepted

With the advance in deep learning, text-to-speech (TTS) using clean speech has witnessed significant performance improvements. As the data collected in real scenes often contain noise and thus needs to be denoised, TTS models trained on the enhanced speech suffer from distortions and residual noises…

Cited by 0SourceScholar
2024

DurIAN-E 2: Duration Informed Attention Network with Adaptive Variational Autoencoder and Adversarial Learning for Expressive Text-to-Speech Synthesis

ICASSP 2024accepted

This paper proposes an improved version of DurIAN-E (DurIAN-E 2), which is also a duration informed attention neural network for expressive and high-fidelity text-to-speech (TTS) synthesis. Similar with the DurIAN-E model, multiple stacked SwishRNN-based Transformer blocks are utilized as linguistic…

Cited by 0SourceScholar
2024

Listen Again and Choose the Right Answer: A New Paradigm for Automatic Speech Recognition with Large Language Models

ACL 2024findings

Recent advances in large language models (LLMs) have promoted generative error correction (GER) for automatic speech recognition (ASR), which aims to predict the ground-truth transcription from the decoded N-best hypotheses. Thanks to the strong language generation ability of LLMs and rich informati…

2024

Multichannel AV-wav2vec2: A Framework for Learning Multichannel Multi-Modal Speech Representation

AAAI 2024technical

Self-supervised speech pre-training methods have developed rapidly in recent years, which show to be very effective for many near-field single-channel speech tasks. However, far-field multichannel speech processing is suffering from the scarcity of labeled multichannel data and complex ambient noise…

2023

Cross-Modal Global Interaction and Local Alignment for Audio-Visual Speech Recognition

IJCAI 2023poster

Audio-visual speech recognition (AVSR) research has gained a great success recently by improving the noise-robustness of audio-only automatic speech recognition (ASR) with noise-invariant visual information. However, most existing AVSR approaches simply fuse the audio and visual features by concaten…

2023

Gradient Remedy for Multi-Task Learning in End-to-End Noise-Robust Speech Recognition

ICASSP 2023accepted

Speech enhancement (SE) is proved effective in reducing noise from noisy speech signals for downstream automatic speech recognition (ASR), where multi-task learning strategy is employed to jointly optimize these two tasks. However, the enhanced speech learned by SE objective may not always yield goo…

Cited by 0SourceScholar