← Search

Deyi Tuo

5 accepted papers

2022

Disentangling Content and Fine-Grained Prosody Information Via Hybrid ASR Bottleneck Features for Voice Conversion

ICASSP 2022accepted

Non-parallel data voice conversion (VC) have achieved considerable breakthroughs recently through introducing bottleneck features (BNFs) extracted by the automatic speech recognition(ASR) model. However, selection of BNFs have a significant impact on VC result. For example, when extracting BNFs from…

Cited by 0SourceScholar
2022

FullSubNet+: Channel Attention Fullsubnet with Complex Spectrograms for Speech Enhancement

ICASSP 2022accepted

Previously proposed FullSubNet has achieved outstanding performance in Deep Noise Suppression (DNS) Challenge and attracted much attention. However, it still encounters issues such as input-output mismatch and coarse processing for frequency bands. In this paper, we propose an extended single-channe…

Cited by 0SourceScholar
2021

The Huya Multi-Speaker and Multi-Style Speech Synthesis System for M2voc Challenge 2020

ICASSP 2021accepted

Text-to-speech systems now can generate speech that is hard to distinguish from human speech. In this paper, we propose the Huya multi-speaker and multi-style speech synthesis system which is based on DurIAN and HiFi-GAN to generate high-fidelity speech even under low-resource condition. We use the…

Cited by 0SourceScholar
2019

Boundary Discriminative Large Margin Cosine Loss for Text-independent Speaker Verification

ICASSP 2019accepted

Deep neural network based speaker embeddings have attracted much attention in text-independent speaker verification task. In addition to the network architecture, an appropriate design of the loss function is crucial for the deep discriminative embedding extractor. Inspired by the success of Large M…

Cited by 0SourceScholar
2019

Quasi-fully Convolutional Neural Network with Variational Inference for Speech Synthesis

ICASSP 2019accepted

Recurrent neural networks, such as gated recurrent units (GRUs) and long short-term memory (LSTM), are widely used on acoustic modeling for speech synthesis. However, such sequential generating processes are not friendly to today’s massively parallel computing devices. We introduce a fully convoluti…

Cited by 0SourceScholar