← Search

Fenglong Xie

2 accepted papers

2025

Speaking from Coarse to Fine: Improving Neural Codec Language Model via Multi-Scale Speech Coding and Generation

ICASSP 2025accepted

The neural codec language model (CLM) has demonstrated remarkable performance in text-to-speech (TTS) synthesis. However, troubled by "recency bias", CLM lacks sufficient attention to coarse-grained information at a higher temporal scale, often producing unnatural or even unintelligible speech. This…

Cited by 0SourceScholar
2020

Improving End-to-End Speech Synthesis with Local Recurrent Neural Network Enhanced Transformer

ICASSP 2020accepted

Although Transformer based neural end-to-end TTS model has demonstrated extreme effectiveness in capturing long-term dependencies and achieved state-of-the-art performance, it still suffers from two problems. 1) limited ability to model sequential and local structures in sequences; 2) heavily rely o…

Cited by 0SourceScholar