← Search

Changfeng Gao

4 accepted papers

2026

MELA-TTS: JOINT TRANSFORMER-DIFFUSION MODEL WITH REPRESENTATION ALIGNMENT FOR SPEECH SYNTHESIS

ICASSP 2026poster

This work introduces MELA-TTS, a novel joint transformer-diffusion framework for end-to-end text-to-speech synthesis. By autoregressively generating continuous mel-spectrogram frames from linguistic and speaker conditions, our architecture eliminates the need for speech tokenization and multi-stage…

Cited by 0SourcePDFScholar
2021

Pre-Training Transformer Decoder for End-to-End ASR Model with Unpaired Text Data

ICASSP 2021accepted

This paper presents a method to pre-train transformer-based encoder-decoder automatic speech recognition (ASR) models using sufficient target-domain text. During pre-training, we train the transformer decoder as a conditional language model with empty or artifical states, rather than the real encode…

Cited by 0SourceScholar
2020

Transformer-Based Online CTC/Attention End-To-End Speech Recognition Architecture

ICASSP 2020accepted

Recently, Transformer has gained success in automatic speech recognition (ASR) field. However, it is challenging to deploy a Transformer-based end-to-end (E2E) model for online speech recognition. In this paper, we propose the Transformer-based online CTC/attention E2E ASR architecture, which contai…

Cited by 0SourceScholar