← Search

Zhengyan Sheng

4 accepted papers

2026

SYNCSPEECH: EFFICIENT AND LOW-LATENCY TEXT-TO-SPEECH BASED ON TEMPORAL MASKED TRANSFORMER

ICASSP 2026oral

Current text-to-speech (TTS) models face a persistent limitation: autoregressive (AR) models suffer from low generation efficiency, while modern non-autoregressive (NAR) models experience high latency due to their unordered temporal nature. To bridge this divide, we introduce SyncSpeech, an efficien…

Cited by 0SourcePDFScholar
2025

UniSpeaker: A Unified Approach for Multimodality-driven Speaker Generation

EMNLP 2025

While recent advances in reference-based speaker cloning have significantly improved the authenticity of synthetic speech, speaker generation driven by multimodal cues such as visual appearance, textual descriptions, and other biometric signals remains in its early stages. To pioneer truly multimoda

2023

Zero-Shot Personalized Lip-To-Speech Synthesis with Face Image Based Voice Control

ICASSP 2023accepted

Lip-to-Speech (Lip2Speech) synthesis, which predicts corresponding speech from talking face images, has witnessed significant progress with various models and training strategies in a series of independent studies. However, existing studies can not achieve voice control under zero-shot condition, be…

Cited by 0SourceScholar
2022

Dementia Detection by Fusing Speech and Eye-Tracking Representation

ICASSP 2022accepted

This paper proposes a method of detecting dementia from the simultaneous speech and eye-tracking recordings of subjects in a picture description task. First, automatic speech recognition (ASR) and regional picture recognition (RPR) models are built to extract content-related bottleneck (BN) features…

Cited by 0SourceScholar