← Search

Xinfa Zhu

6 accepted papers

2026

KALL-E: Autoregressive Speech Synthesis with Next-Distribution Prediction

AAAI 2026technical

We introduce KALL-E, a novel autoregressive (AR) language model for text-to-speech (TTS) synthesis that operates by predicting the next distribution of continuous speech frames. Unlike existing methods, KALL-E directly models the continuous speech distribution conditioned on text, eliminating the ne

Cited by 0SourcePDFScholar
2025

LLaSE-G1: Incentivizing Generalization Capability for LLaMA-based Speech Enhancement

ACL 2025long

Recent advancements in language models (LMs) have demonstrated strong capabilities in semantic understanding and contextual modeling, which have flourished in generative speech enhancement (SE). However, many LM-based SE approaches primarily focus on semantic information, often neglecting the critic…

2025

ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training

ICASSP 2025accepted

Style voice conversion aims to transform the speaking style of source speech into a desired style while keeping the original speaker’s identity. However, previous style voice conversion approaches primarily focus on well-defined domains such as emotional aspects, limiting their practical application…

Cited by 0SourceScholar
2024

SELM: Speech Enhancement using Discrete Tokens and Language Models

ICASSP 2024accepted

Language models (LMs) have recently shown superior performances in various speech generation tasks, demonstrating their powerful ability for semantic context modeling. Given the intrinsic similarity between speech generation and speech enhancement, harnessing semantic information is advantageous for…

Cited by 0SourceScholar
2024

Spontts: Modeling and Transferring Spontaneous Style for TTS

ICASSP 2024accepted

Spontaneous speaking style exhibits notable differences from other speaking styles due to various spontaneous phenomena (e.g., filled pauses, prolongation) and substantial prosody variation (e.g., diverse pitch and duration variation, occasional non-verbal speech like a smile), posing challenges to…

Cited by 0SourceScholar
2023

Multi-Speaker Expressive Speech Synthesis via Multiple Factors Decoupling

ICASSP 2023accepted

This paper aims to synthesize the target speaker’s speech with desired speaking style and emotion by transferring the style and emotion from reference speech recorded by other speakers. We address this challenging problem with a two-stage framework composed of a text-to-style-and-emotion (Text2SE) m…

Cited by 0SourceScholar