← Search

Jingbei Li

6 accepted papers

2026

Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability

ICML 2026poster

Recent advances in Large Audio Language Models (LALMs) have extended Text-to-Speech (TTS) to interactive role-play scenarios, which demand high expressiveness and strict adherence to role-play instructions. However, existing models struggle to maintain stylistic consistency with character profiles a…

Cited by 0SourcecodeScholar
2025

DiffCSS: Diverse and Expressive Conversational Speech Synthesis with Diffusion Models

ICASSP 2025accepted

Conversational speech synthesis (CSS) aims to synthesize both contextually appropriate and expressive speech, and considerable efforts have been made to enhance the understanding of conversational context. However, existing CSS systems are limited to deterministic prediction, overlooking the diversi…

Cited by 0SourceScholar
2022

Enhancing Speaking Styles in Conversational Text-to-Speech Synthesis with Graph-Based Multi-Modal Context Modeling

ICASSP 2022accepted

Comparing with traditional text-to-speech (TTS) systems, conversational TTS systems are required to synthesize speeches with proper speaking style confirming to the conversational context. However, state-of-the-art context modeling methods in conversational TTS only model the textual information in…

Cited by 0SourceScholar
2022

Neufa: Neural Network Based End-to-End Forced Alignment with Bidirectional Attention Mechanism

ICASSP 2022accepted

Although deep learning and end-to-end models have been widely used and shown their superiority in automatic speech recognition (ASR) and text-to-speech (TTS) synthesis, state-of-the-art forced alignment (FA) models are still based on hidden Markov model (HMM). HMM has limited view of contextual info…

Cited by 0SourceScholar
2021

Emotion Controllable Speech Synthesis Using Emotion-Unlabeled Dataset with the Assistance of Cross-Domain Speech Emotion Recognition

ICASSP 2021accepted

Neural text-to-speech (TTS) approaches generally require a huge number of high quality speech data, which makes it difficult to obtain such a dataset with extra emotion labels. In this paper, we propose a novel approach for emotional TTS synthesis on a TTS dataset without emotion labels. Specificall…

Cited by 0SourceScholar
2021

Syntactic Representation Learning For Neural Network Based TTS with Syntactic Parse Tree Traversal

ICASSP 2021accepted

Syntactic structure of a sentence text is correlated with the prosodic structure of the speech that is crucial for improving the prosody and naturalness of a text-to-speech (TTS) system. Nowadays TTS systems usually try to incorporate syntactic structure information with manually designed features b…

Cited by 0SourceScholar