← Search

Xiaoxue Gao

7 accepted papers

2026

Emo-LiPO: Listwise Preference Optimization for Fine-Grained Emotion Intensity Control in LLM-based Text-to-Speech

IJCAI 2026

Large language model (LLM)-based text-to-speech (TTS) systems enable prompt-conditioned emotional control but struggle with fine-grained emotion intensity due to the semantic--acoustic gap between text and speech. To address this challenge, we formulate emotion intensity control in LLM-based TTS as

Cited by 0Scholar
2025

Emo-DPO: Controllable Emotional Speech Synthesis through Direct Preference Optimization

ICASSP 2025accepted

Current emotional text-to-speech (TTS) models pre-dominantly conduct supervised training to learn the conversion from text and desired emotion to its emotional speech, focusing on a single emotion per text-speech pair. These models only learn the correct emotional outputs without fully comprehending…

Cited by 0SourceScholar
2025

MEDSAGE: Enhancing Robustness of Medical Dialogue Summarization to ASR Errors with LLM-generated Synthetic Dialogues

AAAI 2025technical

Automatic Speech Recognition (ASR) systems are pivotal in transcribing speech into text, yet the errors they introduce can significantly degrade the performance of downstream tasks like summarization. This issue is particularly pronounced in clinical dialogue summarization, a low-resource domain whe…

Cited by 0SourcePDFScholar
2024

Beyond Single-Audio: Advancing Multi-Audio Processing in Audio Large Language Models

EMNLP 2024finding

Various audio-LLMs (ALLMs) have been explored recently for tackling different audio tasks simultaneously using a single, unified model. While existing evaluations of ALLMs primarily focus on single-audio tasks, real-world applications often involve processing multiple audio streams simultaneously. T…

2023

Token2vec: A Joint Self-Supervised Pre-Training Framework Using Unpaired Speech and Text

ICASSP 2023accepted

Self-supervised pre-training has been successful in both text and speech processing. Speech and text offer different but complementary information. The question is whether we are able to perform a speech-text joint pre-training on unpaired speech and text. In this paper, we take the idea of self-sup…

Cited by 0SourceScholar
2022

Genre-Conditioned Acoustic Models for Automatic Lyrics Transcription of Polyphonic Music

ICASSP 2022accepted

Lyrics transcription of polyphonic music is challenging not only because the singing vocals are corrupted by the background music, but also because the background music and the singing style vary across music genres, such as pop, metal, and hip hop, which affects lyrics intelligibility of the song i…

Cited by 0SourceScholar