← Search

Yaoxun Xu

4 accepted papers

2025

LeVo: High-Quality Song Generation with Multi-Preference Alignment

NeurIPS 2025poster

Recent advances in large language models (LLMs) and audio language models have significantly improved music generation, particularly in lyrics-to-song generation. However, existing approaches still struggle with the complex composition of songs and the scarcity of high-quality data, leading to limit…

Cited by 0SourcecodeScholar
2025

SongEditor: Adapting Zero-Shot Song Generation Language Model as a Multi-Task Editor

AAAI 2025technical

The emergence of novel generative modeling paradigms, particularly audio language models, has significantly advanced the field of song generation. Although state-of-the-art models are capable of synthesizing both vocals and accompaniment tracks up to several minutes long concurrently, research about…

2024

SECap: Speech Emotion Captioning with Large Language Model

AAAI 2024technical

Speech emotions are crucial in human communication and are extensively used in fields like speech synthesis and natural language understanding. Most prior studies, such as speech emotion recognition, have categorized speech emotions into a fixed set of classes. Yet, emotions expressed in human spee…

2023

CB-Conformer: Contextual Biasing Conformer for Biased Word Recognition

ICASSP 2023accepted

Due to the mismatch between the source and target domains, how to better utilize the biased word information to improve the performance of the automatic speech recognition model in the target domain becomes a hot research topic. Previous approaches either decode with a fixed external language model…

Cited by 0SourceScholar