← Search

Gus Xia

24 accepted papers

2026

Do LLMs “Feel”? Emotion Circuits Discovery and Control

ICML 2026poster

As the demand for emotional intelligence in large language models (LLMs) grows, a key challenge lies in understanding the internal mechanisms that give rise to emotional expression and in controlling emotions in generated text. This study addresses three core questions: (1) Do LLMs contain context-a…

Cited by 0SourceScholar
2026

VITEX: VISUAL TEXTURE CONTROL FOR MULTI-TRACK SYMBOLIC MUSIC GENERATION VIA DISCRETE DIFFUSION MODELS

ICASSP 2026poster

In automatic music generation, a central challenge is to design controls that enable meaningful human-machine interaction. Existing systems often rely on extrinsic inputs such as text prompts or metadata, which do not allow humans to directly shape the composition. While prior work has explored intr…

Cited by 0SourcePDFScholar
2026

YuE: Scaling Open Foundation Models for Long-Form Music Generation

ICLR 2026poster

We tackle the task of long-form music generation, particularly the challenging \textbf{lyrics-to-song} problem, by introducing \textbf{YuE (乐)}, a family of open-source music generation foundation models. Specifically, YuE scales to trillions of tokens and generates up to five minutes of music while…

Cited by 0SourcecodeScholar
2025

CLaMP 3: Universal Music Information Retrieval Across Unaligned Modalities and Unseen Languages

ACL 2025finding

CLaMP 3 is a unified framework developed to address challenges of cross-modal and cross-lingual generalization in music information retrieval. Using contrastive learning, it aligns all major music modalities–including sheet music, performance signals, and audio recordings–with multilingual text in a…

2025

MuPT: A Generative Symbolic Music Pretrained Transformer

ICLR 2025poster

In this paper, we explore the application of Large Language Models (LLMs) to the pre-training of music. While the prevalent use of MIDI in music modeling is well-established, our findings suggest that LLMs are inherently more compatible with ABC Notation, which aligns more closely with their design…

Cited by 10SourcePDFScholar
2025

Music for All: Representational Bias and Cross-Cultural Adaptability of Music Generation Models

NAACL 2025findings

The advent of Music-Language Models has greatly enhanced the automatic music generation capability of AI systems, but they are also limited in their coverage of the musical genres and cultures of the world. We present a study of the datasets and research papers for music generation and quantify the…

2025

Unifying Symbolic Music Arrangement: Track-Aware Reconstruction and Structured Tokenization

NeurIPS 2025poster

We present a unified framework for automatic multitrack music arrangement that enables a single pre-trained symbolic music model to handle diverse arrangement scenarios, including reinterpretation, simplification, and additive generation. At its core is a segment-level reconstruction objective opera…

Cited by 0SourcecodeScholar
2025

Unsupervised Disentanglement of Content and Style via Variance-Invariance Constraints

ICLR 2025poster

We contribute an unsupervised method that effectively learns disentangled content and style representations from sequences of observations. Unlike most disentanglement algorithms that rely on domain-specific labels or knowledge, our method is based on the insight of domain-general statistical differ…

Cited by 0SourcePDFScholar
2024

Arrange, Inpaint, and Refine: Steerable Long-term Music Audio Generation and Editing via Content-based Controls

IJCAI 2024poster

Controllable music generation plays a vital role in human-AI music co-creation. While Large Language Models (LLMs) have shown promise in generating high-quality music, their focus on autoregressive generation limits their utility in music editing tasks. To bridge this gap, To address this gap, we pr…

2024

ChatMusician: Understanding and Generating Music Intrinsically with LLM

ACL 2024findings

While LLMs demonstrate impressive capabilities in musical knowledge, we find that music reasoning is still an unsolved task.We introduce ChatMusician, an open-source large language model (LLM) that integrates intrinsic musical abilities. It is based on continual pre-training and finetuning LLaMA2 on…

2024

MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

ICLR 2024poster

Self-supervised learning (SSL) has recently emerged as a promising paradigm for training generalisable models on large-scale data in the fields of vision, text, and speech. Although SSL has been proven effective in speech and audio, its application to music audio has yet to be thoroughly explored.…

2024

MusicMagus: Zero-Shot Text-to-Music Editing via Diffusion Models

IJCAI 2024poster

Recent advances in text-to-music generation models have opened new avenues in musical creativity. However, the task of editing these generated music remains a significant challenge. This paper introduces a novel approach to edit music generated by such models, enabling the modification of specific a…

2024

Structured Multi-Track Accompaniment Arrangement via Style Prior Modelling

NeurIPS 2024poster

In the realm of music AI, arranging rich and structured multi-track accompaniments from a simple lead sheet presents significant challenges. Such challenges include maintaining track cohesion, ensuring long-term coherence, and optimizing computational efficiency. In this paper, we introduce a novel…

2024

Whole-Song Hierarchical Generation of Symbolic Music Using Cascaded Diffusion Models

ICLR 2024spotlight

Recent deep music generation studies have put much emphasis on long-term generation with structures. However, we are yet to see high-quality, well-structured **whole-song** generation. In this paper, we make the first attempt to model a full music piece under the realization of *compositional hierar…

Cited by 11SourcePDFScholar
2023

Controllable Music Inpainting with Mixed-Level and Disentangled Representation

ICASSP 2023accepted

Music inpainting, which is to complete the missing part of a piece given some context, is an important task of automated music generation. In this study, we contribute a controllable inpainting model by combining the high expressivity of mixed-level, disentangled music representations and the strong…

Cited by 0SourceScholar
2023

Learning Interpretable Low-dimensional Representation via Physical Symmetry

NeurIPS 2023poster

We have recently seen great progress in learning interpretable music representations, ranging from basic factors, such as pitch and timbre, to high-level concepts, such as chord and texture. However, most methods rely heavily on music domain knowledge. It remains an open question what general comput…

2023

Q&A: Query-Based Representation Learning for Multi-Track Symbolic Music re-Arrangement

IJCAI 2023poster

Music rearrangement is a common music practice of reconstructing and reconceptualizing a piece using new composition or instrumentation styles, which is also an important task of automatic music generation. Existing studies typically model the mapping from a source piece to a target piece via superv…

2022

Audio-To-Symbolic Arrangement Via Cross-Modal Music Representation Learning

ICASSP 2022accepted

Could we automatically derive the score of a piano accompaniment based on the audio of a pop song? This is the audio-to-symbolic arrangement problem we tackle in this paper. A good arrangement model should not only consider the audio content but also have prior knowledge of piano composition (so tha…

Cited by 0SourceScholar
2022

Music Phrase Inpainting Using Long-Term Representation and Contrastive Loss

ICASSP 2022accepted

Deep generative modeling has already become the leading technique for music automation. However, long-term generation remains a challenging task as most methods fall short in preserving a natural structure and the overall musicality when the generation scope exceeds several beats. In this study, we…

Cited by 0SourceScholar
2020

Transformer VAE: A Hierarchical Model for Structure-Aware and Interpretable Music Representation Learning

ICASSP 2020accepted

Structure awareness and interpretability are two of the most desired properties of music generation algorithms. Structure-aware models generate more natural and coherent music with long-term dependencies, while interpretable models are more friendly for human-computer interaction and co-creation. To…

Cited by 0SourceScholar