← Search

Xuchen Song

7 accepted papers

2024

InstructME: An Instruction Guided Music Edit Framework with Latent Diffusion Models

IJCAI 2024poster

Music editing primarily entails the modification of instrument tracks or remixing in the whole, which offers a novel reinterpretation of the original piece through a series of operations. These music processing methods hold immense potential across various applications but demand substantial experti…

2023

Efficient Neural Music Generation

NeurIPS 2023poster

Recent progress in music generation has been remarkably advanced by the state-of-the-art MusicLM, which comprises a hierarchy of three LMs, respectively, for semantic, coarse acoustic, and fine acoustic modelings. Yet, sampling with the MusicLM requires processing through these LMs one by one to obt…

2022

Modeling Beats and Downbeats with a Time-Frequency Transformer

ICASSP 2022accepted

Transformer is a successful deep neural network (DNN) architecture that has shown its versatility not only in natural language processing but also in music information retrieval (MIR). In this paper, we present a novel Transformer-based approach to tackle beat and downbeat tracking. This approach em…

Cited by 0SourceScholar
2021

Modeling the Compatibility of Stem Tracks to Generate Music Mashups

AAAI 2021technical

A music mashup combines audio elements from two or more songs to create a new work. To reduce the time and effort required to make them, researchers have developed algorithms that predict the compatibility of audio elements. Prior work has focused on mixing unaltered excerpts, but advances in source…

2021

Supervised Chorus Detection for Popular Music Using Convolutional Neural Network and Multi-Task Learning

ICASSP 2021accepted

This paper presents a novel supervised approach to detecting the chorus segments in popular music. Traditional approaches to this task are mostly unsupervised, with pipelines designed to target some quality that is assumed to define "chorusness," which usually means seeking the loudest or most frequ…

Cited by 0SourceScholar
2020

Source Separation with Weakly Labelled Data: an Approach to Computational Auditory Scene Analysis

ICASSP 2020accepted

Source separation is the task of separating an audio recording into individual sound sources. Source separation is fundamental for computational auditory scene analysis. Previous work on source separation has focused on separating particular sound classes such as speech and music. Much previous work…

Cited by 0SourceScholar
2019

A Neural Network Based Ranking Framework to Improve ASR with NLU Related Knowledge Deployed

ICASSP 2019accepted

This work proposes a new neural network framework to simultaneously rank multiple hypotheses generated by one or more automatic speech recognition (ASR) engines for a speech utterance. Features fed in the framework not only include those calculated from the ASR information, but also involve natural…

Cited by 0SourceScholar