← Search

Shansong Liu

12 accepted papers

2026

YuE: Scaling Open Foundation Models for Long-Form Music Generation

ICLR 2026poster

We tackle the task of long-form music generation, particularly the challenging \textbf{lyrics-to-song} problem, by introducing \textbf{YuE (乐)}, a family of open-source music generation foundation models. Specifically, YuE scales to trillions of tokens and generates up to five minutes of music while…

Cited by 0SourcecodeScholar
2025

Editing Music with Melody and Text: Using ControlNet for Diffusion Transformer

ICASSP 2025accepted

Despite the significant progress in controllable music generation and editing, challenges remain in the quality and length of generated music due to the use of Mel-spectrogram representations and UNet-based model structures. To address these limitations, we propose a novel approach using a Diffusion…

Cited by 0SourceScholar
2024

Humtrans: A Novel Open-Source Dataset for Humming Melody Transcription and Beyond

ICASSP 2024accepted

This paper introduces the HumTrans dataset, which is publicly available and primarily designed for humming melody transcription. The dataset can also serve as a foundation for downstream tasks such as humming melody based music generation. It consists of 500 musical compositions of different genres…

Cited by 0SourceScholar
2024

Music Understanding LLaMA: Advancing Text-to-Music Generation with Question Answering and Captioning

ICASSP 2024accepted

Text-to-music generation (T2M-Gen) faces a major obstacle due to the scarcity of large-scale publicly available music datasets with natural language captions. To address this, we propose the Music Understanding LLaMA (MU-LLaMA), capable of answering music-related questions and generating captions fo…

Cited by 0SourceScholar
2024

Unified Pretraining Target Based Video-Music Retrieval with Music Rhythm and Video Optical Flow Information

ICASSP 2024accepted

Background music (BGM) can enhance the video’s emotion. However, selecting an appropriate BGM often requires domain knowledge. This has led to the development of video-music retrieval techniques. Most existing approaches utilize pretrained video/music feature extractors trained with different target…

Cited by 0SourceScholar
2022

Exploiting Cross Domain Acoustic-to-Articulatory Inverted Features for Disordered Speech Recognition

ICASSP 2022accepted

Articulatory features are inherently invariant to acoustic signal distortion and have been successfully incorporated into automatic speech recognition (ASR) systems for normal speech. Their practical application to disordered speech recognition is often limited by the difficulty in collecting such s…

Cited by 0SourceScholar
2021

Bayesian Transformer Language Models for Speech Recognition

ICASSP 2021accepted

State-of-the-art neural language models (LMs) represented by Transformers are highly complex. Their use of fixed, deterministic parameter estimates fail to account for model uncertainty and lead to over-fitting and poor generalization when given limited training data. In order to address these issue…

Cited by 0SourceScholar
2021

Development of the Cuhk Elderly Speech Recognition System for Neurocognitive Disorder Detection Using the Dementiabank Corpus

ICASSP 2021accepted

Early diagnosis of Neurocognitive Disorder (NCD) is crucial in facilitating preventive care and timely treatment to delay further progression. This paper presents the development of a state-of-the-art automatic speech recognition (ASR) system built on the Dementia-Bank Pitt corpus for automatic NCD…

Cited by 0SourceScholar
2021

Neural Architecture Search for LF-MMI Trained Time Delay Neural Networks

ICASSP 2021accepted

Deep neural networks (DNNs) based automatic speech recognition (ASR) systems are often designed using expert knowledge and empirical evaluation. In this paper, a range of neural architecture search (NAS) techniques are used to automatically learn two types of hyper-parameters of state-of-the-art fac…

Cited by 0SourceScholar
2020

Audio-Visual Recognition of Overlapped Speech for the LRS2 Dataset

ICASSP 2020accepted

Automatic recognition of overlapped speech remains a highly challenging task to date. Motivated by the bimodal nature of human speech perception, this paper investigates the use of audio-visual technologies for overlapped speech recognition. Three issues associated with the construction of audio-vis…

Cited by 0SourceScholar
2019

Bayesian and Gaussian Process Neural Networks for Large Vocabulary Continuous Speech Recognition

ICASSP 2019accepted

The hidden activation functions inside deep neural networks (DNNs) play a vital role in learning high level discriminative features and controlling the information flows to track longer history. However, the fixed model parameters used in standard DNNs can lead to over-fitting and poor generalizatio…

Cited by 0SourceScholar
2018

Limited-Memory BFGS Optimization of Recurrent Neural Network Language Models for Speech Recognition

ICASSP 2018accepted

Recurrent neural network language models (RNNLM) have become an increasingly popular choice for state-of-the-art speech recognition systems. RNNLMs are normally trained by minimizing the cross entropy (CE) using the stochastic gradient descent (SGD) algorithm. The SGD method only uses first-order de…

Cited by 0SourceScholar