Synthetic Dataset Generation for String Ensemble Separation
Minju Kim, Joonhyeon Bae, Eunsik Shin, Kyogu Lee
Abstract
Most studies on music source separation have traditionally concentrated on separating popular music into vocals, bass, drums, and others, with fewer studies focusing on chamber ensemble audio. The task of separating sources from chamber ensemble audio presents greater challenges due to factors like timbral similarity, high synchronization, and spectral overlap. These complexities necessitate more refined and realistic datasets for effective chamber ensemble separation. Yet, the creation of such datasets faces numerous challenges, depending on the methods of dataset formation. Taking these challenges into account, our work seeks to enhance the performance of chamber ensemble separation tasks, with a particular focus on string quartets. Based on dataset generation using a neural synthesis model, we propose an approach that incorporates musical expressions into the dataset generation process using MusicXML files. Furthermore, our framework adapt reverberation to the audio to better match the target mixtures. Our objective evaluation shows an increase in performance of chamber ensemble separation, supported by a subjective listening test to demonstrate improvement in our dataset. We also release our dataset and source code for further usage.
BibTeX
@inproceedings{icassp2025_syntheticdataset,
title = {Synthetic Dataset Generation for String Ensemble Separation},
author = {Minju Kim and Joonhyeon Bae and Eunsik Shin and Kyogu Lee},
booktitle = {ICASSP 2025},
year = {2025}
}