Enhancing Note-Level Singing Transcription Model with Unlabeled and Weakly Labeled Data
Yao Qiu, Jinchao Zhang, Yong Shan, Jie Zhou
Abstract
Note-level automatic singing transcription, involving the extraction of onset, offset, and pitch information from a singing voice, is a crucial process in the field of Music Information Retrieval (MIR), The recent advancements in deep learning models have led to significant progress in this field. However, annotating a training dataset requires professional music expertise, and the entire annotation process is time-consuming and labor-intensive. Therefore, this field suffers from a severe data scarcity problem. To address this issue, we developed a singing transcription model based on wav2vec 2.0, a pretrained speech representation model. The model can learn from unlabeled speech and weakly-labeled singing data and use this knowledge to benefit the transcription task. The experiments showed that our proposed method achieves a significant improvement over previous approaches on various benchmarks. Moreover, additional experiments demonstrate that our method achieves competitive performance even with a small proportion of training data.
BibTeX
@inproceedings{icassp2024_enhancingnotelev,
title = {Enhancing Note-Level Singing Transcription Model with Unlabeled and Weakly Labeled Data},
author = {Yao Qiu and Jinchao Zhang and Yong Shan and Jie Zhou},
booktitle = {ICASSP 2024},
year = {2024}
}