ICASSP 2019accepted0 citations

A Streamlined Encoder/decoder Architecture for Melody Extraction

Tsung-Han Hsieh, Li Su, Yi-Hsuan Yang

Abstract

Melody extraction in polyphonic musical audio is important for music signal processing. In this paper, we propose a novel streamlined encoder/decoder network that is designed for the task. We make two technical contributions. First, drawing inspiration from a state-of-the-art model for semantic pixel-wise segmentation, we pass through the pooling indices between pooling and un-pooling layers to localize the melody in frequency. We can achieve result close to the state-of-the-art with much fewer convolutional layers and simpler convolution modules. Second, we propose a way to use the bottleneck layer of the network to estimate the existence of a melody line for each time frame, and make it possible to use a simple argmax function instead of ad-hoc thresholding to get the final estimation of the melody line. Our experiments on both vocal melody extraction and general melody extraction validate the effectiveness of the proposed model.

BibTeX
@inproceedings{icassp2019_astreamlinedenco,
  title = {A Streamlined Encoder/decoder Architecture for Melody Extraction},
  author = {Tsung-Han Hsieh and Li Su and Yi-Hsuan Yang},
  booktitle = {ICASSP 2019},
  year = {2019}
}
A Streamlined Encoder/decoder Architecture for Melody Extraction · ICASSP 2019