2024
MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training
ICLR 2024poster
Self-supervised learning (SSL) has recently emerged as a promising paradigm for training generalisable models on large-scale data in the fields of vision, text, and speech. Although SSL has been proven effective in speech and audio, its application to music audio has yet to be thoroughly explored.…