SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes
Self-supervised pre-trained audio networks have seen widespread adoption in real-world systems, particularly in multi-modal large language models. These networks are often employed in a frozen state, under the assumption that the self-supervised pre-training has sufficiently equipped them to handle…