Adapting Self-Supervised Models to Multi-Talker Speech Recognition Using Speaker Embeddings
Self-supervised learning (SSL) methods which learn representations of data without explicit supervision have gained popularity in speech-processing tasks, particularly for single-talker applications. However, these models often have degraded performance for multi-talker scenarios — possibly due to t…