2024
Improving Multi-Speaker ASR With Overlap-Aware Encoding And Monotonic Attention
ICASSP 2024accepted
End-to-end (E2E) multi-speaker speech recognition with the serialized output training (SOT) strategy demonstrates good performance in modeling diverse speaker scenarios. However, the E2E architecture doesn’t explicitly address the modeling of overlapping speech areas, potentially limiting the model’…