Hybrid Neural Network with Cross- and Self-Module Attention Pooling for Text-Independent Speaker Verification
Extraction of a speaker embedding vector plays an important role in deep learning-based speaker verification. In this contribution, to extract speaker discriminant utterance level embeddings, we propose a hybrid neural network that employs both cross- and self-module attention pooling mechanisms. Mo…