2025
Neighborhood Attention Transformer with Progressive Channel Fusion for Speaker Verification
ICASSP 2025accepted
Transformer-based architectures for speaker verification typically require more training data than ECAPA-TDNN. Therefore, recent work has generally been trained on VoxCeleb1&2. We propose a backbone network based on self-attention, which can achieve competitive results when trained on VoxCeleb2 alon…