2025
LAVViT: Latent Audio-Visual Vision Transformers for Speaker Verification
ICASSP 2025accepted
Recently, Vision Transformers (ViTs) have shown remarkable success in various computer vision applications. In this work, we have explored the potential of ViTs, pre-trained on visual data, for audio-visual speaker verification. To cope with the challenges of large-scale training, we introduce the L…