← Search

Raghuveer Peri

6 accepted papers

2025

Knowledge Distillation From Ensemble for Spoken Language Identification

ICASSP 2025accepted

Spoken language identification (LID) has seen substantial performance gains with the rise of large-scale models. However, these models are often computationally expensive and impractical for many real-world applications. In this work, we propose a novel knowledge distillation from ensemble framework…

Cited by 0SourceScholar
2024

SpeechGuard: Exploring the Adversarial Robustness of Multi-modal Large Language Models

ACL 2024findings

Integrated Speech and Large Language Models (SLMs) that can follow speech instructions and generate relevant text responses have gained popularity lately. However, the safety and robustness of these models remains largely unclear. In this work, we investigate the potential vulnerabilities of such in…

2021

Adversarial Defense for Deep Speaker Recognition Using Hybrid Adversarial Training

ICASSP 2021accepted

Deep neural network based speaker recognition systems can easily be deceived by an adversary using minuscule imperceptible perturbations to the input speech samples. These adversarial attacks pose serious security threats to the speaker recognition systems that use speech biometric. To address this…

Cited by 0SourceScholar
2021

Disentanglement for Audio-Visual Emotion Recognition Using Multitask Setup

ICASSP 2021accepted

Deep learning models trained on audio-visual data have been successfully used to achieve state-of-the-art performance for emotion recognition. In particular, models trained with multitask learning have shown additional performance improvements. However, such multitask models entangle information bet…

Cited by 0SourceScholar
2020

Robust Speaker Recognition Using Unsupervised Adversarial Invariance

ICASSP 2020accepted

In this paper, we address the problem of speaker recognition in challenging acoustic conditions using a novel method to extract robust speaker-discriminative speech representations. We adopt a recently proposed unsupervised adversarial invariance architecture to train a network that maps speaker emb…

Cited by 0SourceScholar
2020

Speaker Diarization Using Latent Space Clustering in Generative Adversarial Network

ICASSP 2020accepted

In this work, we propose deep latent space clustering for speaker diarization using generative adversarial network (GAN) back-projection with the help of an encoder network. The proposed diarization system is trained jointly with GAN loss, latent variable recovery loss, and a clustering-specific los…

Cited by 0SourceScholar