← Search

Subhadeep Dey

10 accepted papers

2020

Attentive Modality Hopping Mechanism for Speech Emotion Recognition

ICASSP 2020accepted

In this work, we explore the impact of visual modality in addition to speech and text for improving the accuracy of the emotion detection system. The traditional approaches tackle this task by independently fusing the knowledge from the various modalities for performing emotion classification. In co…

Cited by 0SourceScholar
2020

Incremental Semi-Supervised Learning for Multi-Genre Speech Recognition

ICASSP 2020accepted

In this work, we explore a data scheduling strategy for semi-supervised learning (SSL) for acoustic modeling in automatic speech recognition. The conventional approach uses a seed model trained with supervised data to automatically recognize the entire set of unlabeled (auxiliary) data to generate n…

Cited by 0SourceScholar
2019

A Bayesian Approach to Inter-task Fusion for Speaker Recognition

ICASSP 2019accepted

In i-vector based speaker recognition systems, back-end classifiers are trained to factor out nuisance information and retain only the speaker identity. As a result, variabilities arising due to gender, language and accent (among many others) are suppressed. Inter-task fusion, in which such metadata…

Cited by 0SourceScholar
2019

Speech Emotion Recognition Using Multi-hop Attention Mechanism

ICASSP 2019accepted

In this paper, we are interested in exploiting textual and acoustic data of an utterance for the speech emotion classification task. The baseline approach models the information from audio and text independently using two deep neural networks (DNNs). The outputs from both the DNNs are then fused for…

Cited by 0SourceScholar
2018

DNN Based Speaker Embedding Using Content Information for Text-Dependent Speaker Verification

ICASSP 2018accepted

In this paper, we are interested in exploring Deep Neural Network (DNN) based speaker embedding for Random-digit task using content information. To this end, a technique is applied to automatically select common phonetic units between the enrollment and test data to produce speaker verification scor…

Cited by 0SourceScholar
2017

Exploiting sequence information for text-dependent Speaker Verification

ICASSP 2017accepted

Model-based approaches to Speaker Verification (SV), such as Joint Factor Analysis (JFA), i-vector and relevance Maximum-a-Posteriori (MAP), have shown to provide state-of-the-art performance for text-dependent systems with fixed phrases. The performance of i-vector and JFA models has been further e…

Cited by 0SourceScholar
2017

Intra-class covariance adaptation in PLDA back-ends for speaker verification

ICASSP 2017accepted

Multi-session training conditions are becoming increasingly common in recent benchmark datasets for both text-independent and text-dependent speaker verification. In the state-of-the-art i-vector framework for speaker verification, such conditions are addressed by simple techniques such as averaging…

Cited by 0SourceScholar
2016

Deep neural network based posteriors for text-dependent speaker verification

ICASSP 2016accepted

The i-vector and Joint Factor Analysis (JFA) systems for text-dependent speaker verification use sufficient statistics computed from a speech utterance to estimate speaker models. These statistics average the acoustic information over the utterance thereby losing all the sequence information. In thi…

Cited by 0SourceScholar
2016

Information theoretic clustering for unsupervised domain-adaptation

ICASSP 2016accepted

The aim of the domain-adaptation task for speaker verification is to exploit unlabelled target domain data by using the labelled source domain data effectively. The i-vector based Probabilistic Linear Discriminant Analysis (PLDA) framework approaches this task by clustering the target domain data an…

Cited by 0SourceScholar
2015

Employment of Subspace Gaussian Mixture Models in speaker recognition

ICASSP 2015accepted

This paper presents Subspace Gaussian Mixture Model (SGMM) approach employed as a probabilistic generative model to estimate speaker vector representations to be subsequently used in the speaker verification task. SGMMs have already been shown to significantly outperform traditional HMM/GMMs in Auto…

Cited by 29SourceScholar