← Search

Vikramjit Mitra

12 accepted papers

2024

Investigating Salient Representations and Label Variance in Dimensional Speech Emotion Analysis

ICASSP 2024accepted

Representations derived from models such as BERT (Bidirectional Encoder Representations from Transformers) and Hu-BERT (Hidden units BERT), have helped to achieve state-of-the-art performance in dimensional speech emotion recognition. Despite their large dimensionality, and even though these represe…

Cited by 0SourceScholar
2023

Pre-Trained Model Representations and Their Robustness Against Noise for Speech Emotion Analysis

ICASSP 2023accepted

Pre-trained model representations have demonstrated state-of-the-art performance in speech recognition, natural language processing, and other applications. Speech models, such as Bidirectional Encoder Representations from Transformers (BERT) and Hidden units BERT (HuBERT), have enabled generating l…

Cited by 0SourceScholar
2021

SEP-28k: A Dataset for Stuttering Event Detection from Podcasts with People Who Stutter

ICASSP 2021accepted

The ability to automatically detect stuttering events in speech could help speech pathologists track an individual’s fluency over time or help improve speech recognition systems for people with atypical speech patterns. Despite increasing interest in this area, existing public datasets are too small…

Cited by 0SourceScholar
2020

Detecting Emotion Primitives from Speech and Their Use in Discerning Categorical Emotions

ICASSP 2020accepted

Emotion plays an essential role in human-to-human communication, enabling us to convey feelings such as happiness, frustration, and sincerity. While modern speech technologies rely heavily on speech recognition and natural language understanding for speech content understanding, the investigation of…

Cited by 17SourceScholar
2018

Articulatory Information and Multiview Features for Large Vocabulary Continuous Speech Recognition

ICASSP 2018accepted

This paper explores the use of multi-view features and their discriminative transforms in a convolutional deep neural network (CNN) architecture for a continuous large vocabulary speech recognition task. Mel-filterbank energies and perceptually motivated forced damped oscillator coefficient (DOC) fe…

Cited by 18SourceScholar
2018

Interpreting DNN Output Layer Activations: A Strategy to Cope with Unseen Data in Speech Recognition

ICASSP 2018accepted

Unseen data can degrade performance of deep neural net (DNN) acoustic models. To cope with unseen data, adaptation techniques are deployed. For unlabeled unseen data, one must generate some hypothesis given an existing model, which is used as the label for model adaptation. However, assessing the go…

Cited by 0SourceScholar
2017

Joint modeling of articulatory and acoustic spaces for continuous speech recognition tasks

ICASSP 2017accepted

Articulatory information can effectively model variability in speech and can improve speech recognition performance under varying acoustic conditions. Learning speaker-independent articulatory models has always been challenging, as speaker-specific information in the articulatory and acoustic spaces…

Cited by 0SourceScholar
2017

Speech recognition in unseen and noisy channel conditions

ICASSP 2017accepted

Speech recognition in varying background conditions is a challenging problem. Acoustic condition mismatch between training and evaluation data can significantly reduce recognition performance. For mismatched conditions, data-adaptation techniques are typically found to be useful, as they expose the…

Cited by 0SourceScholar
2016

Noise and reverberation effects on depression detection from speech

ICASSP 2016accepted

Speech-based depression detection has gained importance in recent years, but most research has used relatively quiet conditions or examined a single corpus per study. Little is thus known about the robustness of speech cues in the wild. This study compares the effect of noise and reverberation on de…

Cited by 0SourceScholar
2015

Cross-corpus depression prediction from speech

ICASSP 2015accepted

Research on detecting depression from speech has advanced in recent years, but most work has focused on the analysis of one corpus at a time. Given that clinical corpora are typically small, it is important to explore approaches that generalize across corpora and that could ultimately be adapted to…

Cited by 0SourceScholar
2015

Effects of feature type, learning algorithm and speaking style for depression detection from speech

ICASSP 2015accepted

Computational methods for speech-based detection of depression are still relatively new, and have focused on either a standard set of features or on specific additional approaches. We systematically study the effects of feature type, machine learning approach, and speaking style (read versus spontan…

Cited by 0SourceScholar