← Search

Prasanta Kumar Ghosh

38 accepted papers

2025

Improving Dialect Identification in Indian Languages Using Multimodal Features from Dialect Informed ASR

ICASSP 2025accepted

Dialect identification (DID) addresses the challenge of recog-nizing regional variations within a language. The current deep learning approaches focus on audio-only, text-only, or multi-task setups combining automatic speech recognition (ASR) with DID. This work introduces a novel multimodal archite…

Cited by 0SourceScholar
2025

RESPIN-S1.0: A read speech corpus of 10000+ hours in dialects of nine Indian Languages

NeurIPS 2025poster

We introduce **RESPIN-S1.0**, the largest publicly available dialect-rich read-speech corpus for Indian languages, comprising more than 10,000 hours of validated audio across nine major languages: Bengali, Bhojpuri, Chhattisgarhi, Hindi, Kannada, Magahi, Maithili, Marathi, and Telugu. Indian languag…

Cited by 0SourcecodeScholar
2025

Role of the Pretraining and the Adaptation data sizes for low-resource real-time MRI video segmentation

ICASSP 2025accepted

Real-time Magnetic Resonance Imaging (rtMRI) is frequently used in speech production studies as it provides a complete view of the vocal tract during articulation. This study investigates the effectiveness of rtMRI in analyzing vocal tract movements by employing the SegNet and UNet models for Air-Ti…

Cited by 0SourceScholar
2024

An Unsupervised Segmentation of Vocal Breath Sounds

ICASSP 2024accepted

Breathing is essential to human survival, which carries information about a person’s physiological and psychological state. Mostly breath sound boundaries are marked manually before being used for any task such as classification, spectral analysis, etc., which is very tedious. Various techniques hav…

Cited by 0SourceScholar
2024

Spectral Analysis of Vowels and Fricatives at Varied Levels of Dysarthria Severity for Amyotrophic Lateral Sclerosis

ICASSP 2024accepted

Dysarthria due to Amyotrophic Lateral Sclerosis (ALS) affects the acoustic characteristics of different speech sounds. The effects intensify with increasing severity leading to the collapse of the acoustic space of the affected individuals. With an aim to characterize such changes in the acoustic sp…

Cited by 0SourceScholar
2023

Exploring the Role of Fricatives in Classifying Healthy Subjects and Patients with Amyotrophic Lateral Sclerosis and Parkinson's Disease

ICASSP 2023accepted

Dysarthria due to Amyotrophic Lateral Sclerosis (ALS) and Parkinson’s Disease (PD) impairs sustained phoneme productions. Vowels and fricatives get affected differently owing to the differences in their production mechanisms. This paper examines three sustained voiceless fricatives - /s/, /sh/ and /…

Cited by 0SourceScholar
2023

Improved Acoustic-to-Articulatory Inversion Using Representations from Pretrained Self-Supervised Learning Models

ICASSP 2023accepted

In this work, we investigate the effectiveness of pretrained Self-Supervised Learning (SSL) features for learning the mapping for acoustic to articulatory inversion (AAI). Signal processing-based acoustic features such as MFCCs have been predominantly used for the AAI task with deep neural networks.…

Cited by 0SourceScholar
2023

Lightweight, Multi-Speaker, Multi-Lingual Indic Text-to-Speech

ICASSP 2023accepted

The Lightweight, Multi-speaker, Multi-lingual Indic Text-to-Speech (LIMMITS’23) challenge is organized as part of the ICASSP 2023 signal processing grand challenge. LIMMITS’23 aims at the development of a lightweight, multi-speaker, multi-lingual Text to Speech (TTS) model using datasets in Marathi,…

Cited by 0SourceScholar
2023

Real-Time MRI Video Synthesis from Time Aligned Phonemes with Sequence-to-Sequence Networks

ICASSP 2023accepted

Real-Time Magnetic resonance imaging (rtMRI) of the midsagittal plane of the mouth is of interest for speech production research. In this work, we focus on estimating utterance level rtMRI video from the spoken phoneme sequence. We obtain time-aligned phonemes from forced alignment, to obtain frame-…

Cited by 0SourceScholar
2023

Static and Dynamic Source and Filter Cues for Classification of Amyotrophic Lateral Sclerosis Patients and Healthy Subjects

ICASSP 2023accepted

Dysarthria due to Amyotrophic Lateral Sclerosis (ALS) affects speech production. Even the elementary sustained vowel utterances get impaired. For these, the impairments can be in achieving vowel-specific articulatory configurations, reflected in static acoustic cues, and/or in sustaining a configura…

Cited by 0SourceScholar
2022

An Error Correction Scheme for Improved Air-Tissue Boundary in Real-Time MRI Video for Speech Production

ICASSP 2022accepted

The best performance in Air-tissue boundary (ATB) segmentation of real-time Magnetic Resonance Imaging (rtMRI) videos in speech production is known to be achieved by a 3-dimensional convolutional neural network (3D-CNN) model. However, the evaluation of this model, as well as other ATB segmentation…

Cited by 0SourceScholar
2022

Dual Attention Pooling Network for Recording Device Classification Using Neutral and Whispered Speech

ICASSP 2022accepted

In this work, we proposed a method for recording device classification using the recorded speech signal. With the rapid increase in different mobile and professional recording devices, determining the source device has many applications in forensics and in further improving various speech-based appl…

Cited by 0SourceScholar
2022

SegNet-Based Deep Representation Learning for Dysphagia Classification

ICASSP 2022accepted

Swallowing disorders, broadly known as Dysphagia, are difficulties in the process of swallowing food. Many currently available methods for classifying healthy and dysphagic swallows typically use hand-picked acoustic features. This article presents a SegNet-based method for classifying healthy and d…

Cited by 0SourceScholar
2022

The impact of cross language on acoustic-to-articulatory inversion and its influence on articulatory speech synthesis

ICASSP 2022accepted

Estimating articulatory representations (ARs) from acoustic features is known as acoustic-to-articulatory inversion (AAI). Various factors of input acoustic features impact the performance of AAI. In this work, we investigate the effect of unseen language on the AAI performance in both seen and unse…

Cited by 0SourceScholar
2021

Acoustic-to-Articulatory Inversion for Dysarthric Speech by Using Cross-Corpus Acoustic-Articulatory Data

ICASSP 2021accepted

In this work, we focus on estimating articulatory movements from acoustic features, known as acoustic-to-articulatory inversion (AAI), for dysarthric patients with amyotrophic lateral sclerosis (ALS). Unlike healthy subjects, there are two potential challenges involved in AAI on dysarthric speech. D…

Cited by 0SourceScholar
2021

Effect of Noise and Model Complexity on Detection of Amyotrophic Lateral Sclerosis and Parkinson's Disease Using Pitch and MFCC

ICASSP 2021accepted

Dysarthria due to Amyotrophic Lateral Sclerosis (ALS) and Parkinson’s disease (PD) impacts both articulation and prosody in an individual’s speech. Complex deep neural networks exploit these cues for detection of ALS and PD. These are typically done using recordings in laboratory condition. This stu…

Cited by 0SourceScholar
2021

Impact of Speaking Rate on the Source Filter Interaction in Speech: A Study

ICASSP 2021accepted

Source filter interaction (SFI) explains the drop in pitch caused due to the constriction in the vocal tract during voiced consonant production in a vowel-consonant-vowel (VCV) sequence. In this work, we examine how the drop in pitch alters when such a VCV sequence is spoken at three different speak…

Cited by 0SourceScholar
2020

A Comparative Study of Estimating Articulatory Movements from Phoneme Sequences and Acoustic Features

ICASSP 2020accepted

Unlike phoneme sequences, movements of speech articulators (lips, tongue, jaw, velum) and the resultant acoustic signal are known to encode not only the linguistic message but also carry para-linguistic information. While several works exist for estimating articulatory movement from acoustic signals…

Cited by 0SourceScholar
2020

Analysis of Acoustic Features for Speech Sound Based Classification of Asthmatic and Healthy Subjects

ICASSP 2020accepted

Non-speech sounds (cough, wheeze) are typically known to perform better than speech sounds for asthmatic and healthy subject classification. In this work, we use sustained phonations of speech sounds, namely, /α:/, /i:/, /u:/, /eI/, /ou/, /s/, and /z/ from 47 asthmatic and 48 healthy controls. We co…

Cited by 0SourceScholar
2020

Automatic Classification of Volumes of Water Using Swallow Sounds from Cervical Auscultation

ICASSP 2020accepted

The signatures of swallowing vary depending on the volume of bolus swallowed. Among existing instrumental methods, cervical auscultation (CA) captures the acoustic signatures of the swallow sound. Although many features present in the literature can characterize volumes of swallow using CA, they req…

Cited by 0SourceScholar
2020

Automatic Identification of Speakers From Head Gestures in a Narration

ICASSP 2020accepted

In this work, we focus on quantifying speaker identity information encoded in the head gestures of speakers, while they narrate a story. We hypothesize that the head gestures over a long duration have speaker-specific patterns. To establish this, we consider a classification problem to identify spea…

Cited by 0SourceScholar
2020

Pseudo Likelihood Correction Technique for Low Resource Accented ASR

ICASSP 2020accepted

With the availability of large data, ASRs perform well on native English but poorly for non-native English data. Training nonnative ASRs or adapting a native English ASR is often limited by the availability of data, particularly for low resource scenarios. A typical HMM-DNN based ASR decoding requir…

Cited by 0SourceScholar
2020

Voice based classification of patients with Amyotrophic Lateral Sclerosis, Parkinson's Disease and Healthy Controls with CNN-LSTM using transfer learning

ICASSP 2020accepted

In this paper, we consider 2-class and 3-class classification problems for classifying patients with Amyotrophic Lateral Sclerosis (ALS), Parkinson's Disease (PD), and Healthy Controls (HC) using a CNNLSTM network. Classification performance is examined for three different tasks, namely, Spontaneous…

Cited by 30SourceScholar
2019

A Study on Robustness of Articulatory Features for Automatic Speech Recognition of Neutral and Whispered Speech

ICASSP 2019accepted

Traditionally, automatic speech recognition (ASR) systems are trained on acoustic representations of neutral speech. As a result, their performance degrades when tested with whispered speech. In this work, we explore the robustness of articulatory features in ASR of neutral and whispered speech. We…

Cited by 0SourceScholar
2019

Air-tissue Boundary Segmentation in Real Time Magnetic Resonance Imaging Video Using a Convolutional Encoder-decoder Network

ICASSP 2019accepted

In this paper, we propose a convolutional encoder-decoder network (CEDN) based approach for upper and lower Air-Tissue Boundary (ATB) segmentation within vocal tract in real-time magnetic resonance imaging (rtMRI) video frames. The output images from CEDN are processed using perimeter and moving ave…

Cited by 0SourceScholar
2019

An Improved Air Tissue Boundary Segmentation Technique for Real Time Magnetic Resonance Imaging Video Using Segnet

ICASSP 2019accepted

This paper presents an improved methodology for the segmentation of the Air-Tissue boundaries (ATBs) in the upper airway of the human vocal tract using Real-Time Magnetic Resonance Imaging (rtMRI) videos. Semantic segmentation is deployed in the proposed approach using a Deep learning architecture c…

Cited by 0SourceScholar
2019

Formant-gaps Features for Speaker Verification Using Whispered Speech

ICASSP 2019accepted

In this work, we propose a new feature based on formants for whispered speaker verification (SV) task, where neutral data is used for enrollment and whispered recordings are used for test. Such a mismatch between enrollment and test often degrades the performance of whispered SV systems due to the d…

Cited by 0SourceScholar
2019

Representation Learning Using Convolution Neural Network for Acoustic-to-articulatory Inversion

ICASSP 2019accepted

Recent techniques employ end-to-end systems to learn relevant features for several speech related applications, including speech recognition, and speaker verification. In this work, we focus on the task of acoustic-to-articulatory inversion (AAI) for which we propose an end-to-end system that compri…

Cited by 0SourceScholar
2018

A Supervised Air-Tissue Boundary Segmentation Technique in Real-Time Magnetic Resonance Imaging Video Using a Novel Measure of Contrast and Dynamic Programming

ICASSP 2018accepted

This paper introduces a technique for the supervised segmentation of Air-Tissue Boundaries (ATBs) in the upper airway of the vocal tract in the real time magnetic resonance imaging (rtMRI) videos. The proposed technique uses a novel measure of contrast across a boundary using Fisher discriminant fun…

Cited by 0SourceScholar
2018

Binaural Speech Source Localization Using Template Matching of Interaural Time Difference Patterns

ICASSP 2018accepted

In this paper we present a template based algorithm for localizing speech sources from a binaural recording. Binaural recordings are associated with head related transfer functions (HRTFs) for each direction which are specific to the object, say head, in between the two microphones. So, using these…

Cited by 0SourceScholar
2018

Comparison of Speech Tasks for Automatic Classification of Patients with Amyotrophic Lateral Sclerosis and Healthy Subjects

ICASSP 2018accepted

In this work, we consider the task of acoustic and articulatory feature based automatic classification of Amyotrophic Lateral Sclerosis (ALS) patients and healthy subjects using speech tasks. In particular, we compare the roles of different types of speech tasks, namely rehearsed speech, spontaneous…

Cited by 0SourceScholar
2018

Concatenative Articulatory Video Synthesis Using Real-Time MRI Data for Spoken Language Training

ICASSP 2018accepted

Spoken language training benefits from showing a video of native speakers' articulatory movements to train the second language learners. Typically, the articulatory video is prepared in conjunction with the audio which is collected simultaneously with the articulatory recording. Articulatory video r…

Cited by 0SourceScholar
2017

A comparative study of acoustic-to-articulatory inversion for neutral and whispered speech

ICASSP 2017accepted

Whispered speech is known to have different characteristics in acoustics and articulation compared to neutral speech. In this study, we compare the accuracy with which the articulation can be recovered from the acoustics of both types of speech, individually. Acoustic-to-articulatory inversion (AAI)…

Cited by 0SourceScholar
2017

Automatic detection of syllable stress using sonority based prominence features for pronunciation evaluation

ICASSP 2017accepted

Automatic syllable stress detection is useful in assessing and diagnosing the quality of the pronunciation of second language (L2) learners in an automated way. Typically, the syllable stress depends on three prominence measures - intensity level, duration, pitch - around the sound unit with the hig…

Cited by 0SourceScholar
2016

A robust speech rate estimation based on the activation profile from the selected acoustic unit dictionary

ICASSP 2016accepted

A typical solution for the speech rate estimation consists of two stages, which involves first computing a short-time feature contour such that most of peaks of the contour correspond to the syllable nuclei followed by the detection of the peaks of the contour corresponding to the syllable nuclei. T…

Cited by 0SourceScholar
2016

Better acoustic normalization in subject independent acoustic-to-articulatory inversion: Benefit to recognition

ICASSP 2016accepted

In subject independent acoustic-to-articulatory inversion (SII), the training and test subjects are in general different, whereas subject dependent inversion (SDI) uses the same training and test subjects. Thus, acoustic normalization is used to compensate for the mismatch between the training and t…

Cited by 0SourceScholar
2015

Estimation of the invariant and variant characteristics in speech articulation and its application to speaker identification

ICASSP 2015accepted

Speech articulation varies across speakers for producing a speech sound due to the differences in their vocal tract morphologies, though the speech motor actions are executed in terms of relatively invariant gestures [1]. While the invariant articulatory gestures are driven by the linguistic content…

Cited by 0SourceScholar