← Search

S. Pavankumar Dubagunta

6 accepted papers

2024

Multitask Speech Recognition and Speaker Change Detection for Unknown Number of Speakers

ICASSP 2024accepted

Traditionally, automatic speech recognition (ASR) and speaker change detection (SCD) systems have been independently trained to generate comprehensive transcripts accompanied by speaker turns. Recently, joint training of ASR and SCD systems, by inserting speaker turn tokens in the ASR training text,…

Cited by 0SourceScholar
2023

Towards Learning Emotion Information from Short Segments of Speech

ICASSP 2023accepted

Conventionally, speech emotion recognition has been approached by utterance or turn-level modelling of input signals, either through extracting hand-crafted low-level descriptors, bag-of-audio-words features or by feeding long-duration signals directly to deep neural networks (DNNs). While this appr…

Cited by 0SourceScholar
2020

Estimating the Degree of Sleepiness by Integrating Articulatory Feature Knowledge in Raw Waveform Based CNNS

ICASSP 2020accepted

Speech-based degree of sleepiness estimation is an emerging research problem. This paper investigates an end-to-end approach, where given raw waveform as input, a convolutional neural network (CNN) estimates at its output the degree of sleepiness. Within this approach, we investigate constraining th…

Cited by 0SourceScholar
2019

Improving Children Speech Recognition through Feature Learning from Raw Speech Signal

ICASSP 2019accepted

Children speech recognition based on short-term spectral features is a challenging task. One of the reasons is that children speech has high fundamental frequency that is comparable to formant frequency values. Furthermore, as children grow, their vocal apparatus also undergoes changes. This present…

Cited by 0SourceScholar
2019

Learning Voice Source Related Information for Depression Detection

ICASSP 2019accepted

During depression neurophysiological changes can occur, which may affect laryngeal control i.e. behaviour of the vocal folds. Characterising these changes in a precise manner from speech signals is a non trivial task, as this typically involves reliable separation of the voice source information fro…

Cited by 0SourceScholar
2019

Segment-level Training of ANNs Based on Acoustic Confidence Measures for Hybrid HMM/ANN Speech Recognition

ICASSP 2019accepted

We show that confidence measures estimated from local posterior probabilities can serve as objective functions for training ANNs in hybrid HMM based speech recognition systems. This leads to a segment-level training paradigm that overcomes the limitation of frame-level updates ignoring the sequence…

Cited by 0SourceScholar