← Search

Sunil Kumar Kopparapu

9 accepted papers

2025

HamaraAwaz: Advancing Low-Latency Streaming TTS for Multilingual Speech in Indian Languages

ICASSP 2025accepted

We present a multilingual, multi-speaker, low-latency speech synthesis system developed by the HamaraAwaz team for Track 1 of the LIMMITS’25 challenge. To improve speaker similarity and naturalness in Indic languages, we build on ParrotTTS. We utilize disentangled self-supervised speech representati…

Cited by 0SourceScholar
2021

Deep Lung Auscultation Using Acoustic Biomarkers for Abnormal Respiratory Sound Event Detection

ICASSP 2021accepted

Lung Auscultation is a non-invasive process of distinguishing normal respiratory sounds from abnormal ones by analyzing the airflow along the respiratory tract. With developments in the Deep Learning (DL) techniques and wider access to anonymized medical data, automatic detection of specific sounds…

Cited by 0SourceScholar
2020

A Novel Approach for Intelligibility Assessment in Dysarthric Subjects

ICASSP 2020accepted

Dysarthria is a motor speech impairment caused by muscle weakness. Individuals, with this condition, are unable to control rapid movement of the velum leading to reduction in intelligibility, audibility, naturalness and efficiency of vocal communication. Systems that can assess intelligibility of dy…

Cited by 0SourceScholar
2020

Deep Encoded Linguistic and Acoustic Cues for Attention Based End to End Speech Emotion Recognition

ICASSP 2020accepted

An End-to-End model with convolutional layers and multi-head self attention mechanism is proposed for Speech Emotion Recognition (SER) task. As inputs, we propose to use both the deep encoded linguistic features that carry the language related context of emotion and the audio spectrogram that are re…

Cited by 0SourceScholar
2020

Improved Speaker Independent Dysarthria Intelligibility Classification Using Deepspeech Posteriors

ICASSP 2020accepted

Individuals with dysarthria are unable to control rapid movement of the velum leading to reduction in intelligibility, audibility, naturalness and efficiency of vocal communication. Automatic intelligibility assessment of dysarthric patients allows clinicians diagnose the impact of therapy and medic…

Cited by 0SourceScholar
2020

Multi-Conditioning and Data Augmentation Using Generative Noise Model for Speech Emotion Recognition in Noisy Conditions

ICASSP 2020accepted

Degradation due to additive noise is a significant road block in the real-life deployment of Speech Emotion Recognition (SER) systems. Most of the previous work in this field dealt with the noise degradation either at the signal or at the feature level. In this paper, to address the robustness aspec…

Cited by 0SourceScholar
2019

Improving ASR Robustness to Perturbed Speech Using Cycle-consistent Generative Adversarial Networks

ICASSP 2019accepted

Naturally introduced perturbations in audio signal, caused by emotional and physical states of the speaker, can significantly degrade the performance of Automatic Speech Recognition (ASR) systems. In this paper, we propose a front-end based on Cycle-Consistent Generative Adversarial Network (CycleGA…

Cited by 0SourceScholar
2017

Automatic assessment of dysarthria severity level using audio descriptors

ICASSP 2017accepted

Dysarthria is a motor speech impairment, often characterized by speech that is generally indiscernible by human listeners. Assessment of the severity level of dysarthria provides an understanding of the patient's progression in the underlying cause and is essential for planning therapy, as well as i…

Cited by 0SourceScholar