← Search

Sudarsana Reddy Kadiri

10 accepted papers

2025

Enhancing Listened Speech Decoding from EEG via Parallel Phoneme Sequence Prediction

ICASSP 2025accepted

Brain-computer interfaces (BCI) offer numerous human-centered application possibilities, particularly affecting people with neurological disorders. Text or speech decoding from brain activities is a relevant domain that could augment the quality of life for people with impaired speech perception. We…

Cited by 0SourceScholar
2025

Wavelet Scattering Network Features for Intensity Category Classification and Prediction of SPL from Speech

ICASSP 2025accepted

Speakers change vocal intensity in daily life to communicate over long distances and to express vocal emotions. Humans produce speech using different intensity categories (e.g. soft, normal and loud voice) and they can regulate intensity across a wide sound pressure level (SPL) range. Knowing the in…

Cited by 0SourceScholar
2023

Automatic Classification of Vocal Intensity Category from Speech

ICASSP 2023accepted

Regulation of vocal intensity is a fundamental phenomenon in speech communication. Vocal intensity can be quantified using sound pressure level (SPL), which can be measured easily by recording a standard calibration signal with speech and by comparing the energy of the recorded speech signal with th…

Cited by 0SourceScholar
2023

Utilizing Wav2Vec In Database-Independent Voice Disorder Detection

ICASSP 2023accepted

Automatic detection of voice disorders from acoustic speech signals can help to improve reliability of medical diagnosis. However, the real-life environment in which speech signals are recorded for diagnosis can be different from the environment in which the detection system’s training data was orig…

Cited by 0SourceScholar
2023

Wav2vec-Based Detection and Severity Level Classification of Dysarthria From Speech

ICASSP 2023accepted

Automatic detection and severity level classification of dysarthria directly from acoustic speech signals can be used as a tool in medical diagnosis. In this work, the pre-trained wav2vec 2.0 model is studied as a feature extractor to build detection and severity level classification systems for dys…

Cited by 0SourceScholar
2020

Comparison of Glottal Closure Instants Detection Algorithms for Emotional Speech

ICASSP 2020accepted

In production of voiced speech, epochs or glottal closure instants (GCIs) refer to the instants of significant excitation of the vocal tract. Extraction of GCIs is used as a pre-processing stage in many areas of speech technology, such as in prosody modification, speech synthesis and voice source an…

Cited by 10SourceScholar
2020

Study of Formant Modification for Children ASR

ICASSP 2020accepted

The performance of automatic speech recognition systems for children’s speech is known to suffer from the large variation and mismatch in the acoustic and linguistic attributes between children’s and adults’ speech. One of the various identified sources of mismatch is the difference in formant frequ…

Cited by 0SourceScholar
2017

Speech polarity detection using strength of impulse-like excitation extracted from speech epochs

ICASSP 2017accepted

In this paper, we address the issue of speech polarity detection using strength of impulse-like excitation around epoch. The correct detection of speech polarity is a crucial step for many speech processing algorithms to extract suitable information. Occurrence of errors in the detection of speech p…

Cited by 0SourceScholar
2015

Analysis of singing voice for epoch extraction using Zero Frequency Filtering method

ICASSP 2015accepted

Epoch is the instant of significant excitation of the vocal tract system during the production of voiced speech. Estimation of epochs or Glottal closure instants (GCIs) is a well studied topic in the speech analysis. From the recent studies on GCI detection from singing voice with state-of-art metho…

Cited by 0SourceScholar