← Search

Mathew Magimai-Doss

23 accepted papers

2025

Automatic Parkinson's disease detection from speech: Layer selection vs adaptation of foundation models

ICASSP 2025accepted

In this work, we investigate Speech Foundation Models (SFMs) for Parkinson’s Disease (PD) detection. We explore two main approaches: (1) using SFMs as frozen feature extractors and, (2) fine-tuning/adapting SFMs for PD detection. We propose a cross-validation-based layer selection methodology to ide…

Cited by 0SourceScholar
2025

Comparing Self-Supervised Learning Models Pre-Trained on Human Speech and Animal Vocalizations for Bioacoustics Processing

ICASSP 2025accepted

Self-supervised learning (SSL) foundation models have emerged as powerful, domain-agnostic, general-purpose feature extractors applicable to a wide range of tasks. Such models pre-trained on human speech have demonstrated high transferability for bioacoustic processing. This paper investigates (i) w…

Cited by 0SourceScholar
2025

Emotion information recovery potential of wav2vec2 network fine-tuned for speech recognition task

ICASSP 2025accepted

Fine-tuning has become a norm to achieve state-of-the-art performance when employing pre-trained networks like foundation models. These models are typically pre-trained on large-scale unannotated data using self-supervised learning (SSL) methods. The SSL-based pre-training on large-scale data enable…

Cited by 0SourceScholar
2024

Comparing data-Driven and Handcrafted Features for Dimensional Emotion Recognition

ICASSP 2024accepted

Speech Emotion Recognition (SER) has garnered significant attention over the past two decades. In the early stages of SER technology, ’brute force’-based techniques led to a significant expansion in knowledge-based acoustic feature representation (FR) for modeling sparse emotional data. However, as…

Cited by 0SourceScholar
2024

Content-Based Objective Evaluation of Artificially Generated Sign Language Videos

ICASSP 2024accepted

Sign language is vital for communication within the deaf and hard-of-hearing community. Avatar-based methods and deep learning techniques like Generative Adversarial Networks have shown promise in generating sign language video content. One of the challenges in sign language generation is the evalua…

Cited by 0SourceScholar
2023

Towards Learning Emotion Information from Short Segments of Speech

ICASSP 2023accepted

Conventionally, speech emotion recognition has been approached by utterance or turn-level modelling of input signals, either through extracting hand-crafted low-level descriptors, bag-of-audio-words features or by feeding long-duration signals directly to deep neural networks (DNNs). While this appr…

Cited by 0SourceScholar
2022

Modeling of Pre-Trained Neural Network Embeddings Learned From Raw Waveform for COVID-19 Infection Detection

ICASSP 2022accepted

COVID-19 is a respiratory system disorder that can disrupt the function of lungs. Effects of dysfunctional respiratory mechanism can reflect upon other modalities which function in close coupling. Audio signals result from modulation of respiration through speech production system, and hence acousti…

Cited by 0SourceScholar
2021

On The Relationship Between Speech-Based Breathing Signal Prediction Evaluation Measures and Breathing Parameters Estimation

ICASSP 2021accepted

The respiratory system is one of the major components of the speech production system. Any alteration in breathing can result in changes in speech. Specific breathing characteristics, such as breathing rate and tidal volume, can indicate a person’s pathological condition. More recently, neural netwo…

Cited by 0SourceScholar
2020

Detection Of S1 And S2 Locations In Phonocardiogram Signals Using Zero Frequency Filter

ICASSP 2020accepted

Heart auscultation is a widely used technique for diagnosing cardiac abnormalities. In that context, capturing of phonocardiogram (PCG) signals and automatically monitoring of the heart by identifying S1 and S2 complexes is an emerging field. One of the first steps involved for identifying S1-S2 com…

Cited by 0SourceScholar
2020

Estimating the Degree of Sleepiness by Integrating Articulatory Feature Knowledge in Raw Waveform Based CNNS

ICASSP 2020accepted

Speech-based degree of sleepiness estimation is an emerging research problem. This paper investigates an end-to-end approach, where given raw waveform as input, a convolutional neural network (CNN) estimates at its output the degree of sleepiness. Within this approach, we investigate constraining th…

Cited by 0SourceScholar
2019

HMM-based Approaches to Model Multichannel Information in Sign Language Inspired from Articulatory Features-based Speech Processing

ICASSP 2019accepted

Sign language conveys information through multiple channels, such as hand shape, hand movement, and mouthing. Modeling this multichannel information is a highly challenging problem. In this paper, we elucidate the link between spoken language and sign language in terms of production phenomenon and p…

Cited by 0SourceScholar
2019

Improving Children Speech Recognition through Feature Learning from Raw Speech Signal

ICASSP 2019accepted

Children speech recognition based on short-term spectral features is a challenging task. One of the reasons is that children speech has high fundamental frequency that is comparable to formant frequency values. Furthermore, as children grow, their vocal apparatus also undergoes changes. This present…

Cited by 0SourceScholar
2019

Learning Voice Source Related Information for Depression Detection

ICASSP 2019accepted

During depression neurophysiological changes can occur, which may affect laryngeal control i.e. behaviour of the vocal folds. Characterising these changes in a precise manner from speech signals is a non trivial task, as this typically involves reliable separation of the voice source information fro…

Cited by 0SourceScholar
2019

Segment-level Training of ANNs Based on Acoustic Confidence Measures for Hybrid HMM/ANN Speech Recognition

ICASSP 2019accepted

We show that confidence measures estimated from local posterior probabilities can serve as objective functions for training ANNs in hybrid HMM based speech recognition systems. This leads to a segment-level training paradigm that overcomes the limitation of frame-level updates ignoring the sequence…

Cited by 0SourceScholar
2018

Towards Directly Modeling Raw Speech Signal for Speaker Verification Using CNNS

ICASSP 2018accepted

Speaker verification systems traditionally extract and model cepstral features or filter bank energies from the speech signal. In this paper, inspired by the success of neural network-based approaches to model directly raw speech signal for applications such as speech recognition, emotion recognitio…

Cited by 0SourceScholar
2015

An HMM-based formalism for automatic subword unit derivation and pronunciation generation

ICASSP 2015accepted

We propose a novel hidden Markov model (HMM) formalism for automatic derivation of subword units and pronunciation generation using only transcribed speech data. In this approach, the subword units are derived from the clustered context-dependent units in a grapheme based system using maximum-likeli…

Cited by 0SourceScholar
2015

Convolutional Neural Networks-based continuous speech recognition using raw speech signal

ICASSP 2015accepted

State-of-the-art automatic speech recognition systems model the relationship between acoustic speech signal and phone classes in two stages, namely, extraction of spectral-based features based on prior knowledge followed by training of acoustic model, typically an artificial neural network (ANN). In…

Cited by 0SourceScholar
2015

Integrated pronunciation learning for automatic speech recognition using probabilistic lexical modeling

ICASSP 2015accepted

Standard automatic speech recognition (ASR) systems use phoneme-based pronunciation lexicon prepared by linguistic experts. When the hand crafted pronunciations fail to cover the vocabulary of a new domain, a grapheme-to-phoneme (G2P) converter is used to extract pronunciations for new words and the…

Cited by 0SourceScholar
2015

Objective speech intelligibility assessment through comparison of phoneme class conditional probability sequences

ICASSP 2015accepted

Assessment of speech intelligibility is important for the development of speech systems, such as telephony systems and text-to-speech (TTS) systems. Existing approaches to the automatic assessment of intelligibility in telephony typically compare a reference speech signal to a degraded copy, which r…

Cited by 0SourceScholar