← Search

Tilak Purohit

4 accepted papers

2025

Automatic Parkinson's disease detection from speech: Layer selection vs adaptation of foundation models

ICASSP 2025accepted

In this work, we investigate Speech Foundation Models (SFMs) for Parkinson’s Disease (PD) detection. We explore two main approaches: (1) using SFMs as frozen feature extractors and, (2) fine-tuning/adapting SFMs for PD detection. We propose a cross-validation-based layer selection methodology to ide…

Cited by 0SourceScholar
2025

Emotion information recovery potential of wav2vec2 network fine-tuned for speech recognition task

ICASSP 2025accepted

Fine-tuning has become a norm to achieve state-of-the-art performance when employing pre-trained networks like foundation models. These models are typically pre-trained on large-scale unannotated data using self-supervised learning (SSL) methods. The SSL-based pre-training on large-scale data enable…

Cited by 0SourceScholar
2023

Towards Learning Emotion Information from Short Segments of Speech

ICASSP 2023accepted

Conventionally, speech emotion recognition has been approached by utterance or turn-level modelling of input signals, either through extracting hand-crafted low-level descriptors, bag-of-audio-words features or by feeding long-duration signals directly to deep neural networks (DNNs). While this appr…

Cited by 0SourceScholar
2021

Impact of Speaking Rate on the Source Filter Interaction in Speech: A Study

ICASSP 2021accepted

Source filter interaction (SFI) explains the drop in pitch caused due to the constriction in the vocal tract during voiced consonant production in a vowel-consonant-vowel (VCV) sequence. In this work, we examine how the drop in pitch alters when such a VCV sequence is spoken at three different speak…

Cited by 0SourceScholar