← Search

Carol Y. Espy-Wilson

9 accepted papers

2025

CPT-Boosted Wav2vec2.0: Towards Noise Robust Speech Recognition for Classroom Environments

ICASSP 2025accepted

Creating Automatic Speech Recognition (ASR) systems that are robust and resilient to classroom conditions is paramount to the development of AI tools to aid teachers and students. In this work, we study the efficacy of continued pretraining (CPT) in adapting Wav2vec2.0 to the classroom domain. We sh…

Cited by 0SourceScholar
2025

Self-supervised Multimodal Speech Representations for the Assessment of Schizophrenia Symptoms

ICASSP 2025accepted

Multimodal schizophrenia assessment systems have gained traction over the last few years. This work introduces a schizophrenia assessment system to discern between prominent symptom classes of schizophrenia and predict an overall schizophrenia severity score. We develop a Vector Quantized Variationa…

Cited by 0SourceScholar
2023

The Secret Source : Incorporating Source Features to Improve Acoustic-To-Articulatory Speech Inversion

ICASSP 2023accepted

In this work, we incorporated acoustically derived source features, aperiodicity, periodicity and pitch as additional targets to an acoustic-to-articulatory speech inversion (SI) system. We also propose a Temporal Convolution based SI system, which uses auditory spectrograms as the input speech repr…

Cited by 0SourceScholar
2022

Harmonicity Plays a Critical Role in DNN Based Versus in Biologically-Inspired Monaural Speech Segregation Systems

ICASSP 2022accepted

Recent advancements in deep learning have led to drastic improvements in speech segregation models. Despite their success and growing applicability, few efforts have been made to analyze the underlying principles that these networks learn to perform segregation. Here we analyze the role of harmonici…

Cited by 0SourceScholar
2022

Multimodal Depression Classification using Articulatory Coordination Features and Hierarchical Attention Based text Embeddings

ICASSP 2022accepted

Multimodal depression classification has gained immense popularity over the recent years. We develop a multimodal depression classification system using articulatory coordination features extracted from vocal tract variables and text transcriptions obtained from an automatic speech recognition tool…

Cited by 0SourceScholar
2018

Semi-Supervised and Transfer Learning Approaches for Low Resource Sentiment Classification

ICASSP 2018accepted

Sentiment classification involves quantifying the affective reaction of a human to a document, media item or an event. Although researchers have investigated several methods to reliably infer sentiment from lexical, speech and body language cues, training a model with a small set of labeled datasets…

Cited by 0SourceScholar
2018

Smoothing Model Predictions Using Adversarial Training Procedures for Speech Based Emotion Recognition

ICASSP 2018accepted

Training discriminative classifiers involves learning a conditional distribution p(y <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">i</sup> |x <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">i</sub> ), giv…

Cited by 0SourceScholar
2017

Joint modeling of articulatory and acoustic spaces for continuous speech recognition tasks

ICASSP 2017accepted

Articulatory information can effectively model variability in speech and can improve speech recognition performance under varying acoustic conditions. Learning speaker-independent articulatory models has always been challenging, as speaker-specific information in the articulatory and acoustic spaces…

Cited by 0SourceScholar