← Search

Emily Mower Provost

9 accepted papers

2025

Efficient Finetuning for Dimensional Speech Emotion Recognition in the Age of Transformers

ICASSP 2025accepted

Accurate speech emotion recognition is essential for developing human-facing systems. Recent advancements have included finetuning large, pretrained transformer models like Wav2Vec 2.0. However, the finetuning process requires substantial computational resources, including high-memory GPUs and signi…

Cited by 0SourceScholar
2021

Learning Paralinguistic Features from Audiobooks through Style Voice Conversion

NAACL 2021long

Paralinguistics, the non-lexical components of speech, play a crucial role in human-human interaction. Models designed to recognize paralinguistic information, particularly speech emotion and style, are difficult to train because of the limited labeled datasets available. In this work, we present a…

Cited by 2SourcePDFScholar
2019

Exploiting Acoustic and Lexical Properties of Phonemes to Recognize Valence from Speech

ICASSP 2019accepted

Emotions modulate speech acoustics as well as language. The latter influences the sequences of phonemes that are produced, which in turn further modulate the acoustics. Therefore, phonemes impact emotion recognition in two ways: (1) they introduce an additional source of variability in speech signal…

Cited by 0SourceScholar
2019

Muse-ing on the Impact of Utterance Ordering on Crowdsourced Emotion Annotations

ICASSP 2019accepted

Emotion recognition algorithms rely on data annotated with high quality labels. However, emotion expression and perception are inherently subjective. There is generally not a single annotation that can be unambiguously declared "correct." As a result, annotations are colored by the manner in which t…

Cited by 0SourceScholar
2019

Trainable Time Warping: Aligning Time-series in the Continuous-time Domain

ICASSP 2019accepted

DTW calculates the similarity or alignment between two signals, subject to temporal warping. However, its computational complexity grows exponentially with the number of time-series. Although there have been algorithms developed that are linear in the number of time-series, they are generally quadra…

Cited by 0SourceScholar
2018

Improving End-of-Turn Detection in Spoken Dialogues by Detecting Speaker Intentions as a Secondary Task

ICASSP 2018accepted

This work focuses on the use of acoustic cues for modeling turn-taking in dyadic spoken dialogues. Previous work has shown that speaker intentions (e.g., asking a question, uttering a backchannel, etc.) can influence turn-taking behavior and are good predictors of turn-transitions in spoken dialogue…

Cited by 0SourceScholar
2016

Cross-corpus acoustic emotion recognition from singing and speaking: A multi-task learning approach

ICASSP 2016accepted

Emotion is expressed over both speech and song. Previous works have found that although spoken and sung emotion recognition are different tasks, they are related. Classifiers that explicitly utilize this relatedness can achieve better performance than classifiers that do not. Further, research in sp…

Cited by 58SourceScholar
2016

Mood state prediction from speech of varying acoustic quality for individuals with bipolar disorder

ICASSP 2016accepted

Speech contains patterns that can be altered by the mood of an individual. There is an increasing focus on automated and distributed methods to collect and monitor speech from large groups of patients suffering from mental health disorders. However, as the scope of these collections increases, the v…

Cited by 0SourceScholar