← Search

Korin Richmond

11 accepted papers

2025

Cross-Lingual Speech Emotion Recognition: Humans vs. Self-Supervised Models

ICASSP 2025accepted

Utilizing Self-Supervised Learning (SSL) models for Speech Emotion Recognition (SER) has proven effective, yet limited research has explored cross-lingual scenarios. This study presents a comparative analysis between human performance and SSL models, beginning with a layer-wise analysis and an explo…

Cited by 0SourceScholar
2025

Exploring Acoustic Similarity in Emotional Speech and Music via Self-Supervised Representations

ICASSP 2025accepted

Emotion recognition from speech and music shares similarities due to their acoustic overlap, which has led to interest in transferring knowledge between these domains. However, the shared acoustic cues between speech and music, particularly those encoded by Self-Supervised Learning (SSL) models, rem…

Cited by 0SourceScholar
2022

Requirements and Motivations of Low-Resource Speech Synthesis for Language Revitalization

ACL 2022long

This paper describes the motivation and development of speech synthesis systems for the purposes of language revitalization. By building speech synthesis systems for three Indigenous languages spoken in Canada, Kanien’kéha, Gitksan & SENĆOŦEN, we re-evaluate the question of how much data is required…

2021

TaLNet: Voice Reconstruction from Tongue and Lip Articulation with Transfer Learning from Text-to-Speech Synthesis

AAAI 2021technical

This paper presents TaLNet, a model for voice reconstruction with ultrasound tongue and optical lip videos as inputs. TaLNet is based on an encoder-decoder architecture. Separate encoders are dedicated to processing the tongue and lip data streams respectively. The decoder pre…

Cited by 18SourcePDFScholar
2019

Attentive Filtering Networks for Audio Replay Attack Detection

ICASSP 2019accepted

An attacker may use a variety of techniques to fool an automatic speaker verification system into accepting them as a genuine user. Anti-spoofing methods meanwhile aim to make the system robust against such attacks. The ASVspoof 2017 Challenge focused specifically on replay attacks, with the intenti…

Cited by 0SourceScholar
2019

Speaker-independent Classification of Phonetic Segments from Raw Ultrasound in Child Speech

ICASSP 2019accepted

Ultrasound tongue imaging (UTI) provides a convenient way to visualize the vocal tract during speech production. UTI is increasingly being used for speech therapy, making it important to develop automatic methods to assist various time-consuming manual tasks currently performed by speech therapists.…

Cited by 0SourceScholar
2016

Initial investigation of speech synthesis based on complex-valued neural networks

ICASSP 2016accepted

Although frequency analysis often leads us to a speech signal in the complex domain, the acoustic models we frequently use are designed for real-valued data. Phase is usually ignored or modelled separately from spectral amplitude. Here, we propose a complex-valued neural network (CVNN) for directly…

Cited by 0SourceScholar
2016

Testing the consistency assumption: Pronunciation variant forced alignment in read and spontaneous speech synthesis

ICASSP 2016accepted

Forced alignment for speech synthesis traditionally aligns a phoneme sequence predetermined by the front-end text processing system. This sequence is not altered during alignment, i.e., it is forced, despite possibly being faulty. The consistency assumption is the assumption that these mistakes do n…

Cited by 14SourceScholar
2015

Methods for applying dynamic sinusoidal models to statistical parametric speech synthesis

ICASSP 2015accepted

Sinusoidal vocoders can generate high quality speech, but they have not been extensively applied to statistical parametric speech synthesis. This paper presents two ways for using dynamic sinusoidal models for statistical speech synthesis, enabling the sinusoid parameters to be modelled in HMM-based…

Cited by 0SourceScholar