← Search

Catherine Lai

6 accepted papers

2025

Can We "Cherry-Pick"? Investigating Multiple Renditions from a Generative Speech Synthesis Model

ICASSP 2025accepted

Generative Speech Models (GSMs) have seen a surge in popularity due to their ability to generate diverse and high-quality speech. Evaluating models that generate many different renditions for a given input sentence presents a new challenge. Listening tests are still the gold standard for evaluating…

Cited by 0SourceScholar
2025

Revise, Reason, and Recognize: LLM-Based Emotion Recognition via Emotion-Specific Prompts and ASR Error Correction

ICASSP 2025accepted

Annotating and recognizing speech emotion using prompt engineering has recently emerged with the advancement of Large Language Models (LLMs), yet its efficacy and reliability remain questionable. In this paper, we conduct a systematic study on this topic, beginning with the proposal of novel prompts…

Cited by 0SourceScholar
2025

Semi-Supervised Cognitive State Classification from Speech with Multi-View Pseudo-Labeling

ICASSP 2025accepted

The lack of labeled data is a common challenge in speech classification tasks, particularly those requiring extensive subjective assessment, such as cognitive state classification. In this work, we propose a Semi-Supervised Learning (SSL) framework, introducing a novel multi-view pseudo-labeling met…

Cited by 0SourceScholar
2024

Language Technologies as If People Mattered: Centering Communities in Language Technology Development

COLING 2024main

In this position paper we argue that researchers interested in language and/or language technologies should attend to challenges of linguistic and algorithmic injustice together with language communities. We put forward that this can be done by drawing together diverse scholarly and experiential ins…

Cited by 5SourcePDFScholar
2023

Multimodal Dyadic Impression Recognition via Listener Adaptive Cross-Domain Fusion

ICASSP 2023accepted

As a sub-branch of affective computing, impression recognition, e.g., perception of speaker characteristics such as warmth or competence, is potentially a critical part of both human-human conversations and spoken dialogue systems. Most research has studied impressions only from the behaviors expres…

Cited by 0SourceScholar